Cluster-filtering scores

The cluster-filtering system improves sequencing by using SNR and base-call-quality scores across multiple cycles to filter high-quality reads, enhancing throughput and accuracy in nucleotide sequencing.

WO2026006771A1PCT designated stage Publication Date: 2026-01-02ILLUMINA INC
View PDF 36 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035750
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-14
Filing Date
2025-06-27
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing sequencing systems suffer from premature filtering out of high-quality nucleotide reads due to reliance on limited sequencing cycles and noisy chastity metrics, leading to reduced throughput and inaccurate variant calling.

Method used

A cluster-filtering system that uses signal-based features, such as signal-to-noise ratio (SNR) and base-call-quality scores from multiple sequencing cycles to determine a cluster-filtering score, ensuring higher quality reads pass filtering criteria.

Benefits of technology

This approach increases the throughput of high-quality reads by approximately 2% of millions or billions of clusters, enhances variant-calling accuracy, and reduces false positive/negative calls, leading to more precise identification of genetic variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035750_02012026_PF_FP_ABST
    Figure US2025035750_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure describes methods, non-transitory computer readable media, and systems that generate a cluster-filtering score that indicate an expected number of errors within a read based on features for signal values, including signal-to-noise ratios or quality scores from various base positions of a read. The disclosed systems can use the cluster-filtering scores to generate more accurate filtering thresholds that increase the proportion of reads that pass filter relative to a chastity filter without degrading read quality. In particular, the disclosed system can use at least two different methods to generate filtering thresholds, each relying on a different feature. The disclosed systems can use base-call-quality scores and signal-to-noise ratios from both reads of a paired-end read to generate a cluster-filtering score for a cluster. In both methods, the disclosed system can use information from several base positions of the read to inform decisions on which reads pass a cluster-filtering threshold score.
Need to check novelty before this filing date? Find Prior Art

Description

CLUSTER-FILTERING SCORESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 745,188, entitled, “CLUSTER-FILTERING SCORES FOR IDENTIFYING OLIGONUCLEOTIDE CLUSTERS THAT PASS FILTER BASED ON SEQUENCING DATA FROM DIFFERENT BASE POSITIONS WITHIN A NUCLEOTIDE READ,” filed on January 14, 2025 (IP-2774-PRV2) and U.S. Provisional Patent Application No. 63 / 665,910, entitled, “CLUSTER-FILTERING SCORES FOR IDENTIFYING OLIGONUCLEOTIDE CLUSTERS THAT PASS FILTER BASED ON SEQUENCING DATA FROM DIFFERENT BASE POSITIONS WITHIN A NUCLEOTIDE READ,” filed on June 28, 2024 (IP-2774-PRV). Both of the aforementioned applications are hereby incorporated by reference in their entirety.BACKGROUND

[0002] In recent years, biotechnology firms and research institutions have improved hardware and software platforms used for determining a sequence of nucleotide bases (also referred to as “nucleobases”) in a sample. For instance, some existing sequencing devices and sequencing-data- analysis software (together “existing sequencing systems”) determine individual nucleobases of nucleic-acid sequences by using conventional Sanger sequencing or by using sequencing-by- synthesis (SBS). When using SBS, existing sequencing systems can monitor many thousands to billions of oligonucleotides being synthesized in parallel to detect more accurate nucleobase calls. For instance, a camera in SBS platforms can capture images of irradiated fluorescent tags from nucleotide bases incorporated into such synthesized nucleic-acid sequences (often grouped into clusters of oligonucleotides). After capturing the images, a computing device from the existing systems uses sequencing-data-analysis software to determine nucleobases that were detected in a given image based on the light signal (e.g., the corresponding intensity values) captured in the image data. By iteratively incorporating nucleobases into the oligonucleotides and capturing images of the emitted light signals in various sequencing cycles, some existing sequencing systems can determine the sequence of nucleobases present in the samples of nucleic acid.

[0003] In employing SBS or other florescence-based sequencing chemistry, light signals and densely organized oligonucleotide clusters can compromise data quality for nucleotide reads of a sample. By using SBS or other fluorescence-based sequencing chemistry for detecting incorporated nucleobases, for example, existing sequencing systems sometimes produce unclear or noisy signals in part because signals from different nucleobases may overlap or be indistinct from one another. Such existing sequencing systems may also produce noisy signal data due to various factors, such as cluster density, reagent quality, and instrument performance. To take but one example, high oligonucleotide-cluster density on a flow cell or other nucleotide sample slides can cause signalsfrom adjacent clusters to overlap, thereby complicating the resolution required and complexity of correctly determining nucleobase calls.

[0004] To compensate for unclear or noisy signal data and other sequencing issues, some existing sequencing systems implement a chastity filter to filter out oligonucleotide clusters or nucleotide reads that are more likely of poor quality or ambiguous. For example, some sequencing systems apply a chastity filter by evaluating a purity of the signal from each oligonucleotide cluster during initial sequencing cycles. In some cases, existing sequencing systems determine a chastity value as a ratio of a brightest nucleobase intensity divided by a sum of the brightest and a second brightest nucleobase intensities. For instance, clusters may “pass filter” if no more than 1 nucleobase call has a chastity value below a chastity value threshold (e.g., 0.6) in the initial cycles of sequencing.

[0005] Despite the utility and well-tested reliability of such a chastity filter and recent advances in sequencing devices, some existing sequencing systems suffer from technical limitations that reduce the throughput of quality nucleotide reads for a sample. Some existing systems rely on chastity filtering that use data from a limited set of sequencing cycles or limited base locations within a nucleotide read. For instance, some existing sequencing systems determine a chastity value for nucleotide reads based on information from an initial 25 cycles of a nucleotide read. Because a chastity filter typically uses information from a limited set of sequencing cycles or base locations within a nucleotide read, transient or variation issues in signal intensity during early cycles can disproportionately impact the chastity metric. For example, high-quality reads that may exhibit slight signal variability or transient fluctuations in early cycles can be incorrectly discarded by the chastity filter. Consequently, high-quality reads that may have stabilized and produced reliable sequence data in later sequencing cycles do not pass a chastity filter and are not analyzed during downstream analysis. Accordingly, some existing systems that rely solely on a limited set of sequencing cycles to filter reads can result in the unnecessary exclusion of high-quality reads that could have provided accurate data if evaluated across an entire sequencing run or using a more reliable approach.

[0006] Due in part to the premature filtering out of high-quality reads by chastity filters, existing sequencing systems may suffer from less accurate variant calling or other downstream analysis. More particularly, the premature removal of high-quality reads based on a limited set of sequencing cycles may result in data biases that skew data for variant calling or other downstream analyses. For example, existing sequencing systems may exclude high-quality reads that align to specific genomic regions of a reference genome based on a chastity filter, including nucleotide reads that align or map to difficult-to-call regions. Such exclusion may be particularly problematic for genomic regions that already have low coverage or are difficult to sequence. Because manyvariant calling models rely on the depth and quality of coverage to accurately detect variants from a sample’s nucleotide reads, the loss of high-quality nucleotide reads can reduce a confidence in variant calls or lead to missed variant calls or false negative variant calls. Conversely, remaining nucleotide reads that successfully pass filter may disproportionately represent other genomic regions, potentially leading to falsely identified variants or false positive variant calls.

[0007] In addition to erroneously removing some high quality read data or skewing data, a chastity filter depends on data inhibited by technical limitations that negatively impact the accuracy of some existing sequencing systems. For instance, a chastity value can be an inaccurate metric when derived from noisy or ambiguous signal data. To illustrate, a relatively high cluster density on a flow cell or other nucleotide-sample slide can lead to overlapping signals from adjacent clusters, which causes interference and makes it difficult to accurately distinguish a brightest nucleobase intensity from a second brightest nucleobase intensity. Additionally, intensity of fluorescent tags from incorporated nucleobases can vary significantly due to fluctuations in the efficiency of chemical reactions, instrument performance, and other factors. These and other factors contribute to chastity frequently being a noisy feature. Accordingly, some existing sequencing systems erroneously filter out high quality clusters based on inaccurate chastity data.

[0008] These, along with additional problems and issues exist in existing sequencing systems.SUMMARY

[0009] This disclosure describes one or more embodiments of systems, methods, and non- transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. For example, the disclosed systems use features derived from signal values, including signal-to-noise ratio (SNR) or base-call-quality scores (e.g., Q- scores), from sequencing cycles corresponding to various base positions within a nucleotide read to generate cluster-filtering scores. Such cluster-filtering scores measure a quality or reliability of signals from an oligonucleotide cluster as a basis for passing data from the oligonucleotide cluster’s nucleotide reads through for variant calling.

[0010] In contrast to a chastity -based approach to filtering, the disclosed system can use signalbased features from different sequencing cycles of paired-end reads to increase the proportion of nucleotide reads that pass filter. In particular, the disclosed system can access signal information from sequencing cycles representing several base positions of nucleotide reads. For instance, the disclosed system may access and use signal-based features (e.g., features derived from SNR or base-call-quality scores) from such sequencing cycles to determine a cluster-filtering score for a cluster of oligonucleotides. The disclosed systems can determine whether the cluster-filtering score satisfies a cluster-filtering threshold score to determine whether a cluster passes filter. Thedisclosed system may further output base-call data files that indicate which reads or clusters pass filter.

[0011] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The detailed description refers to the drawings briefly described below.

[0013] FIG. 1 illustrates an environment in which a cluster-filtering system can operate in accordance with one or more embodiments of the present disclosure.

[0014] FIG. 2 illustrates the cluster-filtering system generating a cluster-filtering score for a cluster of oligonucleotides in accordance with one or more embodiments of the present disclosure.

[0015] FIG. 3 illustrates the cluster-filtering system determining a cluster filtering threshold score in accordance with one or more embodiments of the present disclosure.

[0016] FIG. 4 further illustrates the cluster-filtering system determining a cluster-filtering score based on a base-call-quality score in accordance with one or more embodiments of the present disclosure.

[0017] FIG. 5 illustrates different methods by which the cluster-filtering system generates cluster-filtering scores based on base-call-quality scores in accordance with one or more embodiments of the present disclosure.

[0018] FIG. 6 illustrates improvements in percentage of clusters passing filter and quality of clusters passing filter when using cluster-filtering scores based on base-call-quality scores in accordance with one or more embodiments of the present disclosure.

[0019] FIG. 7 illustrates the cluster-filtering system determining a signal-to-noise ratio-based cluster-filtering score in accordance with one or more embodiments of the present disclosure.

[0020] FIG. 8 illustrates tradeoffs for the cluster-filtering system selecting different cycles for evaluating SNR for the first nucleotide read or the second nucleotide read in accordance with one or more embodiments of the present disclosure.

[0021] FIG. 9 illustrates the cluster-filtering system improving the percentage of clusters passing filter while maintaining high-quality reads relative to a chastity filter in accordance with one or more embodiments of the present disclosure.

[0022] FIGS. 10A-10B illustrate charts quantifying FP and FN variant calls of the clusterfiltering system in various genomic regions relative to existing sequencing systems in accordance with one or more embodiments of the present disclosure.

[0023] FIG. 11 illustrates a flowchart of a series of acts for generating a base-call data file based on the cluster-filtering score satisfying the cluster-filtering threshold score in accordance with one or more embodiments of the present disclosure.

[0024] FIG. 12 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION

[0025] This disclosure describes embodiments of a cluster-filtering system that can use features derived from signal values, including signal-to-noise ratio (SNR) or base-call-quality scores (e.g., Q-scores), from sequencing cycles corresponding to various base positions of a nucleotide read to generate cluster-filtering scores specific to oligonucleotide clusters. Such cluster-filtering scores measure a quality or reliability of signals from an oligonucleotide cluster as a basis for passing data from the oligonucleotide cluster’s nucleotide reads through for variant calling or other downstream analysis for a sample. In contrast to a chastity-based approach to filtering that uses information from a set of initial sequencing cycles (e.g., the first 25 sequencing cycles), the disclosed clusterfiltering system can use signal-based features from different sequencing cycles of paired-end reads to increase the proportion of quality nucleotide reads that pass filter — without degrading read quality and without decreasing variant-calling accuracy. In particular, the cluster-filtering system can use information from sequencing cycles representing several base positions within nucleotide reads and use signal-based features (e.g., SNR or Q-scores) from such cycles to inform decisions on which nucleotide reads pass filter.

[0026] In some embodiments, for example, the cluster-filtering system accesses, for an oligonucleotide cluster, a set of signal values from a set of sequencing cycles for a sequencing run. The cluster-filtering system may further determine a cluster-filtering score for the oligonucleotide cluster based on a subset of signal values of the set of signal values from a sequencing cycle for a first nucleotide read (e.g., Rl) of the oligonucleotide cluster and a sequencing cycle for a second nucleotide read (e.g., R2) of the oligonucleotide cluster. The cluster-filtering system may determine that the cluster-filtering score for the oligonucleotide cluster satisfies a cluster-filtering threshold score and generate, based on the cluster-filtering score satisfying the cluster-filtering threshold score, a base-call data file comprising base calls for the first or second nucleotide reads of the oligonucleotide cluster.

[0027] The cluster-filtering system can use different methods or features to generate filtering thresholds, each relying on a different signal-based-discriminating feature that differs from limited features of a chastity-based-filtering approach. For instance, the cluster-filtering system can use (i) base-call-quality score(s) for base calls within an oligonucleotide cluster’s paired-end reads as a discriminating feature or (ii) signal-to-noise ratio(s) determined at different sequencing cycles forthe paired-end reads as the discriminating feature. As described further below, base-call-quality scores comprise metrics that measure the accuracy of individual base calls within a sequencing run. By contrast, a signal-to-noise ratio (SNR) can indicate the quality of a signal obtained during the sequencing process. The cluster-filtering system may utilize base-call-quality scores and / or SNRs from both read mates of a paired-end read to generate a cluster-filtering score for each cluster.

[0028] In further contrast to a chastity-based-filtering approach, for instance, the clusterfiltering system may use base-call-quality scores from sequencing cycles of both read mates as the discriminating feature in generating a cluster-filtering score. To illustrate, in some embodiments, the cluster-filtering system evaluates base-call-quality scores for a set of sequencing cycles within a sequencing run to generate a cluster-filtering score specific to an oligonucleotide cluster. In a sequencing run with 302 total sequencing cycles, for example, the cluster-filtering system can evaluate some or all 302 sequencing cycles for base-call-quality scores as a basis for filtering, where each paired-end read (e.g., R1 and R2) corresponds with 151 sequencing cycles and at least one base-call-quality score for a first nucleotide read of a paired-end read and at least one base- call-quality score for a second nucleotide read of the paired-end read form a basis for a clusterfiltering score. Instead of using information from a limited number of sequencing cycles for a single read of a paired-end read, therefore, the cluster-filtering system can gather base-call-quality information for either a full set of sequencing cycles or a base call from each paired-end read reflecting data for regions of a nucleotide read. Based on base-call-quality scores for base calls of both a first nucleotide read and a second nucleotide read, therefore, the cluster-filtering system determines a cluster-filtering score for the relevant oligonucleotide cluster.

[0029] In addition or alternative to the base-call-quality score approach, the cluster-filtering system may determine a first SNR as measured in at least one sequencing cycle for a first nucleotide read (e.g., Rl) and a second SNR as measured in at least one sequencing cycle for a second nucleotide read (e.g., R2) use the first and second SNRs as the discriminating feature in generating a cluster-filtering score. To illustrate, the cluster-filtering system stores the SNRs across sequencing cycles up to a particular sequencing cycle (e.g., cycle 10, cycle 50, cycle 75) for a first nucleotide read of paired-end reads. At the same sequencing cycle for the second nucleotide read of the paired- end reads, the cluster-filtering system accesses the saved SNRs for the first nucleotide read (e.g., Rl) and loads the SNRs for the second nucleotide read (e.g., R2) to calculate an SNR-based clusterfiltering score. Accordingly, the cluster-filtering system can gather signal-to-noise ratio information across different sequencing cycles that are representative of different regions of a nucleotide read to determine an SNR-based cluster-filtering score.

[0030] As mentioned, the cluster-filtering system provides various advantages over existing sequencing systems. The cluster-filtering system incorporates significant technical improvementsto a special-purpose computing device — that is, a sequencing device — by improving throughput of quality reads for a sample without decreasing the quality of those reads or negatively impacting variant-calling accuracy. In contrast to existing sequencing systems that often rely on information from a limited number of sequencing cycles to fdter clusters and base calls based on a chastity value, the cluster-filtering system can collect base-call-quality, signal-to-noise ratio, and / or other signal-based-distinguishing information from at least one sequencing cycle of a first nucleotide read and at least one sequencing cycle of a corresponding second nucleotide read — where such cycles corresponds to different regions of nucleotide reads — and use such information to determine cluster-filtering scores for specific clusters that result in higher throughput of quality nucleotide reads. In some examples, and as set forth in greater detail below, the cluster-filtering system can increase a number of passing-filter quality clusters by approximately 2% of millions or billions of clusters, such as by increasing a number of passing-filter quality clusters within a single tile of a nucleotide-sample slide by approximately 90,000 to 100,000 clusters relative to a chastity-filter- based approach. Such increased throughput is illustrated by FIG. 9 and other portions of this disclosure. The cluster-filtering system can provide a more comprehensive assessment of read quality throughout the entire sequencing. By factoring in the base-call-quality scores or signal-to- noise ratios from (i) at least a sequencing cycle for a first nucleotide read of an oligonucleotide cluster and (ii) at least a sequencing cycle for a second nucleotide read — thereby accounting for signal-based features across greater lengths of a nucleotide read — the cluster-filtering system can determine a cluster-filtering score specific to the oligonucleotide cluster that more accurately distinguishes between high-quality reads and reads that degrade in quality over time. This holistic approach can minimize the premature exclusion of high-quality reads that might temporarily fluctuate in early cycles but stabilize in later cycles, ensuring that more reliable data is retained.

[0031] In addition or alternative to higher quality read throughput, the cluster-filtering system incorporates technical improvements to sequencing devices and enhances the accuracy of variant calling or other downstream analysis. By determining base-call-quality scores or signal-to-noise ratios for an oligonucleotide cluster from (i) at least a sequencing cycle for a first nucleotide read of the oligonucleotide cluster and (ii) at least a sequencing cycle for a second nucleotide read, the cluster-filtering system can determine a cluster-filtering score that increases the number of clusters that pass filter without negatively impacting read quality. Because of an increase in the throughput of quality reads for a sample, the cluster-filtering system recovering quality sequencing data that existing sequencing systems typically discard. The cluster-filtering system can use such recovered sequencing data to improve metrics associated with variant calling or other parts of secondary analysis. More specifically, by increasing the number of high-quality reads that pass filter, the cluster-filtering system can ensure better coverage and depth across genomic regions. By retaininggreater numbers of reads that are consistently high quality throughout their length, the clusterfiltering system can reduce false positive variant calls and false negative variant calls. This, in turn, leads to more precise identification of single nucleotide polymorphisms (SNPs), indels, and structural variants, which enhances the overall reliability of genetic analyses.

[0032] Beyond higher quality read throughput and increased variant-calling accuracy, the cluster-filtering system can more accurately distinguish higher versus lower quality read data relative to existing systems by utilizing more reliable metrics, such as base-call-quality scores and SNR, to filter nucleotide reads. Base-call-quality scores and SNR are often less noisy metrics than chastity signals. More particularly, base-call-quality scores can provide a relatively more precise, probabilistic measure of the accuracy of each base call by evaluating the confidence of nucleotide predictions at a given read position. Base-call-quality scores often offer a detailed and granular view of the quality of sequencing data and can offer a more stable metric for determining read quality than a relative brightness of nucleobase intensity in an initial set of sequencing cycles. Similarly, SNR is a robust metric that evaluates the overall clarity of a signal emitted by oligonucleotides incorporated into a cluster by comparing the intensity of the desired signal to background noise. By determining base-call-quality scores or SNRs for an oligonucleotide cluster from sequencing cycles for at least both a first nucleotide read and a second nucleotide read of the oligonucleotide cluster, therefore, the cluster-filtering system can determine a cluster-filtering score that better identifies quality nucleotide reads and cluster data that satisfies a cluster-filtering threshold score. Unlike chastity filters that primarily rely on early-cycle data and can be susceptible to transient issues affecting those initial cycles, SNR can account for greater portions of a nucleotide read, offering a more consistent and reliable assessment of signal quality.

[0033] As suggested by the foregoing discussion, this disclosure utilizes a variety of terms to describe features and benefits of the cluster-filtering system. Additional detail is hereafter provided regarding the meaning of these terms as used in this disclosure. As used in this disclosure, for instance, the term “sample” refers to a specimen, culture, or the like that is suspected of including a target nucleic acid. In some embodiments, the sample comprises DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acids as targets. The sample can likewise include any biological, clinical, surgical, agricultural- atmospheric, or aquatic-based specimen containing one or more nucleic acids. A sample also includes any isolated or extracted nucleic acid sample from an organism, such a genomic DNA, fresh-frozen, or formalin-fixed paraffin-embedded nucleic acid specimen. In some cases, accordingly, a sample can include a full genome or partial genome that is isolated or extracted (e.g., in whole or in part by a kit) from an organism and that is prepared to undergo sequencing or an assay in a sequencing device. A sample can be from a single individual, a collection of nucleicacid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.

[0034] The sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0035] As further used herein, the term “cluster of oligonucleotides” (or simply “cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-pattemed flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.

[0036] As used herein the term “signal” refers to a signal emitted, reflected, or otherwise communicated from a labeled nucleotide base or a group of labeled nucleotide bases (e.g., labeled nucleotide bases added to a cluster of oligonucleotides). In particular, a signal can refer to a signal indicating the type of base. For example, a signal can include a light signal emitted or reflectedfrom a fluorescent tag of a nucleotide base or fluorescent tags of multiple nucleotide bases incorporated into oligonucleotides. In some implementations, the cluster-fdtering system triggers the signal through an external stimulus, such as a laser or other light source. In some cases, the cluster-filtering system triggers the signal through some internal stimuli. Further, in some embodiments, the cluster-filtering system observes the signal using a filter applied when capturing an image of the nucleotide-sample slide (e.g., section of the nucleotide-sample slide). As suggested above, in certain instances, a signal includes an aggregate of the signals provided by each labeled nucleotide base added to individual oligonucleotides in a cluster of oligonucleotides.

[0037] As used herein, the term “signal value” refers to a value indicating a characteristic or attribute of a signal emitted, reflected, or otherwise communicated from a labeled nucleobase or a group of labeled nucleobases from a cluster of oligonucleotides. In particular, a signal value can refer to a value associated with a color intensity (e.g., wavelength) or a light intensity (e.g., brightness). In some cases, the cluster-filtering system captures several images of a cluster of oligonucleotides with labeled nucleobases using different channels. Thus, a signal value of a signal can correspond to the intensity of the signal as observed through a particular channel.

[0038] As used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample (e.g., genomic sample). In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device (including an imaging device, such as a CCD or CMOS) that incorporate nucleobases into growing oligonucleotides to determine nucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell or other nucleotide-sample slide. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more genomic samples seeded in clusters throughout a flow cell. Upon completing a sequencing run, a sequencing device can generate base-call data in a file, such as a binary base call (BCL) sequence file or a fast-all quality (FASTQ) file.

[0039] As used herein, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample’s sequence (e.g., a genomic or transcriptomic sequence from a sample) or a corresponding adapter sequence. In some cases, a sequencing cycle includes an iteration of both incorporating nucleobases into clusters of oligonucleotides using sequencing chemistry and capturing images of such clusters attached to a nucleotide-sample slide (e.g., a flow cell). Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer (e.g., a sample genomic sequence). For example, in one or more embodiments, each sequencing cycle involves incorporating nucleobases into either a single nucleotide read in which DNA or RNA strands areread in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends but in different cycles. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining a particular nucleobase added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleobases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within an SBS run.

[0040] A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For instance, one cluster of oligonucleotides or a set of clusters of oligonucleotides may be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated and another cluster of oligonucleotides or another set of clusters of oligonucleotides may be concurrently undergoing an indexing cycle in which nucleobases corresponding to an indexing sequence for a nucleotide read are incorporated. In some cases, a sequencing device progresses through sequencing cycles that determine nucleobase calls for a nucleotide read from an oligonucleotide comprising an adapter sequence (e.g., p5 / p7 primers), a first indexing sequence, a sample genomic sequence (e.g., gDNA), a second indexing sequence, and another adapter sequence (e.g., p5 / p7 primers).

[0041] As further used herein, the term “genomic sequencing cycle” refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample genomic sequence (or cDNA sequence). In particular, a genomic sequencing cycle can include an iteration of capturing and analyzing one or more images with data indicating individual nucleobases added or incorporated into an oligonucleotide or to oligonucleotides (in parallel) representing or corresponding to one or more sample genomic sequences. Such image analysis can include analyzing data from signals output from an image sensor (e.g., an area capture sensor or a time delayed integration (TDI) sensor). For example, in one or more embodiments, each genomic sequencing cycle involves capturing and analyzing images to determine either single reads or paired-end reads of DNA (or RNA) strands representing part of a genomic sample (or transcribed sequence from a genomic sample). As suggested above, however, a genomic sequencing cycle, in some cases, is specific to a cluster of oligonucleotides or a set of clusters of oligonucleotides.

[0042] By contrast, the term “indexing cycle” refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to one or more indexing sequences. In particular, an indexing cycle can include an iteration of capturing and analyzing one or more images of clusters of oligonucleotides indicating one or more nucleobasesadded or incorporated into an oligonucleotide or to oligonucleotides (in parallel) representing or corresponding to one or more indexing sequences. An indexing cycle differs from a genomic sequencing cycle in that an indexing cycle includes sequencing of at least a nucleobase (or a majority of nucleobases) from one or more indexing sequences that identify or encode one or more sample library fragments. Because genomic sequencing cycles may be specific to a cluster or clusters of oligonucleotides or other structures of oligonucleotides, an indexing cycle for one cluster of oligonucleotides may be performed at a same time as a genomic sequencing cycle for another cluster of oligonucleotides.

[0043] As used herein, the term “nucleotide read” (or simply “read”) refers to an inferred or predicted sequence of one or more nucleobases (or nucleobase pairs) from all or part of a sample genomic sequence (e.g., a sample genomic sequence, complementary DNA). In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a nucleotide fragment (or group of monoclonal nucleotide fragments) from a sequencing library corresponding to a genomic sample. For example, in some embodiments, the cluster-filtering system determines a nucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a well in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g., assembled nucleotide reads) satisfying a threshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.

[0044] Relatedly, and as used herein, the terms “first nucleotide read” and “second nucleotide read” refer to predicted sequences obtained from forward and reverse strands of a sample genomic sequence. In particular, in the context of paired-end sequencing, the first nucleotide read and the second nucleotide read refer to two nucleotide reads obtained from each end of a nucleotide fragment from a sequencing library corresponding to a genomic sample. In some examples, the cluster-filtering system sequences the nucleotide fragment from an initial end through a first subset of sequencing cycles to produce the first nucleotide read. After generating the first nucleotide read, the cluster-filtering system sequences the nucleotide fragment from the other end through a second subset of sequencing cycles to produce the second nucleotide read. Together, the first nucleotide read and the second nucleotide read can provide a more comprehensive view of a nucleotide fragment. In some examples, a first nucleotide read is expressed or represented as “Rl” and the second nucleotide read is expressed as “R2.”

[0045] As used herein, the term “cluster-filtering score” refers to a measure of quality or reliability of one or more signals from a cluster of oligonucleotides as a basis for filtering nucleotide-read data. In particular, a cluster-filtering score includes a filtering metric that measures a quality or reliability of signals from an oligonucleotide cluster based on features of signal values for (i) at least one sequencing cycle for a first nucleotide read (e.g., Rl) (ii) at least one sequencing cycle for a second nucleotide read (e.g., R2). Such features of signal values can include, but are not limited to, signal-to-noise ratio (SNR) or base-call-quality scores (e.g., Q-scores). For example, a cluster-filtering score can be assigned to (or determined for) a cluster of oligonucleotides and one or more of the cluster’s corresponding nucleotide reads during a sequencing run and helps determine whether a nucleotide read corresponding to a given cluster of oligonucleotides meets criteria to pass a filtering stage for data from the corresponding nucleotide reads to be used for variant calling or other downstream processes. In some examples, the cluster-filtering system determines a cluster-filtering score based on a subset of signal values at particular sequencing cycles. More specifically, the cluster-filtering score can reflect the reliability of signals across a proportion of or all sequencing cycles within a sequencing run. For example, a cluster-filtering score can reflect base-call-quality scores and / or SNR across a proportion of or all sequencing cycles within a sequencing run.

[0046] As used herein, the term “cluster-filtering threshold score” refers to a value used as a benchmark for classifying or filtering data from a particular cluster of oligonucleotides based on cluster-filtering scores. In particular, nucleotide reads having cluster-filtering scores that meet or exceed a cluster-filtering threshold score are considered to pass filter. For example, clusters or their corresponding nucleotide reads that pass filter can be considered of sufficient quality to be used in subsequent data processing and analysis steps, such as mapping, alignment, and variant calling. In some examples, reads having cluster-filtering scores that fall below a cluster-filtering threshold score can be excluded from further analysis due to concerns about accuracy and reliability. Further, in some examples, the cluster-filtering system dynamically determines a cluster-filtering threshold score to meet a target number of clusters that pass filter.

[0047] As further used herein, the term “base-call data file” refers to a digital file or other digital information indicating individual nucleobases or the sequence of nucleobases for a nucleic-acid polymer. In particular, a base-call data file can include nucleotide reads comprising nucleobase calls for particular genomic samples. Base-call data files can include intensity values (e.g., color or light intensity values for individual clusters) from images taken by a camera of a nucleotide-sample slide or other data that indicate individual nucleobases or the sequence of nucleobases for a nucleic- acid polymer. In addition, or in the alternative to intensity values, a base-call data file may include chromatogram peaks or electrical current changes indicating individual nucleobases in a sequence.Additionally, in some embodiments, a base-call data file includes individual nucleobase calls identifying the individual nucleobases (e.g., A, T, C, or G). For example, a base-call data file can comprise data for nucleobase calls in a sequence for a nucleic-acid polymer, the number of nucleobase calls corresponding to a particular base (e.g., adenine, cytosine, thymine, or guanine), as organized in a digital file, such as a Binary Base Call (BCL) file or a Fast- All Q (FASTQ) file. The format of the base-call data file can vary based upon the sequencing technology used and can include BCF, BAM, and QSEQ, as well as other formats. Further, a base-call data file can include error / accuracy information, such as a quality metric associated with each nucleobase call. In some embodiments, the base-call data comprises information from a sequencing device that utilizes sequencing by synthesis (SBS).

[0048] As used herein, the term “nucleobase call” (or “nucleotide-base call” or simply “base call”) refers to a determination or prediction of a particular nucleobase (or nucleobase pair) for an oligonucleotide (e.g., nucleotide read) during a sequencing cycle or for a genomic coordinate of a genomic sample. In particular, a nucleobase call can indicate a determination or prediction of the type of nucleobase that has been incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., read-based nucleobase calls). In some cases, for a nucleotide read, a nucleobase call includes a determination or a prediction of a nucleobase based on intensity values resulting from fluorescent-tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., in a cluster of a flow cell). As suggested above, a single nucleobase call can be an adenine (A) call, a cytosine (C) call, a guanine (G) call, a thymine (T) call, or an uracil (U) call. Note that the terms “nucleobase” and “nucleotide base” are interchangeable.

[0049] As used herein, the term “chastity value” (sometimes referred to as a “chastity metric”) refers to a purity of a signal from a called nucleobase for a cluster of oligonucleotides relative to signals from other nucleobases. In particular, a chastity value refers to the ratio of the brightest nucleobase intensity of a signal and the sum of the brightest and second brightest nucleobase intensities thereof. For example, a chastity value can be determined for a corresponding signal as the ratio of a distance between the intensity associated with the signal and the nearest nucleobase centroid to a distance between the intensity and another centroid (e.g., the second nearest centroid). In some cases, sequencing systems use a chastity filter to retain clusters in which the signal from the called nucleobase is significantly stronger than the signals from the other three nucleobases, ensuring the purity of signal for the called nucleobase. To illustrate, if a chastity filter is set to a threshold of 0.6, only nucleobase calls with a chastity value of 0.6 or higher are considered pure and included in a base-call data file.

[0050] Additionally, as used herein, the term “signal-to-noise ratio” (or “SNR” or “signal-to- noise-ratio metric”) refers to a measure of a target signal compared to a level or content of noise.In particular, a signal-to-noise ratio includes to the strength of a light signal that is detected from labeled nucleotide bases in a cluster of oligonucleotides compared to associated noise. For example, in some implementations, a signal-to-noise ratio likewise includes a ratio of a scaling factor associated with a signal compared to the corresponding noise level. In one or more embodiments, the scaling factor determined for a light signal can be equated to the light signal itself (e.g., the signal purity without the addition of noise). In some embodiments, signal-to-noise ratio can be calculated using signal data from a single sequencing cycle. In other embodiments, signal-to-noise ratio is calculated using signal data from consecutive sequencing cycles.

[0051] As used herein, the term “base-call-quality score” (sometimes referred to as a “quality score”) refers to a specific score or other measurement indicating an accuracy of a nucleobase call. In particular, a base-call-quality score comprises a value indicating a likelihood that one or more predicted nucleobase calls for a genomic coordinate contain errors. For example, in certain implementations, a base-call-quality score can comprise a Q score (e.g., a Phil’s Read Editor (PhRED)-scaled quality score) predicting the error probability of any given nucleobase call. To illustrate, a base-call-quality score (or Q score) may indicate that a probability of an incorrect nucleobase call at a genomic coordinate is equal to 1 in 100 for a Q20 score, 1 in 1,000 for a Q30 score, 1 in 10,000 for a Q40 score, etc. In some cases, the base-call-quality score is generated by a machine-learning model or an algorithm, either of which can be scaled to be consistent with a PhRED scale.

[0052] The following paragraphs describe the cluster-filtering system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which a cluster-filtering system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server device(s) 110, and a client device 114. As shown in FIG. 1, the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114 can communicate with each other via a network 118. The network 118 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 12. While FIG. 1 shows an embodiment of the cluster-filtering system 106, this disclosure describes alternative embodiments and configurations below.

[0053] As indicated by FIG. 1, the sequencing device 102 comprises a computing device and a sequencing device system 104 for sequencing a genomic sample or other nucleic-acid polymer. In some embodiments, by executing the sequencing device system 104 using a processor, the sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods andsystems either directly or indirectly on the sequencing device 102. More particularly, the sequencing device 102 receives nucleotide-sample slides (e.g., flow cells) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments.

[0054] In one or more embodiments, the sequencing device 102 utilizes sequencing-by- synthesis (SBS) techniques to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition or in the alternative to communicating across the network 118, in some embodiments, the sequencing device 102 bypasses the network 118 and communicates directly with the local device 108, and / or the client device 114. By executing the sequencing device system 104, the sequencing device 102 can further store the nucleobase calls as part of a base-call data file that is formatted as a binary base call (BCL) file and / or a FASTQ file and send the BCL file and / or FASTQ file to the local device 108, and / or the server device(s) 110.

[0055] As further indicated by FIG. 1, the local device 108 is located at or near a same physical location of the sequencing device 102. Indeed, in some embodiments, the local device 108 and the sequencing device 102 are integrated into a same computing device. The local device 108 may run the sequencing device system 104 and / or the cluster-filtering system 106 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As shown in FIG. 1, the sequencing device 102 may send (and the local device 108 may receive) base-call data generated during a sequencing run of the sequencing device 102. The local device 108 may also communicate with the client device 114. In particular, the local device 108 can send data to the client device 114, including a binary alignment map (BAM) file, a variant call format (VCF) file, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.

[0056] As further indicated by FIG. 1, the server device(s) 110 are located remotely from the local device 108 and the sequencing device 102. Like the local device 108, in some embodiments, the server device(s) 110 include a version of (or are otherwise able to access or implement) the cluster-filtering system 106. For example, the server device(s) 110 can implement the clusterfiltering system 106 as part of a sequencing system 112. Accordingly, the server device(s) 110 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As indicated above, the sequencing device 102 may send (and the server device(s) 110 may receive) base-call data from the sequencing device 102. The server device(s) 110 may also communicate with the client device 114. In particular, the server device(s) 110 can send data to the client device 114, including BAM files, VCF files, or other sequencing related information.

[0057] In some embodiments, the server device(s) 110 comprise a distributed collection of servers where the server device(s) 110 include a number of server devices distributed across the network 118 and located in the same or different physical locations. Further, the server device(s) 110 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0058] As further illustrated and indicated in FIG. 1, by executing a sequencing application 116, the client device 114 can generate, store, receive, and send digital data. In particular, the client device 114 can receive sequencing data from the local device 108 or receive base-call data files (e.g., BCL and FASTQ) and sequencing metrics from the sequencing device 102. Furthermore, the client device 114 may communicate with the local device 108 or the server device(s) 110 to receive a VCF comprising genotype or variant calls and / or other metrics, such as base-call-quality metrics or pass-filter metrics. The client device 114 can accordingly present or display information pertaining to variant calls or other genotype calls within a graphical user interface of the sequencing application 116 to a user associated with the client device 114. For example, the client device 114 can present nucleobase calls, genotype calls, variant calls, and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of the sequencing application 116.

[0059] Although FIG. 1 depicts the client device 114 as a desktop or laptop computer, the client device 114 may comprise various types of client devices. For example, in some embodiments, the client device 114 includes non -mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 114 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 114 are discussed below with respect to FIG. 12.

[0060] As further illustrated in FIG. 1, the client device 114 includes the sequencing application 116. The sequencing application 116 may be a web application or a native application stored and executed on the client device 114 (e.g., a mobile application, desktop application). The sequencing application 116 can include instructions that (when executed) cause the client device 114 to receive data from the cluster-filtering system 106 and present, for display at the client device 114, base-call data or data from an alignment data file or VCF. Furthermore, the sequencing application 116 can instruct the client device 114 to display summaries for multiple sequencing runs.

[0061] As further illustrated in FIG. 1, a version of the cluster-filtering system 106 may be located and / or implemented (e.g., entirely or in part) on the client device 114 or the sequencing device 102. In yet other embodiments, the cluster-filtering system 106 is implemented by one or more other components of the computing system 100, such as the local device 108. In particular, the cluster-filtering system 106 can be implemented in a variety of different ways across thesequencing device 102, the local device 108, the server device(s) 110, and the client device 114. For example, the cluster-filtering system 106 can be downloaded from the server device(s) 110 to the client device 114 and / or the local device 108 where all or part of the functionality of the clusterfiltering system 106 is performed at each respective device within the computing system 100.

[0062] As mentioned previously, in some embodiments, the cluster-filtering system 106 improves the percentage of clusters passing filter by incorporating information from different sequencing cycles corresponding to various base positions within a nucleotide read. For instance, the cluster-filtering system 106 can determine a cluster-filtering score for an oligonucleotide cluster based on base-call-quality scores and / or signal-to-noise ratios. In accordance with one or more embodiments of the present disclosure, FIG. 2 depicts an overview diagram of the cluster-filtering system 106 generating a cluster-filtering score for a cluster of oligonucleotides. As illustrated in FIG. 2, the cluster-filtering system 106 performs a series of acts 200 comprising an act 202 of accessing, for an oligonucleotide cluster, a set of signal values from a set of sequencing cycles; an act 204 of determining a cluster-filtering score for the oligonucleotide cluster; an act 206 of determining that the cluster-filtering score satisfies a cluster-filtering threshold score; and an act 208 of generating a base-call data file. The following paragraphs further describe the acts depicted in FIG. 2 as part of an overview.

[0063] As shown in FIG. 2, for example, the cluster-filtering system 106 performs the act 202 of accessing, for an oligonucleotide cluster, a set of signal values from a set of sequencing cycles. As illustrated, the cluster-filtering system 106 utilizes a nucleotide-sample slide 210 for sequencing of samples. As described above, the nucleotide-sample slide 210 can include oligonucleotides that receive or incorporate labeled nucleotide bases during a given sequencing cycle. In particular, the nucleotide-sample slide 210 can include a cluster of oligonucleotides with each section (e.g., a tile comprising wells of a flow cell) thereof. When stimulated, the labeled nucleotide bases can emit a signal having characteristics associated with the type of nucleotide base.

[0064] As further shown in FIG. 2, the cluster-filtering system 106 captures a series of images 211 of at least one section of the nucleotide-sample slide 210, such as a section corresponding to an individual cluster of oligonucleotides and / or multiple oligonucleotide clusters. The clusterfiltering system 106 captures the series of images 211 as the labeled nucleotide bases within a cluster of oligonucleotides emit a respective series of signals. As illustrated, the cluster-filtering system 106 captures multiple images for each sequencing cycle in a set of sequencing cycles 212. For example, in some embodiments, the cluster-filtering system 106 utilizes a two-channel implementation or a four-channel implementation to capture two or four different images of the section of the nucleotide-sample slide 210 for each sequencing cycle in the set of sequencing cycles 212.

[0065] To illustrate signals from oligonucleotide clusters, FIG. 2 depicts a set of signals 214 emitted from the labeled nucleotide bases during the set of sequencing cycles 212. As mentioned, the set of signals 214 can indicate the type of nucleotide base that was added to the cluster of oligonucleotides for a given sequencing cycle of the set of sequencing cycles 212. For example, as described in additional detail below, a signal within the set of signals 214 can have one or more corresponding intensity values that indicate the corresponding type of nucleotide base.

[0066] As further illustrated in FIG. 2, the cluster-fdtering system 106 performs the act 204 of determining a cluster-fdtering score for the oligonucleotide cluster. As mentioned previously, the cluster-fdtering system 106 can assess the quality of reads based on signals from at least one sequencing cycle of a first nucleotide read and at least one sequencing cycle of a corresponding second nucleotide read — including an extended range of sequencing cycles to incorporate information from different base positions within paired-end reads. For instance, the cluster-fdtering system 106 identifies a subset of signal values of the set of signals 214 that include signal information from at least one sequencing cycle of a first nucleotide read and at least one sequencing cycle of a corresponding second nucleotide read. Accordingly, in some examples, the subset of signal values comprises a proportion or fraction of the set of signals 214. In these examples, the cluster-fdtering system 106 uses information from select base positions of a nucleotide read to evaluate whether the nucleotide read will pass filter. In some examples, by contrast, the cluster- fdtering system 106 uses signal values from all sequencing cycles within a sequencing run as a basis for a cluster-fdtering score. Regardless of whether all or a subset of sequencing cycles is used, the cluster-fdtering system 106 uses data from different base positions (relative to those positions used for a chastity value) of both a first nucleotide read and a second nucleotide read of paired-end reads to determine whether the nucleotide reads of the oligonucleotide cluster pass filter.

[0067] As indicated above, in some implementations, the cluster-fdtering system 106 determines cluster-fdtering scores for an oligonucleotide cluster and its corresponding nucleotide reads as part of paired-end sequencing. For example, the cluster-filtering system 106 can determine or identify a sequencing cycle for a first nucleotide read 216 at which to collect and evaluate data for a first nucleotide read. The cluster-fdtering system 106 may also determine or identify a sequencing cycle for a second nucleotide read 218 at which to collect and evaluate data for a second nucleotide read. Accordingly, the sequencing cycle for a first nucleotide read 216 can be part of a first subset of sequencing cycles corresponding to the first nucleotide read; and the sequencing cycle for a second nucleotide read 218 can be part of a second subset of sequencing cycles corresponding to the second nucleotide read. For instance, a paired-end sequencing run may comprise a total of 302 sequencing cycles where the first nucleotide read is sequenced during a first subset of 151 sequencing cycles and the second nucleotide read is sequenced during a second subsetof 151 sequencing cycles. While FIG. 2 refers to a single sequencing cycle for each of a first nucleotide read and a second nucleotide read, as described further below, the cluster-filtering system 106 may also determine multiple sequencing cycles for each of a first nucleotide read and a second nucleotide read at which to use signal data as a basis for a cluster-filtering score.

[0068] As shown in FIG. 2, as part of performing the act 204, the cluster-filtering system 106 determines a cluster filtering score 220 based on a subset of signal values from at least the sequencing cycle for the first nucleotide read 216 and the sequencing cycle for the second nucleotide read 218. In some examples, the cluster-filtering system 106 generates a single cluster filtering score for the first nucleotide read and the second nucleotide read. As mentioned, the cluster-filtering system 106 may generate the cluster filtering score 220 based on different metrics. In some implementations, for instance, the cluster-filtering system 106 determines the cluster filtering score 220 based on a base-call-quality scores. As described below, FIG. 3 illustrates an overview of the cluster-filtering system 106 determining a cluster filtering score 220 based on base- call-quality scores in accordance with one or more embodiments of the present disclosure. In some implementations, the cluster-filtering system 106 determines the cluster filtering score 220 based on signal -to-noise ratios. FIG. 6 illustrates an overview of the cluster-filtering system 106 determining the cluster filtering score 220 based on signal-to-noise ratios in accordance with one or more embodiments of the present disclosure. The overview depicted by FIG. 2, however, applies to both quality-score-based and SNR-based implementations.

[0069] As further illustrated in FIG. 2, the series of acts 200 includes the act 206 of determining that the cluster-filtering score satisfies a cluster-filtering threshold score. Generally, the clusterfiltering system 106 uses a cluster-filtering threshold score as a benchmark to identify a cluster with nucleotide reads for secondary analysis. As further described below, the cluster-filtering system 106 uses the cluster-filtering threshold score to improve the throughput of sequencing devices by increasing the useful percentage of clusters that pass filter without degrading read quality. In some implementations, the cluster-filtering system 106 receives the cluster-filtering threshold score based on input from a client device. For example, a user may input a cluster-filtering threshold score. In other implementations, the cluster-filtering system 106 determines the clusterfiltering threshold score. FIG. 3 illustrates the cluster-filtering system 106 determining a clusterfiltering threshold score in accordance with one or more embodiments of the present disclosure.

[0070] After determining that the cluster-filtering score satisfies a cluster-filtering threshold score, as further illustrated in FIG. 2, the cluster-filtering system 106 performs the act 208 of generating a base-call data file. For instance, the cluster-filtering system 106 generates a base-call data file comprising base calls for nucleotide reads with cluster-filtering scores that satisfy the cluster-filtering threshold score. More particularly, the cluster-filtering system 106 generates abase-call data file comprising base calls for the first nucleotide read and / or second nucleotide read of the cluster of oligonucleotides having the cluster-filtering score that satisfies the cluster-filtering threshold score. In some examples, the base-call data file comprises raw intensity data captured during the sequencing process. For instance, the base-call data file can comprise a data file in BCL format. In other implementations, the base-call data file comprises selected information for clusters of oligonucleotides that pass filter. For example, the base-call data file may comprise a data file in a converted FASTQ format.

[0071] In addition or alternative to determining a cluster-filtering score and filtering clusters for a base-call data file based on signal values from at least one sequencing cycle of a first nucleotide read and at least one sequencing cycle of a corresponding second nucleotide read, the cluster-filtering system 106 can determine the cluster-filtering score based on signal values from one or both of (i) indexing cycles and (ii) genomic sequencing cycles. In some implementations, for instance, the cluster-filtering system 106 determines the cluster-filtering score based on a subset of signal values from genomic sequencing cycles. For example, the cluster-filtering system 106 uses data from various parts of a genomic read to determine the cluster-filtering score. The genomic sequencing cycles may occur after one or both indexing cycles for both nucleotide reads. For example, in some implementations, a genomic sequencing cycle for a first nucleotide read occurs after the indexing cycle for the first nucleotide read. In other implementations using an indexing- first workflow, the genomic sequencing cycle for the first nucleotide read occurs after indexing cycles for both the first nucleotide read and the second nucleotide read.

[0072] In addition or in the alternative to using genomic sequencing cycles, in other implementations, the cluster-filtering system 106 determines the cluster-filtering score based on a subset of signal values from indexing cycles. Based on indexing cycles and corresponding indexing sequences, the cluster-filtering system 106 can identify which nucleotide reads belong to which samples. For instance, in some cases, the cluster-filtering system 106 accesses raw sequencing data comprising indexing sequences associated with sample genomic sequences. The indexing sequences can comprise “barcodes” that act as unique identifiers for each sample, allowing for differentiation and sorting of nucleotide reads during demultiplexing. The cluster-filtering system 106 demultiplexes nucleotide reads by utilizing a reference of known indexes. By comparing indexing sequences with known indexing sequences in a reference of registered indexes, the cluster-filtering system 106 can identify genomic samples that correspond with one or more unique barcodes. In some implementations, the cluster-filtering system 106 uses an indexing first workflow in which indexing cycles for both nucleotide reads precede genomic sequencing cycles for both reads. More particularly, the cluster-filtering system 106 may demultiplex samples before sequencing the genomic samples. In other implementations, the cluster-filtering system 106 doesnot use the indexing first workflow. In such examples, the cluster-filtering system 106 performs the genomic sequencing cycles after their respective indexing cycles.

[0073] As mentioned, the cluster-filtering system 106 can determine the cluster-filtering score based on a subset of signal values from indexing cycles in either an indexing-first workflow or a more standard workflow. In some implementations, the cluster-filtering system 106 can determine a first cluster-filtering score for an oligonucleotide cluster based on indexing cycles and a second cluster-filtering score for the oligonucleotide cluster based on genomic sequencing cycles. For example, in indexing-first workflows, the cluster-filtering system 106 performs indexing cycles and demultiplexing as part of pre-processing. The cluster-filtering system 106 can determine a first cluster-filtering core based on data obtained in pre-processing but determine to retain all read data. The cluster-filtering system 106 can further determine a second cluster-filtering score for a nucleotide read based on data obtained during subsequent genomic sequencing cycles. The clusterfiltering system 106 can thus use either one or both of the indexing cycles and the genomic sequencing cycles to generate one or more cluster-filtering scores.

[0074] As mentioned, in some embodiments, the cluster-filtering system 106 determines a cluster-filtering threshold score that optimizes or otherwise improves a balance between increasing a percentage of nucleotide reads that pass filter while also retaining high-quality data that satisfies a cluster-filtering threshold score. FIG. 3 illustrates the cluster-filtering system 106 determining such a cluster-filtering threshold score in accordance with one or more embodiments of the present disclosure.

[0075] As indicated above, in some implementations, the cluster-filtering system 106 receives the cluster-filtering threshold score as input from a client device, such as the client device 114. For example, the cluster-filtering system 106 may provide, to the client device, user interface elements for selecting a desired cluster-filtering threshold score. In other embodiments, the cluster-filtering system 106 automatically determines the cluster-filtering threshold score. In some implementations, the cluster-filtering system 106 automatically both determines the clusterfiltering threshold score and provides user interface elements to the client device for modifying or adjusting the automatically determined cluster-filtering threshold score. Regardless of whether a client device is involved, in some embodiments, the cluster-filtering system 106 by titrating or incrementally adjusting a threshold based on cluster-filtering scores relative to a number of oligonucleotide clusters.

[0076] To illustrate such incremental adjustment, FIG. 3 depicts a graph 300 representing a distribution of clusters across different cluster-filtering scores. The x-axis of the graph 300 represents the range of cluster-filtering scores, and the y-axis of the graph 300 represents a number of clusters having the corresponding cluster-filtering scores. As mentioned, the cluster-filteringsystem 106 may dynamically determine the cluster-filtering threshold score by titrating or incrementally adjusting the cluster-filtering threshold over the cluster-filtering scores across a set of clusters of oligonucleotides relative to a number of clusters from a set of clusters of oligonucleotides satisfying the cluster-filtering threshold score. Generally, titrating the clusterfiltering threshold involves adjusting the cluster-filtering threshold score to achieve an optimal or otherwise relatively improved balance between increasing the percentage of clusters that pass filter while maintaining relatively high quality read data supporting accurate variant calls. For purposes of FIG. 3, such a relatively high quality read data includes read data for a sample that maintains (i) an average read-quality score (e.g., Q40) for nucleotides reads of a sample or given genomic region or (ii) results in a maintained or increased level of accuracy in terms of precision, sensitivity, or recall for variant calling in a given genomic region or whole genome. For example, lowering the cluster-filtering threshold score may allow more clusters to pass filter, increasing the overall read data yield and ensuring sufficient coverage for robust analysis for variant calling. Conversely, raising the cluster-filtering threshold may improve the overall read data quality but also reduce the percentage of clusters that pass filter.

[0077] In some implementations, the cluster-filtering system 106 titrates the cluster-filtering threshold by determining a candidate cluster-filtering threshold score, applying the candidate cluster-filtering threshold scores to the set of clusters of oligonucleotides, and evaluating the performance of the candidate cluster-filtering threshold score. As illustrated in FIG. 3, as part of determining the candidate cluster-filtering threshold scores, the cluster-filtering system 106 establishes a range of candidate threshold score values 302. In some implementations, the clusterfiltering system 106 may predetermine the range of candidate threshold score values 302. For example, the cluster-filtering system 106 may predetermine the range of candidate threshold score values 302 based on an acceptable read data quality level. In some implementations, the clusterfiltering system 106 receives the range of candidate threshold score values 302 as input from a client device, range of candidate threshold score values 302. Furthermore, in some examples, the range of candidate threshold score values 302 may encompass an entire range of cluster-filtering scores (e.g., from 0 to 1).

[0078] As further illustrated in FIG. 3, the cluster-filtering system 106 determines a candidate cluster-filtering threshold score. The cluster-filtering system 106 may determine a candidate cluster-filtering threshold score beginning with a conservative cluster-filtering threshold score that yields a smaller percentage of clusters passing filter with relatively high quality read data. For instance, the cluster-filtering system 106 may determine, as a candidate cluster-filtering threshold score, a candidate cluster-filtering threshold score 304a having the highest value within the rangeof candidate threshold score values 302 rather than a candidate cluster-filtering threshold score 304c having a lowest value within the range of candidate threshold score values 302.

[0079] As mentioned, the cluster-filtering system 106 applies the candidate cluster-filtering threshold score to the set of clusters of oligonucleotides. For example, and as shown in FIG. 3, the cluster-filtering system 106 applies the candidate cluster-filtering threshold score 304a to the set of clusters of oligonucleotides. In some cases, accordingly, the cluster-filtering system 106 identifies clusters from the set of clusters of oligonucleotides having cluster-filtering scores that satisfy the candidate cluster-filtering threshold score 304a.

[0080] The cluster-filtering system 106 further evaluates the performance of the candidate cluster-filtering threshold score 304a. Generally, the cluster-filtering system 106 evaluates the performance of the candidate cluster-filtering threshold score 304a based on performance metrics, such as (i) a number of retained reads or reads that pass filter and (ii) the quality of the retained reads. As mentioned, the cluster-filtering system 106 attempts to increase the number of clusters that pass filter while ensuring the clusters that pass filter meet a basic quality standard. For example, the cluster-filtering system 106 can determine a number of clusters that satisfy the candidate cluster-filtering threshold score. The cluster-filtering system 106 can compare the number of clusters with a target number of passing filter clusters. For instance, the target number of passing filter clusters may equal or exceed a number of clusters passing a chastity filter as indicated by an expected pass-filter indicator. As FIGS. 4 and 7 depict or indicate in further detail, the clusterfiltering system 106 may generate an expected pass-filter indicator by checking clusters of oligonucleotides for chastity failures at an evaluation sequencing cycle. In some examples, the cluster-filtering system 106 automatically determines the target number of passing filter clusters. In other examples, the cluster-filtering system 106 receives the target number of passing filter clusters as input from a client device.

[0081] In addition or in the alternative to a number of retained reads or target clusters passing filter as performance metrics, the cluster-filtering system 106 can further evaluate the performance of the candidate cluster-filtering threshold score 304a by analyzing (ii) the quality of the clusters that satisfy the candidate cluster-filtering threshold score 304a. In some implementations, the cluster-filtering system 106 evaluates the quality of the clusters based on a target passing filter quality. The target passing filter quality comprises an expected quality criteria that clusters passing filter are expected to meet. In some implementations, the cluster-filtering system 106 automatically determines the target passing filter quality or receives the target passing filter quality as input from a client device. In one or more embodiments, the cluster-filtering system 106 determines the target passing filter quality based on an expected pass-filter indicator. The expected pass-filter indicator indicates clusters of oligonucleotides that satisfy a chastity filter. To illustrate, the cluster-filteringsystem 106 applies a chastity filter to the set of clusters and estimates the quality of clusters passing filter. The cluster-filtering system 106 aims to increase the percentage or number of clusters that pass filter while maintaining or improving a level of base calling quality of clusters passing the chastity filter. The expected pass-filter indicator may comprise a quality benchmark expressed as an average base-call-quality score (e.g., PhRED-scaled quality score) or another quality measure.

[0082] Based on evaluating the performance of the candidate cluster-filtering threshold score 304a, the cluster-filtering system 106 may select the candidate cluster-filtering threshold score 304a as the cluster-filtering threshold score. Alternatively, the cluster-filtering system 106 may further evaluate an additional candidate cluster-filtering threshold score. For instance, based on determining that a number of clusters satisfying the cluster-filtering threshold score is lower than a target number of clusters, the cluster-filtering system 106 determines to incrementally select more lenient candidate cluster-filtering threshold scores. Based on determining that a quality of passing filter clusters is below target passing filter quality, the cluster-filtering system 106 may determine to incrementally select stricter (e.g., higher) candidate cluster-filtering threshold scores. The target number of clusters may equal a number of clusters greater than clusters that pass a chastity filter. As illustrated in FIG. 3, the cluster-filtering system 106 determines to evaluate a candidate clusterfiltering threshold score 304b. Based on determining that the candidate cluster-filtering threshold score 304b yields a number of clusters passing filter that is lower than the target number of clusters, the cluster-filtering system 106 can further evaluate the candidate cluster-filtering threshold score 304c, etc.

[0083] As mentioned, the cluster-filtering system 106 may determine a cluster-filtering score based on base-call-quality scores as signal-based-discriminating features. FIG. 4 illustrates an overview of the cluster-filtering system 106 determining a base-call-quality score-based clusterfiltering score in accordance with one or more embodiments of the present disclosure. FIG. 4 illustrates acts performed by the cluster-filtering system 106 at various stages within a sequencing run. As described further below, FIG. 4 illustrates acts performed by the cluster-filtering system 106 with reference to an evaluation sequencing cycle, a sequencing cycle for a first nucleotide read 424, and a sequencing cycle for a second nucleotide read 426.

[0084] As illustrated in FIG. 4, the cluster-filtering system 106 optionally performs an act 402 of checking for chastity failures between cycle 1 of a sequencing run and an evaluation sequencing cycle (e.g., cycle 25) of the sequencing run. More specifically, the cluster-filtering system 106 checks for chastity failures by evaluating the quality of nucleotide reads in the sequencing cycles up to and including the evaluation sequencing cycle. The evaluation sequencing cycle comprises a sequencing cycle within a sequencing run at which the cluster-filtering system 106 performs an act 404 of evaluating a chastity filter. In some embodiments, the evaluation sequencing cyclecomprises cycle 25 of a sequencing run. As mentioned previously, the chastity value can represent a ratio of the brightest base signal to the sum of the brightest and second base brightest signals. If this ratio falls below a predefined chastity threshold value (e.g., 0.6), for example, the read is flagged as a chastity failure.

[0085] As illustrated in FIG. 4, the cluster-filtering system 106 may perform an act 406 of outputting an expected pass-filter indicator. The expected pass-filter indicator indicates an anticipated proportion or percentage of nucleotide reads that likely pass a cluster-filtering threshold score based on the chastity filter. For example, the expected pass-filter indicator may equal a proportion or percentage of nucleotide reads that passed the chastity filter at the evaluation sequencing cycle. In some examples, the cluster-filtering system 106 performs the act 406 as part of a quality control process. The cluster-filtering system 106 can use the expected pass-filter indicator as predictive of quality of a sequencing run. To illustrate, a low expected pass-filter indicator (e.g., below 90%) may indicate issues with the sequencing run, such as poor sample quality or other technical problems. Furthermore, in some embodiments, the cluster-filtering system 106 utilizes the chastity filter to output the expected pass-filter indicator for purposes of backwards compatibility. While the cluster-filtering system 106 implements a novel and different cluster-filtering method, and the cluster-filtering system 106 may continue to evaluate a chastity filter to maintain backward compatibility with sequencing systems that necessitate use of a chastity filter.

[0086] In contrast to existing sequencing systems that typically discard data for clusters that fail the chastity filter, in some embodiments, the cluster-filtering system 106 retains data for clusters that fail the chastity filter. As mentioned, the cluster-filtering system 106 may utilize an expected pass-filter indicator (i) as a quality control measure, (ii) as part of determining a clusterfiltering threshold score, and (iii) to ensure backwards compatibility with existing sequencing systems. The cluster-filtering system 106 can retain the data for clusters that fail the chastity filter for evaluation using a cluster-filtering threshold score. By retaining all read data after the evaluation sequencing cycle, the cluster-filtering system 106 can later determine to retain or ignore read data regardless of whether a given cluster passes a chastity filter.

[0087] As further shown in FIG. 4, the cluster-filtering system 106 further performs an act 408 of determining first base-call-quality score(s) at the sequencing cycle for the first nucleotide read 424. As mentioned previously, the cluster-filtering system 106 can generate a first base-call-quality score that reflects base-call-quality scores for different base positions in a first nucleotide read. More specifically, at the sequencing cycle for the first nucleotide read 424, the cluster-filtering system 106 determines a first base-call-quality score for a base call of the first nucleotide read based on a first set of signal values representing a signal of the cluster of oligonucleotides.

[0088] In some embodiments, the sequencing cycle for the first nucleotide read 424 comprises a last sequencing cycle in the first subset of sequencing cycles corresponding to the first nucleotide read. In particular, the cluster-filtering system 106 uses base-call-quality scores from all sequencing cycles for the first nucleotide read to generate the first base-call-quality score. For example, in a sequencing run comprising 302 total sequencing cycles, the sequencing cycle for the first nucleotide read 424 may comprise the last sequencing cycle for sequencing the first nucleotide read or cycle 151 for Rl. In some implementations, the cluster-filtering system 106 accesses and aggregates base-call-quality scores for all base calls of the first nucleotide read at the sequencing cycle for the first nucleotide read 424.

[0089] Additionally, or alternatively, in some implementations, the sequencing cycle for the first nucleotide read 424 occurs after a fraction of sequencing cycles for the first nucleotide read within a sequencing run. For example, the sequencing cycle for the first nucleotide read 424 can occur after a predetermined fraction of sequencing cycles (e.g., 1 / 4, 1 / 3, 1 / 2, etc.). In another example, the sequencing cycle for the first nucleotide read 424 occurs after a set number of sequencing cycles for the first nucleotide read. For instance, the sequencing cycle for the first nucleotide read 424 can occur after cycle 15, cycle 20, cycle 25, etc.

[0090] In addition to illustrating filtering acts for at least one sequencing cycle of a first nucleotide read, FIG. 4 further illustrates filtering acts that the cluster-filtering system 106 may perform for at least one sequencing cycle for the second nucleotide read 426. In some examples, the sequencing cycle for the second nucleotide read 426 comprises the last sequencing cycle within a sequencing run. For instance, in a sequencing run comprising 302 total sequencing cycles, the sequencing cycle for the second nucleotide read 426 may comprise the last sequencing cycle for sequencing the second nucleotide read, or cycle 151 for R2. In some implementations, upon completing all sequencing cycles, the cluster-filtering system 106 can output base call files.

[0091] Similar to a first nucleotide read, in some implementations, the sequencing cycle for the second nucleotide read 426 occurs after a fraction of sequencing cycles for the second nucleotide read within a sequencing run. For example, the sequencing cycle for the second nucleotide read 426 can occur after a predetermined fraction of sequencing cycles (e.g., 1 / 4, 1 / 3, 1 / 2, etc.). In another example, the sequencing cycle for the second nucleotide read 426 occurs after a set number of sequencing cycles for the second nucleotide read. For instance, the sequencing cycle for the second nucleotide read 426 can occur after cycle 15, cycle 20, cycle 25, etc.

[0092] After determining base-call-quality scores for at least one sequencing cycle of a first nucleotide read, as illustrated in FIG. 4, the cluster-filtering system 106 optionally performs an act 410 of outputting Base Call (BCL) File(s) at the sequencing cycle for the second nucleotide read 426. Generally, BCL files comprise raw base call data produced by a sequencing device. Such BCLfiles contain the raw sequences or base calls detected during a sequencing run along with their associated base-call-quality scores and / or read-quality scores. In some implementations, the BCL files contain unfiltered data for all clusters of oligonucleotides across all sequencing cycles within a sequencing run. As shown in FIG. 4, based on completing the sequencing cycle for the second nucleotide read 426, the cluster-filtering system 106 can output the BCL file(s).

[0093] Additionally, and as illustrated in FIG. 4, the cluster-filtering system 106 can perform an act 412 of determining second base-call-quality score(s) at the sequencing cycle for the second nucleotide read 426. Like how the cluster-filtering system 106 determines first base-call-quality score(s), the cluster-filtering system 106 determines second base-call-quality score(s) for base calls of the second nucleotide based on a second set of signal values representing a signal of the cluster of oligonucleotides. More particularly, the cluster-filtering system 106 may access and use data across one or more sequencing cycles for the second nucleotide read to generate the second base- call-quality score. As described below, FIG. 5 illustrates different methods by which the clusterfiltering system 106 determines first and second base-call-quality score(s) in accordance with one or more embodiments of the present disclosure.

[0094] As suggested above, the sequencing cycle for the first nucleotide read 424 and the sequencing cycle for the second nucleotide read 426 may comprise any sequencing cycle within the first subset of sequencing cycles and the second subset of sequencing cycles, respectively. As previously mentioned, the sequencing cycle for the first nucleotide read 424 may comprise the last sequencing cycle for the first nucleotide read, and the sequencing cycle for the second nucleotide read 426 may comprise the last sequencing cycle for the second nucleotide read. In some implementations, the cluster-filtering system 106 determines that the sequencing cycle for the first nucleotide read 424 occurs after a fraction of sequencing cycles for the first nucleotide read and that the sequencing cycle for the second nucleotide read 426 occurs after a fraction of sequencing cycles for the second nucleotide read within the sequencing run. For example, in some implementations, the sequencing cycle for the first nucleotide read 424 may occur anywhere between a first sequencing cycle for the first nucleotide read (e.g., cycle 1 for Rl) and the last sequencing cycle for the first nucleotide read (e.g., cycle 151 for Rl). Similarly, the sequencing cycle for the second nucleotide read 426 may occur anywhere between the first sequencing cycle for the second nucleotide read (e.g., cycle 1 for R2) and the last sequencing cycle for the second nucleotide read (e.g., cycle 151 for R2). In some examples, the sequencing cycle for the first nucleotide read 424 precedes the evaluation sequencing cycle.

[0095] The cluster-filtering system 106 can perform the act 408 and the act 412 at similar respective sequencing cycles for the first nucleotide read 424 and the second nucleotide read 426. For instance, if the sequencing cycle for the first nucleotide read 424 comprises cycle 1 for the firstnucleotide read, the cluster-filtering system 106 determines the first base-call-quality score(s) based on a base-call-quality score from the first sequencing cycle for the first nucleotide read. Similarly, if the sequencing cycle for the second nucleotide read 426 comprises cycle 1 for the second nucleotide read, the cluster-filtering system 106 determines the second base-call-quality score(s) based on the base-call-quality score from the first sequencing cycle for the second nucleotide read.

[0096] In addition or in the alternative to using a similar and read-specific sequencing cycle, the cluster-filtering system 106 can determine first base-call-quality score(s) and second base-call- quality score(s) based on a single base call of a first nucleotide read and a second nucleotide read, respectively. In some examples, the cluster-filtering system 106 accesses a first base-call-quality score for a base call at the sequencing cycle for the first nucleotide read 424. In other implementations, the cluster-filtering system 106 determines the first base-call-quality score(s) and the second base-call-quality score(s) based on an aggregate of base-call-quality score(s) up to the sequencing cycle for the first nucleotide read 424 and the sequencing cycle for the second nucleotide read 426, respectively. For instance, if the sequencing cycle for the first nucleotide read 424 comprises the cycle 151 for the first nucleotide read, the cluster-filtering system 106 can determine the first base-call-quality score by aggregating base-call-quality scores across base calls of the first nucleotide read up to cycle 151 for the first nucleotide read. The cluster-filtering system 106 may do the same for base calls for the second nucleotide read by aggregating base-call-quality scores across base calls of the second nucleotide read up to cycle 151 for the first nucleotide read.

[0097] As further shown in FIG. 4, the cluster-filtering system 106 may use the first base-call- quality score(s) and the second base-call-quality score(s) to perform an act 416 of determining a cluster-filtering score. In such an embodiment based on base-call-quality scores, the clusterfiltering score reflects the quality of base calls from different base positions in paired-end reads. FIG. 5 and the corresponding discussion further detail how the cluster-filtering system 106 determines the cluster-filtering score based on the first base-call-quality score(s) and the second base-call-quality score(s) in accordance with one or more implementations of the present disclosure.

[0098] After determining such a cluster-filtering score for an oligonucleotide cluster, as further shown in FIG. 4, the cluster-filtering system 106 performs an act 418 of outputting nucleotide reads and cluster-filtering scores. As part of the act 418, the cluster-filtering system 106 generates nucleotide sequence data for a nucleotide read and corresponding cluster-filtering scores. As mentioned previously, in some embodiments, the cluster-filtering scores offer a more accurate measure of quality of different portions of or an entire nucleotide read of a paired-end read when compared with chastity filter that reflect only a limited portion of a nucleotide read of the paired-end read. Accordingly, the cluster-fdtering system 106 may output nucleotide reads corresponding to all clusters of oligonucleotides together with their corresponding cluster-filtering scores.

[0099] As illustrated in FIG. 4, the cluster-filtering system 106 performs an act 420 of applying a cluster-filtering threshold score. More specifically, the cluster-filtering system 106 determines whether the cluster-filtering score for each cluster of oligonucleotides satisfies a cluster-filtering threshold score. In some embodiments, the cluster-filtering system 106 generates a separate filter file that contains information about which nucleotide reads of a sample satisfy the cluster-filtering threshold score. The cluster-filtering system 106 generates the filter file by applying the clusterfiltering threshold score to the determined cluster-filtering scores.

[0100] As further illustrated in FIG. 4, the cluster-filtering system 106 may perform an act 422 of outputting base-call data files. As shown, base-call data files may comprise BCL file(s) and / or FASTQ file(s). As described previously, BCL files comprise raw base call files generated by sequencing devices. FASTQ files can combine sequence data with quality scores in a text-based format and is populated with high-quality reads. In some implementations, the cluster-filtering system 106 uses the aforementioned filter file to create FASTQ files. More specifically, the clusterfiltering system 106 uses filter files to guide the conversion process from BCL files to FASTQ files. When converting BCL files to FASTQ files using tools, such as bcl2fastq software, the cluster-filtering system 106 references the filter file to exclude reads that fail to satisfy the clusterfiltering threshold score. This results in FASTQ files that include only reads that satisfy the clusterfiltering threshold score.

[0101] In some implementations, the cluster-filtering system 106 generates a plurality of filter files. For instance, the cluster-filtering system 106 may generate a first filter file indicating whether a cluster of oligonucleotides satisfies a chastity threshold value. As part of performing the act 406 of outputting an expected pass-filter indicator, in some cases, the cluster-filtering system 106 generates a first filter file corresponding to the chastity filter. The cluster-filtering system 106 may generate a second filter file indicating whether the cluster of oligonucleotides satisfies a clusterfiltering threshold score. More specifically, and as shown in FIG. 4, the cluster-filtering system 106 can generate a base-call-quality score-based cluster-filtering threshold score. The cluster-filtering system 106 may determine to apply either the first filter file or the second filter file as part of outputting the base-call data file(s).

[0102] In some examples, the cluster-filtering system 106 determines which filter file to use based on user input. For example, the cluster-filtering system 106 may receive a user indication to utilize either the chastity filter or the cluster-filtering threshold score to filter clusters of oligonucleotides. In another example, the cluster-filtering system 106 determines whether to use the first filter file or the second filter file based on metrics such as the percent of clusters passingfilter and / or the quality of the clusters passing filter. For instance, the cluster-filtering system 106 can determine to use the filter file corresponding with the highest percent of clusters passing filter and / or the highest quality of clusters passing filter.

[0103] As mentioned, the cluster-filtering system 106 may utilize different methods to determine cluster-filtering scores based on base-call-quality scores. In accordance with one or more implementations of the present disclosure, FIG. 5 illustrates the cluster-filtering system 106 using such different methods, including an error probability method that scales and averages base-call- quality scores for an oligonucleotide cluster to generate a cluster-filtering score and a threshold base-call-quality score method that counts a number of base calls within one or more nucleotide reads satisfying a threshold base-call-quality score to generate a cluster-filtering score. As explained below, the cluster-filtering system 106 can use one or both of an error probability method and a threshold base-call-quality score method to determine a cluster-filtering score for an oligonucleotide cluster based on base-call-quality scores.

[0104] As shown in FIG. 5, for instance, the cluster-filtering system 106 employs an error probability method 502 to determine a cluster-filtering score. In some implementations, the clusterfiltering system 106 converts base-call-quality scores for every base in a read of a paired-end read, including the first nucleotide read and the second nucleotide read, to a linear representation. The cluster-filtering system 106 averages the base-call-quality score(s) across the nucleotide read. Because base-call-quality scores tend to be well calibrated in a PhRED-based scale, the averaged base-call-quality score(s) across the nucleotide read correlates closely with clusters that pass a filtering threshold.

[0105] As illustrated by FIG. 5, the error probability method 502 comprises a series of acts including an act 506 of accessing base-call-quality scores for the first nucleotide read and the second nucleotide read, an act 508 of converting base-call-quality scores to a probability scale, and an act 510 of averaging the set of probability-scaled base-call-quality scores. As shown in FIG. 5, for instance, the cluster-filtering system 106 performs the act 506 of accessing base-call-quality scores for the first nucleotide read and the second nucleotide read. As mentioned previously, the cluster-filtering system 106 can access one or more base-call-quality scores for the first nucleotide read and the second nucleotide read. Depending on which sequencing cycle constitutes the sequencing cycle for the first nucleotide read and the sequencing cycle for the second nucleotide read, the cluster-filtering system 106 retrieves base-call-quality scores for anywhere between one to all base calls for the first nucleotide read and the second nucleotide read.

[0106] The base-call-quality score(s) can be in a Phil’s Read Editor (PhRED) scale and can be referred to as Q-scores. PhRED-scaled base-call-quality scores comprise a numerical value that quantifies the probability of an incorrect base call in sequencing data and can be calculated usingthe formula Q = — 10logwP, where Q is the PhRED score and P is the probability of an incorrect base call. In some embodiments, PhRED-scaled base-call-quality scores are generated from the PhRED algorithm. In other embodiments, PhRED-scaled base-call-quality are generated using methods other than the PhRED algorithm. For example, PhRED-scaled base-call-quality scores can comprise machine-learning- based quality scores that rely on the PhRED scale (e.g., Q10, Q20, etc.) generated by XGBoost or another machine-learning model.

[0107] As shown in FIG. 5, the cluster-filtering system 106 performs the act 508 of converting base-call-quality scores to a probability scale. In some cases, base-call-quality scores for the first nucleotide read and the second nucleotide read are in a log probability domain. For purposes of averaging base-call-quality scores across a read, the cluster-filtering system 106 converts the set of base-call-quality scores from the PhRED scale for quality scores to a probability scale for quality scores. In some examples, the cluster-filtering system 106 inverts the log for PhRED-scaled base- call-quality scores to return the base-call-quality scores to a probability space. The cluster-filtering system 106 thus generates probability-scaled base-call-quality scores for base call(s) within a given nucleotide read.

[0108] After converting base-call-quality scores to a probability scale, as further shown in FIG. 5, the cluster-filtering system 106 performs the act 510 of averaging the set of probability-scaled base-call-quality scores. In particular, the cluster-filtering system 106 averages the set of probability-scaled base-call-quality scores to generate the cluster-filtering score for the cluster of oligonucleotides. As indicated by FIG. 5, in some embodiments, the cluster-filtering system 106 averages probability -scaled base-call-quality scores across all sequencing cycles, including the first nucleotide read and the second nucleotide read. By contrast, in some embodiments, the clusterfiltering system 106 averages probability-scaled base-call-quality scores from a subset of sequencing cycles, such as a faction or percentage of sequencing cycles for the first nucleotide read and a faction or percentage of sequencing cycles for the second nucleotide read.

[0109] In addition or in the alternative to the error probability method 502, the cluster-filtering system 106 employs the threshold base-call-quality score method 504 to generate a cluster-filtering score for a given oligonucleotide cluster. In particular, the cluster-filtering system 106 determines a number of base calls within the first nucleotide read and the second nucleotide read with base- call-quality scores satisfying a threshold base-call-quality score (e.g., Q30, Q35, Q40, Q45). In such an embodiment, the cluster-filtering system 106 uses the number of base calls — from within the first nucleotide read and the second nucleotide read — with base-call-quality scores satisfying the threshold base-call-quality score as the cluster-filtering score for the cluster of oligonucleotides.

[0110] In some examples, and as illustrated in FIG. 5, the cluster-filtering system 106 determines that the threshold base-call-quality score is a PhRED-scaled base-call-quality score of30 (i.e., Q30). The cluster-filtering system 106 counts a number of base calls within the first nucleotide read and the second nucleotide read with base-call-quality scores equaling or exceeding Q30. The cluster-filtering system 106 uses the number of base calls having base-call-scores equaling or exceeding Q30 as the cluster-filtering score.[OHl] As mentioned, the cluster-filtering system 106 improves the percentage of clusters passing filter and accuracy of clusters passing filter relative to existing sequencing systems. In accordance with one or more implementations of the present disclosure, FIG. 6 illustrates improvements in percentage of clusters passing filter and quality of clusters passing filter when the cluster-filtering system 106 determines cluster-filtering scores based on base-call-quality scores. More particularly, FIG. 6 illustrates a chart 600 depicting how filtering by cluster-filtering score improves sequencing quality and percentage of oligonucleotide clusters that pass filter.

[0112] As shown in FIG. 6, the chart 600 depicts differences between chastity filter results 606, error probability results 604, and threshold base-call-quality score results 602. The chastity filter results 606 comprises results for a chastity filter implemented by an existing sequencing system. The error probability results 604 comprises results from using a base-call-quality score determined using the error probability method of the cluster-filtering system 106 described above. The threshold base-call-quality score results 602 reflect results from using a base-call-quality score determined using the threshold base-call-quality score method of the cluster-filtering system 106 described above. The x-axis of the chart 600 comprises %PF or the percentage of clusters that pass respective filters or threshold base-call-quality scores. The y-axis of the chart 600 comprises %Q30 or the percentage of base calls that meet or exceed a PhRED-scaled base-call-quality score of Q30.

[0113] FIG. 6 demonstrates tradeoffs between the percentage of clusters that pass filter and the percentage of bases with a Q30 quality score. As shown in FIG. 6, as the percentage of clusters that pass filter increases, the percentage of clusters that meet or exceed a PhRED-scaled base-call- quality score of Q30 decreases. As further shown in FIG. 6, both the error probability method and the threshold base-call-quality score method described above result in more favorable curves than the chastity filter results 606. More specifically, for the same %PF, the error probability results 604 and the threshold base-call-quality score results 602 demonstrate higher %Q30 than the chastity filter results 606. Similarly, for the same %Q30, the error probability results 604 and the threshold base-call-quality score results 602 demonstrate higher percentages of clusters that pass filter.

[0114] As mentioned previously, in some embodiments, the cluster-filtering system 106 can determine a cluster-filtering score based on signal-to-noise ratios. In accordance with one or more embodiments of the present disclosure, FIG. 7 illustrates an overview of the cluster-filtering system 106 determining a signal-to-noise ratio-based cluster-filtering score. FIG. 7 illustrates acts performed by the cluster-filtering system 106 at various stages and cycles within a sequencing run.For example, FIG. 7 illustrates acts performed by the cluster-filtering system 106 with reference to an evaluation sequencing cycle 702, a sequencing cycle for a first nucleotide read 704, a sequencing cycle for a second nucleotide read 706, and a final sequencing cycle 708 within a sequencing run.

[0115] As shown in FIG. 7, for instance, the cluster-filtering system 106 optionally performs an act 700 of checking for chastity failures from cycle 1 of the sequencing run to an evaluation sequencing cycle. Like performing the act 402 illustrated in FIG. 4, the cluster-filtering system 106 can perform the act 700 of checking for chastity failures. At the evaluation sequencing cycle 702, the cluster-filtering system 106 can perform the act 712 of evaluating the chastity filter. In some embodiments, the evaluation sequencing cycle comprises cycle 25 of a sequencing run. As mentioned previously, the chastity value or metric can represent a ratio of the brightest base signal to the sum of the brightest and second brightest base signals. If this ratio falls below a predefined chastity threshold value (e.g., 0.6), the read is flagged as a chastity failure.

[0116] As illustrated in FIG. 7, the cluster-filtering system 106 may perform an act 714 of outputting an expected pass-filter indicator. The act 714 illustrated in FIG. 7 is like the act 406 illustrated in FIG. 4 and can include all the steps and features of the act 406 in FIG. 4 described above. More specifically, the cluster-filtering system 106 can use the expected pass-filter indicator as predictive of quality of a sequencing run. To illustrate, a relatively lower expected pass-filter indicator (e.g., below 90%) may indicate issues with the sequencing run, such as poor sample quality or other technical problems. Furthermore, in some embodiments, the cluster-filtering system 106 utilizes the chastity filter to output the expected pass-filter indicator for purposes of backwards compatibility. The cluster-filtering system 106 implements a different and novel cluster filtering method and may continue to evaluate a chastity filter to maintain backward compatibility with sequencing systems that necessitate use of a chastity filter.

[0117] As mentioned previously, in some embodiments, the cluster-filtering system 106 retains data for clusters, even if they fail the chastity filter. For example, the cluster-filtering system 106 optionally retains all read data after the evaluation sequencing cycle 702 and can determine later to retain or disregard read data.

[0118] As further illustrated in FIG. 7, the cluster-filtering system 106 performs an act 716 of saving R1 signal-to-noise ratio(s) for a first nucleotide read at the sequencing cycle for the first nucleotide read 704. In particular, the cluster-filtering system 106 stores signal-to-noise ratio data corresponding to the first nucleotide read based on a first set of signal values representing signals of the cluster of oligonucleotides up to the sequencing cycle for the first nucleotide read 704. For instance, in some implementations, the cluster-filtering system 106 saves SNR(s) for the first nucleotide read for all sequencing cycles from the first sequencing cycle until the sequencing cycle for the first nucleotide read 704.

[0119] The cluster-filtering system 106 may determine the sequencing cycle for the first nucleotide read 704 from a first subset of sequencing cycles corresponding to the first nucleotide read. More specifically, the first subset of sequencing cycles corresponding to the first nucleotide read may comprise sequencing cycles used to sequence the first nucleotide read. In some examples, the cluster-filtering system 106 uses the first subset of sequencing cycles to sequence the genomic read. In other examples, the first subset of sequencing cycles is used to sequence an indexing sequence.

[0120] The sequencing cycle for the first nucleotide read 704 can occur after a fraction of sequencing cycles for the first nucleotide read. For instance, the sequencing cycle for the first nucleotide read 704 can occur after a fraction (e.g., 1 / 4, 1 / 3, 1 / 2, etc.) of sequencing cycles or after a number of sequencing cycles (e.g., cycle 15, cycle 25, cycle 50, cycle 75, etc.). In some implementations, the sequencing cycle for the first nucleotide read 704 precedes the evaluation sequencing cycle. For example, in some implementations, the sequencing cycle for the first nucleotide read 704 comprises cycle 10 for the first nucleotide read, whereas the evaluation sequencing cycle occurs on cycle 25. In some implementations, the first sequencing cycle for the first nucleotide read 704 is the same as the evaluation sequencing cycle 702. For instance, both the sequencing cycle for the first nucleotide read 704 and the evaluation sequencing cycle 702 may be cycle 25. In yet other embodiments, the evaluation sequencing cycle 702 precedes the sequencing cycle for the first nucleotide read 704. In an example where the evaluation sequencing cycle 702 is cycle 25 of a sequencing run, the sequencing cycle for the first nucleotide read 704 may comprise any one of cycle 30, cycle 75, cycle 90, and even cycle 151 for the first nucleotide read.

[0121] As further shown in FIG. 7, in addition to saving SNR(s) up to a particular cycle for the first nucleotide read, the cluster-filtering system 106 performs the act 718 of determining a first signal-to-noise ratio. In some cases, the first signal-to-noise ratio indicates a likelihood that signals for the first nucleotide read up to the sequencing cycle for the first nucleotide read 704 have been corrupted by noise. In some implementations, the cluster-filtering system 106 determines the first signal-to-noise ratio based on signal values representing signals of the cluster of oligonucleotides from the first sequencing cycle up to the sequencing cycle for the first nucleotide read 704. For example, the cluster-filtering system 106 may determine an average of signal-to-noise ratios across sequencing cycles from the first sequencing cycle until the sequencing cycle for the first nucleotide read 704. To illustrate, in cases where the sequencing cycle for the first nucleotide read 704 equals cycle 75, the cluster-filtering system 106 can average SNRs across all cycles between cycle 1 and cycle 75 for the first nucleotide read. In some implementations, the cluster-filtering system 106 performs the act 718 of determining the first signal-to-noise ratio on a sequencing device during the sequencing cycle for the first nucleotide read 704. In other examples, the cluster-filtering system106 performs the act 718 of determining the first signal-to-noise ratio later as part of downstream processing or secondary analysis, for example, on a server device, local device, or client device.

[0122] Having determined SNR(s) for the first nucleotide read, in some examples, the clusterfiltering system 106 stores the first signal-to-noise ratio as at least one float number for the first nucleotide read. A float number, short for floating-point number, comprises a type of data representation used to store real numbers that include a fractional part. A float number can represent very large or small numbers, as well as numbers with fractional parts. The cluster-filtering system 106 may store a float number indicating one or more SNR(s) for the first nucleotide read. Accordingly, the cluster-filtering system 106 may, at the sequencing cycle for the first nucleotide read 704, determine the first signal-to-noise ratio and store the first signal-to-noise ratio as a first float number.

[0123] As further illustrated in FIG. 7, the cluster-filtering system 106 performs an act 720 of saving SNR(s) for the second nucleotide read at a sequencing cycle for the second nucleotide read 706. For example, the cluster-filtering system 106 accesses signal values representing signals of the cluster of oligonucleotides up to the sequencing cycle for the second nucleotide read 706. In some implementations, the cluster-filtering system 106 accesses signal values for all sequencing cycles from cycle 1 to the sequencing cycle for the second nucleotide read 706 for a cluster.

[0124] Similar to using a first subset of sequencing cycles for the first nucleotide read, the cluster-filtering system 106 may determine the sequencing cycle for the second nucleotide read 706 from a second subset of sequencing cycles corresponding to the second nucleotide read. More specifically, the second subset of sequencing cycles corresponding to the second nucleotide read may comprise sequencing cycles used to sequence the second nucleotide read. In some examples, the cluster-filtering system 106 uses the second subset of sequencing cycles to sequence the genomic read for the second nucleotide read. In other examples, the first subset of sequencing cycles is used to sequence an indexing sequence attached to the second nucleotide read.

[0125] The sequencing cycle for the second nucleotide read 706 can occur after a fraction of sequencing cycles for the second nucleotide read. For instance, the sequencing cycle for the second nucleotide read 706 can occur after a fraction (e.g., 1 / 4, 1 / 3, 1 / 2, etc.) of sequencing cycles or after a number of sequencing cycles (e.g., cycle 15, cycle 25, cycle 75, cycle 150 etc.).

[0126] As further shown in FIG. 7, in addition to saving SNR(s) for the second nucleotide read, the cluster-filtering system 106 performs the act 722 of determining a second signal-to-noise ratio. In particular, the cluster-filtering system 106 can load the R2 SNR(s) saved as part of the act 720 and uses the saved R2 SNR(s) to determine a second signal-to-noise ratio. In some cases, the second signal-to-noise ratio indicates a likelihood that signals for the second nucleotide read up to the sequencing cycle for the second nucleotide read 706 have been corrupted by noise. For example,in some implementations, the cluster-filtering system 106 determines the second signal-to-noise ratio based on signal values representing signals of the cluster of oligonucleotides from the second sequencing cycle up to the sequencing cycle for the second nucleotide read 706. To illustrate, the cluster-filtering system 106 may determine an average of signal-to-noise ratios across sequencing cycles from the first sequencing cycle until the sequencing cycle for the second nucleotide read 706. In cases where the sequencing cycle for the second nucleotide read 706 equals cycle 75, for instance, the cluster-filtering system 106 can average SNRs across all cycles between cycle 1 and cycle 75 for the second nucleotide read. In some implementations, the cluster-filtering system 106 performs the act 722 of determining the second signal-to-noise ratio on a sequencing device during the sequencing cycle for the second nucleotide read 706. In other examples, the cluster-filtering system 106 performs the act 722 of determining the second signal-to-noise ratio later as part of downstream processing or secondary analysis, for example, on a server device, local device, or client device.

[0127] As further illustrated in FIG. 7, the cluster-filtering system 106 performs the act 724 of determining a cluster-filtering score for an oligonucleotide cluster. In particular, the clusterfiltering system 106 determines a cluster filtering score based on the first signal-to-noise ratio corresponding to the first nucleotide read and the second signal-to-noise ratio corresponding to the second nucleotide read. For example, in some implementations, the cluster-filtering system 106 determines the cluster-filtering score by converting the first signal-to-noise ratio to a first bit error rate and converting the second signal-to-noise ratio to a second bit error rate. In some implementations, the cluster-filtering system 106 utilizes the following complement-of-error function to convert the first signal-to-noise ratio and the second signal-to-noise ratio to the first bit error rate and the second bit error rate, respectively:BER = 1 - erf (Where BER represents the bit error rate for the first nucleotide read or the second nucleotide read, erf represents an error function, and SNR represents the first signal-to-noise ratio or the second signal-to-noise ratio. In some examples, erf is an error function that is related to a Gaussian (normal) distribution and describes the probability that a value in a normally distributed set of data will fall within a certain range around the mean. In some implementations, the cluster-filtering system 106 utilizes other methods to determine the cluster-filtering score based on the first signal- to-noise ratio and the second signal-to-noise ratio.

[0128] After determining cluster-filtering scores, as further illustrated in FIG. 7, the clusterfiltering system 106 performs the act 726 of outputting nucleotide reads and cluster-filtering scores. As part of the act 726, the cluster-filtering system 106 generates nucleotide sequence data for anucleotide read and corresponding cluster-filtering scores. As mentioned previously, in some embodiments, the cluster-filtering scores offer a more accurate measure of quality of different portions of or an entire nucleotide read when compared with chastity filter that reflects only a limited and initial portion of a nucleotide read. Accordingly, the cluster-filtering system 106 may output nucleotide reads corresponding to all clusters of oligonucleotides together with their corresponding cluster-filtering scores.

[0129] At the final sequencing cycle 708, in some cases, the cluster-filtering system 106 performs an act 728 of outputting base call (BCL) file(s). As mentioned, BCL files can comprise raw base call data produced by a sequencing device. The BCL files contain the raw sequences or base calls detected during a sequencing run along with their associated quality scores. In some implementations, the BCL files contain unfiltered data for all clusters of oligonucleotides across all sequencing cycles within a sequencing run. As shown in FIG. 7, the cluster-filtering system 106 may perform the act 728 of outputting unfiltered BCL file(s) comprising all unfiltered raw data for both the first nucleotide read and the second nucleotide read.

[0130] As further illustrated in FIG. 7, the cluster-filtering system 106 performs an act 732 of applying a cluster-filtering threshold score. In particular, the cluster-filtering system 106 determines whether the cluster-filtering score for each cluster of oligonucleotides satisfies a clusterfiltering threshold score. In some implementations, the cluster-filtering system 106 generates a filter file that contains information about which reads satisfy the cluster-filtering threshold score. The cluster-filtering system 106 generates the filter file by applying the cluster-filtering threshold score to the determined cluster-filtering scores.

[0131] As further illustrated in FIG. 7, the cluster-filtering system 106 performs an act 730 of outputting base-call data file(s). As mentioned previously, base-call data fde(s) can comprise file formats including BCL files and FASTQ files. The cluster-filtering system 106 performs the act 730 of outputting the base-call data fde(s) at or after the final sequencing cycle 708.

[0132] The cluster-filtering system 106 may perform the act 730 of outputting the base-call data file(s) after the cluster-filtering system 106 has filtered nucleotide clusters based on their cluster-filtering scores. To illustrate, whereas BCL files comprise raw base call files generated by sequencing devices, FASTQ files combine sequence data with quality scores in a text-based format and is populated with high-quality reads. The cluster-filtering system 106 may use filter files to guide the conversion process from BCL files to FASTQ files. For example, when converting BCL files to FASTQ files using tools, such as bcl2fastq software, the cluster-filtering system 106 references the filter file to exclude reads that fail to satisfy the cluster-filtering threshold score. This results in FASTQ files that include only reads that satisfy the cluster-filtering threshold score.

[0133] In some implementations, the cluster-filtering system 106 generates a plurality of filter files. For instance, the cluster-filtering system 106 may generate a first filter file indicating whether a cluster of oligonucleotides satisfies a chastity threshold value. For instance, as part of performing the act 714 of outputting an expected pass-filter indicator, the cluster-filtering system 106 generates a first filter file corresponding to the chastity filter. The cluster-filtering system 106 may generate a second filter file indicating whether the cluster of oligonucleotides satisfies a cluster-filtering threshold score. More specifically, and as shown in FIG. 7, the cluster-filtering system 106 can generate an SNR-based cluster-filtering threshold score. The cluster-filtering system 106 may determine to apply either the first filter file or the second filter file as part of outputting the basecall data file(s). In some examples, the cluster-filtering system 106 determines which filter file to use based on user input. For example, the cluster-filtering system 106 may receive a user indication to utilize either the chastity filter or the cluster-filtering threshold score to filter clusters of oligonucleotides. In another example, the cluster-filtering system 106 determines whether to use the first filter file or the second filter file based on metrics such as the percent of clusters passing filter and / or the quality of the clusters passing filter. For instance, the cluster-filtering system 106 can determine to use the filter file corresponding with the highest percent of clusters passing filter and / or the highest quality of clusters passing filter.

[0134] As mentioned, the cluster-filtering system 106 may select any sequencing cycle within the first subset of sequencing cycles as the sequencing cycle for the first nucleotide read and any sequencing cycle within the second subset of sequencing cycles as the sequencing cycle for the second nucleotide read. FIG. 8 illustrates tradeoffs for the cluster-filtering system 106 selecting different cycles for evaluating SNR for the first nucleotide read or the second nucleotide read in accordance with one or more embodiments of the present disclosure. In particular, FIG. 8 illustrates a line graph 800 showing how the percent of clusters passing filter varies based on selecting different cycles as the sequencing cycle for the first and second nucleotide reads. FIG. 8 further illustrates violin plots 802 and 804 showing asymmetries of distributions based on selection of various cycles as the sequencing cycle for the first and second nucleotide reads.

[0135] As shown in FIG. 8, the line graph 800 shows how the percent of clusters that pass filter changes as the cluster-filtering system 106 selects different sequencing cycles (cycles 25, 50, 75, 100, 125, and 150) as the sequencing cycle for the first nucleotide read and the second nucleotide read. As mentioned, the cluster-filtering system 106 can store and determine the first signal -to- noise ratio at the sequencing cycle for the first nucleotide read and determine the second signal-to- noise ratio at the sequencing cycle for the second nucleotide read. More specifically, the clusterfiltering system 106 aggregates signal data between cycle 1 to the relevant sequencing cycle for determining SNR(s) corresponding to the first nucleotide read or the second nucleotide read as thebasis for a cluster-filtering score. The x-axis of the line graph 800 indicates the selected sequencing cycle for the first nucleotide read or the second nucleotide read, and the y-axis indicates shows the values of the percentage of clusters that pass filter.

[0136] As shown in FIG. 8, in some embodiments, the highest percent of clusters passing filter occurs when the cluster-filtering system 106 selects cycles 50 or 75 as the sequencing cycle for determining SNR(s) corresponding to the first nucleotide read or the second nucleotide read as the basis for a cluster-filtering score. In some implementations, the highest percent of clusters passing filter occurs at other cycles beside cycle 50 and cycle 75. Accordingly, the cluster-filtering system 106 can dynamically select the sequencing cycle for the first nucleotide read and the second nucleotide read based on the sequencing cycle associated with the highest percentage of clusters passing filter or satisfying the cluster-filtering threshold score.

[0137] FIG. 8 further illustrates the violin plot 802 and the violin plot 804. The violin plot 802 visualizes a distribution SNR in decibels (db) at different sequencing cycles (e.g., 25, 50, 75, 100, 125, and 150) for a first nucleotide read. The violin plot 804 visualizes a distribution SNR in decibels (db) at different sequencing cycles (e.g., 25, 50, 75, 100, 125, and 150) for a second nucleotide read. As shown by both the violin plot 802 and the violin plot 804 in FIG. 8, the distributions of SNRs for cycles 50 and 75 are more skewed toward higher SNRs. In contrast, other cycles correspond with slightly lower average SNRs with more diffused distributions. Due to the asymmetry of distributions at cycles 50 and 75, the cluster-filtering system 106 may determine it is more simple to form a decision boundary or to determine the cluster-filtering threshold score at cycles 50 and 75. The cluster-filtering system 106 may further evaluate distributions of SNRs across cycles for both nucleotide reads to identify which cycle to designate as the sequencing cycle for the first nucleotide read and the sequencing cycle for the second nucleotide read for purposes of determining SNR(s) as a basis of a cluster-filtering score.

[0138] As mentioned, the cluster-filtering system 106 can achieve improved or competitive secondary analysis metrics relative to existing sequencing systems. More particularly, the clusterfiltering system 106 produces numbers of false positive (FP) variant calls and false negative (FN) variant calls that are at least comparable or an improvement to FP and FN numbers generated using a chastity filter. Generally, false positive variant calls occur when a variant calling model incorrectly identifies a genomic coordinate or region of a sample as having a genetic variant. False negative variant calls occur when a variant calling model fails to identify a genomic coordinate or region of a sample that actually has a genetic variant and incorrectly labels the coordinate or region as a non-variant or as matching the reference genome.

[0139] As mentioned, in some implementations, the cluster-filtering system 106 can improve the percentage of clusters that pass filter while also maintaining relatively high read quality incomparison to a chastity filter. FIG. 9 illustrates the cluster-filtering system 106 improving the percentage of clusters passing filter while maintaining high-quality reads relative to a chastity filter in accordance with one or more embodiments of the present disclosure. FIG. 9 illustrates a line plot 900 with a magnified portion 902 showing the relationship between a percentage of bases with a base-call-quality score of 30 or higher (i.e., %Q30) and the percentage of clusters that pass filter (i.e., %PF) using various filtering methods.

[0140] More specifically, FIG. 9 illustrates the following filtering methods: chastity filter (shown as “Chastity” in FIG. 9), an error probability method for cluster-filtering scores based on base-call-quality scores (shown as “Error Probability” in FIG. 9), a threshold base-call-quality score method for cluster-filtering scores based on base-call-quality scores (shown as “Num Q30s” in FIG. 9), and SNR-based cluster-filtering scores as determined at cycle 25 (shown as “SNR Cycle 25” in FIG. 9), and cycle 75 (shown as “SNR Cycle 75” in FIG. 9).

[0141] As shown in FIG. 9, the cluster-filtering system 106 outperforms existing sequencing systems when using virtually any method. In particular, plot line 912 shows the performance of the chastity filter. As shown in FIG. 9, plot lines corresponding to any method employed by the clusterfiltering system 106 yields both a higher percentage of clusters passing filter and a higher percentage of bases with a base-call-quality score of 30 or higher. Plot line 920 represents the performance of SNR-based cluster filtering scores as determined at cycle 25 (shown as “SNR Cycle 75” in FIG. 9); plot line 930 represents the performance of the error probability method (shown as “Error Probability” in FIG. 9); plot line 940 represents the performance of SNR-based cluster filtering scores as determined at cycle 75 (shown as “SNR Cycle 75” in FIG. 9); and plot line 950 represents the performance of the threshold base-call-quality score method (shown as “Num Q30s” in FIG. 9).

[0142] In some embodiments, the cluster-filtering system 106 improves the percentage of clusters that pass filter and the quality of base calls (both R1 and R2) by using the error probability method by using a base-call-quality score-based cluster-filtering score. More specifically, by using the error probability method described above, the cluster-filtering system 106 improves the percentage of clusters that pass filter by approximately 1.8%. In other words, in a region of a nucleotide-sample slide (e.g., a tile) containing a total of 5 million clusters, the cluster-filtering system 106 can increase the number of clusters that pass filter by 90,000 clusters. Furthermore, by employing the error probability method described above, the cluster-filtering system 106 can improve the percentage of base calls that meet or exceed a base-call-quality score of 30 (i.e., Q30) by approximately 0.31%. The error probability method is further associated with a decrease in numbers of false positive and false negative single nucleotide polymorphisms (SNPs) by approximately 1%.

[0143] In some embodiments, the cluster-filtering system 106 improves the percentage of clusters that pass filter and the quality of base calls (both R1 and R2) by using an SNR-based cluster-filtering score when determining signal-to-noise ratios at cycle 75 for the first nucleotide read and the second nucleotide read. To illustrate, by using an SNR-based cluster-filtering score, the cluster-filtering system 106 improves the percentage of clusters that pass filter by approximately 2.1%. For a region of a nucleotide-sample slide (e.g., a tile) containing a total of 5 million clusters, a 2.1% increase translates to an increase of approximately 105,000 clusters that pass filter. Additionally, the cluster-filtering system 106 can improve the percentage of base calls that meet or exceed a base-call-quality score of 30 (i.e., Q30) by approximately 0.01%. Use of the SNR-based cluster-filtering score is further associated with decreases in false positive and false negative SNPs by approximately 0.7% and INDELS by approximately 0.7%.

[0144] As mentioned, the cluster-filtering system 106 can improve the percentage of clusters that pass filter while also decreasing the number of false positive (FP) and false negative (FN) variant calls. In accordance with one or more embodiments, FIGS. 10A-10B depict charts quantifying FP and FN variant calls of the cluster-filtering system 106 in various genomic regions relative to existing sequencing systems. For example, the genomic regions for the charts depicted in FIGS. 10A-10B include dinucleotide regions, 5’-C-phosphate-G-3’ (CpG) Islands, homopolymers, and easy-to-map regions from precision FDA (pFDA) challenge data. As indicated by the charts in FIGS. 10A-10B, the cluster-filtering system 106 decreases the number of FP and FN variant calls in various genomic regions relative to existing sequencing systems.

[0145] In the context of FIGS. 10A-10B, note that data for some such genomic regions are often output in Browser Extensible Data (BED) format following bioinformatics standards and include data annotations for variant calls in target genomic regions. Accordingly, some target genomic regions are referred to as BED regions according to industry standard. Following such a standard, a BED region is sometimes referred to as a genomic region with data or annotations described or stored in a BED file format. The BED file itself often comprises a plain text file that represents genomic data and specifies intervals or target regions of interest within a genome. Accordingly, and for instance, BED regions may comprise genes, variants, regulatory regions, or any other feature of interest. In some examples, and as illustrated in FIGS. 10A-10B, BED regions comprise SNPs. While the following paragraphs refer to BED regions per bioinformatics standards, any alternative data format (e.g., VCF, BGEN) could likewise be used. Reference to BED regions below and in this disclosure is merely for ease of reference and such genomic regions are not limited to being represented or stored in a particular file format.

[0146] Accordingly, FIGS. 10A-10B illustrate a series of bar charts demonstrating how the cluster-filtering system 106 decreases FP and FN variant calls for SNPs in various genomic regionsstored in BED format or, simply, BED regions. In particular, the bar charts illustrated in FIGS. 10A-10B show counts of FP and FN SNPs for existing sequencing systems using the following filtering methods: chastity filter (shown as “Chastity” in FIGS. 10A-10B) and SNR-based clusterfiltering scores as determined at cycle 25 (shown as “SNR at cycle 25” in FIGS. 10A-10B) and cycle 75 (shown as “SNR at cycle 75” in FIGS. 10A-10B).

[0147] As shown in FIG. 10A, the cluster-filtering system 106 outperforms existing sequencing systems when using SNR-based cluster-filtering scores. In particular, bar chart 1010 shows that the chastity filter yields 25 FP and FN variant calls for SNPs in AT Dinucleotide regions. In contrast, the cluster-filtering system 106, when determining SNR-based cluster-filtering scores at cycle 25 and cycle 75, resulted in only 19 and 20 FP and FN variant calls in AT Dinucleotide regions, respectively.

[0148] As further shown by bar chart 1020 in FIG. 10A, the cluster-filtering system 106 outperforms existing sequencing systems and results in fewer FP and FN in CpG Island BED regions. More specifically, the chastity filter yielded significantly more FP and FN variant calls for SNPs in CpG Islands when compared with the cluster-filtering system 106 determining SNR-based cluster-filtering scores at cycle 25 and 75.

[0149] FIG. 10B illustrates a bar chart 1030 portraying decreased FP and FN variant calls for SNPs in 10-25 nucleotide-length homopolymer regions when the cluster-filtering system 106 uses SNR-based cluster-filtering scores relative to chastity filters used by existing sequencing systems. A bar chart 1040 illustrated in FIG. 10B similarly shows a decrease in FP and FN variant calls in easy-to-map regions from Precision FDA (pFDA) Challenge data when the cluster-filtering system 106 uses SNR-based cluster-filtering scores relative to chastity filters used by existing sequencing systems. In addition to the genomic regions depicted in FIGS. 10A and 10B, the cluster-filtering system 106 has further demonstrated decreases in the number of FP and FN variant calls for SNPs in the following BED regions relative to chastity filters: G quads, inverted repeats, mirrored repeats, pFDA Challenge SNPs in a combination of difficult- and easy-to-map regions, pFDA Challenge SNPs in difficult-to-map regions, short tandem repeats, and Z-DNA.

[0150] Turning now to FIG. 11, this figure illustrates an example flowchart of a series of acts for generating a base-call data file based on a cluster-filtering score satisfying a cluster-filtering threshold score in accordance with one or more embodiments of the present disclosure. While FIG. 11 illustrates acts according to particular embodiments, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 11. The acts of FIG. 11 can be performed as part of a method. Alternatively, a non-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device to perform the acts depicted in FIG. 11. In still further embodiments, a system comprising at least oneprocessor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 11.

[0151] As shown in FIG. 11, the series of acts 1100 includes an act 1102 of accessing a set of signal values, an act 1104 of determining a cluster-filtering score based on a subset of signal values from (i) a sequencing cycle for a first nucleotide read and (ii) a sequencing cycle for a second nucleotide read, an act 1106 of determining the cluster-filtering score satisfies a cluster-filtering threshold score, and an act 1108 of generating a base-call data file based on the cluster-filtering score satisfying the cluster-filtering threshold score.

[0152] For example, the series of acts 1100 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A computer-implemented method comprising: accessing, for a cluster of oligonucleotides, a set of signal values from a set of sequencing cycles for a sequencing run; determining a cluster-filtering score for the cluster of oligonucleotides based on a subset of signal values of the set of signal values from (i) a sequencing cycle for a first nucleotide read of the cluster of oligonucleotides and (ii) a sequencing cycle for a second nucleotide read of the cluster of oligonucleotides; determining that the cluster-filtering score for the cluster of oligonucleotides satisfies a cluster-filtering threshold score; and generating, based on the cluster-filtering score satisfying the cluster-filtering threshold score, a base-call data file comprising base calls for the first or second nucleotide read of the cluster of oligonucleotides.CLAUSE 2. The computer-implemented method of clause 1, further comprising determining the cluster-filtering score by determining a single cluster-filtering score for both the first nucleotide read and the second nucleotide read of the cluster of oligonucleotides.CLAUSE 3. The computer-implemented method of clause 1 or 2, further comprising: determining a first estimated chastity value for the first nucleotide read or a second estimated chastity value for the second nucleotide read based on a subset of signal values from an evaluation sequencing cycle preceding (i) the sequencing cycle for the first nucleotide read or (ii) the sequencing cycle for the second nucleotide read; and determining the cluster-filtering score for the cluster of oligonucleotides after determining the first estimated chastity value or the second estimated chastity value.CLAUSE 4. The computer-implemented method of any one of clauses 1-3, wherein: the sequencing cycle for the first nucleotide read is part of a first subset of sequencing cycles corresponding to the first nucleotide read; andthe sequencing cycle for the second nucleotide read is part of a second subset of sequencing cycles corresponding to the second nucleotide read.CLAUSE 5. The computer-implemented method of any one of clauses 1-4, wherein: the sequencing cycle for the first nucleotide read occurs after a first fraction of sequencing cycles for the first nucleotide read within the sequencing run; and the sequencing cycle for the second nucleotide read occur after a first fraction of sequencing cycles for the second nucleotide read within the sequencing run.CLAUSE 6. The computer-implemented method of any one of clauses 1-5, further comprising determining the cluster-filtering score by: determining, for the cluster of oligonucleotides, a first signal-to-noise ratio corresponding to the first nucleotide read based on a first set of signal values representing signals of the cluster of oligonucleotides up to (i) the sequencing cycle for the first nucleotide read; determining, for the cluster of oligonucleotides, a second signal-to-noise ratio corresponding to the second nucleotide read based on a second set of signal values representing signals of the cluster of oligonucleotides up to (ii) the sequencing cycle for the second nucleotide read; and determining the cluster-filtering score based on the first signal-to-noise ratio and the second signal-to-noise ratio.CLAUSE 7. The computer-implemented method of clause 6, further comprising determining the cluster-filtering score by: converting the first signal-to-noise ratio to a first bit error rate; converting the second signal-to-noise ratio to a second bit error rate; and determining an average between the first bit error rate and the second bit error rate.CLAUSE 8. The computer-implemented method of clause 7, further comprising utilizing a complement-of-error function to convert the first signal-to-noise ratio to the first bit error rate and the second signal-to-noise ratio to the second bit error rate.CLAUSE 9. The computer-implemented method of any one of clauses 1-8, further comprising determining the cluster-filtering score by: determining a first base-call-quality score for a base call of the first nucleotide read based on a first set of signal values representing a signal of the cluster of oligonucleotides at (i) the sequencing cycle for the first nucleotide read; determining a second base-call-quality score for a base call of the second nucleotide read based on a second set of signal values representing a signal of the cluster of oligonucleotides at (ii) the sequencing cycle for the second nucleotide read; anddetermining the cluster-filtering score based on the first base-call-quality score and the second base-call-quality score.CLAUSE 10. The computer-implemented method of clause 9, further comprising determining the first base-call-quality score or the second base-call-quality score further based on one or more of chastity values, signal -to-noise ratio, or other quality predictor values corresponding to the base call of the first nucleotide read or the second nucleotide read.CLAUSE 11. The computer-implemented method of clause 9, further comprising determining the cluster-filtering score by: accessing a set of base-call-quality scores for the first nucleotide read and the second nucleotide read from across the set of sequencing cycles, wherein the set of base-call-quality scores comprises the first base-call-quality score for the base call of the first nucleotide read and the second base-call-quality score for the base call of the second nucleotide read; converting the set of base-call-quality scores from a Phil’s Read EDitor (PhRED) scale for quality scores to a probability scale for quality scores to generate a set of probability-scaled base- call-quality scores; and averaging the set of probability-scaled base-call-quality scores to generate the clusterfiltering score for the cluster of oligonucleotides.CLAUSE 12. The computer-implemented method of clause 9, further comprising determining the cluster-filtering score by: determining a number of base calls within the first nucleotide read and the second nucleotide read with base-call-quality scores satisfying a threshold base-call-quality score; and generating, as the cluster-filtering score for the cluster of oligonucleotides, the number of base calls within the first nucleotide read and the second nucleotide read with base-call-quality scores satisfying the threshold base-call-quality score.CLAUSE 13. The computer-implemented method of any one of clauses 1-12, further comprising: determining, from an evaluation sequencing cycle for the first nucleotide read preceding the sequencing cycle for the first nucleotide read, an estimated chastity value for the first nucleotide read; determining that the estimated chastity value for the first nucleotide read satisfies a chastity threshold value; and generating, based on the estimated chastity value satisfying the chastity threshold value, an expected pass-filter indicator for the first nucleotide read or the second nucleotide read.CLAUSE 14. The computer-implemented method of any one of clauses 1-13, further comprising:accessing, for a set of clusters of oligonucleotides, additional sets of signal values from the set of sequencing cycles for the sequencing run; determining cluster-filtering scores for the set of clusters of oligonucleotides based on subsets of signal values of the set of clusters of oligonucleotides from (i) a respective sequencing cycle for respective first nucleotide reads of the set of clusters of oligonucleotides and (ii) a respective sequencing cycle for respective second nucleotide reads of the set of clusters of oligonucleotides; determining whether the cluster-filtering scores satisfy the cluster-filtering threshold score; and generating the base-call data file further comprising base calls for one or more nucleotide reads corresponding to one or more clusters of oligonucleotides from the set of clusters of oligonucleotides having cluster-filtering scores that satisfy the cluster-filtering threshold score.CLAUSE 15. The computer-implemented method of any one of clauses 1-14, further comprising determining the cluster-filtering threshold score by titrating the cluster-filtering threshold score over cluster-filtering scores across a set of clusters of oligonucleotides relative to a number of clusters from the set of clusters of oligonucleotides satisfying the cluster-filtering threshold score.CLAUSE 16. The computer-implemented method of any one of clauses 1-15, further comprising generating the base-call data file comprising the base calls for the first or second nucleotide read of the cluster of oligonucleotides by: generating an initial base-call data file comprising, for the cluster of oligonucleotides, the set of signal values from the set of sequencing cycles; generating a first filter file indicating whether the cluster of oligonucleotides satisfies a chastity threshold value; generating a second filter file indicating that the cluster of oligonucleotides satisfies the cluster-filtering threshold score; and generating a filtered base-call data file by utilizing the first filter file or the second filter file to identify, from the initial base-call data file, base calls for the first or second nucleotide read of the cluster of oligonucleotides that satisfies the chastity threshold value or the cluster-filtering threshold score.CLAUSE 17. The computer-implemented method of any one of clauses 1-16, further comprising determining (i) the sequencing cycle for the first nucleotide read and (ii) the sequencing cycle for the second nucleotide read by: accessing, for a selected set of clusters of oligonucleotides, a set of selected signal values from candidate sequencing cycles of the set of sequencing cycles;determining selected cluster-fdtering scores for the selected set of clusters of oligonucleotides based on a subset of selected signal values for the selected set of clusters of oligonucleotides; and identifying, from the candidate sequencing cycles, a target sequencing cycle for the first nucleotide read and a target sequencing cycle for the second nucleotide read with a highest selected cluster-filtering score of the selected cluster-filtering scores relative to a number of clusters from the selected set of clusters of oligonucleotides satisfying the cluster-filtering threshold score.

[0153] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.

[0154] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.

[0155] SBS can utilize nucleotide monomers that have a terminator moiety or those that lack any terminator moieties. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using y-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is generally variable and dependent upon the template sequence and the mode of nucleotide delivery. For SBS techniques that utilize nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator can be reversible as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).

[0156] SBS techniques can utilize nucleotide monomers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as release ofpyrophosphate; or the like. In embodiments where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods developed by Solexa (now Illumina, Inc.).

[0157] Preferred embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on realtime pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C or G). Images obtained after addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed and analyzed using the methods set forth herein. For example, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different detection channels for reversible terminatorbased sequencing methods.

[0158] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, in WO 04 / 018497 and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference. This approach is being commercialized by Solexa (now Illumina Inc.), and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators in which both the termination can be reversed, and the fluorescent label cleavedfacilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from these modified nucleotides.

[0159] Preferably in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured following incorporation of labels into arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show nucleic acid features that have incorporated nucleotides of a particular type. Different features are present or absent in the different images due the different sequence content of each feature. However, the relative position of the features will remain unchanged in the images. Images obtained from such reversible terminator- SBS methods can be stored, processed and analyzed as set forth herein. Following the image capture step, labels can be removed, and reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.

[0160] In particular embodiments some or all of the nucleotide monomers can include reversible terminators. In such embodiments, reversible terminators / cleavable fluors can include fluor linked to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3' allyl group to block extension but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30 second exposure to long wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of natural termination that ensues after placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modifiednucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.

[0161] Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.

[0162] Some embodiments can utilize detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair. As a second example, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples, is a fluorescent-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).

[0163] Further, as described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so- called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.

[0164] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained following treatment of an array of nucleic acid features with the labeled sequencing reagents. Each image will show nucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based sequencing methods can be stored, processed and analyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No. 6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.

[0165] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147- 151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed and analyzed as set forth herein. Inparticular, the data can be treated as an image in accordance with the exemplary treatment of optical images and other images that is set forth herein.

[0166] Some embodiments can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). The illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed and analyzed as set forth herein.

[0167] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.

[0168] The above SBS methods can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows convenient delivery of sequencing reagents, removal of unreacted reagents and detection of incorporation events in a multiplex manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In anarray format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as, bridge amplification or emulsion PCR as described in further detail below.

[0169] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.

[0170] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprising components such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeqTM platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference. The sequencing system described above sequences nucleic acid polymers present in samples received by a sequencing device, as described further above.

[0171] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples caninclude nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.

[0172] The components of the cluster-filtering system 106 can include software, hardware, or both. For example, the components of the cluster-filtering system 106 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the client device 114, the local device 108, or the server device(s) 110). When executed by the one or more processors, the computer-executable instructions of the cluster-filtering system 106 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the cluster-filtering system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the cluster-filtering system 106 can include a combination of computer-executable instructions and hardware.

[0173] Furthermore, the components of the cluster-filtering system 106 performing the functions described herein with respect to the cluster-filtering system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as acloud-computing model. Thus, components of the cluster-filtering system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the cluster-filtering system 106 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.

[0174] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non- transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0175] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0176] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0177] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly viewsthe connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.

[0178] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0179] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0180] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0181] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0182] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.

[0183] FIG. 12 illustrates a block diagram of a computing device 1200 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1200 may implement the cluster-filtering system 106 and the sequencing device system 104. As shown by FIG. 12, the computing device 1200 can comprise a processor 1202, a memory 1204, a storage device 1206, an I / O interface 1208, and a communication interface 1210, which may be communicatively coupled by way of a communication infrastructure 1212. In certain embodiments, the computing device 1200 can include fewer or more components than those shown in FIG. 12. The following paragraphs describe components of the computing device 1200 shown in FIG. 12 in additional detail.

[0184] In one or more embodiments, the processor 1202 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1202 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1204, or the storage device 1206 and decode and execute them. The memory 1204 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1206 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0185] The I / O interface 1208 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1200. The I / O interface 1208 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, networkinterface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1208 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1208 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.

[0186] The communication interface 1210 can include hardware, software, or both. In any event, the communication interface 1210 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1200 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1210 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.

[0187] Additionally, the communication interface 1210 may facilitate communications with various types of wired or wireless networks. The communication interface 1210 may also facilitate communications using various communication protocols. The communication infrastructure 1212 may also include hardware, software, or both that couples components of the computing device 1200 to each other. For example, the communication interface 1210 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.

[0188] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.

[0189] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders.

[0190] Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

CLAIMS1. A system comprising: at least one processor; and a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to: access, for a cluster of oligonucleotides, a set of signal values from a set of sequencing cycles for a sequencing run; determine a cluster-filtering score for the cluster of oligonucleotides based on a subset of signal values of the set of signal values from (i) a sequencing cycle for a first nucleotide read of the cluster of oligonucleotides and (ii) a sequencing cycle for a second nucleotide read of the cluster of oligonucleotides; determine that the cluster-filtering score for the cluster of oligonucleotides satisfies a cluster-filtering threshold score; and generate, based on the cluster-filtering score satisfying the cluster-filtering threshold score, a base-call data file comprising base calls for the first or second nucleotide read of the cluster of oligonucleotides.

2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by determining a single cluster-filtering score for both the first nucleotide read and the second nucleotide read of the cluster of oligonucleotides.

3. The system of claim 1 or 2, further comprising instructions that, when executed by the at least one processor, cause the system to: determine a first estimated chastity value for the first nucleotide read or a second estimated chastity value for the second nucleotide read based on a subset of signal values from an evaluation sequencing cycle preceding (i) the sequencing cycle for the first nucleotide read or (ii) the sequencing cycle for the second nucleotide read; and determine the cluster-filtering score for the cluster of oligonucleotides after determining the first estimated chastity value or the second estimated chastity value.

4. The system of any one of claims 1-3, wherein: the sequencing cycle for the first nucleotide read is part of a first subset of sequencing cycles corresponding to the first nucleotide read; and the sequencing cycle for the second nucleotide read is part of a second subset of sequencing cycles corresponding to the second nucleotide read.

5. The system of any one of claims 1-4, wherein:the sequencing cycle for the first nucleotide read occurs after a fraction of sequencing cycles for the first nucleotide read within the sequencing run; and the sequencing cycle for the second nucleotide read occur after a fraction of sequencing cycles for the second nucleotide read within the sequencing run.

6. The system of any one of claims 1-5, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by: determining, for the cluster of oligonucleotides, a first signal-to-noise ratio corresponding to the first nucleotide read based on a first set of signal values representing signals of the cluster of oligonucleotides up to (i) the sequencing cycle for the first nucleotide read; determining, for the cluster of oligonucleotides, a second signal-to-noise ratio corresponding to the second nucleotide read based on a second set of signal values representing signals of the cluster of oligonucleotides up to (ii) the sequencing cycle for the second nucleotide read; and determining the cluster-filtering score based on the first signal-to-noise ratio and the second signal-to-noise ratio.

7. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by: converting the first signal-to-noise ratio to a first bit error rate; converting the second signal-to-noise ratio to a second bit error rate; and determining an average between the first bit error rate and the second bit error rate.

8. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to utilize a complement-of-error function to convert the first signal-to-noise ratio to the first bit error rate and the second signal-to-noise ratio to the second bit error rate.

9. The system of any one of claims 1-8, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by: determining a first base-call-quality score for a base call of the first nucleotide read based on a first set of signal values representing a signal of the cluster of oligonucleotides at (i) the sequencing cycle for the first nucleotide read; determining a second base-call-quality score for a base call of the second nucleotide read based on a second set of signal values representing a signal of the cluster of oligonucleotides at (ii) the sequencing cycle for the second nucleotide read; and determining the cluster-filtering score based on the first base-call-quality score and the second base-call-quality score.

10. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to determine the first base-call-quality score or the second base-call-quality score further based on one or more of chastity values, signal-to-noise ratio, or other quality predictor values corresponding to the base call of the first nucleotide read or the second nucleotide read.

11. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by: accessing a set of base-call-quality scores for the first nucleotide read and the second nucleotide read from across the set of sequencing cycles, wherein the set of base-call-quality scores comprises the first base-call-quality score for the base call of the first nucleotide read and the second base-call-quality score for the base call of the second nucleotide read; converting the set of base-call-quality scores from a Phil’s Read EDitor (PhRED) scale for quality scores to a probability scale for quality scores to generate a set of probability-scaled base- call-quality scores; and averaging the set of probability-scaled base-call-quality scores to generate the clusterfiltering score for the cluster of oligonucleotides.

12. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering score by: determining a number of base calls within the first nucleotide read and the second nucleotide read with base-call-quality scores satisfying a threshold base-call-quality score; and generating, as the cluster-filtering score for the cluster of oligonucleotides, the number of base calls within the first nucleotide read and the second nucleotide read with base-call-quality scores satisfying the threshold base-call-quality score.

13. The system of any one of claims 1-12, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, from an evaluation sequencing cycle for the first nucleotide read preceding the sequencing cycle for the first nucleotide read, an estimated chastity value for the first nucleotide read; determine that the estimated chastity value for the first nucleotide read satisfies a chastity threshold value; and generate, based on the estimated chastity value satisfying the chastity threshold value, an expected pass-filter indicator for the first nucleotide read or the second nucleotide read.

14. The system of any one of claims 1-13, further comprising instructions that, when executed by the at least one processor, cause the system to:access, for a set of clusters of oligonucleotides, additional sets of signal values from the set of sequencing cycles for the sequencing run; determine cluster-filtering scores for the set of clusters of oligonucleotides based on subsets of signal values of the set of clusters of oligonucleotides from (i) a respective sequencing cycle for respective first nucleotide reads of the set of clusters of oligonucleotides and (ii) a respective sequencing cycle for respective second nucleotide reads of the set of clusters of oligonucleotides; determine whether the cluster-filtering scores satisfy the cluster-filtering threshold score; and generate the base-call data file further comprising base calls for one or more nucleotide reads corresponding to one or more clusters of oligonucleotides from the set of clusters of oligonucleotides having cluster-filtering scores that satisfy the cluster-filtering threshold score.

15. The system of any one of claims 1-14, further comprising instructions that, when executed by the at least one processor, cause the system to determine the cluster-filtering threshold score by titrating the cluster-filtering threshold score over cluster-filtering scores across a set of clusters of oligonucleotides relative to a number of clusters from the set of clusters of oligonucleotides satisfying the cluster-filtering threshold score.

16. The system of any one of claims 1-15, further comprising instructions that, when executed by the at least one processor, cause the system to generate the base-call data file comprising the base calls for the first or second nucleotide read of the cluster of oligonucleotides by: generating an initial base-call data file comprising, for the cluster of oligonucleotides, the set of signal values from the set of sequencing cycles; generating a first filter file indicating whether the cluster of oligonucleotides satisfies a chastity threshold value; generating a second filter file indicating that the cluster of oligonucleotides satisfies the cluster-filtering threshold score; and generating a filtered base-call data file by utilizing the first filter file or the second filter file to identify, from the initial base-call data file, base calls for the first or second nucleotide read of the cluster of oligonucleotides that satisfies the chastity threshold value or the cluster-filtering threshold score.

17. The system of any one of claims 1-16, further comprising instructions that, when executed by the at least one processor, cause the system to determine (i) the sequencing cycle for the first nucleotide read and (ii) the sequencing cycle for the second nucleotide read by: accessing, for a selected set of clusters of oligonucleotides, a set of selected signal values from candidate sequencing cycles of the set of sequencing cycles;determining selected cluster-fdtering scores for the selected set of clusters of oligonucleotides based on a subset of selected signal values for the selected set of clusters of oligonucleotides; and identifying, from the candidate sequencing cycles, a target sequencing cycle for the first nucleotide read and a target sequencing cycle for the second nucleotide read with a highest selected cluster-filtering score of the selected cluster-filtering scores relative to a number of clusters from the selected set of clusters of oligonucleotides satisfying the cluster-filtering threshold score.

18. A computer-implemented method comprising: accessing, for a cluster of oligonucleotides, a set of signal values from a set of sequencing cycles for a sequencing run; determining a cluster-filtering score for the cluster of oligonucleotides based on a subset of signal values of the set of signal values from (i) a sequencing cycle for a first nucleotide read of the cluster of oligonucleotides and (ii) a sequencing cycle for a second nucleotide read of the cluster of oligonucleotides; determining that the cluster-filtering score for the cluster of oligonucleotides satisfies a cluster-filtering threshold score; and generating, based on the cluster-filtering score satisfying the cluster-filtering threshold score, a base-call data file comprising base calls for the first or second nucleotide read of the cluster of oligonucleotides.

19. The computer-implemented method of claim 18, further comprising determining the cluster-filtering score by determining a single cluster-filtering score for both the first nucleotide read and the second nucleotide read of the cluster of oligonucleotides.

20. The computer-implemented method of claim 18 or 19, further comprising: determining a first estimated chastity value for the first nucleotide read or a second estimated chastity value for the second nucleotide read based on a subset of signal values from an evaluation sequencing cycle preceding (i) the sequencing cycle for the first nucleotide read or (ii) the sequencing cycle for the second nucleotide read; and determining the cluster-filtering score for the cluster of oligonucleotides after determining the first estimated chastity value or the second estimated chastity value.

Citation Information

Patent Citations

  • Method of nucleic acid amplification

    US20050100900A1

  • Labelled nucleotides

    US20060188901A1

  • Modified polymerases for improved incorporation of nucleotide analogues

    US20060240439A1

  • Polymerases

    US20060281109A1

  • Modified nucleotides

    US20070166705A1