Determining offline corrections for sequence specific errors caused by low complexity sequences using a machine learning model
A machine learning-based offline correction system addresses sequence specific errors in base calling, improving accuracy and reducing computational delays by detecting and correcting errors outside the real-time process.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-02
AI Technical Summary
Existing base calling systems fail to accurately and efficiently correct sequence specific errors, particularly those caused by low complexity sequences like homopolymers and repeating dinucleotides, leading to inaccuracies in nucleotide reads and downstream applications such as variant calling.
A machine learning-based offline correction system that detects sequence specific errors during sequencing and performs corrections outside the real-time process, using a neural network to generate accurate base calls and quality scores.
The system significantly reduces computational delays and improves accuracy of base calling, particularly at error-prone positions, enhancing the reliability of downstream applications by correcting sequence specific errors efficiently.
Smart Images

Figure US2025048223_02042026_PF_FP_ABST
Abstract
Description
DETERMINING OFFLINE CORRECTIONS FOR SEQUENCE SPECIFIC ERRORS CAUSED BY LOW COMPLEXITY SEQUENCES USING A MACHINE LEARNING MODELCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 700,018, filed on September 27, 2024, entitled “DETERMINING OFFLINE CORRECTIONS FOR SEQUENCE SPECIFIC ERRORS CAUSED BY LOW COMPLEXITY SEQUENCES USING A MACHINE LEARNING MODEL,” (IP-2817-PRV), which is incorporated herein by reference in its entirety.BACKGROUND
[0002] In recent years, biotechnology firms and research institutions have improved hardware and software platforms for performing base calling for a genomic sample or other nucleic-acid polymer. Indeed, systems and methods have been developed for analyzing nucleotide sequences and identifying the nucleobases contained therein. In some cases, by identifying the nucleobases, these systems and methods identify sequences of repeating nucleotide units or other motifs within the nucleotides. As an example, some existing base calling systems identify individual nucleobases of nucleotide sequences by using conventional Sanger sequencing or sequencing-by-synthesis (SBS). When using SBS, existing systems can monitor millions or billions of oligonucleotides that are grouped into clusters and synthesized in parallel to detect more accurate base calls. For instance, a camera device (or other light detector) of a specialized sequencing device implemented in SBS systems can capture images of irradiated fluorescent tags from nucleobases incorporated into such clustered and synthesized oligonucleotides. After capturing the images, existing SBS platforms send image data to sequencing-data-analysis software within the sequencing device (or on a separate computing device) to determine a nucleotide-base sequence for a genomic sample or other nucleic-acid polymer. For instance, the sequencing-data-analysis software can determine the nucleobases with tags that irradiate in a given image based on the light signal captured in the image data. By cyclically incorporating nucleobases into the oligonucleotides and capturing images of the emitted light signals in various sequencing cycles, the SBS systems can use the sequencing device to determine nucleotide reads corresponding to particular clusters and further determine the sequence of nucleobases (including any particular motifs) that are present in a whole genomic sample or other samples of nucleic-acid polymers.
[0003] Despite these advances, existing base calling systems often suffer from technical limitations that impair their flexibility and the accuracy of their determined base calls. While some of these systems have attempted to improve base call accuracy, these attempts have often led toAttorney Docket No. IP-2817-PCT 1 Patent Applicationcomputational inefficiencies and, in some instances, have resulted in complications that affect the operation of downstream applications, such as variant calling applications.
[0004] Indeed, existing base calling systems typically fail to flexibly accommodate the presence and effect of error-inducing sequences within oligonucleotides as part of their base calling process. To illustrate, certain low complexity sequences (e.g., homopolymers or sequences of repeating dinucleotides) are known to suffer from polymerase slippage during sequencing (e.g., via SBS). Further, certain chemical issues faced during the SBS process can lead to poor signal quality for the low complexity sequences. These problems can cause sequence specific errors, such as insertion-deletion (also referred to as “indel”) errors and / or mismatch errors, after the polymerase exits the low complexity sequence. Many existing base calling systems fail to identify such errorinducing sequences during the base calling process. Further, these systems often fail to implement measures to correct the errors caused by these sequences.
[0005] As existing base calling systems typically fail to detect error-inducing sequences and correct their corresponding errors, these systems often fail to produce an accurate determination of nucleobases within an oligonucleotide. Rather, these systems typically output nucleotide reads that include base calls that have been changed or shifted due to the presence of an error-inducing sequence. Such base calling errors can lead to further inaccuracies in downstream applications, such as variant calling applications.
[0006] As mentioned, some existing base calling systems have attempted solutions for correcting errors caused by certain sequences but have failed to do so in a computationally efficient manner and have even exacerbated problems in some instances. For instance, some existing systems use a type of real-time-based approach to base calling where sequencing images acquired by an imaging system are processed during a sequencing run without being saved to storage. Some of these systems have attempted to incorporate the correction of errors caused by certain sequences into their real-time-based base-calling process. By doing so, however, these systems tend to overcomplicate the real-time analyzing component through the use of multiple buffers that allows certain sequencing data to be retained and used to perform the error correction. This overcomplication often slows the real-time analyzing component so that base calling is either no longer performed in real time or is not performed accurately in real time. In many cases, the real- time-based approach can slow down the real-time analyzing component between 150-200%.
[0007] Additionally, existing base calling systems that attempt to correct errors in real time are often limited in computing power, leading to poor correction results. Indeed, existing sequencing platforms — such as SBS platforms — are sometimes implemented on devices with limited or fixed computing power (e.g., using a Field Programmable Gate Arrays (FPGA)). Additionally, the realtime base calling process used by such platforms are often computationally demanding, leavingAttorney Docket No. IP-2817-PCT 2 Patent Applicationrelatively few computing resources for other processes, such as error correction, to be incorporated. Accordingly, existing systems tend to use error correction models that are relatively unsophisticated with the purpose of fitting the resource demands of the models within the remaining computing budget. While such error correction models do provide some improvements to base calling, their relatively simplified analysis leaves many erroneous base calls improperly corrected or uncorrected altogether.
[0008] Further, existing base calling systems that incorporate error correction into their realtime base calling process often fail to correct a number of errors, which can degrade the accuracy of downstream applications that rely on the base calls, such as variant calling applications. For instance, while existing systems using real-time correction techniques may provide some improvement to base calling accuracy, they often fail to make corrections to a significant number of errors, especially those errors at certain read positions, such as the first position following an error-inducing sequence. A variant calling application, however, may filter out nucleotide reads that contain too many errors to prevent those reads from negatively influencing the variant calling analysis. By correcting some base calling errors for a nucleotide read without correcting all the included errors, existing systems that perform error correction as part of their real-time base calling process tend to increase the number of error-prone nucleotide reads that are used in the variant calling analysis. Accordingly, these systems facilitate errors in the variant calling application, such as by increasing the number of false positives.
[0009] These, along with additional problems and issues, exist in existing base calling systems. Systems and methods for addressing these issues will be described below.SUMMARY
[0010] This disclosure describes one or more embodiments of systems, methods, and non- transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can implement a flexible approach to base calling that involves the detection during a sequencing run and machine learning-based offline correction of sequence specific errors. For instance, the disclosed systems can detect the presence of certain low complexity, error-inducing sequences — such as homopolymers or repeating dinucleotides — within a cluster of oligonucleotides while performing base calling as part of a sequencing run (e.g., SBS sequencing run). The disclosed systems can further write, to storage, read data associated with base calls affected by such present error-inducing sequences and implement error correction for those base calls as a supplementary process. For example, the disclosed systems can perform the error correction using an offline machine learning model based on the stored read data. The disclosed systems can further use a quality model to generate updated quality scores for the base calls that have been corrected. In this manner, theAttorney Docket No. IP-2817-PCT 3 Patent Applicationdisclosed systems can flexibly accommodate the presence and effect of error-inducing sequences within oligonucleotides as part of the base calling process while accurately and efficiently determining their base calls and corresponding quality scores.
[0011] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] This disclosure will describe one or more embodiments of the invention with additional specificity and detail by referencing the accompanying figures. The following paragraphs briefly describe those figures, in which:
[0013] FIG. 1 illustrates a schematic diagram of a computing system in which a base call correction system can operate in accordance with one or more embodiments.
[0014] FIG. 2 illustrates an overview of the base call correction system using a machine learning model to generate base call corrections for base calls associated with a sequence specific error in accordance with one or more embodiments.
[0015] FIGS. 3A-3D each illustrate the base call correction system detecting a base call sequence known to cause a sequence specific error in accordance with one or more embodiments.
[0016] FIG. 4 illustrates using a machine learning model to generate a base call correction and corresponding quality score in accordance with one or more embodiments.
[0017] FIG. 5 illustrates generating corrected intensity data from raw intensity data in accordance with one or more embodiments.
[0018] FIG. 6 illustrates a diagram graph reflecting experimental results regarding the effectiveness of the base call correction system in accordance with one or more embodiments.
[0019] FIG. 7 illustrates graphs reflecting additional experimental results regarding the effectiveness of the base call correction system in accordance with one or more embodiments.
[0020] FIG. 8 illustrates graphs reflecting further experimental results regarding the effectiveness of the base call correction system in accordance with one or more embodiments.
[0021] FIG. 9 illustrates diagrams reflecting yet further experimental results regarding the effectiveness of the base call correction system in generating well-calibrated quality scores in accordance with one or more embodiments.
[0022] FIG. 10 illustrates a graph reflecting additional experimental results regarding the effectiveness of various implementations of the base call correction system in accordance with one or more embodiments.Attorney Docket No. IP-2817-PCT 4 Patent Application
[0023] FIG. 11 illustrates a flowchart of a series of acts for using a machine learning model to generate base call corrections for base calls associated with a sequence specific error in accordance with one or more embodiments.
[0024] FIG. 12 illustrates a block diagram of an example computing device in accordance with one or more embodiments.DETAILED DESCRIPTION
[0025] This disclosure describes one or more embodiments of a base call correction system that corrects sequence specific errors for a nucleotide read offline using a machine learning model based on detecting sequences causing such errors during a sequencing cycle, sequencing run, or other real-time-based approach to base calling. To illustrate, in one or more embodiments, the base call correction system performs base calling for a cluster of oligonucleotides as part of the sequencing process (e.g., in real time). During the base calling process, the base call correction system can detect, within the oligonucleotides, the presence of one or more sequences — such as homopolymers, repeating dinucleotides, or other motifs — that are known to cause sequence specific errors in subsequent base calls. The base call correction system can write the read data gathered for the subsequent base calls (e.g., the raw intensity data from the images captured during sequencing, corrected intensity data, and / or the base calls themselves) to storage and perform offline corrections for those base calls. For instance, the base call correction system can perform the offline corrections using a machine learning model, such as a convolutional neural network or a recurrent neural network. Accordingly, the base call correction system can generate a sequence of base calls for the cluster of oligonucleotides that incorporates the offline corrections.
[0026] To illustrate, in one or more embodiments, the base call correction system determines, from read data of a cluster of oligonucleotides during a sequencing run, a set of base calls associated with a sequence specific error and a set of read data for the set of base calls. The base call correction system additionally writes, to storage, the set of read data for the set of base calls associated with the sequence specific error. Based on the set of read data written to storage, the base call correction system determines one or more base call corrections for the set of base calls using a machine learning model. Further, the base call correction model generates one or more corrected base calls for the read data of the cluster of oligonucleotides based on the one or more base call corrections.
[0027] Indeed, as mentioned above, in one or more embodiments, the base call correction system generates base calls for a cluster of oligonucleotides. In particular, the base call correction system determines base calls for the oligonucleotides during a sequencing cycle, sequencing run, or other real-time-approach to base calling. The base call correction system can further identify, during a sequencing run or other real-time-approach to base calling, a set of base calls that are associated with a sequence specific error (e.g., base calls that may be incorrect due to the error).Attorney Docket No. IP-2817-PCT 5 Patent ApplicationAfter performing corrections for the identified set offline (e.g., externally to the real-time process), the base call correction system produces a final sequence of base calls that can incorporate the corrections and can also incorporate one or more of the base calls determined in real time.
[0028] More specifically, as mentioned, the base call correction system can identify, within a cluster of oligonucleotides, a set of base calls that are associated with a sequence specific error using read data determined for the cluster (e.g., read data determined via real-time base calling). For instance, in some cases, the base call correction system uses the read data to determine, within the oligonucleotides, the presence of a base call sequence (e.g., a homopolymer, sequence of repeating dinucleotides, or other motif of nucleotides) that causes the sequence specific error. The base call correction system can further determine that base calls following the error-inducing sequence are associated with the sequence specific error. In other words, the base call correction system can determine that the following base calls are potentially incorrect due to the presence of the error-inducing sequence.
[0029] In one or more embodiments, the base call correction system further writes a set of read data for the set of base calls associated with the sequence specific error to storage. For instance, in some cases, the base call correction system writes the set of read data to non-volatile storage (including non-transitory storage or transitory storage). Thus, the base call correction system stores the set of read data for use offline (e.g., outside the real-time base calling process). In some cases, the base call correction system further writes model parameters (e.g., neural network parameters or other machine learning parameters) to storage for use in offline correction.
[0030] As further mentioned, the base call correction system can use the stored set of read data to determine (e.g., offline) corrections for those base calls associated with the sequence specific error. In some embodiments, the base call correction system uses a correction model, such as a machine learning model, to determine the corrections based on the stored set of read data. For instance, in some cases, the base call correction system uses a neural network to determine the corrections. In some cases, by determining the corrections, the base call correction system determines correct base calls for the set of base calls associated with the sequence specific error. Further, the base call correction system determines or updates quality scores for those base calls for which corrections are determined.
[0031] To illustrate, in some cases, the base call correction system uses a machine learning model to generate, from the set of read data, a set of probabilities indicating likelihoods that a corresponding nucleobase includes adenine, cytosine, guanine, or thymine and / or a likelihood of a non-call. The base call correction system can determine a base call correction for the nucleobase using the set of probabilities (e.g., determine the base call correction based on the highest probability). Further, the base call correction system can determine a quality score for the base callAttorney Docket No. IP-2817-PCT 6 Patent Applicationcorrection using the set of probabilities output by the machine learning model (e.g., the highest probability).
[0032] Based on the determined corrections, the base call correction system can determine corrected base calls for the read data of the cluster of oligonucleotides. For instance, the base call correction system can integrate the corrections into the read data (e.g., the read data determined via the real-time process) to produce a final set of base calls that corrects erroneous calls that were present in the read data due to the sequence specific error. In some cases, the base call correction system generates a file, such as a fast-all quality (FASTQ) file, having the final set of base calls and quality scores. Indeed, the base call correction system can include quality scores for the final set of base calls within the generated file (including the updated quality scores determined for the base call corrections).
[0033] As indicated above, the base call correction system provides several technical advantages over existing base calling systems. For instance, the base call correction system avoids the computing-time delays and inaccuracies relative to existing systems by implementing a new configuration for base calling that involves (i) detection during a sequencing run of sequence specific errors or other real-time-based detection and (ii) offline correction of such sequence specific errors. Indeed, in contrast to prior attempts of correcting base call errors in real time — which significantly slow down the real-time base calling process — the base call correction system detects the errors during a sequencing cycle or sequencing run but performs the actual corrections outside of that real-time-based approach to base calling. Indeed, the base call correction system can significantly reduce the slow-down of the real-time analyzing component from the 150-200% slow down caused by the real-time-based approach. Thus, the base call correction system can correct for errors while maintaining computational efficiency for real-time base calling — while maintaining a more accurate approach to correcting sequence specific errors.
[0034] As just noted, by performing corrections offline, the base call correction system preserves computing resources and generates more accurate base calls compared to those systems attempting real-time corrections. For instance, by performing corrections offline, the base call correction system frees up one or more computer processors on a sequencing device, memory, and / or other computational resources for the correction process. Indeed, by performing corrections offline, the base call correction system can implement the correction process without competing with the real-time base calling process for compute power, such as the processing required to perform real-time base calling. Thus, in many cases, the base call correction system can perform the corrections using a comparatively more sophisticated model — such as a trained machine learning model — that corrects for error via a more complex analysis of corresponding read data.Attorney Docket No. IP-2817-PCT 7 Patent Application
[0035] In addition to preserving computing resources, the base call correction system improves the operation of a specialized instrument or computing device — that is, a sequencing device — by improving base-calling accuracy. In some embodiments, for example, the base call correction system generates more accurate base calls for oligonucleotide clusters in general and, in particular, more accurate base calls at the first read position that follows an error-inducing sequence and / or other read positions that follow such an error-inducing sequence. Thus, the base call correction system facilitates a more accurate operation of downstream applications, such as variant calling applications, which rely on the accuracy of the base calls.
[0036] In addition or in the alternative to resolving slow-downs and inaccuracies of real-time- approaches to base calling, the base call correction system further improves the operational flexibility of a specialized instrument or computing device — that is, a sequencing device — to basecall correction by implementing an approach to base calling that accommodates the presence and effect of certain error-inducing sequences within the oligonucleotide clusters being analyzed. Indeed, the base call correction system can flexibly detect certain low complexity sequences known to cause sequence specific errors and make the necessary corrections, where many existing systems are susceptible to such errors due to polymerase slippage and other chemistry issues involved in sequencing low complexity sequences. By making the necessary corrections, the base call correction system mitigates sequence specific errors for improved base calling. As explained further below, the base call correction system can detect and correct homopolymers, repeating dinucleotides, repeating trinucleotides, or other repeat nucleotide units known to cause sequence specific errors that go uncorrected by existing systems.
[0037] Further, as mentioned, the base call correction system generates more accurate base calls for oligonucleotide clusters when compared to existing systems. Indeed, by detecting and correcting sequence specific errors within the oligonucleotide clusters, the base call correction system produces an output that provides a more accurate set of base calls. For instance, it is estimated that homopolymer clusters represent about five percent of the clusters that are analyzed during sequencing; thus, correcting for errors within these clusters can lead to a significant improvement over the base calling available under existing systems. In addition to correcting homopolymer clusters, the base call correction system can correct clusters having sequences of dinucleotide repeats or other motifs, leading to further improvements in base calling accuracy. Improvements in accuracy extend to the quality scores determined for the base calls as well. Indeed, through its offline correction process, the base call correction system more accurately measures the quality of produced base calls. For instance, the base call correction system leverages the more sophisticated analysis used for correction to also determine base call quality scores for the corrected base calls.Attorney Docket No. IP-2817-PCT 8 Patent Application
[0038] As illustrated by the foregoing discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the base call correction system. For example, as used herein, the term “oligonucleotide” refers to an oligomer or other polymer of nucleotides or mimetics. In particular, an oligonucleotide can include a synthetic or natural molecule comprising a sequence of covalently linked nucleotides formed by a modified phosphodiester or phosphodiester bond between the 3’ position of the pentose in a nucleotide and the 5’ position of the pentose in a nucleotide adjacent. For example, an oligonucleotide can include a short DNA or RNA molecule annealed to a single-stranded polynucleotide to be analyzed or sequenced as part of SBS sequencing.
[0039] Additionally, as used herein, the term “cluster of oligonucleotides” (or simply “cluster”) refers to a localized group or collection of oligonucleotides on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same oligonucleotide. For example, in some embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-pattemed flow cell. A cluster can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.
[0040] As used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample (e.g., genomic sample). In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device that incorporate nucleobases into growing oligonucleotides to determine nucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell or other nucleotide-sample slide. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more genomic samples seeded in clusters throughout a flow cell or other nucleotide-sample slide. Upon completing a sequencing run, a sequencing device can generate nucleotide-base-call data in a fde, such as a binary base call (BCL) sequence fde or a fast-all quality (FASTQ) file.
[0041] Additionally, as used herein, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating a nucleotide base to an oligonucleotide or an iteration of adding or incorporating nucleotide bases to oligonucleotides in parallel. In particular, a cycle can include an iteration of taking and analyzing one or more signals and / or images with data indicating individual nucleotide bases added or incorporated into an oligonucleotide or to oligonucleotides inAttorney Docket No. IP-2817-PCT 9 Patent Applicationparallel. Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer. For example, in some embodiments, each sequencing cycle involves either single reads in which DNA or RNA strands are read in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining a particular nucleotide base added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleotide bases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within a Sequencing-By-Synthesis (SBS) run.
[0042] As further used herein, the term “base call” (or “nucleobase call”) refers to a determination or prediction of a particular nucleobase (or nucleobase pair) for an oligonucleotide (e.g., nucleotide read) during a sequencing cycle or for a genomic coordinate of a genomic sample. In particular, a nucleobase call can indicate a determination or prediction of the type of nucleobase that has been incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., read-based nucleobase calls). In some cases, for a nucleotide read, a nucleobase call includes a determination or a prediction of a nucleobase based on intensity values resulting from fluorescent-tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., in a cluster of a flow cell). As suggested above, a single nucleobase call can be an adenine (A) call, a cytosine (C) call, a guanine (G) call, a thymine (T) call, or an uracil (U) call. In some embodiments, the type of base (e.g., adenine, cytosine, thymine, or guanine) can be determined based on intensity values for a signal emitted by labeled nucleotide bases in a cluster of oligonucleotides, such as signals in 16 quadrature amplitude modulation (QAM) or pulse amplitude modulation (PAM) 4 format.
[0043] As used herein, the term “nucleotide read” refers to an inferred sequence of one or more nucleotide bases (or nucleobase pairs) from all or part of a sample nucleotide sequence. In particular, a nucleotide read includes a determined or predicted sequence of base calls for a nucleotide fragment (e.g., an oligonucleotide) or group of monoclonal nucleotide fragments (e.g., a cluster of oligonucleotides) from a sequencing library corresponding to a genome sample. For example, a sequencing device can determine a nucleotide read during a sequencing run by generating base calls for nucleotide bases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a cluster in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g.,Attorney Docket No. IP-2817-PCT 10 Patent Applicationassembled nucleotide reads) satisfying a threshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.
[0044] Additionally, as used herein, the term “read data” includes data associated with one or more nucleotide reads. In particular, read data can include data that is generated, extracted, or otherwise determined for a nucleotide sequence (e.g., an oligonucleotide or cluster or oligonucleotides) during a sequencing run. For instance, read data can include data that is determined when generating one or more nucleotide reads, such as the base calls determined for the nucleotide read. Read data can also include other data determined when generating a nucleotide read, such as quality scores associated with the determined base calls, intensity data determined for the sequencing run (e.g., raw intensity data or corrected intensity data), a read length, a determination of whether the nucleotide sequence includes an error-inducing sequence (e.g., a homopolymer or repeating dinucleotides), count data from a counter used in determining the presence of an error-inducing sequence, correction parameters used in correcting intensity data, or base call distribution clouds in determining the base calls for the nucleotide read.
[0045] Further, as used herein, the term “read position” refers to a location or coordinate on nucleotide read. In particular, a read position includes a location along a nucleotide read to which a labeled nucleotide has been added. For example, a read position can indicate a position within a nucleotide read at which a most-recently added labeled nucleotide to corresponding oligonucleotides within a cluster when a camera captures an image of a nucleotide-sample slide or a section of the nucleotide-sample slide. But a read position can also refer to any other position within the nucleotide read, such as those positions preceding the position for the most-recently added labeled nucleotide. Accordingly, in some cases, each read position within a nucleotide read corresponds to a sequencing cycle in a sequencing run for determining the nucleotide read.
[0046] As used herein, the term “homopolymer” refers to a polymer in which all the monomeric units are the same. For instance, a homopolymer can include all or a portion of a nucleotide fragment (e.g., an oligonucleotide) where the nucleotide bases therein are identical. In some cases, a homopolymer includes at least a threshold number of identical nucleotide bases. Indeed, in some instances, the base call correction system determines that all or a portion of a nucleotide fragment includes a homopolymer after determining or predicting that a sequence of identical nucleotide bases included therein equals or exceeds a threshold. At least, in some cases, the base call correction system uses a threshold to determine whether a homopolymer is long enough to cause a sequence specific error within the base calls that follow.
[0047] Additionally, as used herein, the term “sequence of repeating dinucleotides” (or simply “repeating dinucleotides”) refers to a polymer having a repeating pair of monomeric subunits. ForAttorney Docket No. IP-2817-PCT 11 Patent Applicationinstance, a sequence of repeating dinucleotides can include all or a portion of a nucleotide fragment (e.g., an oligonucleotide) where pairs of nucleotide bases therein are identical. In some cases, a sequence of repeating dinucleotides includes a pair of nucleotide bases that is repeated in the same order and without interruption by one or more other nucleotide bases. In some implementations, a sequence of repeating dinucleotides includes at least a threshold number of repetitions. Indeed, in some instances, the base call correction system determines that all or a portion of a nucleotide fragment includes a sequence of repeating dinucleotides after determining or predicting that a number of repetitions of a pair of nucleotide bases included therein equals or exceeds a threshold. At least, in some cases, the base call correction system uses a threshold to determine whether a sequence of repeating dinucleotides is long enough to cause a sequence specific error within the base calls that follow.
[0048] Similarly, as used herein, the term “sequence of repeating trinucleotides” (or simply “repeating trinucleotides”) refers to a polymer having a repeating triplet of monomeric subunits. For instance, a sequence of repeating trinucleotides can include all or a portion of a nucleotide fragment (e.g., an oligonucleotide) where triplets of nucleotide bases therein are identical. In some cases, a sequence of repeating trinucleotides includes a triplet of nucleotide bases that is repeated in the same order and without interruption by one or more other nucleotide bases. In some implementations, a sequence of repeating trinucleotides includes at least a threshold number of repetitions. Indeed, in some instances, the base call correction system determines that all or a portion of a nucleotide fragment includes a sequence of repeating trinucleotides after determining or predicting that a number of repetitions of a triplet of nucleotide bases included therein equals or exceeds a threshold. At least, in some cases, the base call correction system uses a threshold to determine whether a sequence of repeating trinucleotides is long enough to cause a sequence specific error within the base calls that follow.
[0049] As used herein, the term “sequence specific error” refers to a sequencing error caused by a particular sequence (e.g., motif) of nucleotide bases within the nucleotide fragment for which sequencing is performed. For instance, a sequence specific error can include one or more errors that result in erroneous base calls for nucleotide sequences that follow certain error-inducing sequences, such as homopolymers, sequences of repeating dinucleotides, or sequences of repeating trinucleotide. To illustrate, a sequence specific error can include, but is not limited to, an indel error (i.e., an insertion-deletion error), or a mismatch error.
[0050] Additionally, as used herein, the term “intensity data” refers to light or image data. In particular, intensity data can include light data or image data used in determining a base call for a nucleotide fragment. For instance, intensity data can include data corresponding to the intensity and / or wavelength of light captured within a digital image as part of a sequencing cycle. In someAttorney Docket No. IP-2817-PCT 12 Patent Applicationcases, the light is emitted by an irradiated fluorescent tag that is added to an oligonucleotide as part of the sequencing cycle. In some instances, intensity data more generally includes data corresponding to the intensity and / or wavelength of light associated with (e.g., emitted from) a cluster of oligonucleotides during a sequencing cycle. To illustrate, in some cases, intensity data includes data indicating a characteristic or attribute of a signal emitted, reflected, or otherwise communicated from a labeled nucleotide base or a group of labeled nucleotide bases from a cluster of oligonucleotides. For instance, intensity data can include one or more values associated with a color intensity (e.g., wavelength) or a light intensity (e.g., brightness). In some cases, the base call correction system captures several images of a cluster of oligonucleotides with labeled nucleotide bases using different channels. Thus, intensity data can correspond to the intensity of the signal as observed through a particular channel. As used herein, the term “raw intensity data” refers to the intensity data that is initially captured or otherwise determined for a sequencing cycle. Additionally, as used herein, the term “corrected intensity data” includes intensity data that has been derived from other intensity data, such as raw intensity data. Indeed, as will be discussed below, the base call correction system can generate corrected intensity data from raw intensity data to account for certain errors in the sequencing process (e.g., errors in the chemistry or image capturing) that may cause the raw intensity data fail to accurately reflect the corresponding nucleobase. In some instances, the base call correction system uses one or more models and corresponding correction parameters for the model(s) to generate corrected intensity data from raw intensity data. In some cases, the base call correction system generates corrected intensity data for a particular sequencing cycle from the raw intensity data determined for the sequencing cycle and / or other sequencing cycles.
[0051] Additionally, as used herein, the term “quality score” (or “base-call-quality score”) refers to a specific score, value, or other measurement indicating a quality of a base call. In particular, a quality score can include a quantitative value indicating (i) a confidence in a base call (including an initial or corrected base call) or (ii) a probability of an error in the base call. For instance, a quality score can include a numerical value where a relatively higher value indicates a relatively higher confidence in the accuracy of the corresponding base call (i.e., a lower probability of error). To illustrate, in some embodiments, a quality score includes a Q score (e.g., aPHil’s Read EDitor (Phred) quality score) predicting the error probability of a given nucleobase call. To illustrate, a quality score (or Q score) may indicate that a probability of an incorrect base call at a genomic coordinate is equal to 1 in 100 for a Q20 score, 1 in 1,000 for a Q30 score, 1 in 10,000 for a Q40 score, etc. In some cases, the quality score is generated by a model or an algorithm, either of which can be scaled to be consistent with a Phred scale.Attorney Docket No. IP-2817-PCT 13 Patent Application
[0052] Indeed, as used herein, the term “quality model” refers to a computer-implemented model or algorithm that generates quality scores for base calls. For instance, in some cases, a quality model includes a computer-implemented model or algorithm that analyzes a base call and / or data (e.g., read data) associated with the base call and generates a quality score for the base call based on the analysis. In some instances, a quality model includes a machine learning model trained to determine quality scores. In some instances, a quality model includes a logarithmic algorithm for determining quality scores.
[0053] Further, as used herein, the term “base call correction” refers to a correction of an erroneous base call. In particular, in some embodiments, abase call correction includes a correction made to an erroneous base call that is associated with a sequence specific error. For instance, in some cases, a base call correction includes a modified base call. To illustrate, in some cases, a base call correction includes a modified base call generated for an erroneous base call by a correction model (e.g., a machine learning model). In some instances, a base call correction is generated as part of a binary file, such as a .obcl file, that is an intermediary file generated by the base call correction system.
[0054] As used herein, the term “correction model” includes a computer-implemented model or algorithm for correcting erroneous base calls. In particular, in some embodiments, a correction model includes a computer-implemented model or algorithm for generating a base call correction for an erroneous base call. In some cases, a correction model generates a base call correction using a set of read data associated with the erroneous base call. In some embodiments, a correction model generates base call corrections within a binary file, such as a .obcl file, as an intermediary output of the base call correction system. In some implementations, a correction model includes a machine learning model, such as a convolutional neural network or a recurrent neural network. More detail regarding the correction model used by the base call correction system will be provided below.
[0055] Additionally, as used herein, the term “corrected base call” includes an updated or finalized base call of the base calling process. In some cases, a corrected base call corresponds to a base call correction. For instance, in some implementations, a corrected base call includes a modified base call generated by a correction model for a previously determined, erroneous base call. In some implementations, however, a corrected base call corresponds to a base call determined via real time base calling. In other words, a corrected base call can include a previously determined (e.g., uncorrected) base call. For instance, a corrected base call can correspond to a base call previously determined via real time base calling based on determining the base call is not associated with a sequence specific error. In some embodiments, while a base call and a base call correction are included in an intermediate binary file (e.g., a binary base call (BCL) file or a .ocbl file), the base call correction system can include a corrected base call in a different file, such as a fast-allAttorney Docket No. IP-2817-PCT 14 Patent Applicationquality (FASTQ) file. For example, as will be discussed in more detail below, the base call correction system can use a base call data convert module to generate a FASTQ file as the result of a base calling process at least partially based on a BCL file generated using a real time analyzer and / or a .obcl file generated using a correction model.
[0056] As used herein, the term “machine learning model” refers to a computer-implemented model that can be tuned (e.g., trained) based on inputs to approximate unknown functions. In particular, in some embodiments, a machine-learning model includes a model that utilizes algorithms to leam from, and make predictions on, known data by analyzing the known data to leam to generate outputs that reflect patterns and attributes of the known data. For instance, in some cases, a machine-learning model includes, but is not limited to, a neural network (e.g., a convolutional neural network, recurrent neural network or other deep learning network), a decision tree (e.g., a gradient boosted decision tree, such as XGBoost), association rule learning, inductive logic programming, support vector learning, a Bayesian network, a regression-based model (e.g., censored regression), principal component analysis, or a combination thereof.
[0057] Further, as used herein, the term “neural network” refers to a model of interconnected artificial neurons (e.g., organized in layers) that communicate and leam to approximate complex functions and generate outputs based on inputs provided to the model. In some instances, a neural network includes one or more machine learning algorithms. Further, in some cases, a neural network includes an algorithm (or set of algorithms) that implements deep learning techniques that utilize a set of algorithms to model high-level abstractions in data. To illustrate, in some embodiments, a neural network includes a convolutional neural network, a recurrent neural network (e.g., a long short-term memory neural network), a generative adversarial network, a graph neural network, a multi-layer perceptron, or a diffusion neural network. In some embodiments, a neural network includes a combination of neural networks or neural network components.
[0058] The following paragraphs describe the base call correction system with respect to illustrative figures that portray example embodiments and implementations. FIG. 1 illustrates a schematic diagram of a computing system 100 in which a base call correction system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes one or more server device(s) 102 connected to a client device 110 and a sequencing device 114 via a network 108. While FIG. 1 shows one embodiment of the base call correction system 106, this disclosure describes alternative embodiments and configurations below.
[0059] As shown in FIG. 1, the server device(s) 102, the client device 110, and the sequencing device 114 are connected via the network 108. Accordingly, each of the components of the computing system 100 can communicate via the network 108. The network 108 comprises anyAttorney Docket No. IP-2817-PCT 15 Patent Applicationsuitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 12.
[0060] As indicated by FIG. 1, the sequencing device 114 comprises a computing device for sequencing a genomic sample or other nucleic-acid polymer. In some embodiments, the sequencing device 114 analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer-implemented methods and systems either directly or indirectly on the sequencing device 114. More particularly, the sequencing device 114 receives nucleotide-sample slides (e.g., flow cells) comprising nucleotide fragments (e.g., genomic DNA fragments) extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments.
[0061] In addition, or in the alternative to communicating across the network 108, in some embodiments, the sequencing device 114 bypasses the network 108 and communicates directly with the server device(s) 102 or the client device 110. Additionally, as shown in FIG. 1, in one or more embodiments, the sequencing device 114 includes the base call correction system 106.
[0062] As further indicated by FIG. 1, the server device(s) 102 may generate, receive, analyze, store, and transmit digital data, such as data for amino-acid sequences or nucleotide sequences. As shown in FIG. 1, the sequencing device 114 may send (and the server device(s) 102 may receive) various data from the sequencing device 114, including data representing amino-acid sequences or nucleotide sequences (e.g., read data, base calls, and / or corresponding quality scores). The server device(s) 102 may also communicate with the client device 110. In particular, the server device(s) 102 can send data representing amino-acid sequences or nucleotide sequences (e.g., read data, base calls, and / or corresponding quality scores) to the client device 110.
[0063] Additionally, as shown in FIG. 1, the server device(s) 102 can include the base call correction system 106. In one or more embodiments, as explained further below, the base call correction system 106 analyzes a cluster of oligonucleotides (e.g., analyzes images captured during a sequencing run) and generates base calls for the oligonucleotides. During the base calling process, the base call correction system 106 can identify a set of base calls associated with a sequence specific error. For instance, the base call correction system 106 can identify a sequence (e.g., a motif) within the oligonucleotides — such as a homopolymer or a sequence of repeating dinucleotides — that causes the sequence specific error and determine that the base calls following the sequence are associated with the sequence specific error. The base call correction system 106 can further employ a correction model 104 (e.g., a machine learning model 116) to generate base call corrections for the base calls associated with the sequence specific error and output corrected base calls. In some cases, the base call correction system 106 further uses a quality model to generate quality scores that correspond to the base call corrections.Attorney Docket No. IP-2817-PCT 16 Patent Application
[0064] In some embodiments, the server device(s) 102 comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 108 and located in the same or different physical locations. Further, the server device(s) 102 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0065] In some cases, the server device(s) 102 is located at or near a same physical location of the sequencing device 114 or remotely from the sequencing device 114. Indeed, in some embodiments, the server device(s) 102 and the sequencing device 114 are integrated into a same computing device. The server device(s) 102 may run software on the sequencing device 114 or the base call correction system 106 to generate, receive, analyze, store, and transmit digital data, such as by sending or receiving corrected base calls and / or corresponding quality scores. In some embodiments, the sequencing device 114 or the base call correction system 106 store and access a database or table of corrected base calls and / or corresponding quality scores.
[0066] As further illustrated and indicated in FIG. 1, the client device 110 can generate, store, receive, and send digital data. In particular, the client device 110 can receive data for amino-acid sequences or nucleotide sequences (e.g., base calls and / or corresponding quality scores) from the server device(s) 102 and / or the sequencing device 114. The client device 110 can accordingly present data concerning sequenced oligonucleotides, such as by presenting determined base calls (including corrected base calls) and / or corresponding quality scores within a graphical user interface to a user associated with the client device 110.
[0067] The client device 110 illustrated in FIG. 1 may comprise one of various types of client devices. For example, in some embodiments, the client device 110 includes a non-mobile device, such as a desktop computer or server, or another type of client device. In yet other embodiments, the client device 110 includes a mobile device, such as a laptop, a tablet, a mobile telephone, or a smartphone. Additional details with regard to the client device 110 are discussed below with respect to FIG. 12.
[0068] As further illustrated in FIG. 1, the client device 110 includes a sequencing application 112. The sequencing application 112 may be a web application or a native application stored and executed on the client device 110 (e.g., a mobile application or a desktop application). The sequencing application 112 can include instructions that (when executed) cause the client device 110 to receive data from the base call correction system 106 and present data from the sequencing device 114 and / or the server device(s) 102. Furthermore, the sequencing application 112 can instruct the client device 110 to display base call data, such as base calls determined for a cluster of oligonucleotides (including corrected base calls) and / or associated quality scores.Attorney Docket No. IP-2817-PCT 17 Patent Application
[0069] As further illustrated in FIG. 1, the base call correction system 106 may be located on the client device 110 as part of the sequencing application 112 or on the sequencing device 114. Accordingly, in some embodiments, the base call correction system 106 is implemented by (e.g., located entirely or in part on) the client device 110. As mentioned, in yet other embodiments, the base call correction system 106 is implemented by one or more other components of the computing system 100, such as the sequencing device 114 or the server device(s) 102. In particular, the base call correction system 106 can be implemented in a variety of different ways across the server device(s) 102, the network 108, the client device 110, and the sequencing device 114.
[0070] Though FIG. 1 illustrates the components of the computing system 100 communicating via the network 108, in certain implementations, the components of the computing system 100 can also communicate directly with each other, bypassing the network 108. For instance, and as previously mentioned, in some implementations, the client device 110 communicates directly with the sequencing device 114. Additionally, in some embodiments, the client device 110 communicates directly with the base call correction system 106. Moreover, the base call correction system 106 can access one or more databases housed on or accessed by the server device(s) 102 or elsewhere in the computing system 100. For instance, FIG. 1 illustrates a database 118 of motifs known to cause sequence specific errors that is accessible by the server device(s) 102. Indeed, in some cases, the base call correction system 106 maintains / accesses the database 118 for use in monitoring a nucleotide read to detect such error-inducing sequences. Though FIG. 1 illustrates the database 118 as a separate component, the database 118 can be part of the server device(s) 102 or another device within the computing system 100.
[0071] As discussed above, the base call correction system 106 can analyze a cluster of oligonucleotides and generate base calls based on the analysis. Further, the base call correction system 106 can identify base calls associated with a sequence specific error and generate corrected base calls that address the error and updated quality scores based on the corrections. FIG. 2 illustrates the base call correction system 106 generating one or more corrected base calls and corresponding quality scores in accordance with one or more embodiments.
[0072] As shown in FIG. 2, the base call correction system 106 analyzes a cluster of oligonucleotides 202. In particular, the base call correction system 106 performs a sequencing run 204 on the cluster of oligonucleotides 202. In one or more embodiments, the sequencing run 204 includes a real time process. Indeed, as shown, during the sequencing run 204, the base call correction system 106 uses a real time analyzer 206 to analyze the cluster of oligonucleotides 202. FIG. 2 illustrates the real time analyzer 206 employing a base caller 208, a sequence specific error detector 210, and a read data captor 212, though additional, fewer, or alternative components can be used in various implementations.Attorney Docket No. IP-2817-PCT 18 Patent Application
[0073] As used herein the term “real time analyzer” includes a computer-implemented model that generates base calls in real time or near real time. In particular, a real time analyzer can include a computer-implemented model that analyzes one or more digital images captured during a sequencing cycle of a sequencing run (e.g., captured by a camera device or other light detector of a sequencing device) and generates — in real time or near real time — one or more base calls for clusters of one or more oligonucleotides represented within the digital images. For instance, in some cases, the real time analyzer can generate the base call(s) without writing the one or more digital images to disk.
[0074] As indicated in FIG. 2, the base call correction system 106 uses the real time analyzer 206 to generate a base call data fde 214. In particular, in some embodiments, the base call correction system 106 generates the base call data file 214 via a base calling process using the base caller 208 of the real time analyzer 206. To illustrate, the base call correction system 106 can use the base caller 208 to analyze image data captured for the cluster of oligonucleotides 202 and generate base calls 216 and, optionally, quality scores 250 accordingly. For instance, in some embodiments, the base call correction system 106 uses the base caller 208 to generate base calls as part of a sequencing-by-synthesis process where the image data corresponds to images of irradiated fluorescent tags from nucleobases incorporated into the cluster of oligonucleotides 202. As the base call correction system 106 can perform the base calling in real time, the base call correction system 106 can process the image data without writing the image data to storage in some implementations.
[0075] Further, in some implementations, the base call correction system 106 uses the real time analyzer 206 to generate the quality scores 250 via a PHil’s Read EDitor Phred algorithm / model. For instance, in some implementations, the base call correction system 106 uses the real time analyzer 206 to determine Phred quality scores using a Phred algorithm or a modified Phred algorithm, such as the algorithm described in U.S. Patent No. 8,392,126 filed on September 23, 2009, entitled Method and System for Determining the Accuracy of DNA Base Identifications, which is incorporated herein by reference in its entirety. In one or more embodiments, the base call correction system 106 generates the quality scores 250 using a quality score table (or Q-table) that maps certain combinations of values determined from sequencing cycle image data (e.g., signal- to-noise ratio and / or chastity value) or other read data to corresponding quality scores (Q-scores). In some cases, the base call correction system 106 generates the Q-table using the Phred algorithm or modified Phred algorithm. In some instances, the base call correction system 106 updates or recalibrates the Q-table to accommodate changes in the values that are used in creating the table (e.g., based on changes in the values due to signal corrections that are determined as part of the offline correction process).Attorney Docket No. IP-2817-PCT 19 Patent Application
[0076] Thus, the base calls 216 of the base call data file 214 can include those base calls determined during the sequencing run 204 (e.g., determined via real time base calling performed during the sequencing run 204). In particular, the base calls 216 can provide a nucleotide read for the cluster of oligonucleotides 202. In some instances, the base call data file 214 is formatted as a binary file, such as a binary base call (BCL) file, though different formats can be used in various embodiments. For instance, in some cases, the base call data file 214 is formatted as a text file.
[0077] In one or more embodiments, the base call correction system 106 uses the sequence specific error detector 210 of the real time analyzer 206 to determine whether the base calls 216 (e.g., the nucleotide read) determined during the sequencing run 204 include a sequence specific error. In particular, the base call correction system 106 can use the sequence specific error detector 210 to identify abase call sequence (e.g., a motif of nucleobases) known to cause sequence specific errors (i.e., an error-inducing sequence). Further, the base call correction system 106 can use the sequence specific error detector 210 to identify a set of base calls associated with the sequence specific error (e.g., those base calls that are potentially incorrect due to the sequence specific error). For instance, in some embodiments, the base call correction system 106 identifies those base calls that follow the error-inducing sequencing as associated with the sequence specific error.
[0078] To illustrate, in one or more embodiments, the base call correction system 106 uses the sequence specific error detector 210 to detect a homopolymer, a sequence of repeating dinucleotides, a sequence of repeating trinucleotides, or another motif of nucleobases known to causes errors within the cluster of oligonucleotides 202 based on determined base calls. Based on the identification, the base call correction system 106 determines that the base calls 216 includes a sequence specific error. Further, the base call correction system 106 can determine that those base calls that follow the detected homopolymer, repeating dinucleotide sequence, repeating trinucleotide sequence, or other motif are associated with the sequence specific error and are candidates for corrective action. In one or more embodiments, the base call correction system 106 detects the error-inducing sequence(s) and identifies the set of base calls associated with the sequence specific error in real time as part of the sequencing run 204. More detail regarding how an error inducing sequence is detected will be provided below.
[0079] In one or more embodiments, the base call correction system 106 uses the read data captor 212 of the real time analyzer to capture a set of read data 218 for the set of base calls determined to be associated with the sequence specific error. For instance, the base call correction system 106 uses the read data captor 212 to collect (e.g., extract) the set of read data 218 from the image data corresponding to the set of base calls associated with the sequence specific error or to otherwise collect the set of read data 218 from data generated from the image data. The base call correction system 106 further writes the set of read data 218 to storage 220. In one or moreAttorney Docket No. IP-2817-PCT 20 Patent Applicationembodiments, the base call correction system 106 extracts the set of read data 218 and writes the extracted data to storage 220 in real time as part of the sequencing run 204.
[0080] More specifically, the base call correction system 106 can use the read data captor 212 to collect and store the set of read data 218 determined based on the signals (e.g., the light signals) represented in the image data during a sequence run. For instance, upon detecting a sequence specific error via the sequence specific error detector 210, the base call correction system 106 can activate the read data captor 212 to begin collecting and storing the read data associated with subsequent base calls (as well as the base call of the detection cycle). To illustrate, upon determining that an identical base call has been determined for a threshold number of sequencing cycles, the base call correction system 106 can determine — via the sequence specific error detector 210 — that a homopolymer has been detected and then activate the read data captor 212 to begin collecting and storing read data for the subsequent base calls (as well as the base call determined at the detection cycle).
[0081] In one or more embodiments, storage 220 includes non-volatile storage. Further, storage 220 can include transitory storage or non-transitory storage. For instance, storage 220 can include storage within a storage device integrated into a computer device (e.g., a hard disk drive or a solid-state drive) or storage within a removable storage device (e.g., a flash drive, a memory card, or an external hard drive). Thus, in some embodiments, the base call correction system 106 facilitates use of the set of read data 218 offline (e.g., external to the real time operation during the sequencing run 204). It should be understood, however, that the base call correction system 106 can utilize volatile storage (e.g., RAM memory) in some implementations. Indeed, in some cases, the base call correction system 106 operates on or in conjunction with a computing device having sufficient memory to store the set of read data 218.
[0082] As shown in FIG. 2, the base call correction system 106 uses a correction model 222 to process the set of read data 218 written to storage. For instance, the base call correction system 106 can use a machine learning model 230 of the correction model 222 to process the set of read data 218. To illustrate, the base call correction system 106 can use the machine learning model 230 to generate base call corrections 224 for the set of base calls associated with the sequence specific error based on the set of read data 218 written to storage. As indicated, the base call correction system 106 can generate the base call corrections 224 within abase call correction file 226. In some cases, the base call correction file 226 includes a binary file, such as a .obcl file, having the base call corrections 224. Other formats, however, can be used. For instance, in some cases, the base call correction file 226 includes a text file.
[0083] In one or more embodiments, the machine learning model 230 of the correction model 222 includes a neural network. For instance, in some cases, the machine learning model 230Attorney Docket No. IP-2817-PCT 21 Patent Applicationincludes a convolutional neural network or a recurrent neural network. An architecture of the machine learning model 230 used by various implementations of the base call correction system 106 will be discussed in detail below.
[0084] As FIG. 2 further illustrates, in addition to the base call corrections 224, the base call correction system 106 can optionally use the correction model 222 to also generate quality scores 228 for the base call corrections 224. In particular, in certain embodiments, the base call correction system 106 can use a quality model 242 of the correction model 222 to generate the quality scores. Thus, in some instances, the base call correction file 226 can also include the quality scores 228 for the base call corrections 224. In one or more embodiments, the base call correction system 106 uses the quality model 242 to generate the quality scores 228 by generating Phred quality scores based on a Phred algorithm or a modified Phred algorithm developed by Illumina, Inc. For instance, in some implementations, the base call correction system 106 uses the quality model 242 to determine Phred quality scores using a Phred algorithm or a modified Phred algorithm, such as the algorithm described in U.S. Patent No. 8,392,126. As will be described in more detail below, in some implementations, the base call correction system 106 uses one or more values generated by the machine learning model 230 when determining the base call corrections 224 to generate the quality scores 228 via a quality model 242.
[0085] To illustrate, in some implementations, the base call correction system 106 formats the base call correction fde 226 as a .obcl file that includes a binary representation for each base callquality score pair represented therein. For example, the base call correction system 106 can use the first two bits for each pair to represent the base call (i.e., the base call correction) — such as [00, 01, 10, 11] for [A, C, G, T], The base call correction system 106 can use subsequent bits to represent the corresponding quality score. For instance, the base call correction system 106 can use Q number of bits to represent the quality score, where the Q bits provide an unsigned little endian integer representation.
[0086] To more clearly contrast the different files described herein, the base call data file 214 generated using the real time analyzer 206 includes raw (e.g., uncorrected) base calls (e.g., the base calls 216) and raw quality scores (e.g., the quality scores 250) in some embodiments, but the base call correction file 226 generated using the correction model 222 can include the base call corrections 224 and the quality scores 228, which include corrections to the raw quality scores. Further, in some embodiments, the base call data file 214 includes base calls and quality scores for the cluster of oligonucleotides 202 as a whole (e.g., all base calls and quality scores determined for the cluster of oligonucleotides 202 during the sequencing run 204), but the base call correction file 226 only includes base call corrections and quality scores for the set of base calls associated with the sequence specific error. In some implementations, the base call correction file 226 includes anAttorney Docket No. IP-2817-PCT 22 Patent Applicationindication (e.g., in the file header) of where the base call corrections 224 fit within the nucleotide read determined for the cluster of oligonucleotides 202. For instance, in some cases, the base call correction file 226 indicates the sequencing cycle associated with the first base call correction represented therein. It should be understood, however, that various file formats can be used for each file in various embodiments. Further, various types and amounts of data can be included in each file in various embodiments.
[0087] As further shown in FIG. 2, the base call correction system 106 uses a base call data convert module 232 to process the base call data file 214 and the base call correction file 226 and generate a corrected base call file 234 as a result. Indeed, generally, the base call correction system 106 uses the base call data convert module 232 to perform a file format transformation. As such, in some instances, the base call correction system 106 uses the base call data convert module 232 to convert data from the base call data file 214 and the base call correction file 226 into a format that is appropriate for the corrected base call file 234. In particular, as shown, the base call correction system 106 uses the base call data convert module 232 to generate the corrected base call file 234 via a post-sequencing run process 236. Additionally, as shown, the corrected base call file 234 includes corrected base calls 238 and, optionally, quality scores 240 for the corrected base calls 238.
[0088] In one or more embodiments, the base call data convert module 232 generates the corrected base calls 238 using the base calls 216 from the base call data file 214 and the base call corrections 224 from the base call correction file 226. For instance, as mentioned, the base call corrections 224 can include corrections for those base calls associated with the sequence specific error. Further, the base call correction file 226 can identify where the base call corrections 224 fit within the nucleotide read for the cluster of oligonucleotides 202. Accordingly, in some cases, the base call correction system 106 uses the base call correction file 226 to replace the (potentially) erroneous base calls included in the base calls 216 with the base call corrections 224 for inclusion within the corrected base call file 234. In other words, the base call data convert module 232 can include, within the corrected base call file 234, those base calls from the base calls 216 that are not associated with the sequence specific error and the base call corrections 224 for those base calls from the base calls 216 that are associated with the sequence specific error.
[0089] As mentioned, the corrected base call file 234 also includes the quality scores 240 for the corrected base calls 238. As with generating the corrected base calls 238, in one or more embodiments, the base call data convert module 232 generates the quality scores 240 using the quality scores 250 from the base call data file 214 and the quality scores 228 from the base call correction file 226. Indeed, in some cases, the base call correction system 106 uses the base call correction file 226 to replace the (potentially) erroneous quality scores included in the qualityAttorney Docket No. IP-2817-PCT 23 Patent Applicationscores 250 with the quality scores 228 within the corrected base call file 234. In other words, the base call data convert module 232 can include, within the corrected base call file 234, those quality scores from the quality scores 250 that are not associated with the sequence specific error and the quality scores 228 as replacements for those quality scores from the quality scores 250 that are associated with the sequence specific error.
[0090] In one or more embodiments, the corrected base call file 234 includes a text file, such as a FASTQ file, but the base call correction system 106 can use other file formats in various embodiments.
[0091] Thus, the base call correction system 106 provides a new configuration for performing real time detection and offline correction of sequence specific errors detected while sequencing a cluster of oligonucleotides. Accordingly, the base call correction system 106 operates with improved efficiency and accuracy when compared to conventional base calling systems by implementing a base calling process that accommodates sequence specific errors and corrects base calls that may be affected by the errors. Further, the base call correction system 106 provides a computationally efficient approach that facilitates the real time component of the base calling process while allowing for correction. As will be described more fully below, by separating the detection process from the correction process and taking the correction process (which can consume a significant amount of time and resources) offline, the base call correction system 106 takes advantage of sophisticated machine learning models in improving base calling accuracy after error-inducing sequences (e.g., homopolymers, sequences of repeating dinucleotides, or sequences of repeating trinucleotides) while minimizing the induced real-time cost in computation and efficiency.
[0092] As previously mentioned, the base call correction system 106 can detect, within a nucleotide read for a cluster of oligonucleotides, a base call sequence (e.g., a motif) known to cause a sequence specific error (i.e., an error-inducing sequence). For instance, the base call correction system 106 can detect a homopolymer, a sequence or repeating dinucleotides, and / or a sequence of trinucleotides within the nucleotide read. FIGS. 3A-3D illustrate the base call correction system 106 detecting a base call sequence known to cause a sequence specific error in accordance with one or more embodiments. In particular, FIG. 3 A illustrates the base call correction system 106 performing general motif detection, and FIGS. 3B-3D illustrate particular motif detection scenarios in accordance with one or more embodiments.
[0093] In one or more embodiments, the base call correction system 106 detects the base call sequence known to cause a sequence specific error as the nucleotide read is generated. In other words, the base call correction system 106 can detect the error-inducing sequence during the real time base calling process. Indeed, as mentioned above, the base call correction system 106 canAttorney Docket No. IP-2817-PCT 24 Patent Applicationperform base calling during a sequencing run by processing image data for each sequencing cycle in real time without writing the image data to storage. As further mentioned, however, the base call correction system 106 can use read data generated from the image data to generate base call corrections for those base calls associated with (e.g., potentially affected by) the sequence specific error. Accordingly, in some cases, the base call correction system 106 detects any present error inducing sequences during the real time base calling process and determines to write the set of read data for subsequent base calls to storage.
[0094] FIG. 3 A illustrates base calls 360 determined as part of a nucleotide read for a cluster of oligonucleotides in accordance with one or more embodiments. As shown in FIG. 3A, the base calls 360 of the nucleotide read indicate the presence of a motif 362 (e.g., the sequence of ‘GCTAATG’) within the cluster of oligonucleotides. While the motif 362 portrays a particular sequence of nucleobases, this sequence is merely illustrative. The base call correction system 106 can detect motifs of various nucleobase sequences in various implementations.
[0095] In one or more embodiments, the motif 362 includes a motif that has previously been determined to be an error-inducing sequence. Indeed, the base call correction system 106 can perform motif detection to determine whether an error-inducing sequence is present within a cluster of nucleotides. In particular, the base call correction system 106 can perform motif detection to determine the presence of a pattern or sequence of nucleotides known to cause errors within a nucleotide read (e.g., cause errors in base calling positions after the sequence or pattern). As mentioned, the base call correction system 106 can perform the motif detection based on the base calls that have been determined for the cluster of oligonucleotides.
[0096] As shown in FIG. 3 A, the base call correction system 106 uses a motif detector 364 to detect the motif 362 within the cluster of oligonucleotides based on the base calls 360. In some cases, the motif detector 364 includes a buffered counter (e.g., as discussed below with reference to FIGS. 3B-3D) that counts the number of consecutive repeating base calls or base call sequences within the base calls 360 and triggers detection of a corresponding motif when a threshold number of repeats has been reached. In some cases, the motif detector 364 includes a sequence comparison mechanism that compares sequences of the base calls 360 (e.g., the last n base calls) to a motif (e.g., of length M) that is known to cause sequence specific errors and triggers detection of the motif when a match has been determined. Various other detection mechanisms can be used in various other implementations.
[0097] As mentioned, the base call correction system 106 can target motifs that are known to be error-inducing sequences. For instance, in some cases, the base call correction system 106 maintains a list or database of one or more motifs known to be error-inducing sequences and configures the motif detector to detect multiple motifs from the list / database. In some cases, theAttorney Docket No. IP-2817-PCT 25 Patent Applicationbase call correction system 106 configures a particular motif detector to detect a particular motif within oligonucleotide clusters. Thus, in some implementations, the base call correction system 106 employs multiple motif detectors where each motif detector is configured to detect a different motif.
[0098] To illustrate, in some cases, as base calls are being determined, the base call correction system 106 tracks (e.g., buffer) the most recent base call (or the most recent n base calls). The base call correction system 106 can further use each motif detector to monitor the tracked base call(s). The base call correction system 106 can configure each motif detector to raise a flag upon detecting its corresponding motif. Thus, the base call correction system 106 can identify which motif has been detected at a given sequencing cycle of a sequencing run based on which motif detector has raised its flag. Indeed, by associating a motif detector with a particular motif, the base call correction system 106 can determine that the motif has been detected when the motif detector raises its flag.
[0099] FIG. 3B illustrates the base call correction system 106 using a motif detector to detect a particular motif — a homopolymer — within a cluster of oligonucleotides in accordance with one or more embodiments. FIG. 3B also illustrates base calls 302 determined as part of a nucleotide read for a cluster of oligonucleotides. As shown in FIG. 3B, the base calls 302 of the nucleotide read indicate the presence of a homopolymer 304 (e.g., a sequence of base calls for the ‘A’ nucleobase) within the cluster of oligonucleotides.
[0100] As shown, the base call correction system 106 can detect the presence of the homopolymer 304 within the nucleotide read using a length threshold 306. In one or more embodiments, the length threshold 306 indicates the length of a sequence of consecutive identical base calls that qualifies as a homopolymer. In other words, the length threshold 306 can indicate a number of consecutive identical base calls that must be identified to determine that a homopolymer is present (e.g., to determine that a sequence of base calls is part of a homopolymer). The length threshold 306 shown in FIG. 3B includes a length of ten base calls. Various lengths can be used in various embodiments, however. In some cases, the length threshold 306 is configurable via a computing device to adjust the number of base calls based on user input. For instance, in some cases, the base call correction system 106 modifies the length threshold 306 (e.g., in response to user input) to improve performance. Indeed, in some instances, the base call correction system 106 can determine that using a particular length threshold value results in better (e.g., more efficient or more accurate) base calling when compared to another length threshold value. For example, the base call correction system 106 can determine that a relatively small length threshold value leads to the incorrect detection of a sequence specific error and / or leads to an inefficient offloading of operations to the offline correction system (e.g., correcting more base calls than necessary). Thus,Attorney Docket No. IP-2817-PCT 26 Patent Applicationin some cases, the base call correction system 106 adjusts the length threshold 306 to a value that provides improved base calling and base call correction. In some cases, the base call correction system 106 configures, modifies, or otherwise determines the length threshold 306 in real time. For example, in some embodiments, the base call correction system 106 monitors the quality scores determined for nucleotide reads using a particular length threshold and modifies the length threshold for subsequent nucleotide reads to improve the quality scores.
[0101] As suggested, the base call correction system 106 can detect the presence of the homopolymer 304 by determining that the length of a sequence of identical base calls is equal to or exceeds the length threshold 306. For instance, FIG. 3B illustrates the homopolymer 304 beginning at a first read position 308 (labeled C#9, indicating that the first read position 308 corresponds to the ninth sequencing cycle of the sequencing run) and extending past a second read position 310 (labeled C#18, indicating that the second read position 310 corresponds to the eighteenth sequencing cycle of the sequence run), where the number of base calls from the first read position 308 to the second read position 310 equals the number of base calls required by the length threshold 306. FIG. 3B labels the second read position 310 as the homopolymer detection cycle, indicating that the second read position 310 corresponds to the sequencing cycle at which the base call correction system 106 determines the presence of the homopolymer 304.
[0102] As further shown in FIG. 3B, the base call correction system 106 uses a buffered counter 312 to detect the presence of the homopolymer 304 within the nucleotide read. In other words, FIG. 3B illustrates the base call correction system 106 using the buffered counter 312 as a motif detector for detecting a particular motif — a homopolymer. Indeed, as the base call correction system 106 can perform base calling in real time, the base call correction system 106 does not store base calls in memory in some embodiments. Accordingly, in some implementations, for a given sequencing cycle, the base call correction system 106 cannot access the base calls for previous sequencing cycles from memory. As such, the base call correction system 106 employs the buffered counter 312 in one or more embodiments.
[0103] In one or more embodiments, the base call correction system 106 uses the buffered counter 312 to buffer a base call from a sequencing cycle for use during the next sequencing cycle. In some instance, the base call correction system 106 uses the buffered counter 312 to buffer one base call at a time. In some embodiments, however, the base call correction system 106 uses the buffered counter 312 to buffer a greater number of base calls. Further, in some embodiments, the base call correction system 106 uses the buffered counter 312 to maintain a count. The base call correction system 106 can use the count for genomic reads (not index reads).
[0104] In one or more embodiments, the buffered counter 312 includes four components: (i) a homopolymer flag buffer that includes one bit per nucleotide read to indicate whether or not aAttorney Docket No. IP-2817-PCT 27 Patent Applicationhomopolymer has been detected; (ii) a least squares (LSQ) parameter write switch buffer that includes one bit per nucleotide read to indicate whether the LSQ parameters have already been written to storage; (iii) a base call buffer that includes two bits per nucleotide read to store base calls from previous cycles (e.g., each combination of bits corresponding to a different base call) and that is used for comparing with the base call of the current cycle; and (iv) a counting buffer that includes four bits per nucleotide read (e.g., to provide a count from zero to fifteen) used for accumulating homopolymer sequence length. As will be explained below, if the count of the counting buffer equals or exceeds a pre-determined threshold number (e.g., ten), the homopolymer flag is triggered (e.g., switching from false to true), and the base call correction system 106 starts to accumulate intensity-related stats.
[0105] To illustrate, in one or more embodiments, for a current sequencing cycle, the base call correction system 106 determines a base call. The base call correction system 106 further determines whether the base call for the current sequencing cycle is identical to the previous base call from the previous sequencing cycle. In particular, the base call correction system 106 can access the previous base call from the buffer of the buffered counter 312 and compare the previous base call to the base call for the current sequencing cycle to determine whether the two base calls are identical. Upon determining that the two base calls are identical, the base call correction system 106 can increase the count of the buffered counter 312 by one. If the two base calls are not identical, the base call correction system 106 can reset the count of the buffered counter 312. For instance, in some cases, the base call correction system 106 resets the count to a value of zero (or one), indicating that the base call for the current sequencing cycle corresponds to the start of a new sequence count. Upon determining that the count of the buffered counter 312 is equal to the value required by the length threshold 306, the base call correction system 106 can determine that a homopolymer has been detected. Thus, to illustrate detection of the homopolymer 304, in some embodiments, the base call correction system 106 initializes the count of the buffered counter 312 with a value of zero (or one) for the start of the sequencing run.
[0106] Further, the base call correction system 106 can increment or reset the count after each base call following the initial base call 314 based on how that base call compares to its preceding base call. Upon determining that the count of the buffered counter 312 reaches the value required by the length threshold 306 at the second read position 310, the base call correction system 106 can determine that the base calls from the first read position 308 to the second read position 310 indicate the presence of the homopolymer 304.
[0107] In some cases, the base call correction system 106 continues to use the buffered counter 312 even after detecting the presence of the homopolymer 304. For instance, upon detecting the presence of the homopolymer 304, the base call correction system 106 can use the buffered counterAttorney Docket No. IP-2817-PCT 28 Patent Application312 to buffer the associated read data of all cycles from then on, and the base call correction system 106 writes this read data to storage. As the base call correction system 106 has access to all the read data from the detection cycle to the end of the read, the base call correction system 106 can determine the end of the homopolymer 304 during offline correction. In particular, the base call correction system can use the stored data to identify a read position within the nucleotide read that corresponds to a final sequencing cycle for the homopolymer 304 (referred to as an “exit cycle”).
[0108] In some cases, upon detecting the presence of the homopolymer 304, the base call correction system 106 begins collecting read data for the base calls. For example, as indicated in FIG. 3B, the base call correction system 106 can collect read data for the base call at the read position corresponding to the homopolymer detection cycle (i.e., the second read position 310) and further collect read data for all subsequent base calls. As previously mentioned, the base call correction system 106 can write the collected read data to storage for use in generating base call corrections for those base calls associated with the sequence specific error caused by the presence of the homopolymer 304.
[0109] FIG. 3C illustrates the base call correction system 106 using a motif detector to detect a particular motif — a sequence of repeating dinucleotides — within a cluster of oligonucleotides in accordance with one or more embodiments. FIG. 3C illustrates base calls 322 determined as part of a nucleotide read for a cluster of oligonucleotides. As shown in FIG. 3C, the base calls 322 of the nucleotide read indicate the presence of a sequence of repeating dinucleotides 324 (e.g., a sequence of base calls that repeats the ‘AT’ pair nucleobases) within the cluster of oligonucleotides.
[0110] As shown in FIG. 3C, the base call correction system 106 can detect the presence of the sequence of repeating dinucleotides 324 within the nucleotide read using a length threshold 326. In one or more embodiments, the length threshold 326 indicates a length of a sequence of consecutive identical base call pairs that qualifies as a sequence of repeating dinucleotides known to cause sequence specific errors. For instance, the length threshold 326 can indicate a number of consecutive identical base call pairs (e.g., a number of consecutive repetitions of a pair of base calls) that must be identified to determine that a sequence of repeating dinucleotides is present (e.g., to determine that a sequence of base calls is part of a sequence of repeating dinucleotides known to cause sequence specific errors). In some embodiments, however, the length threshold 326 indicates a number of base calls (rather than base call pairs) but requires those base calls to be part of the repeated dinucleotide. FIG. 3C shows the length threshold 326 set to a length of seven dinucleotide repetitions (or fourteen base calls that are part of the repeated dinucleotide), though various lengths can be used in various embodiments. Further, in some cases, the length threshold 326 is configurable via a computing device to adjust the number of base calls based on user input.Attorney Docket No. IP-2817-PCT 29 Patent Application[OHl] As suggested, the base call correction system 106 can detect the presence of the sequence of repeating dinucleotides 324 by determining that the length of a sequence of consecutive identical base call pairs is equal to or exceeds the length threshold 326. For instance, FIG. 3C illustrates the sequence of repeating dinucleotides 324 beginning at a first read position 328 (labeled C#5, indicating that the first read position 328 corresponds to the fifth sequencing cycle of the sequencing run) and extending past a second read position 330 (labeled C#18, indicating that the second read position 330 corresponds to the eighteenth sequencing cycle of the sequence run), where the number of base call pairs (or base calls) from the first read position 328 to the second read position 330 equals the number of base call pairs (or base calls) required by the length threshold 326. FIG. 3C labels the second read position 330 as the dinucleotide repeat detection cycle, indicating that the second read position 330 corresponds to the sequencing cycle at which the base call correction system 106 determines the presence of the sequence of repeating dinucleotides 324.
[0112] As further shown in FIG. 3C, the base call correction system 106 uses a buffered counter 332 to detect the presence of the sequence of repeating dinucleotides 324 within the nucleotide read. In other words, FIG. 3C illustrates the base call correction system 106 using the buffered counter 332 as a motif detector for detecting a particular motif — a sequence or repeating dinucleotides. In one or more embodiments, the base call correction system 106 uses the buffered counter 332 to buffer a base call pair from a pair of sequencing cycles for use during the next sequencing cycle or next pair of sequencing cycles. In some instance, the base call correction system 106 uses the buffered counter 332 to buffer one base call pair at a time. In some embodiments, however, the base call correction system 106 uses the buffered counter 332 to buffer a greater number of base call pairs. As will be discussed further below, the base call correction system 106 can shift base calls into and out of the buffer with each sequencing cycle in some implementations. Further, in some embodiments, the base call correction system 106 uses the buffered counter 332 to maintain a count. The base call correction system 106 can use the count for genomic reads (not index reads).
[0113] In one or more embodiments, the buffered counter 332 includes four components: (i) a dinucleotide flag buffer that includes a bit per nucleotide read to indicate whether or not a sequence of repeating dinucleotides has been detected; (ii) a LSQ parameter write buffer that includes one bit per nucleotide read to indicate whether the LSQ parameters have already been written to storage; (iii) a base call buffer that includes four bits per nucleotide read to store base calls from the two previous cycles (e.g., two bits for each cycle where each combination of bits for each pair corresponds to a different base call) and is used for comparing with the base call of the current cycle; and (iv) a counting buffer that includes four bits per nucleotide read (e.g., to provide a countAttorney Docket No. IP-2817-PCT 30 Patent Applicationfrom zero to fifteen) used for accumulating homopolymer sequence length. Given base calls at cycle #(C-2), cycle #(C-1), and cycle #C (the current cycle), the buffered counter 332 will increase the count if the following conditions are met: (i) the base call at #(C-2) equals the base call at cycle #C and (ii) the base call at #(C-1) does not equal the base call at #C. If both conditions are not met, the buffered counter 332 can reset the count. As will be explained below, if the count of the counting buffer equals or exceeds a pre-determined threshold number (e.g., fourteen), the dinucleotide flag is triggered (e.g., switching from false to true), and the base call correction system 106 starts to accumulate intensity-related stats.
[0114] To illustrate at least one example of using the buffered counter 332, in one or more embodiments, for a current sequencing cycle, the base call correction system 106 determines abase call. The base call correction system 106 further determines whether the base call for the current sequencing cycle is identical to a previous base call made two sequencing cycles prior to the current sequencing cycle. Additionally, the base call correction system 106 determines whether the base call for the current sequencing cycle is different from a previous base call made one sequencing cycle prior to the current sequencing cycle. In particular, the base call correction system 106 can access the base call pair from the buffer of the buffered counter 332 and compare the previous base call from two sequencing cycles prior (e.g., the base call in the first position of the buffer) to the base call for the current sequencing cycle to determine whether the base calls are identical. Upon determining that the two base calls are identical, the base call correction system 106 can increase the count of the buffered counter 332 by one. If the two base calls are not identical, the base call correction system 106 can reset the count of the buffered counter 332. For instance, in some cases, the base call correction system 106 resets the count to a value of zero (or one), indicating that the base call pair including the base call for the current sequencing cycle corresponds to the start of a new sequence count.
[0115] As mentioned, in one or more embodiments, with each sequencing cycle, the base call correction system 106 shifts one base call out of the buffer and shifts in the base call determined for the current sequencing cycle. Thus, the base calls that are buffered can change with each sequencing cycle. This allows the base call correction system 106 to identify the beginning of a repeating dinucleotide regardless of its read position within the nucleotide read. For instance, the first base call of a repeating dinucleotide may appear within the second read position of the nucleotide read (e.g., based on the second sequencing cycle of the sequencing run). Shifting base calls into and out of the buffer in pairs risks failing to appropriately detect that base call as the beginning of the dinucleotide, because the pair may be discarded upon determining that the base call pair for the third and fourth sequencing cycles does not match the base call pair for the first and second sequencing cycles. By shifting base calls within the buffer for each sequencing cycle,Attorney Docket No. IP-2817-PCT 31 Patent Applicationhowever, the base call correction system 106 only needs to determine whether the base call for the current sequencing cycle matches the previous base call from two sequencing cycles prior — which would be the base call in the first position of the buffer (i.e., the next base call to be shifted out) — and determine that it does not match the previous base call from the immediately preceding sequencing cycle. During the next sequencing cycle, the base call correction system 106 shifts the base call that was the second base call to the first position in the buffer and can compare the base call for the next sequencing cycle with that base call in the buffer.
[0116] In one or more embodiments, when shifting base calls into and out of the buffer with each sequencing cycle, the base call correction system 106 uses two counters in the buffered counter 332. For instance, the base call correction system 106 can use a raw counter that increases based on determining that a base call for a current sequencing cycle is identical to a previous base call determined two sequencing cycles prior. Further, the base call correction system 106 can use a pair counter that increases with each dinucleotide that is counted. For example, the base call correction system 106 can increase the pair counter by a value of one for every two-count of the raw counter (where the base call correction system 106 resets the raw counter after reaching a value of two). Accordingly, in some embodiments, the base call correction system 106 resets both counters upon determining that a base call for a current sequencing cycle is not identical to a previous base call determined two sequencing cycles prior. In some cases, rather than using two counters, the base call correction system 106 uses a single counter but the length threshold 326 includes an even number, which corresponds to a number of dinucleotide (e.g., base call pair) repetitions.
[0117] In some cases, the base call correction system 106 continues to use the buffered counter 332 even after detecting the presence of the sequence of repeating dinucleotides 324. For instance, upon detecting the presence of the sequence of repeating dinucleotides 324, the base call correction system 106 can use the buffered counter 332 to buffer the associated read data of all cycles from then on, and the base call correction system 106 writes this read data to storage. As the base call correction system 106 has access to all the read data from the detection cycle to the end of the read, the base call correction system 106 can determine the end of the sequence of repeating dinucleotides 324 (i.e., the exit cycle) during offline correction.
[0118] In some cases, upon detecting the presence of the sequence of repeating dinucleotides 324, the base call correction system 106 begins collecting read data for the base calls. For example, as indicated in FIG. 3C, the base call correction system 106 can collect read data for the base call at the read position corresponding to the dinucleotide repeat detection cycle (i.e., the second read position 330) and further collect read data for all subsequent base calls. In some cases, the base call correction system 106 collects the read data for all base calls corresponding to sequencing cyclesAttorney Docket No. IP-2817-PCT 32 Patent Applicationfollowing the detection cycle. As previously mentioned, the base call correction system 106 can write the collected read data to storage for use in generating base call corrections for those base calls associated with the sequence specific error caused by the presence of the sequence of repeating dinucleotides 324.
[0119] In one or more embodiments, the base call correction system 106 operates similarly to detect sequences of trinucleotides within nucleotide reads. For instance, the base call correction system 106 can utilize a length threshold that indicates a length of a sequence of consecutive identical base call triplets that qualifies as a sequence of repeating trinucleotides. Further, the base call correction system 106 can utilize a buffered counter to buffer base calls and count base calls toward the length threshold.
[0120] Thus, in one or more embodiments, when analyzing a nucleotide read, the base call correction system 106 can employ multiple buffered counters or other motif detectors. In some cases, the base call correction system 106 uses the buffered counters or other motif detectors in parallel. For instance, the base call correction system 106 can employ a first buffered counter to detect homopolymers and a second buffered counter (e.g., in parallel) to detect sequences of repeating dinucleotides. As previously mentioned, the buffered counter used to detect each errorinducing sequence can differ in their operation. For instance, the buffered counter used to detect repeating sequences of dinucleotides can buffer two prior sequencing cycles and increase its count based on determining that the base call of the current sequencing cycle (i) matches the base call from two sequencing cycles prior and (ii) does not match the base call from the immediately preceding sequencing cycle. Indeed, the base call correction system 106 can implement this difference to enable determining a distinction between repeating sequences of dinucleotides from detecting homopolymers. To illustrate, a homopolymer can be considered a special case of a sequence of repeating dinucleotides where both nucleobases within the dinucleotide are the same. As such, the base call correction system 106 can determine whether the base call for the current sequencing cycle does not match the base call for the immediately preceding sequencing cycle to ensure that a sequence of repeating dinucleotides, but not a homopolymer, is being counted.
[0121] FIG. 3D also illustrates base calls 342 determined as part of a nucleotide read for a cluster of oligonucleotides in accordance with one or more embodiments. As shown in FIG. 3D, the base calls 342 include non-calls (e.g., base calls labeled ‘N’) at several read positions.
[0122] As used herein, the term “non-call” refers to an undetermined base call. In particular, a non-call can include a determination for a sequencing cycle that indicates that the nucleobase is uncertain or that the base call correction system 106 was otherwise unable to definitively determine a base call. In some cases, the base call correction system 106 determines a non-call for a sequencing cycle due to the image data captured during the sequencing cycle. For instance, in someAttorney Docket No. IP-2817-PCT 33 Patent Applicationcases, the signal captured by the image data is too weak (e.g., below a threshold in signal strength) to make a definitive base call or provides signals that conflict enough to make the base call correction system 106 sufficiently uncertain as to the correct base call (e.g., the signals indicate multiple equally likely base calls). In some instances, the base call correction system 106 determines a non-call due to a malfunction in the hardware. For instance, the base call correction system 106 can determine that the imaging hardware malfunctioned during a particular sequencing cycle or set of sequencing cycles and was not able to capture images.
[0123] FIG. 3D illustrates that the base calls 342 include a base call 344 (at the read position labeled C#7) that is potentially the beginning of a homopolymer. Indeed, several of the base calls that immediately follow the base call 344 are identical to the base call 344. Accordingly, the base call correction system 106 modifies the count of the buffered counter 350 during those sequencing cycles to represent that there are multiple consecutive identical base calls.
[0124] As further shown in FIG. 3D, however, the base calls 342 also include a series of noncalls that begins with the non-call 346 (at the read position labeled C#11) — determined only a few sequencing cycles after the sequencing cycle in which the base call 344 is determined — and ends with the non-call 348 (at the read position labeled C#18). In one or more embodiments, upon determining the non-call 346, the base call correction system 106 resets the count of the buffered counter 350. Further, the base call correction system 106 can reset the count of the buffered counter 350 for each non-call up to the non-call 348 that ends the series of non-calls.
[0125] Thus, in one or more embodiments, the base call correction system 106 resets the count of the buffered counter 350 upon determining a non-call during the sequencing run. Though FIG. 3D illustrates a series of consecutive non-calls, non-calls may be more sporadically dispersed throughout a nucleotide read in various instances. In such scenarios, the base call correction system 106 similarly resets the count of the buffered counter 350 for each non-call. In other words, the base call correction system 106 can operate the same regardless of how non-calls are determined throughout the sequencing run — reset the count upon determination of a non-call. If a nucleotide read has been identified as including an error-inducing sequence, the base call correction system 106 can reset the count upon encountering a new non-call after detecting the error-inducing sequence but still collect the read data (e.g., and other related stats) associated with those non-calls. In one or more embodiments, rather than resetting the count when determining a non-call, the base call correction system 106 takes no action with respect to the count (neither resets the count nor increases the count). Thus, if base calls following a non-call sufficiently continue the pattern of base calls before the non-call (e.g., the number of base calls after the non-call that follow the pattern with the number of base calls before the non-call that follow the pattern satisfies the relevant lengthAttorney Docket No. IP-2817-PCT 34 Patent Applicationthreshold), the base call correction system 106 can use the continued pattern to determine whether an error-inducing sequence is present.
[0126] As previously mentioned, in one or more embodiments, the base call correction system 106 uses a correction model to perform corrections with respect to base calls determined to be associated with a sequence specific error. For example, the base call correction system 106 can use a machine learning model of the correction model to generate base call corrections for the base calls associated with the sequence specific error. Further, the base call correction system 106 can use a quality model of the correction model to determine quality scores for the base call corrections. FIG. 4 illustrates the base call correction system 106 determining base call corrections and corresponding quality scores for base calls associated with a sequence specific error in accordance with one or more embodiments.
[0127] As discussed above with reference to FIGS. 3A-3D, the base call correction system 106 collects and stores read data for base calls associated with a sequence specific error. In particular, in some cases, upon detecting an error-inducing sequence, the base call correction system 106 collects and stores read data for the sequencing cycle at which the error-inducing sequence was detected (i.e., the detection cycle) and / or sequencing cycles that follow (e.g., until the final sequencing cycle for the nucleotide read). Thus, in one or more embodiments, the base call correction system 106 generates base call corrections for those base calls determined at the detection cycle and / or the following sequencing cycles. In some embodiments, the base call correction system 106 generates base call corrections for those base calls determined after completion of the error-inducing sequence (after the exit cycle), which can occur after the detection cycle. It should be understood, however, that the base call correction system 106 can determine base call corrections for those base calls associated with the set read data that has been written to storage (whether determining base call corrections for all base calls having associated read data written to storage or only a subset of the base calls having associated read data written to storage).
[0128] Indeed, as shown in FIG. 4, the base call correction system 106 provides read data 402 as input to a machine learning model 400. In one or more embodiments, the read data 402 includes the set of read data written to storage for those base calls determined to be associated with a sequence specific error. For instance, the read data 402 can include a set of read data for base calls determined for the range of sequencing cycles (i) from the detection cycle to the end of the nucleotide read, (ii) from the exit cycle to the end of the nucleotide read, or (iii) from a sequencing cycle between the detection and exit cycles to the end of the nucleotide read.
[0129] As shown in FIG. 4, the read data 402 includes one or more of raw intensity data 404, corrected intensity data 406, correction parameters 408, base calls 410, or base call distribution clouds 412. In some cases, the read data 402 includes one or more of the above types of read dataAttorney Docket No. IP-2817-PCT 35 Patent Applicationfor each base call associated with the sequence specific error. For instance, the read data 402 can include a first set of raw intensity data for a first base call, a second set of raw intensity data for a second base call, and so forth. Further, various embodiments of the base call correction system 106 use fewer types of read data, additional types of read data, and / or alternative types of read data.
[0130] In one or more embodiments, the correction parameters 408 include those parameters determined and used for generating the corrected intensity data 406 from the raw intensity data 404. For instance, in some embodiments, the base call correction system 106 uses a model to generate the corrected intensity data 406. To use the model, the base call correction system 106 determines and implements the correction parameters 408. For instance, in some cases, the base call correction system 106 uses a first set of correction parameters to generate a first set of corrected intensity data for a first base call, a second set of correction parameters to generate a second set of corrected intensity data for a second base call, and so forth. In some instances, the base call correction system 106 uses at least a set of shared correction parameters to generate corrected intensity data for each of the base calls associated with the sequence specific error. For instance, in some cases, correction parameters include phasing coefficients (e.g., global phasing coefficients and / or per-cluster phasing coefficients) determined and used for generating the corrected intensity data 406.
[0131] In some cases, the base call distribution clouds 412 include base call distribution clouds used in determining the base calls during the real-time base calling process. For instance, in some cases, the base call correction system 106 plots raw intensity data (or corresponding corrected intensity data) determined during a sequencing cycle and determines a base call for the sequencing cycle based on its position within the plot. In particular, the base call correction system 106 can determine a base call for a sequencing cycle based on the intensity data being positioned near or within a particular base call distribution cloud. Indeed, in some instances, the base call correction system 106 uses a cloud for each potential nucleobase. Thus, the plotting of intensity data for a sequencing cycle can indicate the base call based on its position near or within the corresponding cloud.
[0132] As shown in FIG. 4, the base call correction system 106 uses the machine learning model 400 to generate a base call correction 414 for a base call associated with a sequence specific error. In particular, the base call correction system 106 uses the machine learning model 400 to analyze, from the read data 402, the read data that is associated with the base call and generate the base call correction 414 based on the analysis.
[0133] As illustrated, the machine learning model 400 includes convolutional layers 416a - 416b. In particular, FIG. 4 illustrates each of the convolutional layers 416a - 416b including a onedimensional convolutional layer, but convolutional layers of other dimensions (e.g., twoAttorney Docket No. IP-2817-PCT 36 Patent Applicationdimensions) can be used in various embodiments. In some cases, the machine learning model 400 includes convolutional layers of mixed dimensions. Additionally, while FIG. 4 illustrates a particular number of convolutional layers, the machine learning model 400 can use various numbers of convolutional layers in various implementations. As indicated, the machine learning model 400 can use the convolutional layers 416a - 416b to analyze a local context 418 of the base call based on the corresponding read data. In some cases, the convolutional layers 416a - 416b analyze the local context 418 by analyzing the base call in the context of its local position within the nucleotide read. For instance, in some embodiments, the convolutional layers 416a - 416b analyzes those portions of the read data 402 that are indicative of the base call itself (e.g., the raw intensity data 404, the base call itself from the base calls 410 for the nucleotide read, or the base call distribution clouds 412).
[0134] As further illustrated, the machine learning model 400 includes an attention mechanism 420 that follows the convolutional layers 416a - 416b. In one or more embodiments, the attention mechanism 420 includes a self-attention mechanism, such as a multi-head attention mechanism, but other attention mechanisms can be used in other implementations. While FIG. 4 illustrates a single attention mechanism, the machine learning model 400 can use additional attention mechanisms in some embodiments. As indicated, the machine learning model 400 can use the attention mechanism 420 to analyze a global context 422 of the base call based on the corresponding read data. In some cases, the attention mechanism 420 analyzes the global context 422 by analyzing the base call in the context of other portions of the nucleotide read. For instance, in some embodiments, the attention mechanism 420 analyzes the read data 402 for the base call in the context of read data for other base calls (e.g., preceding base calls) in the nucleotide read or by analyzing portions of the read data 402 that are indicative of other portions of the nucleotide read (e.g., the corrected intensity data 406, global coefficients from the correction parameters 408, or other base calls from the base calls 410 for the nucleotide read).
[0135] As further illustrated, the machine learning model 400 includes one or more dense layers 424 following the attention mechanism 420. In some cases, each of the one or more dense layers 424 includes a fully connected layer. The machine learning model 400 can use various numbers of dense layers in various implementations. As indicated, the machine learning model 400 can use the one or more dense layers 424 to make classification determinations 426 for the base call.
[0136] As shown in FIG. 4, the machine learning model 400 uses a softmax layer 428 to generate a set of probabilities 430 for the base call. In particular, the machine learning model 400 can use the softmax layer 428 to generate probabilities that each indicate a likelihood that the base call for which the correction is being determined is associated with a corresponding nucleobase orAttorney Docket No. IP-2817-PCT 37 Patent Applicationa non-call. In other words, the machine learning model 400 uses the softmax layer 428 to generate probabilities that each indicate a likelihood of the appropriate base call correction. While FIG. 4 shows the set of probabilities 430 including a probability for a non-call, the softmax layer 428 generates probabilities for the four potential nucleobases without generating a probability for a non- call in some cases. In some cases, the machine learning model 400 uses the softmax layer 428 so the set of probabilities 430 sum to a value of one.
[0137] To illustrate, in some cases, the machine learning model 400 uses the convolutional layers 416a - 416b, the attention mechanism 420, and / or the one or more dense layers 424 to generate a set of convolved features for the base call based on the corresponding read data. As used herein, the term “convolved feature” includes a feature that is generated or extracted from a neural network layer. In particular, in some cases, a convolved feature includes a latent feature generated or extracted by a neural network layer based on input to the neural network layer. For instance, in some cases, a convolved feature includes a feature generated or extracted by a convolutional layer based on input to the convolutional layer. Indeed, in some cases, the machine learning model 400 uses the convolutional layers 416a - 416b to generate a set of convolved features for a base call from the corresponding read data. The machine learning model 400 further uses the subsequent layers to process the set of convolved features to generate the set of probabilities 430. For instance, in some cases, the machine learning model 400 uses the softmax layer 428 to generate the set of probabilities 430 from the set of convolved features (or other features generated from the convolved features).
[0138] As illustrated, the base call correction system 106 determines the base call correction 414 based on the set of probabilities 430. For instance, in some cases, the base call correction system 106 determines the base call correction 414 based on a highest probability from the set of probabilities 430. In particular, the base call correction system 106 can determine the base call correction 414 to include the nucleobase associated with the highest probability from the set of probabilities 430.
[0139] As FIG. 4 shows, in addition to using the machine learning model 400 to generate the base call correction 414, the base call correction system 106 can also optionally use a quality model 432 to generate a quality score 434 for the base call correction 414. For instance, in some cases, the base call correction system 106 uses the quality model 432 to generate the quality score 434 based on the highest probability from the set of probabilities 430. To illustrate, in some embodiments, the quality model 432 generates the quality score 434 as — 101og10(l — p(x)) where p(x) represents the highest probability from the set of probabilities 430.
[0140] In one or more embodiments, the base call correction system 106 uses a machine learning model as the quality model 432 to generate the quality score 434. For example, in someAttorney Docket No. IP-2817-PCT 38 Patent Applicationcases, the base call correction system 106 uses an XGBoost model as the quality model 432. In some cases, the base call correction system 106 uses the machine learning model to generate the quality score 434 based on the highest probability from the set of probabilities 430. In certain cases, the base call correction system 106 uses the machine learning model to generate the quality score 434 from one or more internal (e.g., intermediate) values generated by the machine learning model 400 (e.g., one or more of the convolved features). In some instances, the base call correction system 106 uses the machine learning model to generate the quality score 434 directly from the read data 402.
[0141] While FIG. 4 illustrates a particular architecture of the machine learning model 400, the base call correction system 106 can use various architectures for the machine learning model 400 in various embodiments. For instance, in some cases, the base call correction system 106 configures the machine learning model 400 to have the one or more dense layers 424 directly following the convolutional layers 416a - 416b (i.e., the base call correction system 106 omits the attention mechanism 420 from the machine learning model). In certain cases, the base call correction system 106 uses a recurrent neural network rather than a convolution -based neural network.
[0142] As previously suggested, in some embodiments, the machine learning model 400 generates the set of probabilities 430 to include five probabilities — four probabilities corresponding to the four potential nucleobases and one probability corresponding to a non-call. As further suggested, however, in certain cases, the machine learning model 400 generates the set of probabilities 430 to include four probabilities — one corresponding to each potential nucleobase. In other words, in some instances, the machine learning model 400 omits generating a probability indicating a likelihood of a non-call.
[0143] To illustrate, in certain embodiments, the base call correction system 106 trains the machine learning model 400 using training data that includes a set of ground truth base calls. During training, the base call correction system 106 can compare a predicted base call correction generated by the machine learning model 400 to a corresponding ground truth base call (e.g., via a loss function), determine an error of the machine learning model 400 based on the comparison, and update the parameters of the machine learning model 400 based on the error. In some cases, the set of ground truth base calls includes one or more non-calls. In other words, the training data can include a set of training input and a set of ground truth non-calls corresponding to the set of training input.
[0144] To train the machine learning model 400 to generate probabilities corresponding to the potential nucleobases without generating a probability corresponding to a non-call, the base call correction system 106 can use a mask (e.g., an N-mask) to filter out the training data thatAttorney Docket No. IP-2817-PCT 39 Patent Applicationcorresponds to non-calls. For instance, in some cases, the base call correction system 106 uses the mask to filter out the set of training input that corresponds to non-calls so the machine learning model 400 does not see that data during the training process. In some embodiments, the base call correction system 106 uses the mask to filter out the outputs generated by the machine learning model 400 from the input data corresponding to non-calls to avoid adjusting the model parameters based on non-calls. When filtering out the model outputs, the base call correction system 106 can also use the mask to filter out the corresponding ground truth base calls (i.e., filter out the non- calls).
[0145] By using an offline machine learning model to generate base call corrections for base calls associated with a sequence specific error, the base call correction system 106 can operate more flexibly and accurately when compared to many existing systems. For instance, by determining the base call corrections offline, the base call correction system 106 enables the correction to be performed in an environment that is not computationally dominated by the real-time base calling process. This frees up a significant amount of computing resources, enabling flexibility in the approach that is used to perform the corrections. With more computing power available, the base call correction system 106 can implement more computationally demanding models — such as machine learning models — to determine the corrections based on a more sophisticated analysis than was available under conventional systems. As such, the base call correction system 106 can generate base call corrections that more accurately remedy the errors of the real-time base calling process, leading to a more accurate base calling result.
[0146] As mentioned above, the base call correction system 106 can use raw intensity data and / or corrected intensity data in determining base call corrections for base calls associated with a sequence specific error. In one or more embodiments, the base call correction system 106 uses a model to generate corrected intensity data from raw intensity data. For instance, the base call correction system 106 can use a phasing correction model that performs per-cluster phasing correction. FIG. 5 illustrates a phasing correction model used by the base call correction system 106 to generate corrected intensity data 522 from raw intensity data 520 by performing per-cluster phasing correction in accordance with one or more embodiments. In particular, FIG. 5 shows the base call correction system 106 performing global phasing correction 510 and per-cluster phasing correction 512.
[0147] In one or more embodiments, the base call correction system 106 performs the global phasing correction 510 and the per-cluster phasing correction 512 during real-time base calling. For instance, in some cases, the base call correction system 106 generates the corrected intensity data 522 during real-time base calling and uses the corrected intensity data 522 to determine the base calls as part of the real time process. In some implementations, however, the base callAttorney Docket No. IP-2817-PCT 40 Patent Applicationcorrection system 106 performs the global phasing correction 510 and / or the per-cluster phasing correction 512 offline (e.g., outside the real-time process). Thus, in some instances, the base call correction system 106 performs real-time base calling using the raw intensity data 520 or partially corrected intensity data and determines base call corrections offline using the corrected intensity data 522.
[0148] As used herein, the term “per-cluster phasing correction” refers to a process or function that, when applied, adjusts a signal from labeled nucleotides bases within a particular cluster of oligonucleotides to correct for estimated phasing or pre-phasing. In particular, a per-cluster phasing correction can include an algorithm or function by which a signal from a cluster should be adjusted to correct for the estimated effects of estimated phasing or pre-phasing using a Fourier transform.
[0149] As used herein, the term “phasing” refers to an instance of (or rate at which) labeled nucleotide bases are incorporated behind a particular sequencing cycle. Phasing includes an instance of (or rate at which) labeled nucleotide bases within a cluster are asynchronously incorporated behind other labeled nucleotide bases within a cluster for a particular sequencing cycle. In particular, during SBS, each DNA strand in a cluster extends incorporation by one nucleotide base per cycle. One or more oligonucleotide strands within the cluster may become out of phase with the current cycle. Phasing occurs when nucleotide bases for one or more oligonucleotides within a cluster fall behind one or more cycles of incorporation. For example, a nucleotide sequence from a first location to a third location may be CT A. In this example, the C nucleobase should be incorporated in a first cycle, T in the second cycle, and A in the third cycle. When phasing occurs during the second sequencing cycle, one or more labeled C nucleobases are incorporated instead of a T nucleobase. Relatedly, as used herein, the term “pre-phasing” refers to an instance of (or rate at which) one or more nucleotide bases are incorporated ahead of a particular cycle. Pre-phasing includes an instance of (or rate at which) labeled nucleotide bases within a cluster are asynchronously incorporated ahead other labeled nucleotide bases within a cluster for a particular sequencing cycle. To illustrate, when pre-phasing occurs during the second sequencing cycle in the example above, one or more labeled A nucleobases are incorporated instead of a T nucleobase.
[0150] In one or more embodiments, the base call correction system 106 performs the per- cluster phasing correction 512 using one or more per-cluster phasing or pre-phasing coefficients. As used herein, the term “per-cluster phasing coefficient” refers to a factor or value that estimates or measures per-cluster phasing on a signal for a cluster. In particular, a per-cluster phasing coefficient estimates the effects of phasing for a cluster within a given sequencing cycle. For example, a per-cluster phasing coefficient can indicate the effect a nucleotide base for a previous cycle has on a signal from labeled nucleotide bases for a current cycle. To illustrate, in the exampleAttorney Docket No. IP-2817-PCT 41 Patent Applicationdescribed above, a per-cluster phasing coefficient can estimate the effect of phasing from the C nucleobase that is incorporated instead of a T nucleobase during the second sequencing cycle.
[0151] Relatedly, the term “per-cluster pre-phasing coefficient” refers to a factor or value that estimates or measures per-cluster pre-phasing on a signal for a cluster. In particular, a per-cluster pre-phasing coefficient estimates the effects of pre-phasing for a cluster within a given sequencing cycle. For example, a per-cluster pre-phasing coefficient can indicate the effect a nucleotide base for a subsequent cycle has on a signal from labeled nucleotide bases for a current cycle. To illustrate, in the example described above, a per-cluster pre-phasing coefficient estimates the effect of pre-phasing from the A nucleobase that is incorporated instead of a T nucleobase during the second sequencing cycle.
[0152] As mentioned above, the base call correction system 106 can determine a per-cluster phasing correction to correct a signal from a cluster of oligonucleotides (e.g., to correct the raw intensity data determined from the signal). Further, the base call correction system 106 can determine global phasing corrections to correct the signal from the cluster and signals from a set of clusters. As further mentioned, the base call correction system 106 can determine coefficients to perform the corrections. FIG. 5 illustrates a per-cluster phasing coefficient operation 502 and global phasing coefficient operation 504 modeled as two convolution operations implemented by the base call correction system 106 in series. It should be understood, however, that the order of the per- cluster phasing coefficient operation 502 and the global phasing coefficient operation 504 can vary in various embodiments.
[0153] Indeed, in one or more embodiments, the base call correction system 106 performs the global phasing corrections using the real time analyzer 206 discussed above with reference to FIG. 2. Thus, the base call correction system 106 can perform the global phasing coefficient operation 504 within the real time analyzer 206. In some cases, the base call correction system 106 performs the per-cluster phasing correction 512 offline (e.g., outside the real time process). As such, the base call correction system 106 can perform the per-cluster phasing coefficient operation 502 offline as well. In some implementations, however, the base call correction system 106 performs the per- cluster phasing coefficient operation 502 via the real time analyzer 206. In one or more embodiments, the base call correction system 106 determines the coefficients as described in U.S. Pat. App. No. 18 / 059,326 file on November 28, 2022, entitled Generating Cluster-Specific-Signal Corrections for Determining Nucleotide-Base Calls, which is incorporated herein by reference in its entirety.
[0154] In particular, FIG. 5 illustrates the base call correction system 106 using a phasing model 500 for estimating various coefficients as part of generating the per-cluster phasing correction 512 and the global phasing correction 510. The phasing model 500 includes operationsAttorney Docket No. IP-2817-PCT 42 Patent Applicationoccurring on a sequencing device 506 or other sequencing machine as well as operations occurring during signal processing 508 on a sequencing device or one or more servers. For example, in some embodiments, the base call correction system 106 performs the per-cl uster phasing coefficient operation 502 to estimate per-cluster phasing coefficients and the global phasing coefficient operation 504 to estimate global phasing coefficients. The base call correction system 106 can further utilize the per-cluster phasing coefficients and the global phasing coefficients as part of the signal processing 508. More specifically, the base call correction system 106 performs the global phasing correction 510 to adjust a signal based on the global phasing coefficients. Furthermore, the base call correction system 106 performs the per-cluster phasing correction 512 to adjust the signal based on per-cluster phasing coefficients. Accordingly, in some cases, the corrected intensity data 522 includes the adjusted signal and / or data extracted from the adjusted signal (e.g., adjusted intensity values).
[0155] The phasing model 500 can comprise a real-time (or near real-time) computing architecture or an offline computing architecture (e.g., a computing architecture external to the realtime environment). Generally, by utilizing a real-time computing architecture, the base call correction system 106 performs one or more of the operations illustrated in FIG. 5 utilizing a processor of the sequencing device 506 (e.g., the sequencing device 114). In contrast, the base call correction system 106 may also employ an offline computing architecture that involves both a sequencing machine and one or more servers (e.g., the server device(s) 102) or involves a GPU or FPGA. In one example, the base call correction system 106 performs the signal processing 508 at one or more server devices while performing the per-cluster phasing coefficient operation 502 and the global phasing coefficient operation 504 at the sequencing device 506. More specifically, the base call correction system 106 can perform (i) the global phasing correction 510 and (ii) the per- cluster phasing correction 512 at the processor of a server device. As further described, in some cases, the base call correction system 106 perform the global phasing correction 510 as part of the real-time base calling process and performs the per-cluster phasing correction 512 as part of the offline correction process.
[0156] Generally, and as previously described, phasing and pre-phasing refer to phenomenon where a fraction of oligonucleotides in a cluster shift forward or backward by incorporating nucleotide bases corresponding to one or more previous or subsequent cycles, respectively. The base call correction system 106 can produce a corrected signal (the output signal y) based on a convolution of a signal for a cluster (input signal x) and per-cluster phasing coefficients (input coefficients h). In some embodiments, a per-cluster phasing coefficient (h) includes both a per- cluster pre-phasing coefficient and a per-cluster phasing coefficient. The corrected signal can be modeled as a convolution operation yc= i h-ixc- which is written as y = x * h. Assuming noAttorney Docket No. IP-2817-PCT 43 Patent Applicationsignal decay, the per-cluster phasing coefficient h is constrained by Xi hi = , hi> 0. In signal processing and communication systems literature, it is common to use D-transform notation, where Dkindicates a delay of k cycles: / t(D) = — I-+ h04- h^D + h2D2+ ... . As written, h_2D~2+ h_ D~ represents phasing coefficients corresponding to nucleotide bases two and one cycles previous to the current cycle. h^D + h2D2represents pre-phasing coefficients corresponding to nucleotide bases one and two cycles following the current cycle.
[0157] As illustrated in FIG. 5, the base call correction system 106 performs the per-cluster phasing coefficient operation 502 to determine the per-cluster phasing coefficient and / or the per- cluster pre-phasing coefficient for each cluster with read positions following an error-inducing sequence. To illustrate, the base call correction system 106 determines various per-cluster phasing coefficients (h) corresponding to a previous cycle ( _1), a current cycle (0), and a subsequent cycle ( i). The per-cluster phasing coefficients vary independently across clusters and may not be determined for some clusters (e.g., at read positions preceding or within an error-inducing sequence). Most clusters unaffected by estimated phasing or pre-phasing have the values h =
[0010] , However, the base call correction system 106 can determine that the per-cluster phasing coefficients change randomly and abruptly after error-inducing sequences, such as homopolymers. In some embodiments, the per-cluster phasing coefficients sum to unity and are non-negative as represented by the function Xi -i (c) = T ht> 0.
[0158] As further illustrated in FIG. 5, the base call correction system 106 performs the global phasing coefficient operation 504 to determine a global phasing coefficient. The base call correction system 106 can utilize the global phasing coefficient across clusters in a particular section of a nucleotide-sample slide (e.g., tile of a flow cell). The global phasing coefficient values can change gradually from cycle to cycle. These values are simpler to estimate accurately than per- cluster phasing coefficients, because the statistics can be averaged across millions of clusters.
[0159] As shown in FIG. 5, for example, the base call correction system 106 calculates various global phasing coefficients (g) corresponding to a previous cycle (c / -i ), a current cycle (g0), and a subsequent cycle (g . As with the per-cluster phasing coefficients, the global phasing coefficients (g) sum to unity and are non-negative as represented by the function i i (c) = 1,510- As illustrated in FIG. 5, the base call correction system 106 adjusts the signal based on both the per- cluster phasing correction (including per-cluster phasing coefficient) and the global phasing correction (including the global phasing coefficient).
[0160] In some embodiments, the base call correction system 106 applies both the per-cluster phasing coefficient operation 502 and the global phasing coefficient operation 504 to a cluster. Additionally, or alternatively, the base call correction system 106 applies the global phasing coefficient operation 504 but not the per-cluster phasing coefficient operation 502 to some clusters.Attorney Docket No. IP-2817-PCT 44 Patent ApplicationIn particular, in some embodiments, the base call correction system 106 adjusts signals from one or more clusters based on a global phasing correction without a per-cluster phasing correction. For example, as mentioned previously, signals for nucleotide bases preceding an error-inducing sequence may not require per-cluster phasing corrections as the signals have not been affected by the error-inducing sequence. Accordingly, in some embodiments, the base call correction system 106 identifies, for an additional cluster of oligonucleotides, a different read position preceding the error-inducing sequence within a different nucleotide-fragment read. The base call correction system 106 further detects an additional signal from labeled nucleotide bases within the additional cluster of oligonucleotides during a cycle corresponding to the different read position. The base call correction system 106 then adjusts the additional signal based on a global phasing correction without a per-cluster phasing correction for the additional cluster of oligonucleotides.
[0161] In yet other embodiments, the base call correction system 106 applies the per-cluster phasing coefficient operation 502 to a signal for a given cluster without performing the global phasing coefficient operation 504. For example, in some cases, the base call correction system 106 applies a per-cluster phasing coefficient and a per-cluster pre-phasing coefficient (or other parameters) for a given cluster to a signal for the given cluster without applying parameters resulting from global coefficient operations. Accordingly, when processing clusters within a nucleotide-sample slide, the base call correction system 106 can apply a per-cluster phasing correction (without a global phasing correction) to a signal for a given cluster but apply a per- cluster phasing correction and a global phasing correction to a signal for a different cluster.
[0162] As previously mentioned, the base call correction system 106 adjusts the signal based on per-cluster phasing coefficients and global phasing coefficients as part of the signal processing 508. In particular, and as illustrated in FIG. 5, the base call correction system 106 performs the global phasing correction 510 as part of the signal processing 508. The base call correction system 106 utilizes global phasing coefficients generated from the global phasing coefficient operation 504 together with an algorithm (such as a finite impulse response (FIR) algorithm) to perform the global phasing correction 510. For example, the base call correction system 106 adjusts a signal based on corrections (y) corresponding to a previous cycle (y_i), a current cycle (y0), and a subsequent cycle Yi)-
[0163] As further illustrated in FIG. 5, the base call correction system 106 performs per-cluster phasing correction 512 as part of the signal processing 508. In particular, as part of the per-cluster phasing correction 512, the base call correction system 106 utilizes the per-cluster phasing coefficients generated as part of the per-cluster phasing coefficient operation 502 to estimate and apply per-cluster phasing corrections to the signal. In some embodiments, the base call correction system 106 utilizes the per-cluster phasing coefficients together with an algorithm, such as an FIRAttorney Docket No. IP-2817-PCT 45 Patent Applicationalgorithm, to perform the per-cluster phasing correction. Thus, the base call correction system 106 can produce the corrected intensity data 522 that can be used for determining base calls and / or base call corrections for those base calls that are associated with a sequence specific error.
[0164] As previously mentioned, the base call correction system 106 can determine the per- cluster phasing coefficient and the per-cluster pre-phasing coefficient utilizing several models or algorithms. More specifically, the base call correction system 106 can utilize various models to perform the per-cluster phasing coefficient operation 502. In particular, the base call correction system 106 can utilize a Linear Equalizer (LE) model or a Maximum Likelihood Sequence Estimator (MLSE) model to determine a per-cluster phasing coefficient and / or a per-cluster prephasing coefficient. Furthermore, the base call correction system 106 may utilize a machine learning model, such as a multilayer perceptron, to determine the coefficients. For example, in certain embodiments, the base call correction system 106 uses a LE model or a MLSE model as described in U.S. Pat. App. No. 63 / 552,510 filed on February 12, 2024, entitled “Determining Offline Corrections for Sequence Specific Errors Caused by Low Complexity Nucleotide Sequences,” and International Application No. PCT / US2025 / 015427, filed on February 11, 2025, entitled “Determining Offline Corrections for Sequence Specific Errors Caused by Low Complexity Nucleotide Sequences,” which are each incorporated herein by reference in their entirety.
[0165] As previously mentioned, by performing base call corrections, the base call correction system 106 can operate more accurately compared to conventional base calling systems. Indeed, the base call correction system 106 provides more accurate base calls for a cluster of oligonucleotides when compared to conventional systems. As further mentioned, by writing read data for base calls associated with a sequence specific error to storage and performing the base call corrections offline, the base call correction system 106 can also operate more accurately than those conventional systems that implement the base call correction process in real time (e.g., during the real-time base calling process). FIGS. 6-10 illustrate experimental results regarding the effectiveness of the base call correction system 106 performing offline base call correction in accordance with one or more embodiments.
[0166] In particular, FIG. 6 illustrates diagrams comparing base calling results generated by tested models. The tested models include a real-time model (labeled “RT”) that performs produces base calls via real time analysis. In other words, the real-time model does not include an offline component that determines base call corrections outside the real-time base calling process. Rather, the real-time model provides raw, uncorrected base calls. The tested models further include an embodiment of the base call correction system 106 (labeled “ML”) implementing a machine learning model to determine base call corrections outside the real-time base calling.Attorney Docket No. IP-2817-PCT 46 Patent Application
[0167] In particular, FIG. 6 illustrates various nucleotide reads that have been generated by each tested model and aligned in accordance with a reference sequence 604. The pair of vertical lines 602 represent the boundaries of an error-inducing sequence — such as a homopolymer or sequence or repeating dinucleotides — detected within the nucleotide reads.
[0168] As shown by the base calling results of FIG. 6, the real-time model providing the raw, uncorrected base calls produces several nucleotide reads having errors following the read position for the exit cycle of the error-inducing sequence. In other words, the real-time model failed to correct erroneous base calls caused by the error-inducing sequence for many of the nucleotide reads. As further shown, the errors produced by the real-time model range from read positions following the error-inducing sequence (or even within the error inducing sequence) to the ends of the nucleotide reads.
[0169] As further shown in FIG. 6, the embodiment of the base call correction system 106 performing offline base call correction produces nucleotide reads having fewer incorrect base calls when compared to the real-time model providing the raw, uncorrected base calls. In particular, the embodiment of the base call correction system 106 produces nucleotide reads having fewer erroneous base calls following the exit cycle for the error-inducing sequence. Thus, as mentioned above, the base call correction system 106 uses offline correction to perform well at correcting errors. In particular, the base call correction system 106 uses a machine learning model that corrects erroneous base calls using a sophisticated (e.g., deep learning) analysis.
[0170] FIG. 7 illustrates graphs that also compare the performance of the real-time model to an embodiment of the base call correction system 106. In particular, the embodiment of the base call correction system 106 (labeled “CNN+Attn”) represented in FIG. 7 implements a neural network with one or more convolutional layers and an attention mechanism. For example, the embodiment of the base call correction system 106 represented in FIG. 7 can implement a neural network having the architecture described above with reference to FIG. 4
[0171] The graphs of FIG. 7 measures the performance of each tested model using mismatch rate (MMR), which provides a measure of instances in which a determined base call did not match the true base call (e.g., as indicated by a reference sequence). Further, the graphs compare the performance of the tested models on two different reads.
[0172] As shown by the graphs of FIG. 7, the performance of the base call correction system 106 significantly outperforms the performance of the real-time model. Indeed, the mismatch rate of the base call correction system 106 is about a fifth of the mismatch rate of the real-time model for the first read and about a fourth of the mismatch rate of the real-time model for the second read.
[0173] Further, the vertical lines in each graph indicate the location of the error-inducing sequence within the respective read. Thus, the portions of these graphs following these linesAttorney Docket No. IP-2817-PCT 47 Patent Applicationrepresent sequencing cycles that occurred after the exit cycle of the error-inducing sequence. In other words, they represent base calls that follow the error-inducing sequence. As these graphs indicate, the base call that immediately follows the error-inducing sequence (the circled base call) is particularly susceptible to error. For instance, while the average mismatch rate of the real-time model is 0.75% and 1.2% for the first and second reads, respectively, the mismatch rate for the base call immediately following the error-inducing sequence is around 2% and 2.5%. By comparison, however, the mismatch rate of the base call correction system 106 for the base call immediately following the error-inducing sequence is around 0.25% and 0.8% for the first and second reads, respectively. Thus, as indicated by FIG. 7, the base call correction system 106 provides more accurate base calls both generally and for particular positions within the nucleotide reads, such as the position immediately following error-inducing sequences.
[0174] FIG. 8 illustrates bar graphs showing variant calling errors from variant calls determined based on the nucleotide reads provided by the tested models. The bar graphs of FIG. 8 each show the number of false positives (FP) and false negatives (FN) associated with the variant calling errors. As shown, the embodiment of the base call correction system 106 performing offline base call correction facilitates the best variant calling results (i.e., the fewest errors). Thus, as indicated by FIG. 8, by performing base call corrections offline, the base call correction system 106 facilitates the more accurate performance of downstream applications, such as variant calling applications.
[0175] In some cases, because performing the base call corrections offline leads to fewer error in the nucleotide reads, the variant calling application receives better input for its variant calling. Indeed, the input provided to a variant calling application by a real-time base caller that does not perform corrections on erroneous base calls is likely to include input plagued with errors. Even models performing base call corrections in real time have a high chance of providing poor input to the variant calling application. Because some — but significantly less than all — errors are corrected by these models, the resulting nucleotide reads are more likely to pass the filter of the variant calling application, which is meant to prevent erroneous nucleotide reads from being used as input. With more error-filled nucleotide reads being used as input, the variant call application generates undesirable outputs. In comparison, by performing base calling more accurately, the base call correction system 106 generates more accurate nucleotide reads. Thus, the input provided to the variant calling application is of higher quality and leads to better variant calling results.
[0176] FIGS. 9 illustrates graphs comparing the quality scores determined for the base calls generated by the tested models. In particular, the graphs plot the quality scores determined for the base calls (“pred_qs”) against the true quality scores (“emp_qs”). The diagonal line of each graph indicates where the x-value and the y-value are equal. In other words, the diagonal line marks whereAttorney Docket No. IP-2817-PCT 48 Patent Applicationthe quality score for a base call is equal to the true quality score (i.e., the quality score determined for the base call is correct). Thus, having data points that on the line indicate an accuracy of the determined quality scores. As shown in FIG. 9, by performing offline corrections using a machine learning model, the base call correction system 106 determines more accurate quality scores than those determined for base calls generated by a model that does not implement offline corrections.
[0177] FIG. 10 illustrates a graph comparing the performance of various embodiments of the base call correction system 106. In particular, the graph compares the performance of embodiments of the base call correction system 106 using various neural network architectures. Indeed, the graph compares the performance of combinations of various components, including fully connected layers (labeled “FC”), convolutional layers (labeled “CNN”), and attention mechanisms (labeled “Attention”). The graph of FIG. 10 measures each tested model using both floating point operations and mismatch rates. As shown by FIG. 10, the various architectures provide various measures of accuracy in base calling. Further, the various architectures require various amounts of operations to operate in determining base calls and corresponding corrections.
[0178] FIGS. 1-10, the corresponding text and the examples provide a number of different methods, systems, devices, and non-transitory computer-readable media of the base call correction system 106. In addition to the foregoing, one or more embodiments can also be described in terms of flowcharts comprising acts for accomplishing particular results, as shown in FIG. 11. FIG. 11 may be performed with more or fewer acts. Further, the acts may be performed in different orders. Additionally, the acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar acts.
[0179] FIG. 11 illustrates a flowchart of a series of acts 1100 for using a machine learning model to generate base call corrections for base calls associated with a sequence specific error in accordance with one or more embodiments. While FIG. 11 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 11. In some implementations, the acts of FIG. 11 are performed as part of a method, such as a computer-implemented method. In some instances, a non-transitory computer-readable medium stores instructions thereon that, when executed by at least one processor, cause a computing device to perform the acts of FIG. 11. In some implementations, a system performs the acts of FIG. 11. For example, in one or more cases, a system includes at least one processor and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to perform the acts of FIG. 11.
[0180] As shown in FIG. 11, the series of acts 1100 includes an act 1102 of determining a set of base calls associated with a sequence specific error (SSE) and a set of read data for the set of base calls; an act 1104 of writing the set of read data for the set of base calls to storage; an act 1106Attorney Docket No. IP-2817-PCT 49 Patent Applicationof determining, using a machine learning model, base call corrections for the set of base calls based on the set of data; and an act 1108 of generating corrected base calls based on the base call corrections. For example, the series of acts 1100 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A computer-implemented method comprising: determining, from read data of a cluster of oligonucleotides during a sequencing run, a set of base calls associated with a sequence specific error and a set of read data for the set of base calls; writing, to storage, the set of read data for the set of base calls associated with the sequence specific error; determining, using a machine learning model, one or more base call corrections for the set of base calls based on the set of read data written to storage; and generating one or more corrected base calls for the read data of the cluster of oligonucleotides based on the one or more base call corrections.CLAUSE 2. The computer-implemented method of clause 1, wherein determining the one or more base call corrections using the machine learning model comprises determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more fully connected layers.CLAUSE 3. The computer-implemented method of clause 2, wherein determining the one or more base call corrections using the neural network comprising the one or more convolutional layers followed by the one or more fully connected layers comprises determining the one or more base call corrections using the neural network having one or more one-dimensional convolutional layers followed by the one or more fully connected layers.CLAUSE 4. The computer-implemented method of any of clauses 1-3, wherein determining the one or more base call corrections using the machine learning model comprises determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more attention mechanisms.CLAUSE 5. The computer-implemented method of clause 4, wherein determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more attention mechanisms comprises determining the one or more base call corrections using the neural network having the one or more convolutional layers followed by one or more multi-head self-attention mechanisms.CLAUSE 6. The computer-implemented method of any of clauses 1-5, wherein determining the one or more base call corrections based on the set of read data using the machine learning model comprises:Attorney Docket No. IP-2817-PCT 50 Patent Applicationgenerating, for a base call from the set of base calls and based on corresponding read data, a set of probabilities indicating likelihoods that a corresponding nucleobase includes adenine, cytosine, guanine, or thymine using the machine learning model; and determining a base call correction for the base call based on a highest probability from the set of probabilities.CLAUSE 7. The computer-implemented method of clause 6, wherein generating the set of probabilities for the base call based on the corresponding read data and using the machine learning model comprises: generating, for the base call, a set of convolved features based on the corresponding read data using one or more layers of a neural network; and generating the set of probabilities from the set of convolved features using a softmax layer of the neural network.CLAUSE 8. The computer-implemented method of any of clauses 1-7, further comprising generating one or more quality scores for the one or more corrected base calls using outputs of the machine learning model.CLAUSE 9. The computer-implemented method of any of clauses 1-8, wherein determining the set of base calls associated with the sequence specific error comprises detecting a base call sequence indicating a homopolymer present within the cluster of oligonucleotides or indicating a sequence of repeating dinucleotides present within the cluster of oligonucleotides.CLAUSE 10. The computer-implemented method of any of clauses 1-9, wherein determining the one or more base call corrections using the machine learning model comprises determining the one or more base call corrections using a convolutional neural network or a recurrent neural network.CLAUSE 11. The computer-implemented method of any of clauses 1-10, wherein: writing the set of read data for the set of base calls to storage comprises writing to storage one or more of raw intensity data from images for the set of base calls, corrected intensity data for the set of base calls, or the set of base calls; and determining, using the machine learning model, the one or more base call corrections for the set of base calls based on the set of read data written to storage by determining, using the machine learning model, the one or more base call corrections for the set of base calls based on one or more of the raw intensity data from the images for the set of base calls, the corrected intensity data for the set of base calls, or the set of base calls.
[0181] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and whereinAttorney Docket No. IP-2817-PCT 51 Patent Applicationthe array is repeatedly imaged. Implementations in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some implementations, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic-acid polymer) can be an automated process. Preferred implementations include sequencing-by-synthesis (SBS) techniques.
[0182] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0183] SBS can utilize nucleotide monomers that have a terminator moiety or those that lack any terminator moieties. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using y-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is generally variable and dependent upon the template sequence and the mode of nucleotide delivery. For SBS techniques that utilize nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator can be reversible as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).
[0184] SBS techniques can utilize nucleotide monomers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as the release of pyrophosphate; or the like. In implementations, where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods developed by Solexa (now Illumina, Inc.).
[0185] Preferred implementations include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical BiochemistryAttorney Docket No. IP-2817-PCT 52 Patent Application242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on realtime pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to the incorporation of nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C, or G). Images obtained after the addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed, and analyzed using the methods set forth herein. For example, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.
[0186] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, in WO 04 / 018497 and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference. This approach is being commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently- labeled terminators in which both the termination can be reversed, and the fluorescent label cleaved facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from these modified nucleotides.
[0187] Preferably in reversible terminator-based sequencing implementations, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured following the incorporation of labels into arrayed nucleic acid features. In particular implementations, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially and an image of the array can be obtained between each addition step. In such implementations, each image will show nucleic acid features that have incorporated nucleotides ofAttorney Docket No. IP-2817-PCT 53 Patent Applicationa particular type. Different features are present or absent in the different images due to the different sequence content of each feature. However, the relative position of the features will remain unchanged in the images. Images obtained from such reversible terminator- SB S methods can be stored, processed, and analyzed as set forth herein. Following the image capture step, labels can be removed, and reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.
[0188] In particular implementations, some or all of the nucleotide monomers can include reversible terminators. In such implementations, reversible terminators / cleavable fluors can include fluor linked to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3' allyl group to block extension, but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30-second exposure to long- wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of natural termination that ensues after the placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.
[0189] Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100800, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.
[0190] Some implementations can utilize the detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing methods and systemsAttorney Docket No. IP-2817-PCT 54 Patent Applicationdescribed in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes an apparent signal to appear or disappear compared to the signal detected for the other member of the pair. As a second example, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on the presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary implementation that combines all three examples, is a fluorescentbased SBS method that uses a first nucleotide type that is detected in a first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).
[0191] Further, as described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so- called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0192] Some implementations can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained following treatment of an array of nucleic acid featuresAttorney Docket No. IP-2817-PCT 55 Patent Applicationwith the labeled sequencing reagents. Each image will show nucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due to the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No. 6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.
[0193] Some implementations can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147- 151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such implementations, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed, and analyzed as set forth herein. In particular, the data can be treated as an image in accordance with the exemplary treatment of optical images and other images that is set forth herein.
[0194] Some implementations can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). TheAttorney Docket No. IP-2817-PCT 56 Patent Applicationillumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed, and analyzed as set forth herein.
[0195] Some SBS implementations include the detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.
[0196] The above SBS methods can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In particular implementations, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a multiplex manner. In implementations using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle, or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as bridge amplification or emulsion PCR as described in further detail below.
[0197] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.Attorney Docket No. IP-2817-PCT 57 Patent Application
[0198] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprising components such as pumps, valves, reservoirs, fluidic lines, and the like. A flow cell can be configured and / or used in an integrated system for the detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing implementation as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference.
[0199] The sequencing system described above sequences nucleic-acid polymers present in samples received by a sequencing device. As defined herein, a "sample" (and its derivatives) is used in its broadest sense and includes any specimen, culture, and the like that is suspected of including a target. In some implementations, the sample comprises DNA, RNA, PNA, LNA, chimeric or hybrid forms of nucleic acids. The sample can include any biological, clinical, surgical, agricultural, atmospheric, or aquatic-based specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample such a genomic DNA, fresh- frozen, or formalin-fixed paraffin-embedded nucleic acid specimen. It is also envisioned that the sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample, and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some implementations, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.Attorney Docket No. IP-2817-PCT 58 Patent Application
[0200] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0201] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one implementation, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example, derived from a buccal swab, paper, fabric, or other substrates that may be impregnated with saliva, blood, or other bodily fluids. As such, in some implementations, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some implementations, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine, and serum. In some implementations, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some implementations, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some implementations, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant, or entomological DNA. In some implementations, target sequences or amplified target sequences are directed to purposes of human identification. In some implementations, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some implementations, the disclosure relates generally to human identification methods using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer designAttorney Docket No. IP-2817-PCT 59 Patent Applicationcriteria outlined herein. In one implementation, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the targetspecific primers disclosed herein or using the primer criteria outlined herein.
[0202] The components of the base call correction system 106 can include software, hardware, or both. For example, the components of the base call correction system 106 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the server device(s) 102, the client device 110, and / or the sequencing device 114). When executed by the one or more processors, the computer-executable instructions of the base call correction system 106 can cause the computing devices to perform the base call correction methods described herein. Alternatively, the components of the base call correction system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the base call correction system 106 can include a combination of computer-executable instructions and hardware.
[0203] Furthermore, the components of the base call correction system 106 performing the functions described herein with respect to the base call correction system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the base call correction system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the base call correction system 106 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0204] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in anon-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium,Attorney Docket No. IP-2817-PCT 60 Patent Application(e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0205] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0206] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0207] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.
[0208] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0209] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general-purpose computer, special purpose computer, orAttorney Docket No. IP-2817-PCT 61 Patent Applicationspecial purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0210] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0211] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0212] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.Attorney Docket No. IP-2817-PCT 62 Patent Application
[0213] FIG. 12 illustrates a block diagram of a computing device 1200 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1200 may implement the base call correction system 106. As shown by FIG. 12, the computing device 1200 can comprise a processor 1202, a memory 1204, a storage device 1206, an I / O interface 1208, and a communication interface 1210, which may be communicatively coupled by way of a communication infrastructure 1212. In certain embodiments, the computing device 1200 can include fewer or more components than those shown in FIG. 12. The following paragraphs describe components of the computing device 1200 shown in FIG. 12 in additional detail.
[0214] In one or more embodiments, the processor 1202 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1202 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1204, or the storage device 1206 and decode and execute them. The memory 1204 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1206 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0215] The I / O interface 1208 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1200. The I / O interface 1208 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1208 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1208 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0216] The communication interface 1210 can include hardware, software, or both. In any event, the communication interface 1210 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1200 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1210 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.Attorney Docket No. IP-2817-PCT 63 Patent Application
[0217] Additionally, the communication interface 1210 may facilitate communications with various types of wired or wireless networks. The communication interface 1210 may also facilitate communications using various communication protocols. The communication infrastructure 1212 may also include hardware, software, or both that couples components of the computing device 1200 to each other. For example, the communication interface 1210 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.
[0218] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0219] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Attorney Docket No. IP-2817-PCT 64 Patent Application
Claims
CLAIMSWhat is claimed is:
1. A system comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: determine, from read data of a cluster during a sequencing run, a set of base calls associated with a sequence specific error and a set of read data for the set of base calls; write, to storage, the set of read data for the set of base calls associated with the sequence specific error; determine, using a machine learning model, one or more base call corrections for the set of base calls based on the set of read data written to storage; and generate one or more corrected base calls for the read data of the cluster based on the one or more base call corrections.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base call corrections using the machine learning model by determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more fully connected layers.
3. The system of claim 2, wherein determining the one or more base call corrections using the neural network comprising the one or more convolutional layers followed by the one or more fully connected layers comprises determining the one or more base call corrections using the neural network having one or more one-dimensional convolutional layers followed by the one or more fully connected layers.
4. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base call corrections using the machine learning model by determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more attention mechanisms.
5. The system of claim 4, wherein determining the one or more base call corrections using the neural network comprising the one or more convolutional layers followed by the one or more attention mechanisms comprises determining the one or more base call corrections using the neural network having the one or more convolutional layers followed by one or more multi-head self-attention mechanisms.Attorney Docket No. IP-2817-PCT 65 Patent Application6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base call corrections based on the set of read data using the machine learning model by: generating, for a base call from the set of base calls and based on corresponding read data, a set of probabilities indicating likelihoods that a corresponding nucleobase includes adenine, cytosine, guanine, or thymine using the machine learning model; and determining a base call correction for the base call based on a highest probability from the set of probabilities.
7. The system of claim 6, wherein generating the set of probabilities for the base call based on the corresponding read data and using the machine learning model comprises: generating, for the base call, a set of convolved features based on the corresponding read data using one or more layers of a neural network; and generating the set of probabilities from the set of convolved features using a softmax layer of the neural network.
8. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate one or more quality scores for the one or more corrected base calls using outputs of the machine learning model.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the set of base calls associated with the sequence specific error by detecting a base call sequence indicating a homopolymer present within the cluster or indicating a sequence of repeating dinucleotides present within the cluster.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base call corrections using the machine learning model by determining the one or more base call corrections using a convolutional neural network or a recurrent neural network.
11. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: write the set of read data for the set of base calls to storage by writing to storage one or more of raw intensity data from images for the set of base calls, corrected intensity data for the set of base calls, or the set of base calls; and determine, using the machine learning model, the one or more base call corrections for the set of base calls based on the set of read data written to storage by determining, using the machine learning model, the one or more base call corrections for the set of base calls based on one or more of the raw intensity data from the images for the set of base calls, the corrected intensity data for the set of base calls, or the set of base calls.Attorney Docket No. IP-2817-PCT 66 Patent Application12. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to: determine, from read data of a cluster during a sequencing run, a set of base calls associated with a sequence specific error and a set of read data for the set of base calls; write, to storage, the set of read data for the set of base calls associated with the sequence specific error; determine, using a machine learning model, one or more base call corrections for the set of base calls based on the set of read data written to storage; and generate one or more corrected base calls for the read data of the cluster based on the one or more base call corrections.
13. The non-transitory computer-readable medium of claim 12, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the one or more base call corrections using the machine learning model by determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more fully connected layers.
14. The non-transitory computer-readable medium of claim 13, wherein determining the one or more base call corrections using the neural network comprising the one or more convolutional layers followed by the one or more fully connected layers comprises determining the one or more base call corrections using the neural network having one or more one-dimensional convolutional layers followed by the one or more fully connected layers.
15. The non-transitory computer-readable medium of claim 12, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the one or more base call corrections using the machine learning model by determining the one or more base call corrections using a neural network comprising one or more convolutional layers followed by one or more attention mechanisms.
16. The non-transitory computer-readable medium of claim 15, wherein determining the one or more base call corrections using the neural network comprising the one or more convolutional layers followed by the one or more attention mechanisms comprises determining the one or more base call corrections using the neural network having the one or more convolutional layers followed by one or more multi-head self-attention mechanisms.
17. A computer-implemented method comprising: determining, from read data of a cluster during a sequencing run, a set of base calls associated with a sequence specific error and a set of read data for the set of base calls; writing, to storage, the set of read data for the set of base calls associated with the sequence specific error;Attorney Docket No. IP-2817-PCT 67 Patent Applicationdetermining, using a machine learning model, one or more base call corrections for the set of base calls based on the set of read data written to storage; and generating one or more corrected base calls for the read data of the cluster based on the one or more base call corrections.
18. The computer-implemented method of claim 17, wherein determining the one or more base call corrections based on the set of read data using the machine learning model comprises: generating, for a base call from the set of base calls and based on corresponding read data, a set of probabilities indicating likelihoods that a corresponding nucleobase includes adenine, cytosine, guanine, or thymine using the machine learning model; and determining a base call correction for the base call based on a highest probability from the set of probabilities.
19. The computer-implemented method of claim 18, wherein generating the set of probabilities for the base call based on the corresponding read data and using the machine learning model comprises: generating, for the base call, a set of convolved features based on the corresponding read data using one or more layers of a neural network; and generating the set of probabilities from the set of convolved features using a softmax layer of the neural network.
20. The computer-implemented method of claim 17, further comprising generating one or more quality scores for the one or more corrected base calls using outputs of the machine learning model.
21. The computer-implemented method of claim 17, wherein the cluster comprises a cluster of oligonucleotides.Attorney Docket No. IP-2817-PCT 68 Patent Application
Citation Information
Patent Citations
Stencil mask and method of producing the same
US20050100800A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1