Offline correction of sequence-specific errors caused by low-complexity nucleotide sequences
Patent Information
- Application Number
- CN202580010180.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-12
- Filing Date
- 2025-02-11
- Publication Date
- 2026-08-18
AI Technical Summary
因此,这些系统使变体检出应用程序中容易出现错误,诸如通过增加假阳性的数量
[0010]本公开的一个或多个实施方案的附加的特征和优点将在随后的描述中阐述,并且部分地将从该描述中显而易见,或者可以通过此类示例性实施方案的实践获知。
Smart Images

Figure CN122603387A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 552,510, filed February 12, 2024, entitled “DETERMINING OFFLINE CORRECTIONSFOR SEQUENCE SPECIFIC ERRORS CAUSED BY LOW COMPLEXITY NUCLEOTIDE SEQUENCES”, the entire contents of which are incorporated herein by reference. Background Technology
[0002] In recent years, biotechnology companies and research institutions have improved hardware and software platforms for base detection in genomic samples or other nucleic acid polymers. In fact, systems and methods have been developed for analyzing nucleotide sequences (such as repeating nucleotide units or other motifs within oligonucleotides) and identifying the nucleosides contained therein. For example, some existing base detection systems identify individual nucleosides in nucleotide sequences using conventional Sanger sequencing or sequencing-by-synthesis (SBS). When using SBS, existing systems can monitor millions or billions of oligonucleotides grouped into clusters and synthesized in parallel, resulting in more accurate base detection. For example, a camera in an SBS system can capture images of fluorescently tagged nucleosides incorporated into such clusters and synthetic oligonucleotides. After capturing the images, existing SBS platforms send the image data to a computing device with sequencing data analysis software to determine the nucleotide base sequence of the genome or other nucleic acid polymer. For example, the sequencing data analysis software can determine the tagged nucleosides illuminated in a given image based on the light signals captured in the image data. By cyclically incorporating nucleosides into oligonucleotides and capturing images of the emitted light signals in various sequencing cycles, the SBS system can identify nucleotide reads corresponding to specific clusters and determine the nucleobase sequences present in whole-genome samples or other samples of nucleic acid polymers.
[0003] Despite these advances, existing base detection systems are often limited by technology that restricts their flexibility and the accuracy of the bases they identify. While some systems have attempted to improve base detection accuracy, these efforts have typically resulted in computational inefficiencies and, in some cases, introduced complex problems that affect the operation of downstream applications, such as variant detection applications.
[0004] In practice, existing base detection systems are often inflexible in adapting to the presence and influence of error-causing sequences within oligonucleotides as part of their base detection process. For example, certain low-complexity sequences (e.g., homopolymers or repeating dinucleotide sequences) are known to undergo polymerase slippage during sequencing (e.g., via SBS). Furthermore, certain chemical problems encountered during SBS can lead to poor signal quality for low-complexity sequences. These problems can cause sequence-specific errors, such as insertion-deletion (also known as "indel") errors and / or mismatch errors, after the polymerase leaves the low-complexity sequence. Many existing base detection systems fail to recognize such error-causing sequences during base detection. Moreover, these systems typically lack measures to correct errors caused by these sequences.
[0005] Because existing base detection systems are often unable to detect and correct error-causing sequences, they frequently fail to accurately identify nucleobases within oligonucleotides. Instead, these systems typically output nucleotide reads that include bases altered or shifted due to the presence of error-causing sequences. Such base detection errors can lead to further inaccuracies in downstream applications, such as variant detection applications.
[0006] As mentioned, some existing base detection systems have attempted solutions for correcting errors caused by certain sequences, but have failed to do so computationally efficiently and, in some cases, have even exacerbated the problem. For example, some existing systems use a real-time-based approach to base detection, where sequencing images acquired by the imaging system are processed during the sequencing run without being saved to storage. Some of these systems have attempted to incorporate the correction of errors caused by certain sequences into their real-time-based base detection process. However, by doing so, these systems often overcomplicate the real-time analysis component by using multiple buffers to retain some sequencing data for error correction. This overcomplication typically slows down the real-time analysis component, making base detection either impossible to perform in real-time or inaccurately in real-time. In many cases, real-time-based methods can slow down the real-time analysis component by 150% to 200%.
[0007] Furthermore, existing base detection systems that incorporate error correction into their real-time base detection processes often fail to correct many errors, potentially reducing the accuracy of downstream applications that rely on base detection, such as variant detection applications. For example, while existing systems using real-time correction techniques may offer some improvements in base detection accuracy, they typically cannot correct a large number of errors, especially those at certain read positions, such as the first position after a misleading sequence. However, variant detection applications can filter out nucleotide reads containing excessive errors to prevent these reads from negatively impacting variant detection analysis. By correcting only some base detection errors in nucleotide reads without correcting all included errors, existing systems that perform error correction as part of their real-time base detection processes tend to increase the number of error-prone nucleotide reads used in variant detection analysis. Consequently, these systems make variant detection applications prone to errors, such as by increasing the number of false positives.
[0008] These problems and challenges, as well as others, exist in existing base detection systems. Systems and methods for addressing these challenges will be described below. Summary of the Invention
[0009] This disclosure describes one or more embodiments of systems, methods, and nontransitory computer-readable storage media that address one or more of the aforementioned problems or provide other advantages over the prior art. Specifically, the disclosed systems enable flexible base detection methods involving the detection and offline correction of sequence-specific errors during sequencing runs. For example, the disclosed systems can detect the presence of certain low-complexity error-causing sequences (such as homopolymers or repeating dinucleotides) within oligonucleotide clusters when performing base detection as part of a sequencing run (e.g., an SBS sequencing run). The disclosed systems can also write intensity data associated with base detections affected by such currently error-causing sequences to a storage device and perform error correction on those base detections as a supplementary process. For example, the disclosed systems can perform error correction using an offline phasing correction model based on the stored intensity data. In this way, the disclosed systems can flexibly adapt to the presence and influence of error-causing sequences within oligonucleotides as part of the base detection process, while accurately and efficiently determining their base detection.
[0010] Additional features and advantages of one or more embodiments of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practice of such exemplary embodiments. Attached Figure Description
[0011] This disclosure will describe one or more embodiments of the invention in more detail with reference to the accompanying drawings. The following paragraphs briefly describe these drawings, in which:
[0012] Figure 1 A schematic diagram of a computational system in which a base detection correction system according to one or more embodiments may operate is illustrated.
[0013] Figure 2 An overview of a base detection correction system according to one or more embodiments is illustrated, which generates base detection correction for base detections associated with sequence-specific errors.
[0014] Figures 3A to 3C Each example illustrates a base detection correction system according to one or more embodiments for detecting base sequences known to cause sequence-specific errors.
[0015] Figure 4 An example of a phase correction model according to one or more embodiments is illustrated, which the base detection correction system uses to generate base detection corrections by performing per-cluster phase corrections.
[0016] Figure 5A An architecture of a linear equalizer (LE) model according to one or more implementations is illustrated, which the base detection correction system uses to determine the phase-determining coefficients for each cluster.
[0017] Figure 5B An architecture of a maximum likelihood sequence estimator (MLSE) model according to one or more implementations is illustrated, which the base detection correction system uses to determine the phase coefficients for each cluster.
[0018] Figure 6 An example is given of determining the phase coefficients for each cluster and generating base detection corrections for various bases in a nucleotide read, according to one or more embodiments.
[0019] Figure 7 The diagram illustrates experimental results reflecting the effectiveness of the base detection correction system according to one or more implementation schemes.
[0020] Figure 8 Additional experimental results reflecting the effectiveness of the base detection correction system according to one or more implementation schemes are illustrated in graphs.
[0021] Figure 9 The diagram illustrates further experimental results reflecting the effectiveness of the base detection correction system according to one or more implementation schemes.
[0022] Figure 10 Figures illustrating further experimental results reflecting the effectiveness of the base detection correction system according to one or more implementation schemes are shown.
[0023] Figure 11 Additional experimental results reflecting the effectiveness of the base detection correction system according to one or more implementation schemes are illustrated in graphs.
[0024] Figure 12 A flowchart illustrating a series of actions for generating base detection corrections for base detections associated with sequence-specific errors, according to one or more embodiments, is provided.
[0025] Figure 13 A block diagram illustrating an example computing device according to one or more implementation schemes is shown. Detailed Implementation
[0026] This disclosure describes one or more embodiments of a base detection correction system that corrects such errors in nucleotide reads offline based on the detection of sequences that cause sequence-specific errors during sequencing cycles, sequencing runs, or other real-time base detection methods. For illustration, in one or more embodiments, the base detection correction system performs base detection on oligonucleotide clusters as part of a sequencing process (e.g., in real-time). During base detection, the base detection correction system may detect the presence of one or more sequences (such as homopolymers or repeating dinucleotides) within the oligonucleotides that are known to cause sequence-specific errors in subsequent base detections. The base detection correction system may write intensity data collected for subsequent base detections (e.g., intensity data from images captured during sequencing) to a storage device and perform offline correction on those base detections. For example, the base detection correction system may perform per-cluster phase correction on those base detections using a phase correction model. Thus, the base detection correction system can generate base detection sequences incorporating offline correction for oligonucleotide clusters.
[0027] For illustration, in one or more embodiments, a base detection correction system determines a set of base detections associated with sequence-specific errors and intensity data of that set of base detections based on read data of oligonucleotide clusters during sequencing runs. The base detection correction system further writes the intensity data of that set of base detections associated with sequence-specific errors into a storage device. Based on the intensity data written to the storage device, the base detection correction system uses a correction model to determine one or more base detection corrections for that set of base detections. Furthermore, the base detection correction model generates one or more corrected base detections for the oligonucleotide cluster read data based on one or more base detection corrections.
[0028] In practice, as mentioned above, in one or more embodiments, the base detection correction system generates base detections for oligonucleotide clusters. Specifically, the base detection correction system determines the base detections of oligonucleotides during sequencing cycles, sequencing runs, or other real-time base detection methods. The base detection correction system may also identify a set of base detections associated with sequence-specific errors (e.g., base detections that may be incorrect due to the error) during sequencing runs or other real-time base detection methods. After performing corrections on the identified sets offline (e.g., outside of a real-time process), the base detection correction system produces a final base detection sequence that may incorporate these corrections and may also incorporate one or more base detections from the real-time determined base detections.
[0029] More specifically, as mentioned, a base detection correction system can use read data determined for an oligonucleotide cluster (e.g., read data determined via real-time base detection) to identify a set of base detections within that cluster that are associated with a sequence-specific error. For example, in some cases, the base detection correction system uses read data to determine the presence of a base detection sequence (e.g., a homopolymer or repeating dinucleotide sequence) within the oligonucleotide that causes a sequence-specific error. The base detection correction system can further determine that base detections following the erroneous sequence are associated with the sequence-specific error. In other words, the base detection correction system can determine that subsequent base detections may be incorrect due to the presence of the erroneous sequence.
[0030] In one or more embodiments, the base detection correction system also writes intensity data of the set of bases detected that are associated with sequence-specific errors to a storage device. For example, in some cases, the base detection correction system writes the intensity data to a non-volatile storage device (including non-transitory or transient storage devices). Thus, the base detection correction system stores the intensity data for offline use (e.g., outside of the real-time base detection process). In some cases, the base detection correction system also writes model parameters (e.g., the mean, covariance, or weights of a Gaussian mixture model, or least squares channel parameters—such as scaling factors and channel offsets) to a storage device for offline correction.
[0031] As further mentioned, the base detection correction system can use stored intensity data to determine (e.g., offline) the correction of those base detections associated with sequence-specific errors. In some embodiments, the base detection correction system uses a correction model (such as a phase-fixing correction model) based on the stored intensity data to determine the correction. For example, in some cases, the base detection correction system performs per-cluster phase-fixing correction using a phase-fixing correction model. In some cases, by determining the correction, the base detection correction system determines the correct base detection for that group of base detections associated with sequence-specific errors. Furthermore, the base detection (base all) correction system determines or updates quality scores for determining which base detections need to be corrected.
[0032] Based on the determined corrections, a base detection correction system can determine corrected base detections for oligonucleotide cluster read data. For example, the system can integrate these corrections into the read data (e.g., read data determined via a real-time process) to produce a final set of base detections that corrects erroneous detections present in the read data due to sequence-specific errors. In some cases, the base detection correction system generates a file with this final set of base detections and a quality score, such as a Fast Full Quality (FASTQ) file. In practice, the base detection correction system can include a quality score for this set of final base detections within the generated file.
[0033] As indicated above, the base detection correction system offers several technical advantages over existing base detection systems. For example, the base detection correction system avoids the computational time delays and inaccuracies associated with existing systems by implementing a novel base detection configuration involving (i) detecting sequence-specific errors or other real-time-based detections during sequencing runs, and (ii) correcting such sequence-specific errors offline. In fact, compared to previous attempts to correct base detection errors in real time (which significantly slowed down the real-time base detection process), the base detection correction system detects errors during sequencing cycles or sequencing runs, but performs the actual correction outside of real-time-based base detection methods. In practice, the base detection correction system can reduce the slowdown of real-time analysis components from 150% to 200% caused by real-time-based methods to approximately 5% to 10%. Therefore, the base detection correction system can correct errors while maintaining the computational efficiency of real-time base detection—while maintaining a more accurate sequence-specific error correction method.
[0034] As just noted, by performing correction offline, base detection correction systems provide more accurate base detection than systems that attempt real-time correction. Base detection correction systems not only provide more accurate base detection overall, but also at the first read position after the error-causing sequence and / or at other read positions after such an error-causing sequence. Therefore, base detection correction systems facilitate more accurate operation of downstream applications that rely on the accuracy of base detection, such as variant detection applications.
[0035] As a complement or alternative to addressing the slowness and inaccuracy of real-time base detection methods, base detection correction systems improve the flexibility of base detection correction by implementing a method that adapts to the presence and influence of certain error-causing sequences within the analyzed oligonucleotide clusters. In fact, base detection correction systems can flexibly detect and correct certain low-complexity sequences known to cause sequence-specific errors, errors that many existing systems are prone to due to polymerase slip and other chemical issues involved in sequencing low-complexity sequences. By performing the necessary corrections, base detection correction systems mitigate sequence-specific errors, thereby improving base detection. As further explained below, base detection correction systems can detect and correct homopolymers, repeating dinucleotides, repeating trinucleotides, or other repeating nucleotide units known to cause sequence-specific errors that cannot be corrected by existing systems.
[0036] Furthermore, compared to existing systems, the base detection correction system generates more accurate base detections for oligonucleotide clusters. In fact, by detecting and correcting sequence-specific errors within oligonucleotide clusters, the base detection correction system produces an output that provides a more accurate set of base detections. For example, it is estimated that homopolymer clusters account for approximately five percent of the clusters analyzed during sequencing; therefore, correcting errors within these clusters can lead to significant improvements in base detection compared to existing systems. In addition to correcting homopolymer clusters, the base detection correction system can also correct clusters with dinucleotide repeat sequences, further improving base detection accuracy.
[0037] As illustrated in the foregoing discussion, this disclosure uses a variety of terms to describe the features and advantages of the base detection correction system. For example, as used herein, the term "oligonucleotide" refers to an oligomer or other polymer of a nucleotide or analogue. Specifically, oligonucleotides may include synthetic or natural molecules comprising a covalently linked sequence of nucleotides formed by a modified phosphodiester or phosphodiester bond between the 3' position of the pentose sugar in the nucleotide and the 5' position of the pentose sugar in the adjacent nucleotide. For example, oligonucleotides may include short DNA or RNA molecules annealed to single-stranded polynucleotides for analysis or sequencing as part of SBS sequencing.
[0038] Additionally, as used herein, the term "oligonucleotide cluster" (or simply "cluster") refers to a local group or collection of oligonucleotides on a nucleotide sample slide (such as a flow cell) or other solid surface. Specifically, a cluster comprises tens, hundreds, thousands, or more copies of cloned or identical oligonucleotides. For example, in some embodiments, a cluster comprises a group of oligonucleotides immobilized in a portion of a flow cell or other sample slide. In some embodiments, clusters are uniformly spaced or organized into a systematic structure within a patterned flow cell. In contrast, in some cases, clusters are randomly organized within an unpatterned flow cell. Clusters can be imaged using one or more optical signals. For example, images of oligonucleotide clusters can be captured by a camera during sequencing cycles using light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on the flow cell.
[0039] As used herein, the term "sequencing run" refers to an iterative process on which the primary structure of nucleotide sequences from a sample (e.g., a genomic sample) is determined on a sequencing instrument. Specifically, a sequencing run includes cycles of sequencing chemistry and imaging performed by the sequencing instrument, which incorporate nucleotides into growing oligonucleotides to determine nucleotide sequences extracted from the sample (or other sequences within a library fragment) and seed nucleotide reads across a flow-through cell or other nucleotide sample slide. In some cases, a sequencing run involves replicating oligonucleotides obtained or extracted from one or more genomic samples, which are seeded in clusters across a flow-through cell or other nucleotide sample slide. Upon completion of the sequencing run, the sequencing instrument can generate nucleotide base detection data in a file such as a binary base detection (BCL) sequence file or a rapid full-mass (FASTQ) file.
[0040] Additionally, as used herein, the term "sequencing cycle" (or "cycle") refers to the repeated addition or incorporation of nucleotide bases into or into oligonucleotides, or the repeated addition or incorporation of nucleotide bases into or into oligonucleotides in parallel. Specifically, a cycle may include repeated acquisition and analysis of one or more signals and / or images having data indicating individual nucleotide bases added or incorporated into one oligonucleotide or added or incorporated into multiple oligonucleotides in parallel. Thus, cycles may be repeated as part of nucleic acid polymer sequencing. For example, in some embodiments, each sequencing cycle involves reading a single read of a DNA or RNA strand in only a single direction or reading a paired-end read of a DNA or RNA strand from both ends. Furthermore, in some cases, each sequencing cycle involves a camera capturing images of a nucleotide sample slide or multiple portions of a nucleotide sample slide to generate image data for identifying specific nucleotide bases added or incorporated into a particular oligonucleotide. After the image capture phase, the sequencing system may remove certain fluorescent labels from the incorporated nucleotide bases and perform another sequencing cycle until the nucleic acid polymer has been completely sequenced. In one or more implementations, the sequencing cycle includes a cycle within a sequencing-by-synthesis (SBS) run.
[0041] As used further herein, the term "base detection" (or "nucleobase detection") refers to the determination or prediction of a specific nucleobase (or nucleobase pair) of the genomic coordinates of an oligonucleotide (e.g., a nucleotide read) or a genome sample during a sequencing cycle. Specifically, nucleobase detection can indicate the determination or prediction of the type of nucleobase within an oligonucleotide incorporated into a nucleotide sample slide (e.g., read-based nucleobase detection). In some cases, for nucleotide reads, nucleobase detection includes determining or predicting nucleobases based on the intensity values produced by fluorescently tagged nucleotides of oligonucleotides added to a nucleotide sample slide (e.g., in a cluster of flow cells). As proposed above, a single nucleobase detection can be adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U). In some implementations, the type of base (e.g., adenine, cytosine, thymine, or guanine) can be determined based on the intensity value of the signal emitted by the nucleotide bases labeled in the oligonucleotide cluster, such as a signal in 16-quadrature amplitude modulation (QAM) or pulse amplitude modulation (PAM) 4 format.
[0042] As used herein, the term "nucleotide read" refers to a sequence of one or more nucleotide bases (or nucleobase pairs) inferred from all or part of a sample's nucleotide sequence. Specifically, a nucleotide read includes an identified or predicted base detection sequence from a nucleotide fragment (e.g., an oligonucleotide) or a group of monoclonal nucleotide fragments (e.g., an oligonucleotide cluster) from a sequencing library corresponding to a genomic sample. For example, sequencing equipment may identify nucleotide reads during a sequencing run by generating base detection of nucleotide bases passing through nanopores in a nucleotide sample slide; these base detections may be identified via fluorescent tagging or from clusters in a flow cell. In some cases, a nucleotide read may refer to a specific type of read, such as a nucleotide read synthesized from a sample library fragment shorter than a threshold number of nucleobases (e.g., an SBS read). In these or other cases, another type of nucleotide read may refer to (i) an assembled nucleotide read that has been assembled from shorter nucleotide reads to form a continuous sequence of nucleotides that meets a threshold number (e.g., an assembled nucleotide read), (ii) a cycle shared sequencing (CCS) read that meets a threshold number of nucleotides, or (iii) a nanopore long read that meets a threshold number of nucleotides.
[0043] Additionally, as used herein, the term "read data" includes data associated with one or more nucleotide reads. Specifically, read data may include data generated, extracted, or otherwise determined for a nucleotide sequence (e.g., an oligonucleotide or oligonucleotide cluster) during a sequencing run. For example, read data may include data determined when generating one or more nucleotide reads, such as base detections determined for the nucleotide reads. Read data may also include other data determined when generating nucleotide reads, such as quality scores associated with determined base detections, intensity data captured during a sequencing run, read lengths, determination of whether the nucleotide sequence includes error-causing sequences (e.g., homopolymers or repeating dinucleotides), or count data from counters used to determine the presence of error-causing sequences.
[0044] Furthermore, as used herein, the term "read position" refers to a location or coordinate on a nucleotide read. Specifically, a read position includes the location of an added labeled nucleotide along a nucleotide read. For example, a read position may indicate the location within a nucleotide read of a labeled nucleotide of a recently added oligonucleotide corresponding to a cluster when a camera captures an image of a nucleotide sample slide or a portion of a nucleotide sample slide. However, a read position may also refer to any other location within a nucleotide read, such as those preceding the location of the most recently added labeled nucleotide. Thus, in some cases, each read position within a nucleotide read corresponds to a sequencing cycle in a sequencing run used to determine the nucleotide read.
[0045] As used herein, the term "homopolymer" refers to a polymer in which all monomer units are identical. For example, a homopolymer may comprise all or a portion of a nucleotide fragment (e.g., an oligonucleotide) in which the nucleotide bases are identical. In some cases, a homopolymer comprises at least a threshold number of identical nucleotide bases. In fact, in some cases, a base detection correction system determines that all or a portion of a nucleotide fragment comprises a homopolymer after determining or predicting that the sequence of identical nucleotide bases included therein is equal to or exceeds a threshold. In at least some cases, a base detection correction system uses a threshold to determine whether a homopolymer is long enough to cause sequence-specific errors in subsequent base detections.
[0046] Additionally, as used herein, the term "repetitive dinucleotide sequence" (or simply "repetitive dinucleotide") refers to a polymer having repeating monomeric subunit pairs. For example, a repeating dinucleotide sequence may comprise all or a portion of a nucleotide fragment (e.g., an oligonucleotide) in which the nucleotide base pairs are identical. In some cases, a repeating dinucleotide sequence comprises a pair of nucleotide bases repeated in the same order and not interrupted by one or more other nucleotide bases. In some specific embodiments, a repeating dinucleotide sequence comprises at least a threshold number of repetitions. In practice, in some cases, a base detection correction system determines that all or a portion of a nucleotide fragment comprises a repeating dinucleotide sequence after determining or predicting that the number of repetitions of a included pair of nucleotide bases is equal to or exceeds a threshold. In at least some cases, a base detection correction system uses a threshold to determine whether a repeating dinucleotide sequence is long enough to cause sequence-specific errors in subsequent base detections.
[0047] Similarly, as used herein, the term "repeating trinucleotide sequence" (or simply "repeating trinucleotide") refers to a polymer having repeating monomeric subunit triplets. For example, a repeating trinucleotide sequence may comprise all or part of a nucleotide fragment (e.g., an oligonucleotide) in which the nucleotide base triplets are identical. In some cases, a repeating trinucleotide sequence comprises nucleotide base triplets repeated in the same sequence and not interrupted by one or more other nucleotide bases. In some specific embodiments, a repeating trinucleotide sequence comprises repeats at least a threshold number. In practice, in some cases, a base detection correction system determines that all or part of a nucleotide fragment comprises a repeating trinucleotide sequence after determining or predicting that the number of repeats of the included nucleotide base triplets is equal to or exceeds a threshold. In at least some cases, a base detection correction system uses a threshold to determine whether a repeating trinucleotide sequence is long enough to cause sequence-specific errors in subsequent base detections.
[0048] As used herein, the term "sequence-specific error" refers to a sequencing error caused by a specific nucleotide base sequence within a nucleotide fragment to which it is being sequenced. For example, a sequence-specific error may include one or more errors that result in the detection of incorrect bases in the nucleotide sequence following certain error-causing sequences (such as homopolymers, repeating dinucleotide sequences, or repeating trinucleotide sequences). For illustration, sequence-specific errors may include, but are not limited to, insertion / deletion errors (i.e., insertion-deletion errors) or mismatch errors.
[0049] Additionally, as used herein, the term "intensity data" refers to light or image data. Specifically, intensity data may include light or image data used to determine the base detection of a nucleotide fragment. For example, intensity data may include data corresponding to the intensity and / or wavelength of light captured within a digital image as part of a sequencing cycle. In some cases, light is emitted by an irradiated fluorescent tag added to the oligonucleotide as part of a sequencing cycle. In some cases, intensity data more generally includes data corresponding to the intensity and / or wavelength of light associated with (e.g., emitted from) the oligonucleotide cluster during a sequencing cycle. For illustration, in some cases, intensity data indicates the characteristics or properties of a signal emitted, reflected, or otherwise transmitted from a labeled nucleotide base or a group of labeled nucleotide bases from an oligonucleotide cluster. For example, intensity data may include one or more values associated with color intensity (e.g., wavelength) or light intensity (e.g., brightness). In some cases, the base detection correction system uses different channels to capture several images of oligonucleotide clusters with labeled nucleotide bases. Therefore, intensity data may correspond to the intensity of the signal observed through a particular channel.
[0050] Additionally, as used herein, the term "quality score" refers to a value indicating the quality of base detection. Specifically, a quality score may include a quantitative value indicating (i) the confidence level of base detection (including initial or corrected base detection) or (ii) the probability of an error in base detection. For example, a quality score may include a numerical value where a relatively high value indicates a relatively high confidence level in the accuracy of the corresponding base detection. For illustration, in some embodiments, the quality score includes a PHil's Read Editor (PHRED) quality score.
[0051] Furthermore, as used herein, the term "base detection correction" refers to the correction of erroneous base detections. Specifically, in some embodiments, base detection correction includes the correction of erroneous base detections associated with sequence-specific errors. For example, in some cases, base detection correction includes modified base detections. For illustration, in some cases, base detection correction includes modified base detections generated by a correction model for erroneous base detections. In some cases, base detection correction is generated as part of a binary file (such as an .obcl file), which is an intermediate file generated by the base detection correction system.
[0052] As used herein, the term "correction model" includes a computer-implemented model or algorithm for correcting erroneous base detections. Specifically, in some embodiments, the correction model includes a computer-implemented model or algorithm for generating base detection corrections for erroneous base detections. In some cases, the correction model uses intensity data associated with the erroneous base detection to generate the base detection correction. In some embodiments, the correction model generates the base detection correction within a binary file (such as a .obcl file) as an intermediate output of the base detection correction system. In some specific embodiments, the correction model includes a phase-fixing correction model. Further details regarding the correction models used in base detection correction systems are provided below.
[0053] Additionally, as used herein, the term "corrected base detection" includes updates to the base detection process or the final base detection. In some cases, corrected base detection corresponds to base detection correction. For example, in some embodiments, corrected base detection includes modified base detections generated by a correction model for previously identified erroneous base detections. However, in some embodiments, corrected base detection corresponds to base detections determined via real-time base detection. In other words, corrected base detection may include previously identified (e.g., uncorrected) base detections. For example, corrected base detection may correspond to base detections previously determined via real-time base detection based on the determination that the base detection is not associated with sequence-specific errors. In some embodiments, while base detection and base detection correction are included in an intermediate binary file (e.g., a binary base detection (BCL) file or an .ocbl file), the base detection correction system may include the corrected base detection in a separate file (such as a Fast Full Mass (FASTQ) file). For example, as will be discussed in more detail below, a base detection correction system can use a base detection data conversion module to generate a FASTQ file as a result of a base detection process that is at least in part based on a BCL file generated using a real-time analyzer and / or a .obcl file generated using a correction model.
[0054] The following paragraphs describe the base detection correction system with illustrative figures depicting example implementation schemes and specific implementations. Figure 1 A schematic diagram illustrating a computing system 100 operating within a base detection correction system 106 according to one or more embodiments is shown. As illustrated, the computing system 100 includes one or more server devices 102 connected via a network 108 to a client device 110 and a sequencing device 114. Although Figure 1 One embodiment of the base detection correction system 106 is shown, but alternative embodiments and configurations are described below.
[0055] like Figure 1 As shown, server device 102, client device 110, and sequencing device 114 are connected via network 108. Therefore, each component of computing system 100 can communicate via network 108. Network 108 includes any suitable network through which computing devices can communicate. The following section discusses... Figure 13 The example network will be discussed in more detail.
[0056] like Figure 1 As illustrated, sequencing device 114 includes computing devices for sequencing genomic samples or other nucleic acid polymers. In some embodiments, sequencing device 114 analyzes nucleotide fragments or oligonucleotides extracted from the genomic sample to generate nucleotide reads or other data directly or indirectly on sequencing device 114 using computer-implemented methods and systems. More specifically, sequencing device 114 receives a nucleotide sample slide (e.g., a flow cell) comprising nucleotide fragments (e.g., genomic DNA fragments) extracted from the sample, then replicates and determines the nucleobase sequence of such extracted nucleotide fragments.
[0057] As a supplement or alternative to communication across network 108, in some implementations, sequencing device 114 bypasses network 108 and communicates directly with server device 102 or client device 110. Additionally, as... Figure 1 As shown, in one or more embodiments, sequencing device 114 includes a base detection correction system 106.
[0058] like Figure 1 Furthermore, server device 102 can generate, receive, analyze, store, and transmit digital data, such as amino acid sequences or nucleotide sequences. Figure 1As shown, sequencing device 114 can send (and server device 102 can receive) various data from sequencing device 114, including data representing amino acid sequences or nucleotide sequences (e.g., base detection and / or corresponding quality scores). Server device 102 can also communicate with client device 110. Specifically, server device 102 can send data representing amino acid sequences or nucleotide sequences (e.g., base detection and / or corresponding quality scores) to client device 110.
[0059] In addition, such as Figure 1 As shown, server device 102 may include a base detection correction system 106. In one or more embodiments, as further explained below, base detection correction system 106 analyzes oligonucleotide clusters (e.g., analyzes images captured during sequencing runs) and generates base detections for the oligonucleotides. During the base detection process, base detection correction system 106 may identify a set of base detections associated with sequence-specific errors. For example, base detection correction system 106 may identify sequences within oligonucleotides that cause sequence-specific errors, such as homopolymers or repeating dinucleotide sequences, and determine that base detections following that sequence are associated with sequence-specific errors. Base detection correction system 106 may also employ correction model 104 to generate base detection corrections for base detections associated with sequence-specific errors and output the corrected base detections.
[0060] In some implementations, server device 102 includes a distributed collection of servers, comprising a number of server devices distributed across network 108 and located in the same or different physical locations. Furthermore, server device 102 may include a content server, application server, communication server, web hosting server, or another type of server.
[0061] In some cases, server device 102 is located at or near the same physical location as sequencing device 114 or away from sequencing device 114. In fact, in some embodiments, server device 102 and sequencing device 114 are integrated into the same computing device. Server device 102 may run software on sequencing device 114 or base detection correction system 106 to generate, receive, analyze, store, and transmit digital data, such as corrected base detections and / or corresponding quality scores, by sending or receiving them. In some embodiments, sequencing device 114 or base detection correction system 106 stores and accesses a database or table of corrected base detections and / or corresponding quality scores.
[0062] like Figure 1As further illustrated and indicated, client device 110 can generate, store, receive, and transmit digital data. Specifically, client device 110 can receive data on amino acid sequences or nucleotide sequences (e.g., base detections) from server device 102 and / or sequencing device 114. Client device 110 can accordingly present data on the sequenced oligonucleotides, such as by presenting identified base detections (including corrected base detections) and / or corresponding quality scores to a user associated with client device 110 within a graphical user interface.
[0063] Figure 1 The illustrated client device 110 may include one of various types of client devices. For example, in some embodiments, client device 110 includes a non-mobile device (such as a desktop computer or server) or another type of client device. In yet another embodiment, client device 110 includes a mobile device, such as a laptop computer, tablet computer, mobile phone, or smartphone. Further details regarding client device 110 are incorporated herein by reference. Figure 13 discuss.
[0064] like Figure 1 As further illustrated, client device 110 includes sequencing application 112. Sequencing application 112 may be a web application or native application (e.g., a mobile application or desktop application) stored and executed on client device 110. Sequencing application 112 may include instructions that, when executed, cause client device 110 to receive data from base detection correction system 106 and present data from sequencing device 114 and / or server device 102. Furthermore, sequencing application 112 may instruct client device 110 to display base detection data, such as base detections (including corrected base detections) identified for oligonucleotide clusters and / or associated quality scores.
[0065] like Figure 1 As further illustrated, the base detection correction system 106 may be located on the client device 110 or on the sequencing device 114 as part of the sequencing application 112. Thus, in some embodiments, the base detection correction system 106 is implemented (e.g., wholly or partially located) on the client device 110. As mentioned, in yet another embodiment, the base detection correction system 106 is implemented by one or more other components of the computing system 100, such as the sequencing device 114. Specifically, the base detection correction system 106 may be implemented in a variety of different ways across the server device 102, network 108, client device 110, and sequencing device 114.
[0066] although Figure 1Components of a computing system 100 communicating via network 108 are shown; however, in some embodiments, components of the computing system 100 may also bypass network 108 and communicate directly with each other. For example, and as previously mentioned, in some embodiments, client device 110 communicates directly with sequencing device 114. Additionally, in some embodiments, client device 110 communicates directly with base detection and correction system 106. Furthermore, base detection and correction system 106 may access one or more databases hosted or accessed by server device 102 or other locations within the computing system 100.
[0067] As discussed above, the base detection correction system 106 can analyze oligonucleotide clusters and generate base detections based on that analysis. Furthermore, the base detection correction system 106 can identify base detections associated with sequence-specific errors and generate corrected base detections that resolve those errors. Figure 2 An example is a base detection correction system 106 for generating one or more corrected base detections according to one or more embodiments.
[0068] like Figure 2 As shown, the base detection correction system 106 analyzes oligonucleotide clusters 202. Specifically, the base detection correction system 106 performs a sequencing run 204 on the oligonucleotide clusters 202. In one or more embodiments, the sequencing run 204 includes a real-time process. In fact, as shown, during the sequencing run 204, the base detection correction system 106 uses a real-time analyzer 206 to analyze the oligonucleotide clusters 202. Figure 2 An example of a real-time analyzer 206 is shown, which employs a base detector 208, a sequence-specific error detector 210, and an intensity data capture unit 212, but additional, fewer, or alternative components may be used in various specific implementations.
[0069] As used herein, the term "real-time analyzer" includes a computer-implemented model that generates base detections in real-time or near real-time. Specifically, a real-time analyzer may include a computer-implemented model that analyzes one or more digital images captured during a sequencing cycle in a sequencing run and generates one or more base detections of clusters of one or more oligonucleotides represented within the digital images in real-time or near real-time. For example, in some cases, a real-time analyzer may generate base detections without writing one or more digital images to disk.
[0070] like Figure 2As instructed, the base detection correction system 106 uses a real-time analyzer 206 to generate a base detection data file 214. Specifically, in some embodiments, the base detection correction system 106 uses the base detector 208 of the real-time analyzer 206 to generate the base detection data file 214 via a base detection process. For example, the base detection correction system 106 can use the base detector 208 to analyze image data captured for the oligonucleotide cluster 202 and generate a base detection 216 and a quality score 250 accordingly. For instance, in some embodiments, the base detection correction system 106 uses the base detector 208 to generate base detections as part of a sequencing-by-synthesis process, where the image data corresponds to images of irradiated fluorescent tags from nucleobases incorporated into the oligonucleotide cluster 202. Because the base detection correction system 106 can perform base detection in real time, in some specific embodiments, the base detection correction system 106 can process image data without writing the image data to a storage device.
[0071] Furthermore, in some embodiments, the base detection correction system 106 uses a real-time analyzer 206 to generate a quality score 250 via the PHil's Read Editor (PHRED) algorithm / model. For example, in some embodiments, the base detection correction system 106 uses the real-time analyzer 206 to determine the PHRED quality score using the PHRED algorithm or a modified PHRED algorithm, such as the algorithm described in U.S. Patent No. 8,392,126, filed September 23, 2009, entitled "Method and System for Determining the Accuracy of DNA Base Identifications," the entire contents of which are incorporated herein by reference. In one or more embodiments, the base detection correction system 106 generates the quality score 250 using a quality score table (or Q-table) that maps certain combinations of values (e.g., signal-to-noise ratio and / or fidelity values) determined from sequencing cycle image data to corresponding quality scores (Q-scores). In some cases, the base detection correction system 106 uses the PHRED algorithm or a modified PHRED algorithm to generate the Q-table. In some cases, the base detection correction system 106 updates or recalibrates the Q table to accommodate changes in the values used when the table was created (e.g., based on changes in values due to signal correction determined as part of an offline correction process).
[0072] Therefore, the base detections 216 of the base detection data file 214 may include those base detections determined during sequencing run 204 (e.g., via real-time base detection performed during sequencing run 204). Specifically, base detections 216 may provide nucleotide reads for oligonucleotide clusters 202. In some cases, the base detection data file 214 is formatted as a binary file, such as a binary base detection (BCL) file, but different formats may be used in various embodiments. For example, in some cases, the base detection data file 214 is formatted as a text file.
[0073] In one or more embodiments, the base detection correction system 106 uses the sequence-specific error detector 210 of the real-time analyzer 206 to determine whether a base detection 216 (e.g., a nucleotide read) identified during sequencing run 204 includes a sequence-specific error. Specifically, the base detection correction system 106 may use the sequence-specific error detector 210 to identify base detection sequences known to cause sequence-specific errors (i.e., error-causing sequences). Furthermore, the base detection correction system 106 may use the sequence-specific error detector 210 to identify a set of base detections associated with sequence-specific errors (e.g., those base detections potentially incorrect due to sequence-specific errors). For example, in some embodiments, the base detection correction system 106 will identify those base detections following error-causing sequencing as associated with sequence-specific errors.
[0074] For illustration, in one or more embodiments, the base detection correction system 106 uses a sequence-specific error detector 210 to detect homopolymers, repeating dinucleotide sequences, and / or repeating trinucleotide sequences within oligonucleotide clusters 202 based on identified base detections. Based on this identification, the base detection correction system 106 determines that base detections 216 include sequence-specific errors. Furthermore, the base detection correction system 106 can determine which base detections following the detected homopolymers, repeating dinucleotide sequences, and / or repeating trinucleotide sequences are associated with sequence-specific errors and are candidates for correction action. In one or more embodiments, the base detection correction system 106 detects error-causing sequences and identifies, in real time, the set of base detections associated with sequence-specific errors as part of sequencing run 204. Further details on how error-causing sequences are detected are provided below.
[0075] In one or more embodiments, the base detection correction system 106 uses an intensity data capturer 212 of a real-time analyzer to capture intensity data 218 of the set of bases identified as being associated with sequence-specific errors. For example, the base detection correction system 106 uses the intensity data capturer 212 to collect (e.g., extract) intensity data 218 from image data corresponding to the set of bases associated with sequence-specific errors. The base detection correction system 106 also writes the intensity data 218 to a storage device 220. In one or more embodiments, the base detection correction system 106 extracts the intensity data 218 and writes the extracted data to the storage device 220 in real time as part of a sequencing run 204.
[0076] More specifically, the base detection correction system 106 may use the intensity data capture device 212 to collect and store intensity data 218 based on signals (e.g., light signals) represented in image data during sequencing runs. For example, when a sequence-specific error is detected via the sequence-specific error detector 210, the base detection correction system 106 may activate the intensity data capture device 212 to begin collecting and storing intensity data associated with subsequent base detections (and base detections at the detection cycle). To illustrate, when it is determined that the same base detection has been determined within a threshold number of sequencing cycles, the base detection correction system 106 may determine via the sequence-specific error detector 210 that a homopolymer has been detected, and then activate the intensity data capture device 212 to collect and store intensity data for subsequent base detections (and base detections determined at the detection cycle).
[0077] In one or more embodiments, storage device 220 includes a non-volatile storage device. Furthermore, storage device 220 may include transient or non-transitory storage devices. For example, storage device 220 may include storage within a storage device integrated into a computer device (e.g., a hard disk drive or solid-state drive) or within a removable storage device (e.g., a flash drive, memory card, or external hard disk drive). Therefore, in some embodiments, base detection correction system 106 is advantageous for offline use of intensity data 218 (e.g., outside of real-time operation during sequencing run 204). However, it should be understood that in some specific embodiments, base detection correction system 106 may utilize volatile storage devices (e.g., RAM memory). In practice, in some cases, base detection correction system 106 operates on or in conjunction with a computing device having sufficient memory to store intensity data 218.
[0078] like Figure 2As shown, the base detection correction system 106 uses correction model 222 to process the intensity data 218 written to the storage device. Specifically, the base detection correction system 106 uses correction model 222 to generate base detection correction 224 for the set of base detections associated with sequence-specific errors, based on the intensity data 218 written to the storage device. As indicated, the base detection correction system 106 may generate base detection correction 224 within a base detection correction file 226. In some cases, the base detection correction file 226 includes a binary file, such as an .obcl file, containing the base detection correction 224. However, other formats may be used. For example, in some cases, the base detection correction file 226 includes a text file.
[0079] like Figure 2 As illustrated, in addition to base detection correction 224, base detection correction document 226 may also include a quality score 228 for base detection correction 224. In one or more embodiments, base detection correction system 106 uses correction model 222 to generate quality score 228 by generating a PHRED quality score based on the PHRED algorithm developed by Illumina, Inc. or a modified PHRED algorithm. For example, in some specific embodiments, base detection correction system 106 uses correction model 222 to determine the PHRED quality score using the PHRED algorithm or a modified PHRED algorithm (such as the algorithm described in U.S. Patent No. 8,392,126).
[0080] For illustration, in some specific implementations, the base detection correction system 106 formats the base detection correction file 226 as an .obcl file, which includes a binary representation of each base detection-quality score pair represented therein. For example, the base detection correction system 106 may use the first two bits of each pair to represent the base detection (i.e., the base detection correction)—such as [00, 01, 10, 11] corresponding to [A, C, G, T]—and the base detection correction system 106 may use the subsequent bits to represent the corresponding quality score. For example, the base detection correction system 106 may use Q bits to represent the quality score, where the Q bits provide an unsigned little-endian integer representation.
[0081] To more clearly compare the different files described herein, the base detection data file 214 generated using the real-time analyzer 206 includes, in some embodiments, the original base detections (e.g., base detection 216) and the original quality score (e.g., quality score 250). However, the base detection correction file 226 generated using the correction model 222 may include base detection correction 224 and a quality score 228, which includes corrections to the original quality score. Furthermore, in some embodiments, the base detection data file 214 includes base detections and quality scores for the oligonucleotide cluster 202 as a whole (e.g., all base detections and quality scores determined for the oligonucleotide cluster 202 during sequencing run 204), but the base detection correction file 226 only includes base detection corrections and quality scores for that set of base detections associated with sequence-specific errors. In some specific embodiments, the base detection correction file 226 includes an indication (e.g., in the file header) of where the base detection correction 224 is located within the nucleotide reads determined for the oligonucleotide cluster 202. For example, in some cases, base detection correction file 226 indicates the sequencing cycle associated with the first base detection correction represented therein. However, it should be understood that various file formats may be available for each file in various embodiments. Furthermore, in various embodiments, each file may include various types and amounts of data.
[0082] like Figure 2 As indicated, correction model 222 may include a phase correction model, such as a per-cluster phase correction model 230 that performs per-cluster phase correction. More details about per-cluster phase correction will be provided below.
[0083] like Figure 2 As further shown, the base detection correction system 106 uses the base detection data conversion module 232 to process the base detection data file 214 and the base detection correction file 226, and generates a corrected base detection file 234 as the result. In practice, the base detection correction system 106 typically uses the base detection data conversion module 232 to perform file format conversion. Therefore, in some cases, the base detection correction system 106 uses the base detection data conversion module 232 to convert data from the base detection data file 214 and the base detection correction file 226 into a format suitable for the corrected base detection file 234. Specifically, as shown, the base detection correction system 106 uses the base detection data conversion module 232 to generate the corrected base detection file 234 via the post-sequencing process 236. Additionally, as shown, the corrected base detection file 234 includes corrected base detections 238 and a quality score 240 for the corrected base detections 238.
[0084] In one or more embodiments, the base detection data conversion module 232 uses base detection 216 from base detection data file 214 and base detection correction 224 from base detection correction file 226 to generate corrected base detection 238. For example, as mentioned, base detection correction 224 may include corrections for those base detections associated with sequence-specific errors. Furthermore, base detection correction file 226 may identify where base detection correction 224 is located within a nucleotide read of oligonucleotide cluster 202. Therefore, in some cases, base detection correction system 106 uses base detection correction file 226 to replace erroneous base detections included in base detection 216 with base detection correction 224 to be included in the corrected base detection file 234. In other words, the base detection data conversion module 232 can include, within the corrected base detection file 234, those base detections from base detection 216 that are not associated with sequence-specific errors, and base detection correction 224 for those base detections from base detection 216 that are associated with sequence-specific errors.
[0085] As mentioned, the corrected base detection file 234 also includes a quality score 240 for the corrected base detection 238. Similar to generating the corrected base detection 238, in one or more embodiments, the base detection data conversion module 232 uses the quality score 250 from the base detection data file 214 and the quality score 228 from the base detection correction file 226 to generate the quality score 240. In practice, in some cases, the base detection correction system 106 uses the base detection correction file 226 to replace erroneous quality scores included in the quality score 250 with quality scores 228 to include in the corrected base detection file 234. In other words, the base detection data conversion module 232 may include, within the corrected base detection file 234, those quality scores from the quality score 250 that are not associated with sequence-specific errors, and those quality scores from the quality score 250 that are associated with sequence-specific errors, as well as the quality score 228.
[0086] In one or more embodiments, the corrected base detection file 234 includes a text file, such as a FASTQ file, but the base detection correction system 106 may use other file formats in various embodiments.
[0087] Therefore, the base detection correction system 106 provides a novel configuration for performing real-time detection and offline correction of sequence-specific errors detected during oligonucleotide cluster sequencing. Thus, compared to conventional base detection systems, the base detection correction system 106 improves operational efficiency and accuracy by adapting the base detection process to sequence-specific errors and correcting base detections that may be affected by errors. Furthermore, the base detection correction system 106 provides a computationally efficient method that facilitates both the real-time portion of the base detection process and allows for correction. As will be described more fully below, by separating the detection and correction processes and making the correction process (which can be time- and resource-intensive) offline, the base detection correction system 106 utilizes state-of-the-art per-cluster phase correction algorithms to improve base detection accuracy after erroneous sequences (e.g., homopolymers, repeating dinucleotide sequences, or repeating trinucleotide sequences) while minimizing the resulting real-time computational and efficiency costs.
[0088] As previously mentioned, the base detection correction system 106 can detect base sequences known to cause sequence-specific errors (i.e., error-causing sequences) within nucleotide reads of oligonucleotide clusters. For example, the base detection correction system 106 can detect homopolymers, repeating dinucleotide sequences, and / or trinucleotide sequences within nucleotide reads. Figures 3A to 3C An example is illustrated of a base detection correction system 106 according to one or more embodiments for detecting base detection sequences known to cause sequence-specific errors.
[0089] In one or more embodiments, the base detection correction system 106 detects bases known to cause sequence-specific errors when generating nucleotide reads. In other words, the base detection correction system 106 can detect erroneous sequences during the real-time base detection process. In practice, as mentioned above, the base detection correction system 106 can perform base detection during a sequencing run by processing image data from each sequencing cycle in real time without writing the image data to storage. However, as further mentioned, the base detection correction system 106 can use intensity data from the image data to generate base detection corrections for those bases that are associated with sequence-specific errors (e.g., potentially affected by sequence-specific errors). Therefore, in some cases, the base detection correction system 106 detects any present erroneous sequences during the real-time base detection process and determines which intensity data will be written to storage for subsequent base detections.
[0090] Figure 3A Example 302 illustrates the base detection of nucleotides as part of an oligonucleotide cluster, according to one or more embodiments. (e.g.) Figure 3AAs shown, the base detection 302 of the nucleotide read indicates the presence of homopolymer 304 within the oligonucleotide cluster (e.g., the base detection sequence of the 'A' nucleobase).
[0091] As shown in the figure, the base detection correction system 106 can use a length threshold 306 to detect the presence of homopolymer 304 within a nucleotide read. In one or more embodiments, the length threshold 306 indicates the length of a sequence of consecutive identical bases that is identified as a homopolymer. In other words, the length threshold 306 can indicate the number of consecutive identical bases that must be identified to determine the presence of a homopolymer (e.g., to determine that the base detection sequence is part of a homopolymer). Figure 3A The length threshold 306 shown includes a length of ten base detections. However, various lengths can be used in various embodiments. In some cases, the length threshold 306 can be configured via a computing device to adjust the number of base detections based on user input. For example, in some cases, the base detection correction system 106 modifies the length threshold 306 (e.g., in response to user input) to improve performance. In fact, in some cases, the base detection correction system 106 may determine that using a particular length threshold produces better (e.g., more efficient or more accurate) base detections compared to another length threshold. For example, the base detection correction system 106 may determine that a relatively small length threshold will result in incorrect detection of sequence-specific errors and / or lead to inefficiency when transferring operations to an offline correction system (e.g., correcting more base detections than required). Therefore, in some cases, the base detection correction system 106 adjusts the length threshold 306 to a value that provides improved base detection and base detection correction.
[0092] As suggested, the base detection correction system 106 can detect the presence of homopolymer 304 by determining that the length of the same base detection sequence is equal to or exceeds a length threshold 306. For example, Figure 3A The homopolymer 304 is illustrated as starting from a first read position 308 (labeled C#9, indicating that the first read position 308 corresponds to the ninth sequencing cycle of the sequencing run) and extending beyond a second read position 310 (labeled C#18, indicating that the second read position 310 corresponds to the eighteenth sequencing cycle of the sequencing run), wherein the number of bases detected from the first read position 308 to the second read position 310 is equal to the number of bases detected required for the length threshold 306. Figure 3A Marking the second read position 310 as a homopolymer detection cycle indicates that the second read position 310 corresponds to the sequencing cycle in which the base detection correction system 106 determines the presence of homopolymer 304.
[0093] like Figure 3AAs further shown, the base detection correction system 106 uses a buffer counter 312 to detect the presence of homopolymer 304 within nucleotide reads. In practice, because the base detection correction system 106 performs base detection in real time, in some embodiments, the base detection correction system 106 does not store base detections in memory. Therefore, in some specific embodiments, for a given sequencing cycle, the base detection correction system 106 cannot access base detections from previous sequencing cycles from memory. Therefore, in one or more embodiments, the base detection correction system 106 employs a buffer counter 312.
[0094] In one or more embodiments, the base detection correction system 106 uses a buffer counter 312 to buffer base detections from a sequencing cycle for use in the next sequencing cycle. In some cases, the base detection correction system 106 uses the buffer counter 312 to buffer one base detection at a time. However, in some embodiments, the base detection correction system 106 uses the buffer counter 312 to buffer a larger number of base detections. Furthermore, in some embodiments, the base detection correction system 106 uses the buffer counter 312 to maintain a count. The base detection correction system 106 can use the count of genomic reads (non-indexed reads).
[0095] In one or more embodiments, the buffer counter 312 includes four components: (i) a homopolymer flag buffer, which includes one bit per nucleotide read to indicate whether a homopolymer has been detected; (ii) a least squares (LSQ) parameter write switching buffer, which includes one bit per nucleotide read to indicate whether an LSQ parameter has been written to a storage device; (iii) a base detection buffer, which includes two bits per nucleotide read to store base detections from previous cycles (e.g., each combination of bits corresponding to different base detections) and for comparison with base detections in the current cycle; and (iv) a count buffer, which includes four bits per nucleotide read for accumulating the homopolymer sequence length (e.g., to provide a count from zero to fifteen). As will be explained below, if the count in the count buffer is equal to or exceeds a predetermined threshold number (e.g., ten), the homopolymer flag is triggered (e.g., switching from false to true), and the base detection correction system 106 begins accumulating intensity-related statistics.
[0096] For illustration, in one or more embodiments, for the current sequencing cycle, the base detection correction system 106 determines a base detection. The base detection correction system 106 further determines whether the base detection of the current sequencing cycle is the same as a previous base detection from a previous sequencing cycle. Specifically, the base detection correction system 106 may access the previous base detection from the buffer of the buffer counter 312 and compare the previous base detection with the base detection of the current sequencing cycle to determine whether the two base detections are the same. If the two base detections are determined to be the same, the base detection correction system 106 may increment the count of the buffer counter 312 by one. If the two base detections are not the same, the base detection correction system 106 may reset the count of the buffer counter 312. For example, in some cases, the base detection correction system 106 resets the count to zero, indicating that the base detection of the current sequencing cycle corresponds to the start of a new sequence count. When the count of buffer counter 312 equals the value required for length threshold 306, base detection correction system 106 can determine that homopolymer has been detected. Therefore, to illustrate the detection of homopolymer 304, in some embodiments, base detection correction system 106 initializes the count of buffer counter 312 to zero at the start of sequencing run.
[0097] Furthermore, the base detection correction system 106 can increment or reset the count after each base detection following the initial base detection 314, based on a comparison between that base detection and its previous base detections. When the count of the buffer counter 312 reaches the value required for the length threshold 306 at the second read position 310, the base detection correction system 106 can determine the presence of the base detection indicator homopolymer 304 from the first read position 308 to the second read position 310.
[0098] In some cases, the base detection correction system 106 continues to use the buffer counter 312 even after the presence of homopolymer 304 has been detected. For example, upon detection of homopolymer 304, the base detection correction system 106 can use the buffer counter 312 to buffer the relevant intensities of all cycles from the detection cycle in which homopolymer 304 was detected to subsequent cycles. The base detection correction system 106 further writes this data (e.g., intensity data) to a storage device. Because the base detection correction system 106 has access to all data from the detection cycle to the end of the read, the base detection correction system 106 can determine the end of homopolymer 304 during offline correction. Specifically, the base detection correction system 106 can use this data (e.g., intensity data for each cycle) to identify the read position within the nucleotide read that corresponds to the final sequencing cycle (referred to as the "exit cycle") of homopolymer 304 (e.g., Figure 3AThe base detection 316 read position is shown. Therefore, in some embodiments, the base detection correction system 106 can use this data (e.g., intensity data for each cycle) to identify subsequent read positions within a nucleotide read that have base detections (e.g., at the previous read positions) that are different from those at the previous read positions. Figure 3A The base detection at the previously read position shown is 316) different base detection (e.g., Figure 3A The base detection at the read position shown is 318), where such base detection at subsequent read positions is not part of homopolymer 304.
[0099] In some cases, upon detection of the homopolymer 304, the base detection correction system 106 begins collecting base detection intensity data. For example, as... Figure 3A As indicated, the base detection correction system 106 can collect base detection intensity data at the read position corresponding to the homopolymer detection cycle (i.e., the second read position 310), and further collect intensity data for all subsequent base detections. As previously mentioned, the base detection correction system 106 can write the collected intensity data to a storage device for generating base detection corrections for those base detections associated with sequence-specific errors caused by the presence of homopolymer 304.
[0100] Figure 3B Also exemplified is the detection of bases 322 as part of a nucleotide read from an oligonucleotide cluster, according to one or more embodiments. For example... Figure 3B As shown, the base detection 322 of the nucleotide read indicates the presence of a repeating dinucleotide sequence 324 within the oligonucleotide cluster (e.g., a base detection sequence of repeating "AT" pairs of nucleobases).
[0101] like Figure 3B As shown, the base detection correction system 106 can use a length threshold 326 to detect the presence of a repeating dinucleotide sequence 324 within a nucleotide read. In one or more embodiments, the length threshold 326 indicates the length of consecutive identical base detection pairs that are known to cause sequence-specific errors in repeating dinucleotide sequences. For example, the length threshold 326 may indicate the number of consecutive identical base detection pairs (e.g., the number of consecutive repeats of a pair of base detections) that must be identified to determine the presence of a repeating dinucleotide sequence (e.g., to determine that a base detection sequence is part of a repeating dinucleotide sequence known to cause sequence-specific errors). However, in some embodiments, the length threshold 326 indicates the number of base detections (rather than base detection pairs), but requires that those base detections be part of a repeating dinucleotide. Figure 3BA length threshold 326 is shown, which is set to the length of seven dinucleotide repeats (or fourteen bases detected as part of a repeating dinucleotide), but various lengths can be used in various embodiments. Furthermore, in some cases, the length threshold 326 can be configured via a computing device to adjust the number of bases detected based on user input.
[0102] As suggested, the base detection correction system 106 can detect the presence of a repeating dinucleotide sequence 324 by determining that the length of a sequence of consecutive identical base pairs is equal to or exceeds a length threshold 326. For example, Figure 3B The example illustrates a repeating dinucleotide sequence 324 that begins at a first read position 328 (labeled C#5, indicating that the first read position 328 corresponds to the fifth sequencing cycle of the sequencing run) and extends beyond a second read position 330 (labeled C#18, indicating that the second read position 330 corresponds to the eighteenth sequencing cycle of the sequencing run), wherein the number of base pairs (or bases detected) from the first read position 328 to the second read position 330 is equal to the number of base pairs (or bases detected) required for the length threshold 326. Figure 3B Marking the second read position 330 as a dinucleotide repeat detection cycle indicates that the second read position 330 corresponds to the sequencing cycle in which the base detection correction system 106 determines the presence of the repeating dinucleotide sequence 324.
[0103] like Figure 3B As further shown, the base detection correction system 106 uses a buffer counter 332 to detect the presence of repetitive dinucleotide sequences 324 within nucleotide reads. In one or more embodiments, the base detection correction system 106 uses the buffer counter 332 to buffer base detection pairs from a pair of sequencing cycles for use in the next sequencing cycle or the next pair of sequencing cycles. In some cases, the base detection correction system 106 uses the buffer counter 332 to buffer one base detection pair at a time. However, in some embodiments, the base detection correction system 106 uses the buffer counter 332 to buffer a larger number of base detection pairs. As will be further discussed below, in some specific embodiments, the base detection correction system 106 may move base detections into and out of the buffer with each sequencing cycle. Furthermore, in some embodiments, the base detection correction system 106 uses the buffer counter 332 to maintain a count. The base detection correction system 106 may use the count of genomic reads (non-indexed reads).
[0104] In one or more embodiments, the buffer counter 332 includes four components: (i) a dinucleotide flag buffer, which includes one bit per nucleotide read to indicate whether a repeating dinucleotide sequence has been detected; (ii) an LSQ parameter write buffer, which includes one bit per nucleotide read to indicate whether an LSQ parameter has been written to a storage device; (iii) a base detection buffer, which includes four bits per nucleotide read to store base detections from two previous cycles (e.g., two bits per cycle, where each combination of bits in each pair corresponds to a different base detection) and is used for comparison with the base detection of the current cycle; and (iv) a count buffer, which includes four bits per nucleotide read for accumulating the homopolymer sequence length (e.g., to provide a count from zero to fifteen). Given base detections at loop #(C-2), loop #(C-1), and loop #C (the current loop), buffer counter 332 increments if the following conditions are met: (i) the base detection at #(C-2) equals the base detection at loop #C, and (ii) the base detection at #(C-1) does not equal the base detection at #C. If neither condition is met, buffer counter 332 may reset. As will be explained below, if the count in the counting buffer equals or exceeds a predetermined threshold number (e.g., fourteen), a dinucleotide flag is triggered (e.g., switching from false to true), and base detection correction system 106 begins accumulating intensity-related statistics.
[0105] To illustrate at least one example of using buffer counter 332, in one or more embodiments, for the current sequencing cycle, base detection correction system 106 determines a base detection. Base detection correction system 106 further determines whether the base detection of the current sequencing cycle is the same as a previous base detection performed in the two sequencing cycles preceding the current sequencing cycle. Additionally, base detection correction system 106 determines whether the base detection of the current sequencing cycle is different from a previous base detection performed in the sequencing cycle preceding the current sequencing cycle. Specifically, base detection correction system 106 may access base detection pairs from the buffer of buffer counter 332 and compare the previous base detections from the two sequencing cycles prior (e.g., base detections in a first position of the buffer) with the base detection of the current sequencing cycle to determine whether these base detections are the same. If it is determined that two base detections are the same, base detection correction system 106 may increment the count of buffer counter 332 by one. If the two base detections are different, base detection correction system 106 may reset the count of buffer counter 332. For example, in some cases, the base detection correction system 106 resets the count to zero, indicating that the base detection pairs, including the base detections of the current sequencing cycle, correspond to the start of a new sequence count.
[0106] As mentioned, in one or more embodiments, for each sequencing cycle, the base detection correction system 106 removes a base detection from the buffer and moves a base detection determined within the current sequencing cycle into the buffer. Therefore, the buffered base detections can change with each sequencing cycle. This allows the base detection correction system 106 to recognize the start of a repeating dinucleotide regardless of its read position within a nucleotide read. For example, the first base detection of a repeating dinucleotide might appear within the second read position of a nucleotide read (e.g., in the second sequencing cycle based on the sequencing run). Moving base detections in and out of the buffer in pairs carries the risk of not correctly detecting the base detection as the start of a dinucleotide, because the base pair might be discarded if the base pairs determined in the third and fourth sequencing cycles do not match the base pairs in the first and second sequencing cycles. However, by moving base detections into the buffer within each sequencing cycle, the base detection correction system 106 only needs to determine whether the base detection of the current sequencing cycle matches the previous base detections from the two sequencing cycles prior (which would be the base detection in the first position of the buffer (i.e., the next base detection to be moved out)), and whether it matches the previous base detection from the immediately preceding sequencing cycle. During the next sequencing cycle, the base detection correction system 106 moves the base detection that was originally the second base detection to the first position in the buffer and can compare the base detection of the next sequencing cycle with that base detection in the buffer.
[0107] In one or more embodiments, when base detections are moved into and out of the buffer in each sequencing cycle, the base detection correction system 106 uses two counters in the buffer counter 332. For example, the base detection correction system 106 may use a raw counter that increments based on the determination that the base detection in the current sequencing cycle is the same as a previous base detection determined in two previous sequencing cycles. Alternatively, the base detection correction system 106 may use a pair counter that increments with each dinucleotide counted. For example, the base detection correction system 106 may increment the pair counter by one every two counts of the raw counter (whereby the base detection correction system 106 resets the raw counter after reaching a value of two). Therefore, in some embodiments, the base detection correction system 106 resets both counters when it determines that the base detection in the current sequencing cycle is different from a previous base detection determined in two previous sequencing cycles. In some cases, the base detection correction system 106 does not use two counters, but uses a single counter, but the length threshold 326 includes even numbers, which correspond to the number of dinucleotide (e.g., base detection pairs) repeats.
[0108] In some cases, the base detection correction system 106 continues to use the buffer counter 332 even after the presence of the repeating dinucleotide sequence 324 has been detected. For example, upon detection of the repeating dinucleotide sequence 324, the base detection correction system 106 may use the buffer counter 332 to buffer the relevant intensities for all subsequent cycles, and the base detection correction system 106 writes this data to a storage device. Because the base detection correction system 106 has access to all data from the detection cycle to the end of the read, the base detection correction system 106 can determine the sequencing cycle (i.e., the exit cycle) indicating the end of the repeating dinucleotide sequence 324 during offline correction. For example, as indicated above, the base detection correction system 106 may use this data (e.g., intensity data for each cycle) to identify the read position within the nucleotide read that corresponds to the final sequencing cycle (referred to as the "exit cycle") of the repeating dinucleotide sequence 324 (e.g., Figure 3B The reading position of base detection 334 is shown.
[0109] In some cases, upon detecting the presence of the repeating dinucleotide sequence 324, the base detection correction system 106 begins collecting base detection intensity data. For example, as... Figure 3B As indicated, the base detection correction system 106 can collect base detection intensity data at the read position corresponding to the dinucleotide repeat detection cycle (i.e., the second read position 330), and further collect intensity data for all subsequent base detections. In some cases, the base detection correction system 106 collects intensity data for all base detections corresponding to sequencing cycles following the detection cycle. As previously mentioned, the base detection correction system 106 can write the collected intensity data to a storage device for generating base detection corrections for those base detections associated with sequence-specific errors caused by the presence of the repeating dinucleotide sequence 324.
[0110] In one or more embodiments, the base detection correction system 106 similarly operates to detect trinucleotide sequences within a nucleotide read. For example, the base detection correction system 106 may utilize a length threshold indicating the length of consecutive identical base detection triplets identified as repetitive trinucleotide sequences. Furthermore, the base detection correction system 106 may utilize a buffer counter to count buffered base detections and include these detections in the length threshold.
[0111] Therefore, in one or more embodiments, the base detection correction system 106 may employ multiple buffer counters when analyzing nucleotide reads. For example, the base detection correction system 106 may employ a first buffer counter to detect homopolymers and a second buffer counter to detect repetitive dinucleotide sequences. As previously mentioned, the buffer counters used to detect each error-causing sequence may differ in their operation. For example, the buffer counter used to detect repetitive dinucleotide sequences may buffer two previous sequencing cycles and increase its count based on determining that the base detection of the current sequencing cycle (i) matches the base detections of the two preceding sequencing cycles and (ii) does not match the base detections from the immediately preceding sequencing cycle. In practice, the base detection correction system 106 can implement this difference to enable the determination of the distinction between repetitive dinucleotide sequences and the detection of homopolymers. For illustration, homopolymers can be considered a special case of repetitive dinucleotide sequences in which the two nucleobases within the dinucleotide are identical. Therefore, the base detection correction system 106 can determine whether the base detection of the current sequencing cycle does not match the base detection of the immediately preceding sequencing cycle, so as to ensure that the count is of the repeating dinucleotide sequence rather than the homopolymer.
[0112] Figure 3C Also exemplified is the base detection 342, determined as part of a nucleotide read from an oligonucleotide cluster according to one or more embodiments. For example... Figure 3C As shown, base detection 342 includes non-detection at several read positions (e.g., base detection marked "N").
[0113] As used herein, the term "non-detection" refers to an undetermined base detection. Specifically, non-detection may include determination of a sequencing cycle indicating that a nuclear base is indeterminate or that the base detection correction system 106 is unable to definitively determine a base detection for other reasons. In some cases, the base detection correction system 106 determines a sequencing cycle non-detection due to image data captured during the sequencing cycle. For example, in some cases, the signal captured by the image data is too weak (e.g., below a signal strength threshold) to allow for definitive base detection, or the provided signal contains sufficiently large conflicts that cause the base detection correction system 106 to have sufficiently large uncertainty regarding a correct base detection (e.g., the signal indicates multiple equally likely base detections). In some cases, the base detection correction system 106 determines a non-detection due to hardware failure. For example, the base detection correction system 106 may determine that the imaging hardware malfunctioned during a particular sequencing cycle or a set of sequencing cycles and was unable to capture images.
[0114] Figure 3CExample: Base detection 342 includes base detection 344 (at the read position marked C#7), which may be the start of a homopolymer. In fact, several base detections immediately following base detection 344 are identical to base detection 344. Therefore, the base detection correction system 106 modifies the count of the buffer counter 350 during those sequencing cycles to indicate the presence of multiple consecutive identical base detections.
[0115] However, as Figure 3C As further shown, base detection 342 also includes a series of non-detections, beginning with non-detection 346 (at the read position marked C#11) (determined only a few sequencing cycles after the sequencing cycle that determined base detection 344) and ending with non-detection 348 (at the read position marked C#18). In one or more embodiments, upon determining non-detection 346, base detection correction system 106 resets the count of buffer counter 350. Furthermore, base detection correction system 106 may reset the count of buffer counter 350 for each non-detection, up to non-detection 348 at the end of the series of non-detections.
[0116] Therefore, in one or more embodiments, the base detection correction system 106 determines the count of the reset buffer counter 350 during a non-detection event during sequencing runs. Although Figure 3C An example of a series of consecutive non-detections is illustrated, but in various cases, non-detections may be more sporadically scattered throughout the nucleotide read. In such scenarios, the base detection correction system 106 similarly resets the count of the buffer counter 350 for each non-detection. In other words, regardless of how non-detections are determined throughout the sequencing run, the base detection correction system 106 operates in the same manner—resetting the count when a non-detection is determined. If a nucleotide read containing an error-causing sequence has been identified, the base detection correction system 106 may reset the count when a new non-detection is encountered after the error-causing sequence has been detected, but still collects intensity data (e.g., and other relevant statistics) associated with those non-detections. In one or more embodiments, instead of resetting the count when a non-detection is determined, the base detection correction system 106 takes no action on the count (neither resetting nor incrementing the count). Therefore, if the base detection after non-detection sufficiently continues the pattern of the base detection before non-detection (e.g., the number of base detections following the pattern after non-detection plus the number of base detections following the pattern before non-detection satisfies the relevant length threshold), the base detection correction system 106 can use the continuation pattern to determine whether there is an error-causing sequence.
[0117] As mentioned above, the base detection correction system 106 can use a correction model (such as a phase correction model) to generate base detection corrections for base detections that are associated with sequence-specific errors (e.g., affected by sequence-specific errors). For example, the base detection correction system 106 can use a phase correction model that performs phase correction per cluster. Figure 4 An example is illustrated using a phase correction model according to one or more embodiments, which the base detection correction system 106 uses to generate base detection corrections by performing per-cluster phase corrections. Specifically, Figure 4 The diagram illustrates base detection and correction system 106 performing global phase correction 410 and per-cluster phase correction and base detection 412. It should be noted that base detection and correction system 106 can perform global phase correction 410 during a real-time base detection process, and per-cluster phase correction and base detection 412 as part of an offline correction process.
[0118] As used herein, the term "per-cluster phasing correction" refers to a process or function, when applied, to modulate the signal from labeled nucleotide bases within a specific oligonucleotide cluster to correct for estimated phasing or pre-phase. Specifically, per-cluster phasing correction may include an algorithm or function by which the signal from the cluster should be modulated to correct for the estimated effects of the estimated phasing or pre-phase using a Fourier transform.
[0119] As used herein, the term "phase fixation" refers to the incorporation (or rate) of a labeled nucleotide base after a specific sequencing cycle. Phase fixation includes the asynchronous incorporation (or rate) of labeled nucleotide bases within a cluster after other labeled nucleotide bases within the cluster for a given sequencing cycle. Specifically, during SBS, each DNA strand in the cluster is extended by one nucleotide base incorporation per cycle. One or more oligonucleotide strands within the cluster may be out of phase with the current cycle. Phase fixation occurs when the nucleotide base of one or more oligonucleotides within the cluster falls after one or more incorporation cycles. For example, the nucleotide sequence from the first to the third position could be CTA. In this example, the C nucleotide base should be incorporated in the first cycle, the T nucleotide base should be incorporated in the second cycle, and the A nucleotide base should be incorporated in the third cycle. When phase fixation occurs during the second sequencing cycle, one or more labeled C nucleotide bases are incorporated instead of the T nucleotide bases. Relatedly, as used herein, the term "pre-phase" refers to the incorporation (or rate) of one or more nucleotide bases prior to a specific cycle. A predetermined phase refers to the situation (or rate) prior to the asynchronous incorporation of labeled nucleotide bases within a cluster into other labeled nucleotide bases within the cluster for a given sequencing cycle. For example, when a predetermined phase occurs during the second sequencing cycle in the example above, one or more labeled A bases are incorporated instead of T bases.
[0120] In one or more embodiments, the base detection correction system 106 uses one or more per-cluster phasing or pre-phasing coefficients to perform per-cluster phasing correction. As used herein, the term "per-cluster phasing coefficient" refers to a factor or value that estimates or measures the per-cluster phasing of the signal for a cluster. Specifically, the per-cluster phasing coefficient estimates the effect on the phasing of a cluster within a given sequencing cycle. For example, the per-cluster phasing coefficient may indicate the effect of nucleotide bases from a previous cycle on the signal of labeled nucleotide bases from the current cycle. For illustration, in the example above, the per-cluster phasing coefficient may estimate the effect of phasing from C nucleobases incorporated during a second sequencing cycle instead of T nucleobases.
[0121] Relatedly, the term "pre-phase coefficient per cluster" refers to a factor or value that estimates or measures the pre-phase of the signal for a cluster. Specifically, the pre-phase coefficient per cluster estimates the effect of the pre-phase of the cluster within a given sequencing cycle. For example, the pre-phase coefficient per cluster may indicate the effect of a nucleotide base from a subsequent cycle on the signal from a labeled nucleotide base from the current cycle. For illustration, in the example above, the pre-phase coefficient per cluster estimates the effect of the pre-phase from an A nucleobase incorporated during a second sequencing cycle instead of a T nucleobase.
[0122] As mentioned above, the base detection correction system 106 can determine phase corrections for each cluster to correct signals from oligonucleotide clusters. Furthermore, the base detection correction system 106 can determine global phase corrections to correct signals from clusters and signals from a group of clusters. As further mentioned, the base detection correction system 106 can determine coefficients to perform corrections. Figure 4 The per-cluster phase coefficient operation 402 and the global phase coefficient operation 404 are illustrated, which are modeled as two convolution operations implemented in series by the base detection correction system 106. However, it should be understood that the order of the per-cluster phase coefficient operation 402 and the global phase coefficient operation 404 may vary in various embodiments.
[0123] In fact, in one or more embodiments, the base detection correction system 106 uses the reference above. Figure 2The real-time analyzer 206 discussed performs global phase correction. Therefore, the base detection correction system 106 can perform global phase correction coefficient calculation 404 within the real-time analyzer 206. In some cases, the base detection correction system 106 performs per-cluster phase correction offline (e.g., outside of a real-time process). Therefore, the base detection correction system 106 can also perform per-cluster phase correction coefficient calculation 402 offline. However, in some embodiments, the base detection correction system 106 performs per-cluster phase correction coefficient calculation 402 via the real-time analyzer 206. In one or more embodiments, the base detection correction system 106 determines coefficients as described in U.S. Patent Application No. 18 / 059,326, filed November 28, 2022, entitled “Generating Cluster-Specific-Signal Corrections for Determining Nucleotide-Base Calls,” the entire contents of which are incorporated herein by reference.
[0124] Specifically, Figure 4 An example of a base detection correction system 106 is illustrated, which uses a phasing model 400 to estimate various coefficients as part of generating per-cluster phasing corrections and global phasing corrections. The phasing model 400 includes operations occurring on sequencer 406 or other sequencing machines, as well as operations occurring during signal processing 408 on the sequencer or one or more servers. For example, in some embodiments, the base detection correction system 106 performs per-cluster phasing coefficient operation 402 to estimate per-cluster phasing coefficients, and performs per-cluster phasing coefficient operation 402 to estimate global phasing coefficients. The base detection correction system 106 may also utilize per-cluster phasing coefficients and global phasing coefficients as part of signal processing 408. More specifically, the base detection correction system 106 performs global phasing correction 410 to modulate the signal based on the global phasing coefficients. Furthermore, the base detection correction system 106 performs per-cluster phasing correction and base detection 412 to modulate the signal based on the per-cluster phasing coefficients and generate nucleotide base detections based on the modulated signal.
[0125] Phase-determining model 400 may include a real-time (or near-real-time) computing architecture or an offline computing architecture (e.g., a computing architecture outside the real-time environment). Typically, by utilizing a real-time computing architecture, base detection correction system 106 is executed using the processor of sequencer 406 (e.g., sequencing device 114). Figure 4One or more of the operations illustrated. In contrast, the base detection correction system 106 may also employ an offline computing architecture involving both a sequencer and one or more servers (e.g., server device 102) or involving a GPU or FPGA. In one example, the base detection correction system 106 performs signal processing 408 at one or more server devices, while simultaneously performing per-cluster phasing coefficient operation 402 and global phasing coefficient operation 404 at the sequencer 406. More specifically, the base detection correction system 106 may perform (i) global phasing correction 410 and (ii) per-cluster phasing correction and base detection 412 at the processor of the server device. As further described, in some cases, the base detection correction system 106 performs global phasing correction as part of a real-time base detection process and per-cluster phasing correction as part of an offline correction process.
[0126] Typically, and as previously mentioned, phasing and pre-phase refer to the phenomenon where a portion of an oligonucleotide in a cluster moves forward or backward by incorporating nucleotide bases corresponding to one or more previous or subsequent cycles. The base detection correction system 106 can generate a corrected signal (output signal y) based on the convolution of the cluster signal (input signal x) and the phasing coefficients per cluster (input coefficients h). In some embodiments, the phasing coefficients per cluster (h) include both the pre-phase coefficients per cluster and the phasing coefficients per cluster. The corrected signal can be modeled as a convolution operation y. c = Σ i h i x c-i Let y = x * h be the phase coefficient of each cluster. Assuming no signal attenuation, the phase coefficient h of each cluster is affected by Σ. i h i = 1 constraint, h i ≥ 0. In signal processing and communication systems literature, the D transform notation is commonly used, where D... k Indicates the delay for k loops: h(D) = ... + h -2 D -2 + h -1 D -1 + h0 + h1D + h2D 2 +……。 As recorded, h -2 D -2 + h -1 D -1 This represents the phase-fixing coefficient corresponding to the two nucleotide bases preceding the current cycle and the bases of the next cycle. h1D + h2D 2 This represents the predetermined phase coefficient corresponding to the nucleotide bases of the current cycle and the next one or two cycles.
[0127] like Figure 4As shown, the base detection correction system 106 performs a per-cluster phase coefficient calculation 402 to determine the per-cluster phase coefficient and per-cluster pre-phase coefficient for each cluster at the read position following the error-causing sequence. For illustration, the base detection correction system 106 determines the phase coefficient relative to the previous cycle (h... -1 The various per-cluster phase-determining coefficients (h) corresponding to the current cycle (h0) and the next cycle (h1). The per-cluster phase-determining coefficients vary independently between clusters and may be undeterminable for some clusters (e.g., at read positions before or within the error-causing sequence). Most clusters, unaffected by estimated or pre-phased phases, have a value h = [0 10]; however, the base detection correction system 106 can determine that the per-cluster phase-determining coefficients change randomly and abruptly after an error-causing sequence such as a homopolymer. In some embodiments, the per-cluster phase-determining coefficients sum to 1 and are all non-negative, as determined by the function Σ. i h i (c) = 1 means that h i ≥ 0.
[0128] like Figure 4 As further shown, the base detection correction system 106 performs a global phasing coefficient calculation 404 to determine global phasing coefficients. The base detection correction system 106 can utilize global phasing coefficients across clusters in a specific portion of the nucleotide sample slide (e.g., a block of flow cell). Global phasing coefficient values can be gradually varied cycle by cycle. These values are easier to estimate accurately than per-cluster phasing coefficients because the statistics can be averaged over millions of clusters.
[0129] like Figure 4 As shown, for example, the base detection correction system 106 calculates the difference between the base detection correction system 106 and the previous cycle (g -1 The global phasing coefficients (g) correspond to the current loop (g0) and the next loop (g1). Similar to the phasing coefficients of each cluster, the global phasing coefficients (g) sum to 1 and are all non-negative, as shown by the function Σ. i g i (c) = 1 means that g i ≥ 0. For example... Figure 4 As illustrated, the base detection correction system 106 adjusts the signal based on both per-cluster phase correction (including per-cluster phase coefficients) and global phase correction (including global phase coefficients).
[0130] In some embodiments, the base detection correction system 106 applies both per-cluster phasing coefficient operation 402 and global phasing coefficient operation 404 to clusters. Alternatively or additionally, the base detection correction system 106 applies global phasing coefficient operation 404 instead of per-cluster phasing coefficient operation 402 to some clusters. Specifically, in some embodiments, the base detection correction system 106 modulates signals from one or more clusters based on global phasing correction without per-cluster phasing correction. For example, as previously mentioned, the signal of nucleotide bases preceding the error-causing sequence may not require per-cluster phasing correction because the signal is not affected by the error-causing sequence. Therefore, in some embodiments, for additional oligonucleotide clusters, the base detection correction system 106 identifies different read positions preceding the error-causing sequence within different nucleotide fragment reads. The base detection correction system 106 further detects additional signals from labeled nucleotide bases within the additional oligonucleotide clusters during cycles corresponding to the different read positions. Then, the base detection correction system 106 modulates the additional signal based on global phase correction, without requiring per-cluster phase correction for each additional oligonucleotide cluster.
[0131] In yet another embodiment, the base detection correction system 106 applies per-cluster phase-fixing coefficient calculation 402 to the signal of a given cluster without performing global phase-fixing coefficient calculation 404. For example, in some cases, the base detection correction system 106 applies both the per-cluster phase-fixing coefficient and the per-cluster pre-phase coefficient (or other parameters) of a given cluster to the signal of that given cluster, without applying the parameters generated by the global coefficient calculation. Thus, when processing clusters within a nucleotide sample slide, the base detection correction system 106 can apply per-cluster phase-fixing correction (without global phase-fixing correction) to the signal of a given cluster, but apply per-cluster phase-fixing correction and global phase-fixing correction to the signals of different clusters.
[0132] As previously mentioned, the base detection correction system 106 adjusts the signal based on the per-cluster phase coefficient and the global phase coefficient as part of signal processing 408. Specifically, and as... Figure 4 As illustrated, the base detection correction system 106 performs global phase correction 410 as part of signal processing 408. The base detection correction system 106 utilizes global phase coefficients generated from global phase coefficient operation 404 and algorithms (such as the Finite Impulse Response (FIR) algorithm) to perform global phase correction 410. For example, the base detection correction system 106 is based on the previous cycle (γ... -1 The signal is adjusted by the correction (γ) corresponding to the current cycle (γ0) and the next cycle (γ1).
[0133] like Figure 4As further illustrated, the base detection correction system 106 performs per-cluster phase correction and base detection 412 as part of signal processing 408. Specifically, as part of per-cluster phase correction and base detection 412, the base detection correction system 106 estimates per-cluster phase correction and applies it to the signal using per-cluster phase coefficients generated as part of per-cluster phase coefficient operation 402. In some embodiments, the base detection correction system 106 performs per-cluster phase correction using per-cluster phase coefficients and algorithms such as the FIR algorithm. Furthermore, and as... Figure 4 As shown, the base detection correction system 106 also performs base detection. Specifically, the base detection correction system 106 generates base detection based on the adjusted signal.
[0134] As previously mentioned, the base detection correction system 106 can utilize several models or algorithms to determine the phase-determining coefficients and pre-phase-determining coefficients for each cluster. More specifically, the base detection correction system 106 can utilize various models to perform the phase-determining coefficient calculation 402 for each cluster. Specifically, the base detection correction system 106 can utilize a linear equalizer (LE) model or a maximum likelihood sequence estimator (MLSE) model to determine the phase-determining coefficients and / or pre-phase-determining coefficients for each cluster. Furthermore, the base detection correction system 106 can utilize machine learning models such as multilayer perceptrons to determine the coefficients. Figures 5A to 5B The corresponding paragraphs describe in detail how the base detection correction system 106, according to one or more embodiments, utilizes the LE model or the MLSE model.
[0135] Typically, the base detection correction system 106 can use various receiver types and computational architectures to estimate the per-cluster phase coefficients and per-cluster pre-phase coefficients. As indicated above, the base detection correction system 106 can utilize at least one of two models or algorithms as receivers: LE or MLSE. In some embodiments, the base detection correction system 106 uses forward-backward models and / or machine learning models to estimate the per-cluster phase coefficients and per-cluster pre-phase coefficients. Additionally, in some embodiments, the base detection correction system 106 uses least squares error or other optimizations to derive the per-cluster phase coefficients and per-cluster pre-phase coefficients.
[0136] The base detection correction system 106 may also utilize a real-time (or near-real-time) computing architecture or an offline computing architecture. The base detection correction system 106 utilizes a real-time computing architecture to output the final base detection in each cycle without accessing all future cycle data. For example, in some embodiments, the base detection correction system 106 requires only limited signal data to utilize a real-time computing architecture. Additionally or alternatively, the base detection correction system 106 utilizes an offline computing architecture. The base detection correction system 106 utilizes an offline computing architecture by using signal data from all cycles before performing the final base detection. For example, the base detection correction system 106 may utilize an offline computing architecture to generate per-cluster phase corrections based on signal data from all previous and subsequent cycles. The base detection correction system 106 may combine different receiver types with different computing architectures. For example, as previously discussed, the base detection correction system 106 may implement a linear equalizer model or an MLSE model via an offline computing architecture.
[0137] Typically, real-time computing architectures (e.g., for global phasing correction) limit computational complexity by using only real-time (or near-real-time) information. For example, when base detection correction system 106 utilizes a real-time computing architecture, it requires only signal data from one or more previous cycles, the current cycle, and one or more subsequent cycles. In some embodiments, base detection correction system 106 utilizes a set of signaling data from the previous cycle and a set of signaling data from subsequent cycles. Because real-time computing architectures are computationally more efficient, base detection correction system 106 can utilize real-time computing architectures to perform operations using sequencing machines or devices such as sequencing equipment 114.
[0138] In contrast, as specifically referenced above... Figure 2 The base detection correction system 106, as discussed, determines the nucleotide fragment reads of oligonucleotide clusters on a nucleotide sample slide offline after the sequencing equipment has identified them. For example, in some cases using MLSE or machine learning models, the base detection correction system 106 determines the phasing coefficients and pre-phasing coefficients for each cluster on different computing devices after the sequencing equipment has identified the nucleotide fragment reads of a given cluster, and adjusts the signal corresponding to the given cluster. By correcting those base detections associated with sequence-specific errors, the base detection correction system 106 provides an efficient base detection correction method. In fact, the base detection correction system 106 does not correct the entire nucleotide read, but only applies the correction process to those base detections that may have been affected by sequence-specific errors. Therefore, the base detection correction system 106 enables a real-time analyzer to perform some base detections without writing image data to disk, while moving potentially erroneous base detections to offline correction.
[0139] In contrast, offline computing architectures often require more storage resources. However, as suggested above, the base detection correction system 106 can generate more accurate results by utilizing an offline computing architecture. For example, by utilizing an offline computing architecture, the base detection correction system 106 processes a large number of clusters and cycles in parallel. This type of processing requires significant storage, communication, and computing resources for per-cluster phasing and pre-phase estimation. However, utilizing an offline computing architecture can also produce more accurate results because the base detection correction system 106 processes signaling data for all cycles. In some implementations, the base detection correction system 106 performs offline computations when the sequencer or device is online and actively communicating with a central processing system.
[0140] As mentioned, Figure 5A An example of a base detection correction system 106 according to one or more embodiments is illustrated, which utilizes a linear equalizer (LE) model to determine the phasing coefficients and / or pre-phasing coefficients per cluster. Typically, an LE model is a model for implementing a linear filter, which can be designed or optimized to suppress inter-symbol interference (ISI) or filter noise. ISI refers to the form of signal distortion in which one symbol interferes with subsequent symbols. The effects of other symbols can have similar effects to noise, thereby reducing the reliability of communication. The base detection correction system 106 can optimize the LE model to find an appropriate trade-off between suppressing ISI and minimizing noise amplification. In some embodiments, the base detection correction system 106 utilizes a linear equalizer model implemented as an FIR filter. Using such an equalizer, the base detection correction system 106 linearly weights the current and previous values of the input signal through filter coefficients. For example, in some embodiments, the current and previous values include the current signal and the previous signal from the cluster. The base detection correction system 106 also sums the weighted current and previous values to generate the regulated signal.
[0141] Figure 5A The architecture of a linear equalizer 500 according to one or more embodiments is illustrated. Typically, the base detection correction system 106 inputs the input signal x into the linear equalizer 500 to generate the regulated signal x̂. As previously stated, h represents the phase coefficient per cluster. Therefore, h(D) represents the first filter. Additive noise is denoted by n ~ CN(0,σ). 2 () indicates. For example... Figure 5A As further illustrated, w represents the weight, and w(D) represents the second filter. The base detection correction system 106 also utilizes a decision device 502 to process the signal to generate the regulated signal x̂.
[0142] In order to determine Figure 5A In the LE structure shown, h is used to define S(f) as the frequency domain SNR.
[0143]
[0144] Where F(h) represents the Fourier transform of h(D). The base detection correction system 106 can generate a measure of signal quality by determining the signal-to-interference-plus-noise ratio (SINR). Assuming the presence of Gaussian noise, the SINR ratio can be used to derive the error rate for binary signals or other modulation types. For an ideal infinite-length unbiased minimum mean square error linear equalizer (U-MMSE-LE), it can be shown as follows:
[0145] .
[0146] The error rate can be approximately estimated using the following formula:
[0147] .
[0148] Where P 误差 The transmission power representing the error. For example... Figure 5A Given the signal and noise levels in a given frequency band, as indicated by the corresponding function, the base detection correction system 106 calculates the total SNR after receiver processing and then converts the SNR into an error rate estimate.
[0149] In some implementations, the base detection correction system 106 utilizes a 3-tap LE model to generate previous cycle weights, next cycle weights, and current cycle weights. Specifically, the base detection correction system 106 generates previous cycle weights based on the phase-fixing coefficients per cluster to estimate the phase-fixing effect of nucleotide bases used in the previous cycle. The base detection correction system 106 also generates next cycle weights based on the pre-phase coefficients per cluster to estimate the pre-phase effect of nucleotide bases used in the next cycle. Furthermore, the base detection correction system 106 generates current cycle weights based on the phase-fixing coefficients per cluster and the pre-phase coefficients per cluster to estimate the phase-fixing effect and the pre-phase effect.
[0150] In some implementations, the base detection correction system 106 determines the weights of the previous cycle (w). -1 The current loop weight (w0) and the next loop weight (w1) are used. Typically, the base detection correction system 106 can optimize the parameters using optimization algorithms such as least squares error or another optimization algorithm. For example, the base detection correction system 106 can generate decision-oriented minimum least squares estimates.
[0151] After generating decision-oriented minimum least squares estimates or otherwise optimizing the parameters, the base detection correction system 106 can then use intermediate statistics to calculate the phase-determining coefficients (a) and pre-phase-determining coefficients (b) for each cluster. Specifically, the base detection correction system 106 utilizes intermediate statistics that are part of minimizing the squared error across several cycles and across one or more channels. Instead of maintaining all values for each channel per cycle, the base detection correction system 106 efficiently accumulates running statistics.
[0152] As mentioned, in yet another embodiment, the base detection correction system 106 uses a maximum likelihood sequence estimator (MLSE) model to determine the phase coefficients and pre-phase coefficients for each cluster. Figure 5B An example of a maximum likelihood sequence estimator architecture 510 according to one or more implementations is illustrated. The MLSE model is a model that implements a nonlinear estimation technique, replacing the equalization filter with MLSE estimation. Typically, the base detection correction system 106 utilizes the MLSE model to test all possible data sequences (instead of decoding each received signal manually) and selects the output signal with the highest probability as the output. As shown, the MLSE model uses a Viterbi decoder 512 to determine the probability of all possible transmitted sequences. Figure 5B As illustrated, the base detection correction system 106 inputs the input signal x into the maximum likelihood sequence estimator architecture 510 to generate the regulated signal x̂. The maximum likelihood sequence estimator architecture 510 includes filters h(D) corresponding to each cluster of phase coefficients h. The additive noise of the signal x is determined by n ~ CN(0,σ). 2 )express.
[0153] In one or more implementations, the error rate is limited by the matched filter bound (MFB) as follows:
[0154]
[0155] Where SNR represents the signal-to-noise ratio, and P 误差 The transmission power represents the error. Typically, SNR compares the desired signal level to the background noise level. For example... Figure 5B As indicated by the corresponding function, the base detection correction system 106 uses Passevar's theorem to determine the total signal power by summing the responses in the time domain. The total signal power may be the same as or equal to the total power in the frequency domain. Once the base detection correction system 106 determines the SNR, it calculates the error limits. (The above is related to...) Figure 5B In the corresponding function, the number of states is N 长度(h)-1 Given, where N is the number of constellation points. For square constellations with uncorrelated noise, the two SBS channels can be processed independently, thus reducing the number of states.
[0156] As indicated above, except Figures 5A to 5B In addition to the illustrated receivers LE and MLSE, the base detection correction system 106 may utilize other models. For example, the base detection correction system 106 may utilize other Hidden Markov Models (HMMs) besides those listed above. For instance, in some embodiments, the base detection correction system 106 may utilize a forward-backward model to generate a maximum posterior probability (MAP) estimate. The forward-backward model calculates the posterior maximum path probability for each state at a given time. Typically, the forward-backward model utilizes dynamic programming principles to compute the values needed to obtain the posterior marginal distribution over two passes. The first pass is forward in time, while the second pass is backward in time.
[0157] In addition to the models listed above, the base detection correction system 106 may also utilize machine learning models to determine the phase-determining coefficients and pre-phase-determining coefficients for each cluster. Typically, the base detection correction system 106 may use machine learning models to estimate the phase-determining coefficients and pre-phase-determining coefficients for each cluster, modulate the resulting signal, or directly modulate nucleotide base detection. For example, in some embodiments, the base detection correction system 106 utilizes a sequence-to-sequence machine learning model based on convolutional layers. Additionally or alternatively, the base detection correction system 106 may utilize recurrent neural networks (RNNs) such as Long Short-Term Memory (LSTM) to estimate the phase-determining coefficients and pre-phase-determining coefficients for each cluster. In yet another embodiment, the base detection correction system 106 utilizes an attention-based model.
[0158] As discussed, the base detection correction system 106 can determine the per-cluster phasing coefficients for one or more base detections that are associated with (e.g., affected by) sequence-specific errors in nucleotide reads. Additionally, the base detection correction system 106 can use the per-cluster phasing coefficients to generate base detection corrections for the base detections. Figure 6 An example is given of determining the phase coefficients for each cluster and generating base detection corrections for various bases in a nucleotide read, according to one or more embodiments.
[0159] Specifically, Figure 6 The example illustrates the detection of base 602 in nucleotide reads of oligonucleotide clusters including homopolymer 604. It should be noted that the following can also be applied to oligonucleotide clusters including repeating dinucleotide or trinucleotide sequences, but the detection methods differ somewhat from those described above. Figure 6 It also indicated the positions of certain reads within the nucleotide reads. For example, Figure 6Read positions 606 (C#9), 608 (C#18), 610 (C#26), and 651 (C#151) are indicated. Read position 606 corresponds to the read position with the first base detection of homopolymer 604. Read position 608 corresponds to the read position of the homopolymer detection cycle (D) – sequencing cycle, at which the base detection correction system 106 determines the number of consecutive identical bases detected that satisfy the length threshold 614 (or more generally, the read position of the sequencing cycle, at which the base detection correction system 106 identifies sequence-specific errors). Read position 610 corresponds to the read position of the homopolymer exit cycle (X) – sequencing cycle, at which the base detection correction system 106 determines the final base detection of homopolymer 604. Read position 612 corresponds to the final read position of the nucleotide read (E) – the final genomic read.
[0160] like Figure 6 As shown and discussed above, the base detection correction system 106 can perform an action 616 to determine the phase coefficient of one or more bases detected in a nucleotide read per cluster. For example, Figure 6 The base detection correction system 106 is shown to determine the phasing factor for each cluster of bases detected from read position 608 corresponding to the homopolymer detection cycle to read position 612 corresponding to the last genomic read. Alternatively, the base detection correction system 106 can determine the phasing factor for each cluster of bases detected from read position 608 (but before read position 610 corresponding to the homopolymer exit cycle) to read position 612 corresponding to the last genomic read. Furthermore, the base detection correction system 106 can determine the phasing factor for each cluster of bases detected from read position 610 corresponding to the homopolymer exit cycle to read position 612 corresponding to the last genomic read. It should be understood that various other ranges are used in various embodiments.
[0161] In addition, such as Figure 6 As shown and discussed above, the base detection correction system 106 can perform a base detection correction action 618 to determine the detection of one or more bases in a nucleotide read. For example, Figure 6The base detection correction system 106 is shown to determine base detection corrections from read position 608 corresponding to the homopolymer detection cycle to read position 612 corresponding to the last genomic read. In practice, as previously mentioned, sequence-specific errors typically affect base detections determined after the erroneous sequence. However, in some embodiments, the base detection correction system 106 can determine base detection corrections for one or more bases within the erroneous sequence itself. Alternatively, the base detection correction system 106 can determine base detection corrections from read positions after read position 608 (but before read position 610 corresponding to the homopolymer exit cycle) to read position 612 corresponding to the last genomic read. Furthermore, the base detection correction system 106 can determine base detection corrections from read positions exactly after read position 610 corresponding to the homopolymer exit cycle to read position 612 corresponding to the last genomic read. Therefore, in some cases, the base detection correction system 106 only determines the base detection correction for those bases detected after the error-causing sequence. It should be understood that various other scopes are used in different implementations.
[0162] As previously mentioned, by performing base detection correction, the base detection correction system 106 can operate more accurately than conventional base detection systems. In fact, the base detection correction system 106 provides more accurate base detection for oligonucleotide clusters compared to conventional systems. As further mentioned, by writing the intensity data of base detection associated with sequence-specific errors to a storage device and performing base detection correction offline, the base detection correction system 106 can also operate more accurately than those conventional systems that perform base detection correction in real time (e.g., during a real-time base detection process). Figures 7 to 11 Experimental results are illustrated regarding the effectiveness of the base detection correction system 106 in performing offline base detection correction according to one or more embodiments.
[0163] Specifically, Figure 7 A first graph 702 and a second graph 704 illustrate the performance of the tested models in terms of total mismatch rate. MMR can include a measure of the alignment of sequencing reads (e.g., nucleotide reads) with a reference genome. For example, MMR can represent the ratio (or proportion) of bases in a sequencing read that do not match the corresponding bases in the reference genome. Therefore, a lower MMR value indicates that the sequencing reads are more accurate. The first graph 702 shows the MMR values of the tested models performing base detection on polyA sequences across various sequence lengths, and the second graph 704 shows the MMR values of the tested models performing base detection on polyT sequences across various sequence lengths.
[0164] like Figure 7As shown, first graph 702 and second graph 704 compare the performance of a model (e.g., an implementation of base detection correction system 106 or a conventional base detection system) performing base detection correction without correction (e.g., real-time base detection) (labeled "raw") with two implementations of base detection correction system 106 performing base detection correction offline: an implementation using the LE model and an implementation using the MLSE model. As shown, in both cases, the offline implementation of base detection correction system 106 performs more accurately (e.g., lower MMR values) than the other models that provide raw, uncorrected base detection. The implementation of base detection correction system 106 using the MLSE model for offline base detection correction provides the optimal values.
[0165] Figure 8 The first graph 802 and the second graph 804 illustrate the performance of the tested model in terms of the percentage of soft shearing. Soft shearing can refer to the process in which certain portions of a sequencing read are not used for further processing. For example, soft shearing can involve removing certain bases (e.g., terminal bases) from the sequencing read to align with the reference genome when they do not align well. Therefore, a lower percentage of soft shearing indicates better alignment of the sequencing read with the reference genome. Figure 8 As shown, the implementation of the base detection correction system 106 using offline correction performs more accurately (e.g., with a lower soft shear percentage) than other models that provide raw, uncorrected base detection. Similarly, the implementation of the base detection correction system 106 using an MLSE model for offline correction provides optimal performance.
[0166] Figure 9 The first graph 902 and the second graph 904 illustrate the performance of the tested model in terms of average mapping quality (Q). Mapping quality measures the confidence level of a sequenced read aligned with a reference genome. Therefore, a higher average mapping quality corresponds to more accurate nucleotide reads. Figure 9 As shown, the implementation of the base detection correction system 106 using offline correction performs more accurately than other models that provide raw, uncorrected base detection (e.g., higher average mapping quality value). Similarly, the implementation of the base detection correction system 106 using the MLSE model for offline correction provides optimal performance.
[0167] Figure 10 A graph illustrating the base detection results generated by the tested model is shown. Specifically, Figure 10 Examples of various nucleotide reads generated by each tested model and aligned to a reference sequence are shown. The vertical line pair 1002 indicates the boundary of error-causing sequences (such as homopolymers or repeating dinucleotide sequences) detected within the nucleotide read.
[0168] like Figure 10 The base detection results show that models providing raw, uncorrected base detections produce several nucleotide reads with errors (e.g., error 1004) at read positions immediately following the exit sequence of the erroneous sequence. In other words, for many nucleotide reads, the real-time model fails to correct the erroneous base detection caused by the erroneous sequence at that read position. As further shown, many nucleotide reads generated by the model (especially those with erroneous base detections at read positions after the exit sequence) include additional erroneous base detections at read positions further below the nucleotide reads.
[0169] like Figure 10 As further shown, compared to other models that provide raw, uncorrected base detection, the implementation of the base detection correction system 106 performing offline base detection correction consistently produces nucleotide reads with fewer incorrect base detections. Specifically, the offline base detection correction implementation produces nucleotide reads with fewer incorrect base detections at the read position immediately following the exit cycle of the erroneous sequence. The offline base detection correction implementation also produces additional incorrect base detections further below the nucleotide read. Notably, the implementation of the base detection correction system 106 using the MLSE model for offline base correction appears to provide nucleotide reads with minimal errors.
[0170] Therefore, as mentioned above, the base detection correction system 106 not only performs well overall in correcting errors using offline correction, but also provides improved correction more specifically for base detection errors that occur immediately following the error-causing sequence.
[0171] Figure 11 A bar graph is illustrated, showing variant detection errors in variant detection determined based on nucleotide reads provided by the tested model. Specifically, the tested model represents a model that does not implement correction measures (e.g., a conventional base detection system), an implementation of a base detection correction system 106 that performs offline base detection correction, and another model that performs base detection correction in real time (e.g., an implementation of a base detection correction system 106 or a conventional base detection system). Figure 11 The bar graph shows the number of false positives (FP) and false negatives (FN) associated with variant detection errors. As shown, the implementation of the base detection correction system 106, which performs offline base detection correction, not only promotes optimal variant detection results (i.e., the fewest errors), but also the model that performs base detection correction in real time (labeled here as "online correction") causes more errors in variant detection than the model that does not implement any correction measures. Therefore, as Figure 11As indicated, by performing base detection correction offline, the base detection correction system 106 facilitates more accurate execution of downstream applications, such as variant detection applications.
[0172] In some cases, variant detection applications receive better input for their variant detection because offline base detection correction reduces errors in nucleotide reads. However, as previously mentioned, nucleotide reads with some errors may pass through the variant detection application's filter, which attempts to remove erroneous inputs that could cause incorrect output. Since models performing real-time base detection correction have a higher chance of not removing all errors from nucleotide reads compared to models performing offline correction, real-time correction models have a higher chance of providing erroneous nucleotide reads that might pass through the variant detection application's filter, resulting in poorer variant detection results.
[0173] Figures 1 to 11 The accompanying text and examples provide numerous different methods, systems, apparatuses, and non-transitory computer-readable media for the base detection correction system 106. In addition to the foregoing, one or more embodiments may also be described according to flowcharts including actions for achieving specific results, such as… Figure 12 As shown. Figure 12 It can be performed with more or fewer actions. Furthermore, these actions can be performed in different orders. Additionally, the actions described herein can be repeated, performed in parallel with each other, or performed in parallel with different instances of the same or similar actions.
[0174] Figure 12 A flowchart illustrating a series of actions 1200 according to one or more embodiments for generating base detection corrections for base detections associated with sequence-specific errors is shown. Although Figure 12 The examples illustrate actions according to one implementation scheme, but alternative implementation schemes may omit, add, reorder, and / or modify them. Figure 12 Any action shown. In some specific implementations, Figure 12 The actions are performed as part of a method. In some cases, a non-transitory computer-readable medium stores the following instructions thereon: these instructions, when executed by at least one processor, cause the computing device to perform... Figure 12 The action. In some specific implementations, the system executes... Figure 12 The system includes, for example, in one or more cases, at least one processor and a non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform actions. Figure 12 The action.
[0175] like Figure 12As shown, the series of actions 1200 includes actions 1202 of determining a set of base detections and corresponding intensity data associated with a sequence-specific error (SSE); actions 1204 of writing the intensity data of the set of base detections into a storage device; actions 1206 of determining base detection corrections and associated quality scores for the set of base detections based on the intensity data; and actions 1208 of generating corrected base detections based on the base detection corrections. For example, the series of actions 1200 may include actions that perform any of the operations described in the following clauses:
[0176] Clause 1. A computer-implemented method, the computer-implemented method comprising:
[0177] Based on read data of oligonucleotide clusters during sequencing runs, a set of bases detected that are associated with sequence-specific errors and the intensity data of said set of bases detected are determined;
[0178] The intensity data of the set of bases detected in connection with the sequence-specific error is written into a storage device;
[0179] Based on the intensity data written to the storage device, a correction model is used to determine one or more base detection corrections for the set of bases detected; and
[0180] One or more corrected base detections are generated for the read data of the oligonucleotide cluster based on the one or more base detection corrections.
[0181] Clause 2. The computer-implemented method according to Clause 1, wherein determining the set of base detections associated with the sequence-specific error comprises detecting a base detection sequence, the base detection sequence indicating a homopolymer present in the oligonucleotide cluster, indicating a repeating dinucleotide sequence present in the oligonucleotide cluster, or indicating a repeating trinucleotide sequence present in the oligonucleotide cluster.
[0182] Clause 3. The computer-implemented method according to Clause 1, wherein determining the one or more base detection corrections for the set of bases using the correction model based on the intensity data written to the storage device includes determining the one or more base detection corrections for the set of bases using a phase correction model, the phase correction model performing per-cluster phase correction based on the intensity data written to the storage device.
[0183] Clause 4. The computer-implemented method according to Clause 3, wherein determining the one or more base detection corrections for the set of bases using the phase correction model includes determining the one or more base detection corrections for the set of bases using the phase correction model to perform the per-cluster phase correction via a linear equalizer model or a maximum likelihood sequence estimator model.
[0184] Clause 5. The computer-implemented method according to Clause 3, the computer-implemented method further comprising determining one or more per-cluster phasing coefficients via the phasing correction model to perform the per-cluster phasing correction using the intensity data of a set of read positions, the set of read positions ranging from the read position within the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
[0185] Clause 6. The computer-implemented method according to Clause 1, wherein determining the one or more base detection corrections for the set of base detections using the correction model includes determining a set of read positions using the correction model, the set of read positions ranging from the read position following the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
[0186] Clause 7. The computer-implemented method according to Clause 1, wherein determining the one or more base detection corrections for the set of base detections using the correction model includes determining a set of read positions using the correction model, the set of read positions ranging from the read position within the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
[0187] Clause 8. The computer-implemented method according to Clause 1, the computer-implemented method further comprising generating one or more quality scores for the one or more corrected base detections of the read data of the oligonucleotide cluster.
[0188] Clause 9. The computer-implemented method according to Clause 1, the computer-implemented method further comprising detecting base detection sequences known to cause sequence-specific errors within nucleotide reads of the oligonucleotide cluster.
[0189] Clause 10. The computer-implemented method according to Clause 1, wherein writing the intensity data of the set of bases detected in association with the sequence-specific error into a storage device comprises writing the intensity data into a non-transitory computer-readable medium or a transient computer-readable medium.
[0190] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which nucleic acids are attached to fixed positions in the array such that their relative positions do not change, and in which the array is repeatedly imaged. Specific embodiments that obtain images in different color channels (e.g., corresponding to different markers used to distinguish one nucleobase type from another) are particularly suitable. In some embodiments, the process of determining the nucleotide sequence of the target nucleic acid (i.e., the nucleic acid polymer) can be automated. Preferred embodiments include sequencing-by-synthesis (SBS) technology.
[0191] SBS technology typically involves the enzymatic elongation of nascent nucleic acid chains through repeated addition of nucleotides to the template strand. In conventional SBS methods, a single nucleotide monomer is delivered to the target nucleotide in the presence of polymerase during each delivery. However, in the method described herein, more than one type of nucleotide monomer can be delivered to the target nucleic acid in the presence of polymerase during delivery.
[0192] SBS can utilize nucleotide monomers with or without a terminator motif. Methods utilizing nucleotide monomers lacking a terminator include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in further detail below. In methods using nucleotide monomers lacking a terminator, the number of nucleotides added in each cycle is typically variable and depends on the template sequence and the method of nucleotide delivery. For SBS techniques utilizing nucleotide monomers with a terminator motif, the terminator can be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or the terminator can be reversible, as in sequencing methods developed by Solexa (now Illumina, Inc.).
[0193] SBS technology can utilize nucleotide monomers with or without a labeled moiety. Therefore, incorporation events can be detected based on: the characteristics of the label, such as the fluorescence of the label; the characteristics of the nucleotide monomer, such as molecular weight or charge; byproducts of the incorporated nucleotide, such as the release of pyrophosphate; and so on. In specific implementations where two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be indistinguishable under the detection technique used. For example, the different nucleotides present in the sequencing reagent may have different labels, and they can be distinguished using appropriate optics, as exemplified by the sequencing method developed by Solexa (now Illumina, Inc.).
[0194] Preferred specific implementations include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry, Vol. 242, No. 1, pp. 84-89; Ronaghi, M. (2001), “Pyrosequencing sheds light on DNA sequencing.”, Genome Res., Vol. 11, No. 1, pp. 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998), “Asequencing method based on real-time…”). pyrophosphate, Science, Vol. 281, No. 5375, p. 363; U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320 (the full text of which is incorporated herein by reference). In pyrosequencing, the released PPi can be detected by being immediately converted to ATP by adenosine triphosphate (ATP) sulfatase, and the level of ATP produced can be detected by photons generated by luciferase. The nucleic acid to be sequenced can be attached to a feature in an array, and the array can be imaged to capture the chemiluminescent signal generated by the incorporation of nucleotides at the feature in the array. Images can be obtained after the array is treated with a specific type of nucleotide (e.g., A, T, C, or G). The images obtained after adding each type of nucleotide will differ in which feature in the array is detected. These differences in the images reflect the different sequence contents of the feature on the array. However, the relative position of each feature will remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after processing the array with each different nucleotide type can be processed in the same way as illustrated in this paper for images obtained from different detection channels used in reversible terminator-based sequencing methods.
[0195] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, cleavable or photobleachable dye labels, as described in, for example, WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method was commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators (where not only is the termination reversible, but the fluorescent label is cleavable) facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0196] Preferably, in a reversible terminator-based sequencing embodiment, the marker does not substantially inhibit elongation under SBS reaction conditions. However, the detection marker can be removable, for example, by cleavage or degradation. Images can be captured after the marker is incorporated into the arrayed nucleic acid signature. In a particular embodiment, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally distinct marker. Four images are then obtained, each using a detection channel selective for one of the four different markers. Alternatively, different nucleotide types can be added sequentially, and images of the array can be obtained between each addition step. In such embodiments, each image will show a nucleic acid signature incorporating a specific type of nucleotide. Different signatures may or may not be present in different images due to the different sequence contents of each signature. However, the relative positions of the signatures will remain unchanged in the images. Images obtained by such a reversible terminator-SBS method can be stored, processed, and analyzed as described herein. After the image capture step, the marker can be removed, and the reversible terminator portion can be removed for subsequent cycles of nucleotide addition and detection. Removing tags after they have been detected in a specific loop and before subsequent loops can reduce background signal and crosstalk between loops. Examples of available tagging and removal methods are illustrated below.
[0197] In certain specific embodiments, some or all of the nucleotide monomers may include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore linked to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other methods have separated the terminator chemistry from the cleavage of the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. described the development of reversible terminators that use a small 3' allyl group to block elongation but can be easily deblocked by a short treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker that can be easily cleaved by exposure to long-wavelength ultraviolet light for 30 seconds. Therefore, disulfide reduction or photocleavage can be used as the cleavable linker. Another method for reversible termination is to use natural termination, which occurs after a bulk dye is placed on the dNTP. The presence of a charged bulk dye on the dNTP can act as an efficient terminator due to steric and / or electrostatic hindrance. The presence of an incorporation event prevents further incorporation unless the dye is removed. The cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entirety.
[0198] Additional exemplary SBS systems and methods that can be used in conjunction with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100800, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entirety.
[0199] Some specific implementations may use fewer than four different labels to detect four different nucleotides. For example, the methods and systems described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232 may be used to perform SBS. As a first example, a pair of nucleotide types may be detected at the same wavelength, but distinguished based on the intensity difference of one member relative to the other member, or based on a change in the presence or absence of a signal in one member that results in a significantly different signal compared to the detected signal of the other member of that pair (e.g., by chemical modification, photochemical modification, or physical modification). As a second example, three of the four different nucleotide types may be detectable under specific conditions, while the fourth nucleotide type lacks a label that is detectable under those conditions or is minimally detectable under those conditions (e.g., minimal detection due to background fluorescence, etc.). The incorporation of the first three nucleotide types into the nucleic acid may be determined based on the presence of their respective signals, and the incorporation of the fourth nucleotide type into the nucleic acid may be determined based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type may include a label that is detected in two different channels, while other nucleotide types are detected in no more than one channel. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations. An exemplary specific implementation combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type lacking a label or minimally detected in either channel (e.g., unlabeled dGTP).
[0200] Additionally, as described in the incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In this so-called single-dye sequencing method, a first nucleotide type is labeled, but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.
[0201] Some specific implementations may utilize ligation-based sequencing (SBS) techniques. Such techniques utilize DNA ligases to incorporate oligonucleotides and recognize the incorporation of these oligonucleotides. Oligonucleotides typically have distinct labels associated with the identity of specific nucleotides in the sequence to which they hybridize. As with other SBS methods, images are obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show nucleic acid features incorporating a specific type of label. Due to the different sequence content of each feature, different features may be present or absent in different images, but the relative positions of the features will remain unchanged within the images. Images obtained by ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.
[0202] Some implementations utilize nanopore sequencing (Deamer, DW, and Akeson, M., “Nanopores and nucleic acids: prospects for ultrarapid sequencing.”, Trends Biotechnol. 18, 147-151 (2000); Deamer, D., and D. Branton, “Characterization of nucleic acids by nanopore analysis.”, Acc. Chem. Res. 35: 817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and JA Golovchenko, “DNA molecules and configurations in asolid-state nanopore microscope.”, Nat. Mater. 2: 611-615 (2003), the full text of which is incorporated herein by reference). In such implementations, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or a biomembrane protein, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the pore's conductivity. (US Patent 7,001,792; Soni, GV, and Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores.", Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis.", Nanod., 2,459-481 (2007); Cockroft, SL, Chu, J., Amorin, M., and Ghadiri, MR, "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.", J. Am. Chem. Soc. 130, 818-820 (2008), the full text of which is incorporated herein by reference). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein.Specifically, data can be processed like images, based on the exemplary processing of optical images and other images described herein.
[0203] Some specific implementations may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interaction between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414 (each of which is incorporated herein by reference), or by using a zero-mode waveguide, as described, for example, in U.S. Patent No. 7,315,019 (which is incorporated herein by reference), and by using fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). Illumination can be limited to a volume of approximately 1 / 2 psi around the surface-tethered polymerase, allowing observation of the incorporation of fluorescently labeled nucleotides against a low background (Levene, MJ et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.”, Science 299, 682-686 (2003); Lundquist, PM et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures.”, Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the full text of which is incorporated herein by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0204] Some specific implementations of SBS include detecting protons released during nucleotide incorporation into the extension product. For example, sequencing based on the detection of released protons can utilize electrical detectors and related technologies commercially available from Ion Torrent (a subsidiary of Guilford, CT, and Life Technologies) or the ion semiconductor sequencing method described in US 2009 / 0026082 A1; US 2009 / 0127589 A1; US 2010 / 0137143 A1; or US 2010 / 0282617 A1. ™ Sequencing was performed, and each of these references is incorporated herein by reference. The method described in this paper for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates for proton detection. More specifically, the method described in this paper can be used to generate amplicon clonal population for proton detection.
[0205] The SBS method described above can be advantageously performed in a variety of formats, enabling the simultaneous manipulation of multiple different target nucleic acids. In specific implementations, different target nucleic acids can be processed in a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In implementations using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically bind to the surface in a spatially distinguishable manner. Target nucleic acids can bind via direct covalent attachment, attachment to beads or other particles, or binding to polymerases or other molecules attached to the surface. The array can comprise a single copy of the target nucleic acid at each site (also referred to as a feature region), or multiple copies having the same sequence can be present at each site or feature region. Multiple copies can be generated by amplification methods such as, for example, bridging amplification or emulsion PCR, which are described in further detail below.
[0206] The method described herein can use an array of features having any one of a variety of densities, including, for example, at least about 10 features / cm². 2 100 feature parts / cm 2 500 features / cm 2 1,000 features / cm 2 5,000 features / cm 2 10,000 features / cm 2 50,000 features / cm 2 100,000 features / cm 2 1,000,000 features / cm 2 5,000,000 features / cm 2 Or higher.
[0207] The advantages of the methods described herein are that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Therefore, this disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Thus, the integrated system of this disclosure may include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc. A flow cell in the integrated system may be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference. As illustrated with respect to a flow cell, one or more fluid components of the integrated system may be used for both amplification and detection methods. Taking a nucleic acid sequencing implementation as an example, one or more fluid components of the integrated system may be used for the amplification methods described herein and for delivering sequencing reagents in sequencing methods, such as those exemplified above. Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of generating amplified nucleic acids and determining their sequences include, but are not limited to, MiSeq. ™ The platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Patent Serial No. 13 / 273,666 are incorporated herein by reference.
[0208] The sequencing system described above sequences nucleic acid polymers present in samples received by the sequencing equipment. As defined herein, “sample” and its derivatives are used in their broadest sense, including any specimen, culture, etc., suspected of containing a target. In some embodiments, a sample includes nucleic acids in the form of DNA, RNA, PNA, LNA, chimeric, or hybrid forms. A sample may include any biological, clinical, surgical, agricultural, atmospheric, or aquatic plant or animal specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, freshly frozen, or formalin-fixed paraffin-embedded nucleic acid specimens. It is also envisioned that a sample may originate from a single individual, or be a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, (matched) nucleic acid samples from a single individual such as tumor samples and normal tissue samples, or samples from a single source containing two different forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or samples containing plant or animal DNA with contaminating bacterial DNA. In some embodiments, the source of nucleic acid material may include nucleic acids obtained from newborns, such as those commonly used for newborn screening.
[0209] Nucleic acid samples may include high molecular weight substances, such as genomic DNA (gDNA). Samples may include low molecular weight substances, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another embodiment, low molecular weight substances include enzymatically fragmented or mechanically fragmented DNA. Samples may include cell-free circulating DNA. In some embodiments, the sample may include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microscopy, surgical resection, and other clinical or laboratory samples. In some embodiments, the sample may be an epidemiological sample, an agricultural sample, a forensic sample, or a pathogenic sample. In some embodiments, the sample may include nucleic acid molecules obtained from animals, such as humans or mammalian sources. In another embodiment, the sample may include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecules may be archived or extinct samples or species.
[0210] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples containing low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or may include forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel. The nucleic acid sample may be a purified sample or a lysate containing crude DNA, such as derived from oral swabs, paper, fabric, or other substrates that can be impregnated with saliva, blood, or other bodily fluids. Thus, in some embodiments, the nucleic acid sample may contain a small amount of DNA, such as genomic DNA, or fragmented portions of DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from the victim's hair, skin, tissue samples, autopsy, or remains. In some embodiments, the nucleic acid including one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence may include nucleic acids obtained from non-human DNA, such as microbial, plant, or insect DNA. In some embodiments, the target sequence or amplified target sequence is for human identification purposes. In some embodiments, this disclosure relates in its entirety to methods for identifying the characteristics of forensic samples. In some embodiments, this disclosure relates in its entirety to methods for human identification using one or more target-specific primers disclosed herein or designed using one or more target-specific primers with the primer design standards outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence may be amplified using any one or more target-specific primers disclosed herein or using the primer standards outlined herein.
[0211] Components of the base detection and correction system 106 may include software, hardware, or both. For example, components of the base detection and correction system 106 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., server device 102, client device 110, and / or sequencing device 114). When executed by one or more processors, the computer-executable instructions of the base detection and correction system 106 cause the computing device to perform the base detection and correction method described herein. Alternatively, components of the base detection and correction system 106 may include hardware, such as a dedicated processing device for performing a function or group of functions. Additionally or alternatively, components of the base detection and correction system 106 may include a combination of computer-executable instructions and hardware.
[0212] Furthermore, components of the base detection and correction system 106 that perform the functions described herein with respect to the base detection and correction system 106 may, for example, be implemented as part of a standalone application, a module of an application, a plug-in to an application, one or more library functions detectable by other applications, and / or a cloud computing model. Thus, components of the base detection and correction system 106 may be implemented as part of a standalone application on a personal computing device or mobile device. Additionally or alternatively, components of the base detection and correction system 106 may be implemented in any application that provides sequencing services, including but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumina," "BaseSpace," "DRAGEN," and "TruSight" are registered trademarks or trademarks of Illumina, Inc., in the U.S. and / or other countries.
[0213] As discussed in more detail below, embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware such as, for example, one or more processors and system memory. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more processes described herein.
[0214] Computer-readable media can be any available medium accessible by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of this disclosure may include at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0215] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drive (SSD) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that is accessible by a general-purpose or special-purpose computer.
[0216] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately considers that connection as a transmission medium. The transmission medium may include networks and / or data links that can be used to carry desired program code in the form of computer-executable instructions or data structures, and which are accessible to general-purpose or special-purpose computers. Combinations of the above should also be included within the scope of computer-readable media.
[0217] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., NIC) and then eventually transferred to the computer system RAM and / or to a less transient computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0218] Computer-executable instructions include, for example, instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a function or group of functions. In some embodiments, the computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary numbers, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or actions. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims.
[0219] Those skilled in the art will understand that this disclosure can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframes, mobile phones, PDAs, tablets, pagers, routers, switches, and so on. This disclosure can also be practiced in distributed system environments, where tasks are performed by both local and remote computer systems linked via a network (whether via a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links). In a distributed system environment, program modules can reside on both local and remote memory storage devices.
[0220] The embodiments disclosed herein can also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then thus scaled.
[0221] Cloud computing models can be composed of various characteristics, such as on-demand self-service, widespread network access, resource pooling, rapid elasticity, and metered services. Cloud computing models can also exhibit various service models, such as Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). Cloud computing models can also be deployed using different deployment models, such as private clouds, community clouds, public clouds, and hybrid clouds. In this specification and in the claims, a "cloud computing environment" means an environment in which cloud computing is employed.
[0222] Figure 13 A block diagram illustrating a computing device 1300 that can be configured to perform one or more of the processes described above is shown. It will be understood that one or more computing devices, such as computing device 1300, can implement the base detection and correction system 106. Figure 13 As shown, computing device 1300 may include a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310, which can be communicatively coupled via communication infrastructure 1312. In some embodiments, computing device 1300 may include a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310. Figure 13 Those shown have fewer or more components. The following paragraphs describe them in more detail. Figure 13 The components of the computing device 1300 shown.
[0223] In one or more embodiments, processor 1302 includes hardware for executing instructions such as those constituting a computer program. By way of example and not limitation, in order to execute instructions for dynamically modifying a workflow, processor 1302 may retrieve (or obtain) instructions from internal registers, internal cache, memory 1304, or storage device 1306, and decode and execute them. Memory 1304 may be volatile or non-volatile memory for storing data, metadata, and programs executed by the processor. Storage device 1306 includes storage means for storing data or instructions for performing the methods described herein, such as a hard disk, flash drive, or other digital storage device.
[0224] I / O interface 1308 allows a user to provide input to computing device 1300, receive output from the computing device, and otherwise transmit and receive data to and from the computing device. I / O interface 1308 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O interfaces. I / O interface 1308 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O interface 1308 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be servicing a particular implementation.
[0225] Communication interface 1310 may include hardware, software, or both. In any case, communication interface 1310 may provide one or more interfaces for communication (such as, for example, packet-based communication) between computing device 1300 and one or more other computing devices or networks. By way of example and not limitation, communication interface 1310 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi.
[0226] Additionally, the communication interface 1310 facilitates communication with various types of wired or wireless networks. The communication interface 1310 also facilitates communication using various communication protocols. The communication infrastructure 1312 may also include hardware, software, or both, that couples components of the computing device 1300 to each other. For example, the communication interface 1310 may use one or more networks and / or protocols to enable multiple computing devices connected through a specific infrastructure to communicate with each other to perform one or more aspects of the processes described herein. For illustration, a sequencing process may allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0227] In the foregoing description, this disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of this disclosure have been described with reference to the details discussed herein, and various embodiments are illustrated in the accompanying drawings. The above description and figures are illustrative of this disclosure and should not be construed as limiting it. Numerous specific details have been described to provide a thorough understanding of various embodiments of this disclosure.
[0228] This disclosure may be embodied in other specific forms without departing from its substance or essential characteristics. The embodiments described herein should be considered in all respects as exemplary rather than restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other or in parallel with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description. All modifications within the equivalent meaning and scope of the claims are to be included within their scope.
Claims
1. A system comprising: At least one processor; as well as A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: Based on read data of oligonucleotide clusters during sequencing runs, a set of bases detected that are associated with sequence-specific errors and the intensity data of said set of bases detected are determined; The intensity data of the set of bases detected in connection with the sequence-specific error is written into a storage device; The strength data written to the storage device is used to determine one or more base detection corrections for the set of bases using a correction model; as well as One or more corrected base detections are generated for the read data of the oligonucleotide cluster based on the one or more base detection corrections.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the set of base detections associated with the sequence-specific error by: detecting a base detection sequence indicating a homopolymer present in the oligonucleotide cluster, indicating a repeating dinucleotide sequence present in the oligonucleotide cluster, or indicating a repeating trinucleotide sequence present in the oligonucleotide cluster.
3. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to determine the one or more base detection corrections for the set of base detections using the correction model based on the intensity data written to the storage device in such a way that: the one or more base detection corrections for the set of base detections are determined using a phase correction model, the phase correction model performing per-cluster phase correction based on the intensity data written to the storage device.
4. The system of claim 3, further comprising instructions, which, when executed by the at least one processor, cause the system to determine the one or more base detection corrections for the set of bases using the phase correction model in such a way as to perform the per-cluster phase correction via a linear equalizer model or a maximum likelihood sequence estimator model.
5. The system of claim 3, further comprising instructions that, when executed by the at least one processor, cause the system to determine one or more per-cluster phasing coefficients via the phasing correction model to perform the per-cluster phasing correction using the intensity data of a set of read positions, the set of read positions ranging from the read position within the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base detection corrections for the set of base detections using the correction model in such a way as to determine base detection corrections for a set of read positions using the correction model, the set of read positions ranging from read positions following the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more base detection corrections for the set of base detections using the correction model in such a way as to determine base detection corrections for a set of read positions using the correction model, the set of read positions ranging from the read position within the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
8. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate one or more quality scores for the one or more corrected base detections of the read data of the oligonucleotide cluster.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to detect bases known to cause sequence-specific errors within nucleotide reads of the oligonucleotide cluster.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to write the intensity data of the set of bases detected in association with the sequence-specific error to a storage device by writing the intensity data to the non-transitory computer-readable medium or the transient computer-readable medium.
11. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to: Based on read data of oligonucleotide clusters during sequencing runs, a set of bases detected that are associated with sequence-specific errors and the intensity data of said set of bases detected are determined; The intensity data of the set of bases detected in connection with the sequence-specific error is written into a storage device; The strength data written to the storage device is used to determine one or more base detection corrections for the set of bases using a correction model; as well as One or more corrected base detections are generated for the read data of the oligonucleotide cluster based on the one or more base detection corrections.
12. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the set of base detections associated with the sequence-specific error by: detecting a base detection sequence indicating a homopolymer present in the oligonucleotide cluster, indicating a repeating dinucleotide sequence present in the oligonucleotide cluster, or indicating a repeating trinucleotide sequence present in the oligonucleotide cluster.
13. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine, based on the intensity data written to the storage device, the one or more base detection corrections for the set of base detections using the correction model in such a way that: the one or more base detection corrections for the set of base detections are determined using a phase-fixing correction model, the phase-fixing correction model performing per-cluster phase-fixing correction based on the intensity data written to the storage device.
14. The non-transitory computer-readable medium of claim 13, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the one or more base detection corrections for the set of bases using the phase correction model in such a way as to perform the per-cluster phase correction via a linear equalizer model or a maximum likelihood sequence estimator model.
15. The non-transitory computer-readable medium of claim 13, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine one or more per-cluster phasing coefficients via the phasing correction model to perform the per-cluster phasing correction using the intensity data of a set of read positions, the set of read positions ranging from the read position within the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
16. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine, using the correction model, the one or more base detection corrections for the set of base detections in such a way as to determine, using the correction model, a set of base detection corrections for a set of read positions, the set of read positions ranging from read positions following the base detection sequence that caused the sequence-specific error to the final read position of the nucleotide read of the oligonucleotide cluster.
17. A method, the method comprising: Based on read data of oligonucleotide clusters during sequencing runs, a set of bases detected that are associated with sequence-specific errors and the intensity data of said set of bases detected are determined; The intensity data of the set of bases detected in connection with the sequence-specific error is written into a storage device; The strength data written to the storage device is used to determine one or more base detection corrections for the set of bases using a correction model; as well as One or more corrected base detections are generated for the read data of the oligonucleotide cluster based on the one or more base detection corrections.
18. The method of claim 17, wherein determining the set of base detections associated with the sequence-specific error comprises detecting a base detection sequence indicating a homopolymer present in the oligonucleotide cluster, indicating a repeating dinucleotide sequence present in the oligonucleotide cluster, or indicating a repeating trinucleotide sequence present in the oligonucleotide cluster.
19. The method of claim 17, wherein determining the one or more base detection corrections for the set of base detections using the correction model based on the intensity data written to the storage device includes determining the one or more base detection corrections for the set of base detections using a phase correction model, the phase correction model performing per-cluster phase correction based on the intensity data written to the storage device.
20. The method of claim 19, wherein determining the one or more base detection corrections for the set of bases using the phase correction model comprises determining the one or more base detection corrections for the set of bases using the phase correction model to perform the per-cluster phase correction via a linear equalizer model or a maximum likelihood sequence estimator model.
Citation Information
Patent Citations
Stencil mask and method of producing the same
US20050100800A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1