Generating adaptive intensity value decision boundaries for nucleotide sequencing
The adaptive decision-boundary base-calling system enhances nucleobase call accuracy by generating artificial intensity values from genomic cycles to address low diversity and intensity shifts, improving sequencing efficiency and resource conservation.
Patent Information
- Application Number
- PCT/US2025/039737
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-17
- Filing Date
- 2025-07-29
- Publication Date
- 2026-02-05
AI Technical Summary
Existing sequencing systems struggle to accurately determine base-decision boundaries during indexing cycles due to low nucleobase-data diversity in index sequences and variations in intensity values, leading to incorrect nucleobase calls and increased resource consumption.
The adaptive decision-boundary base-calling system generates artificial intensity values based on expected base-specific intensity values from genomic sequencing cycles to supplement intensity values during indexing cycles, determining accurate intensity-value base-decision boundaries.
This approach improves the accuracy of nucleobase calls, conserves computing resources, and increases read-data throughput by correctly distinguishing index sequences, reducing the need for additional sequencing runs.
Smart Images

Figure US2025039737_05022026_PF_FP_ABST
Abstract
Description
GENERATING ADAPTIVE INTENSITY VALUE DECISION BOUNDARIES FOR NUCLEOTIDE SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 772,977, entitled, “GENERATING ADAPTIVE INTENSITY VALUE DECISION BOUNDARIES FOR NUCLEOTIDE SEQUENCING,” filed on March 17, 2025 (IP- 2864-PRV2) and U.S. Provisional Patent Application No. 63 / 677,206, entitled, “GENERATING ADAPTIVE INTENSITY VALUE DECISION BOUNDARIES FOR NUCLEOTIDE SEQUENCING,” filed on July 30, 2024 (IP-2864-PRV). Both of the aforementioned applications is hereby incorporated by reference in its entirety.BACKGROUND
[0002] In recent years, biotechnology firms and research institutions have improved hardware and software platforms used for determining a sequence of nucleotide bases (also referred to as “nucleobases”) in a sample. For instance, some existing sequencing devices and sequencing-data- analysis software (together “existing sequencing systems”) determine individual nucleobases of nucleic-acid sequences by using conventional Sanger sequencing or by using sequencing-by- synthesis (SBS). When using SBS, existing sequencing systems can monitor thousands, millions, or more nucleic-acid polymers being synthesized in parallel to detect more accurate nucleobase calls. For instance, a camera in SBS platforms can capture images of irradiated fluorescent tags from nucleotide bases incorporated into such synthesized nucleic-acid sequences (often grouped into clusters of oligonucleotides). After capturing the images, a computing device from the existing sequencing systems uses sequencing-data-analysis software to determine nucleobases that were detected in a given image based on the light signal (e.g., extracted and corresponding intensity values) captured in the image data. By iteratively incorporating nucleobases into the oligonucleotides and capturing images of the emitted light signals in various sequencing cycles, existing sequencing systems can determine the sequence of nucleobases present in the samples of nucleic acid.
[0003] To increase the efficiency and cost-effectiveness of such high-throughput sequencing technologies, many existing sequencing systems utilize multiplexing to simultaneously sequence multiple samples by tagging each sample with a unique sequence known as a barcode or index sequence. Index sequences, however, often exhibit relatively low nucleobase-data diversity compared to genomic sequences due to one or both of (i) the relative simplicity of index sequences compared to corresponding genomic sequences in terms of nucleobase content and (ii) the lack of a reference / quality control sequence (e.g., PhiX sequence) being base called during indexing cycles. Due at least in part to such a lack of nucleobase-data diversity in indexing cycles, existingsequencing systems often struggle to accurately determine base-decision boundaries for generating nucleobase calls in indexing cycles. In particular, existing sequencing systems often fail to accurately set detection thresholds for certain nucleobases in the absence of sufficient intensity data representing one or more of the four different candidate nucleobases.
[0004] Furthermore, variations in nucleotide-sample-slide (e.g., flow cell) preparations and sequencing equipment can cause the distribution of intensity values observed for one fluorescent nucleobase label to generally shift towards the intensity values of a second fluorescent label, such that the two different nucleobases are more difficult to distinguish from one another during basecalling analysis. For instance, intensity values observed for signals emitted by excited labels attached to cytosine can shift toward — and be difficult to distinguish from — intensity values observed for signals emitted by excited labels attached to adenine. While some existing sequencing systems can account for such shifts in intensity data during genomic sequencing runs comprising a diverse distribution of intensity values across all four candidate bases, many such systems further struggle to compensate for such shifts during indexing cycles due to the aforementioned lack of intensity data for one or more candidate nucleobases. When sequencing devices cannot effectively or do not fully compensate for such intensity-data shifts, the base calls generated by the sequencing devices can become less accurate.
[0005] Incorrect nucleobase calls by existing sequencing systems can result in a sequencing device expending additional computing resources and consumables in longer or additional sequencing runs to compensate for lack of read-data coverage in a given sample. Because incorrect base calls on index sequences can lead to a sequencing device’s failure in recognizing the index sequences or determining the correct sample to which the index sequence of a given oligonucleotide cluster belongs, existing sequencing systems that suffer from incorrectly base-called index sequences may need to discard the nucleotide reads to which such index sequences are attached for one or more samples. If such discarded nucleotide reads lead to read coverage of a genomic region of interest below a threshold (e.g., 20X, 30X), some existing sequencing systems must consume additional computing time, memory, and consumable materials (e.g., fluidic reagents, nucleobases for building nucleotide reads) to compensate for these under-covered genomic regions by performing additional (or top-off) sequencing runs.
[0006] These, along with additional problems and issues exist in existing sequencing systems.SUMMARY
[0007] This disclosure describes embodiments of methods, non-transitory computer-readable media, and systems that adaptively generate and implement artificial intensity values during indexing cycles of a sequencing run to determine intensity-value base-decision boundaries with increased accuracy over existing sequencing systems and methods. For example, the disclosedsystems can analyze base-specific intensity-value distributions of genomic intensity values from a set of genomic sequencing cycles within a sequencing run to determine one or more expected basespecific intensity values for the sequencing run. Based on the one or more expected base-specific intensity values, the disclosed systems can generate artificial base-specific intensity values to supplement intensity values for an indexing cycle of the sequencing run and determine, based at least in part on the artificial base-specific intensity values, intensity-value base-decision boundaries for generating nucleobase calls for the indexing cycle based on the respective intensity values.
[0008] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The detailed description refers to the drawings briefly described below.
[0010] FIG. 1 illustrates an environment in which an adaptive decision-boundary base-calling system can operate in accordance with one or more embodiments of the present disclosure.
[0011] FIG. 2 illustrates an overview of the adaptive decision-boundary base-calling system generating nucleotide reads for a nucleotide-sample slide in accordance with one or more embodiments of the present disclosure.
[0012] FIG. 3 illustrates the adaptive decision-boundary base-calling system determining and utilizing an adaptive intensity -value base-decision boundary to generate a nucleobase call in accordance with one or more embodiments of the present disclosure.
[0013] FIGS. 4A-4B illustrate the adaptive decision-boundary base-calling system generating artificial base-specific intensity values for indexing cycles based on expected base-specific intensity values determined from genomic sequencing cycles in accordance with one or more embodiments of the present disclosure.
[0014] FIG. 5 illustrates comparative experimental results of identifying an expected basespecific intensity value for sequencing cycles of a sequencing run utilizing (i) an existing sequencing system and (ii) the adaptive decision-boundary base-calling system in accordance with one or more embodiments of the present disclosure.
[0015] FIGS. 6A-6B, 7A-7B, and 8 illustrate comparative experimental results of determining nucleobase calls for multiplexed nucleotide-sample slides utilizing (i) an existing sequencing system and (ii) the adaptive decision-boundary base-calling system in accordance with one or more embodiments of the present disclosure.
[0016] FIG. 9 illustrates a flowchart of a series of acts for determining a nucleobase call for an indexing cycle in accordance with one or more embodiments of the present disclosure.
[0017] FIG. 10 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION
[0018] This disclosure describes embodiments of an adaptive decision-boundary base-calling system that can adaptively generate and implement artificial intensity values during indexing cycles of a sequencing run to determine intensity-value base-decision boundaries with increased accuracy over existing sequencing systems and methods. For example, the adaptive decision-boundary basecalling system can analyze base-specific intensity-value distributions of genomic intensity values from a set of genomic sequencing cycles or other sequencing cycles that precede certain indexing cycles within a sequencing run to determine one or more expected base-specific intensity values for the sequencing run. Based on the one or more expected base-specific intensity values, the adaptive decision-boundary base-calling system can (i) generate artificial base-specific intensity values to supplement intensity values received (or otherwise identified) for an indexing cycle of the sequencing run and (ii) determine, based at least in part on the artificial base-specific intensity values, intensity-value base-decision boundaries for generating nucleobase calls for one or more cycles of the indexing cycles based on the respective intensity values.
[0019] In some embodiments, for example, the adaptive decision-boundary base-calling system receives (or otherwise identifies), for an indexing cycle, intensity values corresponding to labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide comprising multiple samples respectively labeled with index sequences. Such index sequences may include fabricated unique sequences attached to oligonucleotides of each respective sample on the nucleotide-sample slide. As mentioned, to increase the quality and accuracy of intensity-value base-decision boundaries for nucleobase calling of indexing cycles and / or increase the accuracy of such nucleobase calls, the adaptive decision-boundary base-calling system can supplement the respective intensity values with artificial base-specific intensity values and determine the intensityvalue base-decision boundaries based at least in part on the artificial base-specific intensity values.
[0020] To facilitate generating nucleobase calls for an indexing cycle, the adaptive decisionboundary base-calling system can determine intensity -value base-decision boundaries that are specific to a particular nucleobase based at least in part on artificial base-specific intensity values implemented according to the disclosed embodiments. In some embodiments, the adaptive decision-boundary base-calling system determines intensity -value base-decision boundaries for a particular indexing cycle based on (i) base-specific distributions of intensity values for the particular indexing cycle corresponding to labeled nucleobases within the nucleotide-sample slide and (ii) the artificial base-specific intensity values generated based on the aforementioned set of genomic sequencing cycles.
[0021] Having determined intensity-value base-decision boundaries, the adaptive decisionboundary base-calling system can utilize these boundaries to generate a nucleobase call for a particular oligonucleotide cluster within the nucleotide-sample slide based on an intensity value corresponding to the particular oligonucleotide cluster. Furthermore, in some embodiments, the adaptive decision-boundary base-calling system utilizes the intensity -value base-decision boundaries to generate nucleobase calls for multiple oligonucleotide clusters within the nucleotide- sample slide based on the intensity values for the indexing cycle. After the indexing cycle, in some embodiments, the adaptive decision-boundary base-calling system determines intensity -value basedecision boundaries for multiple subsequent indexing cycles of a sequencing run and utilizes the respective intensity-value base-decision boundaries to determine nucleobase calls of index sequences comprising nucleobases identified from the multiple subsequent indexing cycles.
[0022] In addition or in the alternative to determining intensity -value base-decision boundaries for an indexing cycle, in some embodiments, the adaptive decision-boundary base-calling system can determine intensity-value base-decision boundaries for a genomic sequencing cycle. For example, the adaptive decision-boundary base-calling system can analyze base-specific intensityvalue distributions of intensity values from a set of genomic sequencing cycles or other sequencing cycles (e.g., indexing cycles) that precede certain genomic sequencing cycles within a sequencing run to determine one or more expected base-specific intensity values. Based on the one or more expected base-specific intensity values, the adaptive decision-boundary base-calling system can (i) generate artificial base-specific intensity values to supplement intensity values received (or otherwise identified) for a genomic sequencing cycle of the sequencing run and (ii) determine, based at least in part on the artificial base-specific intensity values, intensity-value base-decision boundaries for generating nucleobase calls for the genomic sequencing cycle based on the respective intensity values.
[0023] As mentioned, the adaptive decision-boundary base-calling system provides various advantages over existing sequencing systems. For instance, the adaptive decision-boundary basecalling system incorporates significant technical improvements to a special-purpose computing device — that is, a sequencing device — by improving the quality and accuracy of index sequencing and demultiplexing in multi-sample sequencing compared to existing sequencing systems. In contrast to existing sequencing systems, for example, the adaptive decision-boundary base-calling system can supplement intensity values in indexing cycles with adaptively generated and artificial intensity values, resulting in a significant increase to base calling accuracy during indexing cycles. By determining expected base-specific intensity values based on genomic-sequencing-cycle intensity values, further generating artificial base-specific intensity values based on such expected base-specific intensity values — and determining intensity-value base-decision boundaries thatcompensate for intensity-data shifts during indexing sequencing cycles — the adaptive decisionboundary base-calling system can determine accurate base calls for nucleobases within index sequences specific to oligonucleotide clusters of different samples of a sequencing run. Such accurately base-called index sequences facilitate the adaptive decision-boundary base-calling system in correctly distinguishing among (or demultiplexing) one set of index sequences for nucleotide reads of one sample and another set of index sequences for nucleotide reads of another sample.
[0024] In addition to and independent of improving the index sequencing and demultiplexing operations of a sequencing device, in some embodiments, the adaptive decision-boundary basecalling system improves the base-calling operations of the sequencing device. As suggested above and further discussed herein (e.g., in relation to FIG. 2), index sequences from multi-sample nucleotide-sample slides often exhibit relatively less nucleobase-data diversity relative to their corresponding genomic sequences determined in genomic sequencing cycles because of one or both of (i) the relative simplicity of index sequences compared to corresponding genomic sequences in terms of nucleobase content and (ii) the lack of a reference / quality control sequence (e.g., PhiX sequence) being base called during an indexing cycle. The relatively lower nucleobase- data diversity of an index sequence can depend on how many different samples are present in a multiplexed sequencing run because, for example, index sequences for different samples can result in different base calls in a given indexing cycle. Furthermore, as noted above, variations in nucleotide-sample-slide preparations and sequencing equipment can cause the distribution of intensity values observed throughout a sequencing run for one fluorescent nucleobase label (e.g., for cytosine) to generally shift towards the intensity values of another fluorescent nucleobase label (e.g., for adenine). By supplementing intensity values in indexing cycles to (i) compensate for relatively low nucleobase-data diversity inherent to index sequences and (ii) adjust for the drift or drop of intensity values often encountered during sequencing runs, the adaptive decision-boundary base-calling system can determine detection thresholds for certain nucleobases with increased accuracy relative to existing sequencing systems (e.g., as demonstrated by the experimental results shown in FIGS. 5-8). As described herein, the adaptive decision-boundary base-calling system determines, for an indexing cycle, intensity -value base-decision boundaries based on artificial basespecific intensity values derived from genomic or other sequencing cycles. Such improved, intensity-value-drift-compensating decision boundaries better demarcate different intensity values for difference nucleobases and thereby improve the accuracy of nucleobase calls generated by a sequencing device.
[0025] Beyond improved base-calling accuracy, in some embodiments, the adaptive decisionboundary base-calling system also conserves a sequencing device’s computing resources andconsumable materials — and increases quality read-data throughput for biological samples — relative to existing sequencing systems. By improving base-calling quality during indexing cycles, the adaptive decision-boundary base-calling system can significantly improve the reliability and amount of base-call data used for demultiplexing of sequenced samples, further resulting in a greater number of identifiable nucleotide reads for respective samples. Additionally, such resulting increase in nucleotide reads conclusively and reliably associated with their respective samples in turn can provide increased read depth (and other base-call data) for genomic regions of a particular sample. Because the number of nucleotide reads identifiable for samples increases and the readdata coverage of genomic regions of interest increases for samples, the adaptive decision-boundary base-calling system is less likely to require an additional (or top-off) sequencing run performed by some existing systems that would consume additional memory, processing, consumable materials, and biological genomic material for the sample (e.g., blood or spit) to compensate for the previous sequencing cycle’s discarded nucleotide reads or lack of read-data coverage. In addition to such conservation of computing-resource and biological-sample, the adaptive decision-boundary basecalling system’s higher quality and quantity of sample-specific reads also significantly improves downstream analyses, such as mapping, alignment, and variant calling, for the particular sample by increasing the quality of read alignments and the overall depth of nucleobase data.
[0026] As suggested by the foregoing discussion, this disclosure utilizes a variety of terms to describe features and benefits of the adaptive decision-boundary base-calling system. Additional detail is hereafter provided regarding the meaning of these terms as used in this disclosure. As used in this disclosure, for instance, the term “sample” refers to a specimen, culture, or the like that is suspected of including a target nucleic acid. In some embodiments, the sample comprises DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acids as targets. The sample can likewise include any biological, clinical, surgical, agricultural-atmospheric, or aquatic-based specimen containing one or more nucleic acids. A sample also includes any isolated or extracted nucleic acid sample from an organism, such a genomic DNA, fresh-frozen, or formalin-fixed paraffin-embedded nucleic acid specimen. In some cases, accordingly, a sample can include a full genome or partial genome that is isolated or extracted (e.g., in whole or in part by a kit) from an organism and that is prepared to undergo sequencing or an assay in a sequencing device. A sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In someembodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.
[0027] The sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0028] As further used herein, the term “nucleotide-sample slide” (or “nucleotide-sample substrate”) refers to a plate or substrate, such as a flow cell, comprising oligonucleotides for sequencing nucleotide sequences from samples or other sample nucleic-acid polymers. In particular, a nucleotide-sample slide can refer to a substrate containing fluidic channels through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, a flow cell (e.g., a patterned flow cell or non-pattemed flow cell) may comprise small fluidic channels and oligonucleotide samples that can be bound to adapter sequences on the substrate. In other implementations, a nucleotide-sample slide can be an open substrate with one or more regions for oligonucleotide samples to be analyzed and the oligonucleotide samples may be positioned using charged pads or other means. In yet another implementation, the nucleotide- sample slide can be a membrane having a nanopore through which one or more oligonucleotide samples may pass.
[0029] As used herein, a flow cell or other nucleotide-sample slide can (i) include a device having a lid extending over a reaction structure to form a flow channel therebetween that is in communication with a plurality of reaction sites of the reaction structure and (ii) include a detection device that is configured to detect designated reactions that occur at or proximate to the reaction sites. A flow cell or other nucleotide-sample slide may include a solid-state light detection or “imaging” device, such as a Charge-Coupled Device (CCD) or Complementary Metal-Oxide Semiconductor (CMOS) (light) detection device. As one specific example, a flow cell may be configured to fluidically and electrically couple to a cartridge (having an integrated pump), whichmay be configured to fluidically and / or electrically couple to a bioassay system. A cartridge and / or bioassay system may deliver a reaction solution to reaction sites of a flow cell according to a predetermined protocol (e.g., sequencing-by-synthesis), and perform a plurality of imaging events. For example, a cartridge and / or bioassay system may direct one or more reaction solutions through the flow channel of the flow cell, and thereby along the reaction sites. At least one of the reaction solutions may include four types of nucleobases having the same or different fluorescent labels. The nucleobases may bind to the reaction sites of the flow cell, such as to corresponding oligonucleotides at the reaction sites. The cartridge and / or bioassay system may then illuminate the reaction sites using an excitation light source (e.g., solid-state light sources, such as light-emitting diodes (LEDs)). The excitation light may provide emission signals (e.g., light of a wavelength or wavelengths that differ from the excitation light and, potentially, each other) that may be detected by the light sensors of the flow cell.
[0030] In addition, as used herein, the term “cluster of oligonucleotides” (or simply “cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-pattemed flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.
[0031] Additionally, as used herein, the term “labeled nucleobase” (or “labeled nucleotide base”) refers to a nucleobase having a fluorescent or light-based indicator or fluorescent dye indicator of a classification of the nucleobase. In particular, a labeled nucleobase can refer to a nucleobase that incorporates a fluorescent or light-based indicator or fluorescent dye indicator to identify the type of base (e.g., adenine, cytosine, thymine, or guanine). For example, in one or more embodiments, a labeled nucleobase includes a nucleobase having a fluorescent tag that emits a signal that either by itself or together with another fluorescent tag identifies the base type. Accordingly, a nucleobase may be identified by a mixture of dyes (or a mixture of fluorescent tags) that together indicate the nucleobase type (e.g., “ON” / “ON” expected response signals and / or estimated response signals). Based on intensity values for a signal emitted by labeled nucleobases in a cluster of oligonucleotides, such as signals in 16 quadrature amplitude modulation (QAM) orpulse amplitude modulation (PAM) 4 format, the type of base (e.g., adenine, cytosine, thymine, or guanine) can be determined in certain embodiments.
[0032] As further used herein, the term “nucleobase call” (or simply “base call”) refers to a determination or prediction of a particular nucleobase (or nucleobase pair) for an oligonucleotide (e.g., nucleotide read) during a sequencing cycle or for a genomic coordinate of a sample. In particular, a nucleobase call can indicate a determination or prediction of the type of nucleobase that has been incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., read-based nucleobase calls). In some cases, for a nucleotide read, a nucleobase call includes a determination or a prediction of a nucleobase based on intensity values resulting from fluorescent-tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., in a cluster of a flow cell). As suggested above, a single nucleobase call can be an adenine (A) call, a cytosine (C) call, a guanine (G) call, a thymine (T) call, or an uracil (U) call. Note that the terms nucleobase and nucleotide base are interchangeable.
[0033] As further used herein, the term “sample genomic sequence” refers to a nucleotide sequence extracted from, copied from, or complementary to a sample’s chromosome. For example, a sample genomic sequence includes a nucleotide sequence that has been separated or copied from chromosomal DNA of a sample or has been sequenced to be complementary to an extracted or copied nucleotide sequence. Accordingly, a sample genomic sequence includes genomic DNA (gDNA) for a particular unknown sample. Accordingly, as described herein, in some embodiments, the target-sequence-coverage system can use a sample complementary sequence comprising cDNA rather than a sample genomic sequence comprising gDNA in a sample library fragment or wherever suitable cDNA may replace gDNA as understood by a skilled artisan. Indeed, any embodiment or nucleotide read in this disclosure that uses or includes a sample genomic sequence can also use or include a cDNA sequence corresponding to a genomic sample.
[0034] Also, as used herein, the term “nucleotide read” (or simply “read”) refers to an inferred or predicted sequence of one or more nucleobases (or nucleobase pairs) from all or part of a sample genomic sequence (e.g., a sample genomic sequence, complementary DNA). In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a nucleotide fragment (or group of monoclonal nucleotide fragments) from a sequencing library corresponding to a sample. For example, in some embodiments, the adaptive decision-boundary base-calling system determines a nucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a well in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can referto (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g., assembled nucleotide reads) satisfying a threshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.
[0035] As used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample. In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device (including an imaging device, such as a CCD or CMOS) that incorporate nucleobases into growing oligonucleotides to determine nucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more samples seeded in clusters throughout a flow cell. Upon completing a sequencing run, a sequencing device can generate base-call data in a file, such as a binary base call (BCL) sequence file or a fast- all quality (FASTQ) file.
[0036] As used herein, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample’s sequence (e.g., a genomic or transcriptomic sequence from a sample) or a corresponding adapter sequence. In some cases, a sequencing cycle includes an iteration of both incorporating nucleobases into clusters of oligonucleotides using sequencing chemistry and capturing images of such clusters attached to a nucleotide-sample slide (e.g., a flow cell). Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer (e.g., a sample genomic sequence). For example, in one or more embodiments, each sequencing cycle involves incorporating nucleobases into either a single nucleotide read in which DNA or RNA strands are read in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends but in different cycles. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining a particular nucleobase added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleobases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within an SBS run. A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For instance, one cluster of oligonucleotides or a set of clusters of oligonucleotides may be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated and another cluster of oligonucleotides or another set of clusters of oligonucleotides may be concurrently undergoing anindexing cycle in which nucleobases corresponding to an index sequence for a nucleotide read are incorporated.
[0037] As further used herein, the term “genomic sequencing cycle” refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample genomic sequence (or cDNA sequence). In particular, a genomic sequencing cycle can include an iteration of capturing and analyzing one or more images with data indicating individual nucleobases added or incorporated into an oligonucleotide or to oligonucleotides (in parallel) representing or corresponding to one or more sample genomic sequences. Such image analysis can include analyzing data from signals output from an image sensor (e.g., an area capture sensor or a time delayed integration (TDI) sensor). For example, in one or more embodiments, each genomic sequencing cycle involves capturing and analyzing images to determine either single reads or paired-end reads of DNA (or RNA) strands representing part of a sample (or transcribed sequence from a sample). As suggested above, however, a genomic sequencing cycle, in some cases, is specific to a cluster of oligonucleotides or a set of clusters of oligonucleotides.
[0038] By contrast, the term “indexing cycle” refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to one or more index sequences. In particular, an indexing cycle can include an iteration of capturing and analyzing one or more images of clusters of oligonucleotides indicating one or more nucleobases added or incorporated into an oligonucleotide or to oligonucleotides (in parallel) representing or corresponding to one or more index sequences. An indexing cycle differs from a genomic sequencing cycle in that an indexing cycle includes sequencing of at least a nucleobase (or a majority of nucleobases) from one or more index sequences that identify or encode one or more sample library fragments. During a given sequencing run for a nucleotide-sample slide comprising clusters of oligonucleotides for one or more samples, an indexing cycle for one cluster of oligonucleotides on the nucleotide-sample slide may be performed at a same time as a genomic sequencing cycle for another cluster of oligonucleotides on the nucleotide-sample slide.
[0039] Relatedly, as used herein, the term “demultiplexing” refers to a process of sorting nucleotide reads generated during multi-sample sequencing into individual sample data sets based on unique index sequences or barcodes assigned to (or otherwise associated with) individual samples. To illustrate, during preparation of a nucleotide-sample slide, each sample is assigned with one or more unique index sequences. After or during a sequencing run, a sequencing system reads the index sequences in conjunction with the actual sequence data to assign each nucleotide read to its corresponding sample. In some cases, nucleotide reads having inconclusive indexsequences (e.g., due to poor base-call quality) are discarded due to the inability to directly associate such reads with an individual sample.
[0040] Further, as used herein, the term “signal” refers to a signal emitted, reflected, or otherwise communicated from a labeled nucleobase or a group of labeled nucleobases (e.g., labeled nucleobases added to a cluster of oligonucleotides). In particular, a signal can refer to a signal indicating the type of base. For example, a signal can include a light signal emitted or reflected from a fluorescent tag of a nucleotide base or fluorescent tags of multiple nucleotide bases incorporated into oligonucleotides. In some implementations, the adaptive decision-boundary basecalling system triggers the signal through an external stimulus, such as a laser or other light source. In some cases, the adaptive decision-boundary base-calling system triggers the signal through some internal stimuli. Further, in some embodiments, the adaptive decision-boundary base-calling system observes the signal using a filter applied when capturing an image of the nucleotide-sample slide (e.g., section of the nucleotide-sample slide). As suggested above, in certain instances, a signal includes an aggregate of the signals provided by each labeled nucleobase added to individual oligonucleotides in a cluster of oligonucleotides.
[0041] As used herein, the term “sequencing signal channel” (or simply “channel”) refers to a range or filter of light, intensity, or color used to detect and / or measure a signal from a cluster of oligonucleotides. For example, a sequencing signal channel can include a particular range of light, intensity, or color of a laser used to illicit a fluorescent signal from fluorescent tags on nucleobases incorporated into oligonucleotides within a cluster. In some embodiments, the adaptive decisionboundary base-calling system utilizes a two-channel implementation by, for instance, using two different ranges of light, intensities, or colors to illicit signals from clusters per sequencing cycle and capturing two corresponding images of a region of a nucleotide-sample slide per sequencing cycle. The first and second images can capture the intensity values of the emitted signal from the clusters that correspond to first and second light ranges. In some embodiments, the adaptive decision-boundary base-calling system can utilize a single-channel implementation, three-channel implementation, or four-channel implementation.
[0042] As used herein, the term “intensity value” refers to a value indicating a characteristic or attribute of a signal emitted, reflected, or otherwise communicated from a labeled nucleobase or a group of labeled nucleobases from a cluster of oligonucleotides. In particular, an intensity value can refer to a value associated with a color intensity (e.g., wavelength) or a light intensity (e.g., brightness). In some cases, the adaptive decision-boundary base-calling system captures several images of a cluster of oligonucleotides with labeled nucleobases using different sequencing signal channels. Thus, an intensity value of a signal can correspond to the intensity of the signal as observed through a particular sequencing signal channel. In one or more embodiments, the intensityvalue is a measured degree of intensity for a cluster of oligonucleotides at the predicted location, and the adaptive decision-boundary base-calling system can accordingly be applied to 16 quadrature amplitude modulation (QAM) modulation or pulse amplitude modulation (PAM) 4 modulation (e.g., using amplitude to encode base-call information).
[0043] Additionally, as used herein, the term “expected intensity value” refers to an intensity value that is expected to represent or otherwise correspond to a particular nucleobase (A, C, G, T). In particular, an expected intensity value is an intensity value derived from at least one sequencing cycle (e.g., a genomic sequencing cycle or a preceding indexing cycle) that represents an expected value for a nucleobase (A, C, G, T) in another sequencing cycle (e.g., a target indexing cycle). For instance, in some cases, the expected intensity value refers to an average of (or centroid for) intensity values associated with a signal for a nucleobase in a particular sequencing signal channel or a particular combination of sequencing signal channels. In certain implementations, the expected intensity value is an average of intensity values falling within the base-specific intensity-value distribution (e.g., nucleotide clouds) of a certain base (A, C, G, or T). As further suggested above, in certain implementations, an expected intensity value comprises a centroid value of the intensityvalue distributions of one or more sequencing cycles. In some embodiments, the expected intensity value is the same for all clusters of oligonucleotides within the region.
[0044] Additionally, as used herein, the term “intensity-value boundaries” refers to decision boundaries used in generating a nucleobase call for a signal. In particular, intensity -value boundaries can refer to decision boundaries that classify a nucleotide base (e.g., as A, T, C, or G) based on one or more intensity values of the signal. To illustrate, intensity-value boundaries can define or otherwise indicate the boundaries of a nucleotide cloud corresponding to each of the nucleobases or nucleobase types. In some implementations, intensity -value boundaries do not mark the limits at which a signal is classified as a nucleotide base, but rather a point at which the signal can be classified as the nucleotide base with a particular level of accuracy.
[0045] As used herein, the term “base-call-distribution model” refers to a computer model or algorithm that generates intensity -value boundaries. For example, in some implementations, a base- call-distribution model includes, but is not limited to, a Gaussian distribution model, a uniform distribution model, a Bernoulli distribution model, a binomial distribution model, or a Poisson distribution model.
[0046] As used herein, the term “centroid” refers to a center or representative value or data point of a nucleotide cloud defined or otherwise indicated by one or more boundaries (e.g., intensity-value base-decision boundaries). In particular, a centroid includes intensity values that have been averaged or geometrically determined to be representative of a particular base type (e.g., A, T, C, G) over multiple sequencing cycles. Relatedly, as used herein, the term “centroid intensityvalue” refers to an intensity value associated with a centroid. In particular, a centroid intensity value indicates an intensity value that corresponds to the center of a nucleotide cloud.
[0047] As further used herein, the term “base-call data file” refers to a digital file or other digital information indicating individual nucleobases or the sequence of nucleobases for a nucleic- acid polymer. In particular, a base-call data file can include nucleotide reads comprising nucleobase calls for particular samples. Base-call data files can include intensity values (e.g., color or light intensity values for individual clusters) from images taken by a camera of a nucleotide-sample slide or other data that indicate individual nucleobases or the sequence of nucleobases for a nucleic-acid polymer. In addition, or in the alternative to intensity values, a base-call data file may include chromatogram peaks or electrical current changes indicating individual nucleobases in a sequence. Additionally, in some embodiments, a base-call data file includes individual nucleobase calls identifying the individual nucleobases (e.g., A, T, C, or G). For example, a base-call data file can comprise data for nucleobase calls in a sequence for a nucleic-acid polymer, the number of nucleobase calls corresponding to a particular base (e.g., adenine, cytosine, thymine, or guanine), as organized in a digital file, such as a Binary Base Call (BCL) file or a Fast-All Q (FASTQ) file. The format of the base-call data file can vary based upon the sequencing technology used and can include BCF, BAM, and QSEQ, as well as other formats. Further, a base-call data file can include error / accuracy information, such as a quality metric associated with each nucleobase call. In some embodiments, the base-call data comprises information from a sequencing device that utilizes sequencing by synthesis (SBS).
[0048] The following paragraphs describe the adaptive decision-boundary base-calling system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which an adaptive decision-boundary base-calling system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server device(s) 110, and a client device 114. As shown in FIG. 1, the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114 can communicate with each other via a network 118. The network 118 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 10. While FIG. 1 shows an embodiment of the adaptive decision-boundary base-calling system 106, this disclosure describes alternative embodiments and configurations below.
[0049] As indicated by FIG. 1, the sequencing device 102 comprises a computing device and a sequencing device system 104 for sequencing DNA from a sample or other nucleic-acid polymer. In some embodiments, by executing the sequencing device system 104 using a processor, thesequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 102. More particularly, the sequencing device 102 receives nucleotide-sample slides (e.g., flow cells) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments.
[0050] In one or more embodiments, the sequencing device 102 utilizes sequencing-by- synthesis (SBS) techniques to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition or in the alternative to communicating across the network 118, in some embodiments, the sequencing device 102 bypasses the network 118 and communicates directly with the local device 108, the client device 114, and / or the server device(s) 110. By executing the sequencing device system 104, the sequencing device 102 can further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the local device 108, the server device(s) 110, and / or the client device 114.
[0051] As further indicated by FIG. 1, the local device 108 is located at or near a same physical location of the sequencing device 102. Indeed, in some embodiments, the local device 108 and the sequencing device 102 are integrated into a same computing device. The local device 108 may run the sequencing device system 104 and / or the adaptive decision-boundary base-calling system 106 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As shown in FIG. 1, the sequencing device 102 may send (and the local device 108 may receive) base-call data generated during a sequencing run of the sequencing device 102. The local device 108 may also communicate with the client device 114. In particular, the local device 108 can send data to the client device 114, including a binary alignment map (BAM) file, a variant call format (VCF) file, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0052] As further indicated by FIG. 1, the server device(s) 110 are located remotely from the local device 108 and the sequencing device 102. Similar to the local device 108, in some embodiments, the server device(s) 110 include a version of (or are otherwise able to access or implement) the adaptive decision-boundary base-calling system 106. For example, the server device(s) 110 can implement the adaptive decision-boundary base-calling system 106 as part of a sequencing system 112. Accordingly, the server device(s) 110 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining genotype or variant calls based on analyzing such base-call data. As indicated above, the sequencing device 102 may send (and the server device(s) 110 may receive) base-call data from the sequencing device 102. The server device(s) 110 may also communicate with the client device 114. In particular, theserver device(s) 110 can send data to the client device 114, including BAM files, VCF fdes, or other sequencing related information.
[0053] In some embodiments, the server device(s) 110 comprise a distributed collection of servers where the server device(s) 110 include a number of server devices distributed across the network 118 and located in the same or different physical locations. Further, the server device(s) 110 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server. Moreover, while not shown in FIG. 1, the server device(s) 110 can also be in communication, either directly or via the network 118, with a database storing, among other things, nucleobase-call data associated with one or more sequencing runs as generated by the adaptive decision-boundary base-calling system 106.
[0054] As further illustrated and indicated in FIG. 1, by executing a sequencing application 116, the client device 114 can generate, store, receive, and send digital data. In particular, the client device 114 can receive sequencing data from the local device 108 or receive call files (e.g., BCL) and sequencing metrics from the sequencing device 102. Furthermore, the client device 114 may communicate with the local device 108 or the server device(s) 110 to receive a VCF comprising genotype or variant calls and / or other metrics, such as base-call-quality metrics or pass-filter metrics. The client device 114 can accordingly present or display information pertaining to nucleobase calls, variant or genotype calls, or other sequencing information within a graphical user interface of the sequencing application 116 to a user associated with the client device 114. For example, the client device 114 can present nucleobase calls, genotype calls, variant calls, and / or sequencing metrics for a sequenced sample within a graphical user interface of the sequencing application 116.
[0055] Although FIG. 1 depicts the client device 114 as a desktop or laptop computer, the client device 114 may comprise various types of client devices. For example, in some embodiments, the client device 114 includes non -mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 114 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 114 are discussed below with respect to FIG. 10.
[0056] As further illustrated in FIG. 1, the client device 114 includes the sequencing application 116. The sequencing application 116 may be a web application or a native application stored and executed on the client device 114 (e.g., a mobile application, desktop application). The sequencing application 116 can include instructions that (when executed) cause the client device 114 to receive data from the adaptive decision-boundary base-calling system 106 and present, for display at the client device 114, base-call data or data from an alignment data file or VCF.Furthermore, the sequencing application 116 can instruct the client device 114 to display summaries for multiple sequencing runs.
[0057] As further illustrated in FIG. 1, a version of the adaptive decision-boundary basecalling system 106 may be located and / or implemented (e.g., entirely or in part) on the client device 114 or the sequencing device 102. In yet other embodiments, the adaptive decision-boundary basecalling system 106 is implemented by one or more other components of the computing system 100, such as the local device 108. In particular, the adaptive decision-boundary base-calling system 106 can be implemented in a variety of different ways across the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114. For example, the adaptive decisionboundary base-calling system 106 can be downloaded from the server device(s) 110 to the client device 114 and / or the local device 108 where all or part of the functionality of the adaptive decisionboundary base-calling system 106 is performed at each respective device within the computing system 100.
[0058] As indicated above, the adaptive decision-boundary base-calling system 106 can generate and implement adaptive intensity-value base-decision boundaries to more accurately determine index sequences on a nucleotide-sample slide processed by a sequencing device. In accordance with one or more embodiments, FIG. 2 illustrates an overview of the adaptive decisionboundary base-calling system 106 (i) receiving a nucleotide-sample slide comprising index sequences attached to sample library fragments, (ii) running genomic sequencing cycles of a sequencing run to determine nucleotide reads for the nucleotide-sample slide, and (iii) running indexing cycles to determine index sequences associated with the nucleotide reads determined during the genomic sequencing cycles.
[0059] As shown in FIG. 2, the adaptive decision-boundary base-calling system 106 may detect or receive a nucleotide-sample slide 202 comprising sample library fragments deposited within one of wells 204 within the nucleotide-sample slide 202. As suggested above, sample library fragments from multiple samples can be deposited into the wells 204 of the nucleotide-sample slide 202, where each sample library fragment has an index sequence adhered or attached thereto, and each unique index sequence corresponds to a respective sample to which each sample library fragment belongs. The nucleotide-sample slide 202 can subsequently be inserted in and detected by a sequencing device (e.g., the sequencing device 102).
[0060] As further shown in FIG. 2, the adaptive decision-boundary base-calling system 106 performs a sequencing run comprising genomic sequencing cycles 206 and indexing cycles 208 to generate nucleobase calls for nucleotide reads 210. In the illustrated implementation, for example, the nucleotide reads 210 include a first genomic sequence Rl, a first index sequence IIcorresponding to the first genomic sequence Rl, a second index sequence 12 corresponding to a second genomic sequence R2, and the second genomic sequence R2.
[0061] During the genomic sequencing cycles 206, for instance, the adaptive decisionboundary base-calling system 106 uses the sequencing device 102 to incorporate nucleobases of one or more nucleobase types into growing oligonucleotides corresponding to the genomic sequences Rl or R2. Likewise, during the indexing cycles 208, the adaptive decision-boundary base-calling system 106 uses the sequencing device 102 to incorporate nucleobases of one or more nucleobase types into growing oligonucleotides corresponding to the index sequences II or 12. Accordingly, during a given genomic sequencing cycle or indexing cycle, the adaptive decisionboundary base-calling system 106 can capture and analyze images of one or more clusters of oligonucleotides that have incorporated nucleobases with tags (e.g., fluorescent tags) complimenting or mirroring the nucleobases of the nucleotide reads 210.
[0062] Although FIG. 2 depicts and suggests the adaptive decision-boundary base-calling system 106 performing the genomic sequencing cycles 206 before the indexing cycles 208, in some embodiments, the adaptive decision -boundary base-calling system 106 performs one or more of the indexing cycles 208 before one or more of the genomic sequencing cycles 206. The adaptive decision-boundary base-calling system 106 can accordingly generate nucleobase calls for indexing cycles in either an indexing-first workflow or a more standard workflow. In such indexing-first workflows, for example, the adaptive decision-boundary base-calling system 106 performs indexing cycles and demultiplexing as part of pre-processing and before performing genomic sequencing cycles. In such cases, the adaptive decision-boundary base-calling system 106 determines nucleobase calls for a first index sequence, nucleobase base calls for a second index sequence, nucleobase calls for a first nucleotide read, and subsequently nucleobase calls for a second nucleotide read. For instance, the adaptive decision-boundary base-calling system 106 determines base calls for index sequences before determining base calls for nucleotide reads as described by International Patent Application No. PCT / US2024 / 035567, entitled “Modifying Sequencing Cycles or Imaging During a Sequencing Run to Meet Customized Coverage Estimation,” filed June 26, 2024, assigned to Illumina, Inc., the disclosure of which is incorporated herein by reference in its entirety.
[0063] To illustrate such an indexing-first workflow, in some cases, a first index primer is annealed to a primer binding site appended to a sample genomic sequence extracted from a sample. After the first index primer is annealed, the adaptive decision-boundary base-calling system 106 determines base calls for a first index sequence attached to the sample genomic sequence. After determining base calls for a first index sequence, the adaptive decision-boundary base-calling system 106 determines base calls for a second index sequence by annealing a second index primerto the primer binding site appended to the sample genomic sequence. After which, the adaptive decision-boundary base-calling system 106 determines base calls for the second index sequence. In such cases, the second index sequence is appended to the 5’ end of the sample genomic sequence while the first index sequence is appended to the 7’ end of the sample genomic sequence.
[0064] After determining base calls for the first index sequence and the second index sequence, the adaptive decision-boundary base-calling system 106 determines base calls for a first nucleotide read and a second nucleotide read. In a paired-end sequencing run, for instance, the sample genomic sequence is sequenced from both ends, providing complementary information about the sample genomic sequence. As part of determining the base calls for the first nucleotide read, the adaptive decision-boundary base-calling system 106 anneals a first nucleotide read primer to a read primer binding site, and the adaptive decision-boundary base-calling system 106 sequences the first portion of the sample genomic sequence. After sequencing such a first portion, the adaptive decision-boundary base-calling system 106 performs a pair-end turn. During the pair-end turn, the P7 region is cleaved and all fragments are attached by the P5 region. Prior to the pair-end turn, the P7 region is annealed to the surface of the flow cell. Following the pair-end turn, the adaptive decision-boundary base-calling system 106 determines base calls for a second nucleotide read. In such an indexing-first workflow, the adaptive decision-boundary base-calling system 106 can generate artificial base-specific intensity values for a particular indexing cycle based on expected base-specific intensity values derived from one or more preceding indexing cycles in a given indexing-first sequencing run.
[0065] As further illustrated in FIG. 2, the adaptive decision-boundary base-calling system 106 determines the nucleobase reads 210 by determining nucleobase calls from intensity values corresponding to labeled nucleotides within one or more clusters of oligonucleotides. For instance, in some embodiments, the adaptive decision-boundary base-calling system 106 analyzes intensity values for signals from a given oligonucleotide cluster in both sequencing signal channels, or each of multiple sequencing signal channels (e.g., concurrently), to determine a nucleobase call. In some embodiments, based on the intensity values of the signal of the cluster in each sequencing signal channel, the adaptive decision-boundary base-calling system 106 can calculate, utilizing an expectation maximization and Gaussian probability distributions, the probability that the signal falls within the intensity-value boundaries of a certain base (A, C, T, G). The adaptive decisionboundary base-calling system 106 can then call the nucleobase incorporated into the cluster by selecting a nucleobase with intensity -value boundaries having the highest probability of comprising the corresponding intensity values. For example, based on the intensity values emitted by the signal of the cluster, the adaptive decision -boundary base-calling system 106 can determine that theintensity -value boundaries of the nucleobase with the highest probability for the cluster is cytosine (C).
[0066] As mentioned above, variations in nucleotide-sample-slide preparations (e.g., preparation of the nucleotide-sample slide 202) and sequencing equipment (e.g., the sequencing device 102) can cause the distribution of intensity values observed for a first fluorescent nucleobase label (e.g., a label for cytosine) to generally shift towards the intensity values of a second fluorescent nucleobase label (e.g., a label for adenine), such that the two different nucleobases are more difficult to distinguish from one another during base calling analysis. For example, a relative location of an intensity -value distribution for a particular nucleobase on a two-dimensional plot can sometimes drift towards another intensity -value distribution for a different nucleobase on the same plot. As shown in FIG. 2, for example, such a drop or drift of intensity values can be observed as a downward location shift of intensity values (e.g., intensity values for cytosine moving down toward the intensity values for adenine) on two-dimensional plots of intensity values detected by two different sequencing signal channels during the genomic sequencing cycles 206.
[0067] Moreover, while genomic sequencing cycles (e.g., the genomic sequencing cycles 206) having more diverse distributions of intensity values across all four candidate bases often comprise sufficiently diverse information for distinguishing respective candidate nucleobases, indexing cycles (e.g., the indexing cycles 208) often exhibit relatively less nucleobase-data diversity because of one or both of (i) the relative simplicity of index sequences compared to corresponding genomic sequences in terms of nucleobase content and (ii) the lack of a reference / quality control sequence (e.g., PhiX sequence) being base called during a respective indexing cycle. As shown in FIG. 2, for example, in certain implementations, the relative simplicity of index sequences results in a partial or complete absence of intensity values corresponding to a particular candidate nucleobase during respective indexing cycles (e.g., the indexing cycles 208).
[0068] Also, as mentioned above, in some embodiments, the adaptive decision-boundary basecalling system 106 can generate a base-call data file comprising the nucleotide reads 210. Furthermore, in one or more embodiments, the base-call data file includes multiple nucleotide reads identified utilizing index sequences identified by the adaptive decision-boundary base-calling system 106. Accordingly, the adaptive decision-boundary base-calling system 106 can determine nucleotide reads from multiple samples of the nucleotide-sample slide 202 and identify each nucleotide read as belonging to a particular sample utilizing their respective index sequences.
[0069] As mentioned previously, the adaptive decision-boundary base-calling system 106 can supplement intensity values in indexing cycles with adaptively generated and artificial intensity values, resulting in a significant increase to base-calling accuracy during indexing cycles. For example, FIG. 3 illustrates an overview diagram of the adaptive decision-boundary base-callingsystem 106 generating and implementing artificial base-specific intensity values to generate a nucleobase call according to one or more embodiments. As illustrated in FIG. 3, the adaptive decision-boundary base-calling system 106 performs a series of acts 300 comprising an act 302 of accessing intensity values for a series of sequencing cycles, an act 304 of determining expected base-specific intensity value(s) from base-specific intensity -value distribution(s), an act 306 of generating artificial base-specific intensity values for an indexing cycle, an act 308 of determining intensity -value base-decision boundaries for the indexing cycle, and an act 310 of generating a nucleobase call based on the intensity -value base-decision boundaries. The following paragraphs further describe the acts depicted in FIG. 3 as part of an overview.
[0070] As shown in FIG. 3, for example, the adaptive decision-boundary base-calling system 106 performs the act 302 of accessing intensity values for a series of sequencing cycles. As illustrated, the adaptive decision-boundary base-calling system 106 utilizes a nucleotide-sample slide 312 for sequencing sample library fragments derived from samples. As described above, the nucleotide-sample slide 312 can include oligonucleotides, including genomic sequences and associated index sequences, that receive or incorporate labeled nucleobases during a given sequencing cycle. In particular, the nucleotide-sample slide 312 can include a cluster of oligonucleotides within each section (e.g., a tile or sub-tile comprising wells of a flow cell) thereof. When stimulated, the labeled nucleobases can emit a signal having characteristics associated with the type of nucleotide base.
[0071] As further shown in FIG. 3, the adaptive decision-boundary base-calling system 106 receives (or identifies) a series of signals 315 corresponding to at least on section of the nucleotide- sample slide 312, such as a section comprising multiple adjacent oligonucleotide clusters and / or an individual cluster of oligonucleotides. In some embodiments, the adaptive decision-boundary basecalling system 106 captures a series of images as the labeled nucleobases within a cluster of oligonucleotides emit the respective series of signals 315. As illustrated, the adaptive decisionboundary base-calling system 106 captures signals for each sequencing cycle in a set of sequencing cycles 314, including genomic sequencing cycles and indexing cycles (e.g., as described above in relation to FIG. 2). For example, in some embodiments, the adaptive decision-boundary basecalling system 106 utilizes a two-channel implementation or a four-channel implementation to capture two or four different images of the section of the nucleotide-sample slide 312 for each sequencing cycle in the set of sequencing cycles 314.
[0072] As also shown in FIG. 3, the adaptive decision-boundary base-calling system 106 extracts or identifies, from images depicting the signals 315, intensity values 316 corresponding to labeled nucleotides within one or more clusters of oligonucleotides of the nucleotide-sample slide 312. To illustrate the intensity values 316, FIG. 3 depicts a two-dimensional plot of intensity valuesdetected by two separate sequencing signal channels of a two-channel implementation. As mentioned, the intensity values 316 can indicate the type of nucleotide base that was added to the cluster of oligonucleotides for a given sequencing cycle of the set of sequencing cycles 314.
[0073] As further shown in FIG. 3, after accessing the intensity values 316, the series of acts includes an act 304 of determining expected base-specific intensity value(s) from base-specific intensity-value distribution(s). As illustrated, the adaptive decision-boundary base-calling system 106 analyzes base-specific intensity -value distributions of intensity values from a subset of intermediate genomic sequencing cycles (e.g., genomic sequencing cycles 16 - 26 or 13 - 33 of a sequencing run of approximately 150 genomic sequencing cycles) of the set of sequencing cycles 314 to determine at least one expected base-specific intensity value 318 for the sequencing run. Additionally or alternatively, the adaptive decision-boundary base-calling system 106 determines expected base-specific intensity values from one or more base-specific intensity-value distributions from a series of sequencing cycles (e.g., early indexing cycles 4-7, 5-10, etc.) that precede a target indexing cycle to accommodate an indexing first approach to sequencing runs in which nucleobase calls for index sequences precede nucleobase calls for genomic sequences.
[0074] To illustrate such expected base-specific intensity values, in one or more embodiments, the adaptive decision-boundary base-calling system 106 determines a centroid (or other representative set of intensity values) for each of the one or more base-specific intensity-value distributions (e.g., intensity-value distributions for adenine, intensity -value distributions for guanine) across a subset of intermediate genomic sequencing cycles or other series of sequencing cycles. Accordingly, the adaptive decision-boundary base-calling system 106 can determine the at least one expected base-specific intensity value 318 by determining an average centroid of each of one or more base-specific intensity-value distributions across the subset of intermediate genomic sequencing cycles. Alternatively, in some embodiments, the disclosed systems determine an average centroid of each of one or more base-specific intensity -value distributions across the subset of intermediate genomic sequencing cycles by determining, for each base-specific intensity-value distribution, a mean intensity value for a first sequencing signal channel and a second sequencing signal channel.
[0075] Moreover, as shown in FIG. 3, the series of acts 300 includes an act 306 of generating artificial base-specific intensity values for an indexing cycle. For instance, the adaptive decisionboundary base-calling system 106 receives (or identifies) intensity values 322a for an indexing cycle 320, as illustrated by an intensity -value distribution plot in FIG. 3. As also shown, in some embodiments, the adaptive decision-boundary base-calling system 106 supplements the intensity values 322a of the indexing cycle with an artificial base-specific intensity value 334 corresponding to the at least one expected base-specific intensity value 318 (e.g., in some cases, supplementingeach of multiple expected base-specific intensity values with a corresponding artificial basespecific intensity value) to generate supplemented intensity values 322b for the indexing cycle 320. Additionally or alternatively, as shown in FIG. 3, the adaptive decision-boundary base-calling system 106 supplements the intensity values 322a of the indexing cycle with a distribution of artificial base-specific intensity values 336 centered at (or otherwise based on) the at least one expected base-specific intensity value 318 (e.g., in some cases, supplementing each of multiple expected base-specific intensity values with a corresponding distribution of artificial base-specific intensity values) to generate further supplemented intensity values 322c for the indexing cycle 320. In some embodiments, for example, the distribution of artificial base-specific intensity values 336 comprises a Gaussian distribution (e.g., with the distribution of intensity values centered at the at least one expected base-specific intensity value 318).
[0076] As further shown in FIG. 3, after generating artificial base-specific intensity values, the series of acts 300 includes an act 308 of determining intensity -value base-decision boundaries for the indexing cycle. As illustrated, the adaptive decision-boundary base-calling system 106 determines intensity -value base-decision boundaries 324 for the indexing cycle 320 based on the further supplemented intensity values 322c, including the intensity values 322a observed for the indexing cycle 320 and the distribution of artificial base-specific intensity values 336. In some embodiments, for example, the adaptive decision-boundary base-calling system 106 generates the intensity -value base-decision boundaries 324 for differentiating signals corresponding to different nucleotide bases according to a base-call-distribution model (e.g., a segmented Gaussian mixture model). As shown in FIG. 3, in certain implementations, the adaptive decision-boundary basecalling system 106 can determine the intensity -value base-decision boundaries 324 with increased accuracy and reliability due to the addition of the distribution of artificial base-specific intensity values 336. This disclosure describes examples of such increased accuracy and reliability for base calling below with respect to FIGS. 5-8.
[0077] As also shown in FIG. 3, the series of acts 300 includes an act 310 of generating a nucleobase call based on the intensity -value base-decision boundaries. In one or more embodiments, for example, the adaptive decision-boundary base-calling system 106 utilizes the intensity -value base-decision boundaries 324 to generate a nucleobase call for a particular oligonucleotide cluster within the nucleotide-sample slide based on an intensity value corresponding to the particular oligonucleotide cluster. Furthermore, in some embodiments, the adaptive decision-boundary base-calling system 106 utilizes the intensity -value base-decision boundaries 324 to generate nucleobase calls for multiple oligonucleotide clusters within the nucleotide-sample slide based on the intensity values for the indexing cycle 320.
[0078] As mentioned, in some embodiments, the adaptive decision-boundary base-calling system 106 identifies which nucleotide reads belong to which samples in a multi-sample nucleotide-sample slide based on indexing cycles and corresponding index sequences in a process known as demultiplexing. For instance, in some cases, the adaptive decision-boundary base-calling system 106 access index sequences associated with sample genomic sequences, the index sequences acting as unique identifiers for each sample, allowing for differentiation and sorting of nucleotide reads during demultiplexing. Accordingly, the adaptive decision-boundary base-calling system 106 can demultiplex nucleotide reads by utilizing a reference of known index sequences. By comparing index sequences with known index sequences in a reference of registered indexes, the adaptive decision-boundary base-calling system 106 can identify samples that correspond with one or more unique index sequences. In some implementations, the adaptive decision-boundary base-calling system 106 uses an indexing-first workflow, as described above, in which indexing cycles for both nucleotide reads precede genomic sequencing cycles for both reads. More particularly, the adaptive decision-boundary base-calling system 106 may demultiplex samples before sequencing the samples. In other implementations, the adaptive decision-boundary basecalling system 106 does not use the indexing -first workflow. In such examples, the adaptive decision-boundary base-calling system 106 performs the indexing cycles between respective sets of genomic sequencing cycles, as particularly shown in parts of FIGS. 2 and 3.
[0079] As suggested by FIG. 3, the adaptive decision-boundary base-calling system 106 generates a nucleobase call for an index sequence as part of the indexing cycle 320 based on the intensity-value base-decision boundaries 324. In some embodiments, the adaptive decisionboundary base-calling system 106 determines intensity -value base-decision boundaries for multiple subsequent indexing cycles of a sequencing run and utilizes the respective intensity-value basedecision boundaries to determine nucleobase calls of index sequences comprising nucleobases identified from the multiple subsequent indexing cycles. Accordingly, intensity-value basedecision boundaries can be cycle specific.
[0080] As mentioned above, in addition or in the alternative to determining intensity -value base-decision boundaries for an indexing cycle, in some embodiments, the adaptive decisionboundary base-calling system 106 can determine intensity -value base-decision boundaries for a genomic sequencing cycle. Accordingly, in some cases, the adaptive decision-boundary basecalling system 106 performs the act 302 of accessing intensity values for a series of sequencing cycles, the act 304 of determining expected base-specific intensity value(s) from base-specific intensity -value distribution(s), and the act 306 of generating artificial base-specific intensity values for a genomic sequencing cycle — instead of or in addition to generating the artificial base-specific intensity values for the indexing cycle. Given this application to a genomic sequencing cycle, theadaptive decision-boundary base-calling system 106 can further perform the act 308 of determining intensity -value base-decision boundaries for the genomic sequencing cycle and the act 310 of generating a nucleobase call for an oligonucleotide cluster based on a corresponding intensity value and the intensity -value base-decision boundaries.
[0081] As indicated above, the adaptive decision-boundary base-calling system 106 can generate artificial base-specific intensity values for indexing cycles based on expected basespecific intensity values determined from genomic sequencing cycles. In accordance with one or more embodiments, FIGS. 4A-4B illustrate the adaptive decision-boundary base-calling system 106 generating and implementing artificial base-specific intensity values during indexing cycles based on one or more expected base-specific intensity value(s) 408 identified from a set of genomic sequencing cycles.
[0082] As shown in FIG. 4A, for example, the adaptive decision-boundary base-calling system 106 receives (or identifies) intensity values 404a-404n observed during genomic sequencing cycles 402a-402n (e.g., a subset of the genomic sequencing cycles 400 depicted in FIG. 4B). As also shown in FIG. 4A, the adaptive decision-boundary base-calling system 106 determines the expected base-specific intensity value(s) 408 from base-specific intensity -value distributions 406a- 406n of the intensity values 404a-404n, respectively. FIGS. 4A and 4B depict the intensity values 404a-404n as distinguishable clouds of intensity values within respective two-dimensional intensity-value distribution plots identified as the base-specific intensity-value distributions 406a- 406n. As mentioned above, for example, the adaptive decision -boundary base-calling system 106 determines centroid intensity values of a given base-specific intensity-value distribution (e.g., the distribution of intensity values corresponding to cytosine or another nucleobase) to be expected base-specific intensity values.
[0083] As further illustrated in FIG. 4A, the adaptive decision-boundary base-calling system 106 receives (or identifies), for an indexing cycle 412a, intensity values 414a detected from labeled nucleobases of oligonucleotide clusters within a nucleotide-sample. The indexing cycle 412a can be an indexing cycle from indexing cycles 410 within the same sequencing run of the genomic sequencing cycles 400. As mentioned previously, a lack of nucleobase-data diversity commonly part of indexing cycles can often complicate a determination of detection thresholds for certain nucleobases in the absence of sufficient intensity data representing one or more of four different candidate nucleobases. In the illustrated implementation, for example, the indexing cycle 412a exhibits a relative lack of nucleobase-data diversity, as evidenced by a lack of intensity values in the upper-right quadrant of the depicted two-dimension plot of the intensity values 414a.
[0084] Accordingly, as shown in FIG. 4A, the adaptive decision-boundary base-calling system 106 generates a distribution of artificial base-specific intensity values 415a to generatesupplemented intensity values 416a for the indexing cycle 412a. Such a distribution of artificial intensity values can, for example, be centered at the expected intensity value (of the expected basespecific intensity value(s) 408) corresponding to cytosine. Based on the supplemented intensity values 416a — including the intensity values 414a observed during the indexing cycle 412a and the distribution of artificial base-specific intensity values 415a — the adaptive decision-boundary basecalling system 106 determines intensity -value base-decision boundaries. Based on a comparison of the resultant boundaries and respective intensity values of the intensity values 414a, the adaptive decision-boundary base-calling system 106 generates one more nucleobase calls 420 (e.g., as described above in relation to FIG. 3).
[0085] Furthermore, in some embodiments, the adaptive decision-boundary base-calling system 106 generates and implements artificial base-specific intensity values corresponding to more than one of four candidate nucleobases. For example, FIG. 4B illustrates the adaptive decision-boundary base-calling system 106 supplementing intensity values observed during the indexing cycles 410 with different artificial base-specific intensity values covering different candidate nucleobases to varying degrees.
[0086] As illustrated in FIG. 4B, for instance, the adaptive decision-boundary base-calling system 106 determines the expected base-specific intensity values 408 for four candidate nucleobases (A, C, T, G) based on the intensity values 404a-404n from a subset of the genomic sequencing cycles 400 (e.g., the genomic sequencing cycles 402a-402n). Based on the expected base-specific intensity values 408, the adaptive decision-boundary base-calling system 106 generates and implements distributions of artificial base-specific intensity values to supplement intensity values 414a-414n of the indexing cycles 410. As illustrated, the adaptive decisionboundary base-calling system 106 adds distributions of artificial base-specific intensity values corresponding to (e.g., centered at) one or more of the expected base-specific intensity value(s) 408. The adaptive decision-boundary base-calling system 106 adds such distributions of artificial base-specific intensity values to either (a) supplement a given indexing cycle with base-specific intensity values where intensity value data is relatively scarce or non-existent (e.g., as shown in FIG. 4A for indexing cycle 412a and in FIG. 4B to generate supplemented intensity values 416a- 416n for indexing cycles 410) or (b) supplement a given indexing cycle with base-specific intensity values for each of the four candidate nucleobases (e.g., as shown in FIG. 4B to generate supplemented intensity values 418a-418n for the indexing cycles 410).
[0087] As shown in FIG. 4B, for example, the adaptive decision-boundary base-calling system 106 can supplement the intensity values 414a with the artificial base-specific intensity values 415a centered at (or otherwise based on) the intensity value (of the expected base-specific intensity value(s) 408) corresponding to cytosine, which has an expected location in the upper-right quadrantof the illustrated two-dimensional plot, to generate the supplemented intensity values 416a for the respective indexing cycle of the indexing cycles 410. Alternatively, as also shown in FIG. 4B, the adaptive decision-boundary base-calling system 106 can supplement the intensity values 414a with artificial base-specific intensity values 415a-415d respectively centered at (or otherwise based on) each of the four expected base-specific intensity values 408 to generate the supplemented intensity values 418a for the respective indexing cycle of the indexing cycles 410.
[0088] Similarly, as shown in FIG. 4B, the adaptive decision-boundary base-calling system 106 can supplement the intensity values 414n with the artificial base-specific intensity values 415b and 415c centered at (or otherwise based on) the respective intensity values of the expected basespecific intensity values 408 corresponding to adenine (bottom-right quadrant) and thymine (upperleft quadrant) to generate the supplemented intensity values 416n for the respective indexing cycle of the indexing cycles 410. Alternatively, as also shown in FIG. 4B, the adaptive decision-boundary base-calling system 106 can supplement the intensity values 414n with the artificial base-specific intensity values 415a-415d respectively centered at (or otherwise based on) each of the four expected base-specific intensity values 408 to generate the supplemented intensity values 418n for the respective indexing cycle of the indexing cycles 410.
[0089] To ensure that artificial base-specific intensity values for one nucleobase are not too close to the intensity -value base-decision boundaries of another nucleobase, the adaptive decisionboundary base-calling system 106 can utilize a threshold or minimum distance to keep the respective intensity values distinct. In one or more embodiments, for instance, the adaptive decision-boundary base-calling system 106 generates the artificial base-specific intensity values with a predetermined minimum distance between respective expected base-specific intensity values corresponding to four different candidate nucleobases. For example, the adaptive decisionboundary base-calling system 106 can modify the relative location of the distributions of artificial base-specific intensity values to avoid overlap or excess proximity of the respective distributions to one another.
[0090] As mentioned above, in certain embodiments, the adaptive decision-boundary basecalling system 106 implements nucleobase calling with increased accuracy relative to existing sequencing systems. To illustrate, FIGS. 5-8 show experimental results of the adaptive decisionboundary base-calling system 106 generating nucleobase calls during indexing cycles with increased accuracy over an existing sequencing system. For instance, FIG. 5 illustrates a plot 500 of comparative experimental results of identifying an expected base-specific intensity value (e.g., a centroid of observed base-specific intensity-value distributions) for sequencing cycles of a sequencing run utilizing (i) an existing sequencing system (labelled as “Baseline” in FIG. 5) and (ii) the adaptive decision-boundary base-calling system 106 (labelled as “Adaptive” in FIG. 5).
[0091] More specifically, the plot 500 includes comparative results of the existing sequencing system (“Baseline”) and the adaptive decision-boundary base-calling system 106 (“Adaptive”) identifying an expected base-specific intensity value for the candidate nucleotide cytosine (“Normalized Cloud Centroid / 3 C Y”). The plot 500 illustrates such identified values for multiple sequencing cycles (approximately 200 cycles) of a sequencing run, including a complete set of genomic sequencing cycles for a first nucleotide read R1 (e.g., approximately cycles 1-150), complete sets of indexing cycles for a first index sequence II and a second index sequence 12 (e.g., approximately cycles 151-160 for II and cycles 161-170 for 12), and a partial set of genomic sequencing cycles for a second nucleotide read R2 (e.g., approximately cycles 171-200). The first index sequence II corresponds to the first nucleotide read Rl, and the second index sequence 12 corresponds to the second nucleotide read R2.
[0092] As mentioned above and as illustrated in FIG. 5, the adaptive decision-boundary basecalling system 106 (“Adaptive”) identifies expected base-specific intensity values (e.g., a respective centroid of the cytosine cloud for each sequencing cycle) for indexing cycles based on (i) intensity values observed during each indexing cycle and (ii) artificial base-specific intensity values generated by the adaptive decision-boundary base-calling system 106 based on base-specific intensity-value distributions of a subset of genomic sequencing cycles of the sequencing run. In contrast, the existing sequencing system (“Baseline”) identifies the expected base-specific intensity values throughout the sequencing run without reference to artificial base-specific intensity values. For example, FIG. 5 includes an example intensity -value distribution plot 512 indicating where the existing sequencing system (“Baseline”) has erroneously identified a significant portion of the base-specific intensity -value distribution 514 for cytosine (the upper-right cloud of intensity values within the intensity -value distribution plot 512). As a consequence of erroneously identifying a significant portion of the base-specific intensity-value distribution 514 for cytosine, the existing sequencing system miss-estimates expected base-specific intensity values during the indexing cycles relative to the estimated values for the genomic sequencing cycles (as indicated by the highlighted centroid values 510 in the upper right quadrant of the plot 500). As mentioned, many existing sequencing systems fail to accurately identify certain base-specific intensity-value distributions during indexing cycles as observed intensity values for one nucleobase (e.g., cytosine) shift towards observed intensity values for another nucleobase (e.g., adenine), as demonstrated by the results shown in FIG. 5.
[0093] To further illustrate, FIG. 5 also includes an intensity-value distribution plot 522 indicating where the adaptive decision-boundary base-calling system 106 (“Adaptive”) has correctly identified the base-specific intensity-value distribution 524 for cytosine (the upper-right cloud of intensity values within the intensity-value distribution plot 522). As a consequence ofcorrectly identifying significantly more intensity values of the base-specific intensity-value distribution 524 for cytosine, relative to the baseline system, the adaptive decision-boundary basecalling system 106 estimates expected base-specific intensity values during the indexing cycles to constitute values more consistent with the estimated values for the genomic sequencing cycles (as indicated by the highlighted centroid values 520 in the lower right quadrant of the plot 500). Indeed, by identifying base-specific intensity values with increased accuracy over existing sequencing systems, the adaptive decision-boundary base-calling system 106 can generate nucleobase calls with increased accuracy, despite the aforementioned shifts in observed intensity values for one nucleobase (e.g., cytosine) towards the observed intensity values for another nucleobase (e.g., adenine).
[0094] As demonstrated by the experimental results portrayed in FIG. 5, the adaptive decisionboundary base-calling system 106 more accurately identifies expected base-specific intensity values for nucleobases (e.g., cytosine) and generates more accurate intensity-value base-decision boundaries for such nucleobases during indexing cycles by supplementing the expected basespecific intensity values with artificial base-specific intensity values. Because of the more accurate intensity-value base-decision boundaries, in certain instances, the adaptive decision-boundary base-calling system 106 prevents the misidentification of intensity values for one nucleotide as corresponding to another (e.g., the intensity values observed for cytosine being mistakenly identified as intensity values observed for adenine).
[0095] As also mentioned above, in certain described embodiments, the adaptive decisionboundary base-calling system 106 provides for increased overall accuracy in calling nucleobases within index sequences of multi-sample nucleotide-sample slides relative to existing sequencing systems. To illustrate, FIGS. 6A-8 show comparative experimental results of determining nucleobase calls for multiplexed nucleotide-sample slides utilizing (i) an existing sequencing system and (ii) the adaptive decision-boundary base-calling system 106.
[0096] For instance, FIGS. 6A and 6B illustrate six box-plot diagrams 600a-600c and 600d- 600f, respectively, of mismatch rates from six respective sequencing runs performed on nucleotide- sample slides with varying degrees of multiplexing — that is, nucleotide-sample slides populated with varying numbers of individual samples labeled with unique index sequences. Along the vertical axis, each of the six box-plot diagrams 600a — 600f include mismatch rates shown as a percentage of sequencing cycles comprising at least one mismatch between an actual nucleobase and a nucleobase call respectively generated by the existing sequencing system (left box plot of each numbered box-plot pair along the horizontal axis) or the adaptive decision-boundary basecalling system 106 (right box plot of each numbered box-plot pair along the horizontal axis) — for each of eight lanes of a respective nucleotide-sample slide (numbered along the horizonal axis).
[0097] As indicated in the box-plot diagrams 600a-600f of FIGS. 6A and 6B, mismatch rates of sequencing runs utilizing the adaptive decision-boundary base-calling system 106 (right box plot of each box-plot pair along the horizontal axis) are consistently decreased relative to mismatch rates of sequencing runs utilizing the existing sequencing system (left box plot of each box-plot pair along the horizontal axis). Indeed, as demonstrated by the comparative experimental results provided by FIGS. 6A and 6B, by adaptively implementing artificial base-specific intensity values to supplement intensity values observed during indexing cycles, the adaptive decision-boundary base-calling system 106 improves the accuracy of nucleobase calling during indexing cycles relative to existing sequencing systems.
[0098] Moreover, the box-plot diagrams 600a-600f of FIGS. 6A and 6B demonstrate that the adaptive decision-boundary base-calling system 106 improves the accuracy of nucleobase calling for nucleotide-sample slides load with samples of varying degrees of multiplexing — that is, nucleotide-sample slides populated with various quantities of individual samples. To illustrate, as shown in FIG. 6A, box-plot diagram 600a represents results for a nucleotide-sample slide comprising 24 samples per lane (“24plex”), box-plot diagram 600b represents results for a nucleotide-sample slide comprising between 4 and 64 samples per lane (“4 Pl ex,” “8 Pl ex,” “16 Plex,” and “64 Plex”), and box-plot diagram 600c represents results for a nucleotide-sample slide comprising 96 samples (“96plex”). Further, as shown in FIG. 6B, box-plot diagram 600d represents results for a nucleotide-sample slide comprising 96 samples (“96plex”), whereas box-plot diagrams 600e and 600f respectively represent results for nucleotide-sample slides comprising 10 samples (“lOplex”).
[0099] To further illustrate, FIGS. 7A-7B and 8 show graphical plots comparing nucleobase- calling results from multiple sequencing runs performed utilizing the existing sequencing system and the adaptive decision-boundary base-calling system 106 on nucleotide-sample slides with varying degrees of multiplexing. Specifically, FIGS. 7A and 7B illustrate six graph plots 700a- 700c and 700d-700f, respectively, of differences (“Diff”) in mismatch rates (e.g., as described above in relation to FIGS. 6A or 6B) between sequencing runs performed by the existing sequencing model and the adaptive decision-boundary base-calling system 106. In particular, a negative “Diff’ value (vertical axis) represents an overall decrease in mismatch rate relative to the existing sequencing system for each lane (horizontal axis) of a respective nucleotide-sample slide. As demonstrated by each of the graph plots 700a-700f, the adaptive decision-boundary basecalling system 106 generally exhibits increased accuracy in nucleobase calls relative to existing systems.
[0100] Relatedly, FIG. 8, illustrates respective percentages of conclusively identifiable nucleotide reads — that is, reads for which a corresponding index sequence is conclusivelyidentified by accurately calling enough nucleobases to distinguish the index sequence from others — from sequencing runs performed utilizing (i) the existing sequencing system (“Baseline”) and (ii) the adaptive decision-boundary base-calling system 106 (“Adaptive”). As mentioned, by generating nucleobase calls for indexing cycles with increased accuracy relative to existing sequencing systems (e.g., as demonstrated by the results shown in FIGS. 6A-6B and 7A-7B), the adaptive decision-boundary base-calling system 106 can conclusively identify more index sequences in a given sequencing run relative to existing sequencing systems (e.g., as demonstrated by the results shown in FIG. 8). Indeed, as demonstrated by the comparative experimental results shown in FIG. 8, the adaptive decision-boundary base-calling system 106 provides for increased accuracy of nucleobase calling during indexing cycles, resulting in an increased percentage of nucleotide reads available for downstream procedures, such as mapping, alignment, variant calling, and so forth.
[0101] Turning now to FIG. 9, this figure illustrates an example flowchart of a series of acts for generating and implementing artificial base-specific intensity values for determining intensityvalue base-decision boundaries for an indexing cycle according to one or more embodiments. While FIG. 9 illustrates acts according to particular embodiments, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 9. The acts of FIG. 9 can be performed as part of a method. Alternatively, a non-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device to perform the acts depicted in FIG. 9. In still further embodiments, a system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 9.
[0102] As shown in FIG. 9, the series of acts 900 includes an act 902 of receiving, for a sequencing run, intensity values for labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide, an act 904 of determining expected base-specific intensity value(s) based on base-specific intensity -value distribution(s) from at least one sequencing cycle of the sequencing run, an act 906 of generating, for an indexing cycle, artificial base-specific intensity values based on the expected base-specific intensity values, an act 908 of determining, for the indexing cycle, intensity-value base-decision boundaries based on the artificial base-specific intensity values, and an act 910 of generating, for the indexing cycle, a nucleobase call for an oligonucleotide cluster based on a corresponding intensity value and the intensity-value base-decision boundaries.
[0103] For example, the series of acts 900 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A method comprising:receiving, for an indexing cycle, intensity values detected from labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide; determining one or more expected base-specific intensity values based on genomic intensity values from a set of genomic sequencing cycles; generating artificial base-specific intensity values based on the one or more expected basespecific intensity values; determining, for the indexing cycle, intensity-value base-decision boundaries based on the artificial base-specific intensity values; and generating a nucleobase call for an oligonucleotide cluster within the nucleotide-sample slide based on a corresponding intensity value of the intensity values for the indexing cycle and the intensity-value base-decision boundaries.CLAUSE 2. The method of clause 1, further comprising: determining, based on the nucleobase call, an index sequence for the oligonucleotide cluster of the nucleotide-sample slide; and identifying, utilizing the index sequence, a particular sample for the oligonucleotide cluster from a plurality of samples corresponding to one or more oligonucleotide clusters on the nucleotide-sample slide.CLAUSE 3. The method of clause 2, further comprising generating a base-call data file comprising a nucleotide read from the particular sample identified utilizing the index sequence.CLAUSE 4. The method of any of clauses 2-3, further comprising generating, for the indexing cycle, additional nucleobase calls for additional oligonucleotide clusters within the nucleotide-sample slide based on other intensity values and the intensity-value base-decision boundaries.CLAUSE 5. The method of any of clauses 1-4, wherein the set of genomic sequencing cycles comprises a subset of intermediate genomic sequencing cycles within a sequencing run performed on the nucleotide-sample slide.CLAUSE 6. The method of any of clauses 1-5, wherein determining the one or more expected base-specific intensity values comprises: determining, from the set of genomic sequencing cycles, one or more base-specific intensity-value distributions comprising the genomic intensity values; and determining, from the one or more base-specific intensity-value distributions, the one or more expected base-specific intensity values for one or more of four candidate nucleobases.CLAUSE 7. The method of clause 6, wherein determining the one or more expected base-specific intensity values comprises determining an average centroid of each of the one or more base-specific intensity-value distributions across the set of genomic sequencing cycles.CLAUSE 8. The method of clause 7, wherein determining the average centroid of each of the one or more base-specific intensity-value distributions across the set of genomic sequencing cycles comprises determining, for each base-specific intensity-value distribution, a mean intensity value for a first sequencing signal channel and a second sequencing signal channel.CLAUSE 9. The method of any of clauses 1-8, wherein generating the artificial basespecific intensity values comprises: generating a respective distribution of artificial intensity values for each of the one or more expected base-specific intensity values; and generating the artificial base-specific intensity values based on the respective distribution of artificial intensity values.CLAUSE 10. The method of any of clauses 1-9, further comprising determining the one or more expected base-specific intensity values to include a respective expected base-specific intensity value for each of four different candidate nucleobases.CLAUSE 11. The method of clause 10, further comprising generating the artificial basespecific intensity values with a predetermined minimum distance between respective expected base-specific intensity values corresponding to the four different candidate nucleobases.CLAUSE 12. The method of any of clauses 1-11, further comprising determining the intensity-value base-decision boundaries further based on the intensity values received for the indexing cycle.CLAUSE 13. The method of clause 12, further comprising determining additional intensity-value base-decision boundaries for a subsequent indexing cycle based on the artificial base-specific intensity values and respective intensity values received for the subsequent indexing cycle.CLAUSE 14. The method of any of clauses 1-13, wherein generating the nucleobase call for the oligonucleotide cluster within the nucleotide-sample slide comprises: determining respective nucleobase probabilities for the oligonucleotide cluster based on comparing the corresponding intensity value to the intensity-value base-decision boundaries; and generating the nucleobase call for the oligonucleotide cluster based on determining a highest nucleobase probability.CLAUSE 15. A method comprising: receiving, for a series of sequencing cycles within a sequencing run, intensity values for labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide; determining one or more expected base-specific intensity values based on one or more basespecific intensity-value distributions of at least one sequencing cycle of the series of sequencing cycles;generating, for an indexing cycle of the series of sequencing cycles, artificial base-specific intensity values based on the one or more expected base-specific intensity values; determining, for the indexing cycle, intensity-value base-decision boundaries based on the artificial base-specific intensity values; and generating, for the indexing cycle, a nucleobase call for an oligonucleotide cluster within the nucleotide-sample slide based on a corresponding intensity value and the intensity-value basedecision boundaries.CLAUSE 16. The method of clause 15, wherein the at least one sequencing cycle of the series of sequencing cycles comprises: a preceding indexing cycle performed before the indexing cycle; or a subset of preceding indexing cycles performed before the indexing cycle.CLAUSE 17. The method of any of clauses 15-16, further comprising performing the series of sequencing cycles within the sequencing run according to an order of indexing cycles before genomic sequencing cycles by: determining nucleobase calls for a first index sequence appended to a sample genomic sequence of a sample of one or more samples; determining nucleobase calls for a second index sequence appended to the sample genomic sequence of the sample; and after determining the nucleobase calls for the first index sequence and the second index sequence, determining nucleobase calls for a first nucleotide read corresponding to a first portion of the sample genomic sequence and determining base calls for a second nucleotide read corresponding to a second portion of the sample genomic sequence.CLAUSE 18. The method of any of clauses 15-17, further comprising: determining, based on the nucleobase call, an index sequence for the oligonucleotide cluster of the nucleotide-sample slide; and identifying, utilizing the index sequence, a particular sample for the oligonucleotide cluster from a plurality of samples corresponding to one or more oligonucleotide clusters on the nucleotide-sample slide.CLAUSE 19. The method of clause 18, further comprising generating a base-call data fde comprising a nucleotide read from the particular sample identified utilizing the index sequence.CLAUSE 20. The method of any of clauses 15-19, further comprising generating, for the indexing cycle, additional nucleobase calls for additional oligonucleotide clusters within the nucleotide-sample slide based on other intensity values and the intensity-value base-decision boundaries.CLAUSE 21. The method of any of clauses 15 and 18-20, wherein the at least one sequencing cycle of the series of sequencing cycles comprises a subset of genomic sequencing cycles preceding the indexing cycle within the sequencing run.CLAUSE 22. The method of any of clauses 15-21, further comprising determining the one or more expected base-specific intensity values by: determining, from the at least one sequencing cycle of the series of sequencing cycles, the one or more base-specific intensity-value distributions; and determining, from the one or more base-specific intensity-value distributions, the one or more expected base-specific intensity values for one or more of four candidate nucleobases.CLAUSE 23. The method of clause 22, further comprising determining the one or more expected base-specific intensity values by determining an average centroid of each of the one or more base-specific intensity-value distributions determined from the at least one sequencing cycle.CLAUSE 24. The method of clause 23, further comprising determining the average centroid of each of the one or more base-specific intensity-value distributions by determining, for each base-specific intensity-value distribution, a mean intensity value for a first sequencing signal channel and a second sequencing signal channel.CLAUSE 25. The method of any of clauses 15-24, further comprising generating the artificial base-specific intensity values by: generating a respective distribution of artificial intensity values for each of the one or more expected base-specific intensity values; and generating the artificial base-specific intensity values based on the respective distribution of artificial intensity values.CLAUSE 26. The method of any of clauses 15-25, further comprising determining the one or more expected base-specific intensity values to include a respective expected base-specific intensity value for each of four different candidate nucleobases.CLAUSE 27. The method of clause 26, further comprising generating the artificial basespecific intensity values with a predetermined minimum distance between respective expected base-specific intensity values corresponding to the four different candidate nucleobases.CLAUSE 28. The method of any of clauses 15-27, further comprising determining the intensity-value base-decision boundaries further based on the intensity values received for the indexing cycle.CLAUSE 29. The method of clause 28, further comprising determining additional intensity-value base-decision boundaries for a subsequent indexing cycle based on the artificial base-specific intensity values and respective intensity values received for the subsequent indexing cycle.CLAUSE 30. The method of any of clauses 15-29, further comprising generating the nucleobase call for the oligonucleotide cluster within the nucleotide-sample slide by: determining respective nucleobase probabilities for the oligonucleotide cluster based on comparing the corresponding intensity value to the intensity-value base-decision boundaries; and generating the nucleobase call for the oligonucleotide cluster based on determining a highest nucleobase probability.
[0104] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels (e.g., intensity values from one or more sequencing signal channels portrayed as different colors representing different nucleobase labels), for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.
[0105] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0106] SBS can utilize nucleotide monomers that have a terminator moiety or those that lack any terminator moieties. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using y-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is generally variable and dependent upon the template sequence and the mode of nucleotide delivery. For SBS techniques that utilize nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator can be reversible as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).
[0107] SBS techniques can utilize nucleotide monomers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such asmolecular weight or charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; or the like. In embodiments where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods developed by Solexa (now Illumina, Inc.).
[0108] Preferred embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C or G). Images obtained after addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed and analyzed using the methods set forth herein. For example, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different sequencing signal channels for reversible terminator-based sequencing methods.
[0109] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, in WO 04 / 018497 and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference. This approach is being commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators in which both the termination can be reversed, and the fluorescent label cleavedfacilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from these modified nucleotides.
[0110] Preferably in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured following incorporation of labels into arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a sequencing signal channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show nucleic acid features that have incorporated nucleotides of a particular type. Different features are present or absent in the different images due the different sequence content of each feature. However, the relative position of the features will remain unchanged in the images. Images obtained from such reversible terminator- SBS methods can be stored, processed and analyzed as set forth herein. Following the image capture step, labels can be removed, and reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.[OHl] In particular embodiments some or all of the nucleotide monomers can include reversible terminators. In such embodiments, reversible terminators / cleavable fluors can include fluor linked to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3' allyl group to block extension but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30 second exposure to long wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of natural termination that ensues after placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modifiednucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.
[0112] Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.
[0113] Some embodiments can utilize detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair. As a second example, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different sequencing signal channels, whereas other nucleotide types are detected in no more than one of the sequencing signal channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples, is a fluorescent-based SBS method that uses a first nucleotide type that is detected in a first sequencing signal channel (e.g. dATP having a label that is detected in the first sequencing signal channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second sequencing signal channel (e.g. dCTP having a label that is detected in the second sequencing signal channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second sequencing signal channels (e.g. dTTP having at least one label that is detected in both sequencing signal channels when excited by the first and / or second excitation wavelength) and a fourth nucleotidetype that lacks a label that is not, or minimally, detected in either sequencing signal channel (e.g. dGTP having no label).
[0114] Further, as described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single sequencing signal channel. In such so-called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0115] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained following treatment of an array of nucleic acid features with the labeled sequencing reagents. Each image will show nucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based sequencing methods can be stored, processed and analyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No. 6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.
[0116] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147- 151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820(2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed and analyzed as set forth herein. In particular, the data can be treated as an image in accordance with the exemplary treatment of optical images and other images that is set forth herein.
[0117] Some embodiments can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). The illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed and analyzed as set forth herein.
[0118] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.
[0119] The above SBS methods can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows convenient delivery of sequencing reagents, removalof unreacted reagents and detection of incorporation events in a multiplex manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as, bridge amplification or emulsion PCR as described in further detail below.
[0120] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.
[0121] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprising components such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeqTM platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference. The sequencing system described above sequences nucleic acid polymers present in samples received by a sequencing device, as described further above.
[0122] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0123] The components of the adaptive decision-boundary base-calling system 106 can include software, hardware, or both. For example, the components of the adaptive decisionboundary base-calling system 106 can include one or more instructions stored on a computer- readable storage medium and executable by processors of one or more computing devices (e.g., the client device 114, the local device 108, or the server device(s) 110). When executed by the one or more processors, the computer-executable instructions of the adaptive decision-boundary basecalling system 106 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the adaptive decision-boundary base-calling system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the adaptivedecision-boundary base-calling system 106 can include a combination of computer-executable instructions and hardware.
[0124] Furthermore, the components of the adaptive decision-boundary base-calling system 106 performing the functions described herein with respect to the adaptive decision-boundary basecalling system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the adaptive decision-boundary base-calling system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the adaptive decision-boundary base-calling system 106 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0125] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non- transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0126] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0127] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storageor other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0128] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.
[0129] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0130] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0131] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0132] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0133] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.
[0134] FIG. 10 illustrates a block diagram of a computing device 1000 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1000 may implement the adaptive decisionboundary base-calling system 106 and the sequencing device system 104. As shown by FIG. 10, the computing device 1000 can comprise a processor 1002, a memory 1004, a storage device 1006, an I / O interface 1008, and a communication interface 1010, which may be communicatively coupled by way of a communication infrastructure 1012. In certain embodiments, the computing device 1000 can include fewer or more components than those shown in FIG. 10. The following paragraphs describe components of the computing device 1000 shown in FIG. 10 in additional detail.
[0135] In one or more embodiments, the processor 1002 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1002 mayretrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1004, or the storage device 1006 and decode and execute them. The memory 1004 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1006 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0136] The I / O interface 1008 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1000. The I / O interface 1008 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1008 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1008 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0137] The communication interface 1010 can include hardware, software, or both. In any event, the communication interface 1010 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1000 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1010 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.
[0138] Additionally, the communication interface 1010 may facilitate communications with various types of wired or wireless networks. The communication interface 1010 may also facilitate communications using various communication protocols. The communication infrastructure 1012 may also include hardware, software, or both that couples components of the computing device 1000 to each other. For example, the communication interface 1010 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.
[0139] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanyingdrawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0140] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
CLAIMSWe Claim:
1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: receive, for an indexing cycle, intensity values detected from labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide; determine one or more expected base-specific intensity values based on genomic intensity values from a set of genomic sequencing cycles; generate artificial base-specific intensity values based on the one or more expected base-specific intensity values; determine, for the indexing cycle, intensity-value base-decision boundaries based on the artificial base-specific intensity values; and generate a nucleobase call for an oligonucleotide cluster within the nucleotide- sample slide based on a corresponding intensity value of the intensity values for the indexing cycle and the intensity-value base-decision boundaries.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, based on the nucleobase call, an index sequence for the oligonucleotide cluster of the nucleotide-sample slide; and identify, utilizing the index sequence, a particular sample for the oligonucleotide cluster from a plurality of samples corresponding to one or more oligonucleotide clusters on the nucleotide-sample slide.
3. The system of claim 2, further comprising instructions that, when executed by the at least one processor, cause the system to generate a base-call data fde comprising a nucleotide read from the particular sample identified utilizing the index sequence.
4. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate, for the indexing cycle, additional nucleobase calls for additional oligonucleotide clusters within the nucleotide-sample slide based on other intensity values and the intensity-value base-decision boundaries.
5. The system of claim 1, wherein the set of genomic sequencing cycles comprises a subset of intermediate genomic sequencing cycles within a sequencing run performed on the nucleotide-sample slide.
6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values by: determining, from the set of genomic sequencing cycles, one or more base-specific intensity-value distributions comprising the genomic intensity values; and determining, from the one or more base-specific intensity-value distributions, the one or more expected base-specific intensity values for one or more of four candidate nucleobases.
7. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values by determining an average centroid of each of the one or more base-specific intensity-value distributions across the set of genomic sequencing cycles.
8. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to determine the average centroid of each of the one or more base-specific intensity-value distributions across the set of genomic sequencing cycles by determining, for each base-specific intensity-value distribution, a mean intensity value for a first sequencing signal channel and a second sequencing signal channel.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the artificial base-specific intensity values by: generating a respective distribution of artificial intensity values for each of the one or more expected base-specific intensity values; and generating the artificial base-specific intensity values based on the respective distribution of artificial intensity values.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values to include a respective expected base-specific intensity value for each of four different candidate nucleobases.
11. The system of claim 10, further comprising instructions that, when executed by the at least one processor, cause the system to generate the artificial base-specific intensity values with a predetermined minimum distance between respective expected base-specific intensity values corresponding to the four different candidate nucleobases.
12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the intensity-value base-decision boundaries further based on the intensity values received for the indexing cycle.
13. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to determine additional intensity-value base-decisionboundaries for a subsequent indexing cycle based on the artificial base-specific intensity values and respective intensity values received for the subsequent indexing cycle.
14. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the nucleobase call for the oligonucleotide cluster within the nucleotide-sample slide by: determining respective nucleobase probabilities for the oligonucleotide cluster based on comparing the corresponding intensity value to the intensity-value base-decision boundaries; and generating the nucleobase call for the oligonucleotide cluster based on determining a highest nucleobase probability.
15. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: receive, for a series of sequencing cycles within a sequencing run, intensity values for labeled nucleobases of oligonucleotide clusters within a nucleotide-sample slide; determine one or more expected base-specific intensity values based on one or more base-specific intensity-value distributions of at least one sequencing cycle of the series of sequencing cycles; generate, for an indexing cycle of the series of sequencing cycles, artificial basespecific intensity values based on the one or more expected base-specific intensity values; determine, for the indexing cycle, intensity-value base-decision boundaries based on the artificial base-specific intensity values; and generate, for the indexing cycle, a nucleobase call for an oligonucleotide cluster within the nucleotide-sample slide based on a corresponding intensity value and the intensity-value base-decision boundaries.
16. The system of claim 15, wherein the at least one sequencing cycle of the series of sequencing cycles comprises: a preceding indexing cycle performed before the indexing cycle; or a subset of preceding indexing cycles performed before the indexing cycle.
17. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to perform the series of sequencing cycles within the sequencing run according to an order of indexing cycles before genomic sequencing cycles by: determining nucleobase calls for a first index sequence appended to a sample genomic sequence of a sample of one or more samples;determining nucleobase calls for a second index sequence appended to the sample genomic sequence of the sample; and after determining the nucleobase calls for the first index sequence and the second index sequence, determining nucleobase calls for a first nucleotide read corresponding to a first portion of the sample genomic sequence and determining base calls for a second nucleotide read corresponding to a second portion of the sample genomic sequence.
18. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, based on the nucleobase call, an index sequence for the oligonucleotide cluster of the nucleotide-sample slide; and identify, utilizing the index sequence, a particular sample for the oligonucleotide cluster from a plurality of samples corresponding to one or more oligonucleotide clusters on the nucleotide-sample slide.
19. The system of claim 18, further comprising instructions that, when executed by the at least one processor, cause the system to generate a base-call data fde comprising a nucleotide read from the particular sample identified utilizing the index sequence.
20. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to generate, for the indexing cycle, additional nucleobase calls for additional oligonucleotide clusters within the nucleotide-sample slide based on other intensity values and the intensity-value base-decision boundaries.
21. The system of claim 15, wherein the at least one sequencing cycle of the series of sequencing cycles comprises a subset of genomic sequencing cycles preceding the indexing cycle within the sequencing run.
22. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values by: determining, from the at least one sequencing cycle of the series of sequencing cycles, the one or more base-specific intensity-value distributions; and determining, from the one or more base-specific intensity-value distributions, the one or more expected base-specific intensity values for one or more of four candidate nucleobases.
23. The system of claim 22, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values by determining an average centroid of each of the one or more base-specific intensity -value distributions determined from the at least one sequencing cycle.
24. The system of claim 23, further comprising instructions that, when executed by the at least one processor, cause the system to determine the average centroid of each of the one or more base-specific intensity-value distributions by determining, for each base-specific intensityvalue distribution, a mean intensity value for a first sequencing signal channel and a second sequencing signal channel.
25. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to generate the artificial base-specific intensity values by: generating a respective distribution of artificial intensity values for each of the one or more expected base-specific intensity values; and generating the artificial base-specific intensity values based on the respective distribution of artificial intensity values.
26. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to determine the one or more expected base-specific intensity values to include a respective expected base-specific intensity value for each of four different candidate nucleobases.
27. The system of claim 26, further comprising instructions that, when executed by the at least one processor, cause the system to generate the artificial base-specific intensity values with a predetermined minimum distance between respective expected base-specific intensity values corresponding to the four different candidate nucleobases.
28. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to determine the intensity-value base-decision boundaries further based on the intensity values received for the indexing cycle.
29. The system of claim 28, further comprising instructions that, when executed by the at least one processor, cause the system to determine additional intensity-value base-decision boundaries for a subsequent indexing cycle based on the artificial base-specific intensity values and respective intensity values received for the subsequent indexing cycle.
30. The system of claim 15, further comprising instructions that, when executed by the at least one processor, cause the system to generate the nucleobase call for the oligonucleotide cluster within the nucleotide-sample slide by: determining respective nucleobase probabilities for the oligonucleotide cluster based on comparing the corresponding intensity value to the intensity-value base-decision boundaries; and generating the nucleobase call for the oligonucleotide cluster based on determining a highest nucleobase probability.
Citation Information
Patent Citations
Method of nucleic acid amplification
US20050100900A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1