Incremental secondary analysis of nucleic acid sequences
By offloading secondary analysis to a programmable logic unit, nucleic acid sequencers can parallelize sequencing and secondary operations, reducing processing time and downtime, facilitating faster tertiary analysis and treatment decisions.
Patent Information
- Application Number
- JP2025154247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-11
- Filing Date
- 2025-09-17
- Publication Date
- 2026-01-27
AI Technical Summary
Conventional nucleic acid sequencers lack computational resources to perform primary and secondary analysis operations in parallel, leading to prolonged processing times and sequencer downtime, which reduces instrument throughput and impacts revenue streams.
Offload secondary analysis operations to a programmable logic unit with hardwired digital logic, allowing parallelization of sequencing and secondary analysis operations, reducing processing time and conserving reagents.
Dramatically reduces secondary analysis time from 56-99 hours to 2-12 hours, enabling faster tertiary analysis and reducing sequencer downtime, thereby improving throughput and enabling timely treatment decisions.
Smart Images

Figure 2026012680000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 988,374, filed March 11, 2020, the entire contents of which are incorporated herein by reference in their entirety.
[0002] The present disclosure relates to nucleic acid sequence analysis. [Background technology]
[0003] A nucleic acid sequencer is an instrument configured to automate the process of nucleic acid sequencing, which is the process of determining the order of nucleotides in a nucleic acid sequence. Nucleic acids can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
[0004] A nucleic acid sequencer is configured to receive a nucleic acid sample and generate one or more output data, called "reads," that represent the order of nucleotides in the nucleic acid sample. Nucleotides in a DNA sample can include one or more bases, including guanine (G), cytosine (C), adenine (A), and thymine (T), in any combination. Nucleotides in an RNA sample can include one or more bases, including G, C, A, and uracil (U), in any combination.
[0005] The reads generated by the DNA sequencer can be mapped to a known sequence of nucleotides in a reference genome using a mapping and alignment engine. Mapping of the reads to the known sequence of nucleotides in the reference genome can be achieved by the mapping and alignment engine using a hash table index. Summary of the Invention [Means for solving the problem]
[0006] The present disclosure relates to a system, method, and computer program for performing incremental secondary analysis. Incremental secondary analysis relates to a process of performing one or more secondary analysis operations on nucleic acid reads of a sample before nucleic acid sequencing of the sample is completed by a nucleic acid sequencer. The one or more secondary analysis operations may include nucleic acid read mapping, nucleic acid read alignment, variant calling, or any combination thereof.
[0007] According to one innovative aspect of the present disclosure, a method for performing incremental secondary analysis of nucleic acid sequence reads is disclosed. In one aspect, the method includes the actions of: (i) acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval, each of the first reads representing a first ordered sequence of nucleotides; and (ii) acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval, each of the second reads representing a second ordered sequence of nucleotides; and, while the second data is being acquired, (a) providing, by the nucleic acid sequencing device, the first data as input to a mapping and alignment unit, (b) receiving alignment results from the mapping and alignment unit, (c) storing the received alignment results, and then (iii) instructing the mapping and alignment unit to begin aligning the second data representing the second plurality of reads to a reference sequence.
[0008] Other versions include corresponding systems, apparatus, and computer programs for performing the actions of the methods defined by instructions encoded on a computer-readable storage device.
[0009] These and other versions may optionally include one or more of the following features: For example, in some implementations, at least a portion of the mapping and alignment unit is implemented using a programmable logic device.
[0010] In some implementations, the programmable circuit is a field programmable gate array (FPGA).
[0011] In some implementations, at least a portion of the mapping and alignment unit is implemented using an application specific integrated circuit (ASIC).
[0012] In some implementations, the mapping and alignment unit is comprised within a nucleic acid sequencing device.
[0013] In some implementations, one or more of the first reads include data representing a first sample identifier, and one or more of the second reads include data representing a second sample identifier.
[0014] In some implementations, the method may further include, while the second data is being acquired, organizing the one or more first reads into respective groups based on at least the first sample identifier or the second sample identifier, and generating tissue statistics, the tissue statistics indicating the number of first reads corresponding to each sample identifier.
[0015] In some implementations, the method may further include providing output data representing stored alignment results corresponding to the plurality of first reads prior to aligning the second portion of the cluster of reads or during aligning the second portion of the cluster of reads.
[0016] In some implementations, the method may further include instructing the mapping and alignment module to initiate a subsequent alignment of data representing the first plurality of reads to a reference sequence.
[0017] In some implementations, the method may further include, while acquiring the second data, determining a set of possible variants of the first data representing the first plurality of reads aligned to the reference sequence.
[0018] In some implementations, at least a portion of the second data representing the second plurality of reads is aligned while acquiring at least a different portion of the second data representing the second plurality of reads.
[0019] In some implementations, the mapping and alignment unit is instructed to begin aligning the second data representing the second plurality of reads a predetermined number of sequencing cycles before fully acquiring the second data.
[0020] According to another innovative aspect of the present disclosure, another method for performing incremental secondary analysis of nucleic acid sequence reads is disclosed. In one aspect, the method includes: (i) generating a plurality of first entity identifiers, each entity first identifier corresponding to a particular read generated during a first read interval; (ii) generating a plurality of second entity identifiers, each second entity identifier corresponding to a particular read generated during the second read interval; and (iii) acquiring first data describing the plurality of first reads generated by a nucleic acid sequencing device based on a plurality of different samples during the first read interval, each of the plurality of first reads corresponding to at least a first entity identifier or a second entity identifier; while the first data is being acquired, the method includes: organizing the plurality of first reads into organized groups based on the first entity identifier or the second entity identifier associated with each of the first reads; and, (iv) obtaining second data describing a plurality of second reads generated by the nucleic acid sequencing device based on a plurality of different samples during a second read interval performed after the first read interval, wherein each of the plurality of second reads corresponds to at least the first entity identifier or the second entity identifier; and (v) providing, by the nucleic acid sequencing device, the second data to a mapping and alignment unit configured to align the second data to the reference sequence.
[0021] Other versions include corresponding systems, apparatus, and computer programs for performing the actions of the methods defined by instructions encoded on a computer-readable storage device.
[0022] These and other versions may optionally include one or more of the following features: For example, in some implementations, at least a portion of the mapping and alignment unit is implemented using a programmable logic device.
[0023] In some implementations, the programmable circuit is a field programmable gate array (FPGA).
[0024] In some implementations, at least a portion of the mapping and alignment unit is implemented using an application specific integrated circuit (ASIC).
[0025] In some implementations, the mapping and alignment unit is comprised within a nucleic acid sequencing device.
[0026] In some implementations, organizing the plurality of first reads includes generating data indicating the number of reads corresponding to each entity identifier.
[0027] In some implementations, while acquiring the second data, for each organized set of first reads, a set of possible variants of the organized set of first reads aligned to the reference sequence is determined.
[0028] According to another innovative aspect of the present disclosure, another method for performing incremental secondary analysis of nucleic acid sequence reads is disclosed. In one aspect, the method includes the actions of acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval of a first sequencing run, acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval of the first sequencing run performed after the first read interval, starting to perform one or more secondary analysis operations on the first data or the second data while acquiring at least some of the second data, performing a second sequencing run using the nucleic acid sequencing device, continuing to perform one or more secondary analysis operations on at least the first data or the second data while performing the second sequencing run using the nucleic acid sequencing device, and storing result data representing results of the secondary analysis operations.
[0029] Other versions include corresponding systems, apparatus, and computer programs for performing the actions of the methods defined by instructions encoded on a computer-readable storage device.
[0030] According to another innovative aspect of the present disclosure, a method for performing secondary analysis of nucleic acid sequence reads is disclosed. In one aspect, the method includes the actions of obtaining one or more genomic workflow attributes, determining a workflow context switching type for a programmable circuit based on the one or more genomic workflow attributes, the workflow context switching type defining a reconfiguration cycle for the programmable circuit, and instructing a controller of the programmable circuit to perform the secondary analysis using the determined context switching type.
[0031] Other versions include corresponding systems, apparatus, and computer programs for performing the actions of the methods defined by instructions encoded on a computer-readable storage device.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Additionally, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0033] Other features and advantages of the invention will be apparent from the following detailed description, and from the claims. [Brief explanation of the drawings]
[0034] [Figure 1A] FIG. 1 is a schematic diagram illustrating an example of a prior art workflow illustrating a linear sequence of secondary analysis operations. [Figure 1B] FIG. 1 is a context diagram of an example system for performing incremental secondary analysis on one or more samples using a secondary analysis unit located within a nucleic acid sequencer. [Figure 2] 1C is a flowchart of an example process for performing incremental secondary analysis according to the workflow diagram of FIG. 1B. [Figure 3] FIG. 1 is a context diagram of an example system for performing incremental secondary analysis of one or more samples using a secondary analysis unit located remotely from a nucleic acid sequencer. [Figure 4] 4 is a flowchart of an example process for performing incremental secondary analysis according to the workflow diagram of FIG. 3. [Figure 5] FIG. 1 is a context diagram of an example system for performing incremental secondary analysis of one or more samples using a secondary analysis unit within a nucleic acid sequencer. [Figure 6] 6 is a flowchart of an example process for performing incremental secondary analysis according to the workflow diagram of FIG. 5. [Figure 7] FIG. 1 is an example of a workflow diagram illustrating a workflow of operations performed during a process for performing incremental secondary analysis using a secondary analysis unit. [Figure 8] 8 is a flowchart of an example process for performing incremental secondary analysis according to the workflow diagram of FIG. 7. [Figure 9] 1 is a flowchart of an example process for performing dynamic programmable circuit context switching. [Figure 10] FIG. 1 is a block diagram of an example of system components that can be used to implement a system for performing incremental second-order analysis. DETAILED DESCRIPTION OF THE INVENTION
[0035] Nucleic acid sequencing of biological samples using a nucleic acid sequencer is a time-consuming and costly task. Conventional systems use linear workflows, such as the one shown in Figure 1A. Such conventional workflows linearly perform operations including (i) primary analysis to generate nucleic acid sequence reads, (ii) secondary analysis of the generated nucleic acid sequencing reads to generate aligned reads and variants, and, in some cases, (iii) tertiary analysis using the results of the secondary analysis, such as variants identified during variant calling. Tertiary analysis can include, for example, classification of identified variants, determining the relevance of identified variants, determining a diagnosis based on identified variants, or determining a treatment based on identified variants.
[0036] Referring to FIG. 1A, a conventional workflow 170A is depicted for performing a sequencing run 172A of one or more samples. The sequencing run 172A includes a first read interval "Read 1" that includes a clustering operation during time T1, a sequencing operation to generate a first read for the sample during time T2A, and a second read interval "Read 2" that includes a sequencing operation to generate a second read for the sample during another time T2B. During the sequencing run 172A, a first primary analysis 100A processes the data to generate a first read and a second read. The primary analysis 100A may include, for example, processing images to generate a sequence of each nucleotide or base in the read. After the first primary analysis 100A is completed, a secondary analysis 100B begins. In this example of FIG. 1A, secondary analysis 100B is performed using the software resources of a nucleic acid sequencer and includes demultiplexing the reads generated during primary analysis 100A of the first sequencing run 172A, mapping and aligning the demultiplexed reads, and then variant calling, all during time T3. Only after the secondary analysis is complete can the next primary analysis 100C be performed by the nucleic acid sequencer. Thus, using a conventional workflow with conventional secondary analysis software on a nucleic acid sequencer, it takes at least T = T + T + T, potentially approximately 56-99 hours from the start of first primary analysis 100A of the first sequencing run 172A until the second primary analysis 100C of the second sequencing run 172B can be performed. Furthermore, this results in sequencer downtime, potentially for at least 30-48 hours, during which the sequencer does not perform secondary analysis and consumes reagents, reducing instrument throughput (the number of nucleotides processed in a given time) and negatively impacting revenue streams from reagent sales.
[0037] Conventional systems operate in this manner because conventional nucleic acid sequencers lack the computational resources to perform primary and secondary analysis operations in parallel. Instead, the software computational resources of conventional nucleic acid sequencers are dedicated to sequencing operations during primary analysis, and then these same computational resources are dedicated to demultiplexing, mapping, alignment, and variant calling operations during secondary analysis. In some implementations, demultiplexing can include a sorting operation.
[0038] The present disclosure addresses these issues by offloading aspects of secondary analysis operations to a programmable logic unit having hardwired digital logic configured to perform one or more secondary analysis operations using hardware circuitry. This dramatically reduces the time, T3, required to perform the secondary analysis operations. Furthermore, the present disclosure parallelizes sequencing operations, such as clustering, primary analysis, other sequencing operations, or a combination thereof, and the secondary analysis described herein, and reduces the overall processing time, T, from the start of a first sequencing run 172A to the start of a second sequencing run 172B by adapting conventional nucleic acid sequencing devices to perform the parallelized workflow operations described herein.
[0039] The techniques of the present disclosure can be used to obtain several other advantages. First, the present disclosure can be used to conserve reagents used by a nucleic acid sequencer during a sequencing run. For example, by starting a secondary analysis operation during a sequencing run and completing at least a portion of the secondary analysis operation before the sequencing is complete, the present disclosure can generate statistics, such as alignment statistics and demultiplexing statistics, and evaluate the generated statistics to measure the quality of the reads generated during the primary analysis. If the statistics indicate that the quality of the reads generated by the nucleic acid sequencer is insufficient, the primary analysis can be terminated, the input to the sequencer can be reconfigured, and another sequencing run using the nucleic acid sequencer can be started again. Thus, this process can save at least a portion of the reagents that would have been spent to complete the entire initial primary analysis sequencing run by stopping the primary analysis sequencing run without using all the reagents to complete the low-quality sequencing run.
[0040] Second, the parallelized workflow of the present disclosure allows tertiary analysis to begin more quickly than conventional systems, thereby enabling faster identification of specific diagnoses and treatments. For example, conventional workflows using conventional computational architectures can, in some cases, take approximately 56-99 hours to begin tertiary analysis. However, in some implementations of the present disclosure, tertiary analysis can begin in as little as 2-12 hours, or even just a few hours, after sequencing is complete. In some cases, this can be particularly advantageous, for example, providing a faster determination of whether a patient's symptoms are viral or bacterial related. However, there are multiple scenarios in which a treatment decision made in a few hours, as opposed to three to four days in some cases, can provide a significant impact, for example, allowing the patient the opportunity to administer antibiotics (or other types of drugs or treatments) before an infection (or other illness) causes irreversible damage.
[0041] These and other advantages will be apparent from the features described in this disclosure.
[0042] FIG. 1B is a context diagram of an example system 100 for performing incremental secondary analysis on a single sample 105 using a secondary analysis unit 140 located within a nucleic acid sequencer. System 100 includes a nucleic acid sequencer 110, one or more flow cells 120, one or more secondary analysis units 140, one or more processing units 150, and one or more memories 160. In the example of FIG. 1B, secondary analysis unit 140 is located within sequencer 110. However, the present disclosure is not so limited. Alternatively, secondary analysis unit 140 can be located within one or more remote computers communicatively coupled to sequencer 110 using one or more wired or wireless networks, such as a LAN, a WAN, a cellular network, the Internet, or any combination thereof. Secondary analysis unit 140 can include memory 140, a programmable circuit 142, a processing unit 150, a memory 160, or any combination thereof. For purposes of this specification, secondary analysis may include mapping operations, alignment operations, variant calling operations, or any subset or combination thereof. In some implementations, processing unit 150, memory 160, or both may be used by the nucleic acid sequencer to perform other operations not related to secondary analysis.
[0043] One or more processing units 150 of nucleic acid sequencer 110 may include one or more processors configured to execute software instructions to implement functionality defined by the software instructions. For example, one or more processing units 150 may retrieve and execute software instructions defining demultiplexing unit 162 stored in memory 160 to implement the functionality of demultiplexing unit 162. One or more processing units 150 may include one or more central processing units (CPUs), one or more graphical processing units (GPUs), or any combination thereof.
[0044] The term “unit” is used herein to describe a software module, a hardware module, or a combination of both used to perform a specified function. The determination of whether a particular “unit” described herein is hardware, software, or a combination of both can be made based on the context of its use. For example, the “mapping and alignment unit” 142a resident in the programmable circuit 142 is a hardware unit whose functionality is realized by hardwired digital logic gates or blocks. As another example, the “demultiplexing unit” 162 resident in the memory 160 is a software unit whose functionality is realized by the processing unit 150 executing software instructions that define the “demultiplexing unit” 162. As another example, the “processing unit” 150 is a hardware device that realizes its functionality by processing software instructions; therefore, the functionality of the “processing unit” 150 is a combination of hardware and software. Similarly, the “secondary analysis unit” 140 may include a combination of hardware and software used to interact with the hardwired programmable circuit 142a.
[0045] The nucleic acid sequencer 110 is a device configured to perform sequencing operations, such as a primary analysis, which may include receiving a biological sample 105, such as a blood sample, a tissue sample, or sputum, by the nucleic acid sequencer 110, and generating, by the nucleic acid sequencer 110, output data such as one or more reads 130-1, 130-2, 130-3, 130-4, 132-1, 132-2, 132-3, 132-4, 134-1, 134-2, 134-3, 134-4, each representing the order of nucleotides in a nucleic acid sequence of the received biological sample. Sequencing by the nucleic acid sequencer 110 can be performed at multiple read intervals, with a first read interval "Read 1" generating one or more first reads representing a first portion, or nucleotide order from the end, of the nucleic acid sequence fragments (or strands) clonally amplified for the clonal population of template nucleic acid fragments bound to the flow cell 120, and a second read interval "Read 2" generating one or more second reads each representing a second portion, e.g., nucleotide order from the other end, of the nucleic acid sequence fragments clonally amplified for the clonal population of template nucleic acid fragments bound to the flow cell 120. Each clonal population of template nucleic acid fragments bound to the flow cell 120 may be referred to herein as a cluster, such as cluster 1 122-1, cluster 2 122-2, cluster 3 122-3, cluster 4 122-4, cluster 5 122-5, cluster N 122-N, etc.
[0046] As a result, during each read interval, a single read will be generated by the nucleic acid sequencing device 110 for each end of the clonally amplified nucleic acid fragment in each cluster. That is, the first read interval of a sequencing cycle generates "Read 1," and the second read interval of a sequencing cycle generates "Read 2." In some implementations, the nucleic acid sequencer may image the read sequence and determine the read sequence, or sequence multiple clones of nucleic acid fragments in the same cluster to identify the read sequence.
[0047] Thus, each read represents a portion of a particular nucleic acid sequence fragment. For example, assuming a short nucleic acid sequence fragment of approximately 600 nucleotides, a first read may represent 150 ordered nucleotides at a first end of the nucleic acid sequence fragment, and a second read may represent 150 ordered nucleotides at the other end of the nucleic acid sequence fragment. However, these numbers are merely examples, and the nucleic acid sequencer 110 may be configured in a manner consistent with the spirit and scope of the present disclosure to generate shorter nucleic acid sequences and respective reads of different lengths than those mentioned herein. To convey the principles of the present disclosure to those skilled in the art, a simple version of this concept is illustrated with reference to FIGS. 1B, 3, and 5. Specifically, these figures show nucleic acid templates bound to a flow cell 120 and reads generated by the nucleic acid sequencer 110 at each end of clonally amplified clustered nucleic acid sequence fragments.
[0048] In some implementations, the biological sample can include a DNA sample, and the nucleic acid sequencer 110 can process DNA. In such implementations, the order of sequenced nucleotides in reads 130-1, 130-2, 130-3, 130-4, 132-1, 132-2, 132-3, 132-4, 134-1, 134-2, 134-3, and 134-4 generated by the nucleic acid sequencer can include one or more of guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. In other implementations, the nucleic acid sequencer 110 can process RNA, and the biological sample can include an RNA sample. In such RNA implementations, the order of sequenced nucleotides in reads generated by the nucleic acid sequencer can include one or more of G, C, A, and uracil (U) in any combination. Thus, although the example in Figure 1 describes processing reads consisting of G, C, A, and T that are based on a DNA sample, the disclosure is not so limited. Instead, other implementations can process reads consisting of C, G, A, and U that are based on an RNA sample.
[0049] However, RNA sequencing does not require the use of an RNA sequencer. For example, in some implementations, the nucleic acid sequencer 110 can be a DNA sequencer that sequences samples and generated reads having one or more of G, C, A, and T. In such implementations, the nucleic acid sequencer 110 can then transcribe the generated reads into cDNA to represent the RNA of the sequenced sample. In such implementations, the reads are represented using bases including G, C, A, and uracil (U) in any combination.
[0050] In some implementations, the nucleic acid sequencer 110 can include a next generation sequencer (NGS) configured to generate sequence reads 130-1, 130-2, 130-3, 130-4, 132-1, 132-2, 132-3, 132-4, 134-1, 134-2, 134-3, 134-4, etc. for a given sample in a manner that achieves ultra-high throughput, scalability, and speed through the use of massively parallel sequencing technologies. NGS enables rapid sequencing of entire genomes and the ability to zoom in on deeply sequenced target regions or utilize RNA sequencing (RNA-Seq) to discover novel RNA variants and splice sites, or for gene expression analysis, analysis of epigenetic factors such as genome-wide DNA methylation and DNA-protein interactions, sequencing of cancer samples to study rare-body variants and tumor subclones, and quantifying mRNA for the study of microbial diversity in humans or the environment.
[0051] The process of generating nucleic acid sequencing reads includes sample preparation, cluster generation, and sequencing. The first step involves sample preparation, which involves adding adapter sequences to the ends of each DNA fragment. Reduced cycle amplification introduces additional motifs, such as any necessary indexes, that can be used to identify the sample from which the read originated and the region complementary to the oligos on the flow cell 120. One or more examples of sample preparation on a solid support are described in U.S. Patent No. 9,683,230, which is incorporated herein by reference in its entirety. The second step involves clustering, in which each DNA fragment is isothermally amplified, for example, using amplification reagents. One or more examples of isothermal amplification of nucleic acids on a solid support are described in more detail in U.S. Patent No. 7,972,820, which is incorporated herein by reference in its entirety. The flow cell 120 can include a glass slide with multiple lanes, each containing a lawn of two oligos. Hybridization is enabled by the attachment of the first of the two oligos to the complementary oligo on the surface of the flow cell. The polymerase forms complements of the hybridized fragments. The DNA fragments can be clonally amplified using techniques such as bridge amplification. In implementations of system 100 and workflow 170B, the clustering step occurs during time T1 of workflow 170B. However, the present disclosure is not so limited. Instead, in some implementations, clustering may begin and be performed before time T1, off-instrument, or both. In such implementations, time T1 can be removed from the run time calculation, and the sequencing run can begin, for example, at T2A. Such pre-T1 and / or off-instrument clustering can be implemented in system 100 of FIG. 1, system 300 of FIG. 3, system 500 of FIG. 5, system 700 of FIG. 7, or any other implementation of the present disclosure. After bridge amplification, the reverse fragments are cleaved, leaving only the forward fragments.
[0052] The third phase involves the performance of sequencing operations by the nucleic acid sequencer 110 during times T2A and T2B. During time T2A, the nucleic acid sequencer 110 performs X cycles of sequencing operations for a first read interval, "Read 1," to generate first reads corresponding to the first ends of each clonally amplified nucleic acid sequence fragment in each cluster 122-1, 122-2, 122-3, 122-4, 122-5, 122-N, where X and N can be any positive integers greater than zero. The first read for each DNA cluster includes a string of base calls corresponding to the respective DNA portions associated with the particular cluster. For example, read 130-1 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster 1 122-1, read 130-3 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster 2 122-2, read 132-1 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster 3 122-3, read 132-3 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster 4 122-4, read 134-1 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster 5 122-5, and read 134-3 includes a string of base calls corresponding to a first end of a nucleic acid fragment associated with cluster N 122-N. Each base call corresponds to or represents a nucleotide. These reads can be generated using a sequencing process, such as sequencing-by-synthesis. Data representing reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 may be output to memory 160 of nucleic acid sequencer 110, input to memory 144 of secondary analysis unit 140, or both.
[0053] In the implementation of system 100 and FIG. 1B , the first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 sequenced during time T2A of the first read interval of workflow 170B represent the number of nucleotides at the first end of the DNA fragment associated with each cluster. For example, in some implementations, the DNA fragments sequenced by nucleic acid sequencer 110 may contain 600 nucleotides. The first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 clusters may represent, for example, the first 150 nucleotides at the first end of the 600-nucleotide DNA fragments amplified in each cluster. Each read interval is a massively parallel process that simultaneously sequences hundreds of millions of DNA fragment clusters. Once the first read interval is completed at the end of T2A, the nucleic acid sequencer 110 can begin a second read interval during time T2B to sequence the opposite end of each DNA fragment in each cluster, generating second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4. As an example, read 130-2 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster 1 122-1, read 130-4 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster 2 122-2, read 132-2 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster 3 122-3, read 132-4 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster 4 122-4, read 134-2 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster 5 122-5, and read 134-4 includes a string of base calls corresponding to the second end of the nucleic acid fragment associated with cluster N 122-N. In this implementation of system 100 and FIG. 1, the second read interval begins at approximately time T1 + T2A in workflow 170B.
[0054] In the conventional system described with reference to FIG. 1A, secondary analysis operations such as mapping and aligning the first leads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 do not occur until after the end of the second lead interval "Lead 2" at the end of time = T1 + T2A + T2B. However, the system 100 of FIG. 1B described by the present disclosure is configured to start secondary analysis operations of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 at time = T1 + T2A, and during the second read interval "Read 2", secondary analysis of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 begins and is performed, while the nucleic acid sequencer 110 performs sequencing operations of the second read interval "Read 2" to generate second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4.
[0055] System 100 takes advantage of this parallel processing by offloading secondary analysis operations of the first read to programmable circuitry 142a of secondary analysis unit 140. Offloading secondary analysis operations to secondary analysis unit 140 frees up the processing unit 150, memory 160, or both of nucleic acid sequencer 110 to continue performing primary analysis operations for the second read interval, "Read 2," to generate second reads 130-2, 130-4, 132-2, 132-4, 134-2, and 134-4 by sequencing opposite ends of DNA clusters while secondary analysis of one or more of the first reads is being performed. Thus, the present disclosure enables sequencing operations, such as primary analysis, to be performed in parallel with one or more secondary analysis operations.
[0056] The secondary analysis unit 140 includes a programmable circuit 142 that can be dynamically configured to include one or more secondary analysis operation units, such as a mapping and alignment unit 142a, to perform one or more secondary analysis operations. Dynamically configuring the programmable circuit 142 to include a secondary analysis operation unit, such as a mapping and alignment unit 142a, can include, for example, providing one or more instructions to the programmable circuit 142 that cause the programmable circuit 142 to configure its hardware logic gates as hardwired digital logic configurations configured to implement the functionality of the mapping and alignment unit 142a in hardware logic. The hardware logic gates of the programmable circuit 142 can be implemented using compiled hardware description language code, or the like. The initial configuration of the programmable circuit 142 and subsequent reconfiguration of the programmable circuit 142 can be initiated by the execution of a software trigger satisfied by the nucleic acid sequencer 110 or other computer that hosts the programmable circuit 142. 1B , at the end of a read 1 interval cycle, the nucleic acid sequencer 110 or other computer hosting the programmable circuit 142 can execute software instructions that trigger a reconfiguration of the programmable circuit to perform the mapping and alignment operations. Such execution of the aforementioned software trigger can, for example, be performed by the programmable circuit control and can cause the loading of compiled hardware description language code into the memory of the programmable circuit 142, which can cause the reconfiguration of logic gates in the programmable circuit 142.The configured functions of the mapping and alignment unit 142a may include obtaining one or more reads, such as first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3, mapping the obtained first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 to one or more reference sequence positions, and then aligning the mapped first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 to the one or more reference sequence positions. The reference sequence may include an organized series of nucleotides corresponding to a known genome.
[0057] Configuring the hardware logic gates of the programmable circuit 142 in response to one or more instructions can include configuring logic gates, such as AND gates, OR gates, NOR gates, XOR gates, or any combination thereof, to perform the digital logic functions of the mapping and alignment unit 142a. Examples of the use of programmable logic circuits, such as FPGAs, to perform the functions of the mapping and alignment unit are described in more detail in, for example, U.S. Pat. No. 9,679,104 or U.S. Patent Application Publication No. 2020 / 0372031, each of which is incorporated by reference herein in its entirety. Alternatively, or in addition, configuring the hardware logic gates can include dynamically configured logic blocks that include customizable hardware logic units for performing complex computational operations, including addition, multiplication, comparison, etc. The exact configuration of the hardware logic gates, logic blocks, or combinations thereof is defined by the received instructions. The received instructions may include, or may be generated from, compiled hardware description language (HDL) program code that defines a schematic layout of the secondary analysis operational unit to be written and programmed by the entity. The HDL program code may include program code written in a language such as Very High Speed Integrated Circuit Hardware Description Language (VHDL), Verilog, etc. The entity may include one or more human users who drafted the HDL program code, one or more artificial intelligence agents that generated the HDL program code, or a combination thereof.
[0058] In some implementations, the programmable circuitry 142 can include one or more field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), or programmable logic arrays (PLAs), or combinations thereof, which are dynamically configurable and reconfigurable by the nucleic acid sequencer 110 as needed to perform a particular workflow. For example, in some implementations, it may be desirable to use the programmable logic circuitry 142 as a mapping and alignment unit 142a, as described above. However, in other implementations, it may be desirable to use the programmable circuitry 142 to perform variant calling functions or functions to aid in variant calling, such as a Hidden Markov Model (HMM) unit. In yet other implementations, the programmable circuit 142 may be dynamically configured to support common computational tasks such as compression and decompression, because the hardware logic of the programmable circuit 142 can perform these tasks, and other tasks identified above, much faster than performing the same tasks using software instructions executed by one or more processing units 150.
[0059] The programmable circuit 142 is one example of a type of integrated circuit that can provide the benefits of the present disclosure described herein. However, other types of integrated circuits can be used as the hardwired digital logic of the secondary analysis unit 140, which can offload secondary analysis from the nucleic acid sequencer 110 and free up the nucleic acid sequencer 110's resources for primary analysis. For example, in some implementations, the secondary analysis unit 140 can be configured to use one or more application-specific integrated circuits (ASICs). The one or more ASICs are not reprogrammable, but can be designed with custom hardware logic for one or more secondary analysis operation units, such as a mapping and alignment unit, a variant calling unit, or a variant calling computational support unit, to accelerate and parallelize the execution of secondary analysis operations. In some implementations, using an ASIC as the hardwired logic circuit of the secondary analysis unit 140 that implements the functionality of one or more secondary analysis operation units can be even faster than using a programmable circuit. Therefore, those skilled in the art will understand that an ASIC can be used in place of an FPGA in any of the implementations described herein.
[0060] For example, in some implementations, the programmable logic circuit 142 may be implemented using an FPGA dynamically configured as a reconstruction unit to access data representing first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 received from the nucleic acid sequencer and reconstruct the data representing the first reads (e.g., if the reads received from the nucleic acid sequencer are compressed). The reconstruction unit may store the reconstructed reads stored in memory 144 or memory 160. In such implementations, the FPGA may then be dynamically reconfigured as a mapping and alignment unit 142a to perform mapping and alignment of the reconstructed first reads stored in memory 144 or memory 160. The mapping and alignment unit 142a may then store data representing the mapped and aligned reads in memory 144 or memory 160. The FPGA can then be dynamically reconfigured into a variant calling unit, or a unit configured to perform functions supporting a software variant calling unit (e.g., an HMM unit), to perform variant calling operations and generate output data that can be used by the sequencing system 100, such as generating a Variant Calling Format (VCF) file based on stored data representing mapped and aligned reads. The fast execution speed of these hardware modules implemented using FPGAs enables secondary analysis of reads to be performed in minutes, instead of the 30-48 hours required for conventional methods. While a series of operations, including recovery, mapping, alignment, and variant calling operations, is described, the present disclosure is not limited to performing all of these operations. Instead, the programmable circuit 142 can be dynamically configured to execute any operational unit in any order as needed to parallelize secondary analysis offloaded from the nucleic acid sequencer 110.
[0061] 1A , the programmable circuit 142 of the secondary analysis unit 140 of the nucleic acid sequencer 110 can be configured to include a mapping and alignment unit 142a. The nucleic acid sequencer 110 can receive a sample 105, such as a nucleic acid of an entity such as a human, a non-human animal, or a plant. The nucleic acid sequencer 110 can prepare the sample 105 and perform cluster generation during time T1 of workflow 170B. The nucleic acid sequencer 110 can perform a sequencing operation, such as sequencing-by-synthesis, during a first read interval to generate first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 during time T2A, which occurs subsequent to time T1. At the end of time T1+T2A, the nucleic acid sequencer 110 completes sequencing of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 and begins sequencing the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4.
[0062] The nucleic acid sequencer 110 is configured to parallelize secondary analysis operations, such as mapping and alignment of first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3, and sequencing operations, such as sequencing by synthesis of second read intervals to generate second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4 during time T2B. The mapping and alignment unit 142a generates mapping and alignment results 149, which can be stored in memory 160 of the nucleic acid sequencer 110, memory 144, some other memory accessible to the nucleic acid sequencer 110, other memory accessible to a user of the nucleic acid sequencer 110, or a combination thereof. The results 149 may include data describing mapping and alignment statistics, such as, for example, a Mapping Quality (MAPQ) score that provides an indication of mapping quality, an alignment score that provides an indication of alignment quality, and the like.
[0063] 1A, the ultra-fast execution time of the mapping and alignment unit 142a, implemented using hardwired digital logic in the programmable circuit 142, allows the mapping and alignment unit 142a to perform mapping and alignment of first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 and perform second read intervals in a fraction of the time required by the nucleic acid sequencer 110. For example, in some implementations, the programmable circuit 142 can perform mapping and alignment of first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 in a matter of minutes, while sequencing of second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4 can take 12 to 24 hours. Thus, the mapping and alignment results 149 can be evaluated by the nucleic acid sequencer 110, a user of the nucleic acid sequencer 110, or both, and a decision can be made as to whether the nucleic acid sequencer 110 should continue sequencing the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4 based on the quality of the mapping and alignment of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 as indicated by the mapping and alignment statistics.
[0064] This decision as to whether to continue sequencing the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4 can be made automatically by the nucleic acid sequencer 110, manually by a user of the nucleic acid sequencer 110, or based on data describing the decision from both. As an example, the nucleic acid sequencer 110 can be configured to determine whether mapping and alignment statistics, such as alignment scores, of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 meet a predetermined threshold. If one or more alignment scores meet the predetermined threshold, then the nucleic acid sequencer 110 can continue sequencing the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4. Alternatively, if it is determined that one or more alignment scores do not meet a predetermined threshold, the nucleic acid sequencer 110 can terminate sequencing of the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4.
[0065] As another example, in some implementations, the mapping and alignment results 149 can be manually reviewed by a user of the nucleic acid sequencer 110. In such an example, the user can determine whether the nucleic acid sequencer 110 should continue sequencing the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4 based on the quality of the alignment of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 as indicated by the alignment scores.
[0066] As yet another example, both the nucleic acid sequencer 110 and the user can determine whether sequencing of the second read should continue based on the quality of the alignment of the first read as indicated by the alignment score indicated by the mapping and alignment results 149. In such implementations, data describing the decisions of the nucleic acid sequencer 110 and the user can be obtained, and in some implementations, the nucleic acid sequencer 110 ends the second read interval only if both the nucleic acid sequencer 110 and the user agree that the second read interval should end.
[0067] In yet other implementations, a weighted average of the two decisions can be calculated to generate a total score representing the decisions of both the nucleic acid sequencer 110 and the user. In such implementations, the nucleic acid sequencer 110 can terminate only if the total score fails to meet a predetermined quality threshold. In still other implementations, data representing alignment statistics, data representing a user decision regarding whether to continue sequencing the second read interval, data representing one or more of the first reads, other data such as data representing features of the sample 105, or a combination thereof, can be vectorized and input to an artificial intelligence agent, such as a machine learning model, that is trained to determine whether the nucleic acid sequencer 110 should continue primary analysis of the second read interval. In such implementations, the machine learning model can be pre-trained based on labeled training data tagged with "end second read interval" or "continue at second read interval," or their respective synonyms. The labeled training data can include data representing the same input types provided to the machine learning model at runtime. Such input types may include data representing alignment statistics, data representing a user decision as to whether sequencing of the second read interval should continue, data representing one or more of the first reads, other data such as data representing characteristics of the sample 105, or combinations thereof.
[0068] The mapping and alignment results 149 generated based on the mapping and alignment of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 to one or more reference sequences are used to enable conservation of reagents used by the nucleic acid sequencer 110 during the second read interval that generates the second reads 130-2, 130-4, 132-2, 132-4, 134-2, 134-4. For example, poor alignment scores of the first reads 130-1, 130-3, 132-1, 132-3, 134-1, 134-3 can indicate the presence of a number of problems, such as a contaminated sample 105, sequencing errors, or a combination thereof. Thus, in such instances, instead of using potentially expensive reagents to sequence the second read during the second read interval and further delaying the time it would take to begin performing another round of primary analysis, the nucleic acid sequencer 110 can be shut down, reconfigured, and then used to begin primary analysis of another sample in a fraction of the time it would take for the nucleic acid sequencer 110 to complete that poor-quality sequencing run. In some implementations, once the quality of the mapping and alignment of the first read is determined to be satisfactory, the nucleic acid sequencer 110 can discard the mapping and alignment results 149. In other implementations, the mapping and alignment of the first read performed in parallel with the second read interval can be used as the mapping and alignment results for the final data run of the first read.
[0069] 1B , after the mapping and alignment results are determined to be satisfactory, the nucleic acid sequencer 110 can continue to run a second read interval to generate second reads. Once the second reads 130-2, 130-4, 132-2, 132-4, 134-2, and 134-4 are generated, the nucleic acid sequencer 110 can instruct the secondary analysis unit 140 to initiate a final secondary analysis data run of the secondary analysis unit 140. The final secondary analysis data run can include mapping and aligning the first reads 130-1, 130-3, 132-1, 132-3, 134-1, and 134-3 and the second reads 130-2, 130-4, 132-2, 132-4, 134-2, and 134-4 using the secondary analysis unit 140. Because these secondary analysis operations are implemented using programmable circuitry 142a, these secondary analysis operations can be performed in parallel with the second sequencing run in a fraction of the time required to perform the second sequencing run.
[0070] This provides an advantage over conventional systems, which allow a system to move on to a subsequent sequencing run while secondary analysis of the reads from a previous sequencing run is being performed. Specifically, as shown in FIG. 1A, while conventional nucleic acid sequencers require a system to wait approximately 24-48 hours after the completion of a first sequencing run before initiating a second sequencing run, the nucleic acid sequencer 110 can parallelize the secondary analysis of the reads from the first sequencing run with the execution of the second sequencing run using the mapping and alignment unit 142a implemented in the programmable circuit 142. Thus, the nucleic acid sequencer 110 of FIG. 1B can be used to perform more sequencing runs in a shorter period of time than conventional systems using the system and workflow described in FIG. 1A. Therefore, parallelization of sequencing runs and secondary analysis by offloading secondary analysis computational tasks to the programmable circuit 142 of the secondary analysis unit 140 can generate increased revenue from additional reagent sales.
[0071] In some implementations, the nucleic acid sequencer 110 can also have software programs, such as a demultiplexing unit 162 and a variant calling unit 164, stored in the memory 160. One or more processors 150 of the nucleic acid sequencer can process the software instructions of these units to implement the functions of these units. For example, in some implementations, DNA fragments of multiple samples can be sequenced simultaneously using the nucleic acid sequencer 110. In such an example, the demultiplexing unit 162 can be used to implement a demultiplexing technique that organizes reads based on an index, such as a barcode, added to each generated read and identifies the sample associated with each read. As another example, the processor 150 can be used to execute a variant calling unit 164 that can analyze the mapped and aligned reads to identify the occurrence of any variants, such as single nucleotide polymorphisms (SNPs), insertions / deletions (indels), and structural polymorphisms. In some implementations, the programmable circuit 142 can be dynamically reconfigured to assist in variant calling. For example, the programmable circuit 142 can be dynamically reconfigured to include an HMM unit that can be used to perform probability calculations on the likelihood of occurrence of variants at one or more reference positions in the mapped and aligned reads. In some implementations, the variant calling unit 164 can be configured to perform variant calling operations on the mapped and aligned reads from the interval of Read 1 in parallel with the sequencing operations of the second sequencing run.
[0072] The example in FIG. 1B describes an example with reads having 8 nucleotides. However, the present disclosure is not so limited. Instead, this simple example is presented to explain the features of the present disclosure in an easy-to-understand manner. In practice, in some implementations, each DNA fragment of the present disclosure may have, for example, up to 600 nucleotides, up to 1000 nucleotides, or more, and each read of a fragment may have, for example, 50 nucleotides, 75 nucleotides, 150 nucleotides, 200 nucleotides, 300 nucleotides, 500 nucleotides, or more from each end of the DNA fragment. However, implementations of the present disclosure with DNA fragments of different lengths and reads of different lengths can be used. Similarly, FIG. 1B or any other diagram should not be construed as limiting the number of clusters of fragments. For example, the nucleic acid sequencer 110 can perform massively parallel sequencing, in which millions of clusters of multiple fragments are sequenced simultaneously.
[0073] 2 is a flowchart of an example process 200 for performing incremental secondary analysis according to the workflow diagram of FIG. 1B. Generally, process 200 includes acquiring first data representing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval (210), acquiring second data representing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval (220), and while the second data is being acquired in step 220, (I) performing one or more secondary analysis operations on the first data representing the plurality of first reads generated by the nucleic acid sequencer, (II) storing results of the secondary analysis of the first plurality of reads (230), and then performing a secondary analysis of the obtained second data representing the second plurality of reads to reference data. For convenience, these steps are described in more detail below as being performed by a sequencing system such as system 100 of FIG. 1B.
[0074] The sequencing system can begin executing process 200 by acquiring 210 first data representing a plurality of first reads generated by the nucleic acid sequencing device during a first read interval. Acquiring the first data can include storing the first data representing the plurality of first reads in a memory device, such as a memory device of a secondary analysis unit, after the first data are generated by the nucleic acid sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform secondary analysis operations. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof. Each read of the plurality of first reads can consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides can correspond to a nucleotide at a first end of a nucleic acid fragment. The nucleic acid sequencing device can include any nucleic acid sequencing device, including a sequencer capable of sequencing either DNA or RNA.
[0075] The sequencing system can continue executing process 200 by acquiring 220 second data representing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval. Acquiring the second data can include storing the second data representing the plurality of second reads in a memory of the secondary analysis unit after the second data is generated by the sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform the secondary analysis operation. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof. In some implementations, at least a portion of the second data is acquired while another portion of the second data is being generated by the nucleic acid sequencing device. Each read of the plurality of second reads can consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides can correspond to nucleotides at a second end of the nucleic acid fragment opposite the first end of the nucleic acid fragment.
[0076] While the sequencing system is acquiring the second data in step 220, the sequencing system can perform one or more secondary analysis operations on the first data representing the plurality of first reads in step 230. In some implementations, performing the one or more secondary analysis operations on the first data representing the plurality of first reads can include (i) providing, by the nucleic acid sequencing device, the first data to a mapping and alignment unit to align the first data representing the plurality of first reads to a reference sequence, (ii) aligning the first data representing the plurality of first reads to the reference sequence using the mapping and alignment unit, (iii) receiving alignment results from the mapping and alignment unit, and (iv) storing the received alignment results of the alignment of the first data representing the plurality of first reads to the reference sequence before completing acquisition of the second data in step 204. The alignment results can include alignment statistics describing the quality of the alignment of the first data representing the first plurality of reads to the reference sequence. The alignment statistics can include, for example, one or more of a MAPQ score, an alignment score, etc. In other implementations, the alignment results can include mapped and aligned reads that can be provided as input to variant calling for the determination of potential variants.
[0077] In some implementations, the output data describing the alignment results can be provided for review by one or more human users. For example, the output data describing the alignment results can be output, for example, on a display connected to the nucleic acid sequencing device or provided in a separate room or building. Alternatively, or in addition, the output data describing the alignment results can be output to print a report describing the alignment results, for example, using a printer communicatively connected directly or indirectly to the nucleic acid sequencing device.
[0078] In some implementations, at least a portion of the mapping and alignment unit is implemented in an integrated circuit, such as a programmable circuit or ASIC, incorporated into the nucleic acid sequencing device. For example, the programmable circuit or ASIC may implement a table lookup function, a Smith-Waterman algorithm, or quality score determination. However, in other implementations, one or more operations of the mapping and alignment unit may be performed in software executed by the nucleic acid sequencing device. For example, controlling the programmable circuit and sorting of alignment results may be implemented in software. In still other implementations, the mapping and alignment unit may be implemented in a programmable circuit, ASIC, executable software, or a combination thereof, of one or more remote computers communicatively connected to the nucleic acid sequencing device using one or more networks. In such implementations, data representing reads, alignment results, etc. may be communicated between the nucleic acid sequencing device and one or more remote computers hosting the mapping and alignment unit using one or more networks.
[0079] A sequencing system, other processing system, or one or more human users can evaluate the alignment results while the second data is being acquired in stage 220. For example, the alignment results can be evaluated to determine whether the alignment is of sufficient quality to continue acquiring second data in stage 220. In some implementations, if the alignment results of the first plurality of reads fail to meet a predetermined threshold, the nucleic acid sequencer can be instructed to stop acquiring second data in stage 220. Alternatively, if it is determined that the alignment results of the first plurality of reads meet a predetermined threshold, then the nucleic acid sequencer can be allowed to continue acquiring second data in stage 220.
[0080] In other implementations, the mapped and aligned first read can be evaluated for detection of potential variants between the mapped and aligned first read and one or more reference sequences while the second data is being acquired in stage 220. Such implementations allow tertiary analysis of the mapped and aligned first read to be accomplished more quickly than conventional methods that prohibit initiation of tertiary analysis until after both the first read interval and the second read interval are complete. Thus, an initial diagnosis and initiation of treatment can be obtained 12 to 24 hours or more earlier, since there is no need to wait for the second read interval to be completed before proceeding to tertiary analysis.
[0081] The sequencing system can continue execution of process 200 by instructing the execution of a secondary analysis operation on the second data in step 240, for example, by instructing the mapping and alignment unit to begin aligning the second data representing the second plurality of reads to the reference sequence. In some implementations, the sequencing system 200 can always proceed to step 240. Such an implementation further provides the technical advantage of facilitating tertiary analysis and reducing downtime of the nucleic acid sequencing device. However, in other implementations, execution of process 200 may continue by instructing the mapping and alignment unit to begin aligning the second data representing the second plurality of reads to the reference sequence only if the received alignment results describing the quality of the alignment of the first data representing the plurality of first reads are determined to meet a predetermined quality threshold.
[0082] In some implementations, the sequencing system may rely on the results of the secondary analysis of the first data, mapping and alignment, variant calling, or both, performed in stage 220 while the second data is being acquired. In other implementations, these initial secondary analysis results associated with the first data performed in stage 230 may be discarded after they are evaluated to determine the quality of the first read interval. In such examples, the sequencing system may begin the second iteration of the secondary analysis of the first data either before or after performing the secondary analysis of the second data in stage 240.
[0083] FIG. 3 is a context diagram of an example system 300 for performing incremental secondary analysis of one or more samples using a secondary analysis unit 340 located remotely from a nucleic acid sequencer 310. System 300 is generally the same as system 100 described with reference to FIG. 1B, with some modifications. One modification is that secondary analysis unit 340 is located on one or more computers 320 that are remote from nucleic acid sequencer 310. For any reference numbers in FIG. 3 not explicitly stated, the components identified by the reference numbers have the same features as the corresponding features in FIG. 1. For example, each cluster 322-1, 322-2, 322-3, 322-4, 322-5, and 322-N have the same meaning as each cluster 122-1, 122-2, 122-3, 122-4, 122-5, and 122-N in FIG. 1, unless additional or different features are described with reference to FIG. 3.
[0084] Another difference between the example of FIG. 3 and the example of FIG. 1B is that in the example of FIG. 3, multiple samples are processed. As a result, reads generated by the nucleic acid sequencer 310 of the system 300 have an index generated for each read. This index is represented in FIG. 3 by labels S1, S2, and S3 attached to each read. In this example, S2, S2, and S3 are strings used to identify reads generated based on the first sample, the second sample, or the third sample, respectively. While the indexes are described herein using the terms S1, S2, and S3, these terms are used as examples to explain the concept of an index, and the present disclosure is not limited to the use of text strings as sample identifiers. Instead, in some implementations, barcodes or other data can be used as sample identifiers for reads. In some implementations, sample identifiers can be generated by adding synthetic nucleotides representing an index to each generated read.
[0085] Referring to the example of FIG. 3 , the nucleic acid sequencer 310 or the remote computer 320 can configure the programmable circuit 342 of the secondary analysis unit 340 to include a mapping and alignment unit 342a. The nucleic acid sequencer 310 can receive multiple samples 105, 106, and 107. The samples 105, 106, and 107 can include, for example, nucleic acid samples from different entities. The different entities can be different humans, different animals, different plants, etc. The nucleic acid sequencer 310 can prepare the samples 105, 106, and 107 and perform cluster generation during time T1 of the workflow 370. The nucleic acid sequencer 310 can perform a sequencing operation, such as sequencing-by-synthesis of a first read interval, to generate first reads 330-1, 330-3, 332-1, 332-3, 334-1, and 334-3 during time T2A subsequent to time T1. At the end of time T1+T2A, the nucleic acid sequencer 310 completes sequencing of the first reads 330-1, 330-3, 332-1, 332-3, 334-1, and 334-3 and begins indexing for the first read generated during the first read interval during time T3A. At the end of time T1+T2A+T3A, the nucleic acid sequencer 310 completes indexing for the first read cycle and begins indexing for the second read generated during the second read interval during time T3B. At the end of time T1+T2A+T3A+T3B, the nucleic acid sequencer 310 begins sequencing the second reads 330-2, 330-4, 332-2, 332-4, 334-2, and 334-4.
[0086] The nucleic acid sequencer 310 is configured to parallelize secondary analysis operations, such as mapping and alignment, of the first reads 330-1, 330-3, 332-1, 332-3, 334-1, and 334-3 while the nucleic acid sequencer 310 performs sequencing operations, such as sequencing-by-synthesis of second read intervals, to generate second reads 330-2, 330-4, 332-2, 332-4, 334-2, and 334-4 during time T2B. This process is similar to that described with reference to the example of FIG. 1B. However, in the example of FIG. 3, multiple samples are sequenced. Therefore, the multiple first reads need to be demultiplexed into groups based on the index of each read before proceeding to other secondary analysis operations, such as mapping and alignment and variant calling. Once the multiple first reads are demultiplexed, one or more secondary analysis operations can be performed on the demultiplexed groups of first reads. In some implementations, the system 300 can generate demultiplexing statistics based on the demultiplexing operation and can evaluate the stored statistics to determine the quality of the sequenced reads.
[0087] 3, secondary analysis of the first reads cannot begin until the end of time T1+T2A+T3A+T3B because organization of the first reads into demultiplexed groups cannot occur until indexing operations during times T3A and T3B are complete. Once the second indexing is complete at the end of time T1+T2A+T3A+T3B, the nucleic acid sequencer 310 can provide the plurality of first reads to a remote computer 320 on the network 112. The remote computer 320 can receive the plurality of first reads and store the plurality of first reads in memory 344. While the nucleic acid sequencer 310 is performing the second read interval during time T2B, the secondary analysis unit 340 can use the processing unit 350 to access the plurality of first reads in the memory 344 and use the demultiplexing unit 362 to demultiplex the plurality of first reads 330-1, 330-3, 332-1, 332-3, 334-1, 334-3 into groups based on the index or sample identifier of each read. Demultiplexing can be achieved using a demultiplexing operation to organize the first reads based on the index. The demultiplexed first reads can be stored in the memory 344. The mapping and alignment unit 342a can then access the reads stored in the memory 344 and perform mapping and alignment operations on the demultiplexed first reads during the second read interval.
[0088] The secondary analysis unit 340 can generate statistics that can be used to evaluate the quality of the reads generated by the nucleic acid sequencer. In some implementations, the secondary analysis unit can generate demultiplexing statistics based on the demultiplexing operation. The mapping and alignment unit 342a can generate mapping and alignment results and statistics for each group of first reads stored in memory 344. The mapping and alignment unit 342a can store the results 359 in memory 360 or return the results 359 to the nucleic acid sequencer 310.
[0089] The results 359 can include demultiplexing statistics, mapping and alignment results, mapping and alignment statistics, variant call statistics, or any combination thereof. The demultiplexing statistics can include the number of reads corresponding to each sample identifier. The mapping and alignment results can include data representing one or more mapped reads to a reference sequence. The mapping and alignment statistics can include data describing, for example, a MAPQ score that provides an indication of mapping quality, an alignment score that provides an indication of alignment quality, etc. The nucleic acid sequencer 310 can receive the results 359 and store the received results in memory 160.
[0090] In the example of Figure 3, the ultra-fast execution time of the mapping and alignment unit 342a, implemented using hardwired logic in the programmable circuit 342, allows the mapping and alignment unit 342a to perform mapping and alignment of each demultiplexed group of first reads 330-1, 330-3, 332-1, 332-3, 334-1, 334-3 in a fraction of the time required by the nucleic acid sequencer 310 to perform the second read interval. For example, in some implementations, the programmable circuitry 342a can map and align the demultiplexed group of first reads 330-1, 330-3, 332-1, 332-3, 334-1, 334-3 in just a few minutes, while sequencing the second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4 during the second read interval can take 12 to 24 hours. Thus, the results 359 can be evaluated by the nucleic acid sequencer 310, the remote computer 320, a user of the nucleic acid sequencer 310 or a user of the remote computer 320, an artificial intelligence agent or model, or a combination thereof, and a decision can be made based on the quality of the demultiplexing of the first reads 330-1, 330-3, 332-1, 332-3, 334-1, 334-3, the quality of the mapping and alignment of the demultiplexed groups of the first reads 330-1, 330-3, 332-1, 332-3, 334-1, 334-3, or both, whether the nucleic acid sequencer 310 should continue sequencing operations during the second read interval to generate second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4.
[0091] The determination of whether to continue sequencing operations during the second read interval to generate second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4 can be made automatically by the nucleic acid sequencer 310, manually by a user of the nucleic acid sequencer, automatically by an artificial intelligence agent or model, or based on data describing a determination from a combination thereof, as described with reference to the example of Figure 1B. Alternatively, or in addition, the remote computer 320, a user of the computer 320, or an artificial intelligence agent or model, or a combination thereof, can determine based on the results 359 whether to continue sequencing during the second read interval to generate second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4. Such analysis of results 359 can be evaluated by remote computer 320, a user of remote computer 320, an artificial intelligence agent or model, or a combination thereof, similar to that described with respect to evaluation of results 149 by nucleic acid sequencer 310, a user of nucleic acid sequencer 310, or an artificial intelligence agent or model, or a combination thereof, in the description of Figure 1B. In the case of an artificial intelligence agent or model, the artificial intelligence model can also be trained with input data types that include demultiplexing properties in addition to the other input data types described in the description of Figure 1B.
[0092] In some implementations, demultiplexing statistics can be evaluated separately from or together with mapping and alignment statistics to determine the quality of the reads generated by the nucleic acid sequencer 310. For example, the nucleic acid sequencer 310 or the remote computer 320 can store data representing the expected number of reads for each sample identifier. The nucleic acid sequencer 310, the remote computer 320, a user, an artificial intelligence agent, or a combination thereof can then determine whether the demultiplexing statistics include a number of reads corresponding to each sample identifier that is within a threshold amount of error of the expected number of reads for each sample identifier. If the demultiplexing statistics are within the threshold amount of error of the expected number of reads for each sample identifier, the nucleic acid sequencer 310, the remote computer 320, a human user, an artificial intelligence agent, or a combination thereof can determine whether to continue the sequencing operation. Alternatively, if the demultiplexing statistics are determined to be not within a threshold amount of error in the expected number of reads for each sample identifier, the nucleic acid sequencer 310, the remote computer 320, a user, an artificial intelligence agent or model, or a combination thereof, can decide to terminate the sequencing run.
[0093] In some implementations, results 359 may not need to be sent from the remote computer 320 back to the nucleic acid sequencer 310. Instead, the remote computer 320, a user of the remote computer 320, or an artificial intelligence agent or model can send data back to the nucleic acid sequencer 310 indicating whether the nucleic acid sequencer 310 should continue generating second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4, based on the computer's 320 analysis, the computer's 320 user's analysis, or the artificial intelligence agent's or model's analysis of results 359. The nucleic acid sequencer can then determine whether to continue or terminate the second read interval based on the data received from the remote computer 320, without actually receiving results 359.
[0094] In yet another embodiment, the nucleic acid sequencer can also take multiple decisions into account, similar to those described with reference to FIG. 1B. For example, in some implementations, data describing the decisions of the nucleic acid sequencer 310, the user of the nucleic acid sequencer 310, the remote computer 320, the user of the remote computer 320, an artificial intelligence agent or model, or any combination thereof can be acquired, and in such implementations, the nucleic acid sequencer 310 terminates the second read interval only if the nucleic acid sequencer 310, the user of the nucleic acid sequencer 310, the remote computer 320, the user of the remote computer 320, the artificial intelligence agent or model, or any combination thereof agree that the second read interval should be terminated. In other implementations, a total score can be generated based on a weighted average of one or more decisions of the nucleic acid sequencer 310, the user of the nucleic acid sequencer 310, the remote computer 320, the user of the remote computer 320, the artificial intelligence agent, or any combination thereof, and whether the second read interval should be terminated can be determined based on the total score. In such implementations, the second read interval may be terminated if the total score falls below a predetermined threshold, or may be continued if the total score exceeds a predetermined threshold.
[0095] Using these techniques, the system 300 of FIG. 3 provides similar technical advantages as those described with reference to FIG. 1B. That is, the system 300 can conserve reagents used to generate a second read if the result 359 indicates that the alignment of the first read is a low-quality alignment. If the quality of the demultiplexing statistics, the quality of the mapping and alignment results, the quality of the mapping and alignment statistics, or a combination thereof is determined to be sufficient, the nucleic acid sequencer 310 can discard the result 359. In other implementations, the mapping and alignment of the first read, performed in parallel with the second read, can be used as the mapping and alignment of the first read for the final data run.
[0096] Continuing with the example of Figure 3, after determining that result 359 is satisfactory, nucleic acid sequencer 310 can continue to perform a second read. Once second reads 330-2, 330-4, 332-2, 332-4, 334-2, 334-4 are generated, nucleic acid sequencer 310 can send instructions to remote computer 320 using network 112 instructing secondary analysis unit 340 to begin a final secondary analysis data run. The final data run can include using the secondary analysis unit 340 to demultiplex the second reads 330-2, 330-4, 332-2, 332-4, 334-2, and 334-4 into organized groups of second reads based on the sample identifier of each second read, and then mapping and aligning the second reads 330-2, 330-4, 332-2, 332-4, 334-2, and 334-4. In some implementations, if the mapping and alignment results of the organized set of first reads are discarded, the final data run can perform mapping and alignment operations on both the first and second reads. Because these operations are implemented using the programmable circuitry 342a, these operations can be performed in parallel with the second sequencing run 374 and in a fraction of the time required to perform the second sequencing run 374. This provides an advantage over conventional systems in that a subsequent sequencing run can continue while secondary analysis of the previous sequencing run 372 is being performed, thereby reducing the sequencer downtime that occurs in the conventional system shown in FIG. 1A.
[0097] In addition to demultiplexing and mapping and alignment, the secondary analysis unit 340 can also perform variant calling operations. As an example, the processing unit 350 can be used to perform a variant calling unit 364 that can analyze the mapped and aligned reads to identify the occurrence of any variants, such as single nucleotide polymorphisms (SNPs), insertions / deletions (indels), structural polymorphisms, etc. In some implementations, the programmable circuit 342 can be dynamically reconfigured, for example, by the remote computer 320, to assist in the variant calling process. For example, the programmable circuit 342 can be dynamically reconfigured to include an HMM unit that can be used to perform probability calculations on the likelihood of occurrence of a variant at one or more reference positions in the mapped and aligned reads. Examples of the use of programmable circuitry such as FPGAs to perform variant calling operations are described in further detail in, for example, U.S. Patent Application Publication No. 2016 / 0180019, U.S. Patent Application Publication No. 2016 / 0306922, and U.S. Patent Application Publication No. 2019-0259468, the entire contents of each of which are incorporated herein by reference in their entirety.
[0098] The example in Figure 3 describes an example with a read having eight nucleotides and three samples. However, the present disclosure is not so limited. Instead, this simple example is presented to explain the features of the present disclosure in an easy-to-understand manner. In practice, in some implementations, the DNA fragments of the present disclosure may have, for example, up to 600 nucleotides, up to 800 nucleotides, up to 1,000 nucleotides, or more, and each read of a fragment may have, for example, 50 nucleotides, 75 nucleotides, 150 nucleotides, 200 nucleotides, 300 nucleotides, 500 nucleotides, or more from each end of the DNA fragment. Similarly, Figure 3 or any other diagram should not be construed as limiting the number of clusters of fragments. For example, the nucleic acid sequencer 310 can perform massively parallel sequencing, in which millions of clusters of multiple fragments are sequenced simultaneously.
[0099] While the example in FIG. 3 relates to multiple samples used to generate reads with indexes or sample identifiers, the disclosure is not so limited. Alternatively, system 300 can also be used to process a single sample, generating reads that are not indexed because all reads belong to the same sample. In such an implementation, the same processing can be performed using a second read interval, "Read 2," that begins immediately after the first read interval, "Read 1," without generating any indexes. Then, once the first read interval, "Read 1," is completed, the second read interval, "Read 2," can begin, with secondary analysis of the first read paralleled with the second read interval. The only substantial difference between single and multiple sample implementations is that the index generation and demultiplexing steps do not need to be performed on a single sample, because all reads are associated with the same sample.
[0100] Figure 4 is a flowchart of an example process 400 for performing incremental secondary analysis according to the workflow diagram of Figure 3. Generally, process 400 includes acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device from a plurality of different samples during a first read interval (410), acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device from the plurality of different samples during a second read interval performed after the first read interval (410), and, while acquiring the second data in step 420, (I) organizing the plurality of first reads into organized groups based on at least a first sample identifier or a second sample identifier associated with each of the first reads. (II) for each organized group of first reads, performing a secondary analysis operation on the organized group of first reads, and (III) storing results of the secondary analysis for each group of first reads (430), and then instructing a secondary analysis unit to initiate: (A) organizing the plurality of second reads into a plurality of organized groups based on at least the first sample identifier or the second sample identifier (440), and (B) for each organized group of second reads, performing a secondary analysis operation on the organized group of second reads or the organized group of first reads and second reads (450). For convenience, but not by way of limitation, these steps are described in more detail below as performed by a sequencing system such as system 300 of FIG. 3.
[0101] The sequencing system can begin executing process 400 by acquiring 410 first data describing a plurality of first reads generated by the nucleic acid sequencing device from a plurality of different samples during a first read interval. Acquiring the first data can include storing the first data representing the plurality of first reads in a memory device, such as a memory device of a secondary analysis unit, after the first data are generated by the sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform the secondary analysis operation. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof.
[0102] Each read of the plurality of first reads can consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides can correspond to nucleotides at a first end of a nucleic acid fragment. The nucleic acid fragment can be clonally amplified to facilitate sequencing, and in such implementations, the ordered sequence of nucleotides can be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read. Each first read can include data identifying a sample used to generate the first read. In some implementations, the data identifying the sample can include a barcode. The nucleic acid sequencing device can include any nucleic acid sequencing device, including a DNA sequencer or an RNA sequencer.
[0103] The sequencing system can continue executing process 400 by acquiring 420 second data describing a plurality of second reads generated by the nucleic acid sequencing device from a plurality of different samples during a second read interval performed after the first read interval. Acquiring the second data can include storing the second data representing the plurality of first reads in a memory of the secondary analysis unit after the second data is generated by the sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform the secondary analysis operation. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof.
[0104] In some implementations, at least a portion of the second data is acquired while another portion of the second data is being generated by the nucleic acid sequencing device. Each read of the plurality of second reads can consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides can correspond to nucleotides at a second end of the nucleic acid fragment opposite the first end of the nucleic acid fragment. The nucleic acid fragments can be clonally amplified to facilitate sequencing; in such implementations, the ordered sequence of nucleotides can be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read. Each second read can include data identifying the sample from which the second read was generated. In some implementations, the data identifying the sample can include a barcode.
[0105] While the second data has been acquired in step 420, the sequencing system can parallelize additional processing of the plurality of first reads using a secondary analysis unit. In some implementations, the additional parallel processing can include (I) organizing data representing the plurality of first reads into organized groups based on at least a first sample identifier or a second sample identifier associated with each of the first reads, (II) for each organized group of first reads, performing a secondary analysis operation on the organized group of first reads, and (III) storing (430) secondary analysis results for each group of first reads.
[0106] Organizing multiple first reads into organized groups based on sample identifiers is necessary when multiple samples are being sequenced. This may include performing one or more demultiplexing operations to map sets of first reads with different first sample identifiers into respective organized groups, each of which has the same sample identifier. Demultiplexing statistics describing the quality of the demultiplexing operation can be generated. For example, the demultiplexing statistics can indicate the number of first reads corresponding to each sample identifier. In some implementations, the secondary analysis unit can return the result data to a nucleic acid sequencer, provide the result data to one or more artificial intelligence agents or models, or output the result data to one or more human users who describe the demultiplexing statistics. In such examples, the sequencing system can determine whether to continue process 400 or terminate process 400 at this point based on the quality of the demultiplexing operation described by the demultiplexing statistics. Alternatively, such demultiplexing statistics can be returned as result data after mapping and alignment operations are performed, as described below.
[0107] Once the plurality of first reads have been organized, the sequencing system can perform, for each organized group of first reads, one or more secondary analysis operations on the organized group of first reads. Performing the secondary analysis operations on the organized group of first reads can include, for each organized group of first reads, (I) providing, by the nucleic acid sequencing device, the organized group of first reads to a mapping and alignment unit to align the organized group of first reads to a reference sequence, (II) aligning the organized group of first reads to the reference sequence using the mapping and alignment unit, (iii) receiving results from the mapping and alignment unit, and (iv) storing the received result data before completing acquisition of the second data in step 420.
[0108] The result data can include demultiplexing statistics or mapping and alignment statistics. The demultiplexing statistics can include data describing the quality of the demultiplexing operation, such as the number of first reads corresponding to each sample identifier. The mapping and alignment statistics can include data describing the quality of the alignment of each organized group of first reads to each reference sequence. The mapping and alignment statistics can include, for example, one or more of a MAPQ score, an alignment score, etc. In other implementations, the mapping and alignment results can include the mapped and aligned reads of each organized group of first reads, which can be provided as input to a variant caller to determine potential variants between the mapped and aligned reads of each organized group of first reads and the respective reference sequence.
[0109] In some implementations, the output data describing the result data for each organized group of first reads can be provided for review by one or more human users. For example, the output data describing the result data for each organized group of first reads can be output, for example, on a display coupled to the nucleic acid sequencing device or provided in a separate room or building. Alternatively, or in addition, the output data describing the result data for each organized group of first reads can be output, for example, using a printer communicatively connected directly or indirectly to the nucleic acid sequencing device to print a report describing the alignment results for each organized group of first reads.
[0110] In some implementations, the sequencing system, a remote computer, one or more human users, an artificial intelligence agent or model, or a combination thereof can evaluate the result data while the second data is being acquired in stage 420. For example, the result data can be evaluated to determine whether the demultiplexed first reads, the mapping and alignment of the first reads, or both, are of sufficient quality to continue acquiring second data in stage 420. In some implementations, if the result data of the organized group of first reads does not satisfy one or more predetermined rules or thresholds, the nucleic acid sequencer can be instructed to stop acquiring second data in stage 420. Alternatively, if it is determined that the result data of the organized group of first reads does satisfy one or more predetermined rules or thresholds, the nucleic acid sequencer can be allowed to continue acquiring second data in stage 420.
[0111] In some implementations, each organized group of mapped and aligned first reads can be evaluated for potential variant detection while second data is being acquired in step 420. Such implementations allow tertiary analysis of identified variants for each group to be achieved more quickly than conventional methods that prohibit initiation of tertiary analysis until after both the first and second read intervals are completed. Thus, an initial diagnosis can be obtained for initiating treatment 12 to 24 hours earlier than conventional methods, since there is no need to wait to complete the second read interval before proceeding to tertiary analysis.
[0112] The sequencing system can continue executing process 400 by instructing the mapping and alignment unit to begin organizing the plurality of second reads into a plurality of organized groups of second reads based on at least the first sample identifier or the second sample identifier at step 430. Organizing the plurality of second reads into organized groups based on the second sample identifier is necessary to obtain associated secondary analysis processing of the second reads. This may include performing one or more demultiplexing operations to map sets of second reads with different sample identifiers into different organized groups, each organized group of second reads having the same second sample identifier. The sequencing system can continue executing process 400 by performing, for each organized group of second reads, a secondary analysis operation on the organized group of second reads (step 440). In some implementations, the secondary analysis operation can be performed on a combination of the first reads and the second reads.
[0113] In some implementations, the sequencing system can proceed to step 430 and step 440. Such implementations further provide the technical advantage of facilitating tertiary analysis and reducing downtime of the nucleic acid sequencer. However, in other implementations, execution of process 400 by the sequencing system can continue to organize (430) the plurality of second reads into a plurality of organized groups and perform secondary analysis operations, such as mapping and alignment, variant calling, or both, only if the received result data describing the demultiplexing quality of the first reads, the mapping and alignment quality of the first reads, or both, for each organized group of first reads is determined to satisfy one or more predetermined quality rules or thresholds.
[0114] In some implementations, the sequencing system can rely on the results of the secondary analysis of the first organized group of reads performed in stage 420, mapping and alignment, variant calling, or both, while the second data is being acquired. In other implementations, these initial secondary analysis results associated with the first organized group of reads performed in stage 420 can be discarded after they are evaluated to determine the quality of the first read interval. In such examples, the sequencing system can begin the second iteration of the secondary analysis of the first organized group of reads either before or after the secondary analysis of the second organized group of reads in stages 430 and 440 is completed.
[0115] FIG. 5 is a context diagram of an example system 500 for performing incremental secondary analysis of one or more samples using a secondary analysis unit within a nucleic acid sequencer. System 500 is generally the same as system 300 described with reference to FIG. 3, with some differences. One difference is that secondary analysis unit 540 is located within nucleic acid sequencer 510. For any reference numbers in FIG. 5 not explicitly stated, the components identified by the reference numbers have the same features as the corresponding features in FIG. 1 or FIG. 3. By way of example, each of clusters 522-1, 522-2, 522-3, 522-4, 522-5, and 522-N has the same meaning as clusters 122-1, 122-2, 122-3, 122-4, 122-5, and 122-N, respectively, in FIG. 1, unless additional or different features are described with reference to FIG. 5.
[0116] Another difference between the example of FIG. 5 and the example of FIG. 3 is that the nucleic acid sequencer is configured to generate a sample identifier or index for each read before the first read interval. This is shown in workflow 570, which shows that IND1 and IND2 are generated following the clustering stage and before the first read, "Read 1," of the first read interval of workflow 570. This differs from the generation of sample identifiers or indexes in the example of FIG. 3 because the index in FIG. 3 is generated after the first read interval. While the implementations of FIGS. 5 and 6 are described as generating separate sample identifiers or indexes for "Read 1" and "Read 2," the disclosure is not so limited. Instead, implementations of the present disclosure may generate only a single sample identifier or index identifier that refers to both "Read 1" and "Read 2" of a particular fragment.
[0117] An advantage of generating sample identifiers before the first read interval is that as reads are generated, organization of reads into demultiplexed groups with the same sample identifier can be performed at run time. Given the generation of all sample identifiers and the ability to organize reads based on sample identifiers at run time, system 500 can begin secondary analysis of the organized groups of first reads during the first read interval. In such a scenario, secondary analysis result data, including demultiplexing statistics, mapping and alignment statistics, or both, for each organized group of first reads can be obtained and evaluated during the first read interval, thereby enabling the option to terminate the first read interval if the result data does not indicate satisfactory results, thereby conserving reagents.
[0118] Furthermore, the ability to begin performing secondary analysis of organized groups of first reads during the first read interval allows for a faster transition to tertiary analysis operations than the example systems described with reference to Figures 1B and 3. The system of Figure 5 can transition to tertiary analysis more quickly than the systems of Figures 1B and 3 because the initial set of variants based on the mapped and aligned first reads and used as input for tertiary analysis can be identified during the first read interval. This allows for the initiation of tertiary analysis within approximately a few hours from the start of the first read interval. This contrasts with the examples of Figures 1B and 3, which may not begin tertiary analysis using identified variants of the mapped and aligned reads, respectively, as input until after sequencing is complete.
[0119] 5, the programmable circuit 542 of the secondary analysis unit 540 of the nucleic acid sequencer 510 can be configured to include a mapping and alignment unit 542a. The nucleic acid sequencer 510 can receive multiple samples 105, 106, and 107. The samples 105, 106, and 107 can include, for example, nucleic acid samples from different species. The different species can be different humans, different animals, different plants, etc. The nucleic acid sequencer 510 can prepare the samples 105, 106, and 107 and perform cluster generation during time T1 of the workflow 570.
[0120] At the end of the cluster phase, the nucleic acid sequencer 510 begins generating an index, or sample identifier, for each first read generated by the nucleic acid sequencer 510 during time T2A. At the end of time T2A, the nucleic acid sequencer 510 begins generating an index or sample identifier for each second read generated by the nucleic acid sequencer 510 during time T2B. The index or sample identifier for each read can include any data that can be used to create a logical relationship between the read and the sample. Thus, at the end of time T1+T2A+T2B in the example of FIG. 5, an index or sample identifier has been generated for each first read generated by the nucleic acid sequencer 510 during the first read interval, or an index or sample identifier has also been generated for each second read generated by the nucleic acid sequencer 510 during the second read interval.
[0121] The nucleic acid sequencer 510 is configured to parallelize secondary analysis operations, such as mapping and alignment, of at least a portion of the first reads 530-1, 530-3, 532-1, 532-3, 534-1, 534-3, while the nucleic acid sequencer 510 continues to perform sequencing operations, such as sequencing-by-synthesis, of the first read interval during time T3. Initiating secondary analysis of at least a portion of the first reads during the first read interval cannot be accomplished in the example of FIG. 3 because the index or sample identifier for each read was not generated until after the first read interval was completed. In contrast, in the example of FIG. 5, the index or sample identifier index for each read to be generated by the nucleic acid sequencer 510 is created in advance.
[0122] 5, the first read interval does not begin until after completion of time T1+T2A+T2B of workflow 570. After completion of T1+T2A+T2B, nucleic acid sequencer 570 can begin the first read interval. Initiating the first read interval can include initiating a primary analysis sequencing operation, such as sequencing by synthesis, to generate one or more first reads 530-1, 530-3, 532-1, 532-3, 534-1, 534-3. After time TX from the start of first read interval "Read 1," one or more first reads 530-1, 530-3, 532-1 generated during time TX can then be stored in memory 544 of secondary analysis unit 540 or other memory accessible by secondary analysis unit 540, processing unit 150, or both.
[0123] Because the nucleic acid sequencer 510 is sequencing multiple samples, the nucleic acid sequencer 510 needs to perform an organization operation to organize one or more first reads 530-1, 530-3, 532-1 into one or more organized groups of first reads. Organizing the first reads can be achieved using a demultiplexing unit 562. For example, the processing unit 550 can access one or more reads stored in memory 544, memory 560, or other memory and execute the programmed functions of the demultiplexing unit 562 to demultiplex the one or more first reads 530-1, 530-3, 532-1 into one or more organized groups of first reads. Demultiplexing can be achieved using one or more demultiplexing operations to organize the one or more first reads 530-1, 530-3, 532-1 based on an index or sample identifier for each first read. The demultiplexed first reads may be stored in memory 544 or other memory accessible to the mapping and alignment unit 542a.
[0124] The mapping and alignment unit 542a can access the organized first reads stored in the memory 544 and perform real-time mapping and alignment operations on the demultiplexed first reads during the first read interval. The secondary analysis unit 540 can generate results 549 for each group of first reads stored in the memory 544. The results 549 can include demultiplexing statistics, mapping and alignment statistics, mapping and alignment results, or a combination thereof. The secondary analysis unit 540 can store the received results in the memory 560. The demultiplexing statistics can include data describing the demultiplexing quality, such as the number of records corresponding to each sample identifier. For example, mapping and alignment statistics such as a MAPQ score that provides an indication of the mapping quality of each group of first reads and an alignment score that provides an indication of the alignment quality of each group of first reads. The mapping and alignment results 549 can include data describing the mapped and aligned reads. In some implementations, these mapping and alignment results can be dynamically updated as more first reads are generated and mapped and aligned to their respective reference sequences.
[0125] 5, the ultra-fast execution time of the mapping and alignment unit 542a implemented using hardwired logic in the programmable circuit 542 allows the mapping and alignment unit 542a to perform mapping and alignment of each demultiplexed group of first reads 530-1, 530-3, 532-1, 532-3, 534-1, 534-3 in a fraction of the time required by the nucleic acid sequencer 510 to perform the first read interval. For example, in some implementations, the programmable circuit 542a can perform mapping and alignment of the demultiplexed group of first reads generated during time Tx in hardwired logic during the first read interval "Read 1" in a few minutes or less, whereas performing the entire first read interval using software executed by the processing unit 150 may take 12 to 24 hours. Thus, nucleic acid sequencer 510 or one or more human users can evaluate results 549 of the secondary analysis of a first read, such as a first read generated during time TX, while the remainder of the first reads are generated by nucleic acid sequencer 510 during time T3. Nucleic acid sequencer 510, a user of nucleic acid sequencer 510, an artificial intelligence agent or model, or a combination thereof can then make a decision based on the quality of the demultiplexing operation, the quality of the mapping and alignment operation, or both, using results 549 to determine whether nucleic acid sequencer 510 should continue performing sequencing operations during the first read interval. This decision regarding whether to continue sequencing operations during the first read interval can be made automatically by nucleic acid sequencer 510, automatically by an artificial intelligence agent or model, by a user of the nucleic acid sequencer, or based on data describing the decision from each of these entities as described with reference to the example of FIG. 1B.
[0126] Using these techniques, the system 500 of FIG. 5 provides the technical advantages described with reference to FIG. 1B. That is, if the result 549 indicates that the demultiplexing of at least some of the first reads already generated during the first read interval, the alignment of some of the first reads already generated during the first read interval, or both, are of low quality, the system 500 can conserve reagents used to continue generating additional reads during the first read interval. Upon determining that the demultiplexing quality of the already generated first reads, the mapping and alignment quality of the already generated first reads, or both, are satisfactory, the nucleic acid sequencer 510 can discard the mapping and alignment result 549. In other implementations, the mapping and alignment of the already generated first reads, performed in parallel with the first read interval, can be used as the mapping and alignment of the final data run of the first reads.
[0127] In addition to demultiplexing and mapping and alignment, the secondary analysis unit 540 can also perform variant calling operations on one or more groups of the first reads mapped and aligned during the first read interval, "Read 1." As an example, the processing unit 550 can be used to execute a variant calling unit 564, which can analyze the mapped and aligned reads to identify the occurrence of any variants, such as single nucleotide polymorphisms (SNPs), insertions / deletions (indels), structural polymorphisms, etc. In some implementations, the programmable circuit 542 can be dynamically reconfigured, for example, by the nucleic acid sequencer 510, to assist in the variant calling process. For example, the programmable circuit 542 can be dynamically reconfigured to include an HMM unit, which can be used to perform probability calculations on the likelihood of a variant occurring at one or more reference positions of the mapped and aligned reads. Subsequently, the nucleic acid sequencer 510, or other computing device, can perform one or more tertiary analysis operations during the first read interval, "Read 1," using any identified variants. This can be useful in facilitating treatment to entities based on tertiary analysis. Entities can include patients, humans, subjects, plants, animals, etc.
[0128] In the example system 500, if the decision to terminate the first read interval is made based on a determination that the demultiplexing statistics, the mapping and alignment statistics, or both, are of poor quality, the system 500 may also terminate the second read interval, "Read 2." Thus, the system 500 provides an additional advantage over the example systems of Figure 1B or 3 in that it may conserve even more reagents when poor demultiplexing results, mapping and alignment results, or both, are detected.
[0129] However, with reference to the example system 500, if the demultiplexing results, the mapping and alignment results, or both are determined to meet a threshold level of quality, the system 500 can begin executing a second read interval, "Read 2," as shown in workflow 570. In some implementations, the system 500 can generate the second read interval, "Read 2," without parallelizing the secondary analysis of the second read. For example, such execution may be preferable because the system 500 has already assessed sequencing quality during the first read interval, "Read 1." However, in other implementations, the system 500 can parallelize the secondary analysis of the second read in the same way that the secondary analysis of the first read was parallelized with the first read interval.
[0130] The example in Figure 5 describes an example with a read having 8 nucleotides and 3 samples. However, the present disclosure is not so limited. Instead, this simple example is presented to explain the features of the present disclosure in an easy-to-understand manner. In practice, in some implementations, the DNA fragments of the present disclosure may have, for example, up to 600 nucleotides, up to 800 nucleotides, up to 1,000 nucleotides, or more, and each read of a fragment may have, for example, 50 nucleotides, 75 nucleotides, 150 nucleotides, 200 nucleotides, 300 nucleotides, 500 nucleotides, or more from each end of the DNA fragment. Similarly, Figure 5 or any other diagram should not be construed as limiting the number of clusters of fragments. For example, the nucleic acid sequencer 510 can perform massively parallel sequencing, in which millions of clusters of multiple fragments are sequenced simultaneously.
[0131] While the example in FIG. 5 relates to multiple samples used to generate reads with indexes or sample identifiers, the disclosure is not so limited. Alternatively, system 500 may also be used to process a single sample, generating reads that are not indexed because all reads belong to the same sample. In such an implementation, the same process may be performed, with a first read interval beginning immediately after the clustering stage. Once a portion of the first reads are generated during the first read interval "Read 1," system 500 may provide the generated portion of the first reads to mapping and alignment unit 542a for mapping and alignment, while the remaining portion of the first reads are generated during the first read interval without having to perform a demultiplexing stage. In this implementation, the first reads do not need to be demultiplexed because they are all associated with the same sample. Similarly, the mapped and aligned portion of the first reads may then be analyzed for variants using first read interval "Read 1," as described above. A similar determination can be made regarding whether to continue the first read interval and the second read interval, as described with respect to exemplary Figure 5. In essence, the substantial difference between the single sample implementation of system 500 of Figure 5 and the multiple sample implementation of Figure 5 is that the single sample implementation does not require the demultiplexing step to be performed.
[0132] Figure 6 is a flowchart of an example process 600 for performing incremental secondary analysis according to the workflow diagram of Figure 5. Generally, process 600 includes generating (610) a plurality of first sample identifiers, each first sample identifier corresponding to a particular read generated during a first read interval; generating (620) a plurality of second sample identifiers, each second sample identifier corresponding to a particular read generated during the second read interval; and acquiring (630) first data describing the plurality of first reads generated by the nucleic acid sequencing device from a plurality of different samples during the first read interval, each of the plurality of first reads corresponding to at least one of the first sample identifier or the second sample identifier. During the first data acquisition in step 630, (I) classifying the plurality of first reads into a sequence of first reads, each of the first reads being a sequence of first reads, and (II) classifying the plurality of first reads into a sequence of second sample identifiers. (II) for each organized group of first reads, performing a secondary analysis operation on the organized group of first reads, and (III) storing the secondary analysis results for each group of first reads (640); acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device from a plurality of different samples during a second read interval performed after the first read interval, wherein each of the plurality of second reads corresponds to at least one of the first sample identifiers or the second sample identifiers (650); and performing secondary analysis on the acquired second data (660). For convenience, but not by way of limitation, these steps are described in more detail below as performed by a sequencing system such as system 500 of FIG. 5.
[0133] The sequencing system can begin execution of process 600 by generating 610 a plurality of first sample identifiers, each corresponding to a particular read generated during the first read interval. In some implementations, each first sample identifier can include an index tag sequence. The index tag sequence can be attached to the target polynucleotide of each sample before the sample is immobilized for sequencing. The index tag can be a synthetic sequence of nucleotides added to the target as part of the template preparation process. Thus, a library-specific index tag is a nucleic acid sequence tag attached to each target molecule of a sample, the presence of which indicates or identifies the entity from which the target molecule was isolated. In some implementations, the index tag sequence can include a barcode embedded in the synthetic sequence.
[0134] The sequencing system can continue executing process 600 by generating a plurality of second sample identifiers at step 620, each second sample identifier corresponding to a particular read generated during a second read interval that occurs after the first read interval. In some implementations, each second sample identifier can include an index tag sequence. The index tag sequence can be attached to the target polynucleotide of each sample before the respective sample is immobilized for sequencing. The index tag can be a synthetic sequence of nucleotides added to the target as part of the template preparation process. Thus, the library-specific index tag is a nucleic acid sequence tag attached to each target molecule of the sample, the presence of which indicates or identifies the entity from which the target molecule was isolated. In some implementations, the index tag sequence can include a barcode embedded in the synthetic sequence.
[0135] The sequencing system can continue executing process 600 by acquiring, at step 630, first data describing a plurality of first reads generated by the nucleic acid sequencing device from a plurality of different samples during the first read interval, where each of the plurality of first reads corresponds to one of the first sample identifiers. Acquiring the first data can include storing the first data representing one or more first reads in a memory of the secondary analysis unit after the first data are generated by the sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform secondary analysis operations. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof. In some implementations, at least a portion of the first data is acquired while another portion of the first data is generated by the nucleic acid sequencing device. That is, data representing a first set of one or more reads can be acquired and stored in the memory of the secondary analysis unit while one or more other first reads are generated by the nucleic acid sequencing device during the first read interval.
[0136] Each read of the plurality of first reads may consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides may correspond to nucleotides at a first end of a nucleic acid fragment. The nucleic acid fragment may be clonally amplified to facilitate sequencing, and in such implementations, the ordered sequence of nucleotides may be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read. Each first sample identifier of each first read generated before the first read interval corresponds to a specific sample from which the first read occurred. The first sample identifier can be used by the sequencing system to determine the sample associated with any particular first read. In some implementations, the sample-identifying data may include a barcode.
[0137] While acquiring first data in step 630 during the first read interval, the sequencing system can use a secondary analysis unit to parallelize in real time additional processing of one or more of the first reads already generated by the nucleic acid sequencer. In some implementations, the additional processing can include (I) organizing the multiple first reads into organized groups based on at least a first sample identifier or a second sample identifier associated with each of the first reads, (II) for each organized group of first reads, performing a secondary analysis operation on the organized group of first reads, and (III) storing secondary analysis results for each group of first reads (step 640).
[0138] Organizing one or more first reads into organized groups based on sample identifiers is necessary to obtain associated secondary analysis processing when multiple samples are sequenced. This may include performing one or more demultiplexing operations to map one or more first reads with different first sample identifiers to respective organized groups, each of which has the same sample identifier. Demultiplexing statistics describing the quality of the demultiplexing operation can be generated. For example, the demultiplexing statistics can indicate the number of first reads corresponding to each sample identifier. In some implementations, the secondary analysis unit can return the result data to the nucleic acid sequencer, provide the result data to one or more artificial intelligence agents or models, or output the result data to one or more human users who describe the demultiplexing statistics. In such examples, the sequencing system can determine whether to continue process 600 or terminate process 600 at this point based on the quality of the demultiplexing operation described by the demultiplexing statistics. Alternatively, such demultiplexing statistics can be returned as result data after mapping and alignment operations are performed, as described below.
[0139] Once the one or more first reads have been organized, the sequencing system can perform, for each organized group of first reads, one or more secondary analysis operations on the organized group of first reads using a secondary analysis unit in parallel with the remainder of the first read interval. Performing the secondary analysis operations on the organized group of first reads can include, for each organized group of first reads, (I) providing the organized group of first reads to a mapping and alignment unit by the nucleic acid sequencing device to align the organized group of first reads to a reference sequence, (II) aligning the organized group of first reads to the reference sequence using the mapping and alignment unit, (III) receiving result data from the mapping and alignment unit, and (IV) storing the received alignment result data before completing acquisition of the first data in step 630.
[0140] The result data can include demultiplexing statistics or mapping and alignment statistics. The demultiplexing statistics can include data describing the quality of the demultiplexing operation, such as the number of first reads corresponding to each sample identifier. The mapping and alignment statistics can include data describing the quality of the alignment of each organized group of first reads to each reference sequence. The mapping and alignment statistics can include, for example, one or more of a MAPQ score, an alignment score, etc. In other implementations, the mapping and alignment results can include the mapped and aligned reads of each organized group of first reads, which can be provided as input to a variant caller to determine potential variants between the mapped and aligned reads of each organized group of first reads and the respective reference sequence.
[0141] In some implementations, the output data describing the result data for each organized group of first reads can be provided for review by one or more human users. For example, the output data describing the result data for each organized group of first reads can be output, for example, on a display coupled to the nucleic acid sequencing device or provided in a separate room or building. Alternatively, or in addition, the output data describing the alignment results for each organized group of first reads can be output, for example, using a printer communicatively connected directly or indirectly to the nucleic acid sequencing device to print a report describing the alignment results for each organized group of first reads.
[0142] In some implementations, the sequencing system, one or more human users, one or more artificial intelligence agents or models, or a combination thereof, can evaluate the alignment results while the first data is being acquired in step 630. For example, the resulting data can be evaluated to determine whether the demultiplexing of the acquired first reads, the mapping and alignment of the acquired first reads, or a combination of both, is of sufficient quality to continue acquiring the first data in step 630. In some implementations, if the resulting data of the organized group of first reads does not satisfy one or more predetermined rules or thresholds, the nucleic acid sequencer can be instructed to stop acquiring the first data during the first read interval in step 630. Alternatively, if it is determined that the resulting data of the organized group of first reads satisfies one or more predetermined rules or thresholds, the nucleic acid sequencer can be permitted to continue acquiring the first data during the first read interval in step 630.
[0143] In some implementations, each organized group of mapped and aligned first reads can be evaluated for potential variant detection while the first data is being acquired in stage 630. Such implementations allow tertiary analysis of identified variants for each group to be achieved more quickly than conventional methods that prohibit initiation of tertiary analysis until after completion of both the first read interval in stage 630 and the second read interval in stage 650. Thus, an initial diagnosis can be obtained to begin treatment days earlier than the conventional method shown in FIG. 1A by not having to wait for completion of the first read interval, the second read interval, and mapping and alignment of the first and second reads before proceeding to tertiary analysis.
[0144] At the end of stage 630, the sequencing system can continue executing process 600 by acquiring 650 second data describing a plurality of second reads generated by the nucleic acid sequencing device from a plurality of different samples during a second read interval performed after the first read interval, each of the plurality of second reads corresponding to at least one of the first sample identifier or the second sample identifier. Acquiring the second data can include storing the second data representing the one or more second reads generated during the second read interval in a memory of a device of the secondary analysis unit after the second data is generated by the sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform secondary analysis operations. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof. In some implementations, at least a portion of the second data is acquired while another portion of the second data is being generated by the nucleic acid sequencing device. That is, data representing a second set of one or more reads can be obtained and stored in the memory of the sequencing device, while one or more other second reads are generated by the nucleic acid sequencing device during a second read interval.
[0145] Each read of the plurality of second reads can consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides can correspond to nucleotides at a second end of the nucleic acid fragment opposite the first end of the nucleic acid fragment. The nucleic acid fragment can be clonally amplified to facilitate sequencing, and in such implementations, the ordered sequence of nucleotides can be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read. Each second sample identifier of each second read generated before the second read interval corresponds to a specific identifier of the second read. The second sample identifier can be used by the sequencing system to determine the sample associated with any particular second read. In some implementations, the sample-identifying data can include a barcode.
[0146] The sequencing system can continue executing process 600 by performing 660 a secondary analysis of the obtained second data. In some implementations, the sequencing system can proceed to step 660 after completing step 650. In the context of process 600, this can occur while still achieving at least some of the advantages of the present disclosure, such as rapid tertiary analysis and reduced nucleic acid sequencer downtime due to the ability to assess sequencing quality during the first read interval of step 640. However, the present disclosure is not so limited. Instead, in some implementations, the sequencing system can parallelize the secondary analysis of the second read in the same way that the secondary analysis of the first read was parallelized with the first read interval.
[0147] In some implementations, the sequencing system may rely on the results of the secondary analysis of the first organized group of reads, mapping and alignment, variant calling, or both, performed in stage 640 while the first data is being acquired during the first read interval. In other implementations, these initial secondary analysis results associated with the first organized group of reads performed in stage 640 may be discarded after they are evaluated to determine the quality of the first read interval. In such examples, the sequencing system may begin the second iteration of the secondary analysis of the first organized group of reads either before or after the secondary analysis of the second organized group of reads is completed in stage 660.
[0148] Figure 7 is an example workflow diagram 770 illustrating a workflow of operations performed during a process for performing incremental secondary analysis using a secondary analysis unit. Workflow diagram 770 is the same as workflow diagram 370 shown in Figure 3. However, in Figure 7, an additional sequence of operations 710 performed during the final data run is shown superimposed on workflow diagram 770.
[0149] In some implementations, the final data run can include secondary analysis or other additional processing that results in a secondary analysis result with a threshold level of confidence. In conventional sequencing systems, the final data run cannot be achieved by the conventional sequencing system until both the first read interval and the second read interval are completed. Furthermore, such conventional systems also have sequencer downtime between the end of the first sequencing run and the start of the second sequencing run, as shown in FIG. 1A. Although exemplary implementations that use a threshold level of confidence are described, other implementations that do not utilize such a threshold can also be used.
[0150] In the example of FIG. 7, a sequencing system, such as the sequencing system of FIG. 3 or FIG. 5, can be configured to initiate a final data run at time T Y before the end of the second read interval. Time T Y can be, for example, a predetermined number of one or more sequencing cycles from the end of the second read interval, where a cycle refers to the time required to generate a single nucleic acid from a read. In some implementations, the nucleic acid sequencer can be configured to detect when a predetermined number of sequencing cycles has elapsed since the end of the second read interval, "Read 2," and initiate a secondary analysis run on one or more first reads generated during the first read interval, "Read 1." The first reads can include one or more organized sets of reads previously demultiplexed at the end of time T3B in the workflow of FIG. 7. Initiating a secondary analysis run can include, for example, instructing a secondary analysis unit to perform mapping and alignment, variant calling of the mapped and aligned reads, or both.
[0151] Once initiated, the secondary analysis unit can continue to perform secondary analysis operations on reads generated during the first read interval and the second read interval of the first sequencing run until the triggered secondary analysis operation is completed. As shown in FIG. 7 , execution of secondary analysis operations using the secondary analysis unit can begin during the first sequencing run and continue during the second sequencing run, which begins after the first sequencing run is completed. Secondary analysis operations on reads generated during the first sequencing run are completed during the second sequencing run. This parallelization of secondary analysis corresponding to the first sequencing run with the second sequencing run operation thus allows the nucleic acid sequencer to continue the sequencing run with little or no sequencer downtime, thereby increasing reagent consumption and resulting revenue. Second sequencing run operations that overlap with the secondary analysis of the first sequencing run can include, but are not limited to, second sequencing run setup, clustering, or primary analysis.
[0152] In the example of Figure 7, the parallelization of the operations of the secondary analysis of the first sequencing run and the second sequencing run is not performed to assess the quality of the reads generated by the nucleic acid sequencer and determine whether the second read interval should continue. Instead, the parallelization of the operations of the secondary analysis and the second sequencing run is performed as part of a final data run to create final result data suitable for use in subsequent operations, such as during tertiary analysis.
[0153] Figure 8 is a flowchart of an example process 800 for performing incremental secondary analysis according to the workflow diagram of Figure 7. Generally, the process includes acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval of a first sequencing run (810), acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval of the first sequencing run that is performed after the first read interval (820), beginning to perform one or more secondary analysis operations on at least the first data or the second data while acquiring at least a portion of the second data in step 820 (830), performing a second sequencing run using the nucleic acid sequencing device (840), and while performing the second sequencing run using the nucleic acid sequencing device in step 840, (I) continuing to perform one or more secondary analysis operations on the first data or the second data, and (II) storing result data representing results of the secondary analysis operations (850). For convenience, but not by way of limitation, these steps are described in more detail below as performed by a sequencing system such as system 100 of FIG. 1A, system 300 of FIG. 3, or system 500 of FIG. 5, respectively.
[0154] The sequencing system can begin performing process 800 at step 810 by acquiring first data describing a plurality of first reads generated by the nucleic acid sequencing device during a first read interval of a first sequencing run. Acquiring the first data can include storing the first data describing the plurality of first reads in a memory device, such as a memory device of a secondary analysis unit, after the first data are generated by the nucleic acid sequencing device. The memory device of the secondary analysis unit can be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform secondary analysis operations. The integrated circuit can include one or more programmable circuits, one or more ASICs, or a combination thereof.
[0155] Each read of the plurality of first reads may consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides may correspond to nucleotides at a first end of a nucleic acid fragment. The nucleic acid fragment may be clonally amplified to facilitate sequencing; in such implementations, the ordered sequence of nucleotides may be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read. The nucleic acid sequencing device may include any nucleic acid sequencing device, including a DNA sequencer or an RNA sequencer. The first sequencing run may include a complete execution of a primary analysis of one or more biological samples by the nucleic acid sequencing device. An example of the stages of a complete first sequencing run is shown in FIG. 7, which includes a clustering stage, a first read interval, and a second read interval. In some implementations, such as that shown in FIG. 7, the primary analysis may also include one or more indexing stages.
[0156] The sequencing system may continue executing process 800 by acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval of the first sequencing run, which is performed after the first read interval, at step 820. Acquiring the second data may include storing the second data representing the plurality of second reads in a memory of the secondary analysis unit after the second data is generated by the sequencing device. The memory device of the secondary analysis unit may be a memory unit accessible by an integrated circuit of the secondary analysis unit configured to perform the secondary analysis operation. The integrated circuit may include one or more programmable circuits, one or more ASICs, or a combination thereof. In some implementations, at least a portion of the second data is acquired while another portion of the second data is being generated by the nucleic acid sequencing device. Each read of the plurality of second reads may consist of an ordered sequence of nucleotides. In some implementations, the ordered sequence of nucleotides may correspond to nucleotides at a second end of the nucleic acid fragment opposite the first end of the nucleic acid fragment. The nucleic acid fragment may be clonally amplified to facilitate sequencing, and in such implementations, the ordered sequence of nucleotides may be determined by analyzing multiple clones of the nucleic acid fragment to generate the nucleotides of the read.
[0157] While acquiring at least a portion of the second data in step 820, the sequencing system may continue execution of process 800 by initiating execution of one or more secondary analysis operations on the first data or the second data in step 830. Initiating execution of the one or more secondary analysis operations may include dynamically configuring a programmable circuit including hardwired logic to perform the secondary analysis operations and then performing at least one secondary analysis operation on one or more reads generated during the first sequencing run. For example, the sequencing system may dynamically configure the programmable circuit as a mapping and alignment unit and then use hardwired logic of the mapping and alignment unit to perform mapping and alignment of at least one read generated during the first sequencing run. In other implementations, initiating execution of the one or more secondary analysis operations may include instructing an ASIC to execute hardwired digital logic to perform a secondary analysis operation on one or more reads generated during the first sequencing run.
[0158] In some implementations, such as when multiple samples are sequenced during the first sequencing run, the first reads or second reads may need to be organized into demultiplexed groups before mapping and alignment. In such implementations, at least a portion of the organization of the first reads, the second reads, or both, may also be performed during stage 820.
[0159] The sequencing system can continue performing process 800 at step 840 by using the nucleic acid sequencing device to perform a second sequencing run. The second sequencing run can include a complete execution of a primary analysis of one or more biological samples by the nucleic acid sequencing device. In some implementations, the second sequencing run can sequence one or more biological samples that are different from those biological samples sequenced during the first sequencing run. The second sequencing run can include a clustering step, a first read interval, and a second read interval. In some implementations, the primary analysis can also include one or more indexing steps.
[0160] While performing a second sequencing run in step 840 using the nucleic acid sequencing device, (I) continuing to perform one or more secondary analysis operations on the first data or the second data 850 and (II) storing result data representing results of the secondary analysis operations. Continuing to perform one or more secondary analysis operations on the first data or the second data generated during step 810 or step 820, respectively, can include continuing to perform the secondary analysis on the first data and the second data until the secondary analysis on the first data and the second data is complete. For example, a hardwired mapping and alignment unit that may be configured in step 830 during the first sequencing run can continue to perform mapping and alignment operations on the first read, the second read, or both during the second sequencing run until the mapping and alignment operations on the first read, the second read, or both are complete.
[0161] 9 is a flowchart of an example process 900 for performing dynamic programmable circuit context switching. Generally, process 900 may include obtaining one or more genomic workflow attributes (910), determining a workflow context switching type for the programmable circuit based on the one or more genomic workflow attributes, where the workflow context switching type defines a reconfiguration of the programmable circuit (920), and instructing a programmable circuit controller to perform secondary analysis using the determined context switching type (930). For convenience, but not by way of limitation, these steps are described in more detail below as performed by a sequencing system, such as system 100 of FIG. 1A, system 300 of FIG. 3, or system 500 of FIG. 5, respectively.
[0162] The sequencing system can begin executing process 900 at step 910 by obtaining one or more genomic workflow attributes. In some implementations, the one or more workflow attributes can include a workflow identifier that identifies a workflow selected by a user of the nucleic acid sequencer. The genomic workflow can include, for example, a whole genome sequencing workflow, an enrichment workflow, an RNA workflow, an amplicon workflow, a single-cell RNA workflow, etc. Alternatively, or in addition, the one or more workflow attributes can include data describing the number of samples to be sequenced by the nucleic acid sequencer. Alternatively, or in addition, the one or more workflow attributes can include a predetermined time threshold for execution of the workflow. Alternatively, or in addition, the one or more workflow attributes can include an amount of available computational resources available to the nucleic acid sequencer.
[0163] The sequencing system may continue execution of process 900 by determining a workflow context switching type for the programmable circuit based on one or more genomic workflow attributes, the workflow context switching type defining a reconfiguration of the programmable circuit, at stage 920. Determining the workflow context switching type may include selecting a particular workflow context switching type from a plurality of context switching types based on the one or more workflow attributes.
[0164] The context switching type defines how the programmable circuit is dynamically reconfigured at run time. For example, a first programmable circuit context can include programmable circuit interlacing alignment and variant calling operations. In such implementations, the programmable circuit can be configured as a mapping and alignment unit that aligns reads corresponding to a first sample to a reference sequence, dynamically reconfigured as a variant calling unit that performs variant calling operations on reads corresponding to the first aligned sample, dynamically reconfigured as a variant calling unit that maps and aligns reads corresponding to a second sample to a reference sequence, dynamically reconfigured as a variant calling unit that performs variant calling operations on reads corresponding to the second aligned sample, and so on. In this context, the programmable circuit can dynamically switch between mapping and alignment and variant calling operations in both directions. This first programmable circuit context is preferred when there is only one sample or a small number of samples.
[0165] As another example, the second programmable circuit context can include a programmable circuit that performs all necessary alignments and then performs all necessary variant calling operations on the aligned reads. In such an implementation, the programmable circuit can be configured as a mapping and alignment unit, aligning a first sample, aligning a second sample, aligning a third, etc., until all samples are aligned, and then dynamically reconfigured as a variant calling unit that performs variant calling operations on the aligned first sample, performs variant calling operations on the second aligned sample, performs variant calling operations on the third aligned sample, etc. Because context switching is computationally intensive, this second programmable circuit context can be selected if the workflow has a large number of samples.
[0166] In some implementations, the sequencing system can determine between the aforementioned context switching types in several ways. For example, in some implementations, the sequencing system can acquire data such as a workflow identifier indicating a workflow selection by a user of the nucleic acid sequencer. In some implementations, the sequencing system can be programmed to automatically select a particular context switching type that is logically related to the acquired workflow identifier. The logical relationship can include, for example, a one-to-one mapping between the workflow identifier and the context switching type.
[0167] Alternatively, or in addition, the sequencing system can determine between the aforementioned context switching types based on the number of samples. For example, a predetermined threshold number of samples can be set. Thus, if the nucleic acid sequencer determines that a particular workflow exceeds the threshold number of samples, the nucleic acid sequencer can select the second programmable context. Alternatively, if the nucleic acid sequencer determines that the number of samples does not exceed the threshold number of samples, the nucleic acid sequencer can select the first programmable context.
[0168] Alternatively, or in addition, the sequencing system can determine between the aforementioned context switching types based on the estimated secondary analysis execution time. For example, the nucleic acid sequencer can be programmed to analyze the received workflow-descriptor data and estimate the estimated secondary analysis execution time using a default programmable circuit context, where the default programmable circuit context is the first programmable circuit context. In such an implementation, if the estimated secondary analysis execution time is less than a predetermined threshold time, the nucleic acid sequencer can select the first programmable circuit context. Alternatively, if the estimated secondary analysis execution time is greater than a predetermined threshold time, the nucleic acid sequencer can select the second programmable circuit context.
[0169] These aforementioned implementations are merely examples of programmable circuit context types and context switching that may be used in accordance with the present disclosure. None of these examples should be considered as limiting the scope of the present disclosure. Instead, other programmable circuit context types and context switching types are within the scope of the present disclosure.
[0170] The sequencing system may continue execution of process 900 at step 930 by instructing the programmable circuit controller to perform a secondary analysis using the determined context switching type. The programmable circuit controller may include software, hardware, or a combination of both that configures the programmable logic of the programmable circuit. Based on the received instruction, the programmable circuit controller may dynamically configure the programmable circuit to include hardwired digital logic configured to implement the context switching type identified by the instruction.
[0171] FIG. 10 is a block diagram of an example of system components that can be used to implement a system for performing incremental quadratic analysis.
[0172] Computing device 1000 is intended to represent various forms of digital computers (e.g., laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers). In some implementations, computing device 1000 may be a nucleic acid sequencer, such as the nucleic acid sequencers of FIG. 1, FIG. 3, or FIG. 5. Mobile computing device 1050 is intended to represent various forms of mobile equipment (e.g., personal digital assistants, cellular phones, smartphones, mobile embedded radio systems, wireless diagnostic computing devices, and other similar computing devices). The components, their connections and relationships, and their functions shown herein are intended to be exemplary only and not limiting.
[0173] The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 connecting to the memory 1004 and multiple high-speed expansion ports 1010, and a low-speed interface 1012 connecting to a low-speed expansion port 1014 and the storage device 1006. Each of the processor 1002, the memory 1004, the storage device 1006, the high-speed interface 1008, the high-speed expansion port 1010, and the low-speed interface 1012 are interconnected using various buses, which may be mounted on a common motherboard or otherwise as needed. The processor 1002 can process instructions for execution within the computing device 1000, including instructions stored in the memory 1004 or on the storage device 1006, to display graphical information for a GUI on an external input / output device, such as a display 1016 coupled to the high-speed interface 1008. In other implementations, multiple processors and / or multiple buses can be used, along with multiple memories and types of memory, as appropriate. Additionally, multiple computing devices can be connected, each providing a portion of the operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system). In some implementations, processor 1002 is a single-threaded processor. In some implementations, processor 1002 is a multi-threaded processor. In some implementations, processor 1002 is a quantum computer.
[0174] The memory 1004 stores information within the computing device 1000. In some implementations, the memory 1004 is a volatile memory unit(s). In other implementations, the memory 1004 is a non-volatile memory unit(s). The memory 1004 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0175] The storage device 1006 can provide mass storage for the computing device 1000. In one implementation, the storage device 1006 can be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. The instructions can be stored on an information medium. When executed by one or more processing devices (e.g., the processor 1002), the instructions perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as a computer- or machine-readable medium (e.g., the memory 1004, the storage device 1006, or memory on the processor 1002). The high-speed interface 1008 manages bandwidth-intensive operations of the computing device 1000, while the low-speed interface 1012 manages lower-bandwidth-intensive operations. This allocation of functionality is merely an example. In some implementations, the high-speed interface 1008 couples to memory 1004, a display 1016 (e.g., via a graphics processor or accelerator), and to a high-speed expansion port 1010 that can utilize various expansion cards (not shown). In this implementation, the low-speed controller 1012 is coupled to the storage device 1006 and the low-speed expansion port 1014. The low-speed expansion port 1014 (which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet)) can be coupled (e.g., via a network adapter) to one or more input / output devices (e.g., a keyboard, a pointing device, a scanner, or a network device (e.g., a switch or router)).
[0176] Computing device 1000 can be implemented in many different forms as shown. For example, the computing device can be implemented as a standard server 1020, or multiple times in a cluster of such servers. In addition, the computing device can be implemented in a personal computer, such as a laptop computer 1022. The computing device can also be implemented as part of a rack server system 1024. Alternatively, the components of computing device 1000 can be combined with other components of a mobile device (e.g., mobile computing device 1050). Each such device can include one or more of computing device 1000 and mobile computing device 1050, and the overall system can be composed of multiple computing devices communicating with each other.
[0177] The mobile computing device 1050 includes, among other components, a processor 1052, a memory 1064, an input / output device (e.g., a display 1054), a communication interface 1066, and a transceiver 1068. The mobile computing device 1050 may include a storage device (e.g., a microdrive or other device) to provide additional storage. Each of the processor 1052, the memory 1064, the display 1054, the communication interface 1066, and the transceiver 1068 may be interconnected using various buses, with some of the components being mounted on a common motherboard or otherwise as needed.
[0178] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 can be implemented as a chipset of chips including discrete and multiple analog and digital processors. The processor 1052 can provide, for example, coordination of other components of the mobile computing device 1050 (e.g., control of a user interface, execution of applications by the mobile computing device 1050, and wireless communication by the mobile computing device 1050).
[0179] The processor 1052 can communicate with a user through a control interface 1058 and a display interface 1056 coupled to a display 1054. The display 1054 can be, for example, a TFT (thin film transistor liquid crystal) display, an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 1056 can include appropriate circuitry to drive the display 1054 to present graphics and other information to the user. The control interface 1058 can receive commands from the user and translate them for transmission to the processor 1052. Additionally, the external interface 1062 can provide communication with the processor 1052 to enable near-field communication between other devices and the mobile computing device 1050. For example, the external interface 1062 can provide wired communication in some implementations or wireless communication in other implementations, and multiple interfaces can be used.
[0180] The memory 1064 stores information within the mobile computing device 1050. The memory 1064 may be implemented as one or more of a computer-readable medium(s), a volatile memory unit(s), or a non-volatile memory unit(s). Expansion memory 1074 may also be provided and connected to the mobile computing device 1050 via an expansion interface 1072, which may include, for example, a Single In-Line Memory Module (SIMM) card interface. The expansion memory 1074 may provide additional storage space for the mobile computing device 1050 or may also store applications or other information for the mobile computing device 1050. In particular, the expansion memory 1074 may include instructions that perform or complement the processes described above and may also include secure information. Thus, for example, the expansion memory 1074 may be provided as a security module for the mobile computing device 1050 and may be programmed with instructions that enable secure use of the mobile computing device 1050. Additionally, secure applications may be provided via SIMM cards with additional information, such as placing identifying information on the SIMM card in an unhackable manner.
[0181] The memory may include, for example, flash memory and / or NVRAM memory (nonvolatile random access memory), as described below. In some implementations, the instructions are stored on an information carrier such that, when executed by one or more processing devices (e.g., processor 1052), the instructions perform one or more methods, such as those described above. The instructions may also be stored by one or more storage devices, such as one or more computer- or machine-readable media (e.g., memory 1064, expansion memory 1074, or memory on processor 1052). In some implementations, the instructions may be received, for example, by transceiver 1068 or in a propagated signal on external interface 1062.
[0182] Mobile computing device 1050 may communicate wirelessly via communication interface 1066, which in some cases may include digital signal processing circuitry. Communication interface 1066 may provide communication under various modes or protocols, such as GSM voice telephony (Global System for Mobile Communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), LTE, 5G / 6G cellular, among others. Such communication may occur via transceiver 1068 using radio frequencies, for example. Additionally, short-range communication may occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, a Global Positioning System (GPS) receiver module 1070 may provide additional navigation-related and location-related wireless data to the mobile computing device 1050, which may be used as needed by applications running on the mobile computing device 1050.
[0183] The mobile computing device 1050 can also communicate audibly using an audio codec 1060, which can receive voice information from a user and convert it into usable digital information. The audio codec 1060 can also generate audible sounds for the user, such as through a speaker in the handset of the mobile computing device 1050. Such sounds can include sounds from voice telephone calls, can include recorded sounds (e.g., voice messages, music files, among others), and can also include sounds generated by applications running on the mobile computing device 1050.
[0184] The mobile computing device 1050 may be implemented in several different forms as shown in the figure. For example, the computing device may be implemented as a cellular telephone 1080. The computing device may also be implemented as part of a smartphone 1082, personal digital assistant, or other similar mobile device.
[0185] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, various forms of the flows described above may be used, and steps may be reordered, added, or removed.
[0186] All of the embodiments and functional operations of the invention described herein can be implemented in digital electronic circuitry, or computer software, firmware, or hardware, or one or more combinations thereof, including the structures disclosed herein and their structural equivalents. Embodiments of the invention can be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" can encompass all apparatuses, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiving apparatus.
[0187] A computer program (also known as a program, software, software application, script, or code) can be written in any type of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple associated files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer or on multiple computers, located at a single location or distributed across multiple locations and interconnected by a communications network.
[0188] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by processing input data and generating output. The processes and logic flows may also be performed by, and apparatus may also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0189] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes or is operatively coupled to one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, for receiving data from, transmitting data to, or both. However, a computer need not include such devices. Furthermore, a computer can be incorporated into another device, e.g., a tablet computer, a mobile phone, a personal digital assistant (PDA), a portable audio player, or a global positioning system (GPS) receiver, to name a few. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0190] To provide for interaction with a user, embodiments of the present invention can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with a user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic input, speech input, or tactile input.
[0191] Embodiments of the present invention can be implemented in a computing system including back-end components, e.g., as a data server, or middleware components, e.g., an application server, or front-end components, e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the present invention, or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks ("LANs") and wide area networks ("WANs"), e.g., the Internet.
[0192] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0193] While this specification contains many details, these should not be construed as limiting the scope of the invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, even if features may be described above as functioning in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be deleted from the combination, and the claimed combination may relate to a subcombination or a variation of a subcombination.
[0194] Similarly, while operations are depicted in the figures in a particular order, it should not be understood that such operations need to be performed in the particular order shown, or sequentially, or that all of the illustrated operations need to be performed, to achieve desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may typically be integrated together in a single software product or packaged within multiple software products.
[0195] In each instance where a particular file format is mentioned, other file types or formats may be substituted. For example, an HTML file may be replaced by an XML, JSON, plain text, or other type of file. Furthermore, where a particular data structure such as a table or hash table is mentioned, other data structures (such as a spreadsheet, relational database, or structured file) may be used in place of the mentioned data structure.
[0196] Other embodiments While the present invention has been described in conjunction with its detailed description, it should be understood that the above description is intended to be illustrative, but not limiting, of the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
[0197] While specific embodiments of the present invention have been described, other embodiments are within the scope of the following claims. For example, the steps recited in the claims can be performed in a different order and still achieve desirable results.
[0198] Several embodiments have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the present invention. Additionally, the logic flow depicted in the figures does not require the particular order shown, or sequential order, to achieve desired results. Additionally, other steps can be provided or steps can be eliminated from the described flow, and other components can be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the following claims. [Explanation of symbols]
[0199] 105 samples 106 samples 107 samples 110 Nucleic Acid Sequencer 112 Network 120 flow cells 140 Secondary Analysis Unit 142 Programmable Circuits 144 memory 149 Results 150 processing units 160 memory 162 Demultiplex Unit 164 variant calling units 170 Workflow 172 sequencing runs 310 Nucleic Acid Sequencer 320 Remote Computer 340 Secondary Analysis Unit 342 Programmable Circuits 344 memory 350 processing units 359 results 360 memory 362 Demultiplex Unit 364 variant calling units 510 Nucleic Acid Sequencer 540 Secondary Analysis Unit 542 Programmable Circuits 544 memory 549 results 550 processing units 560 memory 562 Demultiplex Unit 564 variant calling units 1000 computing devices 1002 processor 1004 memory 1006 Storage Device 1008 High-Speed Interface 1010 High-Speed Expansion Port 1012 Low-speed interface 1014 Low-Speed Expansion Port 1016 display 1020 Standard Server 1022 laptop computers 1024 rack server system 1050 Mobile Computing Devices 1052 processor 1054 display 1056 Display Interface 1058 Control Interface 1060 Audio Codec 1062 External Interface 1064 memory 1066 Communication Interface 1068 Transceiver 1070 (Global Positioning System) receiver module 1072 Extended Interface 1074 extended memory 1080 cellular phone 1082 smartphones
Claims
1. 1. A method for performing incremental secondary analysis of nucleic acid sequence reads, the method comprising: (i) acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval, each of the first reads representing a first ordered sequence of nucleotides; (ii) acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval, each of the second reads representing a second ordered sequence of nucleotides; (iii) at least in part, while the second data is being acquired; (a) providing the first data as input to a mapping and alignment unit by the nucleic acid sequencing device; (b) receiving alignment results from the mapping and alignment unit, the alignment results including data indicative of alignment quality; (c) storing the received alignment results; and after that, (iv) instructing the mapping and alignment unit to terminate alignment of the second data representing the plurality of second reads to a reference sequence based on determining that the alignment result does not meet a quality threshold; A method comprising:
2. one or more of the first reads includes data representing a first sample identifier; wherein one or more of the second reads includes data representing a second sample identifier, and the method further comprises: While the second data is being acquired, Organizing the one or more first reads into respective groups based on at least the first sample identifier or the second sample identifier; 10. The method of claim 1, further comprising generating tissue statistics, the tissue statistics indicating a number of first reads corresponding to each sample identifier.
3. 10. The method of claim 1, further comprising providing output data representing the stored alignment results corresponding to the plurality of first reads prior to aligning a second portion of a cluster of reads or during aligning the second portion of the cluster of reads.
4. 10. The method of claim 1, further comprising instructing the mapping and alignment module to initiate a subsequent alignment of the data representing the plurality of first reads to the reference sequence.
5. 10. The method of claim 1, further comprising, while acquiring the second data, determining a set of possible variants of the first data representing the plurality of first reads aligned to the reference sequence.
6. 10. The method of claim 1, wherein at least some of the second data representing the plurality of second reads are aligned while at least different portions of second data representing the plurality of second reads are acquired.
7. 2. The method of claim 1, wherein the mapping and alignment unit is instructed to begin aligning the second data representing the plurality of second reads a predetermined number of sequencing cycles before the second data is fully acquired.
8. 1. A system for performing incremental secondary analysis of nucleic acid sequence reads, the system comprising: a nucleic acid sequencing device; one or more memory devices storing instructions that, when executed by one or more processors of the nucleic acid sequencing device, cause the nucleic acid sequencing device to perform operations, the operations including: (i) acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval, each of the first reads representing a first ordered sequence of nucleotides; (ii) acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval, each of the second reads representing a second ordered sequence of nucleotides; (iii) at least in part, while the second data is being acquired; (a) providing the first data as input to a mapping and alignment unit by the nucleic acid sequencing device; (b) receiving alignment results from the mapping and alignment unit, the alignment results including data indicative of alignment quality; (c) storing the received alignment results; and after that, (iv) instructing the mapping and alignment unit to terminate alignment of the second data representing the plurality of second reads to a reference sequence based on determining that the alignment result does not meet a quality threshold.
9. The system of claim 8 , wherein at least a portion of the mapping and alignment unit is implemented using a programmable logic device.
10. 10. The system of claim 9, wherein the programmable logic device is a field programmable gate array (FPGA).
11. The system of claim 8 , wherein at least a portion of the mapping and alignment unit is implemented using an application specific integrated circuit (ASIC).
12. The system of claim 8 , wherein the mapping and alignment unit is included within the nucleic acid sequencing device.
13. one or more of the first reads includes data representing a first sample identifier; one or more of the second reads includes data representing a second sample identifier, and the operation comprises: While the second data is being acquired, Organizing the one or more first reads into respective groups based on at least the first sample identifier or the second sample identifier; 10. The system of claim 8, further comprising generating tissue statistics, the tissue statistics indicating a number of first reads corresponding to each sample identifier.
14. The operation is 10. The system of claim 8, further comprising: providing output data representing the stored alignment results corresponding to the plurality of first reads prior to aligning a second portion of a cluster of reads or during aligning the second portion of the cluster of reads.
15. The operation is 10. The system of claim 8, further comprising instructing the mapping and alignment module to initiate a subsequent alignment of the data representing the plurality of first reads to the reference sequence.
16. The operation is 10. The system of claim 9, further comprising, while acquiring the second data, determining a set of possible variants of the first data representing the plurality of first reads aligned to the reference sequence.
17. 11. The system of claim 10, wherein at least some of the second data representing the plurality of second reads are aligned while at least different portions of second data representing the plurality of second reads are acquired.
18. 12. The system of claim 11, wherein the mapping and alignment unit is instructed to begin aligning the second data representing the plurality of second reads a predetermined number of sequencing cycles before the second data is fully acquired.
19. 1. A computer-readable storage medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations including: (i) acquiring first data describing a plurality of first reads generated by a nucleic acid sequencing device during a first read interval, each of the first reads representing a first ordered sequence of nucleotides; (ii) acquiring second data describing a plurality of second reads generated by the nucleic acid sequencing device during a second read interval performed after the first read interval, each of the second reads representing a second ordered sequence of nucleotides; (iii) at least in part, while the second data is being acquired; (a) providing the first data as input to a mapping and alignment unit by the nucleic acid sequencing device; (b) receiving alignment results from the mapping and alignment unit, the alignment results including data indicative of alignment quality; (c) storing the received alignment results; and after that, (iv) instructing the mapping and alignment unit to terminate alignment of the second data representing the plurality of second reads to a reference sequence based on determining that the alignment result does not meet a quality threshold.
20. one or more of the first reads includes data representing a first sample identifier; one or more of the second reads includes data representing a second sample identifier, and the operation comprises: While the second data is being acquired, Organizing the one or more first reads into respective groups based on at least the first sample identifier or the second sample identifier; 20. The computer-readable storage medium of claim 19, further comprising generating tissue statistics, the tissue statistics indicating a number of first reads corresponding to each sample identifier.
21. The operation is 20. The computer-readable storage medium of claim 19, further comprising: providing output data representing the stored alignment results corresponding to the plurality of first reads prior to aligning a second portion of a cluster of reads or during aligning the second portion of the cluster of reads.
22. The operation is 20. The computer-readable storage medium of claim 19, further comprising instructing the mapping and alignment module to initiate a subsequent alignment of the data representing the plurality of first reads to the reference sequence.
23. The operation is 20. The computer-readable storage medium of claim 19, further comprising: while acquiring the second data, determining a set of possible variants of the first data representing the plurality of first reads aligned to the reference sequence.
24. 20. The computer-readable storage medium of claim 19, wherein at least some of the second data representing the plurality of second reads are aligned while at least different portions of second data representing the plurality of second reads are acquired.
25. 20. The computer-readable storage medium of claim 19, wherein the mapping and alignment unit is instructed to begin aligning the second data representing the plurality of second reads a predetermined number of sequencing cycles before the second data is fully acquired.