A self-learning base collab trained using oligo sequences
Patent Information
- Application Number
- JP2023579783
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-01
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-06-29
AI Technical Summary
Deploying deep convolutional neural networks (CNNs) in portable and embedded systems is challenging due to large data volumes, intensive computations, and frequent memory accesses, leading to inefficiencies in resource utilization and power consumption.
Designing efficient data flow and hardware architectures for CNN acceleration using Field Programmable Gate Arrays (FPGAs) to minimize data communication and maximize resource utilization, enabling high performance and flexibility in inference processes.
The proposed solution achieves high-performance, low-latency, and energy-efficient CNN inference on portable systems by optimizing data flow and hardware architecture, addressing the challenges of resource utilization and power consumption.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Priority Application This application claims priority to U.S. Nonprovisional Patent Application No. 17 / 830,287, entitled "Self-Learned Base Caller, Trained Using Oligo Sequences," filed on June 1, 2022 (Attorney Docket No. ILLM1038-3 / IP-2050-US), which claims the benefit of U.S. Provisional Patent Application No. 63 / 216,419, entitled "Self-Learned Base Caller, Trained Using Oligo Sequences," filed on June 29, 2021 (Attorney Docket No. ILLM1038-1 / IP-2050-PRV), and U.S. Provisional Patent Application No. 63 / 216,404, entitled "Self-Learned Base Caller, Trained Using Organism Sequences," filed on June 29, 2021 (Attorney Docket No. ILLM1038-2 / IP-2094-PRV). The priority application is incorporated herein by reference for all purposes.
[0002] This application claims priority to U.S. Nonprovisional Patent Application No. 17 / 830,316, entitled "Self-Learned Base Caller, Trained Using Organism Sequences," filed on June 1, 2022 (Attorney Docket No. ILLM1038-5 / IP-2094-US), which claims the benefit of U.S. Provisional Patent Application No. 63 / 216,404, entitled "Self-Learned Base Caller, Trained Using Organism Sequences," filed on June 29, 2021 (Attorney Docket No. ILLM1038-2 / IP-2094-PRV), and U.S. Provisional Patent Application No. 63 / 216,419, entitled "Self-Learned Base Caller, Trained Using Oligo Sequences," filed on June 29, 2021 (Attorney Docket No. ILLM1038-1 / IP-2050-PRV). The priority application is incorporated herein by reference for all purposes.
[0003] The disclosed technology relates to artificial intelligence based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data.
[0004] Built-in The following are incorporated by reference as if fully set forth herein: A concurrently filed PCT patent application entitled "SELF-LEARNED BASE CALLER, TRAINED USING ORGANISM SEQUENCES" (Attorney Docket No. ILLM ILLM1038-6 / IP-2094-PCT); U.S. Provisional Patent Application No. 62 / 979,384, entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES,” filed on February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV); U.S. Provisional Patent Application No. 62 / 979,414, entitled “ARTIFICIAL INTELLIGENCE-BASED MANY-TO-MANY BASE CALLING,” filed on February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV); U.S. Nonprovisional Patent Application No. 16 / 825,987, entitled “TRAINING DATA GENERATION FOR ARTIFICIAL INTELLIGENCE-BASED SEQUENCING,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-16 / IP-1693-US); U.S. Nonprovisional Patent Application No. 16 / 825,991, entitled “ARTIFICIAL INTELLIGENCE-BASED GENERATION OF SEQUENCING METADATA,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-17 / IP-1741-US); U.S. Nonprovisional Patent Application No. 16 / 826,126, entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLING,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-18 / IP-1744-US); U.S. Nonprovisional Patent Application No. 16 / 826,134, entitled “ARTIFICIAL INTELLIGENCE-BASED QUALITY SCORING,” filed on March 20, 2020 (Attorney Docket No. ILLM1008-19 / IP-1747-US); and U.S. Patent Application Publication No. 16 / 826,168, entitled “ARTIFICIAL INTELLIGENCE-BASED SEQUENCING,” filed March 21, 2020 (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV-US). [Background technology]
[0005] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to implementations of the claimed technology.
[0006] Rapid improvements in computing power have enabled deep Convolutional Neural Networks (CNNs) to achieve great success in many computer vision tasks in recent years, with significantly improved accuracy. During the inference stage, many applications require low-latency processing of a single image with strict power consumption requirements, which reduces the efficiency of Graphics Processing Units (GPUs) and other general-purpose platforms, which creates an opportunity for specific acceleration hardware, such as Field Programmable Gate Arrays (FPGAs), by customizing digital circuits to be particularly effective for inference of deep learning algorithms. However, deploying CNNs in portable and embedded systems remains challenging due to large data volumes, intensive computations, various algorithm structures, and frequent memory accesses.
[0007] Since convolution provides most of the operations in CNN, the convolution acceleration scheme will greatly affect the efficiency and performance of hardware CNN accelerators. Convolution involves multiply and accumulate (MAC) operations with four levels of loops that slide along the kernel and feature maps. The first loop level calculates the MAC of pixels in one kernel window. The second loop level accumulates the sum of MAC products over various different input feature maps. After completing the first and second loop levels, the final output elements in the output feature map are obtained by adding a bias. The third loop level slides the kernel window in the input feature map. The fourth loop level generates various different output feature maps.
[0008] FPGAs, especially for accelerating inference tasks, have attracted more interest and become more widespread because they (1) have high reconfigurability, (2) are superior to application specific integrated circuits (ASICs) in terms of the development time required to catch up with the rapid evolution of CNNs, (3) have good performance, and (4) are more energy efficient than GPUs. The high performance and efficiency of FPGAs can be achieved by synthesizing circuits customized for specific calculations and directly processing billions of operations with customized memory systems. For example, hundreds to thousands of digital signal processing (DSP) blocks in modern FPGAs support core convolution operations, such as multiply-add operations with high parallelism. Dedicated data buffers between external on-chip memory and on-chip processing engines (PEs) can be designed to realize prioritized data flow by configuring tens of megabytes of on-chip block random access memory (BRAM) on the field programmable gate array (FPGA) chip. Summary of the Invention [Problem to be solved by the invention]
[0009] Efficient data flow and hardware architecture for CNN acceleration is desired to minimize data communication while maximizing resource utilization to achieve high performance. This creates an opportunity to design methodologies and frameworks to accelerate the inference process of various CNN algorithms on acceleration hardware and achieve high performance, high efficiency, and high flexibility. [Brief description of the drawings]
[0010] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Figure 1] 1 shows a cross-sectional view of a biosensor that can be used in various embodiments. [Diagram 2] 1 shows one implementation of a flow cell that includes clusters within its tiles. [Diagram 3] An exemplary flow cell with eight lanes is shown, along with a zoom-in of one tile and its cluster and their surrounding background. [Figure 4] FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output. [Diagram 5] FIG. 2 is a simplified diagram illustrating aspects of a base call operation, including functions of a runtime program executed by a host processor. [Figure 6] 5 is a simplified diagram of a configuration of a configurable processor, such as the configurable processor of FIG. 4. [Figure 7] FIG. 1 is a diagram of a neural network architecture that can be implemented using a configurable or reconfigurable array configured as described herein. [Figure 8A] FIG. 8 is a simplified diagram of an organization of tiles of sensor data for use by a neural network architecture such as that of FIG. [Figure 8B] FIG. 8 is a simplified diagram of a patch of tiles of sensor data used by a neural network architecture such as that of FIG. [Figure 9] 8 illustrates part of the configuration of a neural network such as that of FIG. 7 on a configurable or reconfigurable array such as a field programmable gate array (FPGA). [Figure 10]FIG. 13 is a diagram of another alternative neural network architecture that can be implemented using a configurable or reconfigurable array configured as described herein. [Figure 11] 1 shows one implementation of a dedicated architecture of a neural network-based base caller used to separate the processing of data in different sequencing cycles. [Figure 12] 1 illustrates one implementation of separated layers, each of which may contain convolutions. [Figure 13A] 1 illustrates one implementation of combinational layers, each of which may include convolutions. [Figure 13B] 14 illustrates another implementation of combination layers, each of which may include convolutions. [Figure 14A] We illustrate a base-calling system that operates in a single oligo training stage to train a base-caller that contains a neural network architecture using known synthetic oligo sequences. [Figure 14A1] 1 illustrates a comparison operation between a predicted sequence and the corresponding ground truth sequence. [Figure 14B] FIG. 14B illustrates further details of the base-calling system of FIG. 14A operating in a single oligo training stage to train a base-caller comprising a neural network architecture using known synthetic oligo sequences. [Figure 15A] FIG. 14A illustrates the base calling system operating in the training data generation phase of a two-oligo training stage to generate labeled training data using two known synthetic sequences. [Figure 15B] Illustrates two corresponding exemplary selections of the two oligo sequences discussed with respect to FIG. 15A. [Figure 15C] Illustrates two corresponding exemplary selections of the two oligo sequences discussed with respect to FIG. 15A. [Figure 15D]Illustrated are exemplary mapping operations for either (i) mapping a predicted base call sequence to either the first oligo or the second oligo, or (ii) declaring uncertainty in mapping a predicted base call sequence to either of the two oligos. [Figure 15E] FIG. 15D illustrates labeled training data generated from the mapping of FIG. 15D, which is used by another neural network configuration illustrated in FIG. 16A. [Figure 16A] 14A illustrates the base-calling system of FIG. 14A operating in the training data consumption and training phases of a two-oligo training stage to train a base-caller with another neural network configuration (different from and more complex than the neural network configuration of FIG. 14A ) using two known synthetic oligo sequences. [Figure 16B] 14B illustrates the base calling system of FIG. 14A operating in a second iteration of the training data generation phase of a two-oligo training stage. [Figure 16C] FIG. 16B illustrates labeled training data generated from the illustrated mapping, which is used for further training. [Figure 16D] 14A illustrates the base-calling system of FIG. 14A operating in the second iteration of the "training data consumption and training phase" of the "two-oligo training stage" to train a base-caller with the neural network configuration of FIG. 16A using two known synthetic oligo sequences. [Figure 17A] 1 illustrates a flow chart depicting an exemplary method for iteratively training a neural network configuration for base calling using single oligo and two oligo sequences. [Figure 17B] 17 illustrates exemplary labeled training data generated by the Pth NN configuration at the end of the method 1700 of FIG. 17A. [Figure 18A]14A illustrates the base-calling system of FIG. 14A operating in a first iteration of the "training data consumption and training phase" of the "3-oligo training stage" to train a base-caller with a 3-oligo neural network configuration. [Figure 18B] 14A is illustrated operating in the "training data generation phase" of the "three-oligo training stage" to train a base-caller comprising the three-oligo neural network configuration of FIG. 18A. [Figure 18C] Illustrated is a mapping operation that either (i) maps the predicted base call sequence to any of the three oligos in FIG. 18B, or (ii) declares the mapping of the predicted base call sequence to be uncertain. [Figure 18D] FIG. 18C shows labeled training data generated from the mapping, which is used to train another neural network configuration. [Figure 18E] 1 illustrates a flowchart depicting an exemplary method for iteratively training a neural network configuration for base calling using 3-oligo ground truth sequences. [Figure 19] 1 illustrates a flowchart depicting an exemplary method for iteratively training a neural network configuration for base calling using multiple oligo ground truth sequences. [Figure 20A] 14B illustrates the biological sequences used to train the base caller of FIG. 14A. [Figure 20B] 20A illustrates the base calling system of FIG. 14A operating in the training data generation phase of a first organism training stage to train a base calling system comprising a first organism-level neural network configuration using various subsequences of the first organism sequence of FIG. [Figure 20C] 1 illustrates an example of fading in signal intensity as a function of cycle number in a sequencing run of a base calling operation. [Figure 20D]1 conceptually illustrates the decreasing signal-to-noise ratio as cycles of sequencing progress. [Figure 20E] 20 illustrates base calling of the first L2 bases of the L1 bases of a partial sequence, where the first L2 bases of the partial sequence are used to map the partial sequence to the biological sequence of FIG. 20A. [Figure 20F] 20E illustrates labeled training data generated from the mapping of FIG. 20E, where the labeled training data includes a section of the biological sequence of FIG. 20A as ground truth. [Figure 20G] 14A illustrates the base-calling system of FIG. 14A operating in a "training data consumption and training phase" of an "organism-level training stage" to train a base-caller comprising a first organism-level neural network configuration. [Figure 21] 20A illustrates a flowchart depicting an exemplary method for iteratively training a neural network configuration for base calling using the simple biological sequence of FIG. [Figure 22] 14B illustrates the use of complex biological sequences for training a corresponding NN configuration on the base collaborators of FIG. 14A. [Figure 23A] 1 illustrates a flowchart depicting an exemplary method for iteratively training a neural network configuration for basecalling. [Figure 23B] Illustrates various charts illustrating the effectiveness of the bass colleague training process discussed in this disclosure. [Figure 23C] Illustrates various charts illustrating the effectiveness of the bass colleague training process discussed in this disclosure. [Figure 23D] Illustrates various charts illustrating the effectiveness of the bass colleague training process discussed in this disclosure. [Figure 23E] Illustrates various charts illustrating the effectiveness of the bass colleague training process discussed in this disclosure. [Figure 24] FIG. 1 is a block diagram of a base calling system according to one implementation. [Diagram 25]FIG. 25 is a block diagram of a system controller that can be used in the system of FIG. 24. [Figure 26] FIG. 1 is a simplified block diagram of a computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] As used herein, the term "polynucleotide" or "nucleic acid" refers to deoxyribonucleic acid (DNA); however, where appropriate, one of skill in the art will recognize that the systems and devices herein can also be utilized with ribonucleic acid (RNA). These terms should be understood to include, as equivalents, analogs of either DNA or RNA made from nucleotide analogs. As used herein, these terms also encompass complementary cDNA or copy DNA generated from an RNA template, for example, by the action of reverse transcriptase.
[0012] The single-stranded polynucleotide molecules sequenced by the systems and devices herein can originate in single-stranded form as DNA or RNA, or in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, PCR and amplification products, and the like). Thus, the single-stranded polynucleotide can be the sense or antisense strand of a polynucleotide duplex. Methods for preparing single-stranded polynucleotide molecules suitable for use in the methods of the present disclosure using standard techniques are known in the art. The exact sequence of the primary polynucleotide molecule is generally not critical to the present disclosure and can be known or unknown. The single-stranded polynucleotide molecule can represent a genomic DNA molecule (e.g., human genomic DNA), including both intron and exon sequences (coding sequences), as well as non-coding regulatory sequences such as promoter and enhancer sequences.
[0013] In certain embodiments, the nucleic acid to be sequenced through the use of the present disclosure is immobilized on a substrate (e.g., a substrate in a flow cell, or one or more beads on a substrate such as a flow cell). The term "immobilized" as used herein is intended to encompass direct or indirect, covalent or non-covalent attachment, unless otherwise indicated explicitly or by context. In certain embodiments, covalent attachment may be preferred, but generally what is required is that the molecule (e.g., nucleic acid) remains immobilized or associated with the support under conditions in which the support is intended to be used, e.g., in applications requiring nucleic acid sequencing.
[0014] The term "solid support" (or "substrate" in some usages) as used herein refers to any inert substrate or matrix to which nucleic acids can be attached, such as, for example, a glass surface, a plastic surface, latex, dextran, a polystyrene surface, a polypropylene surface, a polyacrylamide gel, a gold surface, a silicon wafer, etc. In many embodiments, the solid support is a glass surface (e.g., the flat surface of a flow cell channel). In certain embodiments, the solid support may be comprised of an inert substrate or matrix that has been "functionalized" by the application of a layer or coating of an intermediate material that contains reactive groups that allow for covalent attachment to a molecule, such as a polynucleotide. As a non-limiting example, such a support may include a polyacrylamide hydrogel supported on an inert substrate such as glass. In such embodiments, the molecule (polynucleotide) may be covalently attached directly to the intermediate material (e.g., a hydrogel), while the intermediate material may itself be non-covalently attached to the substrate or matrix (e.g., a glass substrate). Covalent attachment to a solid support should be interpreted accordingly to encompass this type of arrangement.
[0015] As mentioned above, the present disclosure includes novel systems and devices for sequencing nucleic acids. As will be clear to those skilled in the art, reference herein to a particular nucleic acid sequence may also refer to a nucleic acid molecule that includes such a nucleic acid sequence, depending on the context. Sequencing a target fragment means that a chronological reading of the bases is established. The bases read do not have to be consecutive, which is preferred, but it is also not necessary that all bases on the entire fragment are sequenced during sequencing. Sequencing can be performed using any suitable sequencing technique in which nucleotides or oligonucleotides are added consecutively to a free 3' hydroxyl group, resulting in the synthesis of a polynucleotide chain in a 5' to 3' direction. The nature of the added nucleotide is preferably determined after each nucleotide addition. Sequencing techniques that use sequencing by ligation, in which not all consecutive bases are sequenced, and techniques such as massively parallel signature sequencing (MPSS), which removes bases rather than adding them to a surface strand, are also suitable for use with the systems and devices of the present disclosure.
[0016] In an embodiment, the present disclosure discloses sequencing-by-synthesis (SBS), in which four fluorescently labeled modified nucleotides are used to sequence high-density clusters (potentially millions of clusters) of amplified DNA present on the surface of a substrate (e.g., a flow cell). Various additional aspects of SBS procedures and methods that can be utilized with the systems and devices herein are disclosed, for example, in WO 04018497, WO 04018493, and U.S. Pat. No. 7,057,026 (nucleotides), WO 05024010 and WO 06120433 (polymerases), WO 05065814 (surface attachment techniques), and WO 9844151, WO 06064199, and WO 07010251, the contents of each of which are incorporated herein by reference in their entirety.
[0017] In a particular use of the system / device herein, a flow cell containing a nucleic acid sample for sequencing is placed in a suitable flow cell holder. The sample for sequencing can take the form of a single molecule, an amplified single molecule in the form of a cluster, or a bead containing molecules of nucleic acid. The nucleic acid is prepared to include an oligonucleotide primer flanking an unknown target sequence. To initiate the first SBS sequencing cycle, one or more differently labeled nucleotides, and a DNA polymerase, etc., are flowed into / through the flow cell by a fluid flow subsystem (various embodiments of which are described herein). A single nucleotide can be added at a time, or the nucleotides used in the sequencing procedure can be specifically designed to have a reversible termination nature, thus allowing each cycle of the sequencing reaction to occur simultaneously in the presence of all four labeled nucleotides (A, C, T, G). When the four nucleotides are mixed together, the polymerase can select and incorporate the correct base, and each sequence is extended by a single base. In such methods of using the system, the natural competition between the four options results in greater accuracy than if only one nucleotide was present in the reaction mixture (so most of the sequence would not be exposed to the correct nucleotide). Sequences in which a particular base is repeated one after the other (e.g., homopolymers) are addressed with high accuracy, as are any other sequences.
[0018] The fluid flow subsystem also flows appropriate reagents to remove blocked 3' ends (if appropriate) and fluorophores from each incorporated base. The substrate can be exposed to either a second round of the four blocked nucleotides, or, optionally, a second round with a different individual nucleotide. Such cycles are then repeated, and the sequence of each cluster is read over multiple chemical cycles. Computer aspects of the present disclosure can optionally align sequence data collected from each single molecule, cluster, or bead to determine the sequence of longer polymers, etc. Alternatively, image processing and alignment can be performed on separate computers.
[0019] The heating / cooling components of the system regulate the reaction conditions within the flow cell channel and the reagent storage areas / containers (and optionally the camera, optics, and / or other components), while the fluid flow components allow the substrate surface to be exposed to the appropriate reagents for incorporation (e.g., appropriate fluorescently labeled nucleotides to be incorporated) while non-incorporated reagents are washed away. An optional movable stage on which the flow cell sits allows the flow cell to be properly oriented for laser (or other light) excitation of the substrate, and optionally allows the flow cell to be moved relative to the objective lens to read different areas of the substrate. In addition, other components of the system are also optionally movable / adjustable (e.g., camera, objective lens, heaters / coolers, etc.). During laser excitation, images / locations of the fluorescence emitted from the nucleic acids on the substrate are captured by the camera component, thereby recording the identity of the first base for each single molecule, cluster, or bead in the computer component.
[0020] The embodiments described herein may be used in a variety of biological or chemical processes and systems for academic or commercial analysis. More specifically, the embodiments described herein may be used in a variety of processes and systems in which it is desirable to detect an event, characteristic, quality, or property indicative of a desired response. For example, the embodiments described herein include cartridges, biosensors, and components thereof, as well as bioassay systems that operate with the cartridges and biosensors. In certain embodiments, the cartridges and biosensors include a flow cell and one or more sensors, pixels, photodetectors, or photodiodes that are bonded together in a substantially single structure.
[0021] The following detailed description of certain embodiments may be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks do not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., a processor or memory) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or random access memory, hard disk, etc.). Similarly, a program may be a stand-alone program, may be incorporated as a subroutine in an operating system, may be a function in an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentalities shown in the figures.
[0022] As used herein, elements or steps described in the singular and followed by the word "a" or "an" should be understood as not excluding a plurality of those elements or steps, unless such exclusion is expressly stated. Furthermore, references to "one embodiment" are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. Furthermore, unless expressly stated to the contrary, embodiments that "comprise" or "have" or "include" an element or elements having a particular characteristic may include the additional elements, whether or not they have that characteristic.
[0023] As used herein, a "desired reaction" includes a change in at least one of the chemical, electrical, physical, or optical properties (or qualities) of the analyte of interest. In certain embodiments, the desired reaction is a positive binding event (e.g., incorporation of a fluorescently labeled biomolecule into the analyte of interest). More generally, the desired reaction may be a chemical conversion, chemical change, or chemical interaction. The desired reaction may also be a change in an electrical property. For example, the desired reaction may be a change in the concentration of an ion in a solution. Exemplary reactions include, but are not limited to, chemical reactions such as reduction, oxidation, addition, elimination, rearrangement, esterification, amidation, etherification, cyclization, or substitution; binding interactions in which a first chemical binds to a second chemical; dissociation reactions in which two or more chemicals separate from each other; fluorescence, luminescence, bioluminescence, chemiluminescence, and biological reactions such as nucleic acid replication, nucleic acid amplification, nucleic acid hybridization, nucleic acid ligation, phosphorylation, enzyme catalysis, receptor binding, or ligand binding. The desired reaction may also be the addition or removal of a proton, which is detectable, for example, as a change in the pH of the surrounding solution or environment. An additional desired response can be the detection of ion flow across a membrane (e.g., a natural or synthetic bilayer membrane); for example, when ions flow through the membrane, the current is perturbed and this perturbation can be detected.
[0024] In certain embodiments, the desired reaction includes incorporation of a fluorescently labeled molecule into the analyte. The analyte may be an oligonucleotide and the fluorescently labeled molecule may be a nucleotide. The desired reaction may be detected when excitation light is directed to the oligonucleotide with the labeled nucleotide and the fluorophore emits a detectable fluorescent signal. In alternative embodiments, the detected fluorescence is the result of chemiluminescence or bioluminescence. The desired reaction may also increase fluorophore (or Forster) Resonance Energy Transfer (FRET), for example, by bringing a donor fluorophore into close proximity with an acceptor fluorophore, decrease FRET by separating the donor and acceptor fluorophores, increase fluorescence by separating the quencher from the fluorophore, or decrease fluorophore by colocalizing the quencher and fluorophore.
[0025] As used herein, "reaction component" or "reactant" includes any substance that can be used to obtain a desired reaction. For example, reaction components include reagents, enzymes, samples, other biomolecules, and buffers. Reaction components are typically delivered to the reaction site in solution and / or immobilized at the reaction site. A reaction component may directly or indirectly interact with another substance, such as an analyte of interest.
[0026] As used herein, the term "reaction site" is a localized area where a desired reaction can occur. A reaction site may include a support surface of a substrate on which a substance may be immobilized. For example, a reaction site may include a substantially planar surface within a channel of a flow cell having a colony of nucleic acids thereon. Typically, but not always, the nucleic acids in the colonies have the same sequence, e.g., are clonal copies of a single-stranded or double-stranded template. However, in some embodiments, a reaction site may contain only a single nucleic acid molecule, e.g., in single-stranded or double-stranded form. Furthermore, multiple reaction sites may be distributed non-uniformly along the support surface or may be arranged in a predetermined manner (e.g., parallel in a matrix such as a microarray). A reaction site may also include a reaction chamber (or well) that at least partially defines a spatial region or volume configured to compartmentalize a desired reaction.
[0027] This application uses the terms "reaction chamber" and "well" interchangeably. As used herein, the term "reaction chamber" or "well" includes a spatial region in fluid communication with a flow channel. A reaction chamber may be at least partially isolated from the surrounding environment or other spatial regions. For example, multiple reaction chambers may be separated from each other by a shared wall. As a more specific example, a reaction chamber may include a cavity defined by an inner surface of the well and have an opening or aperture such that the cavity is in fluid communication with the flow channel. A biosensor including such a reaction chamber is described in more detail in International Application PCT / US2011 / 057111, filed October 20, 2011, the entirety of which is incorporated herein by reference.
[0028] In some embodiments, the reaction chamber is sized and shaped relative to a solid (including a semi-solid) so that the solid can be fully or partially inserted therein. For example, the reaction chamber is sized and shaped to accommodate only one capture bead. The capture bead may have clonally amplified DNA or other material thereon. Alternatively, the reaction chamber is sized and shaped to receive an approximate number of beads or solid substrates. As another example, the reaction chamber may also be filled with a porous gel or material configured to control diffusion or filter fluids that may flow into the reaction chamber.
[0029] In some embodiments, a sensor (e.g., a photodetector, photodiode) is associated with a corresponding pixel area of the sample surface of the biosensor. Thus, a pixel area is a geometric construct that represents an area on the sample surface of the biosensor of one sensor (or pixel). The sensor associated with a pixel area detects luminescence collected from the associated pixel area when a desired reaction occurs at a reaction site or reaction chamber above the associated pixel area. In flat surface embodiments, the pixel areas can overlap. In some cases, multiple sensors can be associated with a single reaction site or a single reaction chamber. In other cases, a single sensor can be associated with a group of reaction sites or a group of reaction chambers.
[0030] As used herein, a "biosensor" includes a structure having multiple reaction sites and / or reaction chambers (or wells). The biosensor may include a solid-state imaging device (e.g., a CCD or CMOS imager) and, optionally, a flow cell attached thereto. The flow cell may include at least one flow channel in fluid communication with the reaction sites and / or reaction chambers. As one particular example, the biosensor is configured to fluidly and electrically couple to a bioassay system. The bioassay system may deliver reactants to the reaction sites and / or reaction chambers according to a predetermined protocol (e.g., sequencing by synthesis) and perform multiple imaging events. For example, the bioassay system may direct solutions to flow along the reaction sites and / or reaction chambers. At least one of the solutions may include four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides located in the reaction sites and / or reaction chambers. The bioassay system may then illuminate the reaction sites and / or reaction chambers using an excitation light source (e.g., a solid-state light source such as a light emitting diode or LED). The excitation light may have a predetermined wavelength or multiple wavelengths including a range of wavelengths. The excited fluorescent label provides a luminescent signal that can be captured by a sensor.
[0031] In alternative embodiments, the biosensor may include electrodes or other types of sensors configured to detect other distinguishable characteristics. For example, the sensor may be configured to detect changes in ion concentration. In another example, the sensor may be configured to detect the flow of ionic current across a membrane.
[0032] As used herein, a "cluster" is a colony of similar or identical molecules or nucleotide sequences or DNA strands. For example, a cluster can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides with the same or similar sequences. In other embodiments, a cluster can be any element or group of elements that occupy a physical area on a sample surface. In embodiments, the cluster is immobilized in a reaction site and / or reaction chamber during the base calling cycle.
[0033] As used herein, the term "immobilized" when used in reference to a biomolecule or biological material or chemical includes substantially attaching the biomolecule or biological material or chemical to a surface at a molecular level. For example, the biomolecule or biological material or chemical may be immobilized to the surface of a substrate material using adsorption techniques including non-covalent bonding (e.g., electrostatic forces, van der Waals, and hydrophobic interfacial dehydration), as well as covalent bonding techniques in which a functional group or linker facilitates attachment of the biomolecule to the surface. Immobilizing the biomolecule or biological material or chemical to the surface of a substrate material may be based on the properties of the substrate surface, the liquid medium carrying the biomolecule or biological material or chemical, and the properties of the biomolecule or biological material or chemical itself. In some cases, the substrate surface may be functionalized (e.g., chemically or physically modified) to facilitate immobilization of the biomolecule (or biological material or chemical) to the surface. The substrate surface may first be modified to have functional groups attached to the surface. The functional groups may then bind to the biomolecule or biological material or chemical to immobilize them thereon. Substances can be immobilized on a surface via a gel, for example, as described in US Patent Application Publication No. 2011 / 0059865(A1), which is incorporated herein by reference.
[0034] In some embodiments, nucleic acids can be attached to a surface and amplified using bridge amplification. Useful bridge amplification methods are described, for example, in U.S. Patent No. 5,641,658, International Publication No. WO 2007 / 010251, U.S. Patent No. 6,090,592, U.S. Patent Application Publication No. 2002 / 0055100 (A1), U.S. Patent No. 7,115,400, U.S. Patent Application Publication No. 2004 / 0096853 (A1), U.S. Patent Application Publication No. 2004 / 0002090 (A1), U.S. Patent Application Publication No. 2007 / 0128624 (A1), and U.S. Patent Application Publication No. 2008 / 0009420 (A1), each of which is incorporated herein in its entirety. Another useful method for amplifying nucleic acids on a surface is Rolling Circle Amplification (RCA), for example, using the methods described in more detail below. In some embodiments, the nucleic acid may be attached to a surface and amplified using one or more primer pairs. For example, one of the primers may be in solution and the other primer may be immobilized (e.g., 5'-attached) on the surface. As an example, a nucleic acid molecule may hybridize to one of the primers on the surface, followed by extension of the immobilized primer to generate a first copy of the nucleic acid. The primer in solution then hybridizes to the first copy of the nucleic acid, which may be extended using the first copy of the nucleic acid as a template. Optionally, after the first copy of the nucleic acid is generated, the original nucleic acid molecule may hybridize to a second immobilized primer on the surface and be extended simultaneously or after the primer in solution is extended. In any embodiment, repeated rounds of extension (e.g., amplification) using the immobilized primer and the primer in solution provide multiple copies of the nucleic acid.
[0035] In certain embodiments, the assay protocols performed by the systems and methods described herein include the use of naturally occurring nucleotides and enzymes configured to interact with the naturally occurring nucleotides. Naturally occurring nucleotides include, for example, ribonucleotides (RNA) or deoxyribonucleotides (DNA). Naturally occurring nucleotides may be in monophosphate, diphosphate, or triphosphate form and may have a base selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it will be understood that non-naturally occurring nucleotides, modified nucleotides, or analogs of the above nucleotides may be used. Some examples of useful non-naturally occurring nucleotides are described below with respect to reversible terminator-based sequencing by synthetic methods.
[0036] In embodiments that include a reaction chamber, an article or solid material (including semi-solid material) may be placed in the reaction chamber. When placed, the article or solid may be physically held or immobilized in the reaction chamber via interference fit, adhesion, or entrapment. Exemplary articles or solids that may be placed in the reaction chamber include polymer beads, pellets, agarose gels, powders, quantum dots, or other solids that may be compressed and / or held in the reaction chamber. In certain embodiments, nucleic acid superstructures such as DNA balls may be placed in or on the reaction chamber, for example, by attaching to the inner surface of the reaction chamber or by residing in a liquid in the reaction chamber. DNA balls or other nucleic acid superstructures may be preformed and then placed in or on the reaction chamber. Alternatively, DNA balls may be synthesized in situ in the reaction chamber. DNA balls may be synthesized by rolling circle amplification to generate concatemers of specific nucleic acid sequences, and the concatemers may be treated under conditions to form relatively compact balls. DNA balls and methods for their synthesis are described, for example, in U.S. Patent Application Publication Nos. 2008 / 0242560(A1) or 2008 / 0234136(A1), each of which is incorporated herein in its entirety. The material held or disposed within the reaction chamber can be in a solid, liquid, or gaseous state.
[0037] As used herein, a "base call" identifies a nucleotide base in a nucleic acid sequence. Base calling refers to the process of determining the base call (A, C, G, T) of every cluster in a particular cycle. As an example, base calling can be performed utilizing the four-channel, two-channel, or one-channel methods and systems described in the incorporated materials of US Patent Application Publication No. 2013 / 0079232. In certain embodiments, a base call cycle is referred to as a "sampling event." In a one-dye and two-channel sequencing protocol, a sampling event includes two illumination stages in chronological order such that a pixel signal occurs at each stage. The first illumination stage induces illumination from a given cluster that indicates nucleotide bases A and T in an AT pixel signal, and the second illumination stage induces illumination from a given cluster that indicates nucleotide bases C and T in a CT pixel signal.
[0038] The disclosed technology, for example, the disclosed base code, can be implemented on processors such as Central Processing Units (CPUs), Graphics Processing Units (GPUs), Field Programmable Gate Arrays (FPGAs), Coarse-Grained Reconfigurable Architectures (CGRAs), Application-Specific Integrated Circuits (ASICs), Application Specific Instruction-set Processors (ASIPs), and Digital Signal Processors (DSPs).
[0039] Biosensors FIG. 1 shows a cross-sectional view of a biosensor 100 that can be used in various embodiments. The biosensor 100 has pixel areas 106', 108', 110', 112', and 114', each of which can retain two or more clusters (e.g., two clusters per pixel area) during a base call cycle. As shown, the biosensor 100 can include a flow cell 102 mounted on a sampling device 104. In the illustrated embodiment, the flow cell 102 is fixed directly to the sampling device 104. However, in alternative embodiments, the flow cell 102 can be removably coupled to the sampling device 104. The sampling device 104 has a sample surface 134 that can be functionalized (e.g., chemically or physically modified in a manner suitable for causing a desired reaction). For example, the sample surface 134 may be functionalized and may include multiple pixel regions 106', 108', 110', 112', and 114' each capable of holding two or more clusters during a base calling cycle (e.g., having corresponding cluster pairs 106A, 106B, cluster pairs 108A, 108B, cluster pairs 110A, 110B, cluster pairs 112A, 112B, and cluster pairs 114A, 114B immobilized thereon). Each pixel region is associated with a corresponding sensor (or pixel or photodiode) 106, 108, 110, 112, and 114, such that light received by the pixel region is captured by the corresponding sensor. The pixel region 106' may also be associated with a corresponding reaction site 106" on the reaction surface 134 that holds the cluster pairs, such that light emitted from the reaction site 106" is received by the pixel region 106' and captured by the corresponding sensor 106. As a result of this sensing structure, if two or more clusters are present in a particular sensor pixel area during a base call cycle (e.g., each having a corresponding cluster pair), the pixel signal in that base call cycle carries information based on all of the two or more clusters.As a result, the signal processing described herein is used to distinguish between clusters where there are more clusters than pixel signals at a given sampling event of a particular base call cycle.
[0040] In the illustrated embodiment, the flow cell 102 includes sidewalls 138, 125 and a flow cover 136 supported by the sidewalls 138, 125. The sidewalls 138, 125 are coupled to the sample surface 134 and extend between the flow cover 136 and the sidewalls 138, 125. In some embodiments, the sidewalls 138, 125 are formed from a curable adhesive layer that bonds the flow cover 136 to the sampling device 104.
[0041] The side walls 138, 125 are sized and shaped such that a flow channel 144 exists between the flow cover 136 and the sampling device 104. The flow cover 136 may include a material that is transparent to excitation light 101 propagating from outside the biosensor 100 to the flow channel 144. In one example, the excitation light 101 approaches the flow cover 136 at a non-orthogonal angle.
[0042] As also shown, the flow cover 136 may include inlet and outlet ports 142, 146 configured to fluidly engage other ports (not shown). For example, these other ports may be from a cartridge or a workstation. The flow channel 144 is sized and shaped to direct fluid along the sample surface 134. The height H1 and other dimensions of the flow channel 144 may be configured to maintain a substantially uniform flow of fluid along the sample surface 134. The dimensions of the flow channel 144 may also be configured to control bubble formation.
[0043] By way of example, the flow cover 136 (or flow cell 102) may comprise a transparent material such as glass or plastic. The flow cover 136 may comprise a substantially rectangular block having a planar outer surface and a planar inner surface that defines the flow channel 144. The block may be attached onto the side walls 138, 125. Alternatively, the flow cell 102 may be etched to define the flow cover 136 and the side walls 138, 125. For example, a recess may be etched into the transparent material. When the etched material is attached to the sampling device 104, the recess may become the flow channel 144.
[0044] The sampling device 104 may be similar to an integrated circuit comprising, for example, multiple stacked substrate layers 120-126. The substrate layers 120-126 may include a base substrate 120, a solid-state imager 122 (e.g., a CMOS image sensor), a filter or light management layer 124, and a passivation layer 126. Note that the above is merely exemplary and other embodiments may include fewer or additional layers. Additionally, each of the substrate layers 120-126 may include multiple sublayers. The sampling device 104 may be fabricated using processes similar to those used in fabricating integrated circuits such as CMOS image sensors and CCDs. For example, the substrate layers 120-126 or portions thereof may be grown, deposited, etched, etc. to form the sampling device 104.
[0045] The passivation layer 126 is configured to shield the filter layer 124 from the fluid environment of the flow channel 144. In some cases, the passivation layer 126 is also configured to provide a solid surface (i.e., the sample surface 134) that allows for biomolecules or other analytes of interest to be immobilized thereon. For example, each of the reaction sites may include a cluster of biomolecules immobilized on the sample surface 134. Thus, the passivation layer 126 may be formed from a material that allows for the reaction sites to be immobilized thereon. The passivation layer 126 may also include a material that is at least transparent to the desired fluorescence. By way of example, the passivation layer 126 may include silicon nitride (Si2N4) and / or silica (SiO2). However, other suitable materials may be used. In the illustrated embodiment, the passivation layer 126 may be substantially planar. However, in alternative embodiments, the passivation layer 126 may include recesses, such as pits, wells, grooves, and the like. In the illustrated embodiment, the passivation layer 126 has a thickness of about 150-200 nm, and more specifically, about 170 nm.
[0046] The filter layer 124 may include various features that affect the transmission of light. In some embodiments, the filter layer 124 may perform multiple functions. For example, the filter layer 124 may be configured to (a) filter unwanted light signals, such as light signals from an excitation light source, (b) direct luminescence signals from the reaction sites toward corresponding sensors 106, 108, 110, 112, and 114 configured to detect luminescence signals from the reaction sites, or (c) block or prevent detection of unwanted luminescence signals from adjacent reaction sites. Thus, the filter layer 124 may also be referred to as a light management layer. In the illustrated embodiment, the filter layer 124 has a thickness of about 1-5 μm, more specifically about 2-4 μm. In alternative embodiments, the filter layer 124 may include an array of microlenses or other optical components. Each of the microlenses may be configured to direct the luminescence signal from an associated reaction site to a sensor.
[0047] In some embodiments, the solid-state imager 122 and the base substrate 120 may be provided together as a previously constructed solid-state imaging device (e.g., a CMOS chip). For example, the base substrate 120 may be a wafer of silicon, and the solid-state imager 122 may be mounted thereon. The solid-state imager 122 includes a layer of semiconductor material (e.g., silicon) and the sensors 106, 108, 110, 112, and 114. In the illustrated embodiment, the sensors are photodiodes configured to detect light. In other embodiments, the sensors include photodetectors. The solid-state imager 122 may be fabricated as a single chip via a CMOS-based manufacturing process.
[0048] The solid-state imager 122 may include a high density array of sensors 106, 108, 110, 112, and 114 configured to detect activity indicative of a desired response from within or along the flow channel 144. In some embodiments, each sensor has a size of about 1-2 micrometers squared (μm 2 ) The array may include 500,000 sensors, 5 million sensors, 10 million sensors, or even 120 million sensors. Sensors 106, 108, 110, 112, and 114 may be configured to detect predetermined wavelengths of light that are indicative of a desired response.
[0049] In some embodiments, the sampling device 104 includes a microcircuit arrangement, such as that described in U.S. Patent No. 7,595,882, which is incorporated herein by reference in its entirety. More specifically, the sampling device 104 may include an integrated circuit having a planar array of sensors 106, 108, 110, 112, and 114. The circuitry formed within the sampling device 104 may be configured for at least one of signal amplification, digitization, storage, and processing. The circuitry may collect and analyze the detected fluorescence and generate a pixel signal (or detection signal) for communicating the detection data to a signal processor. The circuitry may also perform additional analog and / or digital signal processing in the sampling device 104. The sampling device 104 may include conductive vias 130 that perform signal routing (e.g., transmit the pixel signal to a signal processor). The pixel signal may also be transmitted through electrical contacts 132 of the sampling device 104.
[0050] The sampling device 104 is discussed in more detail with respect to U.S. Non-Provisional Patent Application No. 16 / 874,599, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 14, 2020 (Attorney Docket No. ILLM1011-4 / IP-1750-US), which is incorporated by reference as if fully set forth herein. The sampling device 104 is not limited to the above configurations or uses as described above. In alternative embodiments, the sampling device 104 may take other forms. For example, the sampling device 104 may comprise a CCD device, such as a CCD camera, coupled to a flow cell or moved to interface with a flow cell having reaction sites therein.
[0051] Figure 2 shows one implementation of a flow cell 200 that includes clusters within its tiles. The flow cell 200 corresponds to the flow cell 102 of Figure 1, e.g., without the flow cover 136. Additionally, the depiction of the flow cell 200 is symbolic in nature, and the flow cell 200 symbolically shows the various lanes and tiles therein without showing the various other components therein. Figure 2 shows a top view of the flow cell 200.
[0052] In one embodiment, flow cell 200 is divided or split into multiple lanes, such as lanes 202a, 202b, ..., 202P, i.e., P lanes. In the example of Figure 2, flow cell 200 is shown as including eight lanes, i.e., P=8 in this example, although the number of lanes in a flow cell is implementation specific.
[0053] In one embodiment, each lane 202 is further divided into non-overlapping regions called "tiles" 212. For example, Figure 2 shows an expanded view of a section 208 of an exemplary lane. The section 208 is shown to include multiple tiles 212.
[0054] In an embodiment, each lane 202 includes one or more tile columns. For example, in FIG. 2, each lane 202 includes two corresponding tile columns 212, as shown in enlarged section 208. The number of tiles in each tile column in each lane is implementation specific, and in one example, there may be 50 tiles, 60 tiles, 100 tiles, or another suitable number of tiles in each tile column in each lane.
[0055] Each tile contains a corresponding number of clusters. During the sequencing procedure, the clusters on the tile and their surrounding background are imaged. For example, Figure 2 shows an example cluster 216 in an example tile.
[0056] FIG. 3 shows an exemplary Illumina GA-IIx™ flow cell with eight lanes, and also shows a zoom-in of one tile and its clusters and their surrounding background. For example, there are 100 tiles per lane in the Illumina Genome Analyzer II, and 68 tiles per lane in the Illumina HiSeq2000. The tile 212 holds hundreds of thousands to millions of clusters. In FIG. 3, an image generated from a tile with clusters shown as bright spots is shown at 308 (e.g., 308 is a magnified image view of a tile), and an exemplary cluster 304 is labeled. The cluster 304 contains about a thousand identical copies of the template molecule, but the clusters differ in size and shape. The clusters are grown from the template molecules by bridge amplification of the input library prior to a sequencing run. The purpose of the amplification and cluster growth is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense single fluorophores. However, the physical distance between the DNA fragments within the cluster 304 is small, so the imaging device perceives the cluster of fragments as a single spot 304 .
[0057] Clusters and tiles are discussed in further detail in U.S. Non-Provisional Patent Application No. 16 / 825,987, entitled "TRAINING DATA GENERATION FOR ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 20, 2020 (Attorney Docket No. ILLM1008-16 / IP-1693-US).
[0058] FIG. 4 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base calling sensor output (see, e.g., FIG. 1). In the example of FIG. 4, the system includes a sequencing machine 400 and a configurable processor 450. The configurable processor 450 can execute a neural network based base caller in coordination with a runtime program executed by a host processor, such as a central processing unit (CPU) 402. The sequencing machine 400 includes a base call sensor (e.g., as discussed with respect to FIGS. 1-3) and a flow cell 401. The flow cell can include one or more tiles in which clusters of genetic material are exposed to a sequence of analyte flows that are used to trigger reactions in the clusters to identify bases in the genetic material, as discussed with respect to FIGS. 1-3. The sensor senses the reaction of each cycle of the sequence in each tile of the flow cell to provide tile data. An example of this technique is described in more detail below. Genetic sequencing is a data-intensive operation that converts the base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.
[0059] The system in this example includes a CPU 402 that executes a runtime program that coordinates the base calling operation, and memory 403 that stores the sequence of the array of tile data, the base call reads generated by the base calling operation, and other information used in the base calling operation. In this figure, the system also includes a configuration file (or files), e.g., an FPGA bit file, and memory 404 that stores model parameters of a neural network used to configure and reconfigure configurable processor 450 and to run the neural network. Sequencing machine 400 can include programs for configuring the configurable processor, and in some embodiments can include a reconfigurable processor that runs the neural network.
[0060] The sequencing machine 400 is coupled to the configurable processor 450 by a bus 405. The bus 405 may be implemented using a high throughput technology, such as a bus technology compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by the PCI Special Interest Group (PCI-SIG). Also, in this embodiment, a memory 460 is coupled to the configurable processor 450 by a bus 461. The memory 460 may be an on-board memory located on a circuit board having the configurable processor 450. The memory 460 is used for fast access by the configurable processor 450 of working data used in base calling operations. The bus 461 may also be implemented using a high throughput technology, such as a bus technology compatible with the PCIe standard.
[0061] Configurable processors, including Field Programmable Gate Arrays (FPGAs), Coarse Grained Reconfigurable Arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using a general-purpose processor executing a computer program. Configuring a configurable processor involves compiling a functional description to generate a configuration file, sometimes called a bitstream or bitfile, and distributing the configuration file to the configurable elements on the processor.
[0062] The configuration file configures the circuit to set data flow patterns, including the use of distributed memory and other on-chip memory resources, lookup table contents, the operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of the configurable array. A configuration file is reconfigurable if it can be changed in the field by modifying a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, in a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for use in base calling operations as described herein. Examples include commercially available products such as the Xilinx Alveo™ U200, Xilinx Alveo™ U250, Xilinx Alveo™ U280, Intel / Altera Stratix™ GX2800, Intel / Altera Stratix™ GX2800, and Intel Stratix™ GX10M. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.
[0063] The embodiments described herein use a configurable processor 450 to implement a multi-cycle neural network. The configuration file of the configurable processor may be implemented by specifying the logic functions to be performed using a high-level description language (HDL) or register transfer level (RTL) language specification. The specification may be compiled using resources designed by a selected configurable processor to generate a configuration file. The same or similar specifications may be compiled to generate a design for an application specific integrated circuit, which may not be a configurable processor.
[0064] Thus, alternatives to the configurable processor in all of the embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or a system-on-chip SOC device, configured to perform the neural network based base calling operations described herein.
[0065] In general, the configurable and configured processors described herein that are configured to perform neural network operations are referred to herein as neural network processors.
[0066] Configurable processor 450 is configured, in this embodiment, by a configuration file loaded using a program executed by CPU 402, or by other sources that configures an array of configurable elements on configurable processor 454 to perform the base calling function. In this embodiment, the configuration includes data flow logic 451 coupled to buses 405 and 461, which performs the function of distributing data and control parameters among the elements used in the base calling operation.
[0067] Configurable processor 450 is also configured with base call execution logic 452 to execute the multi-cycle neural network. Logic 452 includes a number of multi-cycle execution clusters (e.g., 453), which in this example include multi-cycle cluster 1 through multi-cycle cluster X. The number of multi-cycle clusters can be selected according to tradeoffs involving the desired throughput of operation and available resources on the configurable processor.
[0068] The multi-cycle clusters are coupled to data flow logic 451 by data flow paths 454, implemented using configurable interconnect and memory resources on a configurable processor, and by control paths 455, implemented using, for example, configurable interconnect and memory resources on a configurable processor, that provide control signals indicating available clusters, readiness to provide input units to available clusters for execution of neural network operations, readiness to provide trained parameters for the neural network, readiness to provide output patches of base call classification data, and other control data used in the execution of the neural network.
[0069] The configurable processor is configured to use the trained parameters to perform a multi-cycle neural network operation to generate classification data for a sensing cycle of the base calling operation. The neural network operation is performed to generate classification data for a subject sensing cycle of the base calling operation. The neural network operation operates on an array including a number N of arrays of tile data from each sensing cycle of the N sensing cycles, which in the examples described herein provide sensor data for different base calling operations for one base position per operation in the time series. Optionally, some of the N sensing cycles can be out of the array as needed according to the particular neural network model being performed. The number N can be any number greater than 1. In some examples described herein, the sensing cycle of the N sensing cycles represents a set of sensing cycles for at least one sensing cycle preceding the subject sensing cycle and at least one sensing cycle following the subject sensing cycle in the time series. Examples described herein are in which the number N is an integer greater than or equal to 5.
[0070] The data flow logic 451 is configured to use an input unit for a given operation that includes tile data of an array of N spatially aligned patches to move the tile data and at least some trained parameters of the model from the memory 460 to the configurable processor for operation of the neural network. The input unit can be moved by direct memory access operations in a single DMA operation, or in smaller units that move during available time slots in coordination with the execution of the deployed neural network.
[0071] The tile data of the sensing cycle described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and the sensor. For example, in an embodiment of a base calling operation, the tile data can include information about the alignment of the images with the clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.
[0072] During execution of the multi-cycle neural network as described below, the tile data may also include data generated during execution of the multi-cycle neural network, referred to as intermediate data, which may be reused rather than recomputed during execution of the multi-cycle neural network. For example, during execution of the multi-cycle neural network, the data flow logic may write the intermediate data to memory 460 in place of the sensor data for a given patch of the array of tile data. Such embodiments are described in more detail below.
[0073] As shown, a system for analysis of base calling sensor output is described that includes a memory (e.g., 460) accessible by a runtime program that stores tile data including sensor data for tiles from sensing cycles of a base calling operation. The system also includes a neural network processor, such as a configurable processor 450, having access to the memory. The neural network processor is configured to perform operations of the neural network using trained parameters to generate classification data for the sensing cycles. As described herein, the operations of the neural network operate on an arrangement of N arrays of tile data from each sensing cycle of the N sensing cycles that comprise a subject cycle to generate classification data for the subject cycle. Data flow logic 451 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network using input units including data for spatially aligned patches of the N arrays from each sensing cycle of the N sensing cycles.
[0074] Also described is a system in which a neural network processor has access to a memory and includes a plurality of execution clusters, and an execution logic cluster in the plurality of execution clusters is configured to execute a neural network. The data flow logic includes access to the memory and executes a cluster in the plurality of execution clusters to provide an input unit of tile data to an available execution cluster in the plurality of execution clusters, the input unit including an input unit including a number N of spatially aligned patches of the array of tile data from a respective sensing cycle, and causing the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1.
[0075] FIG. 5 is a simplified diagram illustrating aspects of a base calling operation, including functions of a runtime program executed by a host processor. In this diagram, the output of an image sensor from a flow cell (such as that shown in FIGS. 1 and 2) is provided on line 500 to an image processing thread 501, which can perform processes on the image such as resampling, alignment and positioning of the array of sensor data for individual tiles, which can be used by a process to calculate a tile cluster mask for each tile in the flow cell, which can be used by a process to identify pixels in the array of sensor data that correspond to clusters of genetic material on the corresponding tile of the flow cell. To calculate the cluster mask, one exemplary algorithm is based on a process that detects unreliable clusters in early sequencing cycles using a metric derived from the softmax output, and then data from those wells / clusters is discarded and no output data is generated for those clusters. For example, the process can identify clusters that are highly reliable during the first N1 (e.g., 25) base calls and reject other clusters. Rejected clusters can be polyclonal or very weak intensity or unclear according to the criteria. This procedure can be executed on the host CPU. Alternative implementations could potentially use this information to identify necessary clusters that are to be returned to the CPU, thereby limiting the storage required for intermediate data.
[0076] The output of image processing thread 501 is provided on line 502 to dispatch logic 510 in the CPU, which routes the array of tile data according to the status of the base calling operation to either a data cache 504 on high speed bus 503 or to a multi-cluster neural network processor hardware 520, such as the configurable processor of FIG. 4, on high speed bus 505. The hardware 520 returns the classification data output by the neural network to the dispatch logic 510, which passes the information to the data cache 504 or on line 511 to thread 502, which can use the classification data to perform base calling and quality score calculations and place the data in a standard format for base called reads. The output of thread 502, which performs base calling and quality score calculations, is provided on line 512 to thread 503, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base calling output to a specified destination for consumption by the customer.
[0077] In some embodiments, the host may include a thread (not shown) that performs final processing of the output of the hardware 520 supporting the neural network. For example, the hardware 520 may provide an output of classification data from a final layer of a multi-cluster neural network. The host processor may perform output activation functions, such as a softmax function, over the classification data to populate the data used by the base calling and quality score thread 502. The host processor may also perform input operations (not shown), such as resampling, batch normalization, or other adjustments of the tile data before inputting it to the hardware 520.
[0078] FIG. 6 is a simplified diagram of a configuration of a configurable processor, such as the configurable processor of FIG. 4. In FIG. 6, the configurable processor comprises an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 600 including the data flow logic described with reference to FIG. 1. The wrapper 600 manages interfacing and coordinating with a runtime program in the CPU via a CPU communication link 609, and manages communication with an on-board DRAM 602 (e.g., memory 460) via a DRAM communication link 610. The data flow logic in the wrapper 600 provides patch data obtained by traversing an array of tile data on the on-board DRAM 602 to a cluster 601 for a number N of cycles, and obtains process data 615 from the cluster 601 and delivers it to the on-board DRAM 602. The wrapper 600 also manages the transfer of data between the on-board DRAM 602 and the host memory for both the input array of tile data and the output patch of classification data. The wrapper transfers the patch data on line 613 to the assigned cluster 601. The wrapper provides the cluster 601 with trained parameters such as weights and biases obtained from on-board DRAM 602 on line 612. The wrapper provides the cluster 601 with configuration and control data provided from or generated in response to a runtime program on the host over CPU communication link 609 on line 611. The cluster can also provide state signals to the wrapper 600 on line 616 that are used in conjunction with control signals from the host to manage the traversal of an array of tile data to provide spatially aligned patch data and to run a multi-cycle neural network on the patch data using the resources of the cluster 601.
[0079] As described above, there may be multiple clusters on a single configurable processor managed by wrapper 600 configured to run on corresponding ones of the multiple patches of tile data. Each cluster may be configured to provide classification data for base calls in a subject sensing cycle using the tile data of multiple sensing cycles as described herein.
[0080] In an example system, model data including kernel data such as filter weights and biases can be sent from the host CPU to the configurable processor, so that the model can be updated as a function of cycle number. The base calling operation can include, in a representative example, on the order of hundreds of sensing cycles. The base calling operation can include paired end reading in some embodiments. For example, the model trained parameters may be updated every 20 cycles (or other number of cycles) or according to an update pattern implemented in the particular system and neural network model. In some embodiments where the sequence for a given string in a genetic cluster on a tile includes a first portion extending down (or up) from a first end of the string and a second portion extending up (or down) from a second end of the string, the trained parameters can be updated at the transition from the first portion to the second portion.
[0081] In some implementations, image data for multiple cycles of sensor data for a tile can be sent from the CPU to the wrapper 600. The wrapper 600 can optionally perform some pre-processing and conversion of the sensor data and write the information to the on-board DRAM 602. The input tile data for each sensing cycle can include an array of sensor data including 4000x3000 pixels or more per sensing cycle per tile, with two features representing the colors of the two images of the tile, and including one or two bytes per pixel. In an embodiment where the number N is three sensing cycles used in each operation of the multi-cycle neural network, the array of tile data for each operation of the multi-cycle neural network can consume in the hundreds of megabytes per tile. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and the tile.
[0082] In operation, if a multi-cycle cluster is available, the wrapper assigns the patch to the cluster. The wrapper fetches the next patch of tile data for the cross section of the tile and sends it to the assigned cluster along with the appropriate control and configuration information. The cluster can be configured with enough memory on the configurable processor to have enough memory to hold the patch of data, including the patch, from multiple cycles in some systems being processed in place, and in various embodiments are processed using ping-pong buffer or raster scan techniques.
[0083] When the assigned cluster completes its operation of the neural network of the current patch and generates an output patch, it signals the wrapper. The wrapper will either read the output patch from the assigned cluster or the assigned cluster will push the data to the wrapper. The wrapper will then assemble the output patch for the processed tile in DRAM 602. Once the processing of the entire tile is complete and the output patch of data is transferred to DRAM, the wrapper will send the processed output array back to the host / CPU in a specific format. In some embodiments, the on-board DRAM 602 is managed by memory management logic in the wrapper 600. The runtime program can control the sequencing operations to complete the analysis of the array of all tile data for every cycle running in a continuous flow to provide real-time analysis.
[0084] FIG. 7 is a diagram of a multi-cycle neural network model that can be implemented using the system described herein. The example shown in FIG. 7 can be referred to as a 5-cycle input, 1-cycle output neural network. The input to the multi-cycle neural network model includes five spatially aligned patches (e.g., 700) from the tile data array of five sensing cycles for a given tile. The spatially aligned patches have the same aligned row and column dimensions (x, y) as other patches in the set, so that the information relates to the same cluster of genetic material on the tile in the sequence cycle. In this example, the subject patch is a patch from the array of tile data for cycle K. The set of five spatially aligned patches includes a patch from cycle K-2 that precedes the subject patch by two cycles, a patch from cycle K-1 that precedes the subject patch by one cycle, a patch from cycle K+1 that follows the patch from the subject cycle by one cycle, and a patch from cycle K+2 that follows the patch from the subject cycle by two cycles.
[0085] The model includes a separate stack 701 of layers of a neural network for each of the input patches. Thus, stack 701 receives as input patch tile data from cycle K+2 and is separate from stacks 702, 703, 704, and 705 such that they do not share input data or intermediate data. In some embodiments, all of stacks 710-705 can have the same model and the same trained parameters. In other embodiments, the models and trained parameters can be different in different stacks. Stack 702 receives as input patch tile data from cycle K+1. Stack 703 receives as input patch tile data from cycle K. Stack 704 receives as input patch tile data from cycle K-1. Stack 705 receives as input patch tile data from cycle K-2. Each layer of the separate stacks performs a convolution operation of a kernel that includes multiple filters over the input data of the layer. As in the above example, patch 700 may include three features. The output of layer 710 may include many more features, such as 10-20 features. Similarly, the output of each of layers 711-716 may include any number of features suitable for a particular implementation. The parameters of the filters are the trained parameters of the neural network, such as weights and biases. The output feature sets (intermediate data) from each of stacks 701-705 are provided as inputs to an inverse layer 720 of temporal combination layers, where the intermediate data from multiple cycles are combined. In the illustrated example, the inverse layer 720 includes a first layer including three combination layers 721, 722, 723 that respectively receive intermediate data from three of the separated stacks, and a final layer including one combination layer 730 that receives intermediate data from the three temporal layers 721, 722, 723.
[0086] The output of the final combination layer 730 is an output patch of classification data for the clusters located in the corresponding patch of the tile from cycle K. The output patches can be assembled into an output array classification data for the tile in cycle K. In some embodiments, the output patch can have a different size and dimensions than the input patch. In some embodiments, the output patch can include per-pixel data that can be filtered by the host to select cluster data.
[0087] The output classification data can then be applied to a softmax function 740 (or other output-driven function) optionally executed by the host or on a configurable processor, depending on the particular implementation. An output function different from softmax can be used (e.g., create base call output parameters according to the maximum output, and then use a non-linear mapping learned using the context / network output to give the base quality).
[0088] Finally, the output of the softmax function 740 is provided as the base call probability for cycle K (750) and may be stored in host memory for use in subsequent processing. Other systems may use different functions, e.g., different non-linear models, for the output probability calculation.
[0089] The neural network can be implemented using a configurable processor with multiple execution clusters to complete the evaluation of one tile cycle within or near the duration of the time interval of one sensing cycle, effectively outputting output data in real time. Dataflow logic can be configured to distribute input units of tile data and trained parameters to the execution clusters, and to distribute output patches for aggregation in memory.
[0090] The input unit of data for a 5-cycle input, 1-cycle output neural network similar to that of FIG. 7 is described with reference to FIG. 8A and FIG. 8B for a base calling operation using 2-channel sensor data. For example, for a given base in a genetic sequence, the base calling operation can perform two flows of analyte and two reactions, which generate two channels of signals, such as images, that can be processed to identify which one of the four bases is located at the current position of the genetic sequence for each cluster of genetic material. In other systems, a different number of channels of sensor data can be utilized. For example, base calling can be performed using a 1-channel method and system. The incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232 discuss base calling using various numbers of channels, such as 1-channel, 2-channel, or 4-channel.
[0091] 8A shows an array of 5 cycles of tile data for a given tile, Tile M, used to implement a 5 cycle input, 1 cycle output neural network. The 5 cycle input tile data in this example is written to on-board DRAM or other memory in the system that can be accessed by the dataflow logic and includes channel 1 array 801 and channel 2 array 811 for cycle K-2, channel 1 array 802 and channel 2 array 812 for cycle K-1, channel 1 array 803 and channel 2 array 813 for cycle K, channel 1 array 804 and channel 2 array 814 for cycle K+1, and channel 1 array 805 and channel 2 array 815 for cycle K+2. An array of tile metadata 820 can also be written once to memory, in this case with the DFC file included for use as input to the neural network with each cycle.
[0092] Although Figure 8A discusses a two-channel base calling operation, the use of two channels is merely an example, and base calling can be performed using any other suitable number of channels. For example, the incorporated materials of US Patent Application Publication No. 2013 / 0079232 discuss base calling using various numbers of channels, such as one channel, two channels, or four channels, or another suitable number of channels.
[0093] The dataflow logic configures an input unit, which may be understood with reference to FIG. 8B, of tile data including spatially aligned patches of arrays of tile data for each execution cluster configured to perform a neural network execution on the input patches. The input unit of an assigned execution cluster is configured by the dataflow logic to read spatially aligned patches (e.g., 851, 852, 861, 862, 870) from each of the arrays of tile data 801-805, 811, 815, 820 for five input cycles and deliver them via a datapath (schematically 850) to memory on a configurable processor configured for use by the assigned execution cluster. The assigned execution cluster performs a 5 cycle input / 1 cycle output neural network execution and delivers subject cycle K output patches of classification data for the same patch of tiles for subject cycle K.
[0094] Figure 9 is a simplified representation of a stack of neural networks that can be used in a system such as that of Figure 7 (e.g., 701 and 720). In this example, some functions of the neural network (e.g., 900, 902) run on the host and other parts of the neural network (e.g., 901) run on a configurable processor.
[0095] In one example, the first function may be batch normalization (layer 910) formed on the CPU, however, in another example, batch normalization as a function may be blended into one or more layers, and there may not be a separate batch normalization layer.
[0096] Several spatially separated convolutional layers are implemented as a first set of convolutional layers of a neural network as discussed above for the configurable processor. In this example, the first set of convolutional layers applies spatially 2D convolutions.
[0097] As shown in Figure 9, for a number L / 2 of spatially separated neural network layers in each stack (where L was described with reference to Figure 7), a first spatial convolution 921 is performed, followed by a second spatial convolution 922, followed by a third spatial convolution 923, etc. As shown in 923A, the number of spatial layers can be any practical number, which in context can range from a few to more than 20 in different embodiments.
[0098] For SP_CONV_0, the kernel weights are stored in, for example, a (1, 6, 6, 3, L) structure since this layer has 3 input channels. In this example, the "6" in this structure comes from storing the coefficients in the transformed Winograd domain (kernel size is 3x3 in the spatial domain, but expands in the transform domain).
[0099] For the other SP_CONV layers, the kernel weights are stored in a (1, 6, 6L) structure in this embodiment since there are K(=L) inputs and outputs for each of these layers.
[0100] The output of the stack of spatial layers is provided to a temporal layer, including convolution layers 924, 925 running on an FPGA. Layers 924 and 925 may be convolution layers that apply 1D convolution over cycles. As shown in 924A, the number of temporal layers may be any practical number, which in context may range from a few to more than 20 in different embodiments.
[0101] The first temporal layer, TEMP_CONV_0 layer 824, reduces the number of cycle channels from 5 to 3, as shown in Figure 7. The second temporal layer, layer 925, reduces the number of cycle channels from 3 to 1, as shown in Figure 7, reducing the number of feature maps to 4 outputs per pixel representing the confidence of each base call.
[0102] The outputs of the temporal layers are accumulated in output patches and delivered to the host CPU, where a softmax function 930, or other function, is applied to normalize the base call probabilities.
[0103] Figure 10 shows an alternative implementation illustrating a 10-input, 6-output neural network that can be implemented for base calling operations. In this example, the spatially aligned input patch tile data for cycles 0-9 are applied to separated stacks in the spatial layer, such as stack 1001 for cycle 9. The outputs of the separated stacks are applied to the inverted hierarchical arrangement of the time stack 1020, and outputs 1035(2)-1035(7) provide base call classification data for subject cycles 2-7.
[0104] Figure 11 shows one implementation of the dedicated architecture (e.g., Figure 7) of a neural network-based base caller used to separate the processing of data in different sequencing cycles. The motivation for using the dedicated architecture above is first explained.
[0105] The neural network-based base caller processes data from a current sequencing cycle, one or more preceding sequencing cycles, and one or more subsequent sequencing cycles. Data from additional sequencing cycles provides sequence-specific context. The neural network-based base caller learns sequence-specific contexts during training and base calls them. Additionally, data from pre- and post-sequencing cycles provide secondary contributions of pre-phasing and phasing signals to the current sequencing cycle.
[0106] Images captured in different sequencing cycles and in different image channels are misaligned and have residual registration errors with each other. To account for this misalignment, the dedicated architecture includes a spatial convolution layer that does not mix information between sequencing cycles, but only mixes information within the same sequencing cycle.
[0107] The spatial convolutional layer uses so-called "decoupled convolutions" that operate on separation by processing the data for each of multiple sequencing cycles independently through a "dedicated, non-shared" array of convolutions. Decoupled convolutions convolve on the data and resulting feature maps only within a given sequencing cycle, i.e., the cycle, without convolving on the data and resulting feature maps of any other sequencing cycles.
[0108] For example, consider that the input data includes (i) current data for the current (time t) sequencing cycle to be base called, (ii) previous data for the previous (time t-1) sequencing cycle, and (iii) next data for the next (time t+1) sequencing cycle. The dedicated architecture then starts three separate data processing pipelines (or convolution pipelines), namely, the current data processing pipeline, the previous data processing pipeline, and the next data processing pipeline. The current data processing pipeline receives the current data for the current (time t) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "current spatial convolution representation" as the output of the final spatial convolution layer. The previous data processing pipeline receives the previous data for the previous (time t-1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to generate a so-called "previous spatial convolution representation" as the output of the final spatial convolution layer. The next data processing pipeline receives the next data for the next (time t+1) sequencing cycle as input and processes it independently through multiple spatial convolution layers to produce a so-called “next spatially convolved representation” as the output of the final spatial convolution layer.
[0109] In some implementations, the current pipeline, the one or more previous pipelines, and the one or more next processing pipelines execute in parallel.
[0110] In some implementations, the spatial convolutional layer is part of a spatial convolutional network (or sub-network) within a dedicated architecture.
[0111] The neural network-based base caller further includes temporal convolutional layers that blend information between sequencing cycles, i.e., between cycles. The temporal convolutional layers receive their input from the spatial convolutional network and operate on the spatially convolved representations produced by the final spatial convolutional layer for each data processing pipeline.
[0112] The inter-cycle operational freedom of the temporal convolutional layers arises from the fact that misalignment features present in the image data provided as input to the spatial convolutional network are purged from the spatial convolutional representation by the stack or cascade of separated convolutions performed by the array of spatial convolutional layers.
[0113] The temporal convolutional layer uses so-called "combinatorial convolution" that convolves group-wise on the input channels with the subsequent input on a sliding window basis. In one implementation, the subsequent input is the subsequent output generated by the previous spatial convolutional layer or the previous temporal convolutional layer.
[0114] In some implementations, the temporal convolutional layer is part of a temporal convolutional network (or sub-network) in a dedicated architecture. The temporal convolutional network receives its input from a spatial convolutional network. In one implementation, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles by group. In another implementation, the subsequent temporal convolutional layers of the temporal convolutional network combine successive outputs of previous temporal convolutional layers.
[0115] The output of the final temporal convolutional layer is fed into an output layer that produces outputs that are used to base call one or more clusters in one or more sequencing cycles.
[0116] During forward propagation, the dedicated architecture processes information from multiple inputs in two stages. In the first stage, decoupled convolutions are used to prevent mixing of information between inputs. In the second stage, combined convolutions are used to mix information between inputs. The results from the second stage are used to make a single inference on the multiple inputs.
[0117] This differs from batch-mode techniques, where the convolutional layer processes multiple inputs in a batch simultaneously and makes a corresponding inference for each input in the batch. In contrast, dedicated architectures map multiple inputs to a single inference. A single inference may include two or more predictions, such as a classification score for each of the four bases (A, C, T, and G).
[0118] In one implementation, the inputs have a temporal ordering such that each input occurs at a different time step and has multiple input channels. For example, the multiple inputs may include three inputs: a current input generated by a current sequencing cycle at time step (t), a previous input generated by a previous sequencing cycle at time step (t-1), and a next input generated by a next sequencing cycle at time step (t+1). In another implementation, each input is derived from the current, previous, and next inputs by one or more previous convolutional layers, respectively, and includes k feature maps.
[0119] In one embodiment, each input may include five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another implementation, each input may include k feature maps generated by a previous convolutional layer, and each feature map is treated as an input channel. In yet another example, each input may have just one channel, two channels, or another different number of channels. The incorporated materials in US Patent Application Publication No. 2013 / 0079232 discuss base calling using various numbers of channels, such as one channel, two channels, or four channels.
[0120] FIG. 12 illustrates one implementation of separated layers, each of which may include a convolution. Separate convolution processes multiple inputs at once by applying a convolution filter to each input in parallel. In separated convolution, a convolution filter combines input channels within the same input and does not combine input channels within different inputs. In one implementation, the same convolution filter is applied to each input in parallel. In another implementation, a different convolution filter is applied to each input in parallel. In some implementations, each spatial convolution layer includes a bank of k convolution filters, each of which is applied to each input in parallel.
[0121] FIG. 13A shows one implementation of a combination layer, each of which may include a convolution. FIG. 13B shows another implementation of a combination layer, each of which may include a convolution. A combination convolution mixes information between different inputs by grouping corresponding input channels of the different inputs and applying a convolution filter to each group. The grouping of corresponding input channels and the application of the convolution filter occur on a sliding window basis. In this context, a window spans two or more consecutive input channels, for example representing the output for two consecutive sequencing cycles. Because the window is a sliding window, most input channels are used in two or more windows.
[0122] In some implementations, the distinct inputs come from an output array generated by a preceding spatial or temporal convolutional layer, where the distinct inputs are arranged as successive outputs and are therefore viewed by the next temporal convolutional layer as successive inputs, where a combinatorial convolution then applies a convolutional filter to groups of corresponding input channels in the successive inputs.
[0123] In one implementation, the successive inputs have a temporal ordering such that the current input is generated by a current sequencing cycle at time step (t), the previous input is generated by a previous sequencing cycle at time step (t-1), and the next input is generated by a next sequencing cycle at time step (t+1). In another implementation, each successive input is derived from the current, previous, and next inputs by one or more previous convolutional layers, respectively, and includes k feature maps.
[0124] In one embodiment, each input may include five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another implementation, each input may include k feature maps generated by a previous convolutional layer, where each feature map is treated as an input channel.
[0125] The depth B of the convolution filter depends on the number of consecutive inputs whose corresponding input channels are convolved with the convolution filter on a sliding window basis for each group. In other words, the depth B is equal to the number of consecutive inputs in each sliding window and the group size.
[0126] In Figure 13A, corresponding input channels from two consecutive inputs are combined in each sliding window, so B = 2. In Figure 13B, corresponding input channels from three consecutive inputs are combined in each sliding window, so B = 3.
[0127] In one implementation, the sliding windows share the same convolution filter. In another implementation, a different convolution filter is used for each sliding window. In some implementations, each temporal convolution layer includes a bank of k convolution filters, each of which is applied to successive inputs on a sliding window basis.
[0128] Further details of Figures 4-10 and variations thereof can be found in co-pending U.S. non-provisional patent application Ser. No. 17 / 176,147, entitled "HARDWARE EXECUTION AND ACCELERATION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed Feb. 15, 2021 (Attorney Docket No. ILLM1020-2 / IP-1866-US), which is incorporated by reference as if fully set forth herein.
[0129] Training Basecoa from Scratch The base calling system is trained to predict base calls for an unknown analyte that includes a base sequence. For example, the base calling system may have a base caller that includes a neural network that predicts base calls for bases in the unknown analyte.
[0130] It is difficult to train the neural network of a base calling system. This is especially true when there is no labeled training data to be used to train the base calling system. In some embodiments, a Real Time Analysis (RTA) system can be used to generate labeled training data, which can be used to train the base calling system. An example of an RTA system is discussed in U.S. Patent No. US10304189(B2), entitled "Data processing system and methods," issued May 28, 2019, which is incorporated by reference as if fully set forth herein. However, if the system lacks RTA or cannot fully utilize the functionality of RTA, it may be difficult to initially generate labeled training data for training the neural network of the base calling system.
[0131] This disclosure discusses a self-learning base chorus that generates initial labeled training data, trains itself using the labeled training data, generates further labeled training data using the at least partially trained base chorus, trains itself using the further labeled labeled training data, generates further, more labeled training data, and iteratively repeats this process to fully train the base chorus. This iterative training and labeled training data generation process includes different stages such as a single oligo stage, a multiple oligo stage (2 oligo stage, 3 oligo stage, etc.), followed by a single organism stage, a complex organism stage, a further complex organism stage, etc. Thus, the complexity and / or length of the specimens used to train and generate the labeled training data increases progressively and monotonically with each iteration, along with the complexity of the neural network configuration underlying the base chorus, as discussed in turn in more detail herein. Because the base chorus is progressively self-trained, such a system obviates the need for the use of an RTA to generate labeled training data. Thus, while the basecalling systems described herein may include RTAs, the iterative training processes discussed herein can be used in addition to or instead of RTAs to train the basecallers.
[0132] FIG. 14A shows a base-calling system 1400 operating in a single oligo training stage to train a base-caller 1414 that includes a neural network (NN) configuration 1415 using a known synthetic sequence 1406.
[0133] In the example of Figure 14A, the base calling system 1400 includes a sequencing machine 1404, such as sequencing machine 400 of Figure 4. In an embodiment, the sequencing machine 1404 includes a biosensor (not illustrated in Figure 14A) that includes a flow cell 1405 similar to flow cell 102 of biosensor 100 of Figure 1.
[0134] As discussed with respect to Figures 2, 3, and 6, the flow cell 1405 comprises a number of clusters 1407a, ..., 1407G. Specifically, in an embodiment, the flow cell 1405 comprises a number of lanes of tiles, each tile including a corresponding number of clusters, as discussed with respect to Figure 2. In Figure 14A, the flow cell 1405 is illustrated as including several such exemplary clusters 1407a, ..., 1407G. In the base calling process, a base call (A, C, G, T) for each cluster in a particular cycle is predicted.
[0135] A typical flow cell 1405 can include multiple clusters 1407, such as thousands or millions of clusters. Solely by way of example, without limiting the scope of the present disclosure and to illustrate some of the principles of the present disclosure, we will assume that there are 10,000 (or 10k) clusters 1407 in the flow cell 1405 (i.e., G=10,000), although a practical flow cell will likely have a much larger number of such clusters.
[0136] In an embodiment, the known synthetic sequence 1406 is used as an analyte for base calling operations during the single oligo training phase. In an embodiment, the known synthetic sequence 1406 comprises a synthetically generated oligomer. Oligonucleotides are short DNA or RNA molecules, called oligomers or simply oligos, that have a wide range of applications in genetic testing, research, and forensics. These small amounts of nucleic acids, commonly created in laboratories by solid-phase chemical synthesis, can be manufactured as single-stranded molecules with any user-specified sequence and are therefore of great importance in artificial gene synthesis, polymerase chain reaction (PCR), DNA sequencing, molecular cloning, and as molecular probes. The length of an oligonucleotide is typically indicated by a "mer." For example, a 6 nucleotide (nt) oligonucleotide is a hexamer, while a 25 nt oligonucleotide is typically referred to as a "25-mer." In an embodiment, the size of the oligomer or oligo that comprises the known synthetic sequence 1406 can have any suitable number of bases, such as 8, 10, 12, or more, and is implementation specific. By way of example only, FIG. 14A illustrates an oligo of known synthetic sequence 1406 containing 8 bases.
[0137] The oligo referred to in Figure 14A is labeled as oligo #1 (or oligo number 1). Since only one unique oligo is used in Figure 14A, the same oligo #1 is populated into each cluster 1407. Thus, all 10k clusters 1407 are populated with the same oligo sequence; that is, copies of the same oligo are populated into all clusters 1407.
[0138] The sequencing machine 1404 generates sequence signals 1412a, ..., 1412G for corresponding clusters of the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, the sequencing machine 1404 generates a corresponding sequence signal 1412a indicative of the base sequence populated into cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the sequencing machine 1404 generates a corresponding sequence signal 1412b indicative of the base sequence populated into cluster 1407b for a series of sequencing cycles, and so on. The base caller 1414 receives the sequence signals 1412 and is intended to call (e.g., predict) the corresponding bases. In an embodiment, the base calling 1414, including the NN configuration 1415 (and various other NN configurations described later in this specification), can be stored in memory 404, 403, and / or 406 and can execute on a host CPU (such as CPU 402 in FIG. 4) and / or a configurable processor (such as configurable processor 450 in FIG. 4) that is local to the sequencing machine 400. In another embodiment, the base calling 1414 can be stored remotely from the sequencing machine 400 (e.g., stored in the cloud) and executed by a remote processor (e.g., executed in the cloud). For example, in a remote version of the base calling 1414, the base calling 1414 receives the sequence signal 1412 (e.g., over a network such as the Internet), performs base calling operations, and transmits the base calling results (e.g., over a network such as the Internet) to the sequencing machine 400.
[0139] In an example, the sequence signal 1412 includes an image captured by a sensor (e.g., a photodetector, a photodiode), as previously described herein. Thus, at least some of the examples and embodiments discussed herein relate to iteratively training a base caller (such as base caller 1414) to process sequence signals that include images. However, the principles of the present disclosure are not limited to training any particular type of base caller receiving any particular type of sequence signal. For example, the iterative training discussed herein in this disclosure is independent of the type of base caller trained or the type of sequence signal used. For example, the iterative training discussed herein in this disclosure can be used to train any other suitable type of base caller, such as a base caller configured to call bases based on sequence signals that do not include images. For example, the sequence signal can include an electrical signal (e.g., a voltage signal, a current signal), a pH level, and / or the like, and the iterative training methods discussed herein can be applied to training a base caller receiving any such type of sequence signal.
[0140] Neural network configuration 1415, as discussed in more detail herein, is a convolutional neural network (examples of which are illustrated in FIGS. 7, 9, 10, 11, and 12) that uses a relatively small number of layers and a relatively small number of parameters (e.g., compared to some other neural network configurations described later in this specification, such as neural network configuration 1615 of FIG. 16A).
[0141] An initially untrained base caller 1414 including a neural network configuration 1415 predicts base call sequences 1418a, ..., 1418G for corresponding clusters of a plurality of clusters 1407a, ..., 1407G based on corresponding sequence signals 1412a, ..., 1412G, respectively. For example, for cluster 1407a, the base caller 1414 predicts corresponding base call sequence 1418a including a base call for cluster 1407a for a set of sequencing cycles based on corresponding sequence signal 1412a. Similarly, for cluster 1407b, the base caller 1414 predicts corresponding base call sequence 1418b including a base call for cluster 1407b for a set of sequencing cycles based on corresponding sequence signal 1412b, and so on. Thus, G base call sequences 1418a, ..., 1418G are predicted by the base caller 1414.
[0142] Assume that oligo#1 has 8 bases generally labeled GA1, ..., GA8. Merely by way of example, and without limiting the scope of the present disclosure, assume that the 8 bases of oligo# are A, C, T, T, G, C, A, C. Initially, base caller 1414 is untrained and therefore likely to have errors in base calling. For example, predicted base call sequence 1418a (generally labeled Sa1, ..., Sa8) is C, A, T, C, G, C, A, G, as illustrated in FIG. 14A. Thus, when comparing ground truth base sequence 1406 of oligo#1 (i.e., A, C, T, T, G, C, A, C) with predicted base sequence 1418a (i.e., C, A, T, C, G, C, A, G), there are errors in base calling for base numbers 1, 2, 4, and 8. Thus, in FIG. 14A , the ground truth base sequence 1406 and the predicted base sequence 1418a of oligo #1 are compared in operation 1413a, and the error between these two base sequences is used in a backward pass of the neural network configuration 1415 of the base caller 1414 to train the neural network configuration 1415, such as to update the gradients and weights of the neural network configuration 1415 (symbolically labeled in FIG. 14A as gradient update 1417).
[0143] FIG. 14A1 illustrates in more detail the comparison operation between the predicted base sequence 1418a and the ground truth base sequence 1406 of oligo#1. For example, referring to FIG. 14A and FIG. 14A1, the predicted base sequence 1418a is C, A, T, C, G, C, A, G, and the ground truth base sequence 1406 of oligo#1 is A, C, T, T, G, C, A, C. Thus, when comparing the ground truth base sequence 1406 of oligo#1 (i.e., A, C, T, T, G, C, A, C) with the predicted base sequence 1418a (i.e., C, A, T, C, G, C, A, G), there are errors in the base calls for base numbers 1, 2, 4, and 8. For example, in FIG. 14A1, the error in the base call for base number 1 is given by "C should be A", i.e., base call C should be base call A. Similarly, the error in the base call for base number 2 is given by "A should be C," i.e., base call A should be base call B, etc. There are no errors in the base calls for base numbers 3, 5, 6, and 7 (illustrated as "Match (no error)" in FIG. 14A1). Thus, in FIG. 14A1, during comparison, each base call in the predicted base call sequence 1418a is compared to the corresponding base call in the corresponding ground truth sequence (e.g., base sequence 1406 of oligo #1) to generate a corresponding comparison result, as illustrated in FIG. 14A1.
[0144] Referring again to FIG. 14A, the base calling system 1400 also includes mapping logic 1416, the function of which is described later in this specification. In an embodiment, the mapping logic 1416 can be stored in memory 404, 403, and / or 406, and the mapping logic 1416 can execute on a host CPU (such as CPU 402 in FIG. 4) and / or a configurable processor (such as configurable processor 450 in FIG. 4) that is local to the sequencing machine 400. In another embodiment, the mapping logic 1416 can be stored remotely from the sequencing machine 400 (e.g., stored in the cloud) and executed by a remote processor (e.g., executed in the cloud). For example, in a remote version of the mapping logic 1416, the mapping logic receives the data to be mapped from the sequencing machine 400 (e.g., over a network such as the Internet), performs the mapping operation, and transmits the mapping results to the sequencing machine 400 (e.g., over a network such as the Internet). The mapping operation is discussed in more detail later in this specification.
[0145] FIG. 14A and various other figures, examples, and embodiments of the present disclosure refer to base callers that predict base call sequences. Various examples of such prediction of base call sequences are discussed herein. Further examples of base call prediction can be found in co-pending U.S. Provisional Patent Application No. 63 / 217,644, entitled "IMPROVED ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES," filed July 1, 2021 (Attorney Docket No. ILLM1046-1 / IP-2135-PRV), which is incorporated by reference as if fully set forth herein.
[0146] Figure 14B illustrates further details of the base calling system 1400 of Figure 14A operating in a single oligo training phase to train a base caller 1414 comprising a neural network configuration 1415 using a known synthetic sequence 1406. For example, Figure 14B illustrates the use of predicted base call sequences 1418a, ..., 1418G to train the base caller 1414. For example, each of the predicted base call sequences 1418a, ..., 1418G is compared to the ground truth base sequence 1406 of oligo #1 (see comparison operations 1413a, ..., 1413G), and the resulting error is used for gradient update and resulting parameter (e.g., weights and bias) update (symbolically labeled in Figure 14A as gradient update 1417) by the backpropagation section of the neural network configuration 1415.
[0147] Thus, the neural network configuration 1415 has been trained using the base call sequences 1418 predicted by the neural network configuration 1415, and using the ground truth base sequence of oligo #1 1406. Because the training discussed with respect to Figures 14A and 14B uses a single oligo, this training phase is also referred to as the "single oligo training phase," and Figures 14A and 14B are labeled accordingly.
[0148] In an embodiment, the process of FIG. 14A and FIG. 14B can be repeated iteratively. For example, in the first iteration of FIG. 14A, the NN configuration 1415 is at least partially trained. The at least partially trained NN configuration 1415 is again used during the second iteration to regenerate predicted base call sequences from the sequence signal 1412 (e.g., as discussed with respect to FIG. 14A), and the resulting predicted base call sequences are again compared to the ground truth 1406 (i.e., oligo #1) to generate an error signal, which is used to further train the NN configuration 1415. This process can be repeated iteratively multiple times until the NN configuration 1415 is sufficiently trained. In an embodiment, this process can be repeated iteratively for a certain number of times. In another embodiment, this process can be repeated iteratively until some errors saturate (e.g., errors in successive iterations do not decrease significantly).
[0149] FIG. 15A illustrates the base calling system 1400 of FIG. 14A operating in a training data generation phase of a two-oligo training stage to generate labeled training data using two known synthetic sequences 1501A and 1501B.
[0150] The base calling system 1400 of Figure 15A is the same as the base calling system of Figure 14A, in both figures the base calling system 1400 uses a neural network configuration 1415. Furthermore, two different unique oligo sequences 1501A and 1501B are loaded into various clusters of the flow cell 1405. By way of example only, and without limiting the scope of the present disclosure, assume that of the 10,000 clusters 1407, approximately 5,200 clusters are populated with oligo sequence 1501A and the remaining 4,800 clusters are populated with oligo sequence 1501B (although in alternative embodiments the two oligos can be split substantially equally among the 10,000 clusters).
[0151] The sequencing machine 1404 generates sequence signals 1512a, ..., 1512G for corresponding clusters of the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, the sequencing machine 1404 generates a corresponding sequence signal 1512a indicative of the bases of cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the sequencing machine 1404 generates a corresponding sequence signal 1512b indicative of the bases for cluster 1407b for a series of sequencing cycles, and so on.
[0152] A base caller 1414 comprising at least a partially trained neural network configuration 1415 (e.g., trained by iteratively repeating the operations of Figures 14A and 14B) predicts base call sequences 1518a, ..., 1518G for corresponding ones of the plurality of clusters 1407a, ..., 1407G based on corresponding sequence signals 1512a, ..., 1512G, respectively. For example, for cluster 1407a, the base caller 1414 predicts corresponding base call sequence 1518a comprising a base call for cluster 1407a for a set of sequencing cycles based on corresponding sequence signal 1512a. Similarly, for cluster 1407b, the base caller 1414 predicts corresponding base call sequence 1518b comprising a base call for cluster 1407b for a set of sequencing cycles based on corresponding sequence signal 1512b, and so on. Thus, G base call sequences 1518a, ..., 1518G are predicted by the base caller 1414. Note that the neural network configuration 1415 of Figure 15A was trained earlier during the iterations of the single oligo training phase discussed with respect to Figures 14A and 14B, and thus the predicted base call sequences 1518a,...,1518G are reasonably accurate, but not very accurate (because the base call 1414 has not been fully trained).
[0153] In an embodiment, oligo sequences 1501A and 1501B are selected to have a sufficient edit distance between the bases of the two oligos. Figures 15B and 15C illustrate two corresponding exemplary selections of oligo sequences 1501A and 1501B of Figure 15A. For example, in Figure 15B, oligo 1501A is selected to have bases A, C, T, T, G, C, A, C, while oligo 1501B is selected to have bases C, C, T, A, G, C, A, C. Thus, the first and fourth bases in the two oligos 1510A and 1510B are different, resulting in an edit distance of 2 between the two oligos 1510A and 1510B.
[0154] In contrast, in Figure 15B, oligo 1501A is selected to have bases A, C, T, T, G, C, A, C, while oligo 1501B is selected to have bases C, A, T, G, A, T, A, G. Thus, in the example of Figure 15B, the first, second, fourth, fifth, sixth, and eighth bases in the two oligos 1510A and 1510B are different, resulting in an edit distance of 6 between the two oligos 1510A and 1510B.
[0155] In an embodiment, the two oligos 1501A and 1501B are selected such that the two oligos are separated by at least a threshold edit distance. By way of example only, the threshold edit distance could be 4 bases, 5 bases, 6 bases, 7 bases, or even 8 bases. Thus, the two oligos 1501A and 1501B are selected such that the two oligos are sufficiently different from each other.
[0156] Referring again to FIG. 15A, base caller 1414 does not know which oligo sequences are populated into which clusters. Thus, base caller 1414 does not know the mapping between known oligo sequences 1501A, 1501B and the various clusters. In an embodiment, mapping logic 1416 receives predicted base call sequences 1518 and maps each predicted base call sequence 1518 to either oligo 1501A or oligo 1501B, or declares uncertainty when mapping a predicted base call sequence to either of the two oligos. FIG. 15D illustrates an exemplary mapping operation for (i) mapping a predicted base call sequence to either oligo 1501A or oligo 1501B, or (ii) declaring uncertainty when mapping a predicted base call sequence to either of the two oligos.
[0157] In an embodiment, the greater the edit distance between two oligos, the easier (or more accurate) it is to map an individual prediction to either of the two oligos. For example, referring to FIG. 15B, since the edit distance between the two oligos 1501A and 1501B is only 2, the two oligos are largely similar and it may be relatively difficult to map a base call prediction to either of the two oligos. However, since the edit distance between the two oligos 1501A and 1501B in FIG. 15C is 6, the two oligos are very dissimilar and it may be relatively easy to map a prediction to either of the two oligos. Thus, FIG. 15B, with an edit distance of 2, is labeled "less suitable for training" and FIG. 15C, with an edit distance of 6, is labeled "more suitable for training". Thus, in an embodiment, oligos 1501A and 1501B according to FIG. 15C (but not according to FIG. 15B) are generated and used for training, as discussed in more detail in turn herein.
[0158] Referring again to Figure 15D, exemplary predicted base call sequences 1518a, 1518b, and 1518G are illustrated. Also illustrated are exemplary bases for two oligos 1501A and 1501B (the exemplary bases for the two oligos correspond to the bases illustrated in Figure 15C).
[0159] Because the neural network configuration 1415 has been somewhat trained but not fully trained, the neural network configuration 1415 may be capable of making base call predictions, but such base call predictions are prone to error.
[0160] Predicted base call sequence 1518a includes C, A, G, G, C, T, A, C. It is compared to the base call sequence of oligo 1501A, A, C, T, T, G, C, A, C, and to the base call sequence of oligo 1501B, C, A, T, G, A, T, A, G. Predicted base call sequence 1518a has its seventh and eighth bases matching the corresponding seventh and eighth bases of oligo 1501A, and its first, second, fourth, sixth, and seventh bases matching the corresponding bases of oligo 1501B. Thus, as illustrated in FIG. 15D, predicted base call sequence 1518a has two bases of similarity with oligo 1501A, and predicted base call sequence 1518a has five bases of similarity with oligo 1501B.
[0161] If the predicted base call sequence 1518a is in fact for oligo 1501B (e.g., such that the predicted base call sequence 1518a has 5 bases similarity to oligo 1501B), this means that the neural network configuration 1415 was able to correctly predict 5 bases of the 8 base sequence (i.e., correctly predicting the 1st, 2nd, 4th, 6th, and 7th bases that match the corresponding bases in oligo 1501B). However, because the neural network configuration 1415 was not fully trained, the neural network configuration 1415 made errors in predicting the remaining 3 bases (i.e., the 3rd, 5th, and 8th bases).
[0162] The mapping logic 1416 can map the predicted base call array to the corresponding oligo using appropriate logic. For example, assume that the predicted base call array has a similarity in the number of SAs with oligo 1501A and a similarity in the number of SBs with oligo 1501B. In an example, the mapping logic 1416 maps the predicted base call array to oligo 1501A when SA > ST and SB < ST, where ST is a threshold number. That is, the mapping logic 1416 maps the predicted base call array to oligo 1501A when the similarity level with oligo 1501A is higher than the threshold and the similarity level with oligo 1501B is lower than the threshold.
[0163] Similarly, in another example, the mapping logic 1416 maps the predicted base call array to oligo 1501B when SB > ST and SA < ST.
[0164] In yet another example, the mapping logic 1416 declares that the predicted base call array is uncertain when both SA and SB are less than the threshold ST, or when both SA and SB are greater than the threshold ST.
[0165] The above discussion can be written in the form of equations as follows. For the predicted base call array, If SA > ST and SB < ST, map to oligo 1501A. (Equation 1) If SB > ST and SA < ST, map to oligo 1501B. (Equation 2) If both SA and SB are < ST, declare an uncertain mapping. Or (Equation 3) If both SA and SB are > ST, declare an uncertain mapping. (Equation 4)
[0166] The threshold ST depends on the number of bases in the oligo (which is 8 in the exemplary use case illustrated in the figure), the desired accuracy, and / or is implementation specific. By way of example only, the threshold ST is assumed to be 4 in the exemplary use case illustrated in FIG. 15D. Note that the threshold ST of 4 is merely an example, and the selection of the threshold ST may be implementation specific. By way of example only, during the first iteration of training, the threshold ST may have a relatively low value (e.g., 4). The threshold ST may have a relatively high value (e.g., 6 or 7) during later iterations of training (training iterations are described later in this specification). Thus, the threshold ST may be gradually increased as the NN configuration becomes better trained during later training iterations. However, in another example, the threshold ST may have the same value throughout all iterations of training. While the threshold ST is selected as 4 in the example of FIG. 15D, in other exemplary implementations, the threshold ST may be, for example, 5, 6, or 7. In an example, the threshold ST may also be expressed as a percentage. For example, if the threshold ST is 4 and the total number of bases is 8, the threshold ST can be expressed as (4 / 8)×100, i.e., 50%. The threshold ST can be a user-selectable parameter, and in an embodiment, can be selected to be between 50% and 95%.
[0167] 15D, as described above, predicted base call sequence 1518a has 2 bases of similarity with oligo 1501A, and predicted base call sequence 1518a has 5 bases of similarity with oligo 1501B. Therefore, SA=2, SB=5. According to Equation 2, assuming a threshold ST of 4, predicted base call sequence 1518a maps to oligo 1501B.
[0168] Referring now to predicted base call sequence 1518b, predicted base call sequence 1518b has 2 bases of similarity with oligo 1501A, and predicted base call sequence 1518b has 3 bases of similarity with oligo 1501B. Therefore, SA=2, SB=3. According to Equation 3, assuming a threshold ST of 4, predicted base call sequence 1518b is declared uncertain for mapping to any of the oligo sequences.
[0169] Referring now to predicted base call sequence 1518G, predicted base call sequence 1518G has 6 bases of similarity with oligo 1501A, and predicted base call sequence 1518G has 3 bases of similarity with oligo 1501B. Therefore, SA=6, SB=3. According to Equation 2, assuming a threshold ST of 4, predicted base call sequence 1518G maps to oligo 1501A.
[0170] FIG. 15E illustrates labeled training data 1550 generated from the mapping of FIG. 15D, where the labeled training data 1550 is used by another neural network configuration 1615 (e.g., illustrated in FIG. 16A, where the other neural network configuration 1615 is different and more complex than the neural network configuration 1415 of FIGS. 14A, 14B, and 15A).
[0171] As illustrated in FIG. 15E, some of the predicted base call sequences 1518 and corresponding sequence signals map to the base sequence of oligo 1501A (i.e., ground truth 1506a), some other predicted base call sequences 1518 and corresponding sequence signals map to the base sequence of oligo 1501B (i.e., ground truth 1506b), and the mapping of the remainder of the predicted base call sequences 1518 and corresponding sequence signals is uncertain.
[0172] For example, predicted base call sequences 1518c, 1518d, 1518G and corresponding sequence signals 1512c, 1512d, 1512G are mapped to the base sequence of oligo 1501A (i.e., ground truth 1506a), predicted base call sequences 1518a, 1518f and corresponding sequence signals 1512a, 1512f are mapped to the base sequence of oligo 1501B (i.e., ground truth 1506b), and the remaining mappings of predicted base call sequences 1518b, 1518e, 1518g and corresponding sequence signals 1512b, 1512e, 1512g are uncertain.
[0173] Solely by way of example, assume that 2,600 base call sequences of training data 1550 are mapped to oligo 1501A and 3,000 base call sequences of training data 1550 are mapped to oligo 1501B. As illustrated in Figure 15E, the remaining 4,400 base call sequences are undetermined and do not map to either of the two oligos.
[0174] Note that Figures 15A, 15D, and 15E are referred to as the "training data generation phase" of the "2 oligo training stage" because the labeling training data 1550 is generated using sequences from two oligos and using the neural network configuration 1415.
[0175] FIG. 16A illustrates the base-calling system 1400 of FIG. 14A operating in the "training data consumption and training phase" of the "two oligo training stage" to train a base-caller 1414 with another neural network configuration 1615 (different and more complex than the neural network configuration 1415 of FIG. 14A) using two known synthetic sequences 1501A and 1501B.
[0176] The base calling system 1400 of FIG. 16A is the same as the base calling system of FIG. 14A. However, unlike FIG. 14A (where neural network configuration 1415 was used in base calling 1414), the base calling system 1414 of FIG. 16A uses a different neural network configuration 1615. The neural network configuration 1615 of FIG. 16A is different from the neural network configuration 1415 of FIG. 14A. For example, the neural network configuration 1615 is a convolutional neural network (examples of which are illustrated in FIG. 7, FIG. 9, FIG. 10, FIG. 11, FIG. 12) that uses a greater number of layers and parameters (such as weights and biases) than the neural network configuration 1415. In another example, the neural network configuration 1615 is a convolutional neural network that uses a greater number of convolution filters than the neural network configuration 1415. The configuration, topology, and number of layers and / or filters of the two neural network configurations 1415 and 1615 may differ in some examples.
[0177] In the "Training Data Consumption and Training Phase" of the "2 Oligo Training Stage" illustrated in Figure 16A, the base caller 1414 with neural network configuration 1615 receives the sequence signal 1512 previously generated during the "Training Data Generation Phase" of Figure 15A. That is, the base caller 1414 with neural network configuration 1615 reuses the sequence signal 1512 previously generated. Thus, since the previously generated sequence signal 1512 is reused in the "Training Data Consumption and Training Phase" of the "2 Oligo Training Stage" illustrated in Figure 16A, the sequencing machine 1404 and components therein play no role and are therefore illustrated using dotted lines. Similarly, the mapping logic 1416 plays no role (as no mapping has been performed in Figure 16A) and therefore the mapping logic 1416 is also illustrated using dotted lines.
[0178] 16A, base caller 1414 with neural network configuration 1615 receives previously generated sequence signal 1512 and predicts base call sequence 1618 from sequence signal 1512. Predicted base call sequence 1618 includes predicted base call sequences 1618a, 1618b, ..., 1618G. For example, sequence signal 1512a is used to predict base call sequence 1618a, sequence signal 1512b is used to predict base call sequence 1618b, sequence signal 1512G is used to predict base call sequence 1618G, etc.
[0179] The neural network configuration 1615 has not yet been trained, and therefore the predicted base call sequences 1618a, 1618b, ..., 1618G may have many errors. The mapped training data 1550 of Figure 15E is now used to train the neural network configuration 1615. For example, from the training data 1550, the base call 1414 is (i) sequence signals 1512c, 1512d, 1512G are for the base sequence of oligo 1501A (i.e., ground truth 1506a); (ii) sequence signals 1512a, 1512f are for the base sequence of oligo 1501B (i.e., ground truth 1506b); and (iii) Know that the mapping of constellation signals 1512b, 1512e, 1512g is uncertain.
[0180] Thus, the sequence signals 1512 and predicted base call sequences 1518 are classified into a first category including sequence signals 1512c, 1512d, 1512G (and corresponding predicted base call sequences 1518c, 1518d, 1518G) that can be mapped to the base sequence of oligo 1501A (i.e., ground truth 1506a); (ii) a second category including sequence signals 1512a, 1512f (and corresponding predicted base call sequences 1518a, 1518f) that can be mapped to either the base sequence of oligos 1501A or 1501B; and (iii) a third category including sequence signals 1512b, 1512e, 1512g (and corresponding predicted base call sequences 1518b, 1518e, 1518g) that cannot be mapped to either the base sequence of oligos 1501A or 1501B.
[0181] Thus, based on (iii) above, predicted base call sequences 1618b, 1618e, and 1618g (e.g., corresponding to sequence signals 1512b, 1512e, 1512g) are not used to train neural network configuration 1615. Thus, predicted base call sequences 1618b, 1618e, and 1618g are discarded during the training iterations and are not used for gradient update (symbolically illustrated in FIG. 16A using an "X" or "cross symbol" between predicted base call sequences 1618b, 1618e, and 1618g and gradient update box 1617).
[0182] Based on (i) above, base caller 1414 knows that predicted base call sequences 1618c, 1618d, 1618G (e.g., corresponding to sequence signals 1512c, 1512d, 1512G) are likely to be for oligo 1501A. That is, the base sequence of oligo 1501A is likely the ground truth for these predicted base call sequences 1618c, 1618d, 1618G, but the untrained neural network configuration 1615 may have mispredicted at least some bases of these predicted base call sequences. Thus, the neural network configuration uses comparison function 1613 to compare each of predicted base call sequences 1618c, 1618d, and 1618G to ground truth 1506a (which is the base sequence of oligo 1501A) and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615.
[0183] Similarly, based on (ii) above, the base caller knows that predicted base call sequences 1618a and 1618f (e.g., corresponding to sequence signals 1512a and 1512f, respectively) are likely to be for oligo 1501B. That is, the base sequence of oligo 1501B is likely the ground truth for these predicted base call sequences 1618a and 1618f, but the untrained neural network configuration 1615 may have mispredicted at least some bases of these predicted base call sequences. Thus, the neural network configuration uses comparison function 1613 to compare each of predicted base call sequences 1618a and 1618f to ground truth 1506b (which is the base sequence of oligo 1501B) and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615.
[0184] At the end of the training data consumption and training phase of FIG. 16A, the NN configuration 1615 is at least partially trained.
[0185] FIG. 16B illustrates the base calling system 1400 of FIG. 14A operating in a second iteration of the training data generation phase of the two-oligo training stage. For example, in FIG. 16A, the neural network configuration 1615 was trained using the training data 1550. In FIG. 16B, some or at least a partially trained neural network configuration 1615 is used to generate further training data. For example, the at least a partially trained neural network configuration 1615 predicts a base call sequence 1628 using a previously generated sequence signal 1512. The predicted base call sequence 1628 of FIG. 16B is likely to be relatively more accurate than the predicted base call sequence 1618 of FIG. 16A because the predicted base call sequence 1618 of FIG. 16A was generated using an untrained neural network configuration 1615, whereas the predicted base call sequence 1628 of FIG. 16B was generated at least in part using the neural network configuration 1615.
[0186] Further, mapping logic 1416 maps each of the predicted base call sequences 1628 to either oligo 1501A or oligo 1501B, or declares the mapping of the predicted base call sequence 1628 to be uncertain (e.g., similar to the discussion in relation to Figure 15D).
[0187] FIG. 16C illustrates labeled training data 1650 generated from the mapping of FIG. 16B, which is used for further training.
[0188] As illustrated in FIG. 16C, some of the predicted base call sequences 1628 and corresponding sequence signals 1512 map to the base sequence of oligo 1501A (i.e., ground truth 1506a), some other predicted base call sequences 1628 and corresponding sequence signals 1512 map to the base sequence of oligo 1501B (i.e., ground truth 1506b), and the mapping of the remaining predicted base call sequences 1628 and corresponding sequence signals 1512 is uncertain.
[0189] For example, predicted base call sequences 1628 are sorted into three categories: (i) predicted base call sequences 1628c, 1628d, and 1628G and corresponding sequence signals 1512c, 1512d, and 1512G are mapped to the base sequence of oligo 1501A (i.e., ground truth 1506a); (ii) predicted base call sequences 1628a, 1628b, and 1628f and corresponding sequence signals 1512a, 1512b, and 1512f are mapped to the base sequence of oligo 1501B (i.e., ground truth 1506b); and (iii) the mapping of the remaining predicted base call sequences 1628e and 1628g and corresponding sequence signals 1512e and 1512g is indeterminate.
[0190] Solely by way of example, assume that 3,300 base call sequences of training data 1650 are mapped to oligo 1501A and 3,200 base call sequences of training data 1650 are mapped to oligo 1501B. As illustrated in Figure 16C, the remaining 3,500 base call sequences are undetermined and do not map to either of the two oligos.
[0191] Comparing the number of sequences with unmapped (or uncertain) base calls between the training data of Figures 15E and 16C, it is observed that this number is 4,400 in Figure 15E and 3,500 in Figure 16C. This is because the at least partially trained neural network configuration 1615 of Figure 16B (used to generate the mappings of the training data 1650) is relatively more accurate and / or can be trained better than the at least partially trained neural network configuration 1415 of Figure 15A (used to generate the mappings of the training data 1550). Thus, the number of sequences with uncertain base calls gradually decreases as the base calls become relatively more accurate (e.g., have fewer errors) and thus relatively more correctly mapped.
[0192] FIG. 16D illustrates the base calling system 1400 of FIG. 14A operating in a second iteration of the "training data consumption and training phase" of the "2-oligo training stage" to train the base calling system 1414 comprising the neural network configuration 1615 of FIG. 16A using two known synthetic sequences 1501A and 1501B.
[0193] Figures 16A and 16D are at least partially similar, for example, Figures 16A and 16D are used to train a neural network configuration 1615 using training data 1550 of Figure 15E and training data 1650 of Figure 16C, respectively. Note that in the initial stage of Figure 16A, the neural network configuration 1615 is not trained at all, whereas in the initial stage of Figure 16D, the neural network configuration 1615 is at least partially trained.
[0194] In Figure 16D, base caller 1414, including an at least partially trained neural network configuration 1615, receives sequence signal 1512 previously generated during the "training data generation phase" of Figure 15A and predicts base call sequences 1638 from sequence signal 1512. Predicted base call sequences 1638 include predicted base call sequences 1638a, 1638b, ..., 1638G. For example, sequence signal 1512a is used to predict base call sequence 1638a, sequence signal 1512b is used to predict base call sequence 1638b, sequence signal 1512G is used to predict base call sequence 1638G, etc.
[0195] Neural network configuration 1615 is not fully trained, and thus predicted base call sequences 1638a, 1638b, ..., 1638G contain some errors, but the errors in predicted base call sequence 1638 of Figure 16D are likely to be less than the errors in predicted base call sequence 1618 of Figure 16A and predicted base call sequence 1628 of Figure 16B. Mapped training data 1650 of Figure 16C is now used to further train neural network configuration 1615. For example, from training data 1650, base call 1414 is (i) sequence signals 1512c, 1512d, 1512G are for the base sequence of oligo 1501A (i.e., ground truth 1506a); (ii) sequence signals 1512a, 1512b, 1512f are for the base sequence of oligo 1501B (i.e., ground truth 1506b); and (iii) Know that the mapping of constellation signals 1512e, 1512g is uncertain.
[0196] Thus, based on (iii) above, predicted base call sequences 1638e and 1638g in Figure 16D (e.g., corresponding to sequence signals 1512e and 1512g, respectively) are not used to train neural network configuration 1615. These predicted base call sequences 1638e and 1638g are therefore discarded from the training data and are not used for gradient update (symbolically illustrated in Figure 16D using an "X" or "cross symbol" between the predicted base call sequences 1618e, 1618g and the gradient update box 1617).
[0197] Based on (i) above, base caller 1414 knows that predicted base call sequences 1638c, 1638d, and 1638G (e.g., corresponding to sequence signals 1512c, 1512d, and 1512G, respectively) are likely to be for oligo 1501A. That is, the base sequence of oligo 1501A is likely the ground truth for these predicted base call sequences 1638c, 1638d, and 1638G, but in part, neural network configuration 1615 may have mispredicted at least some bases of these predicted base call sequences. Thus, neural network configuration 1615 uses comparison function 1613 to compare each of predicted base call sequences 1638c, 1638d, and 1638G to ground truth 1506a (which is the base sequence of oligo 1501A) and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615. For example, during comparison, each base call of the predicted base call sequence 1638c is compared to a corresponding base call of the corresponding ground truth sequence to generate a corresponding comparison result, for example, as discussed with respect to FIG. 14A1.
[0198] Similarly, based on (ii) above, the base caller knows that predicted base call sequences 1638a, 1638b, and 1638f (e.g., corresponding to sequence signals 1512a, 1512b, and 1512f, respectively) are likely to be for oligo 1501B. That is, the base sequence of oligo 1501A is likely the ground truth for these predicted base call sequences 1638a, 1638b, and 1638f, but in part, neural network configuration 1615 may have mispredicted at least some bases of these predicted base call sequences. Thus, the neural network configuration uses comparison function 1613 to compare each of predicted base call sequences 1638a, 1638b, and 1638f to ground truth 1506b (which is the base sequence of oligo 1501B) and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615.
[0199] FIG. 17A illustrates a flow chart depicting an exemplary method 1700 for iteratively training neural network configurations for base calling using single oligo and two oligo sequences. Method 1700 progressively trains NN configurations that are progressively and monotonically complex in nature. Increasing the complexity of the NN configurations may include increasing the number of layers of the NN configuration, increasing the number of filters of the NN configuration, increasing the topological complexity in the NN configuration, and / or the like. For example, method 1700 refers to a first NN configuration (which is NN configuration 1415 discussed previously herein with respect to FIG. 14A and other figures), a second NN configuration (which is NN configuration 1615 discussed previously herein with respect to FIG. 16A and other figures), a Pth NN configuration (not specifically discussed with respect to FIG. 14A-FIG. 16D), etc. In an example, as symbolically illustrated in box 1710 of Figure 17A, the complexity of the Pth NN configuration is higher than the complexity of the (P-1)th NN configuration, which is higher than the complexity of the (P-2)th NN configuration, etc. The complexity of the second NN configuration is therefore monotonically increasing (i.e., each later stage NN configuration has a complexity at least as high as the previous stage NN configuration).
[0200] It should be noted that in method 1700, operation 1704a is for iteratively training a first NN configuration and generating labeled training data for a second NN configuration, operations 1704b1-1704bk are for training a second NN configuration and generating labeled training data for a third NN configuration, and operation 1704c is for training a third NN configuration and generating labeled training data for a fourth NN configuration. This process continues with operation 1704P being for training a Pth NN configuration and generating labeled training data for a subsequent NN configuration. Thus, generally speaking, in method 1700, operation 1704i is for training an ith NN configuration and generating labeled training data for an (i+1)th NN configuration, where i=1,...,P.
[0201] The method 1700 includes, at 1704a, (i) iteratively training a first NN configuration with a single oligo sequence, and (ii) generating first two-oligo-labeled training data using the trained first NN configuration. As discussed, the first NN configuration is the NN configuration 1415 of FIG. 14A, and the single oligo sequence includes oligo#1 discussed with respect to FIG. 14A, FIG. 14B. The iterative training of the first NN configuration with the single oligo sequence is discussed with respect to FIG. 14A, FIG. 14B. The generation of the first two-oligo-labeled training data using the trained first NN configuration is discussed with respect to FIG. 15A, FIG. 15D, FIG. 15E, where the first two-oligo-labeled training data is training data 1550 of FIG. 15E.
[0202] Method 1700 then proceeds from 1704a to 1704b. Illustratively, operation 1704b is for training a second NN configuration (e.g., using the first two-oligo-labeled training data generated from operation 1704a) and using the trained second NN configuration to generate further two-oligo-labeled training data for training a third NN configuration. Operation 1704b includes suboperations at blocks 1704b1-1704bk.
[0203] In block 1704b1, (i) a second NN configuration is trained using the first two-oligo-labeled training data generated in 1704a, and (ii) a second two-oligo-labeled training data is generated using the at least partially trained second NN configuration. As discussed, the second NN configuration is NN configuration 1615 of FIG. 16A. Training of the second NN configuration using the first two-oligo-labeled training data is also illustrated in FIG. 16A. Generation of the second two-oligo-labeled training data (e.g., training data 1650 of FIG. 16C) using the at least partially trained second NN configuration is discussed with respect to FIG. 16B and FIG. 16C.
[0204] Method 1700 then proceeds from 1704b1 to 1704b2. At block 1704b2, (i) the second NN configuration is further trained using the second two-oligo-labeled training data, and (ii) the third two-oligo-labeled training data is generated using the further trained second NN configuration. Training the second NN configuration using the second two-oligo-labeled training data is illustrated in FIG. 16D. Generation of the third two-oligo-labeled training data using the further trained second NN configuration is not illustrated, but is similar to the discussion of FIG. 16B and FIG. 16C.
[0205] Note that block 1704b1 is a first iteration of training a second NN configuration, block 1704b2 is a second iteration of training a second NN configuration, and so on, and finally block 1704bk is a kth iteration of training a second NN configuration. As discussed, the operation of block 1704b1 is discussed in detail with respect to Figures 16A, 16B, and 16C. The operation of subsequent blocks 1704b2, ..., 1704bk may be similar to the discussion for block 1704b1.
[0206] Note that the same second NN configuration is used in all of the iterations 1704b1, ..., 1704bk, and thus these k iterations aim to repeatedly train the same second NN configuration without increasing the complexity of the second NN configuration.
[0207] Training of the second NN configuration proceeds with each iteration of blocks 1704b1, 1704b2, ..., 1704bk. As the second neural network is gradually trained at each step of iterations 1704b1, ..., 1704bk, the second neural network makes progressively less error in predicting base call sequences. For example, as shown in block 1704a and also illustrated in FIG. 15E, the first two-oligo-labeled training data (i.e., training data 1550) generated using the trained first NN configuration has 44% (i.e., 4,400 out of 10,000) uncertain mappings. As shown in block 1704b1 and also illustrated in FIG. 16C, the second two-oligo-labeled training data (i.e., training data 1650) generated using the partially trained second NN configuration has 35% (i.e., 3,500 out of 10,000) uncertain mappings. As shown in block 1704b2 and by way of example only, a third two-oligo label training data set generated using the further trained second NN configuration may have 32% (i.e., 3,200 out of 10,000) uncertain mappings. The percentage of uncertain mappings may gradually decrease with each iteration until it reaches, for example, about 20% at block 1704bk.
[0208] The number of iterations "k" for training the second NN configuration may be based on satisfying one or more convergence conditions. Once the convergence conditions are satisfied, the iterations for training the second NN configuration may terminate. The convergence conditions are implementation specific and dictate the number of iterations undergone to train the second NN configuration. In an example, satisfying the convergence conditions indicates that further iterations may not significantly aid in further training of the second NN configuration, and thus, the training iterations for the second NN configuration may terminate. Several examples of convergence conditions and satisfying the same are discussed herein. For example, the second NN configuration may be iteratively trained until the proportion of uncertain mappings is less than a threshold proportion. Here, the convergence conditions are satisfied when the proportion of uncertain mappings is less than a threshold proportion. For example, for the second NN configuration, this threshold may be approximately 20%, by way of example only. Thus, once the threshold is satisfied at iteration k, the convergence conditions are satisfied and the training of the second NN configuration terminates. Thus, the method proceeds to 1704c, where the Kth two-oligo-labeled training data generated in block 1704bk is used to train a third NN configuration, which is more complex than the second NN configuration.
[0209] In another example, the iterations of the second NN configuration continue until the proportion of uncertain mappings saturates to some extent (i.e., does not decrease significantly with successive iterations) and satisfies the convergence condition. That is, in this example, saturation below a threshold indicates sufficient convergence of the iterative training (e.g., indicates satisfaction of the convergence condition) and the iterations of the current model can be terminated because further iterations cannot significantly improve the model. For example, assume that in iteration (k-2) (e.g., in block 1704b(k-2)), the proportion of uncertain mappings is 21%, in iteration (k-1) (e.g., in block 1704b(k-2)), the proportion of uncertain mappings is 20.4%, and in iteration k (e.g., in block 1704bk), the proportion of uncertain mappings is 20%. Thus, in the last two iterations, the decrease in the proportion of uncertain mappings is relatively low (e.g., 0.6% and 0.4%, respectively), suggesting that the training is nearly saturated and further training cannot significantly improve the second NN configuration. Here, the saturation is measured as the difference between the proportion of uncertain mappings in two successive iterations. That is, if two successive iterations have approximately the same proportion of uncertain mappings, further iterations may not help to further reduce this proportion, and therefore the training iterations can be terminated. Thus, at this stage, the iterations for the second NN configuration are terminated, and the method 1700 proceeds to 1704c for the third NN configuration.
[0210] In yet another embodiment, a number of iterations "k" is pre-specified, and completing k iterations satisfies the convergence condition, such that training for the current NN configuration can be terminated and the next NN configuration can be initiated.
[0211] Thus, at the end of the iterations for the second NN configuration (i.e., at the end of block 1704k), the method 1700 proceeds to block 1704c, where a third NN configuration is iteratively trained. The training of the third NN configuration also involves iterations similar to those discussed with respect to operations 1704b1,...,1704bk, and therefore will not be discussed in further detail.
[0212] This process of progressively training more complex NN configurations continues until, at 1704P of method 1700, the Pth NN configuration has been trained and two oligo training data has been generated for training the next NN configuration.
[0213] Note that in an embodiment, as discussed herein, the same two oligo sequences may be used for all iterations of blocks 1704b1, ..., 1704bk, 1704c, ..., 1704P, however, in some other embodiments, not discussed herein, different two oligo sequences may be used for different iterations of method 1700 of FIG.
[0214] As discussed, the more complex the model is, the better the model can be trained to predict base calls. For example, at the end of the training of the second NN configuration, the final label training data generated by the second NN configuration has 20% uncertain mapping. The percentage of uncertain mapping further decreases at the end of the training of the third NN configuration. For example, during the first training iteration of the third NN configuration, the percentage of uncertain mapping is 36% (e.g., because the third NN configuration is hardly trained during the first iteration), and this percentage gradually decreases with the subsequent training iterations of the third NN configuration. As illustrated in FIG. 17A, for example, assume that at the end of the training of the third NN configuration, the final label training data generated by the third NN configuration has 17% uncertain mapping. This percentage of uncertain mapping further decreases with the progress of the iterations of FIG. 17A, for example, at the end of the training of the Pth NN configuration, the final label training data generated by the Pth NN configuration has 12% uncertain mapping. It should be noted that the training ends with 12% uncertain mappings, for example, when the convergence condition (described above in this specification) is met for the Pth NN configuration. Thus, P NN configurations are trained in method 1700. The number "P" can be 3, 4, 5, or more, and is implementation specific and can also be based on meeting one or more corresponding convergence conditions. For example, if the (P-1)th NN configuration results in 12.05% uncertain mappings and the Pth NN configuration results in 12% uncertain mappings, there is a slight improvement of 0.05% uncertain mappings between the two NN configurations. This indicates that the training of the new NN configuration with two oligo sequences is saturated. Here, saturation refers to the difference in the percentage of uncertain mappings between two successive NN configurations. If the saturation is below a threshold value (e.g., 0.1%), the training of the two oligo sequence training is terminated. In another embodiment, the number "P" of NN configurations can be pre-specified by the user, for example, 3, 4, or more. As discussed below, once training with P NN configurations using two-oligo sequences is complete, more complex examples (such as three-oligo sequences) can be used for training.
[0215] FIG. 17B illustrates an exemplary final labeled training data generated by the Pth NN configuration at the end of method 1700 of FIG. 17A. As discussed, at the end of training the Pth NN configuration, the final labeled training data generated by the Pth NN configuration has 12% (or 1,200 out of 10,000) uncertain mappings. The predicted base call sequences are sorted into three categories: (i) a first category including predicted base call sequences that map to oligo 1501A, (ii) a second category including predicted base call sequences that map to oligo 1501B, and (iii) a third category including predicted base call sequences that do not map to either oligo 1501A or 1501B. The training data 1750 of FIG. 17B will become apparent based on discussion of the training data of FIG. 15E and FIG. 16C.
[0216] FIG. 18A illustrates the basecalling system 1400 of FIG. 14A operating in a first iteration of the "training data consumption and training phase" of the "3-oligo training stage" to train a basecaller 1414 with a 3-oligo neural network configuration 1815. The reason for labeling the neural network configuration 1815 as a "3-oligo" neural network configuration 1815 will become clear later in this specification. FIG. 18A is at least partially similar to FIG. 16D. However, unlike FIG. 15D, labeled training data 1750 (see FIG. 17B) generated at the end of the method 1700 (e.g., by the Pth NN configuration using 2-oligo base training) is used during training in FIG. 18A.
[0217] For example, in Figure 18A, base call 1414 including a three-oligo neural network configuration 1815 predicts base call sequences 1838a, 1838b, ..., 1838G. The mapped training data 1750 of Figure 17B is now used to further train the three-oligo neural network configuration 1815, similar to the training discussed with respect to Figure 16D.
[0218] FIG. 18B illustrates the base-calling system 1400 of FIG. 14A operating in the "training data generation phase" of the "three-oligo training stage" to train a base-caller 1414 comprising the three-oligo neural network configuration 1815 of FIG. 18A.
[0219] In Figure 18B, three different oligo sequences 1801A, 1801B, and 1801C are loaded into various clusters of flow cell 1405. By way of example only, and without limiting the scope of the present disclosure, assume that of the 10,000 clusters 1407, approximately 3,200 clusters contain oligo sequence 1801A, approximately 3,300 clusters contain oligo sequence 1801B, and the remaining 3,500 clusters contain oligo sequence 1501C (although in alternative embodiments, the three oligos can be divided substantially equally among the 10,000 clusters).
[0220] The sequencing machine 1404 generates sequence signals 1812a, ..., 1812G for corresponding clusters of the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, the sequencing machine 1404 generates a corresponding sequence signal 1812a indicative of the bases of cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the sequencing machine 1404 generates a corresponding sequence signal 1812b indicative of the bases for cluster 1407b for a series of sequencing cycles, and so on.
[0221] A base caller 1414 including a neural network configuration 1815 predicts base call sequences 1818a, ..., 1818G for corresponding clusters of a plurality of clusters 1407a, ..., 1407G based on corresponding sequence signals 1812a, ..., 1812G, respectively, as discussed, for example, with respect to FIG. 15A.
[0222] In an embodiment, oligo sequences 1801A, 1801B, and 1801C are selected to have a sufficient edit distance between the bases of the three oligos, for example, as will become apparent based on a discussion of Figures 15B and 15C. For example, any of the three oligo sequences 1801A, 1801B, and 1801C is separated from another of the three oligo sequences 1801A, 1801B, and 1801C by at least a threshold edit distance. By way of example only, the threshold edit distance could be 4 bases, 5 bases, 6 bases, 7 bases, or even 8 bases. Thus, the three oligos are selected such that the three oligos are sufficiently different from one another.
[0223] Referring again to FIG. 18B, in an embodiment, base caller 1414 does not know which oligo sequences are populated into which clusters. Thus, base caller 1414 does not know the mapping between known oligo sequences 1801A, 1801B, and 1801C and the various clusters. Mapping logic 1416 receives predicted base call sequences 1818 and maps each predicted base call sequence 1818 to one of oligos 1801A, 1801B, or 1801C, or declares uncertainty in mapping the predicted base call sequence to any of the three oligos. FIG. 18C illustrates the mapping operations to (i) map the predicted base call sequence to any of the three oligos 1801A, 1801B, 1801C, or (ii) declare the mapping of the predicted base call sequence to any of the three oligos as uncertain.
[0224] 18C, predicted base call sequence 1818a has 2 bases of similarity with oligo 1801A, 5 bases of similarity with oligo 1801B, and 1 base of similarity with oligo 1801C. Assuming a threshold similarity ST (e.g., as discussed with respect to Equations 1-4) of 4, predicted base call sequence 1818a maps to oligo 1801B.
[0225] Similarly, in the example of Figure 18C, predicted base call sequence 1818b is mapped to oligo 1801C, and the mapping of predicted base call sequence 1818a is declared uncertain by mapping logic 1416 of Figure 18B.
[0226] FIG. 18D illustrates labeled training data 1850 generated from the mapping of FIG. 18C, which is used to train another neural network configuration. As illustrated in FIG. 18D, some of the predicted base call sequences 1818 and corresponding sequence signals are mapped to the base sequence of oligo 1801A (i.e., ground truth 1806a), some of the predicted base call sequences 1818 and corresponding sequence signals are mapped to the base sequence of oligo 1801B (i.e., ground truth 1806b), some of the predicted base call sequences 1818 and corresponding sequence signals are mapped to the base sequence of oligo 1801C (i.e., ground truth 1506c), and the remaining mappings of the predicted base call sequences 1818 and corresponding sequence signals are undetermined. The training data 1850 of FIG. 18D will be clear based on the discussion of the training data 1550 of FIG. 15E above.
[0227] FIG. 18E illustrates a flowchart depicting an exemplary method 1880 for iteratively training a neural network configuration for base calling using a 3-oligo ground truth sequence. Method 1800 progressively trains a 3-oligo NN configuration that is progressively and monotonically complex in nature. Increasing the complexity of the NN configuration may include increasing the number of layers of the NN configuration, increasing the number of filters of the NN configuration, increasing the topological complexity in the NN configuration, and / or the like, as also discussed with respect to FIG. 17A. For example, method 1880 references a first 3-oligo NN configuration (which is the 3-oligo NN configuration 1815 discussed earlier herein with respect to FIG. 18A), a second 3-oligo NN configuration, a Qth NN configuration, etc. In an embodiment, as symbolically illustrated in box 1890 of FIG. 18E, the complexity of the Qth three-oligoNN configuration is higher than the complexity of the (Q-1)th three-oligoNN configuration, which is higher than the complexity of the (Q-2)th three-oligoNN configuration, and so on, with the complexity of the second three-oligoNN configuration being higher than the complexity of the first three-oligoNN configuration.
[0228] Note that in method 1880 of FIG. 18E, operation 1704P is from the last block of method 1700 of FIG. 17A, and operations 1888a1-1888am are for iteratively training a first 3-oligo NN configuration and generating labeled training data for a second 3-oligo NN configuration, operation 1888b is for iteratively training a second 3-oligo NN configuration and generating labeled training data for a third 3-oligo NN configuration, and so on. This process continues, and operation 1888Q is for training a Qth 3-oligo NN configuration and generating labeled training data for training a subsequent NN configuration. Thus, generally speaking, in method 1880, operation 1888i is for training an i-th 3-oligo NN configuration and generating labeled training data for an (i+1)-th 3-oligo NN configuration, where i=1,...,Q.
[0229] Method 1880 includes, at 1704P, repeating operations 1704b1, ..., 1704bk to train the Pth NN configuration using the 2-oligo ground truth data and generating 2-oligo labeled training data for training the next NN configuration, which is the final block of method 1700 of FIG. 17A.
[0230] Method 1880 then proceeds from 1704P to 1888a1. As illustrated, operation 1888a is for training a first 3-oligo NN configuration (e.g., 3-oligo neural network configuration 1815) using labeled training data (e.g., training data 1750 of FIG. 17B) generated from a previous block (e.g., block 1704P) and using the trained first 3-oligo NN configuration to generate further 3-oligo labeled training data for subsequent training of a second 3-oligo NN configuration. Operation 1888a includes sub-operations at blocks 1888a1-1888am.
[0231] In block 1888a1, (i) a first 3-oligo NN configuration (e.g., 3-oligo NN configuration 1815 of FIG. 18A) is trained using the labeled training data generated in 1704P, and (ii) 3-oligo labeled training data is generated using at least partially trained first 3-oligo NN configuration (e.g., training data 1850 of FIG. 18D).
[0232] Method 1880 then proceeds from 1888a1 to 1888a2. At block 1888a2, (i) a first 3-oligo NN configuration is further trained using the 3-oligo labeled training data generated in a previous step (e.g., generated at block 1888a1), and (ii) new 3-oligo labeled training data is generated using the further trained first 3-oligo NN configuration.
[0233] The operations discussed with respect to block 1888a2 (and block 1888a2) are repeated iteratively at 1888a3, ..., 1888am. Note that blocks 1888a1, ..., 1888am are all for training a first 3-oligo NN configuration. The number of iterations "m" may be implementation specific, and example criteria used to select the number of iterations for training a particular NN model are discussed with respect to method 1700 of FIG. 17A (e.g., selection of the number of iterations "k" in this method).
[0234] After the first three-oligo NN configuration has been fully or satisfactorily trained in 1888am, the method 1888 proceeds to block 1888b, where a second three-oligo NN configuration is iteratively trained. The training of the second three-oligo NN configuration also involves iterations similar to those discussed with respect to operations 1888a1, ..., 1888am, and therefore will not be discussed in further detail.
[0235] This process of progressively training more complex NN configurations continues in 1888Q of method 1888 until the Qth 3-oligo NN configuration has been trained and corresponding 3-oligo training data has been generated for training the next NN configuration.
[0236] FIG. 19 illustrates a flowchart depicting an exemplary method 1900 for iteratively training a neural network configuration for base calling using multiple oligo ground truth sequences. FIG. 19 essentially summarizes the discussion of FIG. 14A-FIG. 18E. For example, FIG. 19 illustrates an iterative training and labeled training data generation process using different oligo stages, such as a single oligo stage, a two oligo stage, a three oligo stage, etc. Thus, the complexity and / or length of the analytes used for training and generation of labeled training data increases progressively and monotonically with each iteration, along with the complexity of the neural network configuration underlying the base calling.
[0237] Method 1900 includes, at 1904a, iteratively training one oligoNN configuration to generate labeled training data, e.g., as discussed with respect to block 1704a of method 1700 of Figures 14A and 14B and 17A.
[0238] Method 1900 further includes, at 1904b, iteratively training one or more 2-oligo NN configurations using the 2-oligo sequences and generating labeled 2-oligo training data, e.g., as discussed with respect to blocks 1704b1-1704P of method 1700 of FIG. 17A.
[0239] Method 1900 further includes, at 1904c, iteratively training one or more 3-oligo NN configurations using the 3-oligo sequences and generating labeled 3-oligo training data, e.g., as discussed with respect to blocks 1888a1-1888Q of method 1880 of FIG. 18E.
[0240] This process can continue, with progressively larger numbers of oligo sequences being used. Finally, at 1904N, one or more N oligoNN configurations are trained using the N oligo sequences to generate corresponding N oligo-labeled training data, where N can be any suitable positive integer greater than or equal to 2. The operations at 1904N will be apparent based on a discussion of the operations at 1904b and 1904c.
[0241] 14A-19 are associated with training NN models using synthetically sequenced simple oligo sequences. For example, the oligo sequences used in these figures are likely to have a smaller number of bases compared to sequences found in the DNA of an organism. In an embodiment, the oligo-based training discussed with respect to FIG. 14A-19 is used to train progressively more complex NN models to generate progressively richer labeled training data sets. For example, FIG. 19 uses an N oligo NN configuration to output an N oligo labeled training data set, where the N oligo labeled training data set may have a much richer, more diverse, and larger labeled training data set than a labeled training data set associated with a number of oligos "less than N".
[0242] In practice, however, the sequencing machine 1404 and base caller 1414 will base call sequences that are much more complex than simple oligo sequences. For example, in practice, the sequencing machine 1404 and base caller 1414 will base call biological sequences that are much more complex than simple oligo sequences. Thus, the base caller 1414 must be trained on base sequences found in biological DNA and RNA that are more complex than oligo sequences.
[0243] FIG. 20A illustrates an organism sequence 2000 used to train the base chorus 1414 of FIG. 14A. The organism sequence can be an organism sequence with a relatively small number of bases, such as phix (also called phiX). The phix bacteriophage is a single-stranded DNA (ssDNA) virus. The phix174 bacteriophage is an ssDNA virus that infects E. coli and was the first DNA-based genome sequenced in 1977. Phix (such as ΦX174) virus particles have also been successfully assembled in vitro. In an embodiment, after training the base chorus 1414 with an oligo sequence (as discussed with respect to FIGS. 14A-19), the base chorus 1414 can be further trained with simple organism DNA, such as phix DNA, although this does not limit the scope of the present disclosure. For example, a more complex organism, such as a bacterium (such as E. coli or E-coli), can be used in place of phix. Thus, biological sequence 2000 can be phix, or another relatively simple biological DNA. Biological sequence 2000 is pre-sequenced, i.e., the base sequence of biological sequence 2000 is known a priori (e.g., sequenced by a sequencing machine and a previously trained base collaborator different from that illustrated in FIG. 14A).
[0244] As illustrated in FIG. 20A, when the biological sequence 2000 is loaded into the sequencing machine 1404 of FIG. 14A, the biological sequence 2000 is divided or partitioned into a number of subsequences 2004a, 2004b, ..., 2004N. Each subsequence is loaded into one or more corresponding clusters. Thus, each cluster 1407 is populated with the corresponding subsequence 2004 and its synthetic copy. The biological sequence 2000 can be partitioned using any suitable criteria, such as, for example, the maximum size of a subsequence that can be populated into a cluster. For example, if each cluster of a flow cell can be populated with subsequences having a maximum of about 150 bases, then each of the subsequences 2004 can be partitioned accordingly such that each has a maximum of 150 bases. In an embodiment, each of the subsequences 2004 can have a substantially equal number of bases, while in another embodiment, each of the subsequences 2004 can have a different number of bases. Subsequence 2004b is used as an example to discuss the teachings of the present disclosure and is assumed to have L1 bases. As an example only, the number L1 can be between 100 and 200, but can have any other suitable value and is implementation specific.
[0245] FIG. 20B illustrates the base calling system 1400 of FIG. 14A operating in the training data generation phase of the first organism training stage to train a base calling system 1414 comprising a first organism-level neural network configuration 2015 using subsequences 2004a, ..., 2004S of the first organism sequence 2000 of FIG. 20A.
[0246] Although not illustrated in Figure 20B, note that the first organism-level NN configuration 2015 is first trained using the N oligo-labeled training data from method 1904 of Figure 19. Thus, the first organism-level NN configuration 2015 is at least partially pre-trained. The base calling system 1400 of Figure 20B is the same as the base calling system of Figure 14A, although in the two figures, the base calling system 1400 uses different neural network configurations and different analytes.
[0247] As described above, the subsequences 2004a, ..., 2004S are loaded into corresponding clusters 1407. For example, subsequence 2004a is loaded into cluster 1407a, subsequence 2004b is loaded into cluster 1407b, etc. Note that each cluster 1407 contains multiple sequenced copies of the same subsequence 2004. For example, the subsequences loaded into a cluster are synthetically replicated such that the cluster has multiple copies of the same subsequence, which serves to generate a corresponding sequence signal 2012 for the cluster.
[0248] Note that the base caller 1414 does not know which subsequences are populated into which clusters. For example, if subsequence 2004a and a synthetic copy of it are loaded into a particular cluster, the base caller 1414 does not know which cluster subsequence 2004a is populated into. As described later in this specification, the mapping logic 1416 is intended to map each subsequence 2004 into a corresponding cluster 1407 to facilitate the training process.
[0249] The sequencing machine 1404 generates sequence signals 2012a, ..., 2012G for corresponding clusters of the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, the sequencing machine 1404 generates a corresponding sequence signal 2012a indicative of the bases of cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the sequencing machine 1404 generates a corresponding sequence signal 2012b indicative of the bases for cluster 1407b for a series of sequencing cycles, and so on.
[0250] In an embodiment, each subsequence 2004 is loaded into a corresponding cluster 1407, but the base caller 1414 does not know which subsequence is loaded into which cluster. Thus, the base caller 1414 does not know the mapping between the subsequences 2004 and the clusters 1407. Because each cluster 1407 generates a corresponding array signal 2012, the base caller 1414 does not know the mapping between the subsequences 2004 and the array signals 2012.
[0251] A base caller 1414 including a neural network configuration 2015 predicts base call sequences 2018a, ..., 2018G for corresponding clusters of a plurality of clusters 1407a, ..., 1407G based on corresponding sequence signals 2012a, ..., 2012G, respectively. For example, for cluster 1407a, the base caller 1414 predicts a corresponding base call sequence 2018a including a base call for cluster 1407a for a set of sequencing cycles based on corresponding sequence signals 2012a. Similarly, for cluster 1407b, the base caller 1414 predicts a corresponding base call sequence 2018b including a base call for cluster 1407b for a set of sequencing cycles based on corresponding sequence signals 2012b, and so on.
[0252] Note that the neural network configuration 2015 is only partially trained, and not fully trained, and therefore, it may be the case that the neural network configuration 2015 fails to accurately predict some or most of the bases of each subsequence.
[0253] Furthermore, as base calling on partial sequences progresses, it becomes increasingly difficult to call bases due to, for example, phasing or pre-phasing fading and / or noise. FIG. 20C illustrates an example of fading where the signal intensity decreases as a function of cycle number in a sequencing run of a base calling operation. Fading is the exponential decay of the fluorescent signal intensity of a cluster as a function of cycle number. As a sequencing run progresses, the specimen strands are washed excessively, exposed to laser emissions that create reactive species, and subjected to harsh environmental conditions. All of this results in the gradual loss of fragments in each specimen, reducing its fluorescent signal intensity. Fading is also referred to as extinction or signal decay. FIG. 20C illustrates an example of fading 2000C. In FIG. 20C, the intensity values of specimen fragments with AC microsatellites show exponential decay.
[0254] FIG. 20D conceptually shows the decreasing signal-to-noise ratio as the cycle of sequencing progresses. For example, as sequencing progresses, the signal intensity decreases and the noise increases, resulting in a substantial decrease in the signal-to-noise ratio, making accurate base calling more difficult. Physically, it has been observed that later synthesis steps attach tags to the sensor at different positions than earlier synthesis steps. When the sensor is below the sequence being synthesized, the signal decay comes from attaching tags to strands further away from the sensor in later sequencing steps than in earlier steps. This causes signal decay as the sequencing cycle progresses. In some designs, when the sensor is above the substrate holding the cluster, the signal can increase as sequencing progresses instead of decreasing.
[0255] In the investigated flow cell designs, noise increases while the signal decays. Physically, phasing and prephasing increase noise as sequencing progresses. Phasing refers to a step in sequencing where the tag cannot advance along the sequence. Prephasing refers to a sequencing step where the tag jumps forward by two positions instead of one position during a sequencing cycle. Phasing and prephasing are both relatively infrequent, occurring on the order of once in 500-1000 cycles. Phasing is slightly more frequent than prephasing. Phasing and prephasing affect individual strands within a cluster that generate intensity data, so the intensity noise distribution from the cluster accumulates in binomial, trinomial, quaternary, etc., expansions as sequencing progresses.
[0256] Further details of fading, signal attenuation, and reduction in signal-to-noise ratio, as well as Figures 20C and 20D, can be found in U.S. Non-Provisional Patent Application No. 16 / 874,599, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 14, 2020 (Attorney Docket No. ILLM1011-4 / IP-1750-US), which is incorporated by reference as if fully set forth herein.
[0257] Thus, during base calling, the reliability or predictability of the base calling decreases as the sequencing cycles progress. For example, with reference to a particular subsequence, such as subsequence 2004b in FIG. 20A, calling bases 1-10 of subsequence 2004b may generally be more reliable than calling bases 10-20 or calling bases 50-60. In other words, the first few bases of the L1 bases of subsequence 2004b are relatively more likely to be predicted accurately than the remaining bases of the L1 bases of subsequence 2004b.
[0258] FIG. 20E illustrates base calling of the first L2 bases of the L1 bases of a subsequence, where the first L2 bases of subsequence 2004b are used to map subsequence 2004b to sequence 2000.
[0259] For example, referring to Figures 20A, 20B, and 20E, sequence determination machine 1404 generates sequence signal 2012b corresponding to partial sequence 2004b (i.e., assume that partial sequence 2004b has been populated into cluster 1407b). However, base caller 1414 does not know where the partial sequence corresponding to sequence signal 2012b fits into sequence 2000. That is, base caller 1414 does not know that partial sequence 2004b has been loaded into cluster 1407b in particular.
[0260] As illustrated in Figure 20E, a partially trained NN configuration 2015 (e.g., trained using N oligo-labeled training data from method 1904 of Figure 19) receives sequence signal 2012b and predicts L1 bases indicated by sequence signal 2012b. The prediction of the L1 bases includes a prediction of the first L2 bases, and the prediction of the first L2 bases of subsequence 2004b is used to map subsequence 2004b to sequence 2000.
[0261] In an embodiment, the number L2 is 10. The number L2 can be any suitable number, such as 8, 10, 12, 13, or the like, so long as L2 is relatively smaller than L1. For example, L2 is less than 10% of L1, less than 25% of L1, or the like.
[0262] For example, the first L2 bases of subsequence 2004b predicted by NN configuration 2015 are A, C, C, T, G, A, G, C, G, A, as illustrated in Figure 20E. The remining (L1-L2) base predictions are generally illustrated as B1, ..., B1 in Figure 20E.
[0263] Here, the NN configuration 2015 may have correctly predicted the first L2 bases, or there may be one or more errors in these L2 base predictions. The mapping logic 1416 attempts to map the predictions of the first L2 bases to the corresponding consecutive L2 bases in the biological sequence 2000. In other words, the mapping logic 1416 attempts to match the predictions of the first L2 bases to the consecutive L2 bases in the biological sequence 2000 so that the subsequence 2004b in the biological sequence 2000 can be identified.
[0264] As illustrated in FIG. 20E, the mapping logic 1416 can find a "substantial" and "unique" match between the first L2 bases predicted for the subsequence 2004b and the consecutive L2 bases in the biological sequence 2000. Note that a "substantial" match means that the match may not be 100% and there may be one or more errors in the match. For example, the first L2 bases of the subsequence 2004b predicted by the NN configuration 2015 are A, C, C, T, G, A, G, C, G, A, while the corresponding substantially matching consecutive L2 bases of the biological sequence 2000 are A, G, C, T, G, A, G, C, G, A. Thus, the second base of these two L2 base sequences does not match, but the remaining bases match. As long as the number of such mismatches is below a threshold percentage, the mapping logic 1416 declares the two L2 base fragments to match. The threshold percentage of mismatch can be 10%, or 20%, or a similar percentage of the number L2. Thus, in an embodiment, L2 is 10, and the matching logic 1416 can tolerate up to 2 mismatches (or 20% mismatches). Thus, the mapping logic 1416 aims to map the first L2 bases predicted for the subsequence 2004b, or a small variation thereof (e.g., the variation means error tolerance during matching), to the consecutive L2 bases of the biological sequence 2000. The value of the threshold percentage can be implementation specific and user configurable. Just as an example, during the first iteration of training, the threshold percentage can be relatively high (e.g., 20%), and the threshold percentage can have a relatively low value (e.g., 10%) during later iterations of training. Thus, in the early stages of training iterations, the threshold percentage can be relatively high because the probability of error in base calling prediction is relatively high. As the NN configuration becomes better trained, it becomes more likely to make better base-calling predictions, and so the threshold percentage can be gradually lowered, however, in another embodiment, the threshold percentage can remain the same throughout all iterations of training.
[0265] Also, in an embodiment, a match between two L2 bases must be unique for proper mapping; a non-unique match may result in the match and mapping being declared inconclusive. Thus, the first L2 bases predicted for subsequence 2004b (or a small variation thereof) can only occur once in organism sequence 2000 for the match and mapping to be valid. Typically, in practical base sequences of simpler organisms, consecutive L2 bases (or small variations thereof) are likely to occur only once in organism sequence 2000.
[0266] For example, referring to the example of FIG. 20E, if there is an occurrence of consecutive bases A, G, C, T, G, A, G, C, G, A in one section of the biological sequence 2000 and another occurrence of consecutive bases A, C, A, T, G, A, G, C, G, A in another section of the biological sequence 2000, both sections of the biological sequence 2000 may match the first L2 bases (A, C, C, T, G, A, G, C, G, A) of the subsequence 2004b predicted by the NN configuration 2015. Thus, in this example, the match is not unique and the mapping logic 1416 does not know which of the two sections of the biological sequence 2000 maps to the L2 bases on the subsequence 2004b. In such a scenario, the mapping logic 1416 declares that there is no reliable match (i.e., declares an indeterminate mapping).
[0267] Referring to the example of FIG. 20E, as illustrated, the first L2 bases of the subsequence 2004b predicted by the NN configuration 2015 match "substantially" and "uniquely" with the corresponding consecutive L2 bases of the biological sequence 2000. Also, given a section 2000B (having L1 bases) of the biological sequence 2000, the first L2 prediction of the subsequence 2004b matches "substantially" and "uniquely" with the first L2 bases of the section B of the biological sequence 2000. Thus, the subsequence 2004b is most likely actually the section 2000B of the biological sequence 2000. In other words, the section 2000B of the biological sequence 2000 was most likely partitioned in FIG. 20A to form the subsequence 2004b.
[0268] Thus, section 2000B of biological sequence 2000 serves as ground truth for sequence signal 2012b corresponding to subsequence 2004b. Figure 20F illustrates labeled training data 2050 generated from the mapping of Figure 20E, where the labeled training data 2050 includes the section of biological sequence 2000 of Figure 20A as ground truth.
[0269] In the labeled training data 2050 of FIG. 20F, by way of example only, the subsequences 2004a, 2004d do not map to any section of the biological sequence 2000 due to indeterminate mapping. For example, as discussed with respect to FIG. 20E, in order for the mapping logic 1416 to declare a final mapping, there must be a substantial and unique match between the first L2 bases of the subsequence and the corresponding section of the biological sequence 2000. The NN configuration 2015 may make a relatively large number of errors in the first L2 bases of each of the subsequences 2004a, 2004d, and as a result, these subsequences may not be able to map to any corresponding section of the biological sequence 2000.
[0270] In the labeled training data 2050 of Figure 20F, subsequence 2004b (and thus sequence signal 2012b) is mapped to section 2000B of biological sequence 2000, as discussed with respect to Figure 20E. Similarly, subsequence 2004c is mapped to section 2000C of biological sequence 2000, and subsequence 2004S is mapped to section 2000S of biological sequence 2000. For example, subsequence 2004c is mapped to section 2000C of biological sequence 2000 (e.g., having the same number of bases as subsequence 2004c) such that the first L2 base predictions of subsequence 2004c "substantially" and "uniquely" match the first L2 bases of section 2000C.
[0271] Figure 20G illustrates the basecalling system 1400 of Figure 14A operating in the "training data consumption and training phase" of the "organism-level training stage" to train a basecaller 1414 comprising a first organism-level neural network configuration 2015. For example, the labeled training data 2050 of Figure 20F is used in the training of Figure 20G.
[0272] For example, L1 bases of the partial sequence 2004b predicted by the base caller 1414 are compared to the section 2000B of the biological sequence 2000. Note that the L1 bases of the partial sequence 2004b predicted by the base caller 1414 have the first L2 bases compared to the biological sequence 2000 to generate the mapping of FIG. 20F. The remaining (L1-L2) bases were not compared while generating the mapping of FIG. 20F because they are likely to contain many errors. This is because bases occurring later in the partial sequence are more likely to be incorrectly predicted due to fading, phasing, and / or pre-phasing, as discussed with respect to FIG. 20C and FIG. 20D. In FIG. 20G, all L1 bases of the partial sequence 2004b predicted by the base caller 1414 are compared to the corresponding L1 bases on the section 2000B of the biological sequence 2000.
[0273] Thus, the mapping of Figure 20F identifies the portion of the biological sequence 2000 (i.e., section 2000B) to which subsequence 2004b is compared in Figure 20G. Once the mapping is complete and labeled training data 2050 is generated, the labeled training data 2050 is used in Figure 20G for comparison and generation of an error signal, which is used for gradient update 2017 in a backward pass of the NN configuration 2015 and training of the resulting NN configuration 2015.
[0274] Note that some of the subsequences (such as subsequences 2004a and 2004d, see Figure 20F) do not ultimately match the corresponding sections of biological sequence 2000, and therefore the base call predictions corresponding to these subsequences are not used in the training of Figure 20G.
[0275] FIG. 21 illustrates a flow chart depicting an exemplary method 2100 for iteratively training a neural network configuration for base calling using the simple organism sequence 2000 of FIG. 20A. The method 2100 progressively trains a NN configuration that is monotonically complex in nature. As previously described herein, increasing the complexity of the NN configuration may include increasing the number of layers of the NN configuration, increasing the number of filters of the NN configuration, increasing the topological complexity in the NN configuration, and / or the like. For example, the method 2100 references a first organism-level NN configuration (which is the NN configuration 2015 previously discussed herein with respect to FIG. 20B, FIG. 20G and other figures), a second organism-level NN configuration, an Rth organism-level NN configuration, etc. In an embodiment, the complexity of the Rth organism-level NN configuration is higher than the complexity of the (R-1)th organism-level NN configuration, which is higher than the complexity of the (R-2)th organism-level NN configuration, etc., the complexity of the second organism-level NN configuration is higher than the complexity of the first organism-level NN configuration.
[0276] It should be noted that in method 2100, operation 2104a (including blocks 2104a1, ..., 2104am) is for training a first organism-level NN configuration and generating labeled training data for a second organism-level NN configuration, operation 2104b is for training a second organism-level NN configuration and generating labeled training data for a third organism-level NN configuration, etc. This process continues, and finally operation 2104R is for training an Rth organism-level NN configuration and generating labeled training data for a next stage NN configuration. Thus, generally speaking, in method 2100, operation 2104i is for training an ith organism-level NN configuration and generating labeled training data for an (i+1)th organism-level NN configuration, where i=1, ..., R.
[0277] Method 2100 includes, at 2104a1, (i) training a first organism-level NN configuration (e.g., organism-level NN configuration 2015 of FIG. 20B, although training of this NN configuration is not illustrated in FIG. 20B) using the N oligo-labeled training data from 1904N of method 1900 of FIG. 19, and (ii) generating labeled training data using the at least partially trained first organism-level NN configuration 2015. The labeled training data is illustrated in FIG. 20F, and its generation is discussed with respect to FIG. 20E and FIG. 20F.
[0278] Method 2100 then proceeds from 2104a1 to 2104a2, during which a second iteration of training a first organism-level NN configuration 2015 is performed. For example, at 2104a2, (i) the first organism-level NN configuration 2015 is further trained using labeled training data from a previous stage, e.g., as discussed with respect to FIG. 20G, and (ii) further labeled training data is generated using the at least partially trained first organism-level NN configuration 2015 (e.g., similar to the discussion with respect to FIG. 20E and FIG. 20F).
[0279] The training and generation operations are repeated iteratively, eventually completing the training of the first organism-level NN configuration 2015 at 2104am. Note that block 2014a1 is the first iteration of training the first organism-level NN configuration 2015, block 2104a2 is the second iteration of training the first organism-level NN configuration 2015, and so on, until finally block 2104am is the mth iteration of training the first organism-level NN configuration 2015. The number of iterations may be based on one or more factors, such as those previously discussed herein with respect to method 1700 of FIG. 17A (e.g., when the criteria for selecting the number of iterations "k" were discussed). The complexity of the first organism-level NN configuration 2015 does not change during the iterations of 2104a1, ..., 2104am.
[0280] At the end of the iterations for the first creature-level NN configuration 2015 (i.e., at the end of block 2104am), the method 2100 proceeds to block 2104b, where a second creature-level NN configuration is iteratively trained. The training of the second creature-level NN configuration and the generation of associated training labeled data will also involve iterations similar to those discussed with respect to operations 2104a1,...,2104am, and therefore will not be discussed in further detail.
[0281] This process of progressively training more complex NN configurations associated with the generation of training labeled data continues until, at 2104R of method 2100, the Rth organism-level NN configuration has been trained and corresponding labeled training data has been generated for training the next NN configuration.
[0282] FIG. 22 illustrates the use of complex biological sequences for training corresponding NN configurations for the base colloquium 1414 of FIG. 14A. For example, as discussed with respect to FIG. 20A-FIG. 21, relatively simple biological sequences 2000 containing approximately L1 bases per subsequence are used to iteratively train R simple biological level NN configurations to generate corresponding labeled training data. For example, method 2100 of FIG. 21 illustrates such iterative learning and generation of labeled training data using simple biological sequences 2000. As discussed, simple biological sequences 2000 can be Phix or another organism with a relatively simple (or relatively small) gene sequence.
[0283] 22 also illustrates the use of a relatively complex biological sequence 2200a. The biological sequence 2200a is more complex than the biological sequence 2000, for example, because the number of bases in the composite biological sequence 2200a is greater than the number of bases in the biological sequence 2000. By way of example only, the biological sequence 2000 may have approximately 1 million bases, and the composite biological sequence 2200a may have 4 million bases. In another embodiment, each subsequence partitioned from the composite biological sequence 2200a has a greater number of bases than each subsequence partitioned from the biological sequence 2000. In yet another embodiment, the number of subsequences partitioned from the composite biological sequence 2200a is greater than the number of subsequences partitioned from the biological sequence 2000. For example, when partitioning composite biological sequence 2200a and biological sequence 2000, the number of subsequences partitioned from composite biological sequence 2200a will be greater than the number of subsequences partitioned from biological sequence 2000 because (i) composite biological sequence 2200a has a greater number of bases than biological sequence 2000, and (ii) each subsequence may have at most a threshold number of bases. In an embodiment, composite biological sequence 2200a includes genetic material from bacteria, such as E-coli, or other suitable biological sequence that is more complex than biological sequence 2000.
[0284] As illustrated in Fig. 22, the composite biological sequence 2200a is used to iteratively train Ra composite biological level NN configurations and generate labeled training data. The training and generation of labeled training data is similar to that discussed with respect to the method 2100 of Fig. 21 (the difference is that the method 2100 is specifically directed to the biological sequence 2000, whereas the composite biological sequence 2200a is used here).
[0285] This iterative process continues, and finally, a relatively more complex biological sequence 2200T is used. The more complex biological sequence 2200T is more complex than the biological sequences 2000 and 2200a. For example, the number of bases in the more complex biological sequence 2200T is greater than the number of bases in each of the biological sequences 2000 and 2200a. In another embodiment, each subsequence partitioned from the more complex biological sequence 2200T has a greater number of bases than each subsequence partitioned from the biological sequences 2000 or 2200a. In yet another embodiment, the number of subsequences partitioned from the more complex biological sequence 2200T is greater than the number of subsequences partitioned from the biological sequences 2000 or 2200a. In an embodiment, the more complex biological sequence 2200T includes genetic material from a complex species, such as genetic material from a human or other mammal.
[0286] As illustrated in Figure 22, the biological sequence 2200T is used to iteratively train R T more complex biological level NN configurations and generate labeled training data. The training and generation of labeled training data is similar to that discussed with respect to method 2100 of Figure 21 (the difference is that method 2100 is specifically directed to biological sequence 2000, whereas here biological sequence 2000T is used).
[0287] FIG. 23A illustrates a flowchart depicting an exemplary method 2300 for iteratively training neural network configurations for base calling. Method 2300 summarizes at least some of the embodiments and examples discussed herein with respect to FIGS. 14A-22. Method 2300 incrementally trains NN configurations that are monotonically complex in nature, as discussed herein. Method 2300 also uses monotonically complex gene sequences as exemplars. Method 2300 is used to train base callers 1414 of various figures discussed herein.
[0288] Method 2300 begins at 2304, where a base chore 1414 including a NN configuration 1415 (see, e.g., FIG. 14A) is iteratively trained using single-oligo ground truth data, as discussed with respect to block 1704 of method 1700 of FIG. 17A. The at least partially trained NN configuration 1415 of FIG. 14A is used to generate labeled training data, as also discussed with respect to block 1704 of method 1700 of FIG. 17A.
[0289] Next, method 2300 proceeds from 2304 to 2308, where one or more NN configurations are iteratively trained using the two oligo sequences and corresponding labeled training data is generated, for example, as discussed with respect to method 1700 of FIG. 17A.
[0290] Next, method 2300 proceeds from 2308 to 2312, where one or more NN configurations are iteratively trained using the three oligo sequences and corresponding labeled training data is generated, for example, as discussed with respect to method 1900 of FIG. 19.
[0291] In 2316, one or more NN configurations are iteratively trained using N oligo sequences, and this process of training the NN configurations using increasing numbers of oligos continues until corresponding labeled training data is generated, e.g., as discussed with respect to method 1900 of FIG. 19 .
[0292] Method 2300 then proceeds to 2320, where training and generating labeled training data involves an organism. At 2320, a simple organism sequence, such as simple organism sequence 2000 of FIG. 20A, is used. One or more NN configurations are trained using the simple organism sequence (see, e.g., method 2100 of FIG. 21) to generate labeled training data.
[0293] As the method 2300 proceeds from 2320, increasingly complex biological sequences are used, e.g., as discussed with respect to Figure 22. Finally, at 2328, one or more NN configurations are iteratively trained using a complex biological sequence (e.g., the further complex biological sequence 2200T of Figure 22) to generate corresponding labeled training data.
[0294] Thus, method 2300 continues until base caller 1414 is "sufficiently trained." "Sufficiently trained" may imply that base caller 1414 can now make base calls with an error rate less than the target error rate. As discussed, the training process can be iteratively continued until sufficient training and a target error rate of base calling is achieved (see, e.g., the "Error Rate" chart in FIG. 23E). At the end of method 2300, base caller 1414 including the final NN configuration of method 2300 is now fully trained. Thus, the trained base caller 1414 including the final NN configuration of method 2300 can now be used for inference, e.g., to sequence unknown gene sequences.
[0295] 23B-23E illustrate various charts illustrating the effectiveness of the base collation training process discussed in this disclosure. Referring to FIG. 23B, illustrated is a chart 2360 depicting the mapping percentage of training data generated by (i) a first two-oligo NN configuration, such as NN configuration 1615, trained using the neural network-based training data generation technique discussed herein, and (ii) a NN configuration trained using a conventional two-oligo training data generation technique. The white bars in the chart 2360 illustrate mapping data from a first two-oligo NN configuration trained using training data generated using a neural network-based model discussed herein. Thus, the white bars in the chart 2360 illustrate mapping data generated using the various techniques discussed herein. The gray bars in the chart 2360 illustrate data associated with a NN configuration trained with training data generated by a conventional non-neural network-based model, such as a Real Time Analysis (RTA) model. Examples of RTA models are discussed in U.S. Patent No. US10304189(B2), entitled "Data processing system and methods," issued May 28, 2019, which is incorporated by reference as if fully set forth herein. Thus, the grey bars in chart 2360 illustrate mapping data generated using conventional techniques. In an example, the white bars in chart 2360 can be generated in operation 1704b1 of method 1700 of FIG. 17A. Chart 2360 illustrates the percentage of base call predictions that are mapped to oligo 1, the percentage of base call predictions that are mapped to oligo 2, and the percentage of base call predictions that cannot ultimately be mapped to either oligo 1 or 2 (i.e., the indeterminate percentage). As can be seen, the indeterminate percentage of training data generated using the techniques discussed herein is slightly higher than the indeterminate percentage of training data generated using conventional techniques.Thus, initially (eg, at the beginning of the training iterations), the conventional techniques slightly outperform the training data generation techniques discussed herein.
[0296] 23C, illustrated is a chart 2365 depicting mapping percentages in training data generated using (i) a first two-oligo NN configuration (such as NN configuration 1615) trained using the neural network-based training data generation techniques discussed herein (white bars), (ii) a second two-oligo NN configuration trained using the neural network-based training data generation techniques discussed herein (dotted bars), and (iii) a NN configuration trained using a conventional two-oligo training data generation technique, such as an RTA-based conventional training data generation technique (gray bars). In an example, the first two-oligo NN configuration (white bars) and the second two-oligo NN configuration (dotted bars) correspond to operations 1704b and 1704c, respectively, of method 1700 of FIG. 17A. Chart 2365 illustrates the percentage of base call predictions that are mapped to oligo 1, the percentage of base call predictions that are mapped to oligo 2, and the percentage of base call predictions that cannot be conclusively mapped to either oligo 1 or 2 (i.e., the percentage of uncertain). As can be seen, the percentage of uncertain for the training data generated using the first two-oligo NN configuration is higher than (i) the training data generated using the second two-oligo NN configuration and (ii) each of the training data generated using the conventional technique. Furthermore, the percentage of uncertain for the training data generated using the second two-oligo NN configuration is approximately comparable to the training data generated using the conventional technique. Thus, with iterations and more complex NN configurations, the training data generated using the NN-based configuration is approximately comparable to the training data generated using the conventional technique.
[0297] Referring now to Figure 23D, a chart 2370 is illustrated depicting the mapping percentage of training data generated by (i) a first 4-oligo NN configuration trained using the neural network-based training data generation technique discussed herein (white bars), and (ii) a NN configuration trained using a conventional 4-oligo training data generation technique, e.g., an RTA-based technique (gray bars). As can be seen, the uncertain percentage of the training data generated using the technique discussed herein is comparable to the uncertain percentage of the training data generated using the conventional technique. Thus, when training is shifted to 4-oligo sequences, the conventional technique and the training data generation technique discussed herein produce comparable results.
[0298] Now referring to FIG. 23E, there is illustrated a chart 2375 depicting the error rate of data generated by (i) a NN configuration trained using the complex biological sequences discussed herein, for example with respect to operation 2328 of method 2300 of FIG. 23A (solid line), and (ii) a NN configuration trained using a conventional complex biological training data generation technique, for example, an RTA-based technique (dashed line). As can be seen, the error rate of the data generated using the techniques discussed herein is comparable to the data generated using the conventional techniques. Thus, the conventional techniques and the training data generation techniques discussed herein produce comparable results. As discussed, the training data generation techniques discussed herein may be used in place of conventional techniques, for example, when conventional techniques are not available or ready for training data generation.
[0299] 24 is a block diagram of a base calling system 2400 according to one implementation. The base calling system 2400 can operate to obtain any information or data related to at least one of biological or chemical substances. In some implementations, the base calling system 2400 is a workstation, which can be similar to a benchtop device or desktop computer. For example, most (or all) of the systems and components for carrying out the desired reactions can be in a common housing 2416.
[0300] In certain implementations, the base calling system 2400 is a nucleic acid sequencing system (or sequencer) configured for a variety of applications, including, but not limited to, de novo sequencing, resequencing of whole genomes or target genomic regions, and metagenomics. Sequencers may also be used for DNA or RNA analysis. In some implementations, the base calling system 2400 may also be configured to generate reaction sites within a biosensor. For example, the base calling system 2400 may be configured to receive a sample and generate surface-attached clusters of clonally amplified nucleic acids from the sample. Each cluster may constitute or be part of a reaction site within a biosensor.
[0301] Exemplary base calling system 2400 may include a system receptacle or interface 2412 configured to interact with biosensor 2402 to effect a desired reaction within biosensor 2402. In the discussion below with respect to FIG. 24, biosensor 2402 is loaded into system receptacle 2412. However, it is understood that a cartridge including biosensor 2402 may be inserted into system receptacle 2412, and that in some conditions, the cartridge may be temporarily or permanently removed. As discussed above, the cartridge may include, among other things, fluid control and fluid storage components.
[0302] In certain implementations, the base calling system 2400 is configured to perform multiple parallel reactions within the biosensor 2402. The biosensor 2402 includes one or more reaction sites where a desired reaction can occur. The reaction sites may be immobilized, for example, on a solid surface of the biosensor or on beads (or other movable substrates) located within corresponding reaction chambers of the biosensor. The reaction sites may include, for example, clusters of clonally amplified nucleic acids. The biosensor 2402 may include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell attached thereto. The flow cell may include one or more flow channels that receive solutions from the base calling system 2400 and direct the solutions toward the reaction sites. Optionally, the biosensor 2402 may be configured to engage a thermal element for transferring thermal energy into and out of the flow channel.
[0303] Base calling system 2400 may include various components, assemblies, and systems (or subsystems) that interact with each other to perform a given method or assay protocol for biological or chemical analysis. For example, base calling system 2400 includes a system controller 2404 that may communicate with the various components, assemblies, and subsystems of base calling system 2400, and also includes biosensor 2402. For example, in addition to system receptacle 2412, base calling system 2400 may also include a fluid control system 2406 for controlling the flow of fluids throughout the fluidic network of base calling system 2400 and biosensor 2402, a fluid reservoir system 2408 configured to hold any fluids (e.g., fluids, gases, or liquids) that may be used by the bioassay system, a temperature control system 2410 that may regulate the temperature of the fluidic network, fluid reservoir system 2408, and / or fluids within biosensor 2402, and an illumination system 2409 configured to illuminate biosensor 2402. As described above, when a cartridge having a biosensor 2402 is loaded into the system receptacle 2412, the cartridge may also include fluid control and fluid storage components.
[0304] The base calling system 2400 may also include a user interface 2414 for interacting with a user. For example, the user interface 2414 may include a display 2413 for displaying or requesting information from a user, and a user input device 2415 for receiving user input. In some implementations, the display 2413 and the user input device 2415 are the same device. For example, the user interface 2414 may include a touch-sensitive display configured to detect the presence of an individual touch and to identify the location of the touch on the display. However, other user input devices 2415, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc. may be used. As described in more detail below, the base calling system 2400 may communicate with various components, including a biosensor 2402 (e.g., in the form of a cartridge), to perform the desired reaction. The base calling system 2400 may also be configured to analyze data obtained from the biosensor to provide the desired information to the user.
[0305] The system controller 2404 may include any processor-based or microprocessor-based system, including systems using microcontrollers, reduced instruction set computers (RISC), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), logic circuits, and any other circuits or processors capable of performing the functions described herein. The above examples are merely exemplary and are therefore not intended to limit the definition and / or meaning of the term system controller. In an exemplary implementation, the system controller 2404 executes a set of instructions stored in one or more storage elements, memories, or modules for at least one of acquiring and analyzing detection data. The detection data may include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) may be detected over many base calling cycles. The storage elements may be in the form of information sources or physical memory elements within the base calling system 2400.
[0306] The set of instructions may include various commands that instruct the base call system 2400 or biosensor 2402 to perform certain operations, such as the methods and processes of various implementations described herein. The set of instructions may be in the form of a software program that may form a part of a tangible non-transitory computer readable medium or medium. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in memory that is executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely exemplary and therefore are not limited to the types of memory that can be used to store a computer program.
[0307] The software may be in various forms, such as system software or application software. Furthermore, the software may be in the form of a collection of separate programs, or a program module or a portion of a program module within a larger program. The software may also include modular programming in the form of object-oriented programming. After obtaining the detection data, the detection data may be automatically processed by the base calling system 2400 processed in response to user input, or may be processed in response to a request made by another processing machine (e.g., a remote request via a communication link). In another implementation shown, the system controller 2404 includes an analysis module 2538 (shown in FIG. 25). In another implementation, the system controller 2404 does not include the analysis module 2538, but instead has access to the analysis module 2538 (e.g., the analysis module 2538 may be separately hosted on the cloud).
[0308] The system controller 2404 may be connected to the biosensor 2402 and other components of the base calling system 2400 via a communication link. The system controller 2404 may also be communicatively connected to an off-site system or server. The communication link may be a wire, a cord, or wireless. The system controller 2404 may receive user input or commands from a user interface 2414 and a user input device 2415.
[0309] The fluid control system 2406 includes a fluid network and is configured to direct the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with the biosensor 2402 and the fluid storage system 2408. For example, fluid may be selected from the fluid storage system 2408 and directed to the biosensor 2402 in a controlled manner, or fluid may be drawn from the biosensor 2402 and directed to a waste reservoir, for example, in the fluid storage system 2408. Although not shown, the fluid control system 2406 may include a flow sensor that detects the flow rate or pressure of the fluid in the fluid network. The sensor may be in communication with the system controller 2404.
[0310] The temperature control system 2410 is configured to regulate the temperature of fluids in different regions of the fluid network, the fluid reservoir system 2408, and / or the biosensor 2402. For example, the temperature control system 2410 may include a thermal cycler that interacts with the biosensor 2402 and controls the temperature of the fluid flowing along a reaction site within the biosensor 2402. The temperature control system 2410 may also regulate the temperature of solid elements or components of the base calling system 2400 or the biosensor 2402. Although not shown, the temperature control system 2410 may include sensors for detecting the temperature of the fluids or other components. The sensors may be in communication with the system controller 2404.
[0311] The fluid storage system 2408 is in fluid communication with the biosensor 2402 and may store various reaction components or reactants used to carry out a desired reaction. The fluid storage system 2408 may also store fluids for washing or cleaning the fluidic network and the biosensor 2402 and for diluting reactants. For example, the fluid storage system 2408 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, and the like. Additionally, the fluid storage system 2408 may also include a waste reservoir for receiving waste from the biosensor 2402. In implementations that include a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Thus, one or more of the components described herein with respect to these systems may be housed within the cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, waste, and the like. Thus, one or more of the fluid reservoir system, fluid control system, or temperature control system may be removably engaged with the bioassay system via a cartridge or other biosensor.
[0312] The illumination system 2409 may include a light source (e.g., one or more LEDs) and multiple optical components for illuminating the biosensor. Examples of light sources include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, and the like. In implementations using an illumination system, the illumination system 2409 may be configured to direct excitation light to the reaction sites. As an example, a fluorophore may be excited by a wavelength of green light, so the wavelength of the excitation light may be about 532 nm. In one implementation, the illumination system 2409 is configured to generate illumination parallel to a surface normal of the surface of the biosensor 2402. In another implementation, the illumination system 2409 is configured to generate illumination that is off-angled to the surface normal of the surface of the biosensor 2402. In yet another implementation, the illumination system 2409 is configured to generate illumination having multiple angles, including some parallel illumination and some off-angle illumination.
[0313] The system receptacle or interface 2412 is configured to engage the biosensor 2402 in at least one of mechanical, electrical, and fluidic ways. The system receptacle 2412 can hold the biosensor 2402 in a desired orientation to facilitate fluid flow through the biosensor 2402. The system receptacle 2412 can also include electrical contacts configured to engage the biosensor 2402 such that the base calling system 2400 can communicate with and / or power the biosensor 2402. Additionally, the system receptacle 2412 can include a fluid port (e.g., a nozzle) configured to engage the biosensor 2402. In some implementations, the biosensor 2402 is removably coupled to the system receptacle 2412 both electrically and fluidically.
[0314] In addition, the base calling system 2400 may communicate remotely with other systems or networks, or with other bioassay systems 2400. Detection data obtained by the bioassay system 2400 may be stored in a remote database.
[0315] FIG. 25 is a block diagram of a system controller 2404 that can be used in the system of FIG. 24. In one implementation, the system controller 2404 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include algorithms (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithms for performing a particular process. The system controller 2404 is conceptually illustrated as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 2404 may be implemented using an off-the-shelf PC with a single processor or multiple processors, with functional operations distributed among the processors. As a further option, the modules described below may be implemented using a hybrid configuration in which certain modular functions are performed using dedicated hardware, while the remaining modular functions are performed using off-the-shelf PCs, etc. The modules may also be implemented as software modules within a processing unit.
[0316] In operation, the communication port 2520 may transmit information (e.g., commands) to the biosensor 2402 (FIG. 24) and / or subsystems 2406, 2408, 2410 (FIG. 24). In implementations, the communication port 2520 may output multiple arrays of pixel signals. The communication port 2520 may receive user input from the user interface 2414 (FIG. 24) and transmit data or information to the user interface 2414. Data from the biosensor 2402 or subsystems 2406, 2408, 2410 may be processed in real-time by the system controller 2404 during a bioassay session. Additionally or alternatively, data may be temporarily stored in system memory during a bioassay session and processed in slower than real-time or offline operation.
[0317] As shown in FIG. 25, the system controller 2404 may include multiple modules 2531-2539 in communication with a main control module 2530. The main control module 2530 may be in communication with a user interface 2414 (FIG. 24). Although the modules 2531-2539 are shown in direct communication with the main control module 2530, the modules 2531-2539 may also be in direct communication with each other, with the user interface 2414, and with the biosensor 2402. The modules 2531-2539 may also be in communication with the main control module 2530 via other modules.
[0318] The plurality of modules 2531-2539 include system modules 2531-2533, 2539 that communicate with the subsystems 2406, 2408, 2410, and 2409, respectively. The fluid control module 2531 may communicate with the fluid control system 2406 to control valves and flow sensors of the fluid network to control the flow of one or more fluids through the fluid network. The fluid storage module 2532 may notify a user when fluid is low or when a waste reservoir is at or near full capacity. The fluid storage module 2532 may also communicate with a temperature control module 2533 so that the fluid may be stored at a desired temperature. The illumination module 2539 may communicate with the illumination system 2409 to illuminate the reaction site at a specified time during a protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some implementations, the illumination module 2539 may communicate with the illumination system 2409 to illuminate the reaction site at a specified angle.
[0319] The plurality of modules 2531-2539 may also include a device module 2534 that communicates with the biosensor 2402 and an identification module 2535 that determines identification information associated with the biosensor 2402. The device module 2534 may, for example, communicate with the system receptacle 2412 to verify that the biosensor has established electrical and fluidic connection with the base calling system 2400. The identification module 2535 may receive a signal that identifies the biosensor 2402. The identification module 2535 may use the identification information of the biosensor 2402 to provide other information to the user. For example, the identification module 2535 may determine and then display the lot number, date of manufacture, or a recommended protocol to operate with the biosensor 2402.
[0320] The plurality of modules 2531-2539 also includes an analysis module 2538 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 2402. The analysis module 2538 includes memory (e.g., RAM or flash) for storing the detection data. The detection data can include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The signal data can be stored for subsequent analysis or transmitted to the user interface 2414 to display desired information to a user. In some implementations, the signal data can be processed by a solid-state imager (e.g., a CMOS image sensor) before the analysis module 2538 receives the signal data.
[0321] The analysis module 2538 is configured to obtain image data from the photodetector in each of a plurality of sequencing cycles, the image data being derived from the luminescence signals detected by the photodetector, and process the image data for each of the plurality of sequencing cycles through a neural network (e.g., a neural network-based template generator 2548, a neural network-based base caller 2558 (see, e.g., Figures 7, 9, and 10), and / or a neural network-based quality scorer 2568) to generate base calls for at least some of the analytes in each of the plurality of sequencing cycles.
[0322] Protocol modules 2536 and 2537 communicate with main control module 2530 to control the operation of subsystems 2406, 2408, and 2410 in carrying out a predetermined assay protocol. Protocol modules 2536 and 2537 may include instruction sets for instructing base calling system 2400 to perform specific operations according to a predetermined protocol. As shown, the protocol module may be a sequencing-by-synthesis (SBS) module 2536 configured to issue various commands to carry out a sequencing-by-synthesis process. In SBS, the extension of a nucleic acid primer along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme) or ligation (e.g., catalyzed by a ligase enzyme). In certain polymer-based SBS implementations, fluorescently labeled nucleotides are added to the primer (thereby extending the primer) in a template-dependent manner such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be delivered into / through a flow cell housing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. These reaction sites can be detected where primer extension incorporates labeled nucleotides that can be detected through an imaging event. During the imaging event, an illumination system 2409 can provide excitation light to the reaction sites. Optionally, the nucleotides can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator portion can be added to the primer such that no further extension occurs until a deblocking agent is delivered to remove the portion. Thus, in another implementation using a reversible termination, a command can be given to deliver a deblocking reagent to the flow cell (before or after detection occurs).One or more commands can be given to effect washing between various delivery steps.The cycle is then repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n.Exemplary sequencing techniques are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, WO 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, and U.S. Patent No. 7,405,281, each of which is incorporated herein by reference.
[0323] In the nucleotide delivery step of the SBS cycle, any one of a single type of nucleotide can be delivered at a time, or multiple different nucleotide types (e.g., A, C, T, and G together) can be delivered. In nucleotide delivery configurations where only a single type of nucleotide is present at a time, different nucleotides do not need to have separate labels because they can be distinguished based on the temporal separation inherent to the individualized delivery. Thus, the sequencing method or device can use a single color detection. For example, the excitation source only needs to provide excitation of a single wavelength or a single wavelength range. In nucleotide delivery configurations where delivery results in multiple different nucleotides being present in the flow cell at a given time, the sites incorporating different nucleotide types can be distinguished based on the different fluorescent labels attached to each nucleotide type in the mixture. For example, four different nucleotides can be used, each with one of four different fluorophores. In one implementation, the four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, less than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to generate a range of different excitation radiation in the flow cell.
[0324] In some implementations, less than four different colors can be detected in a mixture with four different nucleotides. For example, pairs of nucleotides can be detected at the same wavelength, but can be distinguished based on the difference in intensity for one member of the pair, or based on a change to one member of the pair (e.g., via chemical modification, photochemical modification, or physical modification) that causes a distinct signal to appear or disappear compared to the signal detected for the other member of the pair. Exemplary devices and methods for distinguishing four different nucleotides using detection of less than four colors are described, for example, in U.S. Patent Application Nos. 61 / 538,294 and 61 / 619,878, which are incorporated by reference in their entirety. U.S. Patent Application No. 13 / 624,200, filed September 21, 2012, is incorporated by reference in its entirety.
[0325] The multiple protocol modules may also include a sample preparation (or generation) module 2537 configured to issue commands to the fluidic control system 2406 and the temperature control system 2410 to amplify the product in the biosensor 2402. For example, the biosensor 2402 may be engaged to the base calling system 2400. The amplification module 2537 can issue instructions to the fluidic control system 2406 to deliver the necessary amplification components to a reaction chamber in the biosensor 2402. In other implementations, the reaction site may already contain some components for amplification, such as template DNA and / or primers. After delivering the amplification components to the reaction chamber, the amplification module 2537 can instruct the temperature control system 2410 to cycle through different temperature steps according to a known amplification protocol. In some implementations, the amplification and / or incorporation of nucleotides is performed isothermally.
[0326] The SBS module 2536 can issue commands to perform a bridge PCR in which a cluster of clonal amplicons is formed over a localized region within the channel of the flow cell. After generating the amplicons via bridge PCR, the amplicons may be "linearized" to create single-stranded template DNA, and sstDNA and sequencing primers may be hybridized to universal sequences flanking the region of interest. For example, a reversible terminator-based sequencing by synthesis method may be used as described above or as follows.
[0327] Each base calling or sequencing cycle can extend the sstDNA by a single base, which can be accomplished, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. The different types of nucleotides can have unique fluorescent labels, and each nucleotide can further have a reversible terminator that allows only a single base incorporation to occur in each cycle. After the single base is added to the sstDNA, an excitation light can be incident on the reaction site and the fluorescent emission can be detected. After detection, the fluorescent label and the terminator can be chemically cleaved from the sstDNA. Another similar base calling or sequencing cycle can be as follows. In such a sequencing protocol, the SBS module 2536 can instruct the fluid control system 2406 to direct the flow of reagents and enzyme solutions through the biosensor 2402. Exemplary reversible terminator-based SBS methods that can be utilized with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007 / 0166705 (A1), U.S. Patent Application Publication No. 2006 / 0188901 (A1), U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439 (A1), U.S. Patent Application Publication No. 2006 / 02814714709 (A1), WO 05 / 065814, WO 06 / 064199, each of which is incorporated herein by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in U.S. Pat. No. 7,541,444, U.S. Pat. No. 7,057,026, U.S. Pat. No. 7,427,673, U.S. Pat. No. 7,566,537, and U.S. Pat. No. 7,592,435, each of which is incorporated herein by reference in its entirety.
[0328] In some implementations, the amplification and SBS modules may operate in a single assay protocol, for example, where template nucleic acid is amplified and subsequently sequenced within the same cartridge.
[0329] The base calling system 2400 may also allow the user to reconfigure the assay protocol. For example, the base calling system 2400 may provide the user with an option through the user interface 2414 to modify the determined protocol. For example, if it is determined that the biosensor 2402 is to be used for amplification, the base calling system 2400 may request the temperature of the annealing cycle. Additionally, the base calling system 2400 may issue a warning to the user if the user provides user input that is not generally accepted for the selected assay protocol.
[0330] In an implementation, biosensor 2402 includes a million sensors (or pixels), each of which generates a sequence of pixel signals over successive base call cycles. Analysis module 2538 detects the sequences of pixel signals and attributes them to corresponding sensors (or pixels) according to the row-wise and / or column-wise positions of the sensors on the array of sensors.
[0331] Each sensor in the array of sensors can generate sensor data for a tile of a flow cell, where the tile is in an area on the flow cell where a cluster of genetic material is placed during a base calling operation. The sensor data can include image data in an array of pixels. For a given cycle, the sensor data can include two or more images, generating multiple features per pixel as tile data.
[0332] 26 is a simplified block diagram of a computer 2600 system that can be used to implement the disclosed techniques. The computer system 2600 includes at least one central processing unit (CPU) 2672 that communicates with several peripheral devices via a bus subsystem 2655. These peripheral devices can include, for example, a storage subsystem 2610 including memory devices and a file storage subsystem 2636, a user interface input device 2638, a user interface output device 2676, and a network interface subsystem 2674. The input and output devices allow user interaction with the computer system 2600. The network interface subsystem 2674 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0333] The user interface input devices 2638 can include pointing devices such as a keyboard, a mouse, a trackball, a touch pad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into the computer system 2600.
[0334] The user interface output devices 2676 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 2600 to a user or to another machine or computer system.
[0335] The storage subsystem 2610 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the deep learning processor 2678.
[0336] In one implementation, the neural network is implemented using a deep learning processor 2678, which may be a configurable and reconfigurable processor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA) and a graphics processing unit (GPU) or other configured device. The deep learning processor 2678 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of deep learning processors 14978 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX149 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu's DPI™, ARM's DynamicIQ™, IBM's TrueNorth™, and others.
[0337] The memory subsystem 2622 used in the storage subsystem 2610 may include several memories including a main random access memory (RAM) 2634 for storing instructions and data during program execution, and a read only memory (ROM) 2632 in which fixed instructions are stored. The file storage subsystem 2636 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of an embodiment may be stored by the file storage subsystem 2636 in the storage subsystem 2610 or in another machine accessible by the processor.
[0338] The bus subsystem 2655 provides a mechanism for allowing the various components and subsystems of the computer system 2600 to communicate with each other as intended. Although the bus subsystem 2655 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0339] The computer system 2600 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 2600 shown in Figure 26 is intended only as a specific example for purposes of illustrating a preferred implementation of the invention. Many other configurations of computer system 2600 can have more or fewer components than the computer system shown in Figure 26.
[0340] The present inventors disclose the following items:
[0341] section Node set #1 (self-learning base cola trained using oligo sequences). 1. A computer-implemented method for progressively training a base colleague, comprising: iteratively initially training a base caller with samples containing a single oligonucleotide sequence and generating labeled training data using the initially trained base caller; (i) further training the base caller using a sample including a multi-oligonucleotide sequence, and generating labeled training data using the further trained base caller; and further training the base collaborator by repeating step (i) while increasing the complexity of the neural network configuration loaded into the base collaborator during at least one iteration, wherein the labeled training data generated during an iteration is used to train the base collaborator during an immediately subsequent iteration. 1a. Further training the base call with an exemplar containing the multi-oligonucleotide sequence during at least one iteration, further comprising increasing the number of unique oligonucleotide sequences of the multi-oligonucleotide sequence within the exemplar; Method of section 1. 2. First, iteratively train the base classifier with samples containing a single oligonucleotide sequence. During the first iteration of the first training of the base cola, Injecting a single known oligonucleotide sequence into multiple clusters of a flow cell; generating a plurality of sequence signals corresponding to the plurality of clusters, each sequence signal of the plurality of sequence signals representing a base sequence loaded into a corresponding cluster of the plurality of clusters; predicting a corresponding base call for the known single oligonucleotide sequence based on each sequence signal of the plurality of sequence signals, thereby generating a plurality of predicted base calls; generating, for each sequence signal of the plurality of sequence signals, a corresponding error signal based on a comparison of (i) the corresponding predicted base call and (ii) the bases of the known single oligonucleotide sequence, thereby generating a plurality of error signals corresponding to the plurality of sequence signals; and initially training a base colleague during a first iteration based on the plurality of error signals. 2a. Training the base colleague first during the first iteration 3. The method of claim 2, comprising: updating weights and / or biases of the neural network configuration based on the plurality of error signals using a backpropagation path of the neural network configuration loaded into the base colleague. 3. First, iteratively train the base classifier with samples containing a single oligonucleotide sequence. During the second iteration of the first training of the base cola, which takes place after the first iteration of the first training, predicting corresponding further base calls for the known single oligonucleotide sequence based on each sequence signal of the plurality of sequence signals using the base calls partially trained during the first iteration of the initial training, thereby generating a plurality of further predicted base calls; generating, for each sequence signal of the plurality of sequence signals, a corresponding further error signal based on a comparison of (i) the corresponding further predicted base call and (ii) the bases of the known single oligonucleotide sequence, thereby generating a plurality of further error signals corresponding to the plurality of sequence signals; The method of clause 2, further comprising: further initially training the base colleague during a second iteration based on a plurality of further error signals. 4. First, iteratively train the base classifier with samples containing a single oligonucleotide sequence. The method of clause 3, further comprising: repeating the second iteration of the initial training of the base collaborator with exemplars containing a single oligonucleotide sequence for the plurality of instances until a convergence condition is met. 5. The method of clause 4, where the convergence condition is met if, during two successive iterations of the second iteration of the initial training of the base colleague, the reduction in the error signal by more than a threshold value. 6. The method of section 4, where the convergence condition is met if the second iteration of the initial training of the base collab is repeated for at least a threshold number of instances. 7. The sequence signals corresponding to the clusters generated during the first iteration of the initial training of the base collaborator are reused for the second iteration of the initial training of the base collaborator; Method of section 3. 8. Comparing (i) the corresponding predicted base calls with (ii) the bases of a known single oligo sequence; 3. The method of claim 2, comprising, for a first predicted base call, (i) comparing a first base of the first predicted base call to a first base of the known single oligo sequence, and (ii) comparing a second base of the first predicted base call to a second base of the known single oligo sequence to generate a corresponding first error signal. 9. Further training of the base chore repetitively further training the base caller for N1 iterations using an exemplar containing two known unique oligonucleotide sequences; further training the base caller for N2 iterations using an exemplar containing three known unique oligonucleotide sequences; The method of clause 1, in which N1 iterations are performed before N2 iterations. 10. A method of iteratively training a base collaborator using samples including a single oligonucleotide sequence, the method comprising: loading a first neural network configuration into the base collaborator; and iteratively further training the base collaborator; further training the base caller for N1 iterations using an exemplar containing two known unique oligonucleotide sequences, thereby (i) a second neural network configuration is loaded into the base collaborator for a first subset of N1 iterations; (ii) for a second subset of N1 iterations occurring after the first subset of N1 iterations, a third neural network configuration is loaded into the base collaborator, the first, second, and third neural network configurations being distinct from one another. 11. The method of clause 10, wherein the second neural network configuration is more complex than the first neural network configuration, and the third neural network configuration is more complex than the second neural network configuration. 12. The method of clause 10, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. 13. The method of clause 10, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. 14. The method of clause 10, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. 15. The method of clause 10, wherein the third neural network configuration has a greater number of layers than the second neural network configuration. 16. The method of clause 10, wherein the third neural network configuration has a greater number of weights than the second neural network configuration. 17. The method of clause 10, wherein the third neural network configuration has a greater number of parameters than the second neural network configuration. 18. Further training the base caller for N1 iterations with a sample containing two known unique oligonucleotide sequences, during one of the N1 iterations. (i) dispensing a first known oligonucleotide sequence of two known unique oligonucleotide sequences into a first plurality of clusters of a flow cell, and (ii) dispensing a second known oligonucleotide sequence of the two known unique oligonucleotide sequences into a second plurality of clusters of a flow cell; predicting a corresponding base call for each cluster of the first and second plurality of clusters, such that a plurality of predicted base calls is generated; (i) mapping a first predicted base call of the plurality of predicted base calls to a first known oligonucleotide sequence, and (ii) mapping a second predicted base call of the plurality of predicted base calls to a second known oligonucleotide sequence, while refraining from mapping a third predicted base call of the plurality of predicted base calls to either the first or second known oligonucleotide sequence; (i) generating a first error signal based on comparing the first predicted base call to the first known oligo base sequence, and (ii) generating a second error signal based on comparing the second predicted base call to the second known oligo base sequence; and further training the base colleague based on the first and second error signals. 19. Mapping the first predicted base call to a first known oligonucleotide sequence of two known unique oligonucleotide sequences, comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has similarity to the first known oligo base sequence by at least a threshold number of bases and similarity to the second known oligo base sequence by less than a threshold number of bases; and mapping the first predicted base call to the first known oligobase sequence based on determining that the first predicted base call has similarity to the first known oligobase sequence by at least a threshold number of bases. 20. Refraining from mapping the third predicted base call to either the first or second known oligonucleotide sequences; comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has a similarity to each of the first and second known oligonucleotide sequences by less than a threshold number of bases; and refraining from mapping the third predicted base call to either the first or second known oligobase sequences based on determining that the first predicted base call has similarity to each of the first and second known oligobase sequences by less than a threshold number of bases. 21. Refraining from mapping the third predicted base call to either the first or second known oligonucleotide sequences; comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has a similarity to each of the first and second known oligo base sequences by more than a threshold number of bases; and refraining from mapping the third predicted base call to either the first or second known oligobase sequences based on determining that the first predicted base call has similarity to each of the first and second known oligobase sequences by more than a threshold number of bases. 22. Generating labeled training data using a further trained base colleague for one of the N1 iterations; re-predicting corresponding base calls after further training the base caller during one of the N1 iterations, such that for each cluster of the first and second plurality of clusters, another plurality of predicted base calls is generated; (i) remapping a first subset of the other plurality of predicted base calls to a first known oligonucleotide sequence, and (ii) remapping a second subset of the other plurality of predicted base calls to a second known oligonucleotide sequence, while refraining from mapping a third subset of the other plurality of predicted base calls to either the first or second known oligonucleotide sequences; and generating the labeled training data based on the remapping such that the labeled training data includes (i) a first subset of the other plurality of predicted base calls, where the first known oligobase sequence forms ground truth data for the first subset of the other plurality of predicted base calls, and (ii) a second subset of the other plurality of predicted base calls, where the second known oligobase sequence forms ground truth data for the second subset of the other plurality of predicted base calls. 23. The labeled training data generated during one of the N1 iterations is used to train the base collaborator during the immediately following iteration of the N1 iterations. Section 22 method. 24. The neural network configuration of the base collab is the same during one of the N1 iterations as during the immediately following one of the N1 iterations. Section 23 method. 25. The neural network configuration of the base collaborator during a immediately subsequent iteration of the N1 iterations is different and more complex than the neural network configuration of the base collaborator during one iteration of the N1 iterations. Section 23 method. 26. Further training of the base chorus repeatedly The method of claim 1, comprising monotonically increasing the number of unique oligonucleotide sequences in the multi-oligonucleotide-containing sample as the iterations progress during the iterative further training. 27. Using base caller to predict base call sequences for unknown samples sequenced with known sequences of oligos; labeling each unknown analyte with a ground truth sequence that matches a known sequence; training a base collaborator using the labeled unknown analytes; Computer-implemented method. 28. The computer-implemented method of clause 27, further comprising repeating the using, labeling, and training until convergence is satisfied. 29. Using base caller to predict base call sequences for a population of unknown specimens that have been sequenced to have two or more known sequences of two or more oligos; Sorting an unknown sample from the population of unknown samples based on classification of the base call sequences of the selected unknown samples into known sequences; labeling each subset of the selected unknown samples with a respective ground truth sequence that matches each of the known sequences based on the classification; training a base classifier using each labeled subset of the selected unknown analytes; Computer-implemented method. 30. The computer-implemented method of clause 29, further comprising repeating the using, screening, labeling, and training until convergence is satisfied. 31. A non-transitory computer-readable storage medium having stored thereon computer program instructions for progressively training a base caller, the instructions, when executed on a processor, performing: iteratively initially training a base caller with samples containing a single oligonucleotide sequence and generating labeled training data using the initially trained base caller; (i) further training the base caller using a sample including a multi-oligonucleotide sequence, and generating labeled training data using the further trained base caller; and further training the base collaborator by repeating step (i) while increasing the complexity of the neural network configuration loaded into the base collaborator during at least one iteration, wherein the labeled training data generated during an iteration is used to train the base collaborator during an immediately subsequent iteration. 31a. The order is The computer-readable storage medium of clause 31, further comprising increasing the number of unique oligonucleotide sequences of the multi-oligonucleotide sequence within the exemplar during at least one iteration of further training the base call with an exemplar including the multi-oligonucleotide sequence. 32. First, iteratively training a base collaborator with samples containing a single oligonucleotide sequence, During the first iteration of the first training of the base cola, Injecting a single known oligonucleotide sequence into multiple clusters of a flow cell; generating a plurality of sequence signals corresponding to the plurality of clusters, each sequence signal of the plurality of sequence signals representing a base sequence loaded into a corresponding cluster of the plurality of clusters; predicting a corresponding base call for the known single oligonucleotide sequence based on each sequence signal of the plurality of sequence signals, thereby generating a plurality of predicted base calls; generating, for each sequence signal of the plurality of sequence signals, a corresponding error signal based on a comparison of (i) the corresponding predicted base call and (ii) the bases of the known single oligonucleotide sequence, thereby generating a plurality of error signals corresponding to the plurality of sequence signals; and initially training a base colleague during a first iteration based on the plurality of error signals. 32a. Training the base colleague first during the first iteration and updating weights and / or biases of the neural network configuration based on the plurality of error signals using a backpropagation path of the neural network configuration loaded into the base code. 33. First iteratively train the base caller with samples containing a single oligonucleotide sequence, During the second iteration of the first training of the base cola, which takes place after the first iteration of the first training, predicting corresponding further base calls for the known single oligonucleotide sequence based on each sequence signal of the plurality of sequence signals using the base calls partially trained during the first iteration of the initial training, thereby generating a plurality of further predicted base calls; generating, for each sequence signal of the plurality of sequence signals, a corresponding further error signal based on a comparison of (i) the corresponding further predicted base call and (ii) the bases of the known single oligonucleotide sequence, thereby generating a plurality of further error signals corresponding to the plurality of sequence signals; and further initially training the base colleague during a second iteration based on the plurality of further error signals. 34. First, iteratively train the base caller with samples containing a single oligonucleotide sequence. and repeating the second iteration of the initial training of the base collaborator with exemplars including a single oligonucleotide sequence for the plurality of instances until a convergence condition is met. 35. The computer-readable storage medium of clause 34, wherein the convergence condition is met if a decrease in a plurality of further error signals during two successive iterations of the second iteration of the initial training of the base colleague is less than a threshold value. 36. The computer-readable storage medium of clause 34, wherein the convergence condition is met if the second iteration of the initial training of the base collaborator is repeated for at least a threshold number of instances. 37. A plurality of sequence signals corresponding to a plurality of clusters generated during a first iteration of the initial training of the base collaborator are reused for a second iteration of the initial training of the base collaborator; 33. The computer-readable storage medium of claim 33. 38. Comparing (i) the corresponding predicted base calls with (ii) the bases of a known single oligo sequence, 33. The computer-readable storage medium of claim 32, comprising, for the first predicted base call, (i) comparing a first base of the first predicted base call to a first base of the known single oligo sequence, and (ii) comparing a second base of the first predicted base call to a second base of the known single oligo sequence to generate a corresponding first error signal. 39. Further training of the base chore repetitively further training the base caller for N1 iterations using an exemplar containing two known unique oligonucleotide sequences; further training the base caller for N2 iterations using an exemplar containing three known unique oligonucleotide sequences; 32. The computer-readable storage medium of claim 31, wherein N1 iterations are performed before N2 iterations. 40. During an initial iterative training of a base collaborator using samples including a single oligonucleotide sequence, a first neural network configuration is loaded into the base collaborator, and the base collaborator is further iteratively trained. further training the base caller for N1 iterations using an exemplar containing two known unique oligonucleotide sequences, thereby (i) for a first subset of N1 iterations, a second neural network configuration is loaded into the base collab; (ii) for a second subset of the N1 iterations occurring after the first subset of the N1 iterations, a third neural network configuration is loaded into the base collaborator, wherein the first, second, and third neural network configurations are distinct from one another. 41. The computer-readable storage medium of clause 40, wherein the second neural network configuration is more complex than the first neural network configuration, and the third neural network configuration is more complex than the second neural network configuration. 42. The computer-readable storage medium of clause 40, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. 43. The computer-readable storage medium of clause 40, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. 44. The computer-readable storage medium of clause 40, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. 45. The computer-readable storage medium of clause 40, wherein the third neural network configuration has a greater number of layers than the second neural network configuration. 46. The computer-readable storage medium of clause 40, wherein the third neural network configuration has a greater number of weights than the second neural network configuration. 47. The computer-readable storage medium of clause 40, wherein the third neural network configuration has a greater number of parameters than the second neural network configuration. 48. Further training the base caller for N1 iterations using a sample containing two known unique oligonucleotide sequences, during one of the N1 iterations. (i) dispensing a first known oligonucleotide sequence of two known unique oligonucleotide sequences into a first plurality of clusters of a flow cell, and (ii) dispensing a second known oligonucleotide sequence of the two known unique oligonucleotide sequences into a second plurality of clusters of a flow cell; predicting a corresponding base call for each cluster of the first and second plurality of clusters, such that a plurality of predicted base calls is generated; (i) mapping a first predicted base call of the plurality of predicted base calls to a first known oligonucleotide sequence, and (ii) mapping a second predicted base call of the plurality of predicted base calls to a second known oligonucleotide sequence, while refraining from mapping a third predicted base call of the plurality of predicted base calls to either the first or second known oligonucleotide sequence; (i) generating a first error signal based on comparing the first predicted base call to the first known oligo base sequence, and (ii) generating a second error signal based on comparing the second predicted base call to the second known oligo base sequence; and further training the base caller based on the first and second error signals. 49. Mapping the first predicted base call to a first known oligonucleotide sequence of two known unique oligonucleotide sequences, comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has similarity to the first known oligo base sequence by at least a threshold number of bases and similarity to the second known oligo base sequence by less than a threshold number of bases; and mapping the first predicted base call to the first known oligobase sequence based on determining that the first predicted base call has similarity to the first known oligobase sequence by at least a threshold number of bases. 50. Refraining from mapping the third predicted base call to either the first or second known oligonucleotide sequences. comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has a similarity to each of the first and second known oligonucleotide sequences by less than a threshold number of bases; and refraining from mapping the third predicted base call to either the first or second known oligobase sequences based on determining that the first predicted base call has similarity to each of the first and second known oligobase sequences of less than a threshold number of bases. 51. Refraining from mapping the third predicted base call to either the first or second known oligonucleotide sequences, comparing each base of the first predicted base call to a corresponding base of the first and second known oligo base sequences; determining that the first predicted base call has a similarity to each of the first and second known oligo base sequences by more than a threshold number of bases; and refraining from mapping the third predicted base call to either the first or second known oligobase sequences based on determining that the first predicted base call has similarity to each of the first and second known oligobase sequences by more than a threshold number of bases. 52. Generating labeled training data using a base colleague further trained for one of the N1 iterations; re-predicting corresponding base calls after further training the base caller during one of the N1 iterations, such that for each cluster of the first and second plurality of clusters, another plurality of predicted base calls is generated; (i) remapping a first subset of the other plurality of predicted base calls to a first known oligonucleotide sequence, and (ii) remapping a second subset of the other plurality of predicted base calls to a second known oligonucleotide sequence, while refraining from mapping a third subset of the other plurality of predicted base calls to either the first or second known oligonucleotide sequences; and generating the labeled training data based on the remapping such that the labeled training data includes (i) a first subset of the other plurality of predicted base calls, where the first known oligobase sequence forms ground truth data for the first subset of the other plurality of predicted base calls, and (ii) a second subset of the other plurality of predicted base calls, where the second known oligobase sequence forms ground truth data for the second subset of the other plurality of predicted base calls. 53. The labeled training data generated during one of the N1 iterations is used to train the base collaborator during the immediately following iteration of the N1 iterations. The computer-readable storage medium of clause 52. 54. The neural network configuration of the base cola is the same during one of the N1 iterations as during the immediately following one of the N1 iterations. 53. The computer-readable storage medium of claim 53. 55. The neural network configuration of the base collaborator during a immediately subsequent iteration of the N1 iterations is different and more complex than the neural network configuration of the base collaborator during one iteration of the N1 iterations. 53. The computer-readable storage medium of claim 53. 56. Further training of the base chorus repeatedly 32. The computer-readable storage medium of claim 31, comprising monotonically increasing the number of unique oligonucleotide sequences in a sample that includes multiple oligonucleotide sequences as the iterations progress during the iterative further training.
[0342] Node set #2 (self-learning base cola trained using biological sequences) A1. A computer-implemented method for progressively training a base colleague, comprising: first training a base colleague and generating labeled training data using the first trained base colleague; (i) further training the base caller with samples including biological base sequences, and generating labeled training data using the further trained base caller; iteratively further training the base collaborators by repeating step (i) for N iterations, further training the base caller for N1 iterations of the N iterations using samples including the first biological base sequence selected into the first plurality of base subsequences; and and iteratively further training the base collaborator for N2 iterations out of the N iterations using a sample including a second biological sequence selected into a second plurality of base subsequences; The complexity of the neural network configuration loaded into the base collab increases monotonically with N iterations, The indicator generated during one of N iterations In one embodiment, the collected training data is used to train a base collaborator during an iteration immediately following the N iterations. A1a. Training the base cola first is the best way to The method of clause A1, comprising initially training a base collaborator with exemplars comprising one or more oligonucleotide sequences, and generating labeled training data using the initially trained base collaborator. A2. The method of clause A1, wherein N1 iterations are performed before N2 iterations, and the second biosequence has a greater number of bases than the first biosequence. A3. Further training the base colleague for N1 repetitions is performed during one of the N1 repetitions. (i) inputting a first base partial sequence of the first plurality of base partial sequences of the first organism into a first cluster of the plurality of clusters of the flow cell, (ii) inputting a second base partial sequence of the first plurality of base partial sequences of the first organism into a second cluster of the plurality of clusters of the flow cell, and (iii) inputting a third base partial sequence of the first plurality of base partial sequences of the first organism into a third cluster of the plurality of clusters of the flow cell; (i) receiving a first sequence signal from a first cluster indicating a partial base sequence input to the first cluster, (ii) a second sequence signal from a second cluster indicating a partial base sequence input to the second cluster, and (iii) a third sequence signal from a third cluster indicating a partial base sequence input to the third cluster; (i) generating a first predicted base subsequence based on the first sequence signal, (ii) generating a second predicted base subsequence based on the second sequence signal, and (iii) generating a third predicted base subsequence based on the third sequence signal; (i) mapping the first predicted base subsequence to a first section of the first organism base sequence, and (ii) mapping the second predicted base subsequence to a second section of the first organism base sequence, while not mapping the third predicted base subsequence to any section of the first organism base sequence; The method of clause A1, comprising: generating labeled training data comprising: (i) a first predicted base subsequence mapped to a first section of the first biosequence, where the first section of the first biosequence is a ground truth for the first predicted base subsequence; and (ii) a second predicted base subsequence mapped to a second section of the first biosequence, where the second section of the first biosequence is a ground truth for the second predicted base subsequence. A3a. Further training the base colleague for N1 repetitions is performed during one of the N1 repetitions. The method of clause A3, including training the base caller using the labeled training data generated during an initial training of the base caller prior to generating the first, second, and third predicted base subsequences. A4. The first predicted base subsequence has L1 bases; one or more of the L1 bases of the first predicted base subsequence do not match corresponding bases of the first section of the first biological base sequence due to an error in base calling prediction by the base caller; Method of section A3. A5. A first predicted base subsequence has L1 bases, and the L1 bases of the first predicted base subsequence include the first L2 bases followed by the subsequent L3 bases, and mapping the first predicted base subsequence to a first section of a first biological base sequence; substantially and uniquely matching the first L2 bases of the first predicted sequence with the L2 consecutive bases of the first biological sequence; identifying a first section of a first biological sequence, the first section comprising (i) L2 consecutive bases as an initial base, and (ii) L1 consecutive bases; and mapping the first predicted base subsequence to the identified first section of the first organism base sequence. A6. The method is The method of A5, further comprising substantially and uniquely matching the first L2 bases of the first predicted base sequence, while refraining from seeking to match the subsequent L3 bases of the first predicted base sequence with any bases of the first biological sequence. A7. The method of A5, wherein the first L2 bases of the first predicted sequence substantially match L2 consecutive bases of the first biological sequence, whereby at least a threshold number of bases of the first L2 bases of the first predicted sequence match L2 consecutive bases of the first biological sequence. A8. The method of A5, wherein the first L2 bases of the first predicted base sequence uniquely match L2 consecutive bases of the first biosequence, such that the first L2 bases of the first predicted base sequence substantially match only the L2 consecutive bases of the first biosequence and do not match any other L2 consecutive bases of the first biosequence. A9. The third predicted base subsequence has L1 bases, and the third predicted base subsequence does not map to any of the base subsequences of the first plurality of base subsequences; (i) not substantially and uniquely matching the first L2 bases of the L1 bases of the third predicted sequence to the first consecutive L2 bases of the first biological sequence. A10. One iteration of the N1 iterations is a first iteration of the N1 iterations, and further training the base collaborator during a second iteration of the N1 iterations; training a base colleague using the labeled training data generated during a first iteration of the N1 iterations; using a base colleague trained with the labeled training data generated during a first iteration of the N1 iterations to generate (i) a first further predicted base subsequence based on the first sequence signal, (ii) a second further predicted base subsequence based on the second sequence signal, and (iii) a third further predicted base subsequence based on the third sequence signal; (i) mapping the first additional predicted base subsequence to a first section of the first organism base sequence, (ii) mapping the second additional predicted base subsequence to a second section of the first organism base sequence, and (iii) mapping the third additional predicted base subsequence to a third section of the first organism base sequence; and generating further labeled training data comprising: (i) a further first predicted base subsequence mapped to a first section of the first biosequence, the first section of the first biosequence being a ground truth for the further first predicted base subsequence; (ii) a further second predicted base subsequence mapped to a second section of the first biosequence, the second section of the first biosequence being a ground truth for the further second predicted base subsequence; and (iii) a further third predicted base subsequence mapped to a third section of the first biosequence, the third section of the first biosequence being a ground truth for the further third predicted base subsequence. A11. (i) generating a first error between a first predicted base subsequence generated during a first iteration of the N1 iterations and (ii) a first section of a first biological base sequence; (i) generating a second error between a further first predicted base subsequence generated during a second iteration of the N1 iterations and (ii) the first section of the first organism base sequence; The second error is less than the first error because the base collaborator is better trained during the second iteration compared to the first iteration. Method of clause A10. A12. The first, second, and third sequence signals generated during the first iteration are reused in a second iteration to generate a first further predicted base subsequence, a second further predicted base subsequence, and a third further predicted base subsequence, respectively. Method of clause A10. A13. The neural network configuration of the base collab is the same between the first iteration of the N1 iterations and the second iteration of the N1 iterations. Method of clause A10. A13a. The neural network configuration of the base classifier is reused for multiple iterations until a convergence condition is met. Method of clause A13. A14. The neural network configuration of the base collaborator during a first iteration of the N1 iterations is different and more complex than the neural network configuration of the base collaborator during a second iteration of the N1 iterations; Method of clause A10. A15. Further training the base caller for N1 iterations of the N iterations using a sample including a first biological base sequence; further training the base collaborator using the first neural network configuration loaded into the base collaborator for a first subset of N1 iterations; The method of clause A1, including further training the base collaborator with a second neural network configuration loaded into the base collaborator for a second subset of the N1 iterations, the second neural network configuration being different from the first neural network configuration. A16. The method of clause A15, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. A17. The method of clause A15, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. A18. The method of clause A15, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. A19. It is important to further train the base chore repeatedly. loading a first neural network configuration into the base collaborator for one or more of the N1 iterations using a sample that includes a first biological sequence; For one or more of the N2 iterations using a sample containing a second biological sequence, loading a second neural network configuration into the base collaborator, the second neural network configuration being different from the first neural network configuration. A20. The method of clause A19, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. A21. The method of clause A19, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. A22. The method of clause A19, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. A23. Further training the base caller for N1 iterations of the N iterations using a sample including a first biological base sequence; The method of clause A1, including repeating further training with the first biosequence until a convergence condition is met after N1 iterations. A24. The method of clause A23, wherein the convergence condition is met when a decrease in the generated error signal during two successive iterations of the N1 iterations is less than a threshold. A25. The method of clause A23, wherein the convergence condition is satisfied after completion of N1 iterations.
[0343] B1. A non-transitory computer readable storage medium having stored thereon computer program instructions for progressively training a base caller, the instructions, when executed on a processor, first training a base colleague and generating labeled training data using the first trained base colleague; (i) further training the base caller with samples including biological base sequences, and generating labeled training data using the further trained base caller; iteratively further training the base collaborators by repeating step (i) for N iterations, further training the base caller for N1 iterations of the N iterations using samples including the first biological base sequence selected into the first plurality of base subsequences; and and iteratively further training the base collaborator for N2 iterations out of the N iterations using a sample including a second biological sequence selected into a second plurality of base subsequences; The complexity of the neural network configuration loaded into the base collab increases monotonically with N iterations, A non-transitory computer-readable storage medium, wherein the labeled training data generated during an iteration of the N iterations is used to train a base colleague during an iteration immediately following the N iterations. B1a. Further training of the base chorus repetitively The computer-readable storage medium of clause B1, comprising initially training a base collaborator with analytes comprising one or more oligonucleotide sequences, and generating labeled training data using the initially trained base collaborator. B2. The computer-readable storage medium of clause B1, wherein N1 iterations are performed before N2 iterations, and the second biosequence has a greater number of bases than the first biosequence. B3. Further training the base collaborator for N1 repetitions is performed during one of the N1 repetitions. (i) inputting a first base partial sequence of the first plurality of base partial sequences of the first organism into a first cluster of the plurality of clusters of the flow cell, (ii) inputting a second base partial sequence of the first plurality of base partial sequences of the first organism into a second cluster of the plurality of clusters of the flow cell, and (iii) inputting a third base partial sequence of the first plurality of base partial sequences of the first organism into a third cluster of the plurality of clusters of the flow cell; (i) receiving a first sequence signal from a first cluster indicating a partial base sequence input to the first cluster, (ii) a second sequence signal from a second cluster indicating a partial base sequence input to the second cluster, and (iii) a third sequence signal from a third cluster indicating a partial base sequence input to the third cluster; (i) generating a first predicted base subsequence based on the first sequence signal, (ii) generating a second predicted base subsequence based on the second sequence signal, and (iii) generating a third predicted base subsequence based on the third sequence signal; (i) mapping the first predicted base subsequence to a first section of the first organism base sequence, and (ii) mapping the second predicted base subsequence to a second section of the first organism base sequence, while not mapping the third predicted base subsequence to any section of the first organism base sequence; and generating labeled training data including (i) a first predicted base subsequence mapped to a first section of a first biosequence, the first section of the first biosequence being a ground truth for the first predicted base subsequence, and (ii) a second predicted base subsequence mapped to a second section of the first biosequence, the second section of the first biosequence being a ground truth for the second predicted base subsequence. B3a. Further training the base collaborator for N1 repetitions is performed during one of the N1 repetitions. The computer-readable storage medium of clause B3, including training the base collaborator using the labeled training data generated during an initial training of the base collaborator prior to generating the first, second, and third predicted base subsequences. B4. The first predicted base subsequence has L1 bases; one or more of the L1 bases of the first predicted base subsequence do not match corresponding bases of the first section of the first biological base sequence due to an error in base calling prediction by the base caller; A computer-readable storage medium according to clause B3. B5. A first predicted base subsequence has L1 bases, the L1 bases of the first predicted base subsequence include the first L2 bases followed by the subsequent L3 bases, and mapping the first predicted base subsequence to a first section of a first biological base sequence; substantially and uniquely matching the first L2 bases of the first predicted sequence with the L2 consecutive bases of the first biological sequence; identifying a first section of a first biological sequence, the first section comprising (i) L2 consecutive bases as an initial base, and (ii) L1 consecutive bases; and mapping the first predicted base subsequence to the identified first section of the first biological base sequence. B6. substantially and uniquely matching the first L2 bases of the first predicted sequence, while refraining from seeking to match the subsequent L3 bases of the first predicted sequence with any bases of the first biological sequence; B5 computer readable storage medium. B7. The computer-readable storage medium of B5, wherein the first L2 bases of the first predicted base sequence substantially match L2 consecutive bases of the first biological base sequence, whereby at least a threshold number of bases of the first L2 bases of the first predicted base sequence match L2 consecutive bases of the first biological base sequence. B8. The computer-readable storage medium of B5, wherein the first L2 bases of the first predicted base sequence uniquely match L2 consecutive bases of the first biological base sequence, whereby the first L2 bases of the first predicted base sequence substantially match only L2 consecutive bases of the first biological base sequence and do not match other L2 consecutive bases of the first biological base sequence. B9. The third predicted base subsequence has L1 bases, and the third predicted base subsequence does not map to any of the base subsequences of the first plurality of base subsequences; (i) not substantially and uniquely matching a first L2 bases of the L1 bases of the third predicted base sequence to a first L2 consecutive bases of the first biological sequence. B10. One iteration of the N1 iterations is a first iteration of the N1 iterations, and further training the base collaborator during a second iteration of the N1 iterations; training a base colleague using the labeled training data generated during a first iteration of the N1 iterations; using a base colleague trained with the labeled training data generated during a first iteration of the N1 iterations to generate (i) a first further predicted base subsequence based on the first sequence signal, (ii) a second further predicted base subsequence based on the second sequence signal, and (iii) a third further predicted base subsequence based on the third sequence signal; (i) mapping the first additional predicted base subsequence to a first section of the first organism base sequence, (ii) mapping the second additional predicted base subsequence to a second section of the first organism base sequence, and (iii) mapping the third additional predicted base subsequence to a third section of the first organism base sequence; and generating further labeled training data including: (i) a further first predicted base subsequence mapped to a first section of the first biosequence, the first section of the first biosequence being a ground truth for the further first predicted base subsequence; (ii) a further second predicted base subsequence mapped to a second section of the first biosequence, the second section of the first biosequence being a ground truth for the further second predicted base subsequence; and (iii) a further third predicted base subsequence mapped to a third section of the first biosequence, the third section of the first biosequence being a ground truth for the further third predicted base subsequence. B11. (i) generating a first error between a first predicted base subsequence generated during a first iteration of the N1 iterations and (ii) a first section of the first biological base sequence; (i) generating a second error between a further first predicted base subsequence generated during a second iteration of the N1 iterations and (ii) the first section of the first biological base sequence; The second error is less than the first error because the base collaborator is better trained during the second iteration compared to the first iteration. The computer-readable storage medium of clause B10. B12. The first, second, and third sequence signals generated during the first iteration are reused in a second iteration to generate a first further predicted base subsequence, a second further predicted base subsequence, and a third further predicted base subsequence, respectively. The computer-readable storage medium of clause B10. B13. The neural network configuration of the base collab is the same between the first iteration of N1 iterations and the second iteration of N1 iterations. The computer-readable storage medium of clause B10. B13a. The base neural network configuration is reused for multiple iterations until a convergence condition is met. A computer-readable storage medium according to clause B13. B14. The neural network configuration of the base collaborator during a first iteration of the N1 iterations is different and more complex than the neural network configuration of the base collaborator during a second iteration of the N1 iterations. The computer-readable storage medium of clause B10. B15. Further training the base caller for N1 iterations of the N iterations using a sample including the first biological sequence; for a first subset of N1 iterations, further training the base collaborator using the first neural network configuration loaded into the base collaborator; and for a second subset of the N1 iterations, further training the base collaborator with a second neural network configuration loaded into the base collaborator, the second neural network configuration being different from the first neural network configuration. B16. The computer-readable storage medium of clause B15, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. B17. The computer-readable storage medium of clause B15, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. B18. The computer-readable storage medium of clause B15, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. B19. Further training of the base chorus repeatedly loading a first neural network configuration into the base collaborator for one or more of the N1 iterations using a sample that includes a first biological sequence; and for one or more of the N2 iterations with a specimen including a second biosequence, loading a second neural network configuration into the base collaborator, the second neural network configuration being different from the first neural network configuration. B20. The computer-readable storage medium of clause B19, wherein the second neural network configuration has a greater number of layers than the first neural network configuration. B21. The computer-readable storage medium of clause B19, wherein the second neural network configuration has a greater number of weights than the first neural network configuration. B22. The computer-readable storage medium of clause B19, wherein the second neural network configuration has a greater number of parameters than the first neural network configuration. B23. Further training the base caller for N1 iterations of the N iterations using a sample including the first biological sequence; The computer-readable storage medium of clause B1, comprising repeating further training with the first biosequence until a convergence condition is met after N1 iterations. B24. The computer-readable storage medium of clause B23, wherein the convergence condition is met when a decrease in the generated error signal between two successive iterations of the N1 iterations is less than a threshold. B25. The computer-readable storage medium of clause B23, wherein the convergence condition is satisfied after completion of N1 iterations.
[0344] 1. A computer-implemented method for progressively training a base colleague, comprising: (i) using a base caller to predict single oligo base call sequences for a population of single oligo unknown analytes (i.e., unknown target sequences) that have been sequenced to have a known sequence of oligos; (ii) labeling each single oligo unknown analyte in the population of single oligo unknown analytes with a single oligo ground truth sequence that matches the known sequence; and (iii) starting with a single oligo training phase in which the labeled population of single oligo unknown analytes is used to train the base caller; (i) using the base caller to predict multi-oligo base call sequences for a population of multi-oligo unknown analytes that have been sequenced to have two or more known sequences of two or more oligos; (ii) sorting multi-oligo unknown analytes from the population of multi-oligo unknown analytes based on the classification of the multi-oligo base call sequences of the sorted multi-oligo unknown analytes into known sequences; (iii) labeling each subset of the sorted multi-oligo unknown analytes with a respective multi-oligo ground truth sequence that matches the respective known sequence based on the classification; and (iv) continuing with one or more multi-oligo training phases to further train the base caller using each labeled subset of the sorted multi-oligo unknown analytes. 11. A computer-implemented method comprising: (i) predicting organism-specific base call sequences for a population of organism-specific unknown analytes that have been sequenced to have one or more known subsequences of a reference sequence for the...
Claims
1. A computer-implemented method for progressively training a base caller, comprising: first iteratively training the base caller with a sample of a single oligonucleotide sequence having a single known oligonucleotide sequence, and generating labeled training data using the first-trained base caller, wherein the base caller has an initial neural network configuration including several layers and parameters; (i) further training the base caller with a sample containing a multi-oligonucleotide sequence including at least two known oligonucleotide sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and generating labeled training data using the further-trained base caller; while iteratively further training the base caller by repeating step (i), during at least one iteration, increasing the complexity of the initial neural network configuration of the base caller by increasing the number of the layers and parameters with respect to the initial neural network configuration in order to adjust the initial neural network configuration for the multi-oligonucleotide sequence, and using the labeled training data generated during the iteration to train the base caller during the immediately following iteration; A computer-implemented method comprising the above steps.
2. During at least one iteration of further training the base caller with the sample containing the multi-oligonucleotide sequence, increasing the number of unique oligonucleotide sequences of the multi-oligonucleotide sequence in the sample; further increasing the number of the layers and parameters and further adjusting the initial neural network configuration for the increased number of unique oligonucleotide sequences; The computer-implemented method according to claim 1, further comprising the above steps.
3. The step of first iteratively training the base caller with the sample containing the single oligonucleotide sequence is during the first iteration of first iteratively training the base caller, loading the single oligonucleotide sequence into a plurality of clusters of a flow cell; generating a plurality of array signals corresponding to the plurality of clusters, wherein each array signal of the plurality of array signals represents a base sequence loaded into the corresponding cluster of the plurality of clusters. Based on each of the plurality of array signals, predicting a corresponding base call for the single known oligo base sequence, thereby generating a plurality of predicted base calls; For each of the plurality of array signals, generating a corresponding error signal based on a comparison between (i) the corresponding predicted base call and (ii) the base of the single known oligo base sequence, thereby generating a plurality of error signals corresponding to the plurality of array signals; Based on the plurality of error signals, initially training the base caller during the first iteration; The computer-implemented method according to claim 1 or 2, comprising:
4. Repeatedly and initially training the base caller using the sample containing the single oligo base sequence, During a second iteration of repeatedly and initially training the base caller, which is performed after the first iteration of repeatedly and initially training the base caller, Using the base caller partially trained during the first iteration, based on each of the plurality of array signals, predicting a corresponding further base call for the single known oligo base sequence, thereby generating a plurality of further predicted base calls; For each of the plurality of array signals, generating a corresponding further error signal based on a comparison between (i) the corresponding further predicted base call and (ii) the base of the single known oligo base sequence, thereby generating a plurality of further error signals corresponding to the plurality of array signals; Based on the plurality of further error signals, further initially training the base caller during the second iteration; The computer-implemented method according to claim 3, further comprising:
5. The plurality of array signals corresponding to the plurality of clusters, generated during the first iteration of repeatedly and initially training the base caller, are reused for the second iteration of repeatedly and initially training the base caller. The computer-implemented method according to claim 4.
6. Comparing (i) the corresponding predicted base call with (ii) the base of the single known oligo sequence, For the first predicted base call, (i) comparing the first base of the first predicted base call with the first base of the single known oligo sequence, and (ii) comparing the second base of the first predicted base call with the second base of the single known oligo sequence to generate a corresponding first error signal, the computer-implemented method according to claim 4.
7. iteratively further training the base caller, further training the base caller for N1 iterations using a sample containing two known unique oligo base sequences, further training the base caller for N2 iterations using a sample containing three known unique oligo base sequences, comprising the computer-implemented method according to claim 1, wherein the N1 iterations are performed before the N2 iterations.
8. A system, at least one processor, a non-transitory computer-readable storage medium storing computer program instructions, wherein when the computer program instructions are executed on the at least one processor, first training the base caller iteratively using a sample of a single oligo base sequence having a single known oligo base sequence, and generating labeled training data using the first trained base caller, wherein the base caller has an initial neural network configuration including several layers and parameters, generating; (i) further training the base caller using a sample containing a multi-oligo base sequence including at least two known oligo base sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and generating labeled training data using the further trained base caller; while iteratively further training the base caller by repeating step (i), during at least one iteration, increasing the complexity of the initial neural network configuration of the base caller by increasing the number of the layers and parameters with respect to the initial neural network configuration in order to adjust the initial neural network configuration for the multi-oligo base sequence, and using the labeled training data generated during the iteration to train the base caller during the immediately subsequent iteration, increasing. A system comprising a non - transitory computer - readable storage medium that performs an action and includes the same. **Claim 9** While initially training the base caller iteratively using the sample containing the single oligonucleotide sequence, a first neural network configuration is loaded into the base caller, and further training the base caller iteratively Includes further training the base caller during N1 iterations using a sample containing two known unique oligonucleotide sequences, whereby (i) During a first subset of the N1 iterations, a second neural network configuration is loaded into the base caller, (ii) During a second subset of the N1 iterations that occurs after the first subset of the N1 iterations, a third neural network configuration is loaded into the base caller, and the first, second, and third neural network configurations are different from each other. The system according to claim 8. **Claim 10** The second neural network configuration is more complex than the first neural network configuration, and the third neural network configuration is more complex than the second neural network configuration. The system according to claim 9. **Claim 11** The second neural network configuration has a greater number of layers, a greater number of weights, or a greater number of parameters than the first neural network configuration. The system according to claim 9 or 10. **Claim 12** The third neural network configuration has a greater number of layers, a greater number of weights, or a greater number of parameters than the second neural network configuration. The system according to claim 9. **Claim 13** Further training the base caller during the N1 iterations using the sample containing two known unique oligonucleotide sequences includes, during one of the N1 iterations, (i) Loading a first known oligonucleotide sequence of the two known unique oligonucleotide sequences into a first plurality of clusters of the flow cell, and (ii) loading a second known oligonucleotide sequence of the two known unique oligonucleotide sequences into a second plurality of clusters of the flow cell, Predicting corresponding base calls such that a plurality of predicted base calls are generated for each of the first and second pluralities of clusters. while mapping (i) a first predicted base call among the plurality of predicted base calls to the first known oligo base sequence and (ii) a second predicted base call among the plurality of predicted base calls to the second known oligo base sequence, refraining from mapping a third predicted base call among the plurality of predicted base calls to either the first or the second known oligo base sequence; generating (i) a first error signal based on comparing the first predicted base call to the first known oligo base sequence and (ii) a second error signal based on comparing the second predicted base call to the second known oligo base sequence; further training the base caller based on the first and second error signals; The system according to any one of claims 9, comprising:
14. mapping the first predicted base call to the first known oligo base sequence of the two known unique oligo base sequences; comparing each base of the first predicted base call to the corresponding bases of the first and second known oligo base sequences; determining that the first predicted base call has a similarity of at least a threshold number of bases with the first known oligo base sequence and a similarity of less than the threshold number of bases with the second known oligo base sequence; mapping the first predicted base call to the first known oligo base sequence based on determining that the first predicted base call has a similarity of at least the threshold number of bases with the first known oligo base sequence; The system according to claim 13, comprising:
15. refraining from mapping the third predicted base call to either the first or the second known oligo base sequence; comparing each base of the first predicted base call to the corresponding bases of the first and second known oligo base sequences; determining that the first predicted base call has a similarity of less than a threshold number of bases with each of the first and second known oligo base sequences; Pending the mapping of the third predicted base call to either the first or second known oligo base sequence based on a determination that the first predicted base call has a similarity of less than the threshold number of bases with each of the first and second known oligo base sequences, The system according to claim 13, comprising.
16. Pending the mapping of the third predicted base call to either the first or second known oligo base sequence, Comparing each base of the first predicted base call with the corresponding bases of the first and second known oligo base sequences, Determining that the first predicted base call has a similarity of more than the threshold number of bases with each of the first and second known oligo base sequences, Pending the mapping of the third predicted base call to either the first or second known oligo base sequence based on a determination that the first predicted base call has a similarity of more than the threshold number of bases with each of the first and second known oligo base sequences, The system according to claim 13, comprising.
17. Pending the generation of labeled training data using the further trained base caller during the one iteration out of the N1 iterations, After further training the base caller during the one iteration out of the N1 iterations, re-predicting the corresponding base calls such that a plurality of additional predicted base calls are generated for each of the first and second plurality of clusters, Pending the remapping of (i) a first subset of the plurality of additional predicted base calls to the first known oligo base sequence and (ii) a second subset of the plurality of additional predicted base calls to the second known oligo base sequence, while pending the mapping of a third subset of the plurality of additional predicted base calls to either the first or second known oligo base sequence, generate labeled training data based on the remapping such that the labeled training data includes (i) a first subset of the additional plurality of predicted basecalls, wherein the first known oligo base sequence forms ground truth data for the first subset of the additional plurality of predicted basecalls, and (ii) a second subset of the additional plurality of predicted basecalls, wherein the second known oligo base sequence forms the ground truth data for the second subset of the additional plurality of predicted basecalls; The system of claim 13, comprising. **Claim 18** the labeled training data generated during the one iteration out of the N1 iterations is used to train the basecaller during the immediately subsequent iteration out of the N1 iterations; the neural network configuration of the basecaller remains unchanged during the one iteration out of the N1 iterations and during the immediately subsequent iteration out of the N1 iterations; or the neural network configuration of the basecaller during the immediately subsequent iteration out of the N1 iterations is different from and more complex than the neural network configuration of the basecaller during the one iteration out of the N1 iterations. The system of claim 17. **Claim 19** A non-transitory computer-readable storage medium storing computer program instructions for progressively training a basecaller, which, when executed on a processor, first trains the basecaller iteratively using a single oligo base sequence specimen having a single known oligo base sequence, and generates labeled training data using the first trained basecaller, wherein the basecaller has an initial neural network configuration including several layers and parameters; (i) further trains the basecaller using a specimen including a multi-oligo base sequence including at least two known oligo base sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and generates labeled training data using the further trained basecaller; By repeating step (i) to further train the base caller, while during at least one iteration, in order to adjust the initial neural network configuration for the multi-oligonucleotide sequence, by increasing the number of said layers and parameters with respect to the initial neural network configuration, increasing the complexity of the initial neural network configuration of the base caller, and using the labeled training data generated during the iteration to train the base caller during the immediately following iteration, the increasing; A non-transitory computer-readable storage medium that performs actions including
20. Actions including further training the base caller by repeating it, when executed on the processor In response to increasing the number of said layers and parameters with respect to the initial neural network configuration to adjust the initial neural network configuration for the multi-oligonucleotide sequence before further training the base caller using a sample containing the multi-oligonucleotide sequence, further training the base caller using the labeled training data generated during the last iteration of initially training the base caller by repeating it, the non-transitory computer-readable storage medium according to claim 19, further storing computer program instructions for performing