Self-learning base caller trained using oligonucleotide sequences

By using FPGA accelerators on embedded systems, an efficient dataflow and CNN acceleration hardware architecture was designed, solving the problem of low computational efficiency of CNNs on portable systems and achieving high-performance CNN inference acceleration.

CN117546249BActive Publication Date: 2026-08-04ILLUMINA INC
View PDF 45 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ILLUMINA INC
Filing Date
2022-06-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Deploying deep convolutional neural networks (CNNs) on portable and embedded systems presents challenges such as computational intensity, large data volume, frequent changes in algorithm structure, and frequent memory access, leading to low hardware efficiency and performance.

Method used

Using a Field Programmable Gate Array (FPGA) as an accelerator, an efficient dataflow and CNN acceleration hardware architecture is designed. Convolution operations are optimized by synthesizing custom circuits, and the high parallelism and flexibility of FPGA are used to accelerate the inference of CNN algorithms.

Benefits of technology

This improves the hardware efficiency and performance of CNN algorithms on portable and embedded systems, achieving high-performance and efficient CNN inference acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117546249B_ABST
    Figure CN117546249B_ABST
Patent Text Reader

Abstract

A method of progressively training a base caller is disclosed. The method includes iteratively initially training a base caller with analytes comprising single oligonucleotide base sequences and generating labeled training data using the initially trained base caller. In operation (i), the base caller is further trained with analytes comprising multiple oligonucleotide base sequences and labeled training data is generated using the further trained base caller. Operation (i) is iteratively repeated to further train the base caller. In an example, during at least one iteration, the complexity of a neural network configuration loaded within the base caller is increased. In an example, the labeled training data generated during an iteration is used to train the base caller during an immediately subsequent iteration.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority application

[0002] This application claims priority to U.S. non-provisional patent application No. 17 / 830,287 (Attorney’s File No. ILLM 1038-3 / IP-2050-US), filed June 1, 2022, entitled “Self-Learned Base Caller, Trained Using Oligo Sequences,” which claims priority to U.S. provisional patent application No. 63 / 216,419 (Attorney’s File No. ILLM 1038-1 / IP-2050-PRV), filed June 29, 2021, entitled “Self-Learned Base Caller, Trained Using Oligo Sequences,” and U.S. provisional patent application No. 63 / 216,404 (Attorney’s File No. ILLM 1038-2 / IP-2094-PRV), filed June 29, 2021, entitled “Self-Learned Base Caller, Trained Using Organism Sequences.” This priority claim is incorporated by reference for all purposes.

[0003] This application claims priority to U.S. non-provisional patent application No. 17 / 830,316 (Attorney’s File No. ILLM1038-5 / IP-2094-US), filed June 1, 2022, entitled “Self-Learned Base Caller, Trained Using Organism Sequences,” which claims priority to U.S. provisional patent application No. 63 / 216,404 (Attorney’s File No. ILLM 1038-2 / IP-2094-PRV), filed June 29, 2021, entitled “Self-Learned Base Caller, Trained Using Organism Sequences,” and U.S. provisional patent application No. 63 / 216,419 (Attorney’s File No. ILLM 1038-1 / IP-2050-PRV), filed June 29, 2021, entitled “Self-Learned Base Caller, Trained Using Oligo Sequences.” This priority claim is incorporated by reference for all purposes. Technical Field

[0004] The technologies disclosed in this invention relate to artificial intelligence-type computers and digital data processing systems, as well as corresponding data processing methods and products for simulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems); and include systems for uncertainty inference (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technologies relate to using deep neural networks, such as deep convolutional neural networks, for data analysis.

[0005] Literature merged

[0006] The following references are incorporated herein by reference as if they were shown in their entirety in this article:

[0007] The PCT patent application filed at the same time is entitled "SELF-LEARNED BASE CALLER, TRAINED USING ORGANISM SEQUENCES" (Agent's file number ILLM 1038-6 / IP-2094-PCT);

[0008] U.S. Provisional Patent Application No. 62 / 979,384, entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLINGOF INDEX SEQUENCES”, filed on February 20, 2020 (Attorney’s File No. ILLM 1015-1 / IP-1857-PRV);

[0009] U.S. Provisional Patent Application No. 62 / 979,414, entitled “ARTIFICIAL INTELLIGENCE-BASED MANY-TO-MANYBASE CALLING”, filed on February 20, 2020 (Attorney’s File No. ILLM 1016-1 / IP-1858-PRV);

[0010] U.S. non-provisional patent application No. 16 / 825,987, entitled “TRAINING DATA GENERATION FOR ARTIFICIALINTELLIGENCE-BASED SEQUENCING”, filed on March 20, 2020 (Attorney’s File No. ILLM 1008-16 / IP-1693-US);

[0011] U.S. non-provisional patent application No. 16 / 825,991, entitled “ARTIFICIAL INTELLIGENCE-BASED GENERATION OF SEQUENCING METADATA”, filed on March 20, 2020 (Attorney’s File No. ILLM 1008-17 / IP-1741-US);

[0012] U.S. non-provisional patent application number 16 / 826,126, entitled “ARTIFICIAL INTELLIGENCE-BASED BASE CALLING”, filed on March 20, 2020 (Attorney’s file number ILLM 1008-18 / IP-1744-US);

[0013] U.S. non-provisional patent application No. 16 / 826,134 (Attorney's File No. ILLM 1008-19 / IP-1747-US), filed on March 20, 2020, entitled "ARTIFICIAL INTELLIGENCE-BASED QUALITYSCORING"; and

[0014] U.S. non-provisional patent application number 16 / 826,168, entitled “ARTIFICIAL INTELLIGENCE-BASED SEQUENCING”, filed on March 21, 2020 (Attorney’s file number ILLM 1008-20 / IP-1752-PRV-US). Background Technology

[0015] The topics discussed in this section should not be considered prior art simply because they are mentioned here. Similarly, problems mentioned in this section or related to the topics provided as background art should not be assumed to have been previously recognized in the prior art. The topics in this section merely represent different methods, which themselves may correspond to specific implementations of the technology protected by the claims.

[0016] In recent years, the rapid increase in computing power has enabled deep convolutional neural networks (CNNs) to achieve great success in many computer vision tasks with significantly improved accuracy. During the inference phase, many applications require low latency processing of an image with stringent power consumption requirements, which reduces the efficiency of graphics processing units (GPUs) and other general-purpose platforms. This has created opportunities for specific acceleration hardware (e.g., field-programmable gate arrays (FPGAs)) by customizing digital circuitry specifically for deep learning algorithm inference. However, deploying CNNs on portable and embedded systems remains challenging due to large data volumes, computationally intensive operations, varying algorithmic architectures, and frequent memory accesses.

[0017] Since convolution contributes a large portion of the computation in CNNs, convolution acceleration schemes significantly impact the efficiency and performance of hardware CNN accelerators. Convolution involves a multiplication and accumulation (MAC) operation with four recurrent stages sliding along the kernel and feature maps. The first recurrent stage computes the MAC of pixels within the kernel window. The second recurrent stage accumulates the sum of the products of MACs across different input feature maps. After completing the first and second recurrent stages, a bias is added to obtain the final output element in the output feature map. The third recurrent stage slides the kernel window within the input feature map. The fourth recurrent stage generates different output feature maps.

[0018] FPGAs have gained increasing attention and popularity, particularly in accelerating inference tasks, due to their (1) high reconfigurability, (2) faster development time compared to application-specific integrated circuits (ASICs) to keep pace with the rapid development of CNNs, (3) good performance, and (4) superior energy efficiency compared to GPUs. The high performance and efficiency of FPGAs can be achieved by synthesizing circuits tailored to specific computations to directly process billions of operations using customized memory systems. For example, hundreds to thousands of digital signal processing (DSP) blocks on modern FPGAs support core convolution operations such as multiplication and addition with high parallelism. Dedicated data buffers between external on-chip memory and on-chip processing engines (PEs) can be designed to achieve optimized data flow by configuring tens of megabytes of on-chip block random access memory (BRAM) on the FPGA chip.

[0019] High-performance CNNs require efficient data flow and accelerated hardware architectures to minimize data communication while maximizing resource utilization. This presents an opportunity to design methods and frameworks for accelerating the inference process of various CNN algorithms on high-performance, efficient, and highly flexible accelerated hardware. Attached Figure Description

[0020] In the accompanying drawings, similar reference numerals generally refer to similar parts in all different views. Furthermore, the drawings are not necessarily drawn to scale, but rather emphasize the principles of the disclosed technology. In the following description, various specific embodiments of the disclosed technology are described with reference to the following drawings, wherein:

[0021] Figure 1 A cross-section of a biosensor that can be used in various implementation schemes is shown.

[0022] Figure 2 A specific implementation of a circulation pool containing clusters in its blocks is shown.

[0023] Figure 3 An exemplary flow pool with eight channels is shown, and a magnified view of a block, its clusters, and their surrounding background is also shown.

[0024] Figure 4 This is a simplified block diagram of a system used to analyze sensor data from a sequencing system (such as the output of a base detection sensor).

[0025] Figure 5 This is a simplified diagram illustrating an aspect of the base detection operation, which includes the functionality of a runtime program executed by the host processor.

[0026] Figure 6 It is a configurable processor (such as, Figure 4 A simplified diagram of the configuration of the configurable processor.

[0027] Figure 7 It is a graph of neural network architectures that can be executed using configurable or reconfigurable arrays as described in this document.

[0028] Figure 8A It is as follows Figure 7 A simplified diagram illustrating the organization of blocks of sensor data using the same neural network architecture.

[0029] Figure 8B It is as follows Figure 7 A simplified illustration of patches of sensor data used in the same neural network architecture.

[0030] Figure 9 This illustrates a configurable or reconfigurable array (such as a field-programmable gate array (FPGA)) as shown in the image. Figure 7 It is part of the same neural network configuration.

[0031] Figure 10 It is a diagram of another alternative neural network architecture that can be executed using a configurable or reconfigurable array as described in this article.

[0032] Figure 11 A specific implementation of a specialized architecture for a neural network-based base detector is shown, which is used to isolate the processing of data from different sequencing cycles.

[0033] Figure 12 A specific implementation of isolation layers is shown, each of which may include convolutions.

[0034] Figure 13A A specific implementation of a combination layer is shown, where each combination layer may include convolutions.

[0035] Figure 13B Another specific implementation of the combined layers is shown, each of which may include convolutions.

[0036] Figure 14AA base detection system operating during the single oligonucleotide training phase is illustrated to train a base detector, including a neural network configuration, using known synthetic oligonucleotide sequences. Figure 14A1 The comparison operation between the predicted base sequence and the corresponding benchmark true base sequence is shown.

[0037] Figure 14B It shows Figure 14A Further details of the base detection system, which operates in a single oligonucleotide training phase to train a base detector, including a neural network configuration, using known synthetic oligonucleotide sequences.

[0038] Figure 15A It shows Figure 14A The base detection system operates in the training data generation phase of the di-oligonucleotide training phase to generate labeled training data using two known synthetic sequences.

[0039] Figure 15B and Figure 15C It shows relative to Figure 15A Two corresponding example choices of the diolithonucleotide sequences discussed.

[0040] Figure 15D Example mapping operations are shown for (i) mapping a predicted base detection sequence to either a first oligonucleotide or a second oligonucleotide, or (ii) declaring an uncertainty in mapping a predicted base detection sequence to either of two oligonucleotides.

[0041] Figure 15E It shows from Figure 15D The labeled training data generated by the mapping, wherein the training data is composed of... Figure 16A Another neural network configuration is shown.

[0042] Figure 16A It shows Figure 14A A base detection system that operates during the training data consumption and training phase of a dual oligonucleotide training phase, to train another neural network configuration (which differs from the one used in the training phase) using two known synthetic oligonucleotide sequences. Figure 14A The neural network configuration, and a more complex base detector.

[0043] Figure 16B It shows Figure 14A The base detection system operates in the second iteration of the training data generation phase during the dioliponucleotide training phase.

[0044] Figure 16C It shows from Figure 16BThe mapping shown generates labeled training data, which will be used for further training.

[0045] Figure 16D It shows Figure 14A The base detection system operates in the second iteration of the "training data consumption and training phase" of the "dual oligonucleotide training phase" to train a sequence containing two known synthetic oligonucleotide sequences. Figure 16A A base detector configured with a neural network.

[0046] Figure 17A A flowchart is shown illustrating an example method for iteratively training a neural network configuration for base detection using single and double oligonucleotide sequences.

[0047] Figure 17B It shows in Figure 17A Method 1700 ends with example labeled training data generated by the Pth NN configuration.

[0048] Figure 18A It shows Figure 14A The base detection system operates in the first iteration of the "training data consumption and training phase" of the "tri-oligonucleotide training phase" to train a base detector including a 3-oligoneuron configuration.

[0049] Figure 18B It shows Figure 14A The base detection system operates in the "training data generation phase" of the "trioliponucleotide training phase" to train a set of bases including... Figure 18A A base detector configured with a 3-oligonucleotide neural network.

[0050] Figure 18C The mapping operation is shown for (i) mapping the predicted base detection sequence to Figure 18B (ii) Any of the three oligonucleotides, or (ii) a statement that the mapping of the predicted base detection sequence is uncertain.

[0051] Figure 18D It shows from Figure 18C The mapping generates labeled training data, which is used to train another neural network configuration.

[0052] Figure 18E A flowchart is shown depicting an example method for iteratively training a neural network configuration for base detection using a 3-oligonucleotide benchmark truth sequence.

[0053] Figure 19A flowchart is shown depicting an example method for iteratively training a neural network configuration for base detection using a polynucleotide benchmark truth sequence.

[0054] Figure 20A The training method is shown. Figure 14A The biological sequence of the base detector.

[0055] Figure 20B It shows Figure 14A The base detection system operates during the training data generation phase of the first organism training phase, in order to use... Figure 20A Various subsequences of the first organism sequence are used to train a base detector that includes a first organism-level neural network configuration.

[0056] Figure 20C An example of fading is shown, where the signal strength decreases with the number of cycles of sequencing runs as a base detection operation.

[0057] Figure 20D The signal-to-noise ratio decreases as the sequencing cycle progresses, conceptually illustrating this.

[0058] Figure 20E This illustrates the base detection of the first L2 bases out of the L1 number of bases in a subsequence, where the first L2 number of bases in the subsequence are used to map the subsequence to... Figure 20A The biological sequence.

[0059] Figure 20F It shows from Figure 20E The mapped labeled training data, wherein the labeled training data includes the benchmark ground values. Figure 20A A portion of the biological sequence.

[0060] Figure 20G It shows Figure 14A The base detection system operates in the "training data consumption and training phase" of the "biological level training phase" to train a base detector including a first biological level neural network configuration.

[0061] Figure 21 A depiction for use is shown Figure 20A A flowchart of an example method for iteratively training a neural network configuration for base detection using a simple biological sequence.

[0062] Figure 22 The training method is shown. Figure 14A The use of complex biological sequences in the corresponding NN configuration of the base detector.

[0063] Figure 23AA flowchart illustrating an example method for iteratively training a neural network configuration for base detection is shown, and Figures 23B to 23E Various diagrams are shown illustrating the effectiveness of the base detector training process discussed in this disclosure.

[0064] Figure 24 It is a block diagram based on a specific implementation of a base detection system.

[0065] Figure 25 It is possible Figure 24 A block diagram of the system controller used in the system.

[0066] Figure 26 It is a simplified block diagram of a computer system that can be used to implement the disclosed technology. Detailed Implementation

[0067] As used herein, the terms “polynucleotide” or “nucleic acid” refer to deoxyribonucleic acid (DNA), but where appropriate, those skilled in the art will recognize that the systems and apparatus described herein can also be applied to ribonucleic acid (RNA). It should be understood that the term includes analogs of DNA or RNA formed from nucleotide analogs as equivalents. As used herein, the term also covers cDNA, i.e., complementary or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase.

[0068] Single-stranded polynucleotide molecules sequenced by the systems and apparatus described herein can originate in single-stranded form, such as DNA or RNA, or in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, PCR and amplification products, etc.). Therefore, a single-stranded polynucleotide can be the sense or antisense strand of a polynucleotide double helix. Methods for preparing single-stranded polynucleotide molecules suitable for the methods of this disclosure using standard techniques are well known in the art. The precise sequence of the primary polynucleotide molecule is generally not important to this disclosure and can be known or unknown. A single-stranded polynucleotide molecule can represent a genomic DNA molecule (e.g., human genomic DNA) comprising intron and exon sequences (coding sequences) and non-coding regulatory sequences, such as promoter and enhancer sequences.

[0069] In some embodiments, nucleic acids to be sequenced using this disclosure are immobilized on a substrate (e.g., a substrate within a flow cell or a substrate such as one or more beads on a flow cell). Unless otherwise stated or clearly indicated by the context, the term "immobilized" as used herein is intended to cover direct or indirect, covalent or non-covalent binding. In some embodiments, covalent attachment may be preferred, but generally all that is required is that the molecule (e.g., nucleic acid) remains immobilized or attached to the vector under conditions intended for use with the vector (e.g., in applications requiring nucleic acid sequencing).

[0070] As used herein, the term "solid carrier" (or "substrate" in some uses) refers to any inert substrate or matrix to which nucleic acids can be attached, such as, for example, glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers. In many embodiments, the solid carrier is a glass surface (e.g., a flat surface of a flow cell channel). In some embodiments, the solid carrier may comprise an inert substrate or matrix that has been "functionalized," for example by applying a layer or coating of an intermediate material that includes reactive groups that allow covalent attachment to molecules such as polynucleotides. As a non-limiting example, such a carrier may comprise a polyacrylamide hydrogel loaded on an inert substrate such as glass. In such embodiments, the molecule (polynucleotide) may be directly covalently attached to the intermediate material (e.g., the hydrogel), but the intermediate material itself may be non-covalently attached to the substrate or matrix (e.g., a glass substrate). Covalent attachment to a solid carrier should be interpreted accordingly to encompass this type of arrangement.

[0071] As noted above, this disclosure includes novel systems and apparatuses for sequencing nucleic acids. It will be apparent to those skilled in the art that, depending on the context, references herein to a particular nucleic acid sequence may also refer to nucleic acid molecules containing such sequences. Sequencing a target fragment means establishing a temporal sequence of reads of bases. The bases read need not be sequential, although this is preferred, but it is not necessary to sequence every base on the entire fragment during sequencing. Sequencing can be performed using any suitable sequencing technology, wherein nucleotides or oligonucleotides are sequentially added to a free 3′ hydroxyl group, resulting in the synthesis of a polynucleotide chain in the 5′ to 3′ direction. Preferably, the nature of the added nucleotide is determined after each nucleotide addition. Sequencing techniques using ligation-while-sequencing (where not every sequential base is sequenced) and techniques such as massively parallel signature sequencing (MPSS) (where bases are removed from the surface of the chain rather than added to it) are also suitable for the systems and apparatuses of this disclosure.

[0072] In some embodiments, this disclosure discloses sequencing-by-synthesis (SBS). In SBS, four fluorescently labeled modified nucleotides are used to sequence dense clusters (potentially millions) of amplified DNA present on the surface of a substrate (e.g., a flow cell). Various additional aspects of SBS procedures and methods that can be used with the systems and apparatus described herein are disclosed, for example, in WO04018497, WO04018493 and U.S. Patent Nos. 7,057,026 (nucleotides), WO05024010 and WO06120433 (polymerases), WO05065814 (surface attachment technology), and WO 9844151, WO06064199 and WO07010251, the contents of each of which are incorporated herein by reference in their entirety.

[0073] In the specific application of the system / apparatus described herein, a flow cell containing a nucleic acid sample for sequencing is placed within a suitable flow cell holder. The sample for sequencing may be in single-molecule form, amplified single-molecule form in clusters, or in the form of beads containing nucleic acid molecules. The nucleic acids are prepared such that they include oligonucleotide primers adjacent to an unknown target sequence. To initiate the first SBS sequencing cycle, one or more different labeled nucleotides and a DNA polymerase, etc., flow into / through the flow cell via a fluid flow subsystem (various embodiments of which are described herein). A single nucleotide may be added at a time, or the nucleotides used in the sequencing process may be specifically designed to have reversible termination properties, such that each cycle of the sequencing reaction occurs simultaneously in the presence of all four labeled nucleotides (A, C, T, G). With the four nucleotides mixed together, the polymerase is able to select the correct base to incorporate, and each sequence is extended by a single base. In such methods using this system, the natural competition among all four selections results in higher accuracy than if only one nucleotide is present in the reaction mixture (where most sequences are therefore not exposed to the correct nucleotide). A sequence that repeats a specific base one after another (e.g., a homopolymer) can be addressed like any other sequence with high accuracy.

[0074] The fluid flow subsystem also allows suitable reagents to flow to remove the blocked 3′ end (if appropriate) and fluorophore from each incorporated base. The substrate may be exposed to a second round of four blocked nucleotides, or optionally to a second round of different individual nucleotides. Such cycles are then repeated, and the sequence of each cluster is read during these multiple chemical cycles. The computer aspects of this disclosure may optionally align sequence data collected from each monomolecule, cluster, or bead to determine the sequence of longer polymers, etc. Alternatively, image processing and alignment may be performed on a separate computer.

[0075] The system's heating / cooling components regulate reaction conditions within the flow cell channel and reagent storage area / container (and optionally, camera, optics, and / or other components), while the fluid flow components allow the substrate surface to be exposed to suitable reagents for incorporation (e.g., appropriately fluorescently labeled nucleotides to be incorporated), while unincorporated reagents are washed away. An optional movable stage on which the flow cell is placed allows the flow cell to be positioned appropriately for laser (or other light) excitation of the substrate and is optionally movable relative to the lens objective to allow reading of different regions of the substrate. Additionally, other components of the system may optionally be movable / adjustable (e.g., camera, lens objective, heater / cooler, etc.). During laser excitation, the image / position of emission fluorescence from nucleic acids on the substrate is captured by the camera component, thereby recording the identity of the first base of each monomolecule, cluster, or bead in the computer component.

[0076] The embodiments described herein can be used in a variety of biological processes and systems, or chemical processes and systems, for academic or commercial analysis. More specifically, the embodiments described herein can be used in a variety of processes and systems where it is desirable to detect events, properties, quality, or characteristics indicating a desired reaction. For example, the embodiments described herein include cartridges, biosensors, and components thereof, as well as bioassay systems that operate in conjunction with cartridges and biosensors. In certain embodiments, the cartridges and biosensors include a flow cell and one or more sensors, pixels, photodetectors, or photodiodes coupled together in a substantially single structure.

[0077] The following detailed description of certain embodiments will be better understood when read in conjunction with the accompanying drawings. With regard to the diagrams illustrating functional blocks of various embodiments, functional blocks do not necessarily indicate a division between hardware circuits. Thus, for example, one or more functional blocks (e.g., a processor or memory) may be implemented in monolithic hardware (e.g., a general-purpose signal processor or random access memory, hard disk, etc.). Similarly, a program may be a standalone program, may be incorporated as a subroutine into an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and means shown in the drawings.

[0078] As used herein, an element or step described in the singular and preceded by the words "an" or "a" should be understood to not exclude multiple said elements or steps unless such exclusion is expressly indicated. Furthermore, references to "an embodiment" are not intended to be construed as excluding the existence of additional embodiments that also incorporate the described features. Moreover, unless expressly stated to the contrary, embodiments that "comprising," "have," or "including" one or more elements having a particular attribute may include additional elements, whether or not they possess that attribute.

[0079] As used herein, “desired reaction” includes a change in at least one of the chemical, electrical, physical, or optical (or mass) properties of the analyte of interest. In a particular embodiment, the desired reaction is a positive binding event (e.g., the binding of a fluorescently labeled biomolecule to the analyte of interest). More generally, the desired reaction can be a chemical transformation, chemical change, or chemical interaction. The desired reaction can also be a change in electrical properties. For example, the desired reaction can be a change in the concentration of ions in solution. Exemplary reactions include, but are not limited to, chemical reactions such as reduction, oxidation, addition, elimination, rearrangement, esterification, amidation, etherification, cyclization, or substitution; binding interactions of a first chemical substance to a second chemical substance; dissociation reactions in which two or more chemical substances separate from each other; fluorescence; luminescence; bioluminescence; chemiluminescence; and biological reactions such as nucleic acid replication, nucleic acid amplification, nucleic acid hybridization, nucleic acid ligation, phosphorylation, enzyme catalysis, receptor binding, or ligand binding. The desired reaction can also be the addition or elimination of protons, for example, detectable as a change in the pH of the surrounding solution or environment. Additional desired reactions could include detecting ion flow across a membrane (e.g., a natural or synthetic bilayer membrane), where, for example, the current is interrupted as ions flow through the membrane, and this interruption can be detected.

[0080] In a particular embodiment, the desired reaction involves binding a fluorescently labeled molecule to an analyte. The analyte may be an oligonucleotide, and the fluorescently labeled molecule may be a nucleotide. The desired reaction can be detected when excitation light is directed to the oligonucleotide with the labeled nucleotide, and the fluorophore emits a detectable fluorescent signal. In an alternative embodiment, the detected fluorescence is the result of chemiluminescence or bioluminescence. The desired reaction may also be enhanced, for example, by bringing the donor fluorophore closer to the acceptor fluorophore to increase fluorescence (or...). Resonance energy transfer (FRET) reduces FRET by separating the donor fluorophore and the acceptor fluorophore, increases fluorescence by separating the quencher group and the fluorophore, or reduces fluorescence by co-locating the quencher group and the fluorophore.

[0081] As used herein, "reaction component" or "reactant" includes any substance that can be used to obtain the desired reaction. For example, reaction components include reagents, enzymes, samples, other biomolecules, and buffer solutions. Reaction components can typically be delivered to and / or immobilized at reaction sites in solution. Reaction components can interact directly or indirectly with another substance, such as the analyte of interest.

[0082] As used herein, the term "reaction site" is a localized region on which a desired reaction may occur. A reaction site may include a supporting surface of a substrate on which material can be immobilized. For example, a reaction site may include a substantially planar surface in a channel of a flow cell containing a population of nucleic acids. Typically, but not always, the nucleic acids in the population have the same sequence, such as clones of single-stranded or double-stranded templates. However, in some embodiments, the reaction site may contain only a single nucleic acid molecule, such as in single-stranded or double-stranded form. Furthermore, multiple reaction sites may be unevenly distributed along a supporting surface or arranged in a predetermined manner (e.g., side-by-side in a matrix, such as in a microarray). A reaction site may also include a reaction chamber (or pore) that at least partially defines a spatial region or volume configured to separate the desired reaction.

[0083] The terms "reaction chamber" and "orifice" are used interchangeably in this application. As used herein, the terms "reaction chamber" or "orifice" include a spatial region in fluid communication with a flow channel. A reaction chamber may be at least partially isolated from the surrounding environment or other spatial regions. For example, multiple reaction chambers may be separated from each other by a common wall. As a more specific example, a reaction chamber may include a cavity defined by the inner surface of an orifice and may have an opening or pore such that the cavity is in fluid communication with a flow channel. Biosensors including such reaction chambers are described in more detail in International Application No. PCT / US201 I / 057111, filed October 20, 2011, the entire contents of which are incorporated herein by reference.

[0084] In some embodiments, the size and shape of the reaction chamber relative to a solid (including a semi-solid) are configured such that the solid can be fully or partially inserted therein. For example, the size and shape of the reaction chamber may be configured to accommodate only one capture bead. This capture bead may have clonal amplified DNA or other material thereon. Alternatively, the size and shape of the reaction chamber may be configured to receive approximately a number of beads or a solid substrate. Furthermore, the reaction chamber may also be filled with a porous gel or material configured to control diffusion or filter fluids flowing into the reaction chamber.

[0085] In some implementations, a sensor (e.g., a photodetector, photodiode) is associated with a corresponding pixel region on the sample surface of the biosensor. Therefore, a pixel region is a geometry representing the area of ​​a sensor (or pixel) on the biosensor sample surface. When the desired reaction occurs at a reaction site or reaction chamber covering the associated pixel region, the sensor associated with the pixel region detects light emission collected from the associated pixel region. In flat surface implementations, pixel regions may overlap. In some cases, multiple sensors may be associated with a single reaction site or a single reaction chamber. In other cases, a single sensor may be associated with a set of reaction sites or a set of reaction chambers.

[0086] As used herein, a “biosensor” includes a structure having multiple reaction sites and / or reaction chambers (or pores). A biosensor may include a solid-state imaging device (e.g., a CCD or CMOS imaging device) and a flow cell optionally mounted thereon. The flow cell may include at least one flow channel in fluid communication with the reaction sites and / or reaction chambers. As a specific example, a biosensor is configured to be fluidly and electrically coupled to a bioassay system. The bioassay system may deliver reactants to the reaction sites and / or reaction chambers according to a predetermined protocol (e.g., sequencing-by-synthesis) and perform multiple imaging events. For example, the bioassay system may guide a solution to flow along the reaction sites and / or reaction chambers. At least one solution in the solution may include four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides located at the reaction sites and / or reaction chambers. The bioassay system may then illuminate the reaction sites and / or reaction chambers using an excitation light source (e.g., a solid-state light source, such as a light-emitting diode (LED)). The excitation light may have one or more predetermined wavelengths, including a wavelength range. The excited fluorescent label provides an emission signal that can be captured by the sensor.

[0087] In alternative embodiments, the biosensor may include electrodes or other types of sensors configured to detect other identifiable properties. For example, the sensor may be configured to detect changes in ion concentration. In another example, the sensor may be configured to detect transmembrane ion currents.

[0088] As used herein, a “cluster” is a group of similar or identical molecular or nucleotide sequences or DNA strands. For example, a cluster can be an amplified oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequences. In other embodiments, a cluster can be any element or group of elements occupying a physical region on a sample surface. In embodiments, clusters are immobilized to reaction sites and / or reaction chambers during base detection cycles.

[0089] As used herein, when referring to biomolecules or biological or chemical substances, the term "fixed" includes the attachment of biomolecules or biological or chemical substances to a surface substantially at the molecular level. For example, adsorption techniques can be used to fix biomolecules or biological or chemical substances to the surface of a substrate material. These adsorption techniques include non-covalent interactions (e.g., electrostatic forces, van der Waals forces, and dehydration at hydrophobic interfaces) and covalent bonding techniques, where functional groups or connectors facilitate the attachment of biomolecules to the surface. The fixation of biomolecules or biological or chemical substances to the surface of a substrate material can be based on the properties of the substrate surface, the liquid medium carrying the biomolecules or biological or chemical substances, and the properties of the biomolecules or biological or chemical substances themselves. In some cases, the substrate surface can be functionalized (e.g., chemically or physically modified) to facilitate the fixation of biomolecules (or biological or chemical substances) to the substrate surface. The substrate surface can be modified first to allow functional groups to bind to the surface. The functional groups can then bind to the biomolecules or biological or chemical substances to fix them thereto. Substances can be fixed to surfaces via gels, for example, as described in U.S. Patent Publication US2011 / 0059865A1, which is incorporated herein by reference.

[0090] In some embodiments, nucleic acids may be attached to a surface and amplified using bridging amplification. Useful bridging amplification methods are described, for example, in U.S. Patent 5,641,658; WO 2007 / 010251; U.S. Patent 6,090,592; U.S. Patent Publication 2002 / 0055100A1; U.S. Patent 7,115,400; U.S. Patent Publication 2004 / 0096853A1; U.S. Patent Publication 2004 / 0002090A1; U.S. Patent Publication 2007 / 0128624A1; and U.S. Patent Publication 2008 / 0009420A1, each of which is incorporated herein by reference in its entirety. Another useful method for amplifying nucleic acids on a surface is rolling circle amplification (RCA), for example, using methods further detailed below. In some embodiments, nucleic acids may be attached to a surface and amplified using one or more primer pairs. For example, one primer may be in solution, and the other primer may be immobilized on the surface (e.g., 5'-attached). By way of example, a nucleic acid molecule can hybridize with one of the primers on a surface, followed by extension of the immobilized primers to produce a first copy of the nucleic acid. The primers in solution then hybridize with the first copy of the nucleic acid, which can be extended using the first copy of the nucleic acid as a template. Optionally, after producing the first copy of the nucleic acid, the original nucleic acid molecule can hybridize with a second immobilized primer on the surface, and this extension can occur simultaneously with or after the extension of the primers in solution. In any embodiment, repeating one cycle (e.g., amplification) using immobilized primers and primers in solution provides multiple copies of the nucleic acid.

[0091] In certain embodiments, the assay protocol performed by the systems and methods described herein includes the use of natural nucleotides and enzymes configured to interact with the natural nucleotides. Natural nucleotides include, for example, ribonucleotides (RNA) or deoxyribonucleotides (DNA). Natural nucleotides may be in the form of monophosphate, diphosphate, or triphosphate, and may have bases selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it should be understood that non-natural nucleotides, modified nucleotides, or analogs of the aforementioned nucleotides may be used. Some examples of useful non-natural nucleotides are listed below regarding sequencing based on reversible terminators performed by synthetic methods.

[0092] In embodiments including a reaction chamber, articles or solid substances (including semi-solid substances) may be disposed within the reaction chamber. When disposed, the articles or solids may be physically held or fixed within the reaction chamber by interference fit, adhesion, or trapping. Exemplary articles or solids that may be disposed within the reaction chamber include polymer beads, microspheres, agarose gels, powders, quantum dots, or other solids that can be compressed and / or retained within the reaction chamber. In certain embodiments, nucleic acid superstructures (such as DNA spheres) may be disposed within or at the reaction chamber, for example, by attaching to the inner surface of the reaction chamber or by remaining in a liquid within the reaction chamber. DNA spheres or other nucleic acid superstructures may be constructed and then disposed within or at the reaction chamber. Alternatively, DNA spheres may be synthesized in situ at the reaction chamber. DNA spheres can be synthesized by rolling circle amplification to produce polymers of specific nucleic acid sequences, and polymers can be processed using conditions that form relatively compact spheres. DNA spheres and methods for their synthesis are described, for example, in U.S. Patent Publications 2008 / 0242560A1 or 2008 / 0234136A1, each of which is incorporated herein by reference in its entirety. The substance held or disposed in the reaction chamber may be in a solid, liquid, or gaseous state.

[0093] As used herein, "base detection" identifies nucleotide bases in a nucleic acid sequence. Base detection refers to the process of determining the base detection (A, C, G, T) for each cluster in a specific cycle. As an example, base detection can be performed using a four-channel method and system, a two-channel method and system, or a one-channel method and system described in the combined material of U.S. Patent Application Publication 2013 / 0079232. In a particular embodiment, the base detection cycle is referred to as a "sampling event." In a dye and two-channel sequencing protocol, a sampling event comprises two illumination phases in a time series, such that pixel signals are generated at each phase. A first illumination phase induces illumination from a given cluster indicating nucleotide bases A and T in the AT pixel signal, and a second illumination phase induces illumination from a given cluster indicating nucleotide bases C and T in the CT pixel signal.

[0094] The disclosed techniques (e.g., the disclosed base detector) can be implemented on processors such as central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), coarse-grained reconfigurable architectures (CGRAs), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), and digital signal processors (DSPs).

[0095] Biosensors

[0096] Figure 1 A cross-section of a biosensor 100 that can be used in various embodiments is shown. The biosensor 100 has pixel regions 106′, 108′, 110′, 112′, and 114′, each of which can maintain more than one cluster (e.g., two clusters per pixel region) during base detection cycles. As shown, the biosensor 100 may include a flow cell 102 mounted to a sampling device 104. In the illustrated embodiment, the flow cell 102 is directly attached to the sampling device 104. However, in an alternative embodiment, the flow cell 102 may be removably coupled to the sampling device 104. The sampling device 104 has a sample surface 134 that can be functionalized (e.g., chemically or physically modified in a manner suitable for carrying out a desired reaction). For example, the sample surface 134 may be functionalized and may include multiple pixel regions 106′, 108′, 110′, 112′, and 114′, each of which may maintain more than one cluster during a base detection cycle (e.g., each pixel region has a corresponding cluster pair 106A, 106B; 108A, 108B; 110A, 110B; 112A, 112B; and 114A, 114B fixed thereon). Each pixel region is associated with a corresponding sensor (or pixel or photodiode) 106, 108, 110, 112, and 114, such that light received by the pixel region is captured by the corresponding sensor. Pixel region 106′ may also be associated with a corresponding reaction site 106″ on the sample surface 134 that maintains a cluster pair, such that light emitted from reaction site 106″ is received by pixel region 106′ and captured by the corresponding sensor 106. Due to this sensing structure, the pixel signal in the base detection cycle carries information based on all of the two or more clusters in the following case: where, during the base detection cycle, there are two or more clusters in a pixel region of a particular sensor (e.g., each pixel region has a corresponding cluster pair). Therefore, signal processing as described herein is used to distinguish each cluster where there are more clusters than pixel signals in a given sampling event of a particular base detection cycle.

[0097] In an illustrated embodiment, flow cell 102 includes sidewalls 138, 125 and a flow hood 136 supported by the sidewalls 138, 125. The sidewalls 138, 125 are coupled to the sample surface 134 and extend between the flow hood 136 and the sidewalls 138, 125. In some embodiments, the sidewalls 138, 125 are formed of a curable adhesive layer that bonds the flow hood 136 to the sampling device 104.

[0098] The dimensions and shapes of the sidewalls 138, 125 are configured such that a flow channel 144 exists between the shroud 136 and the sampling device 104. The shroud 136 may include a material transparent to excitation light 101 propagating from the outside of the biosensor 100 into the flow channel 144. In this example, the excitation light 101 approaches the shroud 136 at a non-orthogonal angle.

[0099] Additionally, as shown in the figure, the flow shroud 136 may include inlet and outlet ports 142, 146, which are configured to fluidly engage with other ports (not shown). These other ports may, for example, originate from a cartridge or workstation. The dimensions and shape of the flow channel 144 are configured to guide fluid along the sample surface 134. The height H1 and other dimensions of the flow channel 144 may be configured to maintain substantially uniform fluid flow along the sample surface 134. The dimensions of the flow channel 144 may also be configured to control bubble formation.

[0100] By way of example, the flow shroud 136 (or flow cell 102) may comprise a transparent material, such as glass or plastic. The flow shroud 136 may be formed as a substantially rectangular block having a planar outer surface and a planar inner surface defining a flow channel 144. This block may be mounted onto sidewalls 138, 125. Alternatively, the flow cell 102 may be etched to define the flow shroud 136 and the sidewalls 138, 125. For example, grooves may be etched into the transparent material. When the etched material is mounted onto the sampling device 104, the grooves may become the flow channel 144.

[0101] The sampling device 104 may resemble, for example, an integrated circuit comprising multiple stacked substrate layers 120 to 126. Substrate layers 120 to 126 may include a substrate 120, a solid-state imaging device 122 (e.g., a CMOS image sensor), a filter or light control layer 124, and a passivation layer 126. It should be noted that the above is illustrative only, and other embodiments may include fewer layers or additional layers. Furthermore, each of the substrate layers 120 to 126 may include multiple sublayers. The sampling device 104 may be fabricated using processes similar to those used in manufacturing integrated circuits such as CMOS image sensors and CCDs. For example, substrate layers 120 to 126, or portions thereof, may be grown, deposited, etched, etc., to form the sampling device 104.

[0102] The passivation layer 126 is configured to shield the fluid environment of the flow channel 144 from the filter layer 124. In some cases, the passivation layer 126 is also configured to provide a solid surface (i.e., sample surface 134) that allows biomolecules or other analytes of interest to be immobilized thereon. For example, each reaction site may include a cluster of biomolecules immobilized to the sample surface 134. Therefore, the passivation layer 126 may be formed of a material that allows reaction sites to be immobilized thereon. The passivation layer 126 may also include a material that is at least transparent to the desired fluorescence. By way of example, the passivation layer 126 may comprise silicon nitride (Si2N4) and / or silicon dioxide (SiO2). However, other suitable materials may be used. In the illustrated embodiment, the passivation layer 126 may be substantially planar. However, in an alternative embodiment, the passivation layer 126 may include grooves, such as pits, holes, trenches, etc. In the illustrated embodiment, the passivation layer 126 has a thickness of about 150 nm to 200 nm, and more specifically about 170 nm.

[0103] Filter layer 124 may include various features that affect the transmission of light. In some embodiments, filter layer 124 may perform multiple functions. For example, filter layer 124 may be configured to (a) filter unwanted optical signals, such as those from an excitation source; (b) direct emission signals from reaction sites to corresponding sensors 106, 108, 110, 112, and 114, which are configured to detect emission signals from reaction sites; or (c) block or prevent the detection of unwanted emission signals from adjacent reaction sites. Thus, filter layer 124 may also be referred to as a light control layer. In an illustrated embodiment, filter layer 124 has a thickness of about 1 μm to 5 μm, more specifically about 2 μm to 4 μm. In an alternative embodiment, filter layer 124 may include an array of microlenses or other optical elements. Each microlens may be configured to direct emission signals from the associated reaction site to the sensor.

[0104] In some embodiments, the solid-state imaging device 122 and the substrate 120 may be provided together with a previously constructed solid-state imaging device (e.g., a CMOS chip). For example, the substrate 120 may be a silicon wafer, and the solid-state imaging device 122 may be mounted thereon. The solid-state imaging device 122 includes a semiconductor material (e.g., silicon) layer and sensors 106, 108, 110, 112, and 114. In an illustrated embodiment, the sensor is a photodiode configured to detect light. In other embodiments, the sensor includes a photodetector. The solid-state imaging device 122 may be fabricated as a single chip using a CMOS-based manufacturing process.

[0105] Solid-state imaging device 122 may include a dense array of sensors 106, 108, 110, 112, and 114, which are configured to detect activity indicating a desired response from or along the flow channel 144. In some embodiments, each sensor has a size of approximately 1 square micrometer to 2 square micrometers (μm). 2 The array can include 500,000 sensors, 5 million sensors, 10 million sensors, or even 120 million sensors. Sensors 106, 108, 110, 112, and 114 can be configured to detect light of a predetermined wavelength that indicates a desired response.

[0106] In some embodiments, sampling device 104 includes a microcircuit arrangement, such as that described in U.S. Patent No. 7,595,882, which is incorporated herein by reference in its entirety. More specifically, sampling device 104 may include an integrated circuit having a planar array of sensors 106, 108, 110, 112, and 114. Circuitry formed within sampling device 104 may be configured for at least one of signal amplification, digitization, storage, and processing. The circuitry may collect and analyze detected fluorescence and generate pixel signals (or detection signals) for transmitting detection data to a signal processor. The circuitry may also perform additional analog and / or digital signal processing within sampling device 104. Sampling device 104 may include conductive vias 130 that perform signal routing (e.g., transmitting pixel signals to a signal processor). Pixel signals may also be transmitted via electrical contacts 132 of sampling device 104.

[0107] The sampling device 104 is discussed in further detail in relation to U.S. Nonprovisional Patent Application No. 16 / 874,599 (Attorney General's File No. ILLM 1011-4 / IP-1750-US), filed May 14, 2020, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," which is incorporated herein by reference as if fully set forth herein. The sampling device 104 is not limited to the construction or use described above. In alternative embodiments, the sampling device 104 may take other forms. For example, the sampling device 104 may include a CCD device (such as a CCD camera) coupled to or movable to interact with a flow cell containing reaction sites.

[0108] Figure 2 A specific implementation of a circulation pool 200 containing clusters in its blocks is shown. Circulation pool 200 corresponds to... Figure 1The flow pool 102, for example, does not have a flow cover 136. Furthermore, the depiction of the flow pool 200 is symbolic in nature, and the flow pool 200 symbolically depicts the various channels and blocks within it, while the various other components within it are not shown. Figure 2 A top view of the flow cell 200 is shown.

[0109] In one implementation, the flow pool 200 is divided or partitioned into multiple channels, such as channels 202a, 202b, ..., 202P, i.e., P channels. Figure 2 In the example, flow cell 200 is shown as including 8 channels, that is, P = 8 in this example, but the number of channels in the flow cell is implementation-specific.

[0110] In one implementation, each channel 202 is further partitioned into non-overlapping regions referred to as “blocks” 212. For example, Figure 2 An enlarged view of segment 208 of an exemplary channel is shown. Segment 208 is shown as comprising multiple blocks 212.

[0111] In the example, each channel 202 includes one or more block columns. For example, in Figure 2 In this configuration, each channel 202 includes two corresponding block columns 212, as shown in the enlarged section 208. The number of blocks in each block column within each channel is implementation-specific, and in one example, each block column within each channel may contain 50 blocks, 60 blocks, 100 blocks, or another suitable number of blocks.

[0112] Each block comprises multiple corresponding clusters. During sequencing, the clusters within the block and their surrounding background are imaged. For example, Figure 2 Example cluster 216 is shown within the example block.

[0113] Figure 3 An exemplary Illumina GA-IIx with eight channels is shown. TM The flow pool is shown, along with a magnified view of a block, its clusters, and their surrounding background. For example, each channel in the Illumina Genome Analyzer II contains one hundred blocks, and each channel in the Illumina HiSeq2000 contains sixty-eight blocks. Block 212 accommodates hundreds of thousands to millions of clusters. Figure 3In the image, at 308 (e.g., 308 is a magnified image view of the block), an image generated from a block having clusters shown as bright spots is shown, with an exemplary cluster 304 marked. Cluster 304 comprises approximately one thousand identical copies of the template molecule, but the clusters differ in size and shape. Clusters are generated from the template molecule by bridging amplification of the input library prior to sequencing. The purpose of amplification and cluster growth is to increase the intensity of the emission signal, as imaging devices cannot reliably sense individual fluorophores. However, the DNA fragments within cluster 304 are physically close together, so the imaging device perceives the cluster of fragments as a single point 304.

[0114] Clusters and blocks are discussed in further detail compared to U.S. non-provisional patent application No. 16 / 825,987 (Attorney’s File No. ILLM 1008-16 / IP-1693-US) entitled “TRAINING DATA GENERATION FORARTIFICIAL INTELLIGENCE-BASED SEQUENCING” filed on March 20, 2020.

[0115] Figure 4 It is used to analyze sensor data from sequencing systems (such as base detection sensor outputs (e.g., see...)). Figure 1 A simplified block diagram of the system. Figure 4 In this example, the system includes a sequencing machine 400 and a configurable processor 450. The configurable processor 450 can execute a neural network-based base detector in coordination with a runtime program executed by a host processor, such as a central processing unit (CPU) 402. The sequencing machine 400 includes a base detection sensor and a flow cell 401 (e.g., relative to...). Figures 1 to 3 (As discussed). A flow cell may include one or more blocks in which clusters of genetic material are exposed to a sequence of analytical streams, the sequence of which is used to elicit a reaction in the clusters to recognize bases in the genetic material, such as relative to... Figures 1 to 3 The sensor senses the response of each cycle of the sequence in each block of the flow cell to provide block data. An example of this technique is described in more detail below. Genetic sequencing is a data-intensive operation that converts base detection sensor data into base detection sequences for each cluster of genetic material sensed during the base detection operation.

[0116] The system in this example includes a CPU 402 that executes runtime programs to coordinate base detection operations, a memory 403 for storing sequences of block data arrays, base detection reads generated by the base detection operations, and other information used in the base detection operations. Additionally, in this illustration, the system includes a memory 404 for storing a configuration file (or multiple files) such as an FPGA bit file, and model parameters for configuring and reconfiguring a configurable processor 450 and executing a neural network. The sequencing machine 400 may include programs for configuring the configurable processor and, in some embodiments, a reconfigurable processor to execute the neural network.

[0117] Sequencing machine 400 is coupled to configurable processor 450 via bus 405. Bus 405 can be implemented using high-throughput technologies, such as, in one example, bus technologies compatible with the PCIe standard (Rapid Peripheral Component Interconnect) currently maintained and developed by PCI-SIG (PCI Special Interest Group). Additionally, in this example, memory 460 is coupled to configurable processor 450 via bus 461. Memory 460 can be onboard memory disposed on a circuit board having configurable processor 450. Memory 460 is used by configurable processor 450 for high-speed access to working data used in base detection operations. Bus 461 can also be implemented using high-throughput technologies such as bus technologies compatible with the PCIe standard.

[0118] Configurable processors, including field-programmable gate arrays (FPGAs), coarse-grained reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to perform a variety of functions more efficiently or faster than what might be possible using a general-purpose processor that executes a computer program. Configuring a configurable processor involves compiling a function description to produce a configuration file, sometimes called a bitstream or bit file, and distributing the configuration file to configurable elements on the processor.

[0119] The configuration file defines the logical functions to be executed by the configurable processor by configuring the circuitry to set data flow modes, the use of distributed memory and other on-chip memory resources, lookup table contents, and the operation of configurable logic blocks and configurable execution units (such as multiply-accumulate units, configurable interconnects, and other elements of the configurable array). A configurable processor is reconfigurable if the configuration file can be changed in the field by changing the loaded configuration file. For example, the configuration file can be stored in volatile SRAM elements, non-volatile read-write memory elements, and combinations thereof, distributed across an array of configurable elements on the configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for base detection operations as described herein. Examples include commercially available products such as Xilinx Alveo. TM U200, Xilinx AlveoTM U250, Xilinx Alveo TM U280, Intel / Altera Stratix TM GX2800, Intel / AlteraStratix TM GX2800 and Intel Stratix TM GX10M. In some examples, the host CPU can be implemented on the same integrated circuit as the configurable processor.

[0120] The implementation described herein uses a configurable processor 450 to implement a multi-recurrent neural network. The configuration file for the configurable processor can be implemented by specifying the logic functions to be performed using a high-level description language (HDL) or register-transfer level (RTL) language specification. The specification can be compiled to generate the configuration file using resources designed for the selected configurable processor. The same or similar specifications can be compiled to generate designs for application-specific integrated circuits (ASICs) that may not be configurable processors.

[0121] Therefore, in all embodiments described herein, alternatives to configurable processors include configured processors comprising application-specific ASICs or application-specific integrated circuits or integrated circuit groups, or system-on-a-chip (SoC) devices, configured to perform neural network-based base detection operations as described herein.

[0122] Generally speaking, the configurable processor and the configured processor described herein, which are configured to execute the operation of a neural network, are referred to herein as neural network processors.

[0123] In this example, the configurable processor 450 is configured using a configuration file loaded by a program executed by CPU 402 or another source, which configures an array of configurable elements on the configurable processor 454 to perform a base detection function. In this example, the configuration includes data flow logic 451 coupled to buses 405 and 461, and performs functions for distributing data and control parameters among the elements used in the base detection operation.

[0124] Additionally, the configurable processor 450 is configured with base detection execution logic 452 to execute a multi-recurrent neural network. Logic 452 includes multiple multi-recurrent execution clusters (e.g., 453), which in this example include multi-recurrent cluster 1 to multi-recurrent cluster X. The number of multi-recurrent clusters can be selected based on a trade-off between the required throughput of the operation and the available resources on the configurable processor.

[0125] The multiple cyclic clusters are coupled to the data flow logic 451 via data flow path 454, implemented using configurable interconnects and memory resources on a configurable processor. Additionally, the multiple cyclic clusters are coupled to the data flow logic 451 via control path 455, implemented, for example, using configurable interconnects and memory resources on a configurable processor. These control paths provide control signals indicating available clusters, preparing to provide input units to the available clusters for executing the neural network, preparing to provide trained parameters for the neural network, preparing to provide output patches for base detection classification data, and other control data for executing the neural network.

[0126] A configurable processor is configured to execute the operation of a multi-recurrent neural network using trained parameters to generate classification data for sensing cycles of a base stream operation. The neural network is executed to generate classification data for a subject sensing cycle used for a base detection operation. The neural network operates on a sequence comprising a number of N arrays of block data from corresponding sensing cycles in N sensing cycles, where the N sensing cycles, in the examples described herein, provide sensor data for a different base detection operation at a base position for each operation in the time series. Optionally, some of the N sensing cycles may be out of order, depending on the specific neural network model being executed, if desired. The number N can be any number greater than 1. In some examples described herein, the sensing cycles in the N sensing cycles represent a set of sensing cycles that precede at least one sensing cycle and follow at least one sensing cycle in the time series. Examples described herein where the number N is an integer equal to or greater than five are given.

[0127] Data flow logic 451 is configured to move block data and at least some trained parameters of the model from memory 460 to a configurable processor for running the neural network using input units (including block data of spatially aligned patches of N arrays) for a given run. Input units can be moved via a direct memory access operation in a DMA operation, or within smaller units that move in coordination with the execution of the deployed neural network during available time slots.

[0128] Block data for sensing cycles, as described herein, may include an array of sensor data having one or more features. For example, the sensor data may include two images, which are analyzed to identify one of four bases at a base position in the genetic sequence of DNA, RNA, or other genetic material. The block data may also include metadata about the images and the sensors. For example, in an embodiment of a base detection operation, the block data may include information about the alignment of the image with the clusters, such as distance from the center, indicating the distance of each pixel in the sensor data array from the center of the cluster of genetic material on the block.

[0129] During the execution of the multi-recurrent neural network as described below, block data may also include data generated during the execution of the multi-recurrent neural network, referred to as intermediate data, which can be reused rather than recomputed during the operation of the multi-recurrent neural network. For example, during the execution of the multi-recurrent neural network, dataflow logic may write intermediate data into memory 460 instead of sensor data for a given patch of the block data array. An embodiment similar to this is described in more detail below.

[0130] As shown in the figure, a system for analyzing the output of a base detection sensor is described. This system includes a memory (e.g., 460) accessible by a runtime program, which stores block data, including sensor data from blocks of sensing loops of a base detection operation. Additionally, the system includes a neural network processor, such as a configurable processor 450 with access to the memory. The neural network processor is configured to execute the operation of a neural network using trained parameters to generate classification data for the sensing loops. As described herein, the operation of the neural network operates on a sequence of N arrays of block data from the corresponding sensing loops (including the subject loop) of N sensing loops to generate classification data for the subject loop. Data flow logic 451 is provided to move the block data and trained parameters from the memory to the neural network processor for the operation of the neural network using input units (including data from spatially aligned patches of the N arrays of the corresponding sensing loops of the N sensing loops).

[0131] Additionally, a system is described in which a neural network processor is capable of accessing memory and includes multiple execution clusters, among which execution logic clusters are configured to execute a neural network. Dataflow logic is capable of accessing memory and the execution clusters within the multiple execution clusters to provide input units of block data to available execution clusters within the multiple execution clusters. These input units include N digital spatial alignment patches from an array of block data from a corresponding sensing loop (including a subject sensing loop), and cause the execution clusters to apply the N spatial alignment patches to the neural network to produce output patches of classification data for the spatial alignment patches of the subject sensing loop, where N is greater than 1.

[0132] Figure 5 This is a simplified diagram illustrating one aspect of the base detection operation, which includes the functionality of the runtime program executed by the host processor. In this diagram, data from the flow cell (such as...) Figures 1 to 2The output of the image sensor (shown in the flow cell) is provided on line 500 to image processing thread 501, which performs image processing such as resampling, alignment, and arrangement in the sensor data array of individual blocks, and can be used by a process that calculates a cluster mask for each block in the flow cell, identifying pixels in the sensor data array corresponding to clusters of genetic material on the corresponding block of the flow cell. To calculate the cluster mask, an exemplary algorithm is based on a process for detecting unreliable clusters in early sequencing cycles using a metric derived from the softmax output, then discarding data from those traps / clusters and not generating output data for those clusters. For example, the process could identify clusters with high reliability during the first N1 (e.g., 25) base detections and reject other clusters. Rejected clusters might be polyclonal, very weak, or due to reference point ambiguity. This procedure can be executed on the host CPU. In an alternative embodiment, this information could potentially be used to identify the necessary clusters of interest to be returned to the CPU, thereby limiting the storage required for intermediate data.

[0133] Based on the status of the base detection operation, the output of image processing thread 501 is provided on line 502 to scheduling logic 510 in the CPU. This scheduling logic routes the block data array to data cache 504 on high-speed bus 503, or to multi-cluster neural network processor hardware 520 on high-speed bus 505, such as... Figure 4 The configurable processor 520 returns the classification data output by the neural network to scheduling logic 510, which passes the information to data cache 504, or to thread 502 on line 511 to perform base detection and quality score calculation using the classification data, and can arrange the data for base detection reads in a standard format. On line 512, the output of thread 502 performing base detection and quality score calculation is provided to thread 503, which aggregates the base detection reads, performs other operations such as data compression, and writes the resulting base detection output to a specified destination for client use.

[0134] In some implementations, the host may include final processing of the output of hardware 520 to support threads (not shown) of the neural network. For example, hardware 520 may provide output of classification data from the final layer of a multi-cluster neural network. The host processor may perform output activation functions such as a softmax function on the classification data to configure the data for use by the base detection and quality scoring threads 502. Additionally, the host processor may perform input operations (not shown), such as resampling, batch normalizing, or other adjustments to the block data before it is input to hardware 520.

[0135] Figure 6 It is a configurable processor (such as, Figure 4A simplified diagram of the configuration of the configurable processor. Figure 6 The configurable processor includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 600, which includes a reference... Figure 1 The data flow logic is described. The package 600 manages the interface and coordination with the runtime program in the CPU via CPU communication link 609, and manages communication with the onboard DRAM 602 (e.g., memory 460) via DRAM communication link 610. The data flow logic in the package 600 provides patch data retrieved by traversing a block data array of N cycles on the onboard DRAM 602 to cluster 601, and retrieves process data 615 from cluster 601 for delivery back to the onboard DRAM 602. The package 600 also manages data transfer between the onboard DRAM 602 and the host memory for both the input array of block data and the output patches of classified data. The package transfers patch data on line 613 to the assigned cluster 601. The package provides trained parameters such as weights and biases on line 612 to the cluster 601 retrieved from the onboard DRAM 602. On line 611, the packager provides configuration and control data to cluster 601, which is provided from or generated in response to a runtime program on the host via CPU communication link 609. The cluster can also provide status signals to packager 600 on line 616, which cooperate with control signals from the host to manage the traversal of the block data array, thereby providing spatially aligned patch data and using the resources of cluster 601 to perform multi-recurrent neural networks on the patch data.

[0136] As described above, multiple clusters may exist on a single configurable processor managed by the wrapper 600, and these clusters are configured to execute on corresponding patches of multiple patches of block data. Each cluster may be configured to provide classification data for base detection in a subject sensing cycle using block data from multiple sensing cycles as described herein.

[0137] In system examples, model data (including kernel data such as filter weights and biases) can be sent from the host CPU to a configurable processor, allowing the model to be updated based on the number of loops. As a representative example, a base detection operation may include approximately several hundred sensing loops. In some implementations, the base detection operation may include paired-end reads. For example, model training parameters may be updated every 20 loops (or another number of loops), or according to an update pattern implemented for a particular system and neural network model. In some implementations that include paired-end reads, where the sequence of a given string in a genetic cluster on a block comprises a first portion extending down (or up) along the string from a first end and a second portion extending up (or down) along the string from a second end, trained parameters may be updated during the transition from the first portion to the second portion.

[0138] In some examples, image data from multiple cycles of sensing data for a block can be sent from the CPU to the packager 600. The packager 600 may optionally perform some preprocessing and transformation on the sensing data and write the information to onboard DRAM 602. The input block data for each sensing cycle may include a sensor data array comprising approximately 4000 × 3000 pixels or more per block per sensing cycle, where two features represent the colors of two images of the block, and each feature is one or two bytes per pixel. For an implementation where the number N is three sensing cycles to be used in each run of the multi-recurrent neural network, the block data array for each run of the multi-recurrent neural network may consume approximately several hundred megabytes per block. In some implementations of the system, the block data may also include an array of DFC data stored once per block, or other types of metadata about the sensor data and the block.

[0139] In operation, when a multi-cycle cluster becomes available, the wrapper allocates a patch to the cluster. The wrapper retrieves the next patch of block data during block traversal and sends it, along with appropriate control and configuration information, to the allocated cluster. The cluster can be configured to have sufficient memory on a configurable processor to hold data patches, including patches from multiple cycles in several systems that are being processed in-situ, as well as data patches to be processed when processing of the current patch is completed using ping-pong buffering or raster scanning techniques in various implementations.

[0140] When an assigned cluster completes its neural network operation on the current patch and produces an output patch, it signals the packager. The packager reads the output patch from the assigned cluster, or alternatively, the assigned cluster pushes data to the packager. The packager then assembles the output patch from the processed block in DRAM 602. When the processing of the entire block is complete and the output patch of data has been transferred to DRAM, the packager sends the processed output array of the block back to the host / CPU in a specified format. In some implementations, the onboard DRAM 602 is managed by memory management logic in the packager 600. The runtime program can control the sequencing operation to perform analysis of all arrays of block data in all loops during operation in a continuous streaming manner, thereby providing real-time analysis.

[0141] Figure 7 This is a graph of a multi-recurrent neural network model that can be executed using the system described in this article. Figure 7 The example shown can be called a five-loop-input, one-loop-output neural network. The input to a multi-loop neural network model consists of five spatially aligned patches (e.g., 700) of block data arrays from five sensing loops of a given block. The spatially aligned patches have the same alignment row and column dimensions (x, y) as the other patches in the set, such that the information pertains to the same clusters of genetic material on the blocks in the sequence loops. In this example, the subject patch is a patch from the block data array of loop K. A set of five spatially aligned patches includes a patch from loop K-2, two loops before the subject patch; a patch from loop K-1, one loop before the subject patch; a patch from loop K+1, one loop after the patch from the subject loop; and a patch from loop K+2, two loops after the patch from the subject loop.

[0142] The model includes an isolated stack 701 of layers of the neural network for each input patch. Therefore, stack 701 receives block data from the patch of cycle K+2 as input and is isolated from stacks 702, 703, 704, and 705 such that they do not share input or intermediate data. In some embodiments, all stacks 710 to 705 may have the same model and the same trained parameters. In other embodiments, the model and trained parameters may differ in different stacks. Stack 702 receives block data from the patch of cycle K+1 as input. Stack 703 receives block data from the patch of cycle K as input. Stack 704 receives block data from the patch of cycle K-1 as input. Stack 705 receives block data from the patch of cycle K-2 as input. Each layer of the isolated stack performs a convolution operation of a kernel, which includes multiple filters on the layer's input data. As in the example above, patch 700 may include three features. The output of layer 710 may include more features, such as 10 to 20 features. Similarly, the output of each layer in layers 711 to 716 may include any number of features suitable for a particular implementation. The parameters of the filters are the trained parameters of the neural network, such as weights and biases. The output feature sets (intermediate data) from each stack in stacks 701-705 are provided as input to the inverse hierarchy 720 of the temporal combination layer, where intermediate data from multiple loops are combined. In the illustrated example, the inverse hierarchy 720 includes: a first layer comprising three combination layers 721, 722, and 723, each combination layer receiving intermediate data from the three isolated stacks in the isolated stacks; and a final layer comprising a combination layer 730 that receives intermediate data from the three temporal layers 721, 722, and 723.

[0143] The output of the final combining layer 730 is an output patch of classification data for clusters located in the corresponding patches of the blocks from cycle K. The output patches can be assembled into an output array of classification data for the blocks of cycle K. In some embodiments, the output patches may have a different size and dimension than the input patches. In some embodiments, the output patches may include pixel-wise data that can be host-filtered to select cluster data.

[0144] Depending on the specific implementation, the output classification data can then be applied to a softmax function 740 (or other output activation function) optionally executed by the host or on a configurable processor. Output functions different from softmax can be used (e.g., generating base detection output parameters based on the maximum output, then using a learned nonlinear mapping with context / network output to give base quality).

[0145] Finally, the output of the softmax function 740 can be provided as the base detection probability (750) of loop K and stored in host memory for use in subsequent processing. Other systems can use another function for calculating the output probability, such as another nonlinear model.

[0146] A neural network can be implemented using a configurable processor with multiple execution clusters to complete the evaluation of a block loop over a duration equal to or close to the time interval of a sensing loop, thereby efficiently providing output data in real time. Dataflow logic can be configured to distribute block data and input units with trained parameters to the execution clusters, and to distribute output patches for aggregation in memory.

[0147] refer to Figure 8A and Figure 8B The following describes the base detection operation using dual-channel sensor data: Figure 7 The same five-loop input, one-loop output neural network data input unit. For example, for a given base in a gene sequence, a base detection operation can perform two analytical streams and two reactions that generate two signal (such as images) channels. These images can be processed to identify which of the four bases is located at the current position of the genetic sequence in each cluster of genetic material. In other systems, different numbers of channels of sensed data can be utilized. For example, a one-channel method and system can be used to perform base detection. The combined material of U.S. Patent Application Publication No. 2013 / 0079232 discusses base detection using various numbers of channels (such as one-channel, two-channel, or four-channel).

[0148] Figure 8A An array of block data for a given block (block M) is shown, consisting of five loops, used for executing a five-loop input, one-loop output neural network. The five-loop input block data in this example can be written to onboard DRAM or other memory in the system accessible by data flow logic. For loop K-2, arrays 801 for channel 1 and 811 for channel 2 are included; for loop K-1, arrays 802 for channel 1 and 812 for channel 2 are included; for loop K, arrays 803 for channel 1 and 813 for channel 2 are included; for loop K+1, arrays 804 for channel 1 and 814 for channel 2 are included; and for loop K+2, arrays 805 for channel 1 and 815 for channel 2 are included. Additionally, an array 820 of the block's metadata can be written to memory once, in this case including a DFC file to be used as input to the neural network along with each loop.

[0149] although Figure 8ATwo-channel base detection operations have been discussed, but the use of two channels is merely an example, and any other suitable number of channels can be used to perform base detection. For example, the combined material of U.S. Patent Application Publication No. 2013 / 0079232 discusses base detection using various numbers of channels (such as one channel, two channels, or four channels, or another suitable number of channels).

[0150] The data flow logic constitutes the input units of the block data, and these input units can be referenced. Figure 8B Understanding that the block data comprises spatially aligned patches of block data arrays for each execution cluster, each execution cluster being configured to run a neural network on input patches. The input units for the allocated execution clusters are configured by data flow logic to read spatially aligned patches (e.g., 851, 852, 861, 862, 870) from each of the five input loop block data arrays 801-805, 811, 815, 820, and deliver them via a data path (illustratively, 850) to memory on a configurable processor configured for use by the allocated execution clusters. The allocated execution clusters run a five-loop input / one-loop output neural network and deliver output patches of classification data for the same patches of blocks in subject loop K for subject loop K.

[0151] Figure 9 Is it like this? Figure 7 A simplified representation of a stack of neural networks that can be used in a system like (e.g., 701 and 720). In this example, some functions of the neural network (e.g., 900, 902) are executed on the host, while other parts of the neural network (e.g., 901) are executed on a configurable processor.

[0152] In one example, the first function could be a batch normalization (layer 910) formed on the CPU. However, in another example, the batch normalization as a function could be merged into one or more layers, and there might not be a separate batch normalization layer.

[0153] As discussed above regarding configurable processors, multiple spatially isolated convolutional layers are executed as the first set of convolutional layers in the neural network. In this example, the first set of convolutional layers applies 2D convolutions spatially.

[0154] like Figure 9 As shown, for each number in the stack L / 2 (L is the reference value), Figure 7The described spatially isolated neural network layers perform a first spatial convolution 921, followed by a second spatial convolution 922, then a third spatial convolution 923, and so on. As indicated at 923A, the number of spatial layers can be any actual number, which, depending on the context, can range from a few to more than 20 in different implementations.

[0155] For SP_CONV_0, the kernel weights are stored, for example, in a (1, 6, 6, 3, L) structure because there are 3 input channels for this layer. In this example, the "6" in the structure is attributed to storing the coefficients in the Winograd domain of the transform (the kernel size is 3×3 in the spatial domain but extended in the transform domain).

[0156] For this example, for other SP_CONV layers, the kernel weights are stored in a (1, 6, 6L) structure because for each of these layers there are K (=L) inputs and outputs.

[0157] The output of the stacked spatial layers is provided to the temporal layers, including convolutional layers 924 and 925 executed on the FPGA. Layers 924 and 925 can be convolutional layers that apply 1D convolutions between cycles. As indicated at 924A, the number of temporal layers can be any actual number, which, depending on the context, can range from a few to more than 20 in different implementations.

[0158] The first time layer TEMP_CONV_0 layer 824 reduces the number of loop channels from 5 to 3, such as Figure 7 As shown. The second time layer (layer 925) reduces the number of loop channels from 3 to 1, as... Figure 7 As shown, the number of feature maps is reduced to four outputs for each pixel, thus representing the confidence level in each base detection.

[0159] The output of the time layer is accumulated in the output patch and delivered to the host CPU to apply, for example, the softmax function 930 or other functions to normalize the base detection probability.

[0160] Figure 10 An alternative implementation of a 10-input, six-output neural network that can be performed for base detection operations is shown. In this example, block data from spatially aligned input patches of cycles 0 through 9 are applied to an isolated stack of spatial layers, such as stack 1001 of cycle 9. The output of the isolated stack is applied to an inverse hierarchical arrangement of a temporal stack 1020 with outputs 1035(2) through 1035(7), thereby providing base detection classification data for subjects in cycles 2 through 7.

[0161] Figure 11A neural network-based base detector (e.g., Figure 7 This is a specific implementation of a specialized architecture, where a neural network-based base detector is used to isolate the processing of data from different sequencing cycles. The motivation for using a specialized architecture is described first.

[0162] A neural network-based base detector processes data from the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. Data from the additional sequencing cycles provides sequence-specific context. The neural network-based base detector learns this sequence-specific context during training and performs base detection within that context. Furthermore, data from the previous and subsequent sequencing cycles provide second-order contributions to the pre-phase and phasing signals of the current sequencing cycle.

[0163] Images captured at different sequencing cycles and in different image channels are misaligned relative to each other and have residual registration errors. To address this misalignment, a specialized architecture includes spatial convolutional layers that do not mix information between sequencing cycles and only mix information within a sequencing cycle.

[0164] Spatial convolutional layers use so-called "isolated convolutions," which achieve isolation by processing data from each of multiple sequencing cycles independently via "dedicated, non-shared" convolutional sequences. This isolated convolution convolves only the data and resulting feature maps from a given sequencing cycle (i.e., within the cycle), without convolving the data and resulting feature maps from any other sequencing cycles.

[0165] For example, consider input data including (i) current data from the current (time t) sequencing cycle for which base detection is to be performed, (ii) previous data from a previous (time t-1) sequencing cycle, and (iii) subsequent data from a previous (time t+1) sequencing cycle. The specialized architecture then initiates three separate data processing pipelines (or convolutional pipelines): a current data processing pipeline, a previous data processing pipeline, and a subsequent data processing pipeline. The current data processing pipeline receives the current data from the current (time t) sequencing cycle as input and processes it independently through multiple spatial convolutional layers to produce a so-called "current spatial convolutional representation" as the output of the final spatial convolutional layer. The previous data processing pipeline receives the previous data from a previous (time t-1) sequencing cycle as input and processes it independently through multiple spatial convolutional layers to produce a so-called "previous spatial convolutional representation" as the output of the final spatial convolutional layer. The subsequent data processing pipeline receives subsequent data from the next (time t+1) sequencing cycle as input and processes this subsequent data independently through multiple spatial convolutional layers to produce a so-called "subsequent spatial convolutional representation" as the output of the final spatial convolutional layer.

[0166] In some implementations, the current pipeline, one or more previous pipelines, and one or more subsequent processing pipelines are executed in parallel.

[0167] In some implementations, spatial convolutional layers are part of a spatial convolutional network (or subnetwork) within a specialized architecture.

[0168] The neural network-based base detector also includes a temporal convolutional layer that mixes information between sequencing cycles (i.e., between cycles). The temporal convolutional layer receives its input from the spatial convolutional network and operates on the spatial convolutional representation produced by the final spatial convolutional layer of the corresponding data processing pipeline.

[0169] The inter-cycle operational freedom of temporal convolutional layers stems from the fact that misaligned properties are cleared from spatial convolutional representations by stacking or cascading isolated convolutions performed by a sequence of spatial convolutional layers. These misaligned properties exist in the image data fed into the spatial convolutional network as input.

[0170] Temporal convolutional layers use so-called "combined convolutions," which convolve input channels in subsequent inputs group by group on a sliding window basis. In one specific implementation, these subsequent inputs are the subsequent outputs produced by previous spatial or temporal convolutional layers.

[0171] In some implementations, temporal convolutional layers are part of a temporal convolutional network (or subnetwork) within a specialized architecture. The temporal convolutional network receives its input from a spatial convolutional network. In one implementation, the first temporal convolutional layer of the temporal convolutional network combines spatial convolutional representations between sequencing cycles, group by group. In another implementation, subsequent temporal convolutional layers of the temporal convolutional network combine subsequent outputs from previous temporal convolutional layers.

[0172] The output of the final temporal convolutional layer is fed into the output layer that produces the final output. The output is used to detect bases in one or more clusters at one or more sequencing cycles.

[0173] During forward propagation, the specialized architecture processes information from multiple inputs in two stages. In the first stage, isolating convolutions are used to prevent information from mixing between inputs. In the second stage, combining convolutions are used to mix the information between inputs. The result from the second stage is then used to perform a single inference on those multiple inputs.

[0174] This differs from batch processing techniques where convolutional layers process multiple inputs in a batch simultaneously and make a corresponding inference for each input in that batch. In contrast, the specialized architecture maps the multiple inputs to a single inference. This single inference can include more than one prediction, such as a classification score for each of the four bases (A, C, T, and G).

[0175] In one implementation, these inputs are temporally ordered, such that each input is generated at different time steps and has multiple input channels. For example, these multiple inputs could include three inputs: the current input generated by the current sequencing cycle at time step (t), the previous input generated by a previous sequencing cycle at time step (t-1), and the subsequent input generated by a subsequent sequencing cycle at time step (t+1). In another implementation, each input is derived from the current output, previous output, and subsequent output generated by one or more previous convolutional layers, and includes k feature maps.

[0176] In one implementation, each input may include five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another implementation, each input may include a k-feature map generated by a previous convolutional layer, and each feature map is considered an input channel. In yet another example, each input may have only one channel, two channels, or another different number of channels. The combined material of U.S. Patent Application Publication No. 2013 / 0079232 discusses base detection using various numbers of channels, such as one, two, or four channels.

[0177] Figure 12 An embodiment of isolation layers is illustrated, each of which may include convolutions. Isolation convolution processes multiple inputs by applying convolutional filters synchronously to each input once. Using isolation convolution, the convolutional filters combine input channels from the same input and do not combine input channels from different inputs. In one embodiment, the same convolutional filters are applied synchronously to each input. In another embodiment, different convolutional filters are applied synchronously to each input. In some embodiments, each spatial convolutional layer includes a set of k convolutional filters, where each convolutional filter is applied synchronously to each input.

[0178] Figure 13A A specific implementation of a combination layer is shown, where each combination layer may include convolutions. Figure 13B Another specific implementation of a combining layer is shown, where each combining layer may include convolutions. Combining convolutions mix information between different inputs by grouping corresponding input channels of different inputs and applying convolutional filters to each group. The grouping of these corresponding input channels and the application of convolutional filters occur on a sliding window basis. In this context, the window spans two or more subsequent input channels, representing, for example, the output of two subsequent sequencing cycles. Because this window is a sliding window, most input channels are used within two or more windows.

[0179] In some implementations, the different inputs originate from the output sequence produced by a previous spatial convolutional layer or a previous temporal convolutional layer. Within this output sequence, these different inputs are arranged as subsequent outputs and thus treated as subsequent inputs by a subsequent temporal convolutional layer. Then, in this subsequent temporal convolutional layer, these combined convolutions apply convolutional filters to the corresponding input channel groups in these subsequent inputs.

[0180] In one implementation, these subsequent inputs are temporally ordered such that the current input is generated by the current sequencing cycle at time step (t), the previous input is generated by the preceding sequencing cycle at time step (t-1), and the subsequent input is generated by the following sequencing cycle at time step (t+1). In another implementation, each subsequent input is derived from the current output, the previous output, and the subsequent output generated by one or more previous convolutional layers, and includes k feature maps.

[0181] In one implementation, each input may include the following five input channels: a red image channel (red), a red distance channel (yellow), a green image channel (green), a green distance channel (purple), and a scaling channel (blue). In another implementation, each input may include a k-feature map generated by the previous convolutional layer, and each feature map is considered as an input channel.

[0182] The depth B of a convolutional filter depends on the number of subsequent inputs, whose corresponding input channels are convolved by the filter group by group on a sliding window basis. In other words, the depth B is equal to the number of subsequent inputs and the group size in each sliding window.

[0183] exist Figure 13A In this case, the corresponding input channels from two subsequent inputs are combined in each sliding window, and therefore B = 2. Figure 13B In this case, the corresponding input channels from the three subsequent inputs are combined in each sliding window, and therefore B = 3.

[0184] In one implementation, the sliding windows share the same convolutional filters. In another implementation, different convolutional filters are used for each sliding window. In some implementations, each temporal convolutional layer comprises a set of k convolutional filters, where each convolutional filter is applied to subsequent inputs based on the sliding window.

[0185] Figures 4 to 10Further details and variations thereof can be found in co-pending U.S. nonprovisional patent application No. 17 / 176,147 (Attorney’s File No. ILLM 1020-2 / IP-1866-US), filed February 15, 2021, entitled “HARDWARE EXECUTION AND ACCELERATION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER”, which is incorporated herein by reference as if fully set forth herein.

[0186] De novo training of base detectors

[0187] A base detection system is trained to predict the base detection of an unknown analyte containing a base sequence. For example, the base detection system has a base detector that includes a neural network, which predicts the base detection of bases for an unknown analyte.

[0188] Training a neural network for a base detection system is challenging, especially when labeled training data for training the system is unavailable. In some examples, real-time analysis (RTA) systems can be used to generate labeled training data that can be used to train the base detection system. An example of an RTA system is discussed in U.S. Patent No. 10304189B2, entitled “Dataprocessing system and methods,” published May 28, 2019, which is incorporated herein by reference as if fully set forth herein. However, generating initial labeled training data for training a neural network for a base detection system can be challenging if the system lacks RTA or cannot fully utilize its capabilities.

[0189] This disclosure discusses a self-learning base detector that generates initial labeled training data, trains itself using the labeled training data, uses the at least partially trained base detector to generate additional labeled training data, trains itself again using the additional labeled training data, generates even more labeled training data, and iteratively repeats this process to fully train the base detector. This iterative training and labeled training data generation process includes different phases, such as a mono-oligonucleotide phase, a poly-oligonucleotide phase (such as a di-oligonucleotide phase, a tri-oligonucleotide phase, and so on), followed by a simple organism phase, a complex organism phase, another complex organism phase, and so on. Therefore, the complexity and / or length of the analytes used for training and generating labeled training data progressively and monotonically increases with iteration and the complexity of the underlying neural network configuration of the base detector, as will be discussed in further detail herein. Because the base detector is progressively self-trained, this system avoids the use of RTA to generate labeled training data. Therefore, although the base detection system discussed in this paper may include RTA, the iterative training process discussed in this paper can be used as a supplement to or alternative to RTA for training base detectors.

[0190] Figure 14A A base detection system 1400 is shown, which operates in a single oligonucleotide training phase to train a base detector 1414, including a neural network (NN) configuration 1415, using a known synthetic sequence 1406.

[0191] exist Figure 14A In the example, the base detection system 1400 includes a sequencing machine 1404, such as Figure 4 The sequencing machine 400. In this embodiment, the sequencing machine 1404 includes a biosensor ( Figure 14A (Not shown in the image), the biosensor includes components similar to... Figure 1 The biosensor 100 has a flow cell 102 and a flow cell 1405.

[0192] As relative to Figure 2 , Figure 3 and Figure 6 The circulation pool 1405 discussed here comprises multiple clusters 1407a, ..., 1407G. Specifically, in the example, circulation pool 1405 comprises multiple block channels, each block comprising corresponding multiple clusters, as relative to... Figure 2 The subject of discussion. Figure 14A In the diagram, flow cell 1405 is shown to include several such example clusters 1407a, ..., 1407G. During the base detection process, the base detection (A, C, G, T) for each cluster in a specific cycle is predicted.

[0193] A typical circulation pool 1405 may include multiple clusters 1407, such as thousands or even millions of clusters. This is merely an example and is not intended to limit the scope of this disclosure. In order to explain some of the principles of this disclosure, it is assumed that there are 10,000 (or 10k) clusters 1407 in circulation pool 1405 (i.e., G = 10,000), although a real circulation pool may have a higher number of such clusters.

[0194] In this example, the known synthetic sequence 1406 is used as an analyte for base detection during the single oligonucleotide training phase. In this example, the known synthetic sequence 1406 comprises a synthetically generated oligomer. Oligonucleotides are short DNA or RNA molecules, referred to as oligomers or simply oligonucleotides, which have wide applications in gene detection, research, and forensics. Typically prepared in the laboratory via solid-phase chemical synthesis, these small amounts of nucleic acids can be manufactured as single-stranded molecules with any user-specified sequence, and are therefore essential for artificial gene synthesis, polymerase chain reaction (PCR), DNA sequencing, molecular cloning, and as molecular probes. The length of an oligonucleotide is typically expressed as a “mer”. For example, an oligonucleotide of six nucleotides (nt) is a hexamer, while one of 25 nucleotides is typically called a “25-mer.” In this example, the size of the oligomer or oligonucleotide containing the known synthetic sequence 1406 can have any suitable number of bases, such as 8, 10, 12, or higher, and is implementation-specific. This is only an example. Figure 14A An oligonucleotide containing the known synthetic sequence 1406 of 8 bases is shown.

[0195] Figure 14A The oligonucleotides mentioned are labeled as oligonucleotide #1 (or oligonucleotide number 1). Because in Figure 14A Only one unique oligonucleotide is used, so the same oligonucleotide #1 is filled in a single cluster 1407. Therefore, all 10k clusters 1407 are filled with the same oligonucleotide sequence. That is, copies of the same oligonucleotide are filled in all clusters 1407.

[0196] Sequencing machine 1404 generates sequence signals 1412a, ..., 1412G for corresponding clusters among the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, sequencing machine 1404 generates a corresponding sequence signal 1412a, which indicates the base sequence filled within cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, sequencing machine 1404 generates a corresponding sequence signal 1412b, which indicates the base sequence filled within cluster 1407b for a series of sequencing cycles, and so on. Base detector 1414 receives sequence signal 1412 and is designed to detect (e.g., predict) the corresponding bases. In the example, base detector 1414, including NN configuration 1415 (and various other NN configurations discussed later herein), may be stored in memories 404, 403, and / or 406, and in a host CPU (such as...) Figure 4 On the CPU 402) and / or on a configurable processor native to the sequencing machine 400 (such as Figure 4 The base detector 1414 may be stored remotely from the sequencing machine 400 (e.g., in the cloud) and may be executed by a remote processor (e.g., in the cloud). For example, in a remote version of the base detector 1414, the base detector 1414 receives (e.g., via a network such as the Internet) a sequence signal 1412, performs a base detection operation, and transmits the base detection result (e.g., via a network such as the Internet) to the sequencing machine 400.

[0197] In the example, sequence signal 1412 includes an image captured by a sensor (e.g., a photodetector, a photodiode), as discussed previously herein. Therefore, at least some of the examples and embodiments discussed herein involve iteratively training a base detector (such as base detector 1414) that processes sequence signals including images. However, the principles of this disclosure are not limited to training any particular type of base detector that receives a particular type of sequence signal. For example, the iterative training discussed herein is independent of the type of base detector to be trained or the type of sequence signal used. For example, the iterative training discussed herein can be used to train any other suitable type of base detector, such as a base detector configured to detect bases based on a sequence signal that does not include an image. For example, the sequence signal may include an electrical signal (e.g., a voltage signal, a current signal), a pH level, etc., and the iterative training method discussed herein can be applied to train a base detector that receives any such type of sequence signal.

[0198] Neural network configuration 1415 is a convolutional neural network (examples of which are in...). Figure 7 , Figure 9 , Figure 10 , Figure 11 , Figure 12 As shown in the diagram, this convolutional neural network uses a relatively small number of layers and a relatively small number of parameters (e.g., compared to some other neural network configurations discussed later in this paper, such as...). Figure 16A The neural network configuration (1615) will be discussed in further detail in this paper.

[0199] The initial, untrained base detector 1414, including the neural network configuration 1415, predicts the base detection sequences 1418a, ..., 1418G for the corresponding clusters in the plurality of clusters 1407a, ..., 1407G based on the corresponding sequence signals 1412a, ..., 1412G. For example, for cluster 1407a, the base detector 1414 predicts the corresponding base detection sequence 1418a based on the corresponding sequence signal 1412a, including the base detection for cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the base detector 1414 predicts the corresponding base detection sequence 1418b based on the corresponding sequence signal 1412b, including the base detection for cluster 1407b for a series of sequencing cycles, and so on. Therefore, the G base detection sequences 1418a, ..., 1418G are predicted by the base detector 1414.

[0200] Assume oligonucleotide #1 has 8 bases, typically labeled GA1, ..., GA8. As an example only and not to limit the scope of this disclosure, assume the 8 bases of oligonucleotide # are A, C, T, T, G, C, A, C. Initially, the base detector 1414 is untrained and therefore likely to contain errors in base detection. For example, the predicted base detection sequence 1418a (typically labeled Sa1, ..., Sa8) is C, A, T, C, G, C, A, G, as... Figure 14A As shown. Therefore, comparing the baseline true base sequence 1406 (i.e., A, C, T, T, G, C, A, C) of oligonucleotide #1 with the predicted base sequence 1418a (i.e., C, A, T, C, G, C, A, G), errors exist in the detection of bases numbered 1, 2, 4, and 8. Therefore, in Figure 14A In operation 1413a, the baseline true base sequence 1406 of oligonucleotide #1 is compared with the predicted base sequence 1418a, and the error between these two base sequences is used in the backward pathway of the neural network configuration 1415 of the base detector 1414 to train the neural network configuration 1415, such as to update the gradients and weights of the neural network configuration 1415. Figure 14A This is symbolically labeled as gradient update 1417.

[0201] Figure 14A1The comparison operation between the predicted base sequence 1418a and the baseline true base sequence 1406 of oligonucleotide #1 is shown in further detail. For example, refer to... Figure 14A and Figure 14A1 The predicted base sequence 1418a is C, A, T, C, G, C, A, G, and the baseline true base sequence 1406 of oligonucleotide #1 is A, C, T, T, G, C, A, C. Therefore, comparing the baseline true base sequence 1406 of oligonucleotide #1 (i.e., A, C, T, T, G, C, A, C) with the predicted base sequence 1418a (i.e., C, A, T, C, G, C, A, G), errors exist in the detection of bases numbered 1, 2, 4, and 8. For example, in Figure 14A1 In the code, the error in detecting base number 1 is given by the following: "C should be A", that is, base detection C should be base detection A. Similarly, the error in detecting base number 2 is given by the following: "A should be C", that is, base detection A should be base detection B, and so on. There are no errors in detecting base numbers 3, 5, 6, and 7 (in... Figure 14A1 The text indicates a "match (no error)" message. Therefore, in... Figure 14A1 During the comparison, each base detection of the predicted base detection sequence 1418a is compared with the corresponding base detection of the corresponding benchmark true sequence (e.g., the base sequence 1406 of oligonucleotide #1) to generate the corresponding comparison result, such as... Figure 14A1 As shown in the image.

[0202] Refer again Figure 14A The base detection system 1400 also includes mapping logic 1416, the function of which will be discussed later herein. In the example, mapping logic 1416 may be stored in memories 404, 403, and / or 406, and mapping logic 1416 may be available on a host CPU (such as...) Figure 4 On the CPU 402) and / or on a configurable processor native to the sequencing machine 400 (such as Figure 4 The mapping logic 1416 is executed on a configurable processor 450. In another example, the mapping logic 1416 may be stored remotely from the sequencing machine 400 (e.g., in the cloud) and may be executed by a remote processor (e.g., in the cloud). For example, in a remote version of the mapping logic 1416, the mapping logic receives data to be mapped from the sequencing machine 400 (e.g., via a network such as the Internet), performs the mapping operation, and transmits the mapping result (e.g., via a network such as the Internet) to the sequencing machine 400. The mapping operation is discussed in further detail later in this document.

[0203] Figure 14AVarious other figures, examples, and embodiments of this disclosure relate to base detectors that predict base detection sequences. Various examples of such prediction of base detection sequences have been discussed herein. Further details of base detection prediction can be found in co-pending U.S. Provisional Patent Application No. 63 / 217,644 (Attorney’s File No. ILLM 1046-1 / IP-2135-PRV), filed July 1, 2021, entitled “IMPROVED ARTIFICIAL INTELLIGENCE-BASED BASECALLING OF INDEX SEQUENCES,” which is incorporated herein by reference as if fully set forth herein.

[0204] Figure 14B It shows Figure 14A Further details of the base detection system 1400, which operates in a single oligonucleotide training phase to train a base detector 1414, including a neural network configuration 1415, using a known synthetic sequence 1406. For example, Figure 14B The diagram illustrates training a base detector 1414 using predicted base detection sequences 1418a, ..., 1418G. For example, each base detection sequence in the predicted base detection sequences 1418a, ..., 1418G is compared to the baseline true base sequence 1406 of oligonucleotide #1 (see comparison operations 1413a, ..., 1413G), and the resulting error is used for gradient updates and subsequent updates to the parameters (such as weights and biases) of the neural network configuration 1415 by the backpropagation portion of the neural network configuration 1415. Figure 14A (This is symbolically labeled as gradient update 1417).

[0205] Therefore, the neural network configuration 1415 is trained using the base detection sequence 1418 predicted by the neural network configuration 1415 and the benchmark true base sequence 1406 of oligonucleotide #1. Because relative to... Figure 14A and Figure 14B The training discussed uses a single oligonucleotide, so this training phase is also known as the "single oligonucleotide training phase," and Figure 14A and Figure 14B It has been marked accordingly.

[0206] In the example, Figure 14A and Figure 14B The process can be repeated iteratively. For example, in Figure 14A In the first iteration, the NN configuration 1415 is trained at least partially. During the second iteration, the at least partially trained NN configuration 1415 is used again to regenerate the predicted base detection sequence from the sequence signal 1412 (e.g., as relative to...). Figure 14A(As discussed), and again compare the resulting predicted base detection sequence with the baseline true value 1406 (i.e., oligonucleotide #1) to generate error signals, which are used to further train the NN configuration 1415. This process can be repeated iteratively multiple times until the NN configuration 1415 is sufficiently trained. In the example, the process can be repeated iteratively a specific number of times. In another example, the process can be repeated iteratively until there is saturation with a certain number of errors (e.g., the errors do not decrease significantly in successive iterations).

[0207] Figure 15A It shows Figure 14A The base detection system 1400 operates in the training data generation phase of the dioliponucleotide training phase to generate labeled training data using two known synthetic sequences 1501A and 1501B.

[0208] Figure 15A The base detection system 1400 and Figure 14A The base detection system is the same, and in both figures, base detection system 1400 uses neural network configuration 1415. Furthermore, two distinct oligonucleotide sequences 1501A and 1501B are loaded into different clusters in flow-through pool 1405. By way of example only and not to limit the scope of this disclosure, it is assumed that in 10,000 clusters 1407, approximately 5,200 clusters are filled with oligonucleotide sequence 1501A and the remaining 4,800 clusters are filled with oligonucleotide sequence 1501B (although in another example, the two oligonucleotides may substantially equally divide the 10,000 clusters).

[0209] Sequencing machine 1404 generates sequence signals 1512a, ..., 1512G for the corresponding clusters among the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, sequencing machine 1404 generates a corresponding sequence signal 1512a, which indicates the bases of cluster 1407a used in a series of sequencing cycles. Similarly, for cluster 1407b, sequencing machine 1404 generates a corresponding sequence signal 1512b, which indicates the bases of cluster 1407b used in a series of sequencing cycles, and so on.

[0210] This includes at least partially trained neural network configuration 1415 (e.g., which is trained by iteratively repeating...). Figure 14A and Figure 14BThe base detector 1414 (trained by the operation of the sequence signal 1512a, ..., 1512G) predicts the base detection sequences 1518a, ..., 1518G for the corresponding clusters in the plurality of clusters 1407a, ..., 1407G, respectively, based on the corresponding sequence signals 1512a, ..., 1512G. For example, for cluster 1407a, the base detector 1414 predicts the corresponding base detection sequence 1518a based on the corresponding sequence signal 1512a, including the base detection for cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the base detector 1414 predicts the corresponding base detection sequence 1518b based on the corresponding sequence signal 1512b, including the base detection for cluster 1407b for a series of sequencing cycles, and so on. Therefore, the G base detection sequences 1518a, ..., 1518G are predicted by the base detector 1414. Note that... Figure 15A The neural network configuration 1415 is relative to Figure 14A and Figure 14B The single oligonucleotide training phase discussed was trained relatively early during the iterations. Therefore, the predicted base detection sequences 1518a, ..., 1518G will be slightly accurate, but not very highly accurate (because base detector 1414 was not fully trained).

[0211] In the implementation scheme, oligonucleotide sequences 1501A and 1501B were selected to have sufficient edit distance between the bases of the two oligonucleotides. Figure 15B and Figure 15C It shows Figure 15A Two corresponding example choices of oligonucleotide sequences 1501A and 1501B. For example, in Figure 15B In this study, oligonucleotide 1501A was selected to have the bases A, C, T, T, G, C, A, C, while oligonucleotide 1501B was selected to have the bases C, C, T, A, G, C, A, C. Therefore, the first and fourth bases of the two oligonucleotides 1510A and 1510B are different, resulting in an edit distance of two between the two oligonucleotides 1510A and 1510B.

[0212] On the contrary, Figure 15B In this study, oligonucleotide 1501A was selected to have the bases A, C, T, T, G, C, A, C, while oligonucleotide 1501B was selected to have the bases C, A, T, G, A, T, A, G. Therefore, in Figure 15B In the example, the first, second, fourth, fifth, sixth, and eighth bases of the two oligonucleotides 1510A and 1510B are different, resulting in an edit distance of six between the two oligonucleotides 1510A and 1510B.

[0213] In this example, the two oligonucleotides 1501A and 1501B are chosen such that the two oligonucleotides are separated by at least a threshold edit distance. For example only, the threshold edit distance could be 4, 5, 6, 7, or even 8 bases. Therefore, the two oligonucleotides 1501A and 1501B are chosen such that the two oligonucleotides are sufficiently different from each other.

[0214] Refer again Figure 15A The base detector 1414 does not know which oligonucleotide sequence is filled in which cluster. Therefore, the base detector 1414 does not know the mapping between the known oligonucleotide sequences 1501A, 1501B and the various clusters. In the example, the mapping logic 1416 receives the predicted base detection sequence 1518 and maps each predicted base detection sequence 1518 to oligonucleotide 1501A or oligonucleotide 1501B, or declares the uncertainty of mapping the predicted base detection sequence to either of the two oligonucleotides. Figure 15D Example mapping operations are shown for (i) mapping a predicted base detection sequence to either oligonucleotide 1501A or oligonucleotide 1501B, or (ii) declaring an uncertainty in mapping a predicted base detection sequence to either of the two oligonucleotides.

[0215] In the example, the higher the edit distance between the two oligonucleotides, the easier (or more accurate) it is to map a single prediction to either of the two oligonucleotides. For example, see reference... Figure 15B Because the edit distance between the two oligonucleotides 1501A and 1501B is only two, the two oligonucleotides are almost similar, and it may be relatively difficult to map base detection predictions to either of the two oligonucleotides. However, due to Figure 15C The edit distance between the two oligonucleotides 1501A and 1501B is six, so these two oligonucleotides are very dissimilar, and predictions can be mapped to either of the two oligonucleotides relatively easily. Therefore, an edit distance of two... Figure 15B Marked as "not very suitable for training", with an edit distance of six. Figure 15C It is marked as "more suitable for training". Therefore, in the example, according to Figure 15C (and not based on) Figure 15B Oligonucleotides 1501A and 1501B were generated and used for training, as will be discussed in further detail in this paper.

[0216] Refer again Figure 15D Example predicted base detection sequences 1518a, 1518b, and 1518G are shown. Example bases for two oligonucleotides 1501A and 1501B are also shown (the example bases for the two oligonucleotides correspond to...). Figure 15C (The bases shown).

[0217] Because neural network configuration 1415 is only slightly trained, but not fully trained, it may be able to make base detection predictions, but such base detection predictions will be prone to errors.

[0218] The predicted base detection sequence 1518a includes C, A, G, G, C, T, A, C. This is compared to the base detection sequence A, C, T, T, G, C, A, C of oligonucleotide 1501A, and also to the base detection sequence C, A, T, G, A, T, A, G of oligonucleotide 1501B. The predicted base detection sequence 1518a has the seventh and eighth bases matching the corresponding seventh and eighth bases of oligonucleotide 1501A, and the first, second, fourth, sixth, and seventh bases matching the corresponding bases of oligonucleotide 1501B. Therefore, as Figure 15D As shown, the predicted base detection sequence 1518a has a 2-base similarity to oligonucleotide 1501A, and the predicted base detection sequence 1518a has a 5-base similarity to oligonucleotide 1501B.

[0219] If the predicted base detection sequence 1518a is indeed for oligonucleotide 1501B (e.g., because the predicted base detection sequence 1518a has a 5-base similarity to oligonucleotide 1501B), then this means that the neural network configuration 1415 is able to correctly predict five bases of the 8-base sequence (i.e., it is able to correctly predict the first, second, fourth, sixth, and seventh bases that match the corresponding bases of oligonucleotide 1501B). However, when the neural network configuration 1415 is not fully trained, it makes errors in predicting the remaining three bases (i.e., the third, fifth, and eighth bases).

[0220] Mapping logic 1416 can use appropriate logic to map the predicted base detection sequence to the corresponding oligonucleotide. For example, suppose the predicted base detection sequence has a similarity of the number of SAs to oligonucleotide 1501A and a similarity of the number of SBs to oligonucleotide 1501B. In this example, if SA > ST and SB < ST, then mapping logic 1416 maps the predicted base detection sequence to oligonucleotide 1501A, where ST is the threshold number. That is, if the similarity level with oligonucleotide 1501A is higher than the threshold, and if the similarity level with oligonucleotide 1501B is lower than the threshold, then mapping logic 1416 maps the predicted base detection sequence to oligonucleotide 1501A.

[0221] Similarly, in another example, if SB > ST and SA < ST, then mapping logic 1416 maps the predicted base detection sequence to oligonucleotide 1501B.

[0222] In yet another example, if both SA and SB are less than the threshold ST, or if both SA and SB are greater than the threshold ST, then mapping logic 1416 declares the predicted base detection sequence to be indeterminate.

[0223] The above discussion can be written in equation form:

[0224] For the predicted base detection sequence:

[0225] If SA > ST and SB < ST, then it maps to oligonucleotide 1501A; Equation 1

[0226] If SB > ST and SA < ST, then it maps to oligonucleotide 1501B; Equation 2

[0227] If both SA and SB < ST, then declare an indeterminate mapping; or Equation 3.

[0228] If both SA and SB > ST, then an indeterminate mapping is declared. Equation 4

[0229] The threshold ST depends on the number of bases in the oligonucleotide (eight in the example use case shown in the figure), the desired accuracy, and / or implementation specificity. This is only an example. Figure 15D In the example usage shown, it is assumed that the threshold ST is 4. Note that the threshold ST of 4 is merely an example, and the choice of threshold ST can be implementation-specific. As an example only, during the initial iterations of training, the threshold ST may have a relatively low value (e.g., 4); and during later iterations of training, the threshold ST may have a relatively high value (e.g., 6 or 7) (training iterations have been discussed later in this document). Therefore, the threshold ST can be gradually increased as the NN configuration becomes better trained during later training iterations. However, in another example, the threshold ST may have the same value throughout all iterations of training. Although in Figure 15D In this example, the threshold ST is chosen to be 4, but in other example implementations, the threshold ST may be, for example, 5, 6, or 7. In this example, the threshold ST may also be expressed as a percentage. For example, when the threshold ST is 4 and the total number of bases is 8, the threshold ST may be expressed as (4 / 8) × 100, i.e., 50%. The threshold ST can be a user-selectable parameter and in this example, it may be selected between 50% and 95%.

[0230] Now refer to it again Figure 15DAs discussed above, the predicted base detection sequence 1518a shares 2 bases of similarity with oligonucleotide 1501A, and the predicted base detection sequence 1518a shares 5 bases of similarity with oligonucleotide 1501B. Therefore, SA = 2 and SB = 5. Assuming a threshold ST of 4, according to Equation 2, the predicted base detection sequence 1518a is mapped to oligonucleotide 1501B.

[0231] Now, referring to the predicted base detection sequence 1518b, the predicted base detection sequence 1518b has a 2-base similarity to oligonucleotide 1501A and a 3-base similarity to oligonucleotide 1501B. Therefore, SA = 2 and SB = 3. Assuming the threshold ST is 4, according to Equation 3, it is declared that the predicted base detection sequence 1518b is indeterminate for any of the oligonucleotide sequences mapped to.

[0232] Now, referring to the predicted base detection sequence 1518G, the predicted base detection sequence 1518G has 6 bases of similarity to oligonucleotide 1501A, and 3 bases of similarity to oligonucleotide 1501B. Therefore, SA = 6 and SB = 3. Assuming the threshold ST is 4, according to Equation 2, the predicted base detection sequence 1518G is mapped to oligonucleotide 1501A.

[0233] Figure 15E It shows from Figure 15D The mapped labeled training data 1550 is used by another neural network configuration 1615 (e.g., in...). Figure 16A As shown, one of the neural network configurations 1615 differs from... Figure 14A , Figure 14B , Figure 15A The neural network configuration is 1415, and is more complex than that.

[0234] like Figure 15E As shown, some predicted base detection sequences 1518 and their corresponding sequence signals are mapped to the base sequence of oligonucleotide 1501A (i.e., the true reference 1506a), some other predicted base detection sequences 1518 and their corresponding sequence signals are mapped to the base sequence of oligonucleotide 1501B (i.e., the true reference 1506b), and the mapping of the remaining predicted base detection sequences 1518 and their corresponding sequence signals is uncertain.

[0235] For example, the predicted base detection sequences 1518c, 1518d, 1518G and the corresponding sequence signals 1512c, 1512d, 1512G are mapped to the base sequence of oligonucleotide 1501A (i.e., the baseline truth value 1506a); the predicted base detection sequences 1518a, 1518f and the corresponding sequence signals 1512a, 1512f are mapped to the base sequence of oligonucleotide 1501B (i.e., the baseline truth value 1506b); and the mapping of the remaining predicted base detection sequences 1518b, 1518e, 1518g and the corresponding sequence signals 1512b, 1512e, 1512g is uncertain.

[0236] As an example only, suppose that the 2,600 base detection sequences from training data 1550 are mapped to oligonucleotide 1501A, and the 3,000 base detection sequences from training data 1550 are mapped to oligonucleotide 1501B. Figure 15E As shown, the remaining 4,400 bases detected are indeterminate and do not map to either of the two oligonucleotides.

[0237] Notice, Figure 15A , Figure 15D and Figure 15E The “training data generation phase” is called the “double oligonucleotide training phase” because it uses sequences from two oligonucleotides and uses a neural network configuration 1415 to generate labeled training data 1550.

[0238] Figure 16A It shows Figure 14A The base detection system 1400 operates in the "training data consumption and training phase" of the "dual oligonucleotide training phase" to train another neural network configuration 1615 (which differs from the original configuration) using two known synthetic sequences 1501A and 1501B. Figure 14A The neural network configuration 1415, and the base detector 1414 (which is more complex than its neural network configuration 1415).

[0239] Figure 16A The base detection system 1400 and Figure 14A The base detection system is the same. However, it is different from... Figure 14A (Where neural network configuration 1415 is used in base detector 1414) is different. Figure 16A The base detector 1414 in the middle uses a different neural network configuration 1615. Figure 16A The neural network configuration 1615 is different Figure 14A Neural network configuration 1415. For example, neural network configuration 1615 is a convolutional neural network (examples of which are shown in...). Figure 7 , Figure 9 , Figure 10 , Figure 11 , Figure 12 As shown in the figure, it uses a larger number of layers and parameters (such as weights and biases) than neural network configuration 1415. In another example, neural network configuration 1615 is a convolutional neural network that uses a larger number of convolutional filters than neural network configuration 1415. In some examples, the configurations, topologies, and the number of layers and / or filters of the two neural network configurations 1415 and 1615 may differ.

[0240] exist Figure 16A In the "Training Data Consumption and Training Phase" of the "Dual Oligonucleotide Training Phase" shown, the base detector 1414 of the neural network configuration 1615 receives sequence signals 1512, which were previously in... Figure 15A The training data is generated during the "training data generation phase". That is, the base detector 1414, including the neural network configuration 1615, reuses the previously generated sequence signal 1512. Therefore, since the previously generated sequence signal 1512... Figure 16A The "Training Data Consumption and Training Phase" of the "Dual Oligonucleotide Training Phase" shown is reused, so sequencing machine 1404 and its components are not active, hence they are shown with dashed lines. Similarly, mapping logic 1416 does not play any role (because...). Figure 16A (No mapping is performed in the process), so mapping logic 1416 is also shown using dashed lines.

[0241] Therefore, in Figure 16A In this configuration, a base detector 1414, including a neural network configuration 1615, receives a previously generated sequence signal 1512 and predicts a base detection sequence 1618 from the sequence signal 1512. The predicted base detection sequence 1618 includes predicted base detection sequences 1618a, 1618b, ..., 1618G. For example, sequence signal 1512a is used to predict base detection sequence 1618a, sequence signal 1512b is used to predict base detection sequence 1618b, sequence signal 1512G is used to predict base detection sequence 1618G, and so on.

[0242] The neural network configuration 1615 has not been trained, so the predicted base detection sequences 1618a, 1618b, ..., 1618G will have many errors. Figure 15E The mapping training data 1550 is now used to train the neural network configuration 1615. For example, based on the training data 1550, the base detector 1414 knows:

[0243] (i) Sequence signals 1512c, 1512d, and 1512G are the base sequences used for oligonucleotide 1501A (i.e., the reference true value 1506a);

[0244] (ii) Sequence signals 1512a and 1512f are the base sequences used for oligonucleotide 1501B (i.e., the reference true value 1506b); and

[0245] (iii) The mapping of sequence signals 1512b, 1512e, and 1512g is uncertain.

[0246] Therefore, sequence signal 1512 and the predicted base detection sequence 1518 were selected into three categories: (i) the first category, which includes sequence signals 1512c, 1512d, and 1512G (and the corresponding predicted base detection sequences 1518c, 1518d, and 1518G) that can be mapped to oligonucleotide 1501A (i.e., the baseline truth 1506a); and (i) the second category, which includes sequences that can be mapped to oligonucleotide 1501B. (iii) The third category includes sequence signals 1512a, 1512f (and corresponding predicted base detection sequences 1518a, 1518f) of the base sequence (i.e., the baseline truth value 1506b); and (iii) the third category, which includes sequence signals 1512b, 1512e, 1512g (and corresponding predicted base sequences 1518b, 1518e, 1518g) of any base detection sequence that cannot be mapped to oligonucleotides 1501A or 1501B.

[0247] Therefore, based on (iii) above, the predicted base detection sequences 1618b, 1618e, and 1618g (e.g., corresponding to sequence signals 1512b, 1512e, and 1512g) are not used to train the neural network configuration 1615. Thus, the predicted base detection sequences 1618b, 1618e, and 1618g are discarded during training iterations and are not used for gradient updates (in... Figure 16A The predicted base detection sequences 1618b, 1618e, and 1618g, along with the gradient update box 1617, are symbolically indicated by an "X" or "cross".

[0248] Based on (i) above, base detector 1414 knows that the predicted base detection sequences 1618c, 1618d, 1618G (e.g., corresponding to sequence signals 1512c, 1512d, 1512G) may be for oligonucleotide 1501A. That is, the base sequence of oligonucleotide 1501A may be the benchmark truth for these predicted base detection sequences 1618c, 1618d, 1618G, even though the untrained neural network configuration 1615 may have incorrectly predicted at least some bases of these predicted base detection sequences. Therefore, the neural network configuration uses comparison function 1613 to compare each of the predicted base detection sequences 1618c, 1618d, and 1618G with the benchmark truth 1506a (which is the base sequence of oligonucleotide 1501A), and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615.

[0249] Similarly, based on (ii) above, the base detector knows that the predicted base detection sequences 1618a and 1618f (e.g., corresponding to sequence signals 1512a and 1512f, respectively) may be for oligonucleotide 1501B. That is, the base sequence of oligonucleotide 1501B may be the benchmark truth for these predicted base detection sequences 1618a and 1618f, even though the untrained neural network configuration 1615 may have incorrectly predicted at least some bases of these predicted base detection sequences. Therefore, the neural network configuration uses comparison function 1613 to compare each of the predicted base detection sequences 1618a and 1618f with the benchmark truth 1506b (which is the base sequence of oligonucleotide 1501B), and uses the generated errors for gradient update 1617 and training of the resulting neural network configuration 1615.

[0250] exist Figure 16A The training data consumption and the end of the training phase result in at least partial training of the NN configuration 1615.

[0251] Figure 16B It shows Figure 14A The base detection system 1400 operates in the second iteration of the training data generation phase during the dioliponucleotide training phase. For example, in Figure 16A In this example, training data 1550 is used to train neural network configuration 1615. Figure 16B In this process, a slightly or at least partially trained neural network configuration 1615 is used to generate further training data. For example, the at least partially trained neural network configuration 1615 uses a previously generated sequence signal 1512 to predict a base detection sequence 1628. Figure 16B The predicted base detection sequence 1628 may be more accurate than expected. Figure 16AThe predicted base detection sequence 1618 is relatively more accurate because Figure 16A The predicted base detection sequence 1618 was generated using an untrained neural network configuration 1615, while Figure 16B The predicted base detection sequence 1628 was generated using at least a portion of the neural network configuration 1615.

[0252] Furthermore, mapping logic 1416 maps each base detection sequence in the predicted base detection sequence 1628 to oligonucleotide 1501A or oligonucleotide 1501B, or declares that the mapping of the predicted base detection sequence 1628 is indeterminate (e.g., similar to relative to...). Figure 15D (Discussion).

[0253] Figure 16C It shows from Figure 16B The mapping generates labeled training data 1650, which will be used for further training.

[0254] like Figure 16C As shown, some predicted base detection sequences 1628 and corresponding sequence signals 1512 are mapped to the base sequence of oligonucleotide 1501A (i.e., the true reference 1506a), some other predicted base detection sequences 1628 and corresponding sequence signals 1512 are mapped to the base sequence of oligonucleotide 1501B (i.e., the true reference 1506b), and the mapping of the remaining predicted base detection sequences 1628 and corresponding sequence signals 1512 is uncertain.

[0255] For example, the predicted base detection sequence 1628 is selected into three categories: (i) the predicted base detection sequences 1628c, 1628d, and 1628G, and their corresponding sequence signals 1512c, 1512d, and 1512G are mapped to the base sequence of oligonucleotide 1501A (i.e., the baseline truth value 1506a); (ii) the predicted base detection sequences 1628a, 1628b, and 1628f, and their corresponding sequence signals 1512a, 1512b, and 1512f are mapped to the base sequence of oligonucleotide 1501B (i.e., the baseline truth value 1506b); and (iii) the mapping of the remaining predicted base detection sequences 1628e and 1628g, and their corresponding sequence signals 1512e and 1512g, is uncertain.

[0256] As an example only, suppose that the 3,300 base detection sequences from training data 1650 are mapped to oligonucleotide 1501A, and the 3,200 base detection sequences from training data 1650 are mapped to oligonucleotide 1501B. Figure 16C As shown, the remaining 3,500 bases detected are indeterminate and do not map to either of the two oligonucleotides.

[0257] Compare Figure 15E and Figure 16C The number of unmapped (or indeterminate) base detection sequences between the training data was observed to be in [a certain range]. Figure 15E The middle is 4,400, and in Figure 16C The middle is 3,500. This is because... Figure 16B The at least partially trained neural network configuration 1615 (which is used to generate the mapping of training data 1650) is comparable to Figure 15A The at least partially trained neural network configuration 1415 (which is used to generate the mappings for training data 1550) is relatively more accurate and / or more trained. Therefore, the number of uncertain sequences detected by base detection gradually decreases as base detection becomes relatively more accurate (e.g., less prone to errors), and thus the mapping is now relatively more correct.

[0258] Figure 16D It shows Figure 14A The base detection system 1400 operates in the second iteration of the "training data consumption and training phase" of the "dual oligonucleotide training phase" to train the system using two known synthetic sequences 1501A and 1501B. Figure 16A The neural network configuration 1615 has a base detector 1414.

[0259] Figure 16A and Figure 16D At least partially similar. For example, Figure 16A and Figure 16D Used respectively Figure 15E Training data 1550 and Figure 16C The training data 1650 is used to train the neural network configuration 1615. Note that in Figure 16A In the initial stage, neural network configuration 1615 was completely untrained; while Figure 16D In the initial stage, neural network configuration 1615 was at least partially trained.

[0260] exist Figure 16D In the process, the base detector 1414, including at least partially trained neural network configuration 1615, receives previously... Figure 15A Sequence signal 1512 is generated during the "training data generation phase," and base detection sequence 1638 is predicted based on sequence signal 1512. The predicted base detection sequence 1638 includes predicted base detection sequences 1638a, 1638b, ..., 1638G. For example, sequence signal 1512a is used to predict base detection sequence 1638a, sequence signal 1512b is used to predict base detection sequence 1638b, sequence signal 1512G is used to predict base detection sequence 1638G, and so on.

[0261] Neural network configuration 1615 was not fully trained, therefore the predicted base detection sequences 1638a, 1638b, ..., 1638G will include some errors, although... Figure 16D The error in the predicted base detection sequence 1638 may be less than Figure 16A The predicted base detection sequence 1618 and Figure 16B Errors were detected in the predicted base sequence 1628. Figure 16C The mapping training data 1650 is now used to further train the neural network configuration 1615. For example, based on the training data 1650, the base detector 1414 knows:

[0262] (i) Sequence signals 1512c, 1512d, and 1512G are the base sequences used for oligonucleotide 1501A (i.e., the reference true value 1506a);

[0263] (ii) Sequence signals 1512a, 1512b, and 1512f are the base sequences used for oligonucleotide 1501B (i.e., the reference true value 1506b); and

[0264] (iii) The mapping of sequence signals 1512e and 1512g is uncertain.

[0265] Therefore, based on (iii) above, Figure 16D The predicted base detection sequences 1638e and 1638g (e.g., corresponding to sequence signals 1512e and 1512g, respectively) are not used to train the neural network configuration 1615. Therefore, these predicted base detection sequences 1638e and 1638g are discarded from the training data and are not used for gradient updates (in... Figure 16D (This is symbolically indicated by an "X" or "cross" between the predicted base detection sequences 1618e, 1618g and gradient update box 1617).

[0266] Based on (i) above, base detector 1414 knows that the predicted base detection sequences 1638c, 1638d, and 1638G (e.g., corresponding to sequence signals 1512c, 1512d, and 1512G, respectively) may be for oligonucleotide 1501A. That is, the base sequence of oligonucleotide 1501A may be the benchmark truth for these predicted base detection sequences 1638c, 1638d, and 1638G, even though part of the neural network configuration 1615 may have incorrectly predicted at least some bases of these predicted base detection sequences. Therefore, the neural network configuration uses comparison function 1613 to compare each of the predicted base detection sequences 1638c, 1638d, and 1638G with the benchmark truth 1506a (which is the base sequence of oligonucleotide 1501A), and uses the generated errors for gradient update 1617 and the resulting training of neural network configuration 1615. For example, during the comparison, each base detection of the predicted base detection sequence 1638c is compared with the corresponding base detection of the corresponding benchmark true sequence to generate a corresponding comparison result, such as, relative to Figure 14A1 The subject of discussion.

[0267] Similarly, based on (ii) above, the base detector knows that the predicted base detection sequences 1638a, 1638b, and 1638f (e.g., corresponding to sequence signals 1512a, 1512b, and 1512f, respectively) may be for oligonucleotide 1501B. That is, the base sequence of oligonucleotide 1501A may be the benchmark truth for these predicted base detection sequences 1638a, 1638b, and 1638f, even though part of the neural network configuration 1615 may have incorrectly predicted at least some bases on these predicted base detection sequences. Therefore, the neural network configuration uses comparison function 1613 to compare each of the predicted base detection sequences 1638a, 1638b, and 1638f with the benchmark truth 1506b (which is the base sequence of oligonucleotide 1501B), and uses the generated errors for gradient update 1617 and training of the resulting neural network configuration 1615.

[0268] Figure 17A A flowchart depicting an example method 1700 for iteratively training a neural network configuration for base detection using single and double oligonucleotide sequences is shown. Method 1700 progressively trains an inherently progressive and monotonically complex NN configuration. Increasing the complexity of the NN configuration may include increasing the number of layers in the NN configuration, increasing the number of filters in the NN configuration, increasing the topological complexity in the NN configuration, etc. For example, method 1700 refers to the first NN configuration (which is previously described in this paper relative to...) Figure 14A Other NN configurations discussed in conjunction with graphs 1415), and a second NN configuration (which is the one previously discussed in this paper relative to...) Figure 16AOther NN configurations discussed in relation to graphs 1615), the Pth NN configuration (which has no relation to...) Figures 14A to 16D (This will be discussed in detail), and so on. In the example, the complexity of the Pth NN configuration is higher than that of the (P-1)th NN configuration, which is higher than that of the (P-2)th NN configuration, and so on, with the second NN configuration being more complex than the first NN configuration. Figure 17A The symbolic representation is shown within box 1710. Therefore, the complexity of the NN configuration increases monotonically (i.e., the NN configuration in later stages has at least similar complexity to or greater complexity than the NN configuration in earlier stages).

[0269] Note that in method 1700, operation 1704a is used to iteratively train the first NN configuration and generate labeled training data for the second NN configuration, operations 1704b1-1704bk are used to train the second NN configuration and generate labeled training data for the third NN configuration, and operation 1704c is used to train the third NN configuration and generate labeled training data for the fourth NN configuration. This process continues, and operation 1704P is used to train the Pth NN configuration and generate labeled training data for subsequent NN configurations. Therefore, in general, in method 1700, operation 1704i is used to train the i-th NN configuration and generate labeled training data for the (i+1)-th NN configuration, where i = 1, ..., P.

[0270] Method 1700 includes, at 1704a, (i) iteratively training a first NN configuration using a single oligonucleotide sequence, and (ii) generating labeled training data for a first 2-oligonucleotide using the trained first NN configuration. As discussed, the first NN configuration is Figure 14A The NN configuration is 1415, and the single oligonucleotide sequence contains relative to Figure 14A , Figure 14B Oligonucleotides #1 are under discussion. (Relative to...) Figure 14A , Figure 14B Iterative training of the first neural network configuration with a single oligonucleotide sequence is discussed. (Relative to...) Figure 15A , Figure 15D , Figure 15E The generation of labeled training data for the first 2-oligonucleotide using the first trained NN configuration is discussed, where the labeled training data for the first 2-oligonucleotide is... Figure 15E The training data is 1550.

[0271] Method 1700 then proceeds from 1704a to 1704b. As shown, operation 1704b is used to train a second NN configuration (e.g., using labeled training data of the first 2-oligonucleotide generated from operation 1704a), and the trained second NN configuration is used to generate additional labeled training data of 2-oligonucleotides for training a third NN configuration. Operation 1704b includes suboperations at boxes 1704b1-1704bk.

[0272] At box 1704b1, (i) the labeled training data of the first 2-oligonucleotide generated at 1704a is used to train the second NN configuration, and (ii) the labeled training data of the second 2-oligonucleotide is generated using the at least partially trained second NN configuration. As discussed, the second NN configuration is... Figure 16A The NN configuration is 1615. The second NN configuration, trained using the labeled training data from the first 2-oligonucleotide, is also shown in [the diagram / example]. Figure 16A In the middle. Relative to Figure 16B and Figure 16C The generation of labeled training data for a second 2-oligonucleotide using a second NN configuration that has been at least partially trained is discussed (e.g., it is...). Figure 16C (Training data 1650).

[0273] Method 1700 then proceeds from 1704b1 to 1704b2. At box 1704b2, (i) a second NN configuration is further trained using the labeled training data of the second 2-oligonucleotide, and (ii) labeled training data of the third 2-oligonucleotide is generated using the further trained second NN configuration. Training the second NN configuration using the labeled training data of the second 2-oligonucleotide is shown in... Figure 16D The generation of training data labeled with a third 2-oligonucleotide using a second NN configuration that has been further trained is not shown, but will be similar to that relative to... Figure 16B and Figure 16C The discussion.

[0274] Note that box 1704b1 represents the first iteration of training the second NN configuration, box 1704b2 represents the second iteration, and so on, with box 1704bk representing the k-th iteration. As discussed, relative to... Figure 16A , Figure 16B , Figure 16C The operation of box 1704b1 is discussed in detail. The operation of subsequent boxes 1704b2, ..., 1704bk will be similar to the discussion of box 1704b1.

[0275] Note that the same second NN configuration is used in all iterations 1704b1, ..., 1704bk. Therefore, these k iterations aim to iteratively train the same second NN configuration without increasing its complexity.

[0276] The training of the second neural network configuration is performed with each iteration of boxes 1704b1, 1704b2, ..., 1704bk. Because the second neural network is trained progressively at each step of iterations 1704b1, ..., 1704bk, it gradually produces relatively fewer errors in predicting base detection sequences. For example, as shown in box 1704a and also as... Figure 15E As shown, the first labeled training data of 2-oligonucleotides generated using the first NN configuration trained (i.e., training data 1550) has 44% (i.e., 4,400 out of 10,000) uncertain mappings. This is illustrated in box 1704b1 and also as... Figure 16C As shown, the labeled training data for the second 2-oligonucleotide generated using a partially trained second NN configuration (i.e., training data 1650) has 35% (i.e., 3,500 out of 10,000) uncertain mappings. As shown in box 1704b2 and only as an example, the labeled training data for the third 2-oligonucleotide generated using a further trained second NN configuration can have 32% (i.e., 3,200 out of 10,000) uncertain mappings. The percentage of uncertain mappings can gradually decrease with each iteration until it reaches approximately 20%, for example, at box 1704bk.

[0277] The number of iterations “k” used to train the second NN configuration can be based on the satisfaction of one or more convergence conditions. Once a convergence condition is satisfied, the iterations used to train the second NN configuration can end. Convergence conditions are implementation-specific and indicate the number of iterations required to train the second NN configuration. In the example, satisfying a convergence condition is an indication that further iterations may not significantly contribute to further training of the second NN configuration, and therefore the training iterations for the second NN configuration can be terminated. This paper discusses convergence conditions and some examples of their satisfaction. For example, the second NN configuration can be trained iteratively until the percentage of uncertain mappings is less than a threshold percentage. Here, the convergence condition is satisfied once the percentage of uncertain mappings becomes less than the threshold percentage. For example, for the second NN configuration, this threshold could be approximately 20%, just as an example. Therefore, at iteration k, once the threshold is satisfied, the convergence condition is satisfied, and the training of the second NN configuration ends. Thus, the method proceeds to 1704c, where the labeled training data of the Kth 2-oligonucleotide generated at box 1704bk is used to train a third NN configuration that is more complex than the second NN configuration.

[0278] In another example, iterations of the second NN configuration continue until the percentage of uncertain mappings slightly saturates (i.e., no longer significantly decreasing with successive iterations), which satisfies the convergence condition. That is, in this example, saturation below a threshold level indicates sufficient convergence of the iterative training (e.g., indicating the satisfaction of the convergence condition), and further iterations do not significantly improve the model, so iterations of the current model can end. For example, suppose that at iteration (k-2) (e.g., at box 1704b(k-2)), the percentage of uncertain mappings is 21%; at iteration (k-1) (e.g., at box 1704b(k-2)), the percentage of uncertain mappings is 20.4%; and at iteration k (e.g., at box 1704bk), the percentage of uncertain mappings is 20%. Therefore, for the last two iterations, the reduction in the percentage of uncertain mappings is relatively small (e.g., 0.6% and 0.4%, respectively), meaning that training has almost saturated and further training does not significantly improve the second NN configuration. Here, saturation is measured as the difference between the percentage of uncertain mappings during two consecutive iterations. That is, if two consecutive iterations have nearly the same percentage of uncertain mapping, further iterations may not help to further reduce that percentage, so the training iterations can be terminated. Therefore, at this stage, the iterations for the second NN configuration are terminated, and method 1700 proceeds to 1704c for the third NN configuration.

[0279] In yet another example, the number of iterations "k" is pre-specified, and completing k iterations satisfies the convergence condition, allowing training for the current NN configuration to end and the next NN configuration to begin.

[0280] Therefore, at the end of the iteration for the second NN configuration (i.e., at the end of box 1704k), method 1700 proceeds to box 1704c, where the third NN configuration is trained iteratively. The training of the third NN configuration will also include iterations similar to those discussed relative to operations 1704b1, ..., 1704bk, and will therefore not be discussed further in detail.

[0281] The process of progressively training more complex NN configurations continues until, at 1704P of method 1700, the P-th NN configuration is trained, and 2-oligonucleotide training data is generated for training the next NN configuration.

[0282] Note that in the examples and as discussed herein, the same 2-oligonucleotide sequence can be used for all iterations of boxes 1704b1, ..., 1704bk, 1704c, ..., 1704P. However, in some other examples and although not discussed herein, different 2-oligonucleotide sequences can also be used for different iterations of method 1700 in Figure 17.

[0283] As discussed, the more complex the model, the better it is trained to predict base detection. For example, at the end of training the second NN configuration, the final labeled training data generated by the second NN configuration has a 20% uncertain mapping. At the end of training the third NN configuration, the percentage of uncertain mapping decreases further. For example, during the first training iteration of the third NN configuration, the percentage of uncertain mapping can be 36% (e.g., because the third NN configuration was just trained during the first iteration), and this percentage can gradually decrease with subsequent training iterations of the third NN configuration. Assume, as Figure 17A As shown, for example, at the end of training the third NN configuration, the final labeled training data generated by the third NN configuration has a 17% uncertain mapping. This percentage of uncertain mapping increases with... Figure 17A The iterations further reduce the uncertainty, and for example, at the end of training the Pth NN configuration, the final labeled training data generated by the Pth NN configuration has a 12% uncertainty mapping. Note that training ends at, for example, a 12% uncertainty mapping when the convergence criterion (discussed earlier in this paper) is met for the Pth NN configuration. Therefore, P number of NN configurations are trained in method 1700. The number “P” can be three, four, five or more, and is implementation-specific, and can also be based on the satisfaction of one or more corresponding convergence criterions. For example, if the (P-1)th NN configuration results in a 12.05% uncertainty mapping, and if the Pth NN configuration results in a 12% uncertainty mapping, then there is a marginal improvement of 0.05% uncertainty mapping between the two NN configurations. This indicates that training new NN configurations with 2-oligonucleotide sequences is saturated. Here, saturation refers to the difference in the percentage of uncertainty mapping between two consecutive NN configurations. If the saturation is equal to or below a threshold (such as 0.1%), training of the 2-oligonucleotide sequence is terminated. In another example, the number "P" of the NN configuration can be pre-specified by the user as, for example, three, four, or more. As will be discussed later in this article, once training is complete with the NN configuration using the number P of 2-oligonucleotide sequences, additional complex analytical materials (such as 3-oligonucleotide sequences) can be used for training.

[0284] Figure 17B It shows in Figure 17AMethod 1700 concludes with example final labeled training data 1750 generated by the Pth NN configuration. As discussed, at the end of training the Pth NN configuration, the final labeled training data generated by the Pth NN configuration has 12% (or 1,200 out of 10,000) uncertain mappings. The predicted base detection sequences are selected into three categories: (i) a first category containing predicted base detection sequences mapped to oligonucleotide 1501A, (ii) a second category containing predicted base detection sequences mapped to oligonucleotide 1501B, and (iii) a third category containing predicted base detection sequences not mapped to either oligonucleotide 1501A or 1501B. Based on relative to Figure 15E and Figure 16C Discussion of training data, Figure 17B The training data 1750 will be obvious.

[0285] Figure 18A It shows Figure 14A A base detection system 1400 operates in the first iteration of the "training data consumption and training phase" of the "tri-oligonucleotide training phase" to train a base detector 1414 including a 3-oligonucleotide neural network configuration 1815. The reason for labeling the neural network configuration 1815 as "3-oligonucleotide" neural network configuration 1815 will become apparent later in this document. Figure 18A At least partially similar to Figure 16D However, with Figure 15D The difference lies in Figure 18A During training, the training data 1750 (see [reference]) generated at the end of method 1700 (e.g., by using the Pth NN configuration based on 2-oligonucleotide training) is used to label the training data. Figure 17B ).

[0286] For example, in Figure 18A The base detector 1414, which includes a 3-oligonucleotide neural network configuration 1815, predicts base detection sequences 1838a, 1838b, ..., 1838G. Figure 17B The mapping training data 1750 is now being used to further train the 3-oligonucleotide neural network configuration 1815, similar to that relative to Figure 16D Training through discussion.

[0287] Figure 18B It shows Figure 14A The base detection system 1400 operates in the "training data generation phase" of the "trioliponucleotide training phase" to train data including... Figure 18A The 3-oligonucleotide neural network configuration 1815 has a base detector 1414.

[0288] exist Figure 18BIn this embodiment, three distinct oligonucleotide sequences, 1801A, 1801B, and 1801C, are loaded into various clusters in flow-through pool 1405. By way of number only and without limiting the scope of this disclosure, it is assumed that in 10,000 clusters 1407, approximately 3,200 clusters include oligonucleotide sequence 1801A, approximately 3,300 clusters include oligonucleotide sequence 1801B, and the remaining 3,500 clusters include oligonucleotide sequence 1501C (although in another example, the three oligonucleotides may substantially equally divide the 10,000 clusters).

[0289] Sequencing machine 1404 generates sequence signals 1812a, ..., 1812G for the corresponding clusters among the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, sequencing machine 1404 generates a corresponding sequence signal 1812a, which indicates the bases of cluster 1407a used in a series of sequencing cycles. Similarly, for cluster 1407b, sequencing machine 1404 generates a corresponding sequence signal 1812b, which indicates the bases of cluster 1407b used in a series of sequencing cycles, and so on.

[0290] The base detector 1414, including the neural network configuration 1815, predicts the base detection sequences 1818a, ..., 1818G of the corresponding clusters in the plurality of clusters 1407a, ..., 1407G based on the corresponding sequence signals 1812a, ..., 1812G, for example, as relative to Figure 15A The subject of discussion.

[0291] In the implementation scheme, oligonucleotide sequences 1801A, 1801B, and 1801C were selected to have sufficient edit distances between the bases of the three oligonucleotides, for example, based on relative to Figure 15B and Figure 15C The discussion will be obvious. For example, any one of the three oligonucleotide sequences 1801A, 1801B, and 1801C is separated from the other three oligonucleotide sequences 1801A, 1801B, and 1801C by at least a threshold edit distance. As an example only, the threshold edit distance could be 4, 5, 6, 7, or even 8 bases. Therefore, the three oligonucleotides are chosen such that the three oligonucleotides are sufficiently different from each other.

[0292] Refer again Figure 18BIn this example, base detector 1414 does not know which oligonucleotide sequence is filled in which cluster. Therefore, base detector 1414 is unaware of the known oligonucleotide sequences 1801A, 1801B, and 1801C, or the mappings between the various clusters. Mapping logic 1416 receives the predicted base detection sequence 1818 and maps each predicted base detection sequence 1818 to one of the oligonucleotides 1801A, 1801B, or 1801C, or declares the uncertainty of mapping the predicted base detection sequence to any of the three oligonucleotides. Figure 18C The mapping operation is shown for (i) mapping a predicted base detection sequence to any one of the three oligonucleotides 1801A, 1801B, 1801C, or (ii) declaring that mapping a predicted base detection sequence to any one of the three oligonucleotides is indeterminate.

[0293] like Figure 18C As shown, the predicted base detection sequence 1818a has 2 bases of similarity to oligonucleotide 1801A, 5 bases of similarity to oligonucleotide 1801B, and 1 base of similarity to oligonucleotide 1801C. Assuming a threshold similarity ST of 4 (e.g., relative to those discussed in Equations 1 to 4), the predicted base detection sequence 1818a is mapped to oligonucleotide 1801B.

[0294] Similarly, in Figure 18C In the example, the predicted base detection sequence 1818b was mapped to oligonucleotide 1801C, and the mapping of the predicted base detection sequence 1818a was... Figure 18B The mapping logic 1416 is declared as indeterminate.

[0295] Figure 18D It shows from Figure 18C The mapping generates labeled training data 1850, which is used to train another neural network configuration. For example... Figure 18D As shown, some predicted base detection sequences 1818 and their corresponding sequence signals are mapped to the base sequence of oligonucleotide 1801A (i.e., true reference 1806a), some predicted base detection sequences 1818 and their corresponding sequence signals are mapped to the base sequence of oligonucleotide 1801B (i.e., true reference 1806b), some predicted base detection sequences 1818 and their corresponding sequence signals are mapped to the base sequence of oligonucleotide 1801C (i.e., true reference 1506c), and the mapping of the remaining predicted base detection sequences 1818 and their corresponding sequence signals is uncertain. Based on an earlier version of this paper... Figure 15E Discussion of training data 1550 Figure 18D The training data 1850 will be obvious.

[0296] Figure 18E A flowchart depicting an example method 1880 for iteratively training a neural network configuration for base detection using a 3-oligonucleotide benchmark truth sequence is shown. Method 1880 progressively trains an inherently progressive and monotonically complex 3-oligonucleotide NN configuration. Increasing the complexity of the NN configuration may include increasing the number of layers in the NN configuration, increasing the number of filters in the NN configuration, increasing the topological complexity in the NN configuration, etc., as relative to... Figure 17A The discussion focuses on, for example, method 1880, which involves the first 3-oligonucleotide NN configuration (which is earlier in this paper compared to...). Figure 18A The discussion covers the 3-oligonucleotide NN configurations 1815, 23-oligonucleotide NN configurations, Q-th NN configurations, and so on. In the examples, the complexity of the Q-th 3-oligonucleotide NN configuration is higher than that of the (Q-1)-th 3-oligonucleotide NN configuration, which is higher than that of the (Q-2)-th 3-oligonucleotide NN configuration, and so on. Furthermore, the complexity of the second 3-oligonucleotide NN configuration is higher than that of the first 3-oligonucleotide NN configuration. Figure 18E It is symbolically shown within the frame of 1890.

[0297] Note that in Figure 18E In method 1880, operation 1704P comes from Figure 17A In the final box of method 1700, operations 1888a1-1888am are used to iteratively train the first 3-oligonucleotide NN configuration and generate labeled training data for the second 3-oligonucleotide NN configuration, and operation 1888b is used to iteratively train the second 3-oligonucleotide NN configuration and generate labeled training data for the third 3-oligonucleotide NN configuration, and so on. This process continues, and operation 1888Q is used to train the Q-th 3-oligonucleotide NN configuration and generate labeled training data for training subsequent NN configurations. Therefore, generally, in method 1880, operation 1888i is used to train the i-th 3-oligonucleotide NN configuration and generate labeled training data for the (i+1)-th 3-oligonucleotide NN configuration, where i = 1, ..., Q.

[0298] Method 1880 includes repeating operations 1704b1, ..., 1704bk at 1704P to train the P-th NN configuration using 2-oligonucleotide benchmark ground truth data, and generating labeled training data for 2-oligonucleotides to train the next NN configuration. Figure 17A The last box of method 1700.

[0299] Method 1880 then proceeds from 1704P to 1888a1. As shown in the figure, operation 1888a is used to generate labeled training data (e.g., from the previous box, box 1704P) using the data generated from the previous box. Figure 17B The training data 1750 is used to train the first 3-oligonucleotide NN configuration (e.g., 3-oligonucleotide neural network configuration 1815), and the trained first 3-oligonucleotide NN configuration is used to generate additional labeled training data of 3-oligonucleotides for subsequent training of the second 3-oligonucleotide NN configuration. Operation 1888a includes the suboperations at boxes 1888a1-1888am.

[0300] At box 1888a1, (i) the labeled training data generated at 1704P is used to train the first 3-oligonucleotide NN configuration (e.g., Figure 18A (ii) using the first 3-oligonucleotide NN configuration (which is at least partially trained) to generate labeled training data for 3-oligonucleotides (such as...). Figure 18D (Training data 1850).

[0301] Method 1880 then proceeds from 1888a1 to 1888a2. At box 1888a2, (i) the first 3-oligonucleotide NN configuration is further trained using the labeled training data of the 3-oligonucleotides generated in the previous stage (e.g., generated at box 1888a1), and (ii) new labeled training data of the 3-oligonucleotides is generated using the further trained first 3-oligonucleotide NN configuration.

[0302] The operations discussed relative to box 1888a2 (and box 1888a2) are iteratively repeated at 1888a3, ..., 1888am. Note that boxes 1888a1, ..., 1888am are all used to train the first 3-oligonucleotide NN configuration. The number of iterations "m" can be implementation-specific and has been relative to... Figure 17A Method 1700 discusses example criteria for selecting the number of iterations used to train a particular NN model (e.g., the selection of the number of iterations "k" in this method).

[0303] After the first 3-oligonucleotide NN configuration is sufficiently or satisfactorily trained in 1888am, method 1888 proceeds to box 1888b, where the second 3-oligonucleotide NN configuration is trained iteratively. The training of the second 3-oligonucleotide NN configuration will also include iterations similar to those discussed relative to operations 1888a1, ..., 1888am, and therefore will not be discussed further in detail.

[0304] The process of progressively training more complex NN configurations continues until, at 1888Q in method 1888, the Q-th 3-oligonucleotide NN configuration is trained, and corresponding 3-oligonucleotide training data is generated for training the next NN configuration.

[0305] Figure 19 A flowchart depicting an example method 1900 for iteratively training a neural network configuration for base detection using polynucleotide benchmark truth sequences is shown. Essentially, Figure 19 Summary relative to Figures 14A to 18E The discussion. For example, Figure 19 An iterative training and labeled training data generation process using different oligonucleotide stages (such as mono-oligonucleotide stages, di-oligonucleotide stages, tri-oligonucleotide stages, and so on) is illustrated. Therefore, the complexity and / or length of the analytes used for training and generating labeled training data progressively and monotonically increases with the iterations and the complexity of the underlying neural network configuration of the base detector.

[0306] Method 1900 includes, at 1904a, iteratively training a 1-oligonucleotide neural network configuration and generating labeled training data, for example, as relative to... Figure 14A and Figure 14B as well as Figure 17A The method discussed in box 1700a of method 1704.

[0307] Method 1900 further includes, at 1904b, iteratively training one or more 2-oligonucleotide NN configurations using a dioligonucleotide sequence, and generating labeled 2-oligonucleotide training data, for example, as relative to... Figure 17A The method discussed in boxes 1704b1-1704P of 1700.

[0308] Method 1900 further includes, at 1904c, iteratively training one or more 3-oligonucleotide NN configurations using a trioligonucleotide sequence, and generating labeled 3-oligonucleotide training data, for example, as relative to... Figure 18E The method discussed in the boxes 1888a1-1888Q of 1880.

[0309] The process continues, and can be progressively increased using a higher number of oligonucleotide sequences. Finally, at 1904N, one or more N-oligonucleotide NN configurations are trained using the N-oligonucleotide sequence, and labeled training data for the corresponding N-oligonucleotides are generated, where N can be a suitable positive integer greater than or equal to 2. Based on the discussion relative to the operations at 1904b and 1904c, the operation at 1904N will be self-evident.

[0310] Figures 14A to 19This is associated with training NN models using synthetically sequenced simple oligonucleotide sequences. For example, the oligonucleotide sequences used in these diagrams may have fewer bases compared to sequences found in the DNA of organisms. In the implementation, relative to Figures 14A to 19 The oligonucleotide-based training discussed is used to progressively train complex neural network models and generate progressively rich labeled training datasets. For example, Figure 19 The N-Oligonucleotide NN configuration outputs a labeled training dataset of N-oligonucleotides, where the labeled training dataset of N-oligonucleotides can be richer, more diverse, and larger than the labeled training dataset associated with "less than N" number of oligonucleotides.

[0311] However, in practice, the sequencing machine 1404 and base detector 1414 are used for base detection of sequences far more complex than simple oligonucleotide sequences. For example, in practice, the sequencing machine 1404 and base detector 1414 are used for base detection of biological sequences far more complex than simple oligonucleotide sequences. Therefore, the base detector 1414 must be trained on base sequences that are more complex than oligonucleotide sequences found in the DNA and RNA of organisms.

[0312] Figure 20A The training method is shown. Figure 14A The base detector 1414 is a biological sequence 2000. The biological sequence can be an organism with relatively few bases, such as phix (also known as phi X). Phages are single-stranded DNA (ssDNA) viruses. Phage 174 is an ssDNA virus that infects E. coli and was the first DNA-based genome sequenced in 1977. Phage (such as ΦX174) virus particles have also been successfully assembled in vitro. In the embodiment, oligonucleotide sequences (such as relative to...) are used... Figures 14A to 19 After training the base detector 1414 (as discussed), it can be further trained with simple biological DNA (such as phix DNA), although this does not limit the scope of this disclosure. For example, instead of phix, more complex organisms, such as bacteria (such as Escherichia coli or Escherichia coli), can be used. Thus, the organism sequence 2000 can be phix or another relatively simple biological DNA. The organism sequence 2000 is pre-sequentially sequenced, that is, the base sequence of the organism sequence 2000 is known a priori (e.g., by a different...). Figure 14A The sequencing machines shown are those with already trained base detectors for sequencing.

[0313] like Figure 20A As shown, when the organism sequence 2000 is loaded into... Figure 14AIn the sequencing machine 1404, the biological sequence 2000 is partitioned or segmented into multiple subsequences 2004a, 2004b, ..., 2004N. Each subsequence is loaded into one or more corresponding clusters. Thus, each cluster 1407 is filled with the corresponding subsequence 2004 and its synthetic copy. The biological sequence 2000 can be partitioned using any suitable criteria, such as the maximum size of the subsequences that can be filled into a cluster. For example, if a single cluster in the flow cell can be filled with a subsequence having a maximum of about 150 bases, then partitioning can be performed accordingly, such that a single subsequence in subsequence 2004 has a maximum of 150 bases. In one example, single subsequences 2004 may have substantially the same number of bases; in another example, single subsequences 2004 may have different numbers of bases. Subsequence 2004b, used as an example for discussing the teachings of this disclosure, is assumed to have an L1 number of bases. As an example only, the quantity L1 can be between 100 and 200, although it can have any other suitable value and is implementation-specific.

[0314] Figure 20B It shows Figure 14A The base detection system 1400 operates during the training data generation phase of the first organism training phase to use... Figure 20A The first organism sequence 2000 subsequences 2004a, ..., 2004S are used to train the base detector 1414, which includes the first organism-level neural network configuration 2015.

[0315] Note that, although not in Figure 20B As shown, but the first organism-level NN configuration initially used in 2015 was from... Figure 19 The method 1904 uses labeled training data of N-oligonucleotides for training. Therefore, the first organism-level NN configuration 2015 is at least partially pre-trained. Figure 20B The base detection system 1400 and Figure 14A The base detection systems are the same, although the base detection system 1400 uses different neural network configurations and different analytes in the two figures.

[0316] As discussed, subsequences 2004a, ..., 2004S are loaded into the corresponding clusters 1407. For example, subsequence 2004a is loaded into cluster 1407a, subsequence 2004b into cluster 1407b, and so on. Note that each cluster 1407 will include multiple sequencing copies of the same subsequence 2004. For example, the subsequence loaded into a cluster will be synthetically replicated so that the cluster has multiple copies of the same subsequence, which helps to generate the corresponding sequence signal 2012 for that cluster.

[0317] Note that base detector 1414 is unaware of which cluster is filled with which subsequence. For example, if subsequence 2004a and its synthetic copy are loaded into a particular cluster, base detector 1414 will not know which cluster is filled by subsequence 2004a. As will be discussed later in this document, mapping logic 1416 is designed to map a single subsequence 2004 to its corresponding cluster 1407 to facilitate the training process.

[0318] Sequencing machine 1404 generates sequence signals 2012a, ..., 2012G for the corresponding clusters in the plurality of clusters 1407a, ..., 1407G. For example, for cluster 1407a, sequencing machine 1404 generates a corresponding sequence signal 2012a, which indicates the bases of cluster 1407a used in a series of sequencing cycles. Similarly, for cluster 1407b, sequencing machine 1404 generates a corresponding sequence signal 2012b, which indicates the bases of cluster 1407b used in a series of sequencing cycles, and so on.

[0319] In the example, although a single subsequence 2004 is loaded into the corresponding cluster 1407, the base detector 1414 does not know which subsequence is loaded into which cluster. Therefore, the base detector 1414 does not know the mapping between subsequence 2004 and cluster 1407. When each cluster 1407 generates a corresponding sequence signal 2012, the base detector 1414 does not know the mapping between subsequence 2004 and sequence signal 2012.

[0320] The base detector 1414, including the neural network configuration 2015, predicts the base detection sequences 2018a, ..., 2018G of the corresponding clusters 1407a, ..., 1407G based on the corresponding sequence signals 2012a, ..., 2012G. For example, for cluster 1407a, the base detector 1414 predicts the corresponding base detection sequence 2018a based on the corresponding sequence signal 2012a, including base detection for cluster 1407a for a series of sequencing cycles. Similarly, for cluster 1407b, the base detector 1414 predicts the corresponding base detection sequence 2018b based on the corresponding sequence signal 2012b, including base detection for cluster 1407b for a series of sequencing cycles, and so on.

[0321] Note that the neural network configuration 2015 is only partially trained, and not fully trained. Therefore, the neural network configuration 2015 cannot accurately predict some or most bases of a single subsequence.

[0322] Furthermore, as base detection is performed in subsequences, bases become increasingly difficult to detect, for example, due to fading and / or noise-induced phasing or pre-phasing. Figure 20CAn example of fading is shown, where the signal intensity decreases with the number of cycles in a sequencing run as a base detection operation. Fading is the exponential decay of fluorescence signal intensity with increasing cycle number. As sequencing runs proceed, analyte strands are overwashed, exposed to laser radiation that produces reactive substances, and subjected to harsh environmental conditions. All of these lead to the gradual loss of fragments in each analyte, thus reducing its fluorescence signal intensity. Fading is also known as darkening or signal attenuation. Figure 20C An example of fading 2000C is shown. Figure 20C In the analysis, the intensity values ​​of the analyte fragments with AC microsatellites exhibit exponential decay.

[0323] Figure 20D This conceptually illustrates the decrease in signal-to-noise ratio (SNR) as sequencing cycles progress. For example, accurate base detection becomes increasingly difficult as sequencing proceeds due to decreasing signal strength and increasing noise, resulting in a significant decrease in SNR. Physically, it is observed that later synthesis steps attach tags at different locations relative to the sensor compared to earlier synthesis steps. When the sensor is located below the sequence being synthesized, signal attenuation occurs because, in later sequencing steps, the tags attach to strands farther from the sensor compared to earlier steps. This leads to signal attenuation as sequencing cycles progress. In some designs, where the sensor is located above a substrate holding the clusters, the signal may increase rather than decrease as sequencing progresses.

[0324] In the flow cell design studied, noise increases as the signal decays. Physically, phasing and pre-phasing increase noise as sequencing progresses. Phosing refers to the step in sequencing where the tag fails to advance along the sequence. Pre-phasing refers to a sequencing step where the tag jumps forward two positions instead of one during a sequencing cycle. Phosing and pre-phasing are relatively infrequent, occurring approximately once every 500 to 1000 cycles. Phosing is slightly more frequent than pre-phasing. Phosing and pre-phasing affect individual strands within clusters that produce intensity data; therefore, as sequencing progresses, the intensity noise distribution from the clusters accumulates into binomial, trinomial, tetranomial, and other expansions.

[0325] Fading, signal attenuation, and reduced signal-to-noise ratio, and Figure 20C and Figure 20D Further details can be found in U.S. non-provisional patent application No. 16 / 874,599 (Attorney’s File No. ILLM 1011-4 / IP-1750-US), filed May 14, 2020, entitled “Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing,” which is incorporated herein by reference as if fully set forth herein.

[0326] Therefore, during base detection, the reliability or predictability of base detection decreases as the sequencing cycle progresses. For example, referring to a specific subsequence, such as Figure 20A For the subsequence 2004b, detection of bases 1 to 10 is generally more reliable than detection of bases 10-20 or 50-60. In other words, the first few bases of the L1 sequence of 2004b are likely to be predicted more accurately than the remaining bases of the L1 sequence.

[0327] Figure 20E The base detection of the first L2 number of bases in the L1 number of bases of subsequence 2004b is shown, wherein the first L2 number of bases of subsequence 2004b are used to map subsequence 2004b to sequence 2000.

[0328] For example, refer to Figure 20A , Figure 20B and Figure 20E The sequencing machine 1404 generates a sequence signal 2012b corresponding to subsequence 2004b (i.e., assuming subsequence 2004b is filled in cluster 1407b). However, the base detector 1414 does not know the position of the subsequence corresponding to sequence signal 2012b in sequence 2000. That is, the base detector 1414 does not know that a specific subsequence 2004b is loaded into cluster 1407b.

[0329] like Figure 20E As shown, a partially trained NN configuration from 2015 (e.g., using from...) Figure 19 Method 1904 (trained on labeled training data of N-oligonucleotides) receives sequence signal 2012b and predicts L1 bases indicated by sequence signal 2012b. The prediction of L1 bases includes the prediction of the first L2 bases, wherein the prediction of the first L2 number of bases of subsequence 2004b is used to map subsequence 2004b to sequence 2000.

[0330] In the example, the quantity L2 is 10. The quantity L2 can be any suitable quantity, such as 8, 10, 12, 13, etc., as long as L2 is relatively less than L1. For example, L2 is less than 10% of L1, less than 25% of L1, etc.

[0331] For example, the first L2 bases of the subsequence 2004b predicted by NN configuration 2015 are A, C, C, T, G, A, G, C, G, A, as shown in the figure. Figure 20E As shown. The prediction of further mining (L1-L2) bases is in Figure 20E In Chinese, these are usually represented as B1, ..., B1.

[0332] Now, it's possible that NN configuration 2015 has correctly predicted the first L2 number of bases, or that there may be one or more errors in these L2 number of base predictions. Mapping logic 1416 attempts to map the first L2 number of base predictions to the corresponding consecutive L2 bases in the organism sequence 2000. In other words, mapping logic 1416 attempts to match the first L2 number of base predictions with the consecutive L2 bases in the organism sequence 2000, so that the subsequence 2004b within the organism sequence 2000 can be identified.

[0333] like Figure 20E As shown, mapping logic 1416 can find “basic” and “unique” matches between the predicted first L2 number of bases for subsequence 2004b and the consecutive L2 bases in organism sequence 2000. Note that a “basic” match means that the match may not be 100%, and there may be one or more errors in the match. For example, the first L2 number of bases predicted by NN configuration 2015 for subsequence 2004b is A, C, C, T, G, A, G, C, G, A, while the corresponding basic match in organism sequence 2000 is A, G, C, T, G, A, G, C, G, A. Therefore, the second base in these two L2 base sequences does not match, but the remaining bases do. Mapping logic 1416 declares a match between the two L2 number of base segments as long as the number of such mismatches is less than a threshold percentage. The threshold percentage for mismatches can be 10% or 20% of the number of L2 bases, or some similar percentage. Therefore, in this example, L2 is 10 and matching logic 1416 tolerates up to 2 mismatches (or 20% mismatches). Thus, mapping logic 1416 is designed to map the first L2 number of bases predicted for subsequence 2004b, or a slight variation thereof (e.g., where the variation implies an error tolerance during matching), to consecutive L2 bases in biological sequence 2000. The value of the threshold percentage can be implementation-specific and can be user-configurable. As an example only, the threshold percentage can have a relatively high value (e.g., 20%) during the initial iterations of training; and the threshold percentage can have a relatively low value (e.g., 10%) during later iterations of training. Therefore, the threshold percentage can be relatively high in the early stages of training iterations because the probability of errors in base detection prediction is relatively high. As the NN configuration is better trained, they are likely to make better base detection predictions, so the threshold percentage can gradually decrease. However, in another example, the threshold percentage can be the same throughout all iterations of training.

[0334] Furthermore, in the example, the match between two L2-number bases must be unique for a proper mapping, and a non-unique match could lead to the match and mapping being declared indeterminate. Therefore, the predicted first L2-number bases for subsequence 2004b (or a slight variation thereof) must occur only once in organism sequence 2000 for the match and mapping to be valid. Typically, for the actual base sequences of simpler organisms, the probability that consecutive L2 bases (or minor variations thereof) will occur only once in organism sequence 2000 is high.

[0335] For example, refer to Figure 20E For example, if the consecutive bases A, G, C, T, G, A, G, C, G, A, G, A, G, G, C, G, A appear in one part of organism sequence 2000, and the consecutive bases A, C, A, T, G, A, G, C, G, A appear in another part of organism sequence 2000, then both parts of organism sequence 2000 can match the first L2 number of bases of subsequence 2004b predicted by NN configuration 2015 (which is A, C, C, T, G, A, G, C, G, A, G). Therefore, in this example, the match is not unique, and mapping logic 1416 does not know which of the two parts of organism sequence 2000 is mapped to the L2 number of bases of subsequence 2004b. In this scenario, mapping logic 1416 declares that there is no reliable match (i.e., declares an indeterminate mapping).

[0336] refer to Figure 20E As shown in the figure, the first L2 number of bases of the subsequence 2004b predicted by NN configuration 2015 "substantially" and "uniquely" match the corresponding L2 number of consecutive bases of the organism sequence 2000. It is also assumed that a portion 2000B of the organism sequence 2000 (which has L1 bases) matches the first L2 predictions of the subsequence 2004b "substantially" and "uniquely" the first L2 bases of a portion B of the organism sequence 2000. Therefore, most likely, the subsequence 2004b is actually a portion 2000B of the organism sequence 2000. In other words, most likely, in Figure 20A The 2000B portion of the biological sequence 2000 is split to form the subsequence 2004b.

[0337] Therefore, part 2000B of the organism sequence 2000 serves as the reference true value of the sequence signal 2012b corresponding to the subsequence 2004b. Figure 20F It shows from Figure 20E The labeled training data 2050 generated by the mapping includes... Figure 20A A portion of the biological sequence 2000 was used as the baseline truth value.

[0338] exist Figure 20F In the labeled training data 2050, as an example only, due to uncertain mapping, subsequences 2004a and 2004d do not map to any part of the organism sequence 2000. For example, as relative to Figure 20E The discussion posits that a fundamental and unique match must exist between the first L2 bases of a subsequence and its corresponding portion of the organism sequence 2000 for mapping logic 1416 to declare a conclusive mapping. NN configuration 2015 may produce a relatively high number of errors in the first L2 bases of each subsequence 2004a, 2004d, resulting in these subsequences not being able to be mapped to any corresponding portion of the organism sequence 2000.

[0339] exist Figure 20F In the labeled training data 2050, subsequence 2004b (and therefore sequence signal 2012b) is mapped to portion 2000B of the biological sequence 2000, as relative to... Figure 20E The discussion continues. Similarly, subsequence 2004c is mapped to portion 2000C of biological sequence 2000, and subsequence 2004S is mapped to portion 2000S of biological sequence 2000. For example, subsequence 2004c is mapped to portion 2000C of biological sequence 2000 (e.g., having the same number of bases as subsequence 2004c) such that the first L2 bases of subsequence 2004c are predicted to match the first L2 bases of portion 2000C "substantially" and "uniquely".

[0340] Figure 20G It shows Figure 14A A base detection system 1400 operates in the "training data consumption and training phase" of a "biological level training phase" to train a base detector 1414 including a first biological level neural network configuration 2015. For example, Figure 20F 2050 labeled training data were used Figure 20G Training.

[0341] For example, the L1 bases of the subsequence 2004b predicted by base detector 1414 are compared with a portion 2000B of the biological sequence 2000. Note that the L1 bases of the subsequence 2004b predicted by base detector 1414 are compared with the biological sequence 2000 to generate... Figure 20F The first L2 bases of the mapping. When generated Figure 20F During mapping, the remaining (L1-L2) bases are not compared because they may contain many errors. This is because, relative to... Figure 20C and Figure 20DThe bases discussed here have a higher chance of misprediction due to decay, phasing, and / or pre-phasing. Figure 20G In this process, all L1 bases of the subsequence 2004b predicted by the base detector 1414 are compared with the corresponding L1 bases on part 2000B of the biological sequence 2000.

[0342] therefore, Figure 20F The mapping specifies a portion of organism sequence 2000 (i.e., portion 2000B), and subsequence 2004b will be mapped to that portion. Figure 20G The comparison is performed. Once the mapping is complete and labeled training data 2050 is generated, then... Figure 20G Labeled training data 2050 is used for comparison and generation of error signals, which are used for gradient updates 2017 in the backward channel of NN configuration 2015 and the training of the results of NN configuration 2015.

[0343] Note that some subsequences (such as subsequences 2004a and 2004d, see...) Figure 20F There was no corresponding part in the final matching organism sequence 2000, therefore, in Figure 20G The training does not use base detection predictions corresponding to these subsequences.

[0344] Figure 21 A depiction for use is shown Figure 20A A flowchart of an example method 2100 for iteratively training a neural network configuration for base detection using a simple biological sequence 2000 is provided. Method 2100 progressively trains an inherently monotonically complex NN configuration. As discussed earlier in this paper, increasing the complexity of the NN configuration can include increasing the number of layers in the NN configuration, increasing the number of filters in the NN configuration, increasing the topological complexity in the NN configuration, etc. For example, method 2100 involves a first biological-level NN configuration (which is the one discussed earlier in this paper relative to...) Figure 20B , 20G (The following diagrams are discussed in conjunction with other diagrams: NN configuration 2015), second organism-level NN configuration, R-th organism-level NN configuration, and so on. In the example, the complexity of the R-th organism-level NN configuration is higher than that of the (R-1)-th organism-level NN configuration, which is higher than that of the (R-2)-th organism-level NN configuration, and so on, with the second organism-level NN configuration being more complex than the first organism-level NN configuration.

[0345] Note that in method 2100, operation 2104a (which includes boxes 2104a1, ..., 2104am) is used to train the first organism-level NN configuration and generate labeled training data for the second organism-level NN configuration, operation 2104b is used to train the second organism-level NN configuration and generate labeled training data for the third organism-level NN configuration, and so on. This process continues, and finally operation 2104R is used to train the Rth organism-level NN configuration and generate labeled training data for the next stage NN configuration. Therefore, in general, in method 2100, operation 2104i is used to train the i-th organism-level NN configuration and generate labeled training data for the (i+1)-th organism-level NN configuration, where i = 1, ..., R.

[0346] Method 2100 includes, at 2104a1, (i) using from Figure 19 The method uses labeled training data of N-oligonucleotides from 1900 to 1904N to train the first organism-level NN configuration (e.g., Figure 20B The biological-level NN configuration 2015, although the training of this NN configuration was not in Figure 20B (as shown in the diagram), and (ii) using at least partially trained first-organism-level NN configuration 2015 to generate labeled training data. Labeled training data in Figure 20F As shown in the figure, its generation is relative to Figure 20E and Figure 20F discuss.

[0347] Method 2100 then proceeds from 2104a1 to 2014a2, during which a second iteration of training the first organism-level NN configuration 2015 is performed. For example, at 2104a2, (i) the first organism-level NN configuration 2015 is further trained using labeled training data from the previous stage, for example, as relative to Figure 20G The discussed; and (ii) using the first biological-level NN configuration 2015, which was at least partially trained, to generate further labeled training data (e.g., similar to that relative to...). Figure 20E and 20F (Discussion).

[0348] The training and generation operations are repeated iteratively, culminating in the training of the first organism-level NN configuration 2015 at box 2104am. Note that box 2014a1 represents the first iteration of training the first organism-level NN configuration 2015, box 2104a2 represents the second iteration, and so on, with the final box 2104am representing the m-th iteration of training the first organism-level NIN configuration 2015. The number of iterations can be based on one or more factors, such as those previously mentioned in this paper relative to... Figure 17A The factors discussed in Method 1700 (e.g., the criteria for selecting the number of iterations "k" are discussed). The complexity of the first biological-level NN configuration 2015 does not change during iterations 2104a1, ..., 2104am.

[0349] At the end of the iteration of the first organism-level NN configuration 2015 (i.e., at the end of box 2104am), method 2100 proceeds to box 2104b, where the second organism-level NN configuration is now trained iteratively. The training of the second organism-level NN configuration and the associated generation of training labeled data will also involve iterations similar to those discussed relative to operations 2104a1, ..., 2104am, and therefore will not be discussed further in detail.

[0350] The process of progressively training more complex NN configurations associated with the generation of training labeled data continues until, at 2104R of method 2100, the Rth organism-level NN configuration is trained, and corresponding labeled training data is generated for training the next NN configuration.

[0351] Figure 22 The training method is shown. Figure 14A The base detector 1414 is used in the NN configuration corresponding to complex biological sequences. For example, as relative to Figures 20A to 21 The discussed relatively simple biological sequences 2000, each containing approximately L1 bases, are used to iteratively train a simple biological-level neural network configuration of R bases and generate corresponding labeled training data. For example, Figure 21 Method 2100 illustrates this iterative training and labeling of training data using a simple organism sequence 2000. As discussed, the simple organism sequence 2000 can be a Phix or another organism with a relatively simple (or relatively small) genetic sequence.

[0352] Figure 22The use of a relatively complex biological sequence 2200a is also illustrated. Biological sequence 2200a is more complex than biological sequence 2000 because, for example, the number of bases in complex biological sequence 2200a is greater than the number of bases in biological sequence 2000. By way of example only, biological sequence 2000 may have approximately 1 million bases, while complex biological sequence 2200a may have 4 million bases. In another example, each subsequence segmented from complex biological sequence 2200a has a higher number of bases than each subsequence segmented from biological sequence 2000. In yet another example, the number of subsequences segmented from complex biological sequence 2200a is greater than the number of subsequences segmented from biological sequence 2000. For example, when segmenting complex organism sequence 2200a and organism sequence 2000, the number of subsequences segmented from complex organism sequence 2200a will be greater than the number of subsequences segmented from organism sequence 2000 because (i) complex organism sequence 2200a has a higher number of bases than organism sequence 2000, and (ii) each subsequence can have at most a threshold number of bases. In the example, complex organism sequence 2200a contains genetic material from bacteria, such as Escherichia coli, or another suitable organism sequence that is more complex than organism sequence 2000.

[0353] like Figure 22 As shown, complex organism sequences 2200a are used to iteratively train a complex organism-level neural network configuration of Ra number and generate labeled training data. The training and generation of labeled training data is similar to that relative to... Figure 21 The methods discussed in Method 2100 (the difference being that Method 2100 is specifically for biological sequences 2000, while complex biological sequences 2200a are used here).

[0354] The iterative process continues, culminating in the use of a relatively more complex biological sequence, 2200T. This additional complex biological sequence, 2200T, is more complex than biological sequences 2000 and 2200a. For example, the additional complex biological sequence 2200T has a higher base count than either biological sequences 2000 or 2200a. In another example, each subsequence segmented from the additional complex biological sequence 2200T has a higher base count than each subsequence segmented from biological sequences 2000 or 2200a. In yet another example, the number of subsequences segmented from the additional complex biological sequence 2200T is greater than the number of subsequences segmented from biological sequences 2000 or 2200a. In these examples, the additional complex biological sequence 2200T contains genetic material from complex species, such as those from humans or other mammals.

[0355] like Figure 22As shown, biological sequences 2200T are used to iteratively train additional complex biological-level neural network configurations with a number of RTs and to generate labeled training data. The training and generation of labeled training data is similar to that of... Figure 21 The methods discussed in Method 2100 (the difference being that Method 2100 is specifically for biological sequence 2000, while biological sequence 2000T is used here).

[0356] Figure 23A A flowchart depicting an example method 2300 for iteratively training a neural network configuration for base detection is shown. Method 2300 summarizes the paper's approach relative to... Figures 14A to 22 At least some implementation schemes and examples are discussed. Method 2300 trains inherently monotonously complex NN configurations, as discussed herein. Method 2300 also monotonously uses complex genetic sequences as analytes. Method 2300 is used to train base detectors 1414 for the various graphs discussed herein.

[0357] Method 2300 begins at 2304, where, relative to Figure 17A The method discussed in box 1704 of method 1700 uses a single oligonucleotide benchmark ground truth data to iteratively train a base detector 1414 including NN configuration 1415 (e.g., see...). Figure 14A ). Figure 14A At least part of the trained NN configuration 1415 is used to generate labeled training data, also as relative to Figure 17A The method discussed in box 1704 of 1700.

[0358] Method 2300 then proceeds from 2304 to 2308, where one or more NN configurations are iteratively trained using 2-oligonucleotide sequences, and corresponding labeled training data is generated, for example, as relative to Figure 17A The method discussed in 1700.

[0359] Method 2300 then proceeds from 2308 to 2312, where one or more NN configurations are iteratively trained using 3-oligonucleotide sequences, and corresponding labeled training data is generated, for example, as relative to Figure 19 The method discussed in 1900.

[0360] The process of training the NN configurations using progressively higher numbers of oligonucleotides continues until, at position 2316, one or more NN configurations are iteratively trained using N-oligonucleotide sequences, and corresponding labeled training data is generated, for example, as relative to... Figure 19 The method discussed in 1900.

[0361] Method 2300 then transitions to 2320, where training and labeling of training data generation involve organisms. At 2320, simple organism sequences, such as... Figure 20A Simple biological sequences 2000. One or more NN configurations can be trained using simple biological sequences (see, for example, see...). Figure 21 Method 2100), and generated labeled training data.

[0362] As method 2300 proceeds from 2320, it uses progressively more complex biological sequences, for example, as relative to... Figure 22 The discussion continues. Finally, at 2328, a complex organism sequence (e.g., Figure 22 Another complex biological sequence (2200T) is used to iteratively train one or more NN configurations and generate corresponding labeled training data.

[0363] Therefore, method 2300 continues until the base detector 1414 is "fully trained". "Fully trained" implies that the base detector 1414 can now detect bases at an error rate less than the target error rate. As discussed, the training process can continue iteratively until sufficient training and the target error rate for base detection are achieved (e.g., see [link to relevant documentation]). Figure 23E (See the "error rate" chart). At the end of Method 2300, the base detector 1414, which includes the final NN configuration of Method 2300, is now fully trained. Therefore, the trained base detector 1414, which includes the final NN configuration of Method 2300, can now be used for inference, for example, for sequencing unknown genetic sequences.

[0364] Figures 23B to 23E Various diagrams illustrating the effectiveness of the base detector training process discussed in this disclosure are shown. References Figure 23BChart 2360 shows the percentage of mappings generated from training data generated by: (i) a first 2-oligonucleotide NN configuration (such as NN configuration 1615) trained using the neural network-based training data generation techniques discussed herein and (ii) NN configurations trained using conventional 2-oligonucleotide training data generation techniques. The white bars in Chart 2360 show mapping data from the first 2-oligonucleotide NN configuration trained using training data generated using the neural network-based models discussed herein. Therefore, the white bars in Chart 2360 show mapping data generated using the various techniques discussed herein. The gray bars in Chart 2360 show data associated with NN configurations trained using data generated from conventional non-neural network-based models (such as Real-Time Analytics (RTA) models). An example of an RTA model is discussed in U.S. Patent No. US10304189B2, entitled “Data processing system and methods,” published May 28, 2019, which is incorporated herein by reference as if fully set forth herein. Therefore, the gray bars in Chart 2360 show mapping data generated using conventional techniques. In the example, the white bars in chart 2360 can be seen... Figure 17A Method 1700 is generated at operation 1704b1. Figure 2360 shows the percentage of base detection predictions mapped to oligonucleotide 1, the percentage of base detection predictions mapped to oligonucleotide 2, and the percentage of base detection predictions that cannot ultimately be mapped to either oligonucleotide 1 or 2 (i.e., the uncertainty percentage). As shown, the uncertainty percentage of training data generated using the technique discussed herein is slightly higher than that of training data generated using conventional techniques. Therefore, initially (e.g., at the start of a training iteration), conventional techniques are slightly superior to the training data generation technique discussed herein.

[0365] Now for reference Figure 23C The chart shown is 2365, illustrating the percentage of mappings in training data generated using the following: (i) a first 2-oligonucleotide NN configuration (such as NN configuration 1615) trained using the neural network-based training data generation techniques discussed herein (white bars), (ii) a second 2-oligonucleotide NN configuration trained using the neural network-based training data generation techniques discussed herein (dashed bars), and (iii) an NN configuration trained using conventional 2-oligonucleotide training data generation techniques (such as conventional training data generation techniques based on RTA) (grey bars). In the example, the first 2-oligonucleotide NN configuration (white bars) and the second 2-oligonucleotide NN configuration (dashed bars) correspond to... Figure 17AMethods 1700 are described in operations 1704b and 1704c. Figure 2365 shows the percentage of base detection predictions mapped to oligonucleotide 1, the percentage of base detection predictions mapped to oligonucleotide 2, and the percentage of base detection predictions that cannot ultimately be mapped to either oligonucleotide 1 or 2 (i.e., the uncertainty percentage). As shown, the uncertainty percentage of the training data generated using the first 2-oligonucleotide NN configuration is higher than each of (i) the training data generated using the second 2-oligonucleotide NN configuration and (ii) the training data generated using conventional techniques. Furthermore, the uncertainty percentage of the training data generated using the second 2-oligonucleotide NN configuration is almost comparable to that of the training data generated using conventional techniques. Therefore, through iteration and more complex NN configurations, the training data generated using NN-based configurations is almost comparable to the training data generated using conventional techniques.

[0366] Now for reference Figure 23D The figure shown is Graph 2370, illustrating the percentage of mappings of training data generated from: (i) a first 4-oligonucleotide NN configuration trained using the neural network-based training data generation techniques discussed herein (white bars) and (ii) an NN configuration trained using conventional 4-oligonucleotide training data generation techniques (e.g., RTA-based techniques) (grey bars). As shown, the percentage of uncertainty in the training data generated using the techniques discussed herein is comparable to the percentage of uncertainty in the training data generated using conventional techniques. Therefore, when training is converted to 4-oligonucleotide sequences, the conventional techniques and training data generation techniques discussed herein produce comparable results.

[0367] Now for reference Figure 23E The figure shown is 2375, which illustrates the error rates in data generated from: (i) NN configurations trained using the complex biological sequences discussed herein, for example, relative to Figure 23A Method 2300 describes operation 2328 (solid line), and (ii) describes the NN configuration trained using conventional complex biological training data generation techniques, such as RTA-based techniques (dashed line). As shown, the error rate of data generated using the techniques discussed herein is comparable to that generated using conventional techniques. Therefore, the conventional techniques and training data generation techniques discussed herein produce comparable results. As discussed, when, for example, conventional techniques are unavailable or not ready for training data generation, the training data generation techniques discussed herein can be used instead of conventional techniques.

[0368] Figure 24This is a block diagram of a specific implementation of a base detection system 2400. The base detection system 2400 is operable to obtain any information or data relating to at least one of a biological substance or a chemical substance. In some implementations, the base detection system 2400 is a workstation that may resemble a desktop device or a desktop computer. For example, most (or all) of the systems and components used to perform the desired reaction may be located within a common housing 2416.

[0369] In certain specific implementations, the base detection system 2400 is a nucleic acid sequencing system (or sequencer) configured for a variety of applications, including but not limited to de novo sequencing, resequencing of whole genomes or target genome regions, and metagenomics. The sequencer can also be used for DNA or RNA analysis. In some implementations, the base detection system 2400 can also be configured to generate reaction sites in a biosensor. For example, the base detection system 2400 can be configured to receive a sample and generate surface-attached clusters of clonal amplified nucleic acids derived from the sample. Each cluster can constitute a reaction site in the biosensor or be part of it.

[0370] An exemplary base detection system 2400 may include a system socket or interface 2412 configured to interact with a biosensor 2402 to perform a desired reaction within the biosensor 2402. (The following is in contrast to...) Figure 24 In the description, the biosensor 2402 is loaded into the system socket 2412. However, it should be understood that a cartridge including the biosensor 2402 can be inserted into the system socket 2412, and in some states, the cartridge can be temporarily or permanently removed. As described above, among other things, the cartridge may also include fluid control components and fluid storage components.

[0371] In a particular implementation, the base detection system 2400 is configured to perform a large number of parallel reactions within a biosensor 2402. The biosensor 2402 includes one or more reaction sites where the desired reaction can occur. These reaction sites may be, for example, fixed to a solid surface of the biosensor or to a bead (or other movable substrate) located within a corresponding reaction chamber of the biosensor. The reaction sites may include, for example, clusters of cloned and amplified nucleic acids. The biosensor 2402 may include a solid-state imaging device (e.g., a CCD or CMOS imaging device) and a flow cell mounted thereon. The flow cell may include one or more flow channels that receive a solution from the base detection system 2400 and direct the solution to the reaction sites. Optionally, the biosensor 2402 may be configured to incorporate a thermal element for transferring heat energy into or out of the flow channels.

[0372] The base detection system 2400 may include various components, parts, and systems (or subsystems) that interact with each other to perform a predetermined method or assay protocol for biological or chemical analysis. For example, the base detection system 2400 includes a system controller 2404 that can communicate with the various components, parts, and subsystems of the base detection system 2400 and the biosensor 2402. For example, in addition to the system socket 2412, the base detection system 2400 may also include a fluid control system 2406 to control the flow of fluids throughout the fluid network of the base detection system 2400 and the biosensor 2402; a fluid storage system 2408 configured to hold all fluids (e.g., gases or liquids) usable by the bioassay system; a temperature control system 2410 that regulates the temperature of the fluids in the fluid network, the fluid storage system 2408, and / or the biosensor 2402; and an irradiation system 2409 configured to illuminate the biosensor 2402. As described above, if a cartridge with a biosensor 2402 is loaded into a system socket 2412, the cartridge may also include fluid control components and fluid storage components.

[0373] As also shown in the figure, the base detection system 2400 may include a user interface 2414 for user interaction. For example, the user interface 2414 may include a display 2413 for displaying or requesting information from the user and a user input device 2415 for receiving user input. In some embodiments, the display 2413 and the user input device 2415 are the same device. For example, the user interface 2414 may include a touch-sensitive display configured to detect the presence of an individual touch and also identify the location of the touch on the display. However, other user input devices 2415 may be used, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc. As will be discussed in more detail below, the base detection system 2400 may communicate with various components, including a biosensor 2402 (e.g., in the form of a cartridge), to perform a desired reaction. The base detection system 2400 may also be configured to analyze data obtained from the biosensor to provide the user with the desired information.

[0374] System controller 2404 may include any processor-based or microprocessor-based system, including those using microcontrollers, reduced instruction set computers (RISCs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), logic circuits, and any other circuitry or processor capable of performing the functions described herein. The examples above are merely illustrative and are therefore not intended to limit the definition and / or meaning of the term "system controller" in any way. In an exemplary specific implementation, system controller 2404 executes a set of instructions stored in one or more storage elements, memories, or modules to perform at least one of acquiring detection data and analyzing the detection data. Detection data may include multiple pixel signal sequences, such that a pixel signal sequence from each of millions of sensors (or pixels) can be detected within a plurality of base detection cycles. Storage elements may be in the form of information sources or physical memory elements within base detection system 2400.

[0375] The instruction set may include various commands instructing the base detection system 2400 or the biosensor 2402 to perform specific operations (such as the methods and processes of various specific embodiments described herein). The instruction set may be in the form of a software program that forms part of one or more tangible, non-transitory computer-readable media. As used herein, the terms “software” and “firmware” are interchangeable and include any computer program stored in memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are merely exemplary and therefore do not limit the types of memory that can be used to store computer programs.

[0376] Software can take various forms, such as system software or application software. Furthermore, software can be a collection of independent programs, or a program module or part of a program module within a larger program. Software can also include modular programming in the form of object-oriented programming. After obtaining the detection data, the detection data can be processed automatically by the base detection system 2400 in response to user input, or in response to a request from another processing machine (e.g., a remote request via a communication link). In the illustrated specific implementation, the system controller 2404 includes an analysis module 2538 (in... Figure 25 (As shown in the diagram). In other implementations, the system controller 2404 does not include the analytics module 2538, but rather has access to the analytics module 2538 (e.g., the analytics module 2538 may be hosted separately in the cloud).

[0377] System controller 2404 can be connected via a communication link to biosensor 2402 and other components of base detection system 2400. System controller 2404 can also be communicatively connected to off-site systems or servers. The communication link can be hardwired, wired, or wireless. System controller 2404 can receive user input or commands from user interface 2414 and user input device 2415.

[0378] The fluid control system 2406 includes a fluid network and is configured to guide and regulate the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with a biosensor 2402 and a fluid storage system 2408. For example, selected fluid may be drawn from the fluid storage system 2408 and guided in a controlled manner to the biosensor 2402, or fluid may be drawn from the biosensor 2402 and directed toward, for example, a waste reservoir in the fluid storage system 2408. Although not shown, the fluid control system 2406 may include a flow sensor that detects the velocity or pressure of the fluid within the fluid network. The sensor may communicate with a system controller 2404.

[0379] Temperature control system 2410 is configured to regulate the temperature of fluids at different regions of the fluid network, fluid storage system 2408, and / or biosensor 2402. For example, temperature control system 2410 may include a thermal circulator that interfaces with biosensor 2402 and controls the temperature of fluid flowing along reaction sites in biosensor 2402. Temperature control system 2410 may also regulate the temperature of solid elements or components of base detection system 2400 or biosensor 2402. Although not shown, temperature control system 2410 may include sensors for detecting the temperature of fluids or other components. The sensors may communicate with system controller 2404.

[0380] The fluid storage system 2408 is in fluid communication with the biosensor 2402 and can store various reaction components or reactants for carrying out the desired reaction therein. The fluid storage system 2408 can also store fluids for washing or cleaning the fluid network and biosensor 2402, as well as for diluting reactants. For example, the fluid storage system 2408 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, and nonpolar solutions, etc. Furthermore, the fluid storage system 2408 may include a waste reservoir for receiving waste from the biosensor 2402. In embodiments including a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Therefore, one or more components relating to those systems described herein may be housed within a cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, and nonpolar solutions, waste, etc. Therefore, one or more of the fluid storage system, fluid control system, or temperature control system may be removably coupled to the bioassay system via the cartridge or other biosensor.

[0381] The illumination system 2409 may include a light source (e.g., one or more LEDs) and multiple optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. Optical components may be, for example, reflectors, dichroic mirrors, beam splitters, collimators, lenses, filters, wedge mirrors, prisms, mirrors, detectors, etc. In a specific embodiment using the illumination system, the illumination system 2409 may be configured to direct excitation light to the reaction site. As an example, the fluorophore may be excited by light of a green wavelength, thus the wavelength of the excitation light may be approximately 532 nm. In one embodiment, the illumination system 2409 is configured to produce illumination parallel to the surface normal of the surface of the biosensor 2402. In another embodiment, the illumination system 2409 is configured to produce illumination at an angle relative to the surface normal of the surface of the biosensor 2402. In yet another embodiment, the illumination system 2409 is configured to produce illumination with multiple angles, including some parallel illumination and some angled illumination.

[0382] System socket or interface 2412 is configured to engage biosensor 2402 in at least one of mechanical, electrical, and fluidic methods. System socket 2412 can hold biosensor 2402 in a desired orientation to facilitate fluid flow through biosensor 2402. System socket 2412 may also include electrical contacts configured to engage biosensor 2402, enabling base detection system 2400 to communicate with and / or power biosensor 2402. Furthermore, system socket 2412 may include a fluid port (e.g., a nozzle) configured to engage biosensor 2402. In some embodiments, biosensor 2402 is removably coupled to system socket 2412 in mechanical, electrical, and fluidic methods.

[0383] Furthermore, the base detection system 2400 can communicate remotely with other systems or networks or with other bioassay systems 2400. Detection data obtained by the bioassay system 2400 can be stored in a remote database.

[0384] Figure 25 It is possible Figure 24 The block diagram of the system controller 2404 used in the system is shown below. In one specific implementation, the system controller 2404 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include an algorithm (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or a sub-algorithm for performing a particular process. The system controller 2404 is conceptually exemplified as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 2404 may be implemented using an off-the-shelf PC with a single processor or multiple processors, wherein functional operations are distributed among the processors. As a further alternative, the modules described below may be implemented using a hybrid configuration, wherein some modular functions are executed using dedicated hardware, while other modular functions are executed using an off-the-shelf PC, etc. Modules may also be implemented as software modules within a processing unit.

[0385] During operation, communication port 2520 can communicate with biosensor 2402 ( Figure 24 ) and / or subsystems 2406, 2408, 2410 ( Figure 24 It can transmit information (e.g., commands) or receive information (e.g., data) from or from a user interface 2414. In a specific implementation, communication port 2520 can output a sequence of multiple pixel signals. Communication port 2520 can be accessed from user interface 2414. Figure 24The system receives user input and transmits data or information to the user interface 2414. Data from the biosensor 2402 or subsystems 2406, 2408, 2410 can be processed in real time by the system controller 2404 during a bioassay session. Alternatively, data can be temporarily stored in the system memory during a bioassay session and processed at a slower rate than in real-time or offline operation.

[0386] like Figure 25 As shown, the system controller 2404 may include multiple modules 2531-2539 that communicate with the main control module 2530. The main control module 2530 may communicate with the user interface 2414. Figure 24 Communication. Although modules 2531-2539 are shown communicating directly with the main control module 2530, modules 2531-2539 can also communicate directly with each other, and directly with the user interface 2414 and the biosensor 2402. Additionally, modules 2531-2539 can communicate with the main control module 2530 through other modules.

[0387] Multiple modules 2531-2539 include system modules 2531-2533, 2539 that communicate with subsystems 2406, 2408, 2410, and 2409, respectively. Fluid control module 2531 can communicate with fluid control system 2406 to control valves and flow sensors in the fluid network, thereby controlling the flow of one or more fluids through the fluid network. Fluid storage module 2532 can notify the user when the fluid volume is low or when the waste storage tank is at or near its capacity. Fluid storage module 2532 can also communicate with temperature control module 2533 to allow fluid to be stored at a desired temperature. Illumination module 2539 can communicate with illumination system 2409 to illuminate the reaction site at a specified time during the protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some embodiments, illumination module 2539 can communicate with illumination system 2409 to illuminate the reaction site at a specified angle.

[0388] Multiple modules 2531-2539 may also include a device module 2534 that communicates with the biosensor 2402 and an identification module 2535 that determines identification information associated with the biosensor 2402. The device module 2534 may, for example, communicate with the system socket 2412 to confirm that the biosensor has established electrical and fluid connections with the base detection system 2400. The identification module 2535 may receive signals identifying the biosensor 2402. The identification module 2535 may use the identity of the biosensor 2402 to provide additional information to the user. For example, the identification module 2535 may determine and subsequently display a batch number, manufacturing date, or suggested protocols for operation with the biosensor 2402.

[0389] Multiple modules 2531-2539 also include an analysis module 2538 (also referred to as a signal processing module or signal processor) for receiving and analyzing signal data (e.g., image data) from biosensor 2402. Analysis module 2538 includes memory (e.g., RAM or flash memory) for storing detection data. Detection data may include multiple pixel signal sequences, enabling the detection of pixel signal sequences from each of millions of sensors (or pixels) within numerous base detection cycles. Signal data may be stored for subsequent analysis or transmitted to user interface 2414 to display desired information to the user. In some embodiments, signal data may be processed by a solid-state imaging device (e.g., a CMOS image sensor) before being received by analysis module 2538.

[0390] Analysis module 2538 is configured to acquire image data from a photodetector at each sequencing cycle of multiple sequencing cycles. The image data originates from the emission signal detected by the photodetector and is processed by a neural network (e.g., a neural network-based template generator 2548, a neural network-based base detector 2558 (e.g., see...). Figure 7 , Figure 9 and Figure 10 (and / or a neural network-based quality scorer 2568) processes image data for each of the multiple sequencing cycles and generates base detections for at least some of the analytes at each of the multiple sequencing cycles.

[0391] Protocol modules 2536 and 2537 communicate with main control module 2530 to control the operation of subsystems 2406, 2408, and 2410 during a predetermined assay protocol. Protocol modules 2536 and 2537 may include a set of instructions for instructing base detection system 2400 to perform specific operations according to a predetermined protocol. As shown, the protocol module may be sequencing-by-synthesis (SBS) module 2536, configured to issue various commands for performing the sequencing-by-synthesis process. In SBS, the extension of nucleic acid primers along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase) or ligation (e.g., catalyzed by a ligase). In a specific polymerase-based SBS implementation, fluorescently labeled nucleotides are added to primers in a template-dependent manner (causing primer extension), such that detection of the sequence and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, to initiate a first SBS cycle, a command can be issued to deliver one or more labeled nucleotides, DNA polymerase, etc., to / through a flow cell containing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. Those reaction sites where primer extension results in the incorporation of labeled nucleotides can be detected by an imaging event. During the imaging event, illumination system 2409 can provide excitation light to the reaction sites. Optionally, the nucleotides may also include a reversible termination property that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible termination moiety can be added to the primer, such that subsequent extension does not occur until a deblocking agent is delivered to remove that moiety. Thus, in a specific implementation using reversible termination, a command can be issued to deliver a deblocking agent to the flow cell (before or after detection). One or more commands can be issued to perform washing between the various delivery steps. This cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing techniques are described, for example, in Bentley et al., Nature, Vol. 456: pp. 53–59 (2008), WO 04 / 018497, US 7,057,026, WO 91 / 06678, WO 07 / 123744, US 7,329,492, US 7,211,414, US 7,315,019 and US 7,405,281, each of which is incorporated herein by reference.

[0392] For the nucleotide delivery step in the SBS cycle, a single type of nucleotide can be delivered at once, or multiple different nucleotide types (e.g., A, C, T, and G together) can be delivered. For nucleotide delivery configurations where only a single type of nucleotide is present at a time, the different nucleotides do not need to have different labels, as they can be distinguished based on the inherent time intervals in individualized delivery. Therefore, sequencing methods or apparatus can use monochromatic detection. For example, the excitation source only needs to provide excitation at a single wavelength or within a single wavelength range. For nucleotide delivery configurations where delivery results in the simultaneous presence of multiple different nucleotides in the flow cell, the sites incorporating different nucleotide types can be distinguished based on different fluorescent labels attached to the corresponding nucleotide types in the mixture. For example, four different nucleotides, each with one of four different fluorophores, can be used. In one specific implementation, excitation in four different regions of the spectrum can be used to distinguish the four different fluorophores. For example, four different excitation radiation sources can be used. Alternatively, fewer than four different excitation sources can be used, but optical filtering of excitation radiation from a single source can be used to generate different ranges of excitation radiation at the flow cell.

[0393] In some specific implementations, fewer than four different colors can be detected in a mixture containing four different nucleotides. For example, nucleotide pairs can be detected at the same wavelength, but distinguished based on the intensity difference of one member relative to the other member, or based on a change in the presence or absence of a signal in one member that results in a signal that is significantly different from the signal of the other member of the pair (e.g., by chemical modification, photochemical modification, or physical modification). Exemplary apparatuses and methods for distinguishing four different nucleotides using detection with fewer than four colors are described, for example, in U.S. Patent Application Serial Nos. 61 / 538,294 and 61 / 619,878, the entire contents of which are incorporated herein by reference. U.S. Application 13 / 624,200, filed September 21, 2012, is also incorporated herein by reference in its entirety.

[0394] Multiple protocol modules may also include a sample preparation (or generation) module 2537, configured to command a fluid control system 2406 and a temperature control system 2410 to amplify the product within the biosensor 2402. For example, the biosensor 2402 may be coupled to a base detection system 2400. The amplification module 2537 may instruct the fluid control system 2406 to deliver necessary amplification components to the reaction chamber within the biosensor 2402. In other embodiments, the reaction site may already contain components for amplification, such as template DNA and / or primers. After the amplification components are delivered to the reaction chamber, the amplification module 2537 may instruct the temperature control system 2410 to cycle through different temperature phases according to a known amplification protocol. In some embodiments, amplification and / or nucleotide incorporation are performed isothermally.

[0395] The SBS module 2536 can issue commands to perform bridged PCR, in which clusters of cloned amplicones are formed on localized regions within the channels of the flow cell. After amplicon generation via bridged PCR, the amplicones can be "linearized" to prepare single-stranded template DNA or sstDNA, and sequencing primers can be hybridized to universal sequences side-linked to regions of interest. For example, sequencing-by-synthesis methods based on reversible terminators can be used as described above or as follows.

[0396] Each base detection or sequencing cycle can be performed by extending sstDNA with a single base, which can be accomplished, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. The different types of nucleotides can have unique fluorescent labels, and each nucleotide can also have a reversible terminator that allows only single-base incorporation to occur in each cycle. After adding a single base to the sstDNA, excitation light can be incident on the reaction site and fluorescence emission can be detected. After detection, the fluorescent label and terminator can be chemically cleaved from the sstDNA. This can then be followed by another similar base detection or sequencing cycle. In such sequencing protocols, the SBS module 2536 instructs the fluid control system 2406 to guide the flow of reagents and enzyme solutions through the biosensor 2402. Exemplary SBS methods based on reversible terminators that can be used with the apparatus and methods described herein are described in U.S. Patent Application Publication 2007 / 0166705 A1, U.S. Patent Application Publication 2006 / 0188901 A1, U.S. Patent 7,057,026, U.S. Patent Application Publication 2006 / 0240439 A1, U.S. Patent Application Publication 2006 / 02814714709 A1, PCT Publication WO05 / 065814, and PCT Publication WO 06 / 064199, each of which is incorporated herein by reference in its entirety. Exemplary reagents based on reversible terminator SBS are described in U.S. 7,541,444; U.S. 7,057,026; U.S. 7,427,673; U.S. 7,566,537; and U.S. 7,592,435, each of which is incorporated herein by reference in its entirety.

[0397] In some implementations, the amplification module and the SBS module can operate in a single assay protocol, where, for example, template nucleic acids are amplified and subsequently sequenced within the same kit.

[0398] The base detection system 2400 can also allow the user to reconfigure the assay protocol. For example, the base detection system 2400 can provide the user with options for modifying the determined protocol via the user interface 2414. For example, if it is determined that the biosensor 2402 will be used for amplification, the base detection system 2400 can request the temperature of the annealing cycle. Furthermore, if the user has provided user input that is generally unacceptable for the selected assay protocol, the base detection system 2400 can issue a warning to the user.

[0399] In a specific implementation, the biosensor 2402 comprises millions of sensors (or pixels), each sensor (or pixel) generating multiple pixel signal sequences in a subsequent base detection cycle. The analysis module 2538 detects the multiple pixel signal sequences based on the row-by-row and / or column-by-column positions of the sensors on the sensor array and assigns them to the corresponding sensors (or pixels).

[0400] Each sensor in the sensor array can generate sensor data for a block of the flow cell, where the block is located in a region on the flow cell where clusters of genetic material are set during the base detection operation. The sensor data may include image data from a pixel array. For a given cycle, the sensor data may include more than one image, thus producing multi-feature per pixel as block data.

[0401] Figure 26 This is a simplified block diagram of a computer system 2600 that can be used to implement the disclosed techniques. The computer system 2600 includes at least one central processing unit (CPU) 2672 that communicates with a plurality of peripheral devices via a bus subsystem 2655. These peripheral devices may include a storage subsystem 2610 (including, for example, memory devices and a file storage subsystem 2636), a user interface input device 2638, a user interface output device 2676, and a network interface subsystem 2674. The input and output devices allow users to interact with the computer system 2600. The network interface subsystem 2674 provides an interface to an external network, including interfaces to corresponding interface devices in other computer systems.

[0402] User interface input device 2638 may include: a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen integrated into a display; audio input devices such as a speech recognition system and a microphone; and other types of input devices. Generally, the term "input device" is intended to encompass all possible types of devices and methods for inputting information into computer system 2600.

[0403] User interface output device 2676 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include an LED display, a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other mechanisms for producing visible images. The display subsystem may also provide non-visual displays, such as audio output devices. Generally, the term "output device" is intended to encompass all possible types of devices and methods for outputting information from computer system 2600 to a user or to another machine or computer system.

[0404] The storage subsystem 2610 stores the programming and data structures that provide the functionality of some or all of the modules and methods described herein. These software modules are typically executed by the deep learning processor 2678.

[0405] In one specific implementation, the neural network is implemented using a deep learning processor 2678. These deep learning processors can be configurable and reconfigurable processors, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and / or coarse-grained reconfigurable architectures (CGRAs) and graphics processing units (GPUs), or other configured devices. The deep learning processor 2678 can be powered by deep learning cloud platforms such as Google Cloud Platform. TM Xilinx TM and Cirrascale TM Hosted. Examples of deep learning processors include Google's Tensor Processing Unit (TPU). TM Rackmount solutions (such as GX4 Rackmount Series) TM GX149 Rackmount Series TM NVIDIA DGX-1 TM Microsoft's Stratix V FPGA TM , Graphcore’s Intelligent Processor Unit (IPU) TM Qualcomm's Snapdragon processors TM Zeroth Platform TM NVIDIA's Volta TM NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2 MODULE TM Intel's Nirvana TMMovidius VPU TM ,Fujitsu DPI TM ARM's DynamicIQ TM IBM TrueNorth TM wait.

[0406] The memory subsystem 2622 used in the storage subsystem 2610 may include multiple memories, including a main random access memory (RAM) 2634 for storing instructions and data during program execution and a read-only memory (ROM) 2632 for storing fixed instructions. The file storage subsystem 2636 can provide persistent storage for program files and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media enclosures. Modules implementing certain specific functionalities may be stored by the file storage subsystem 2636 within the storage subsystem 2610 or on other machines accessible to the processor.

[0407] Bus subsystem 2655 provides mechanisms for enabling various components and subsystems of computer system 2600 to communicate with each other as intended. Although bus subsystem 2655 is schematically shown as a single bus, alternative implementations of this bus subsystem may use multiple buses.

[0408] The computer system 2600 itself can be of different types, including personal computers, portable computers, workstations, computer terminals, network computers, televisions, mainframes, server clusters, a loosely networked group of widely distributed computers, or any other data processing system or user equipment. Due to the constantly evolving nature of computers and networks, [the following applies]. Figure 26 The description of the computer system 2600 depicted herein is intended only as a specific example to illustrate a preferred embodiment of the invention. Many other configurations of the computer system 2600 are possible, which have... Figure 26 The computer system depicted in the text has more or fewer components.

[0409] This invention discloses the following provisions:

[0410] Terms and Conditions

[0411] Item Set #1 (Self-learning base detector trained using oligonucleotide sequences)

[0412] 1. A computer-implemented method for progressively training a base detector, the method comprising:

[0413] The base detector is initially trained iteratively using analytes containing single oligonucleotide base sequences, and labeled training data is generated using the initially trained base detector.

[0414] (i) further training the base detector using analytes containing multiple oligonucleotide base sequences, and generating labeled training data using the further trained base detector; and

[0415] The base detector is further trained iteratively by repeating step (i), while increasing the complexity of the neural network configuration loaded within the base detector during at least one iteration, wherein labeled training data generated during the iteration is used to train the base detector in the immediately following subsequent iteration.

[0416] 1a. The method according to Clause 1, wherein the method further comprises:

[0417] During at least one iteration of further training the base detector using the analyte containing multiple oligonucleotide base sequences, the number of unique oligonucleotide base sequences within the analyte is increased.

[0418] 2. The method according to Clause 1, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence comprises:

[0419] During the first iteration of the initial training of the base detector:

[0420] The known single oligonucleotide base sequence is filled into multiple clusters of the flow cell;

[0421] Generate multiple sequence signals corresponding to the plurality of clusters, each of the plurality of sequence signals representing a base sequence loaded in the corresponding cluster among the plurality of clusters;

[0422] Based on each of the plurality of sequence signals, the corresponding base detection of the known single oligonucleotide base sequence is predicted, thereby generating a plurality of predicted base detections;

[0423] For each of the plurality of sequence signals, based on (i) the corresponding predicted base detection and (ii) the comparison of the bases in the known single oligonucleotide sequence, a corresponding error signal is generated, thereby generating a plurality of error signals corresponding to the plurality of sequence signals; and

[0424] The base detector is initially trained during the first iteration based on the multiple error signals.

[0425] 2a. The method according to Clause 2, wherein initial training of the base detector during the first iteration comprises:

[0426] Based on the multiple error signals, the weights and / or biases of the neural network configuration are updated using the backpropagation path of the neural network configuration loaded in the base detector.

[0427] 3. The method according to Clause 2, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence further comprises:

[0428] During the second iteration of the initial training of the base detector, which occurs after the first iteration of the initial training:

[0429] Using the base detector that has been partially trained during the first iteration of the initial training, additional base detections corresponding to the known single oligonucleotide base sequence are predicted based on each of the plurality of sequence signals, thereby generating a plurality of additional predicted base detections.

[0430] For each of the plurality of sequence signals, based on (i) a corresponding additional predicted base detection and (ii) a comparison of the bases of the known single oligonucleotide sequence, a corresponding additional error signal is generated, thereby generating a plurality of additional error signals corresponding to the plurality of sequence signals; and

[0431] Based on the additional error signals, the base detector is further initially trained during the second iteration.

[0432] 4. The method according to Clause 3, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence comprises:

[0433] For multiple instances, the second iteration of the initial training of the base detector is repeated with an analyte containing the single oligonucleotide base sequence until the convergence condition is met.

[0434] 5. The method according to Clause 4, wherein the convergence condition is satisfied when the reduction of the plurality of additional error signals is less than a threshold between two consecutive repetitions of the second iteration of the initial training of the base detector.

[0435] 6. The method according to Clause 4, wherein the convergence condition is satisfied when the second iteration of the initial training of the base detector is repeated for at least a threshold number of instances.

[0436] 7. The method according to Clause 3, wherein:

[0437] The plurality of sequence signals corresponding to the plurality of clusters generated during the first iteration of the initial training of the base detector are repeatedly used in the second iteration of the initial training of the base detector.

[0438] 8. The method according to Clause 2, wherein comparing (i) the corresponding predicted base detection with (ii) the base of the known single oligonucleotide sequence comprises:

[0439] For a first predicted base detection, (i) the first base of the first predicted base detection is compared with the first base of the known single oligonucleotide sequence, and (ii) the second base of the first predicted base detection is compared with the second base of the known single oligonucleotide sequence, thereby generating a corresponding first error signal.

[0440] 9. The method according to Clause 1, wherein iteratively further training of the base detector comprises:

[0441] The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations; and

[0442] The base detector was further trained using an analyte containing three known unique oligonucleotide base sequences for N² iterations.

[0443] The N1 iterations are performed before the N2 iterations.

[0444] 10. The method according to Clause 1, wherein during the iterative initial training of the base detector with the analyte comprising the single oligonucleotide base sequence, a first neural network configuration is loaded within the base detector, and wherein iteratively further training of the base detector comprises:

[0445] The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations, resulting in...

[0446] (i) For the first subset of the N1 iterations, a second neural network configuration is loaded into the base detector, and

[0447] (ii) For the second subset of the N1 iterations that occurs after the first subset of the N1 iterations, a third neural network configuration is loaded in the base detector, wherein the first neural network configuration, the second neural network configuration, and the third neural network configuration are different from each other.

[0448] 11. The method according to Clause 10, wherein the second neural network configuration is more complex than the first neural network configuration, and wherein the third neural network configuration is more complex than the second neural network configuration.

[0449] 12. The method according to Clause 10, wherein the second neural network configuration has a greater number of layers than the first neural network configuration.

[0450] 13. The method according to Clause 10, wherein the second neural network configuration has a larger number of weights than the first neural network configuration.

[0451] 14. The method according to Clause 10, wherein the second neural network configuration has a larger number of parameters than the first neural network configuration.

[0452] 15. The method according to Clause 10, wherein the third neural network configuration has a greater number of layers than the second neural network configuration.

[0453] 16. The method according to Clause 10, wherein the third neural network configuration has a larger number of weights than the second neural network configuration.

[0454] 17. The method according to Clause 10, wherein the third neural network configuration has a larger number of parameters than the second neural network configuration.

[0455] 18. The method according to Clause 10, wherein for one of the N1 iterations, the analyte comprising two known unique oligonucleotide base sequences is used for…

[0456] Training the base detector in one step for the N1 iterations includes:

[0457] (i) filling a first plurality of clusters in a flow cell with a first known oligonucleotide sequence of the two known unique oligonucleotide sequences, and (ii) filling a second plurality of clusters in a flow cell with a second known oligonucleotide sequence of the two known unique oligonucleotide sequences;

[0458] For each of the first plurality of clusters and the second plurality of clusters, predict the corresponding base detection, thereby generating multiple predicted base detections;

[0459] Map the first predicted base detection among the plurality of predicted base detections in (i) to the first known oligonucleotide base sequence, and map the second predicted base detection among the plurality of predicted base detections in (ii) to the second known oligonucleotide base sequence, while avoiding mapping the third predicted base detection among the plurality of predicted base detections to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence;

[0460] Generate (i) a first error signal based on the comparison of the first predicted base detection with the first known oligonucleotide base sequence, and (ii) a second error signal based on the comparison of the second predicted base detection with the second known oligonucleotide base sequence; and

[0461] The base detector is further trained based on the first error signal and the second error signal.

[0462] 19. The method according to Clause 18, wherein mapping the first predicted base detection to the first known oligonucleotide sequence in the two known unique oligonucleotide sequences comprises:

[0463] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0464] The first predicted base detection is determined to have at least a threshold number of base similarities to the first known oligonucleotide sequence, and to have less than the threshold number of base similarities to the second known oligonucleotide sequence; and

[0465] Based on the determination that the first predicted base detection has at least the threshold number of base similarities to the first known oligonucleotide base sequence, the first predicted base detection is mapped to the first known oligonucleotide base sequence.

[0466] 20. The method according to Clause 18, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises:

[0467] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0468] The first predicted base detection is determined to have less than a threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; and

[0469] Based on the determination that the first predicted base detection has less than the threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

[0470] 21. The method according to Clause 18, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises:

[0471] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0472] The first predicted base detection is determined to have a base similarity greater than a threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; and

[0473] Based on the determination that the first predicted base detection has a base similarity greater than the threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

[0474] 22. The method according to Clause 18, wherein generating labeled training data by performing one of the N1 iterations using the further trained base detector comprises:

[0475] After further training the base detector during one of the N1 iterations, the corresponding base detection is re-predicted for each of the first plurality of clusters and the second plurality of clusters, thereby generating another plurality of predicted base detections.

[0476] Remap (i) the first subset of the other multiple predicted base detections to the first known oligonucleotide sequence, and (ii) remap the second subset of the other multiple predicted base detections to the second known oligonucleotide sequence, while avoiding mapping the third subset of the other multiple predicted base detections to either the first known oligonucleotide sequence or the second known oligonucleotide sequence; and

[0477] The remapping is used to generate labeled training data, such that the labeled training data includes (i) the first subset of the other plurality of predicted base detections, wherein the first known oligonucleotide base sequence forms the benchmark ground truth data of the first subset of the other plurality of predicted base detections, and (ii) the second subset of the other plurality of predicted base detections, wherein the second known oligonucleotide base sequence forms the benchmark ground truth data of the second subset of the other plurality of predicted base detections.

[0478] 23. The method described according to Clause 22, wherein:

[0479] The labeled training data generated during the first iteration of the N1th iteration is used to train the base detector during the subsequent iterations immediately following the N1th iteration.

[0480] 24. The method described according to Clause 23, wherein:

[0481] The neural network configuration of the base detector is the same during the first iteration of the N1th iteration and the immediately following iteration of the N1th iteration.

[0482] 25. The method described according to Clause 23, wherein:

[0483] The neural network configuration of the base detector during the immediately following iteration of the N1th iteration is different from that of the base detector during the first iteration of the N1th iteration, and is more complex than the neural network configuration of the base detector during the first iteration of the N1th iteration.

[0484] 26. The method according to Clause 1, wherein iteratively further training of the base detector comprises:

[0485] During the iterative training, as the iterations proceed, the number of unique oligonucleotide base sequences in the analyte containing the polynucleotide base sequence is monotonically increased.

[0486] 27. A computer-implemented method, the method comprising:

[0487] Base detectors are used to predict the base detection sequence of an unknown analyte that has been sequenced as a known sequence of oligonucleotides.

[0488] Each unknown analyte in the unknown analytes is labeled with a benchmark truth sequence that matches the known sequence; and

[0489] The base detector is trained using the labeled unknown analyte.

[0490] 28. The computer-implemented method according to Clause 27, the method further comprising iterating the use, the labeling, and the training until convergence is satisfied.

[0491] 29. A computer-implemented method, the method comprising:

[0492] Base detectors are used to predict the base detection sequence of an unknown analyte population that has been sequenced as two or more known sequences of two or more oligonucleotides.

[0493] Based on classifying the base detection sequence of the selected unknown analytes into the known sequences, unknown analytes are selected from the group of unknown analytes;

[0494] Based on the classification, the selected subsets of unknown analytes are labeled with corresponding benchmark truth sequences that match the known sequences; and

[0495] The base detector is trained using a labeled subset of the selected unknown analytes.

[0496] 30. The computer-implemented method according to Clause 29, the method further comprising iterating the use, the selection, the labeling, and the training until convergence is satisfied.

[0497] 31. A non-transitory computer-readable storage medium printed with computer program instructions for progressively training a base detector, said instructions, when executed on a processor, implementing a method comprising the following:

[0498] The base detector is initially trained iteratively using analytes containing single oligonucleotide base sequences, and labeled training data is generated using the initially trained base detector.

[0499] (i) further training the base detector using analytes containing multiple oligonucleotide base sequences, and generating labeled training data using the further trained base detector; and

[0500] The base detector is further trained iteratively by repeating step (i), while increasing the complexity of the neural network configuration loaded within the base detector during at least one iteration, wherein labeled training data generated during the iteration is used to train the base detector in the immediately following subsequent iteration.

[0501] 31a. The computer-readable storage medium according to clause 31, wherein the instructions further include the method of:

[0502] During at least one iteration of further training the base detector using the analyte containing multiple oligonucleotide base sequences, the number of unique oligonucleotide base sequences within the analyte is increased.

[0503] 32. The computer-readable storage medium method according to claim 31, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence comprises:

[0504] During the first iteration of the initial training of the base detector:

[0505] The known single oligonucleotide base sequence is filled into multiple clusters of the flow cell;

[0506] Generate multiple sequence signals corresponding to the plurality of clusters, each of the plurality of sequence signals representing a base sequence loaded in the corresponding cluster among the plurality of clusters;

[0507] Based on each of the plurality of sequence signals, the corresponding base detection of the known single oligonucleotide base sequence is predicted, thereby generating a plurality of predicted base detections;

[0508] For each of the plurality of sequence signals, based on (i) the corresponding predicted base detection and (ii) the comparison of the bases in the known single oligonucleotide sequence, a corresponding error signal is generated, thereby generating a plurality of error signals corresponding to the plurality of sequence signals; and

[0509] The base detector is initially trained during the first iteration based on the multiple error signals.

[0510] 32a. The computer-readable storage medium according to clause 32, wherein initial training of the base detector during the first iteration comprises:

[0511] Based on the multiple error signals, the weights and / or biases of the neural network configuration are updated using the backpropagation path of the neural network configuration loaded in the base detector.

[0512] 33. The computer-readable storage medium according to claim 32, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence further comprises:

[0513] During the second iteration of the initial training of the base detector, which occurs after the first iteration of the initial training:

[0514] Using the base detector that has been partially trained during the first iteration of the initial training, additional base detections corresponding to the known single oligonucleotide base sequence are predicted based on each of the plurality of sequence signals, thereby generating a plurality of additional predicted base detections.

[0515] For each of the plurality of sequence signals, based on (i) a corresponding additional predicted base detection and (ii) a comparison of the bases of the known single oligonucleotide sequence, a corresponding additional error signal is generated, thereby generating a plurality of additional error signals corresponding to the plurality of sequence signals; and

[0516] Based on the additional error signals, the base detector is further initially trained during the second iteration.

[0517] 34. The computer-readable storage medium according to clause 33, wherein iteratively initial training of the base detector with the analyte comprising the single oligonucleotide base sequence further comprises:

[0518] For multiple instances, the second iteration of the initial training of the base detector is repeated with an analyte containing the single oligonucleotide base sequence until the convergence condition is met.

[0519] 35. The computer-readable storage medium according to Clause 34, wherein the convergence condition is satisfied when the reduction of the plurality of additional error signals is less than a threshold between two consecutive repetitions of the second iteration of the initial training of the base detector.

[0520] 36. The computer-readable storage medium according to Clause 34, wherein the convergence condition is satisfied when the second iteration of the initial training of the base detector is repeated for at least a threshold number of instances.

[0521] 37. The computer-readable storage medium as described in Clause 33, wherein:

[0522] The plurality of sequence signals corresponding to the plurality of clusters generated during the first iteration of the initial training of the base detector are repeatedly used in the second iteration of the initial training of the base detector.

[0523] 38. The computer-readable storage medium according to clause 32, wherein comparing (i) the corresponding predicted base detection with (ii) the base of the known single oligonucleotide sequence comprises:

[0524] For a first predicted base detection, (i) the first base of the first predicted base detection is compared with the first base of the known single oligonucleotide sequence, and (ii) the second base of the first predicted base detection is compared with the second base of the known single oligonucleotide sequence, thereby generating a corresponding first error signal.

[0525] 39. The computer-readable storage medium according to clause 31, wherein iteratively further training of the base detector comprises:

[0526] The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations; and

[0527] The base detector was further trained using an analyte containing three known unique oligonucleotide base sequences for N² iterations.

[0528] The N1 iterations are performed before the N2 iterations.

[0529] 40. The computer-readable storage medium according to claim 31, wherein during the iterative initial training of the base detector with the analyte comprising the single oligonucleotide base sequence, a first neural network configuration is loaded within the base detector, and wherein iteratively further training of the base detector comprises:

[0530] The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations, resulting in...

[0531] (i) For the first subset of the N1 iterations, a second neural network configuration is loaded into the base detector, and

[0532] (ii) For the second subset of the N1 iterations that occurs after the first subset of the N1 iterations, a third neural network configuration is loaded in the base detector, wherein the first neural network configuration, the second neural network configuration, and the third neural network configuration are different from each other.

[0533] 41. The computer-readable storage medium according to clause 40, wherein the second neural network configuration is more complex than the first neural network configuration, and wherein the third neural network configuration is more complex than the second neural network configuration.

[0534] 42. The computer-readable storage medium according to clause 40, wherein the second neural network configuration has a greater number of layers than the first neural network configuration.

[0535] 43. The computer-readable storage medium according to Clause 40, wherein the second neural network configuration has a greater number of weights than the first neural network configuration.

[0536] 44. The computer-readable storage medium according to Clause 40, wherein the second neural network configuration has a larger number of parameters than the first neural network configuration.

[0537] 45. The computer-readable storage medium according to clause 40, wherein the third neural network configuration has a greater number of layers than the second neural network configuration.

[0538] 46. ​​The computer-readable storage medium according to Clause 40, wherein the third neural network configuration has a greater number of weights than the second neural network configuration.

[0539] 47. The computer-readable storage medium according to Clause 40, wherein the third neural network configuration has a larger number of parameters than the second neural network configuration.

[0540] 48. The computer-readable storage medium according to clause 40, wherein further training of the base detector for the N1 iterations using the analyte comprising two known unique oligonucleotide base sequences comprises: for one of the N1 iterations,

[0541] (i) filling a first plurality of clusters in a flow cell with a first known oligonucleotide sequence of the two known unique oligonucleotide sequences, and (ii) filling a second plurality of clusters in a flow cell with a second known oligonucleotide sequence of the two known unique oligonucleotide sequences;

[0542] For each of the first plurality of clusters and the second plurality of clusters, predict the corresponding base detection, thereby generating multiple predicted base detections;

[0543] Map the first predicted base detection among the plurality of predicted base detections in (i) to the first known oligonucleotide base sequence, and map the second predicted base detection among the plurality of predicted base detections in (ii) to the second known oligonucleotide base sequence, while avoiding mapping the third predicted base detection among the plurality of predicted base detections to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence;

[0544] Generate (i) a first error signal based on the comparison of the first predicted base detection with the first known oligonucleotide base sequence, and (ii) a second error signal based on the comparison of the second predicted base detection with the second known oligonucleotide base sequence; and

[0545] The base detector is further trained based on the first error signal and the second error signal.

[0546] 49. The computer-readable storage medium of claim 38, wherein the first known oligonucleotide sequence that maps the first predicted base detection to the two known unique oligonucleotide sequences comprises:

[0547] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0548] The first predicted base detection is determined to have at least a threshold number of base similarities to the first known oligonucleotide sequence, and to have less than the threshold number of base similarities to the second known oligonucleotide sequence; and

[0549] Based on the determination that the first predicted base detection has at least the threshold number of base similarities to the first known oligonucleotide base sequence, the first predicted base detection is mapped to the first known oligonucleotide base sequence.

[0550] 50. The computer-readable storage medium according to clause 48, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises:

[0551] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0552] The first predicted base detection is determined to have less than a threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; and

[0553] Based on the determination that the first predicted base detection has less than the threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

[0554] 51. The computer-readable storage medium according to Clause 48, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises:

[0555] Each base detected by the first predicted base is compared with the corresponding base of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence;

[0556] The first predicted base detection is determined to have a base similarity greater than a threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; and

[0557] Based on the determination that the first predicted base detection has a base similarity greater than the threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

[0558] 52. The computer-readable storage medium according to Clause 48, wherein generating labeled training data by performing the one of the N1 iterations using the further trained base detector comprises:

[0559] After further training the base detector during one of the N1 iterations, the corresponding base detection is re-predicted for each of the first plurality of clusters and the second plurality of clusters, thereby generating another plurality of predicted base detections.

[0560] Remap (i) the first subset of the other multiple predicted base detections to the first known oligonucleotide sequence, and (ii) remap the second subset of the other multiple predicted base detections to the second known oligonucleotide sequence, while avoiding mapping the third subset of the other multiple predicted base detections to either the first known oligonucleotide sequence or the second known oligonucleotide sequence; and

[0561] The remapping is used to generate labeled training data, such that the labeled training data includes (i) the first subset of the other plurality of predicted base detections, wherein the first known oligonucleotide base sequence forms the benchmark ground truth data of the first subset of the other plurality of predicted base detections, and (ii) the second subset of the other plurality of predicted base detections, wherein the second known oligonucleotide base sequence forms the benchmark ground truth data of the second subset of the other plurality of predicted base detections.

[0562] 53. The computer-readable storage medium as described in Clause 52, wherein:

[0563] The labeled training data generated during the first iteration of the N1th iteration is used to train the base detector during the subsequent iterations immediately following the N1th iteration.

[0564] 54. The computer-readable storage medium as described in Clause 53, wherein:

[0565] The neural network configuration of the base detector is the same during the first iteration of the N1th iteration and the immediately following iteration of the N1th iteration.

[0566] 55. The computer-readable storage medium as described in Clause 53, wherein:

[0567] The neural network configuration of the base detector during the immediately following iteration of the N1th iteration is different from that of the base detector during the first iteration of the N1th iteration, and is more complex than the neural network configuration of the base detector during the first iteration of the N1th iteration.

[0568] 56. The computer-readable storage medium according to clause 31, wherein iteratively further training of the base detector comprises:

[0569] During the iterative training, as the iterations proceed, the number of unique oligonucleotide base sequences in the analyte containing the polynucleotide base sequence is monotonically increased.

[0570] Item set #2 (Self-learning base detector trained using biological sequences)

[0571] A1. A computer-implemented method for progressively training a base detector, the method comprising:

[0572] The base detector is initially trained, and labeled training data is generated using the initially trained base detector;

[0573] (i) Further training the base detector using analytes containing the base sequences of a biological organism, and generating labeled training data using the further trained base detector; and

[0574] The base detector is further trained iteratively by repeating step (i) N times, including:

[0575] The base detector is further trained using an analyte containing a first biological base sequence selected from a first plurality of base subsequences for N1 iterations in the N iterations, and

[0576] The base detector is further trained using an analyte containing a second biological base sequence selected from a second plurality of base subsequences for N2 iterations out of the N iterations.

[0577] The complexity of the neural network configuration loaded in the base detector increases monotonically with the N iterations, and

[0578] The labeled training data generated during the N iterations is used to train the base detector during the subsequent iterations immediately following the N iterations.

[0579] A1a. The method according to clause A1, wherein initial training of the base detector comprises:

[0580] The base detector is initially trained using an analyte containing one or more oligonucleotide base sequences, and labeled training data is generated using the initially trained base detector.

[0581] A2. The method according to clause A1, wherein the N1 iterations are performed before the N2 iterations, and wherein the second organism base sequence has a higher number of bases than the first organism base sequence.

[0582] A3. The method according to clause A1, wherein further training the base detector for the N1 iterations comprises, during one iteration of the N1 iterations:

[0583] (i) filling a first cluster of a plurality of clusters in a flow cell with a first base sequence of the first plurality of base sequences of the first organism, (ii) filling a second cluster of a plurality of clusters in a flow cell with a second base sequence of the first plurality of base sequences of the first organism, and (iii) filling a third cluster of a plurality of clusters in a flow cell with a third base sequence of the first plurality of base sequences of the first organism.

[0584] Receive (i) a first sequence signal from the first cluster indicating the base sequence filled in the first cluster, (ii) a second sequence signal from the second cluster indicating the base sequence filled in the second cluster, and (iii) a third sequence signal from the third cluster indicating the base sequence filled in the third cluster;

[0585] Generate (i) a first predicted base sequence based on the first sequence signal, (ii) a second predicted base sequence based on the second sequence signal, and (iii) a third predicted base sequence based on the third sequence signal;

[0586] Mapping (i) the first predicted base sequence to a first portion of the first organism's base sequence, and (ii) the second predicted base sequence to a second portion of the first organism's base sequence, while failing to map the third predicted base sequence to any portion of the first organism's base sequence; and

[0587] Generate labeled training data, the labeled training data comprising (i) a first predicted subsequence mapped to a first portion of a first organism's base sequence, wherein the first portion of the first organism's base sequence is a benchmark true value of the first predicted subsequence, and (ii) a second predicted subsequence mapped to a second portion of a first organism's base sequence, wherein the second portion of the first organism's base sequence is a benchmark true value of the second predicted subsequence.

[0588] A3a. The method according to clause A3, wherein further training the base detector for the N1 iterations comprises, during one iteration of the N1 iterations:

[0589] Before generating the first predicted base sequence, the second predicted base sequence, and the third predicted base sequence, the base detector is trained using labeled training data generated during the initial training of the base detector.

[0590] A4. The method described according to clause A3, wherein:

[0591] The first predicted base sequence has an L1 number of bases; and

[0592] One or more bases in the L1 bases of the first predicted base subsequence do not match the corresponding bases of the first part of the first organism base sequence, due to an error in the base detection prediction of the base detector.

[0593] A5. The method according to clause A3, wherein the first predicted subsequence has an L1 number of bases, wherein the L1 number of bases in the first predicted subsequence comprises an initial L2 bases, followed by a subsequent L3 bases, and wherein mapping the first predicted subsequence to the first portion of the first organism's base sequence comprises:

[0594] The initial L2 bases of the first predicted base sequence are substantially and uniquely matched with the consecutive L2 bases of the first organism base sequence;

[0595] Identify the first portion of the base sequence of the first organism such that the first portion (i) includes the consecutive L2 bases as initial bases and (ii) includes L1 bases; and

[0596] Map the first predicted base sequence to the first identified portion of the first organism's base sequence.

[0597] A6. The method according to A5 further includes:

[0598] When the initial L2 bases of the first predicted base sequence are substantially and uniquely matched, avoid matching the subsequent L3 bases of the first predicted base sequence with any bases of the first organism's base sequence.

[0599] A7. The method according to A5, wherein the initial L2 bases of the first predicted base sequence substantially match the consecutive L2 bases of the first organism base sequence, such that at least a threshold number of bases of the initial L2 bases of the first predicted base sequence match the consecutive L2 bases of the first organism base sequence.

[0600] A8. The method according to A5, wherein the initial L2 bases of the first predicted base sequence uniquely match the consecutive L2 bases of the first organism base sequence, such that the initial L2 bases of the first predicted base sequence substantially match only the consecutive L2 bases of the first organism base sequence, and do not match any other consecutive L2 bases of the first organism base sequence.

[0601] A9. The method according to clause A3, wherein the third predicted base sequence has an L1 number of bases, and wherein failure to map the third predicted base sequence to any of the first plurality of base sequences comprises:

[0602] (i) The initial L2 bases of the L1 bases of the third predicted base sequence failed to match substantially and uniquely with the consecutive L2 bases of the first organism base sequence.

[0603] A10. The method according to clause A3, wherein the first iteration in the N1 iterations is the first iteration in the N1 iterations, and wherein further training the base detector to perform the second iteration in the N1 iterations comprises:

[0604] The base detector is trained using the labeled training data generated during the first iteration of the N1 iterations;

[0605] Using the base detector trained with the labeled training data generated during the first iteration of the N1 iterations, (i) a further first predicted base subsequence based on the first sequence signal, (ii) a further second predicted base subsequence based on the second sequence signal, and (iii) a further third predicted base subsequence based on the third sequence signal.

[0606] Mapping (i) the additional first predicted base sequence to the first portion of the first organism's base sequence, mapping (ii) the additional second predicted base sequence to the second portion of the first organism's base sequence, and mapping (iii) the additional third predicted base sequence to the third portion of the first organism's base sequence; and

[0607] Generate additional labeled training data, the additional labeled training data comprising (i) the additional first predicted base subsequence mapped to the first portion of the first organism's base sequence, wherein the first portion of the first organism's base sequence is the ground truth of the additional first predicted base subsequence, (ii) the additional second predicted base subsequence mapped to the second portion of the first organism's base sequence, wherein the additional second portion of the first organism's base sequence is the ground truth of the additional second predicted base subsequence, and (iii) the additional third predicted base subsequence mapped to the third portion of the first organism's base sequence, wherein the additional third portion of the first organism's base sequence is the ground truth of the additional third predicted base subsequence.

[0608] A11. The method according to clause A10, the method further comprising:

[0609] A first error is generated between (i) the first predicted base sequence generated during the first iteration of the N1 iterations and (ii) the first portion of the first organism base sequence; and

[0610] A second error is generated between (i) the additional first predicted base sequence generated during the second iteration of the N1 iterations and (ii) the first portion of the first organism base sequence.

[0611] The second error is smaller than the first error because the base detector was better trained during the second iteration compared to the first iteration.

[0612] A12. The method described according to clause A10, wherein:

[0613] In the second iteration, the first sequence signal, the second sequence signal, and the third sequence signal generated during the first iteration are reused to generate the additional first predicted base sequence, the additional second predicted base sequence, and the additional third predicted base sequence, respectively.

[0614] A13. The method described according to clause A10, wherein:

[0615] The neural network configuration of the base detector is the same during the first iteration and the second iteration of the N1 iterations.

[0616] A13a. The method described according to clause A13, wherein:

[0617] The neural network configuration of the base detector is repeated for multiple iterations until the convergence condition is met.

[0618] A14. The method described according to clause A10, wherein:

[0619] The neural network configuration of the base detector during the first iteration of the N1th iteration is different from that during the second iteration of the N1th iteration, and is more complex than the neural network configuration of the base detector during the second iteration of the N1th iteration.

[0620] A15. The method according to clause A1, wherein further training the base detector using the analyte containing the base sequence of the first organism for the N1 iterations of the N iterations comprises:

[0621] For the first subset of the N1 iterations, the base detector is further trained using the first neural network configuration loaded in the base detector;

[0622] For the second subset of the N1 iterations, the base detector is further trained using a second neural network configuration loaded in the base detector, which is different from the first neural network configuration.

[0623] A16. The method according to clause A15, wherein the second neural network configuration has a greater number of layers than the first neural network configuration.

[0624] A17. The method according to clause A15, wherein the second neural network configuration has a larger number of weights than the first neural network configuration.

[0625] A18. The method according to clause A15, wherein the second neural network configuration has a larger number of parameters than the first neural network configuration.

[0626] A19. The method according to clause A1, wherein iteratively further training of the base detector comprises:

[0627] For one or more iterations of the N1 iterations of the anal...

Claims

1. A computer-implemented method for progressively training a base detector, the method comprising: A base detector is iteratively initially trained using an analyte with a single known oligonucleotide base sequence, and labeled training data is generated using the initially trained base detector, which has an initial neural network configuration including multiple layers and parameters. (i) The base detector is further trained using an analyte containing multiple oligonucleotide base sequences, the multiple oligonucleotide base sequences comprising at least two known oligonucleotide base sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and labeled training data is generated using the further trained base detector. as well as The base detector is further trained iteratively by repeating step (i), while during at least one iteration, the complexity of the initial neural network configuration of the base detector is increased by adjusting the number of layers and parameters relative to the initial neural network configuration for the oligonucleotide base sequence, wherein the labeled training data generated during the iteration is used to train the base detector in the immediately following subsequent iteration.

2. The computer-implemented method according to claim 1, the method further comprising: During at least one iteration of further training of the base detector using the analyte containing the oligonucleotide base sequence: Increase the number of unique oligonucleotide base sequences of the poly-oligonucleotide base sequence in the analyte; as well as The initial neural network configuration is further adjusted by increasing the number of layers and parameters to provide a unique number of oligonucleotide base sequences for the increased number of layers.

3. The computer-implemented method according to claim 1 or 2, wherein iteratively initial training of the base detector using the analyte comprising the single oligonucleotide base sequence comprises: During the first iteration of iteratively initial training of the base detector: The single oligonucleotide base sequence is filled into multiple clusters in the flow cell; Generate multiple sequence signals corresponding to the plurality of clusters, each of the plurality of sequence signals representing a base sequence loaded in the corresponding cluster among the plurality of clusters; Based on each of the plurality of sequence signals, the corresponding base detection of the single known oligonucleotide base sequence is predicted, thereby generating a plurality of predicted base detections; For each of the plurality of sequence signals, based on (i) the corresponding predicted base detection and (ii) the comparison of the bases of the single known oligonucleotide base sequence, a corresponding error signal is generated, thereby generating a plurality of error signals corresponding to the plurality of sequence signals; as well as The base detector is initially trained during the first iteration based on the multiple error signals.

4. The computer-implemented method of claim 3, wherein iteratively initial training of the base detector using the analyte comprising the single oligonucleotide base sequence further comprises: During the second iteration of iteratively initial training of the base detector, which occurs after the first iteration of iteratively initial training of the base detector: Using the base detector that has been partially trained during the first iteration, additional base detections corresponding to the single known oligonucleotide base sequence are predicted based on each of the plurality of sequence signals, thereby generating a plurality of additional predicted base detections. For each of the plurality of sequence signals, based on (i) the comparison of the corresponding additional predicted base detection with (ii) the bases of the single known oligonucleotide base sequence, a corresponding additional error signal is generated, thereby generating a plurality of additional error signals corresponding to the plurality of sequence signals; as well as Based on the additional error signals, the base detector is further initially trained during the second iteration.

5. The computer-implemented method according to claim 4, wherein: The plurality of sequence signals corresponding to the plurality of clusters generated during the first iteration of iteratively initial training of the base detector are repeatedly used in the second iteration of iteratively initial training of the base detector.

6. The computer-implemented method according to any one of claims 4 to 5, wherein comparing (i) the corresponding predicted base detection with (ii) the bases of a single known oligonucleotide sequence comprises: For a first predicted base detection, (i) the first base of the first predicted base detection is compared with the first base of the single known oligonucleotide base sequence, and (ii) the second base of the first predicted base detection is compared with the second base of the single known oligonucleotide base sequence, thereby generating a corresponding first error signal.

7. The computer-implemented method according to claim 1 or 2, wherein iteratively further training of the base detector comprises: The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations. as well as The base detector was further trained using an analyte containing three known unique oligonucleotide base sequences for N² iterations. The N1 iterations are performed before the N2 iterations.

8. A system for progressively training a base detector, the system comprising: At least one processor; as well as A non-transitory computer-readable storage medium having printed computer program instructions that, when executed on the at least one processor, perform actions including: A base detector is iteratively initially trained using an analyte with a single known oligonucleotide base sequence, and labeled training data is generated using the initially trained base detector, which has an initial neural network configuration including multiple layers and parameters. (i) The base detector is further trained using an analyte containing multiple oligonucleotide base sequences, the multiple oligonucleotide base sequences comprising at least two known oligonucleotide base sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and labeled training data is generated using the further trained base detector. as well as The base detector is further trained iteratively by repeating step (i), while during at least one iteration, the complexity of the initial neural network configuration of the base detector is increased by adjusting the number of layers and parameters relative to the initial neural network configuration for the oligonucleotide base sequence, wherein the labeled training data generated during the iteration is used to train the base detector in the immediately following subsequent iteration.

9. The system of claim 8, wherein during the iterative initial training of the base detector using the analyte comprising the single oligonucleotide base sequence, a first neural network configuration is loaded within the base detector, and wherein iteratively further training of the base detector comprises: The base detector was further trained using an analyte containing two known unique oligonucleotide base sequences for N1 iterations, resulting in... (i) For the first subset of the N1 iterations, a second neural network configuration is loaded into the base detector, and (ii) For a second subset of the N1 iterations that occurs after the first subset of the N1 iterations, a third neural network configuration is loaded within the base detector, wherein the first neural network configuration, the second neural network configuration, and the third neural network configuration are different from each other.

10. The system of claim 9, wherein the second neural network configuration is more complex than the first neural network configuration, and wherein the third neural network configuration is more complex than the second neural network configuration.

11. The system of claim 9 or 10, wherein the second neural network configuration has a greater number of layers, weights, or parameters than the first neural network configuration.

12. The system of claim 9 or 10, wherein the third neural network configuration has a greater number of layers, weights, or parameters than the second neural network configuration.

13. The system of claim 9, wherein, for one of the N1 iterations, further training the base detector using the analyte comprising two known unique oligonucleotide base sequences comprises: (i) filling a first plurality of clusters in the flow cell with a first known oligonucleotide sequence of the two known unique oligonucleotide sequences, and (ii) filling a second plurality of clusters in the flow cell with a second known oligonucleotide sequence of the two known unique oligonucleotide sequences; For each of the first plurality of clusters and the second plurality of clusters, predict the corresponding base detection, thereby generating multiple predicted base detections; Map the first predicted base detection among the plurality of predicted base detections in (i) to the first known oligonucleotide base sequence, and map the second predicted base detection among the plurality of predicted base detections in (ii) to the second known oligonucleotide base sequence, while avoiding mapping the third predicted base detection among the plurality of predicted base detections to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence; Generate (i) a first error signal based on the comparison of the first predicted base detection with the first known oligonucleotide base sequence, and (ii) a second error signal based on the comparison of the second predicted base detection with the second known oligonucleotide base sequence; as well as The base detector is further trained based on the first error signal and the second error signal.

14. The system of claim 13, wherein mapping the first predicted base detection to the first known oligonucleotide sequence in the two known unique oligonucleotide sequences comprises: Each base detected by the first predicted base is compared with the corresponding bases of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; The first predicted base detection is determined to have at least a threshold number of base similarities with the first known oligonucleotide base sequence, and to have less than the threshold number of base similarities with the second known oligonucleotide base sequence; as well as Based on the determination that the first predicted base detection has at least the threshold number of base similarities to the first known oligonucleotide base sequence, the first predicted base detection is mapped to the first known oligonucleotide base sequence.

15. The system of claim 13 or 14, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises: Each base detected by the first predicted base is compared with the corresponding bases of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; The first predicted base detection is determined to have less than a threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; as well as Based on the determination that the first predicted base detection has less than the threshold number of base similarities to each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

16. The system of claim 13 or 14, wherein avoiding mapping the third predicted base detection to either the first known oligonucleotide sequence or the second known oligonucleotide sequence comprises: Each base detected by the first predicted base is compared with the corresponding bases of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; The first predicted base detection is determined to have a base similarity greater than a threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence; as well as Based on the determination that the first predicted base detection has a base similarity greater than the threshold number with each of the first known oligonucleotide base sequence and the second known oligonucleotide base sequence, the mapping of the third predicted base detection to either the first known oligonucleotide base sequence or the second known oligonucleotide base sequence is avoided.

17. The system of claim 13 or 14, wherein generating labeled training data by performing one of the N1 iterations using the further trained base detector comprises: After further training the base detector during one of the N1 iterations, the corresponding base detection is re-predicted for each of the first plurality of clusters and the second plurality of clusters, thereby generating a plurality of additional predicted base detections. (i) remap the first subset of the additional predicted bases detected to the first known oligonucleotide sequence, and (ii) remap the second subset of the additional predicted bases detected to the second known oligonucleotide sequence, while avoiding mapping the third subset of the additional predicted bases detected to either the first known oligonucleotide sequence or the second known oligonucleotide sequence. as well as The remapping is used to generate labeled training data, such that the labeled training data includes (i) the first subset of the additional plurality of predicted base detections, wherein the first known oligonucleotide base sequence forms the benchmark ground truth data of the first subset of the additional plurality of predicted base detections, and (ii) the second subset of the additional plurality of predicted base detections, wherein the second known oligonucleotide base sequence forms the benchmark ground truth data of the second subset of the additional plurality of predicted base detections.

18. The system according to claim 17, wherein: The labeled training data generated during one of the N1 iterations is used to train the base detector during the subsequent iterations immediately following the N1 iterations. as well as The neural network configuration of the base detector remains unchanged during the first iteration of the N1 iterations and the immediately following iterations of the N1 iterations; or The neural network configuration of the base detector during the immediately following iteration of the N1th iteration is different from that of the base detector during the first iteration of the N1th iteration, and is more complex than the neural network configuration of the base detector during the first iteration of the N1th iteration.

19. A non-transitory computer-readable storage medium having printed computer program instructions for progressively training a base detector, the computer program instructions performing actions including the following when executed on a processor: A base detector is iteratively initially trained using an analyte with a single known oligonucleotide base sequence, and labeled training data is generated using the initially trained base detector, which has an initial neural network configuration including multiple layers and parameters. (i) The base detector is further trained using an analyte containing multiple oligonucleotide base sequences, the multiple oligonucleotide base sequences comprising at least two known oligonucleotide base sequences that differ from each other by a threshold edit distance between their respective nucleotide bases, and labeled training data is generated using the further trained base detector. as well as The base detector is further trained iteratively by repeating step (i), while during at least one iteration, the complexity of the initial neural network configuration of the base detector is increased by adjusting the number of layers and parameters relative to the initial neural network configuration for the oligonucleotide base sequence, wherein the labeled training data generated during the iteration is used to train the base detector in the immediately following subsequent iteration.

20. The non-transitory computer-readable storage medium of claim 19, further comprising computer program instructions, which, when executed on the processor, perform actions including iteratively further training the base detector by means of the following steps: Before further training the base detector using analytes containing oligonucleotide base sequences, and in response to adjusting the initial neural network configuration for the oligonucleotide base sequences by increasing the number of layers and parameters relative to the initial neural network configuration, the base detector is further trained using labeled training data generated during the final iteration of iteratively initial training of the base detector.