Artificial intelligence-based base calling of index sequences
Normalization of index images across adjacent sequencing cycles addresses the low diversity and contrast issues in index images, enhancing the accuracy of neural network-based base callers in multiplexed sequencing by introducing signal diversity and reducing overfitting.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-02-16
- Publication Date
- 2026-03-03
AI Technical Summary
Existing neural network-based base callers for index sequences in next-generation sequencing (NGS) suffer from degraded performance when index images are not normalized, leading to reduced accuracy in base calling due to low nucleotide diversity and signal contrast in index images.
Implementing a normalization process for index images using intensity values from adjacent sequencing cycles to enhance signal diversity, followed by processing through a neural network-based base caller, which includes preprocessing techniques to improve accuracy.
Enhances the base calling accuracy of index sequences by introducing signal diversity and reducing overfitting, thereby improving the performance of neural network-based base callers in multiplexed sequencing runs.
Smart Images

Figure 0007822944000002 
Figure 0007822944000003 
Figure 0007822944000004
Abstract
Description
[Technical Field]
[0001] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data.
[0002] Priority application This PCT application claims priority to and the benefit of U.S. Provisional Patent Application No. 62 / 979,384, entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES," filed February 20, 2020 (Attorney Docket No. ILLM 1015-1 / IP-1857-PRV), and U.S. Patent Application No. 17 / 175,546, entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING OF INDEX SEQUENCES," filed February 12, 2021 (Attorney Docket No. ILLM 1015-2 / IP-1857-US), both of which are incorporated by reference for all purposes as if fully set forth herein. Built-in
[0003] The following are incorporated by reference as if fully set forth herein:
[0004] U.S. Provisional Patent Application No. 62 / 979,414, entitled "ARTIFICIAL INTELLIGENCE-BASED MANY-TO-MANY BASE CALLING," filed February 20, 2020 (Attorney Docket No. ILLM 1016-1 / IP-1858-PRV);
[0005] U.S. Provisional Patent Application No. 62 / 979,385, entitled "KNOWLEDGE DISTILLATION-BASED COMPRESSION OF ARTIFICIAL INTELLIGENCE-BASED BASE CALLER," filed February 20, 2020 (Attorney Docket No. ILLM 1017-1 / IP-1859-PRV);
[0006] U.S. Provisional Patent Application No. 63 / 072,032, entitled "DETECTING AND FILTERING CLUSTERS BASED ON ARTIFICIAL INTELLIGENCE-PREDICTED BASE CALLS," filed August 28, 2020 (Attorney Docket No. ILLM 1018-1 / IP-1860-PRV);
[0007] U.S. Provisional Patent Application No. 62 / 979,412, entitled "MULTI-CYCLE CLUSTER BASED REAL TIME ANALYSIS SYSTEM," filed February 20, 2020 (Attorney Docket No. ILLM 1020-1 / IP-1866-PRV);
[0008] U.S. Provisional Patent Application No. 62 / 979,411, entitled "DATA COMPRESSION FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed February 20, 2020 (Attorney Docket No. ILLM 1029-1 / IP-1964-PRV);
[0009] U.S. Provisional Patent Application No. 62 / 979,399, entitled "SQUEEZING LAYER FOR ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed February 20, 2020 (Attorney Docket No. ILLM 1030-1 / IP-1982-PRV);
[0010] U.S. Patent Application No. 16 / 825,987, entitled "TRAINING DATA GENERATION FOR ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-16 / IP-1693-US);
[0011] U.S. Provisional Patent Application No. 16 / 825,991, entitled "ARTIFICIAL INTELLIGENCE-BASED GENERATION OF SEQUENCING METADATA," filed March 20, 2020 (Attorney Docket No. ILLM 1008-17 / IP-1741-US);
[0012] U.S. Patent Application No. 16 / 826,126, entitled "ARTIFICIAL INTELLIGENCE-BASED BASE CALLING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-18 / IP-1744-US);
[0013] U.S. Patent Application No. 16 / 826,134, entitled "ARTIFICIAL INTELLIGENCE-BASED QUALITY SCORING," filed March 20, 2020 (Attorney Docket No. ILLM 1008-19 / IP-1747-US); and
[0014] U.S. Patent Application No. 16 / 826,168, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2020 (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV-US). [Background technology]
[0015] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to embodiments of the claimed technology.
[0016] Improvements in next-generation sequencing (NGS) technology have dramatically increased sequencing speed and data output, resulting in the massive sample throughput of current sequencing platforms. Approximately 10 years ago, the Illumina Genome Analyzer™ was capable of generating up to 1 gigabyte of sequence data per run. Today, the Illumina NovaSeq™ series of systems can generate up to 2 terabytes of data in two days, representing a more than 2000-fold increase in capacity. Summary of the Invention [Means for solving the problem]
[0017] The key to harnessing this increased capacity is multiplexing, which allows for the pooling and sequencing of multiple libraries simultaneously during a single sequencing run by adding a unique index sequence ("barcode") to each DNA fragment during library preparation. Sequencing reads are sorted into their respective samples during demultiplexing, allowing for proper alignment.
[0018] Opportunities arise for using artificial intelligence and neural networks to base call index sequences. Higher base calling throughput and higher base calling accuracy can result.
[0019] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR (patent application information retrieval) via the Supplemental Content tab.
[0020] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 1 shows one embodiment of sequencing polynucleotides from an indexed library. [Figure 2] FIG. 1 shows an embodiment in which a target sequence is sequenced to generate target reads and an index sequence is sequenced to generate index reads. [Figure 3] FIG. 1 illustrates one implementation of index image normalization. [Figure 4] FIG. 1 illustrates one implementation of processing normalized index images through a neural network-based base caller for base calling. [Figure 5] FIG. 10 illustrates one implementation of extending index image normalization to non-current index sequencing cycles. [Figure 6] FIG. 10 illustrates one embodiment of index image normalization using at least one index image showing one or more nucleotides in a detectable signal state. [Figure 7] FIG. 1 shows one embodiment of base calling of a target sequence and an index sequence. [Figure 8]FIG. 1 illustrates one embodiment of pre-processing using augmentation. [Figure 9] FIG. 1 shows pixel intensity histograms of red and green images of two targeted sequencing cycles (cycles 1 and 151) of the first targeted read (read 1). [Figure 10] FIG. 1 shows pixel intensity histograms of red and green images of two targeted sequencing cycles (cycles 1 and 151) of the first targeted read (read 1). [Figure 11] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 12] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 13] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 14] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 15] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 16]FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 17] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 18] FIG. 1 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (Index Read 1). [Figure 19] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 20] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 21] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 22] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 23]FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 24] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 25] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 26] FIG. 10 shows pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2). [Figure 27] FIG. 10 shows pixel intensity histograms of red and green images of two targeted sequencing cycles (cycles 168 and 169) of the second targeted read (read 2). [Figure 28] FIG. 10 shows pixel intensity histograms of red and green images of two targeted sequencing cycles (cycles 168 and 169) of the second targeted read (read 2). [Figure 29] FIG. 1 shows that in a sequencing run using four index arrays to multiplex four samples, the index base calling performance of a neural network-based base caller degrades when the index images are not normalized. [Figure 30]FIG. 1 shows that in a sequencing run using two index arrays to multiplex two samples, the index base calling performance of a neural network-based base caller degrades when the index images are not normalized. [Figure 31] FIG. 10 shows that in a sequencing run using a single index array to sequence a single sample, the index base calling performance of a neural network-based base caller degrades when the index image is not normalized. [Figure 32] A computer system that can be used to implement the disclosed techniques. [Figure 33] FIG. 1 shows another embodiment of base calling of target and index sequences. [Figure 34] 1 is one embodiment of a flowchart of an artificial intelligence-based method for base calling samples in the index sequencing cycle of a sequencing run. [Figure 35] 1 is an embodiment of a flowchart of an artificial intelligence-based method for base calling a target sequence and an index sequence. DETAILED DESCRIPTION OF THE INVENTION
[0022] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0023] multiplexing Figure 1 shows one embodiment of sequencing polynucleotides from an indexed library. When polynucleotides from different libraries are pooled or multiplexed for sequencing, polynucleotides from each library are modified to include a library-specific index sequence. During sequencing, the index sequence is sequenced along with a target polynucleotide sequence from the library. The index sequence is associated with the target polynucleotide sequence so that the library from which the target sequence originated can be identified.
[0024] Further details regarding multiplexing, index sequences, and demultiplexing can be found in Illumina, "Indexed Sequencing Overview Guide," Document No. 15057455, v.5, March 2019, and in Illumina's U.S. Patent Application Publication Nos. 2018 / 0305751, 2018 / 0334712, 2016 / 0110498, 2018 / 0334711, and WO 2019 / 090251, each of which is incorporated herein by reference.
[0025] Panel A shows an indexed library 102, where unique index sequences ("indexes") are added to two different libraries during library preparation. The first index sequence (index 1) has a barcode of "CATTCG." The second index sequence (index 2) has a barcode of "AACTGA."
[0026] Panel B shows pooling 104, where indexed libraries 102 are pooled together and loaded into the same flow cell lane.
[0027] Panel C shows sequencing 106 and sequencing output 116, where indexed libraries 102 are sequenced together during a single run of the instrument. All sequences are then exported to output file 116. Output file 116 contains sequence reads (green) joined to corresponding index reads (blue and magenta).
[0028] Panel D shows demultiplexing 108, where the demultiplexing algorithm sorts sequence reads into different files according to their indexes.
[0029] Panel E shows alignment 110, where each set of demultiplexed sequence reads is aligned to an appropriate reference sequence.
[0030] Target sequence and index sequence FIG. 2 illustrates one embodiment in which a target sequence 222 is sequenced to generate a target read 202 ("GTCCGATA") and an index sequence 232 is sequenced to generate an index read 204 ("AACTGA"). The index sequence 232 can be a synthetic sequence of nucleotides attached to the target sequence 222 during a template preparation step. The target sequence 222 can be naturally occurring DNA, RNA, or some other biological molecule. The length of the index sequence 232 can range from 2 to 20 nucleotides. For example, the index sequence 232 can be 1 to 10 nucleotides long or 4 to 6 nucleotides long. A 4-nucleotide index sequence allows for multiplexing 256 samples on the same array. A 6-nucleotide index sequence allows for processing 4096 samples on the same array.
[0031] During sequencing 106, target primer 212 traverses target sequence 222 to generate target read 202 ("GTCCGATA"), and index primer 224 traverses index sequence 232 to generate index read 204 ("AACTGA"). In some embodiments, sequencing 106 is Illumina single-indexed sequencing. In other embodiments, sequencing 106 is Illumina dual-indexed sequencing.
[0032] Base calling is the process of determining the nucleotide composition of target sequence 222 and index sequence 232, i.e., generating target read 202 ("GTCCGATA") and index read 204 ("AACTGA"). Base calling involves analyzing image data, i.e., analyzing sequencing images generated during sequencing 106 by a sequencing instrument such as Illumina's iSeq, HiSeqX, HiSeq 3000, HiSeq 4000, HiSeq 2500, NovaSeq 6000, NextSeq, NextSeqDx, MiSeq, and MiSeqDx. The following description outlines how sequencing images are generated and what they depict, according to one embodiment.
[0033] Base calling decodes the raw signal of the sequencing instrument, i.e., the intensity data extracted from the sequencing image, into a nucleotide sequence. In one embodiment, the Illumina platform employs cyclic reversible termination (CRT) chemistry for base calling. This process relies on extending a nascent strand complementary to the template strand with fluorescently labeled nucleotides while tracking the emission signal of each newly added nucleotide. The fluorescently labeled nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide.
[0034] Sequencing 106 is performed in repetitive cycles, each of which includes three steps: (a) extending the nascent strand (e.g., target sequence 222, index sequence 232) by adding fluorescently labeled nucleotides; (b) exciting fluorophores using one or more lasers in the sequencing instrument's optical system and generating a sequencing image by imaging through different filters in the optical system; and (c) cleaving the fluorophores and removing the 3' block in preparation for the next sequencing cycle. The capture and imaging cycles are repeated for a specified number of sequencing cycles, which defines the read length. Using this approach, each cycle locates a new position along the template strand.
[0035] The enormous power of the Illumina platform stems from its ability to simultaneously perform and sense CRT reactions of millions or even billions of analytes (e.g., clusters). A cluster contains approximately 1,000 identical copies of a template strand, but the size and shape of the cluster vary. Prior to a sequencing run, clusters are extended from the template strand by bridge amplification of the input library. The purpose of amplification and cluster extension is to increase the intensity of the emitted signal, since imaging devices cannot reliably sense the fluorophore signal of a single strand. However, because the physical distance between strands within a cluster is small, imaging devices perceive the cluster of strands as a single spot.
[0036] Sequencing 106 occurs in a flow cell, a small glass slide that holds the input strand. The flow cell is connected to an optical system that includes microscope imaging, an excitation laser, and fluorescence filters. The flow cell contains multiple chambers called lanes. The lanes are physically separated from one another and can contain distinct tagged sequencing libraries that can be distinguished without sample cross-contamination. The sequencing instrument's imaging device (e.g., a solid-state imager such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) sensor) captures snapshots at multiple positions along the lane in a series of non-overlapping regions called tiles. For example, Illumina's Genome Analyzer II has 100 tiles per lane, and Illumina's HiSeq2000 has 68 tiles per lane. Tiles hold hundreds of thousands to millions of clusters.
[0037] The output of sequencing 106 are sequencing images, each showing the intensity radiation of a cluster and its surrounding background. A sequencing cycle of sequencing 106 that sequences a target sequence 222 is referred to as a "target sequencing cycle," and a sequencing cycle of sequencing 106 that sequences an index sequence 232 is referred to as an "index sequencing cycle." A sequencing image generated during a target sequencing cycle is referred to as a "target image," and a sequencing image generated during an index sequencing cycle is referred to as an "index image."
[0038] The target image shows the intensity radiation generated as a result of nucleotide incorporation into the target sequence during sequencing 106. The index image shows the intensity radiation generated as a result of nucleotide incorporation into the index sequence during sequencing 106. The intensity radiation is from the relevant analytes and their surrounding background.
[0039] (Neural network-based base calling) The discussion now turns to neural network-based base calling, where a neural network, i.e., neural network-based base caller 430, is trained to map sequencing images to base calls 432.
[0040] The description is structured as follows: First, the input to neural network-based base caller 430 is described according to one embodiment. Then, an example of the structure and configuration of neural network-based base caller 430 is provided. Finally, the output of neural network-based base caller 430 is described according to one embodiment.
[0041] Further details regarding the neural network-based base collation 430 can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2019 (Attorney Docket No. ILLM 1008-9 / IP-1752-PRV), which is incorporated herein by reference.
[0042] In one embodiment, image patches are extracted from the target image and the index image. The extracted image patches are provided to a neural network-based base caller 430 as "input image data" for base calling. The image patches have dimensions w x h, where w (width) and h (height) are any number ranging from 1 to 10,000 (e.g., 3 x 3, 5 x 5, 7 x 7, 10 x 10, 15 x 15, 25 x 25). In some embodiments, w and h are the same. In other embodiments, w and h are different.
[0043] Sequencing 106 generates m images per sequencing cycle for m corresponding image channels. In one embodiment, each image channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each image channel corresponds to one of a plurality of imaging events in a sequencing cycle. In yet another embodiment, each image channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.
[0044] To prepare input image data for a particular sequencing cycle, image patches are extracted from each of the m images. In different embodiments, such as for 4-, 2-, and 1-channel chemistries, m is 4 or 2. In other embodiments, m is greater than 1, 3, or 4. The input image data is in the optical pixel domain in some embodiments and in the upsampled sub-pixel domain in other embodiments.
[0045] For example, consider the case where sequencing 106 uses two different image channels, i.e., a red channel and a green channel. In this case, in each sequencing cycle, sequencing 106 generates a red image and a green image. In this way, for a series of k sequencing cycles, a sequence having k pairs of red and green images is generated as output.
[0046] The input image data includes an array of cycle-by-cycle image patches generated for a series of k sequencing cycles of a sequencing run. The cycle-by-cycle image patches include intensity data for associated analytes and their surrounding background in one or more image channels (e.g., red and green channels). In one embodiment, when a single target analyte (e.g., a cluster) is base called, the cycle-by-cycle image patch is centered on a central pixel that includes intensity data for the target-associated analyte, and the off-center pixels of the cycle-by-cycle image patch include intensity data for associated analytes adjacent to the target-associated analyte.
[0047] The input image data includes data from multiple sequencing cycles (e.g., a current sequencing cycle, one or more preceding sequencing cycles, and one or more consecutive sequencing cycles). In one embodiment, the input image data includes data from three sequencing cycles, such that the data from the current (time t) sequencing cycle to be base-called is accompanied by data from (i) the left adjacent / context / previous / preceding / previous (time t-1) sequencing cycle, and (ii) the right adjacent / context / next / consecutive / subsequent (time t+1) sequencing cycle. In other embodiments, the input image data includes data from a single sequencing cycle. In still other embodiments, the input image data includes data from 58, 75, 92, 130, 168, 175, 209, 225, 230, 275, 318, 325, 330, 525, or 625 sequencing cycles.
[0048] In one embodiment, neural network-based base caller 430 is a multilayer perceptron (MLP). In another embodiment, neural network-based base caller 430 is a feed-forward neural network. In yet another embodiment, neural network-based base caller 430 is a fully connected neural network. In a further embodiment, neural network-based base caller 430 is a fully convolutional neural network. In yet another embodiment, neural network-based base caller 430 is a semantic segmentation neural network.
[0049] In one embodiment, the neural network-based base caller 430 is a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, it is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, the neural network-based base caller includes both a CNN and an RNN.
[0050] In yet other embodiments, the neural network-based base caller 430 may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, pointwise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It may use one or more loss functions such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. Neural network-based base callers can use any parallelism, efficiency, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharding, parallel calls to map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous SGD. Neural network-based base callers can include upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), sigmoid, and hyperbolic tangent function (tanh)), batch normalization layers, normalization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, and attention mechanisms.
[0051] In one embodiment, the neural network-based base caller 430 outputs a base call for a single target analyte in a particular sequencing cycle. In another embodiment, the neural network-based base caller outputs a base call for each target analyte of a plurality of target analytes in a particular sequencing cycle. In yet another embodiment, the neural network-based base caller generates a base call sequence for each target analyte by outputting a base call for each target analyte of a plurality of target analytes in each sequencing cycle of a plurality of sequencing cycles.
[0052] Pretreatment In one embodiment, the image data from the target image and the index image are not directly provided as input to the neural network-based base caller 430. Instead, the target image and the index image are first preprocessed. However, the index image is preprocessed in a different way than the target image.
[0053] The base calling logic described herein accounts for the observation that index images show nucleotides with low complexity patterns in which some of the four bases A, C, T, and G are represented at frequencies less than 15%, 10%, or 5% of all nucleotides, because for any given index sequencing cycle, an index image shows (1) the intensity emissions of multiple analytes that originate from the same sample and share the same index sequence, and (2) the intensity emissions of analytes that belong to different samples and have different index sequences.
[0054] The first type of specimen has the same index base for every index sequencing cycle. As a result, the index image shows the same nucleotide for multiple specimens. This reduces the nucleotide diversity of the index image.
[0055] The nucleotide diversity of an index image is even lower when a second type of specimen has the same index base for a particular index sequencing cycle. This occurs for two reasons. First, index sequences are short sequences with 2 to 20 index bases and therefore do not have enough positions where significant mismatches can occur between different index sequences. Second, in many cases, up to 20 samples are pooled for simultaneous sequencing. As a result, the number of different index sequences that can be depicted by a single index image is not substantial. These factors result in different index sequences having matching index bases at the same positions (base collisions), which causes specimens with different index sequences to have the same index base for a particular index sequencing cycle.
[0056] Low nucleotide diversity in the index image creates an intensity pattern lacking signal diversity (contrast). In contrast, the target image exhibits a highly complex pattern of nucleotides, with each of the four bases A, C, T, and G represented at a frequency of at least 20%, 25%, or 30% of all nucleotides. This is because target sequences are often long (e.g., 150 bases) and unique to each specimen, regardless of the original sample. Therefore, unlike the index image, the target image has adequate signal diversity.
[0057] The convolution kernels and filters of the neural network-based base caller 430 are trained primarily on the target image, so that during inference, when the trained neural network-based base caller 430 is presented with raw index images that have not undergone preprocessing, the base calling accuracy of the index reads will decrease because the convolution kernels and filters are trained to detect intensity patterns based on contrast.
[0058] Bypassing preprocessing by training a neural network-based base caller 430 on a large number of raw index images to introduce signal diversity is not feasible because only a very large number of index sequences are published and publicly available. Also, users often design custom index sequences and use them instead of published index sequences. As a result, a neural network-based base caller 430 trained only on raw index images tends to generalize poorly and overfit during inference.
[0059] One solution is to preprocess the index image using normalization: the index image from the current index sequencing cycle is normalized based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle.
[0060] The intensity values measure the chemiluminescent signal generated due to the incorporation of the nucleotide. The intensity values are encoded into an "image" and represent an "optical signal" that includes a "specific signal." As used herein, the term "image" is intended to mean a representation of all or a portion of an object. The representation may be an optically detected reproduction. For example, an image can be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of an object present in an image may be the surface or other x-y plane of the object. An image is a two-dimensional representation, but in some cases, information in an image can be derived from three or more dimensions. An image need not include optically detected signals. Signals other than light may instead be present (such as voltage, pH, or ion data). An image can be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein. As used herein, the term "optical signal" is intended to include, for example, fluorescence, luminescence, scattering, or absorption signals. Optical signals can be detected in the ultraviolet (UV) range (approximately 200-390 nm), visible (VIS) range (approximately 391-770 nm), infrared (IR) range (approximately 0.771-25 micrometers), or other ranges of the electromagnetic spectrum. Optical signals can be detected in a manner that excludes all or part of one or more of these ranges. As used herein, the term "specific signal" is intended to mean detected energy or encoded information that is selectively observed over other energy or information, such as background energy or information. For example, a specific signal can be an optical signal detected at a specific intensity, wavelength, or color; an electrical signal detected at a specific frequency, power, or field strength; or other signals known in the art related to spectroscopy and analytical detection. In one embodiment, intensity values are extracted from a sequencing image in two different color / intensity channels. The identities of the four different nucleotide types / bases A, C, T, and G are encoded as a combination of intensity values in two color images, i.e., the first and second intensity channels.For example, a nucleic acid can be sequenced by providing a first nucleotide type (e.g., base T) that is detected in a first intensity channel, a second nucleotide type (e.g., base C) that is detected in a second intensity channel, a third nucleotide type (e.g., base A) that is detected in both the first and second intensity channels, and a fourth nucleotide type (e.g., base G) that lacks a label and is not or only minimally detected in either intensity channel. In some embodiments, four intensity distributions (e.g., Gaussian distributions) are iteratively fitted to the intensity values in the first and second intensity channels. The four intensity distributions correspond to the four bases A, C, T, and G. The intensity values in the first intensity channel are plotted (e.g., as a scatter plot) against the intensity values in the second intensity channel, and the intensity values are separated into the four intensity distributions.
[0061] Normalization across index sequencing cycles also includes normalization across image channels within the image data of the index sequencing cycles. For example, consider a case where there are three index sequencing cycles, i.e., a first index sequencing cycle, a second index sequencing cycle, and a third index sequencing cycle. Also consider a case where each of the first, second, and third index sequencing cycles has two index images: a first index image (e.g., red index image) in a first image channel (e.g., red channel) and a second index image (e.g., green index image) in a second image channel (e.g., green channel). The red index image from the second index sequencing cycle is normalized based on (i) the intensity values of the red and green images from the first index sequencing cycle, (ii) the intensity values of the red and green images from the third index sequencing cycle, and (iii) the intensity values of the red and green images from the second index sequencing cycle. The green index image from the second index sequencing cycle is normalized based on (i) the intensity values of the red and green images from the first index sequencing cycle, (ii) the intensity values of the red and green images from the third index sequencing cycle, and (iii) the intensity values of the red and green images from the second index sequencing cycle.
[0062] The normalization includes index images from adjacent index sequencing cycles because the nucleotides represented by the index images from the current, preceding, and subsequent index sequencing cycles are cumulatively more diverse than the nucleotides represented by the index image from the current index sequencing cycle alone. Extending the normalization to index images from adjacent index sequencing cycles also includes at least one index image from the preceding and / or subsequent index sequencing cycles that represents one or more nucleotides in a detectable signal state. Further details are provided below.
[0063] Normalizing indexed images Figure 3 shows one implementation of normalization 344 of the index image.
[0064] The percentile calculation unit 302 calculates (312) the lower percentile of (i) the intensity values of index images 322, 332 from the preceding (time t-1) index sequencing cycle, (ii) the intensity values of index images 326, 336 from the subsequent (time t+1) index sequencing cycle, and (iii) the intensity values of index images 324, 334 from the current (time t) index sequencing cycle.
[0065] The percentile calculation unit 302 is configured with percentile calculation logic for calculating percentile intensity values of an image. The percentile calculation unit 302 may include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which implements certain techniques described herein and the software modules are stored on a computer-readable storage medium (or multiple such media).
[0066] As described above, each index sequencing cycle can have two, three, four, or more index images. Thus, the intensity values of the index images in the respective index image sets from each of the preceding (time t-1) index sequencing cycle, the subsequent (time t+1) index sequencing cycle, and the current (time t) index sequencing cycle are used to normalize the intensity values of the index images in the index image set from the current (time t) index sequencing cycle.
[0067] In the illustrated embodiment, each index sequencing cycle has two index images, one for the first image channel (eg, the red channel) and one for the second image channel (eg, the green channel).
[0068] In a preferred embodiment, the normalization of the index image of a first image channel (e.g., the red channel) uses the index image of the first image channel and also one or more index images of other image channels (e.g., the green channel).
[0069] In other embodiments, normalization of an index image for a particular image channel uses only the index image for that particular image channel and not an index image for a different image channel. For example, in such an embodiment, the current normalized index image for the first channel 364 is generated from only the intensity values of the preceding index image for the first channel 322 and the intensity values of the succeeding index image for the first channel 326. Similarly, the current normalized index image for the second channel 374 is generated from only the intensity values of the preceding index image for the second channel 332 and the intensity values of the succeeding index image for the second channel 336.
[0070] The percentile calculation unit 302 also calculates (312) the upper percentiles of (i) the intensity values of index images 322, 332 from the preceding (time t-1) index sequencing cycle, (ii) the intensity values of index images 326, 336 from the subsequent (time t+1) index sequencing cycle, and (iii) the intensity values of index images 324, 334 from the current (time t) index sequencing cycle.
[0071] An image normalizer 354 then generates normalized versions 364, 374 of the index images 324, 334 based on the lower and upper percentiles so that a first percentage of normalized intensity values are below the lower percentile, a second percentage of normalized intensity values are above the upper percentile, and a third percentage of normalized intensity values are between the lower and upper percentiles.
[0072] In one example, the lower percentile may be the 5th percentile and the upper percentile may be the 95th percentile. The normalized intensity value for the 5th percentile may be zero, and the normalized intensity value for the 95th percentile may be 1. Thus, in a normalized version 364, 374 of an index image 324, 334, (i) 5 percent of the normalized intensity values are less than zero, (ii) another 5 percent of the normalized intensity values are greater than 1, and (iii) the remaining 90 percent of the normalized intensity values are between zero and 1. The intensity values may be pixel intensity values, sub-pixel intensity values, or super-pixel intensity values.
[0073] The normalization function can be expressed mathematically as follows:
number
[0074] Thus, in one example, if the intensity value is the 95th percentile intensity value, the normalized intensity value is 1, and if the intensity value is the 5th percentile, the normalized intensity value is zero.
[0075] In other embodiments, the lower percentile may be the 10th percentile and the upper percentile may be the 90th percentile. In still other embodiments, the lower percentile may be any number between 1 and 100, and the upper percentile may be 100 - the lower percentile. The normalized intensity values assigned to the lower and upper percentiles may also be different, such as between -1 and 1, between 0.5 and 1, between 1 and 10, between 1 and 99, etc.
[0076] FIG. 4 shows one embodiment in which normalized index images are processed through a neural network-based base caller 430 for base calling.
[0077] In one embodiment, normalized index images 404, 414 from the current (time t) index sequencing cycle are accompanied by normalized index images 402, 412 from the preceding (time t-1) index sequencing cycle and normalized index images 406, 416 from the following (time t+1) index sequencing cycle, which are normalized based on the intensity values of the index images in corresponding adjacent index sequencing cycles and their own respective intensity values, as described above.
[0078] According to one embodiment, the neural network-based base caller 430 processes the normalized index images 402, 412, 404, 414, 406, 416 through its convolutional layers to generate alternative representations. The alternative representations are then used by an output layer (e.g., a softmax layer) to generate base calls for only the current (time t) index sequencing cycle or for each of the index sequencing cycles, i.e., the current (time t) index sequencing cycle, the preceding (time t-1) index sequencing cycle, and the subsequent (time t+1) index sequencing cycle. The generated base calls form the index read.
[0079] In one implementation, a patch extraction process 424 extracts patches from the normalized index images 402, 412, 404, 414, 406, 416 to generate input image data 426 as described above. The extracted image patches in the input image data 426 are then provided as input to a neural network-based base caller 430.
[0080] In one embodiment, the index images are normalized during training and inference of the neural network-based base caller 430.
[0081] Further details regarding how the neural network-based base caller 430 performs the base calling and patch extraction process 424 can be found in U.S. Provisional Patent Application No. 62 / 821,766, entitled "ARTIFICIAL INTELLIGENCE-BASED SEQUENCING," filed March 21, 2019 (Attorney Docket No. ILLM 1008-9 / IP-1752-PRV), which is incorporated herein by reference.
[0082] FIG. 5 illustrates one embodiment that extends index image normalization to non-current index sequencing cycles.
[0083] In other embodiments, the index image from the current index sequencing cycle can be normalized based on (i) the intensity values of the index image from one or more non-current index sequencing cycles and (ii) the intensity values of the index image from the current index sequencing cycle. The index image from the non-current index sequencing cycle can be selected by the image selector 522 and provided to the percentile calculator 302 and the image normalizer 354 for normalization.
[0084] That is, normalization 344 can extend beyond just adjacent index sequencing cycles and does not necessarily have to use the immediately preceding or succeeding index sequencing cycle. For example, a non-current index sequencing cycle can include an initial index sequencing cycle 502 (e.g., the first 2, 3, 5, 10, or 20-index sequencing cycle). A non-current index sequencing cycle can include an intermediate index sequencing cycle 512 (e.g., the intermediate 2, 3, 5, 10, or 20-index sequencing cycle). A non-current index sequencing cycle can include a terminal index sequencing cycle 532 (e.g., the last 2, 3, 5, 10, or 20-index sequencing cycle).
[0085] Additionally, the non-current index sequencing cycles can include combinations of early index sequencing cycles, intermediate index sequencing cycles, and late index sequencing cycles (e.g., the 1st and 5th index sequencing cycles, the 15th and 23rd index sequencing cycles, and the 18th and 149th index sequencing cycles).
[0086] FIG. 6 illustrates one embodiment of index image normalization using at least one index image showing one or more nucleotides in a detectable signal state (ie, on / detectable).
[0087] One way to distinguish between different strategies for detecting nucleotide incorporation in sequencing reactions using a single fluorescent dye (or two or more dyes with the same or similar excitation / emission spectra) in terms of detectable signal states is by characterizing the incorporation in terms of the presence or relative absence of, or levels between, fluorescence transitions that occur during a sequencing cycle. Thus, sequencing strategies can be illustrated by their fluorescence profiles over a sequencing cycle. For the strategies disclosed herein, "1" or "on" and "0" or "off" refer to a fluorescence state (1 / on) in which the nucleotide is in a "detectable signal state" (e.g., detectable by fluorescence) or a fluorescence state (0 / off) in which the nucleotide is in a dark state (e.g., not detected or minimally detected in an imaging step). A "0" or "off" state does not necessarily refer to a complete lack or absence of signal. However, in some embodiments, a complete lack or absence of signal (e.g., fluorescence) may occur. Minimal or reduced fluorescence signals (e.g., background signals) are also considered to be included in the "0" or "off" state range, as long as the change in fluorescence from the first image to the second image (or vice versa) can be reliably distinguished.
[0088] In the illustrated two-channel embodiment of Figure 6, nucleotide "G" is dark / off in both index images, nucleotide "A" is on / detectable in both index images, nucleotide "C" is dark / off in the first index image and on / detectable in the second index image, and nucleotide "T" is on / detectable in the first index image and dark / off in the second index image.
[0089] In one embodiment, the image selector 522 selects 622 index images from non-current index sequencing cycles that are in a detectable signal state and passes them to the percentile calculator 302 and image normalizer 354 to generate a normalized image 632. The on / detectable index images can come from non-current index sequencing cycles in which all index images are in a detectable signal state (e.g., the t+3 index sequencing cycle) or from non-current index sequencing cycles in which only some index images are in a detectable signal state (e.g., the t-2 index sequencing cycle).
[0090] In some implementations, index images of many detectable signal states can be used to normalize the index images.
[0091] In a preferred embodiment, on / detectable index images are selected across multiple channels such that an index image for a first image channel (e.g., a red channel) is normalized using one or more on / detectable index images for the first image channel and one or more on / detectable index images for other image channels (e.g., a green channel).
[0092] In other implementations, on / detectable index images are selected for each channel such that an index image for a particular image channel is normalized using one or more on / detectable index images for that particular image channel only, rather than for a different image channel. For example, index image 604 for a first image channel can be normalized using on / detectable index image 602 for the first image channel (t-3 index sequencing cycle). Similarly, index image 614 for a second image channel can be normalized using on / detectable index image 612 for the second image channel (t-2 index sequencing cycle).
[0093] Target image normalization 7 illustrates one embodiment of base calling of target sequences and index sequences. The target sequences are derived from multiple samples and combined with index sequences to form target index sequences. Each index sequence is uniquely associated with a respective sample of the multiple samples. The target index sequences are pooled for sequencing during a sequencing run 702. The target sequences are sequenced during a target sequencing cycle of the sequencing run, and the index sequences are sequenced during an index sequencing cycle of the sequencing run.
[0094] The disclosed techniques normalize target images differently than they normalize index images: target images depict intensity radiation generated as a result of nucleotide incorporation into target sequences, and index images depict intensity radiation generated as a result of nucleotide incorporation into index sequences.
[0095] The disclosed technology preprocesses the target image 714 using a first normalization function 724 that generates a normalized version 734 of the target image 714 from the current target sequencing cycle based solely on the intensity values of the target image 714. The first normalization function 724 calculates a lower percentile of the intensity values of the target image 714 and an upper percentile of the intensity values of the target image 714. In the normalized version 734 of the target image 714, a first percentage of the normalized intensity values are below the lower percentile, a second percentage of the normalized intensity values are above the upper percentile, and a third percentage of the normalized intensity values are between the lower and upper percentiles.
[0096] To preprocess the index image 712, the disclosed technique uses a second normalization function 722 that generates a normalized version 732 of the index image 712 from the current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle.
[0097] The second normalization function 722 calculates a lower percentile of (i) the index image intensity values from one or more preceding index sequencing cycles, (ii) the index image intensity values from one or more subsequent index sequencing cycles, and (iii) the index image intensity values from the current index sequencing cycle, and an upper percentile of (i) the index image intensity values from one or more preceding index sequencing cycles, (ii) the index image intensity values from one or more subsequent index sequencing cycles, and (iii) the index image intensity values from the current index sequencing cycle. In a normalized version 732 of index image 712, a first percentage of the normalized intensity values are below the lower percentile, a second percentage of the normalized intensity values are above the upper percentile, and a third percentage of the normalized intensity values are between the lower and upper percentiles.
[0098] The disclosed technology generates targeted reads for the target sequence by processing a normalized version of the target image through a neural network-based base caller 430 and generating base calls for each of the targeted sequencing cycles.
[0099] The disclosed technology generates index reads for the index sequence by processing a normalized version of the index image through a neural network-based base caller 430 to generate base calls for each of the index sequencing cycles.
[0100] The disclosed technology performs demultiplexing 742 by classifying each target read of a target sequence as belonging to a particular sample among multiple samples based on the corresponding index read of an index sequence bound to the target sequence.
[0101] Augmentation 8 illustrates one embodiment of pre-processing using enhancement. An image intensifier 812 processes the index image 802 and the target image 804 using an enhancement function. In one embodiment, the image intensifier 812 multiplies the intensity values of the index image 802 and the target image 804 by a scaling factor and adds an offset value to the result of the multiplication. In another embodiment, the image intensifier 812 changes the contrast between the index image 802 and the target image 804. In yet another embodiment, the image intensifier 812 changes the focus of the index image 802 and the target image 804.
[0102] The image enhancer 812 is comprised of image enhancement logic that multiplies the intensity values of the image by a scaling factor and adds an offset value to the result of the multiplication operation. The image enhancer 812 may include (i) hardware modules, (ii) software modules running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which implements certain techniques described herein and the software modules are stored on a computer-readable storage medium (or multiple such media).
[0103] In one embodiment, augmentation of the index image 802 and the target image 804 is performed only during training of the neural network-based base coder, and not during inference.
[0104] The augmented index image 822 and the augmented target image 824 are processed through a neural network-based base caller 830 to generate index reads for the index sequence by generating base calls for each index sequencing cycle, and to generate target reads for the target sequence by generating base calls for each target sequencing cycle.
[0105] The disclosed technology performs demultiplexing 832 by classifying each target read of a target sequence as belonging to a particular sample among multiple samples based on the corresponding index read of an index sequence bound to the target sequence.
[0106] Example of preprocessing results 9 and 10 show pixel intensity histograms of the red and green images of two targeted sequencing cycles (cycles 1 and 151) of the first targeted read (read 1).
[0107] Figures 11, 12, 13, 14, 15, 16, 17, and 18 show pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 152, 153, 154, 155, 156, 157, 158, and 159) of the first index read (index read 1).
[0108] Figures 19, 20, 21, 22, 23, 24, 25, and 26 show pixel intensity histograms of red and green images for eight index sequencing cycles (cycles 160, 161, 162, 163, 164, 165, 166, and 167) of the second index read (index read 2).
[0109] Figures 27 and 28 show pixel intensity histograms of the red and green images of two targeted sequencing cycles (cycles 168 and 169) of the second targeted read (read 2).
[0110] So, Read 1 is followed by Index Read 1, followed by Index Read 2, followed by Read 2.
[0111] Here, each figure has two pixel intensity histograms, one for the red image (left) and the other for the green image (right), for a given target or index sequencing cycle. The x-axis of the pixel intensity histogram indicates pixel intensity. The y-axis of the pixel intensity histogram indicates pixel number or pixel density. Thus, for example, if an image has 10,000 pixels, the corresponding pixel intensity histogram indicates how frequently a particular pixel intensity is found in the image.
[0112] The legend points out the names of the seven different sequencing runs (e.g., A00240_0175, A00276_0125, A00675_0021, etc.) along with their corresponding color codes. The color codes convey how the pixel intensity distribution varies across different sequencing runs.
[0113] The series of pixel intensity histograms in Figures 9-28 show that the pixel intensity distribution across the target and index sequencing cycles does not vary significantly, meaning that the pixel intensity values can be blended to calculate normalization parameters with confidence that they do not deviate significantly from the appropriate values. Technical effect and performance results as objective indicators of inventiveness
[0114] The following discussion demonstrates that normalizing and augmenting index images improves the base calling accuracy of neural network-based base caller 430 for index sequences. In particular, the following performance results provide an objective indication of the inventive step of the disclosed technology, whereby base calling errors increase when neural network-based base caller 430 does not use the disclosed normalization and augmentation techniques compared to when neural network-based base caller 430 uses the disclosed normalization and augmentation techniques.
[0115] The graphs shown in Figures 29, 30, and 31 have four types of lines: a cyan line, a yellow line, a green line, and a black line.
[0116] The cyan line represents the index base calling performance of the neural network-based base caller 430 when the index images are not normalized ("DeepRTA (No Normalization)").
[0117] The yellow line represents the index base calling performance of the neural network-based base caller 430 when the index images are normalized ("DeepRTA(normalized)").
[0118] The green line represents the index base calling performance of neural network-based base calling 430 when the index image is augmented ("DeepRTA(Augmented)").
[0119] The black line represents the index base calling performance of Illumina's non-neural network-based base calling system, called Real Time Analysis ("RTA"). Further details regarding RTA can be found in U.S. Patent Application Publication No. 2012 / 0020537, entitled "DATA PROCESSING SYSTEM AND METHODS," filed January 13, 2011 (Attorney Docket No. ILLINC.174A), which is incorporated herein by reference.
[0120] RTA is known to have good base calling accuracy for index sequences and can therefore be used as a baseline for comparison.
[0121] In the graph, the x-axis represents the error rate, which is an index of base calling accuracy, and the y-axis represents the number of index sequencing cycles. The graph also shows two index reads, Read:1 and Read:2, each with seven index sequencing cycles.
[0122] Figure 29 shows that in a sequencing run using four index arrays to multiplex four samples, the index base calling performance of neural network-based base caller 430 degrades when the index image is not normalized (e.g., cyan line for index reads: 2).
[0123] The error rate is relatively low when the index image is normalized (yellow line) and enhanced (green line), as shown by the dotted rectangles. Furthermore, the error rate of the normalized and enhanced implementation lies along the error rate line of RTA.
[0124] Figure 30 shows that in a sequencing run using two index arrays to multiplex two samples, the index base calling performance of neural network-based base caller 430 degrades when the index image is not normalized (e.g., cyan line for index reads: 2).
[0125] The error rate is relatively low when the index image is normalized (yellow line) and enhanced (green line). Furthermore, the error rate of the normalized and enhanced implementation is in line with the error rate of RTA.
[0126] Figure 31 shows that in a sequencing run using a single index array to sequence a single sample, the index base calling performance of neural network-based base caller 430 degrades when the index image is not normalized (e.g., cyan line for index reads: 2).
[0127] The error rate is relatively low when the index image is normalized (yellow line) and enhanced (green line). Furthermore, the error rate of the normalized and enhanced implementation is in line with the error rate of RTA.
[0128] Base calling using target and index images 7 illustrates one embodiment of base calling of target sequences and index sequences. The target sequences are derived from multiple samples and combined with index sequences to form target index sequences. Each index sequence is uniquely associated with a respective sample of the multiple samples. The target index sequences are pooled for sequencing during a sequencing run 702. The target sequences are sequenced during a target sequencing cycle of the sequencing run, and the index sequences are sequenced during an index sequencing cycle of the sequencing run.
[0129] In another embodiment, the disclosed technique normalizes the target image and the index image in the same manner: the target image represents the intensity radiation generated as a result of nucleotide incorporation into the target sequence, and the index image represents the intensity radiation generated as a result of nucleotide incorporation into the index sequence.
[0130] To preprocess the index image 712, the disclosed technique uses a second normalization function 722 that generates a normalized version 732 of the index image 712 from the current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle.
[0131] The second normalization function 722 calculates a lower percentile of (i) the index image intensity values from one or more preceding index sequencing cycles, (ii) the index image intensity values from one or more subsequent index sequencing cycles, and (iii) the index image intensity values from the current index sequencing cycle, and an upper percentile of (i) the index image intensity values from one or more preceding index sequencing cycles, (ii) the index image intensity values from one or more subsequent index sequencing cycles, and (iii) the index image intensity values from the current index sequencing cycle. In a normalized version 732 of index image 712, a first percentage of the normalized intensity values are below the lower percentile, a second percentage of the normalized intensity values are above the upper percentile, and a third percentage of the normalized intensity values are between the lower and upper percentiles.
[0132] The disclosed technology also uses a second normalization function 722 to preprocess the target image 714, which generates a normalized version 732 of the target image 714 from the current target sequencing cycle based on (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle.
[0133] The second normalization function 722 calculates a lower percentile of (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle, and an upper percentile of (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle. In the normalized version 732 of the target image 714, a first percentage of the normalized intensity values are below the lower percentile, a second percentage of the normalized intensity values are above the upper percentile, and a third percentage of the normalized intensity values are between the lower and upper percentiles.
[0134] In one embodiment, the normalization across target sequencing cycles also includes normalization across image channels within the image data of the target sequencing cycles. For example, consider three target sequencing cycles, namely, a first target sequencing cycle, a second target sequencing cycle, and a third target sequencing cycle. Also, consider a case where each of the first, second, and third target sequencing cycles has two target images: a first target image (e.g., red target image) in a first image channel (e.g., red channel) and a second target image (e.g., green target image) in a second image channel (e.g., green channel). The red target image from the second target sequencing cycle is normalized based on (i) the intensity values of the red and green images from the first target sequencing cycle, (ii) the intensity values of the red and green images from the third target sequencing cycle, and (iii) the intensity values of the red and green images from the second target sequencing cycle. The green target image from the second target sequencing cycle is normalized based on (i) the intensity values of the red and green images from the first target sequencing cycle, (ii) the intensity values of the red and green images from the third target sequencing cycle, and (iii) the intensity values of the red and green images from the second target sequencing cycle.
[0135] The disclosed technology generates targeted reads for the target sequence by processing a normalized version of the target image through a neural network-based base caller 430 and generating base calls for each of the targeted sequencing cycles.
[0136] The disclosed technology generates index reads for the index sequence by processing a normalized version of the index image through a neural network-based base caller 430 to generate base calls for each of the index sequencing cycles.
[0137] In one embodiment, preprocessing of the target and index images using the second regularization function 722 is performed during training and inference of the neural network-based base caller.
[0138] The disclosed technology performs demultiplexing 742 by classifying each target read of a target sequence as belonging to a particular sample among multiple samples based on the corresponding index read of an index sequence bound to the target sequence.
[0139] (Computer Systems) 32 illustrates a computer system 3200 that can be used to implement the disclosed techniques. Computer system 3200 includes at least one central processing unit (CPU) 3272 that communicates with a number of peripheral devices via a bus subsystem 3255. These peripheral devices may include, for example, a storage subsystem 3210, including memory devices and a file storage subsystem 3236, a user interface input device 3238, a user interface output device 3276, and a network interface subsystem 3274. The input and output devices enable user interaction with computer system 3200. Network interface subsystem 3274 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0140] In one embodiment, the percentile calculator 302, the image normalizer 354, and the neural network-based base coder 430 are communicatively linked to the storage subsystem 3210 and the user interface input device 3238.
[0141] The user interface input devices 3238 may include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and ways of inputting information into the computer system 3200.
[0142] The user interface output devices 3276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 3200 to a user or to another machine or computer system.
[0143] The storage subsystem 3210 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the deep learning processor 3278.
[0144] The deep learning processor 3278 may be a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRAs). The deep learning processor 3278 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of deep learning processors 3278 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX32 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu's DPI™, ARM's DynamicIQ™, IBM's TrueNorth™, and others.
[0145] The memory subsystem 3222 used in the storage subsystem 3210 may include multiple memories, including a main random access memory (RAM) 3232 for storing instructions and data during program execution, and a read-only memory (ROM) 3234 in which fixed instructions are stored. The file storage subsystem 3236 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 3236 in the storage subsystem 3210 or in another machine accessible by the processor.
[0146] Bus subsystem 3255 provides a mechanism for allowing the various components and subsystems of computer system 3200 to communicate with each other as intended. Although bus subsystem 3255 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0147] The computer system 3200 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the varying nature of computers and networks, the description of computer system 3200 shown in Figure 32 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 3200 can have more or fewer components than the computer system shown in Figure 32.
[0148] Specific Implementations Various embodiments of artificial intelligence-based base calling of index sequences are described. One or more features of an embodiment can be combined with a base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following implementations.
[0149] In one embodiment, an artificial intelligence-based method for base calling an index sequence is disclosed, the method comprising accessing an index image generated for the index sequence during an index sequencing cycle of a sequencing run, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence during the sequencing run.
[0150] The method includes preprocessing the index image using a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) intensity values of the index image from one or more preceding index sequencing cycles, (ii) intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) intensity values of the index image from the current index sequencing cycle.
[0151] The method further includes generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles.
[0152] The methods described in this and other sections of the disclosed technology may include one or more of the following features and / or characteristics described in connection with additional disclosed methods. For purposes of brevity, combinations of features disclosed in this application are not individually listed and are not repeated for each base set of features. The reader will understand how the features identified in these embodiments can be readily combined with the sets of basic features specified in other embodiments.
[0153] In one embodiment, the normalization function calculates a lower percentile of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, and an upper percentile of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, such that in the normalized version of the index image, a first percentage of normalized intensity values are below the lower percentile, a second percentage of normalized intensity values are above the upper percentile, and a third percentage of normalized intensity values are between the lower and upper percentiles.
[0154] In one embodiment, the nucleotides represented by the index images from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle are collectively cumulatively more diverse than the nucleotides represented by the index image from the current index sequencing cycle alone. In some embodiments, at least one of the index images from the preceding index sequencing cycle and the subsequent index sequencing cycle represents one or more nucleotides in a detectable signal state.
[0155] In one embodiment, the nucleotides represented by the index image from the current index sequencing cycle are a low complexity pattern in which some of the four bases A, C, T, and G are represented at a frequency of less than 15%, 10%, or 5% of all nucleotides.
[0156] In one embodiment, the nucleotides represented by the index images from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle collectively cumulatively form a high complexity pattern in which each of the four bases A, C, T, and G is represented at a frequency of at least 20%, 25%, or 30% of all nucleotides.
[0157] In one embodiment, the method includes preprocessing the index images using a regularization function during training and inference of the neural network-based base caller.
[0158] In one embodiment, the method includes preprocessing the index image using an augmentation function that generates an augmented version of the index image by multiplying intensity values of the index image by a scaling factor and adding an offset value to the result of the multiplication. The method further includes generating index reads for the index sequence by processing the augmented version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles.
[0159] In one embodiment, the method includes preprocessing the index images with the augmentation function only during training, not during inference, of the neural network-based base call.
[0160] In one embodiment, the method includes preprocessing the index image using a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) intensity values of the index image from one or more non-current index sequencing cycles and (ii) intensity values of the index image from the current index sequencing cycle. In some embodiments, the non-current index sequencing cycle comprises an initial index sequencing cycle of sequencing. In other embodiments, the non-current index sequencing cycle comprises an intermediate index sequencing cycle of sequencing. In some other embodiments, the non-current index sequencing cycle comprises a terminal index sequencing cycle of sequencing. In still other embodiments, the non-current index sequencing cycle comprises a combination of an initial index sequencing cycle, an intermediate index sequencing cycle, and a terminal index sequencing cycle.
[0161] In one embodiment, at least one index image from a non-current index sequencing cycle shows one or more nucleotides in a detectable signal state.
[0162] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0163] 34 is an embodiment of a flowchart of an artificial intelligence-based method for base calling analytes in an index sequencing cycle of a sequencing run. The method includes, in operation 3402, preprocessing index images generated during an index sequencing cycle using a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) intensity values of the index image from one or more preceding index sequencing cycles, (ii) intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) intensity values of the index image from the current index sequencing cycle.
[0164] The method includes, in operation 3412, for a particular analyte being base called in the current index sequencing cycle, extracting index image patches from normalized versions of index images from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle, such that each normalized index image patch shows intensity emissions of the particular analyte, adjacent analytes, and their surrounding background generated as a result of nucleotide incorporation in the corresponding index sequences of the particular analyte and several adjacent analytes during the current index sequencing cycle.
[0165] The method further includes, at operation 3422, convolving the normalized index image patch through a convolutional neural network to generate a convolved representation.
[0166] The method further includes, at operation 3432, base calling the particular analyte in the current index sequencing cycle based on the convoluted representation.
[0167] Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As noted above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features identified in these embodiments can be readily combined with the basic feature sets specified in other embodiments. Other embodiments of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another embodiment of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0168] 35 is an embodiment of a flowchart of an artificial intelligence-based method for base calling target sequences and index sequences. The target sequences are derived from multiple samples and are combined with index sequences to form target index sequences. Each index sequence is uniquely associated with a respective sample of the multiple samples. The target index sequences are pooled for sequencing during a sequencing run. The target sequences are sequenced during a target sequencing cycle of the sequencing run, and the index sequences are sequenced during an index sequencing cycle of the sequencing run.
[0169] The method includes accessing a target image generated for the target sequence during a target sequencing cycle, in operation 3502. The target image shows intensity radiation generated as a result of nucleotide incorporation into the target sequence.
[0170] The method further includes, at operation 3512, preprocessing the target image using a first normalization function that generates a normalized version of the target image from the current target sequencing cycle based solely on the intensity values of the target image.
[0171] The method further includes, in operation 3522, generating target reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles.
[0172] The method further includes accessing an index image generated for the index sequence during the index sequencing cycle, in operation 3532. The index image shows intensity radiation generated as a result of nucleotide incorporation into the index sequence.
[0173] The method further includes, at operation 3542, preprocessing the index image using a second normalization function that generates a normalized version of the index image from the current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle.
[0174] The method further includes, at operation 3552, generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles.
[0175] The method further includes, at operation 3562, classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence.
[0176] Each of the features described in the specific embodiment sections for other embodiments applies equally to this embodiment. As noted above, all other features are not repeated here and should be repeated by reference. The reader will understand how the features identified in these embodiments can be easily combined with the basic feature sets specified in other embodiments.
[0177] In one embodiment, the first normalization function calculates a lower percentile of intensity values of the target image and an upper percentile of intensity values of the target image such that in a normalized version of the target image, a first percentage of normalized intensity values are below the lower percentile, a second percentage of normalized intensity values are above the upper percentile, and a third percentage of normalized intensity values are between the lower and upper percentiles.
[0178] In one embodiment, the second normalization function calculates a lower percentile of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, and an upper percentile of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, such that in the normalized version of the index image, a first percentage of normalized intensity values are below the lower percentile, a second percentage of normalized intensity values are above the upper percentile, and a third percentage of normalized intensity values are between the lower and upper percentiles.
[0179] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0180] The embodiments disclosed herein may be embodied as a method, apparatus, system, or article of manufacture using standard programming or engineering techniques to generate software, firmware, hardware, or any combination thereof. As used herein, the term "article of manufacture" refers to code or logic implemented in hardware or computer-readable media, such as optical storage devices, as well as volatile or non-volatile memory devices. Such hardware includes, but is not limited to, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microprocessors, or other similar processing devices. In certain embodiments, the information or algorithms described herein reside in non-transitory storage media.
[0181] One or more embodiments of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer-usable program code for performing the illustrated method steps. Furthermore, one or more embodiments of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor, coupled to the memory, operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments of the disclosed technology, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) hardware modules, (ii) software modules executing on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which (i)-(iii) implements particular technology described herein, and the software modules are stored on a computer-readable storage medium (or multiple such media).
[0182] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative position. An individual analyte can contain one or more molecules of a particular type. For example, an analyte can contain a single target nucleic acid molecule having a particular sequence, or an analyte can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.
[0183] Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.
[0184] The terms "analyte," "nucleic acid," "nucleic acid molecule," and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) as provided herein for certain types of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. Nucleic acids in certain embodiments include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester chains, or deoxyribonucleic acid (DNA), e.g., single- and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester or other linkages, such as ribonucleic acid (RNA), including single- and double-stranded RNA, messenger (mRNA), copy or complementary RNA (cRNA), alternatively spliced mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piwi RNA (piRNA), or any form of synthetic or modified RNA. Nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, nucleic acids may carry one or more detectable labels, as described elsewhere herein.
[0185] The terms "specimen," "cluster," "nucleic acid cluster," "nucleic acid colony," and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement attached to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster can be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster can have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a labeled moiety. Corresponding positions can also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as uracil and thymine.
[0186] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies can optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence can be present in a single nucleic acid molecule, such as concatemers generated using rolling circle amplification procedures.
[0187] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can be substantially circular, polyhedral, donut-shaped, or ring-shaped. The diameter of the nucleic acid cluster can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intermediate diameter. In certain embodiments, the diameter of the nucleic acid cluster is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid cluster can be affected by numerous parameters, including, but not limited to, the number of amplification cycles performed in producing the cluster, the length of the nucleic acid template, or the density of primers attached to the surface on which the cluster is formed. The density of the nucleic acid clusters can typically be designed to be in the range of 0.1 / mm, 1 / mm, 10 / mm, 100 / mm, 1,000 / mm, 10,000 / mm to 100,000 / mm. The present invention, in part, further contemplates higher density nucleic acid clusters, for example, 100,000 / mm to 1,000,000 / mm, and 1,000,000 / mm to 10,000,000 / mm.
[0188] As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupy a physical region on a sample. For example, an analyte can be a parcel of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a single pixel.
[0189] The distance between specimens can be described in any number of ways. In some embodiments, the distance between specimens can be described as from the center of one specimen to the center of another specimen. In other embodiments, the distance can be described from the edge of one specimen to the edge of another specimen, or between the outermost identifiable points of each specimen. The edge of the specimen can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the specimen. In other embodiments, the distance can be described with respect to a fixed point on the specimen, or an image of the specimen.
[0190] item The following items are part of this disclosure. Index Read 1. An artificial intelligence-based method for base calling index sequences, comprising: accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence during the sequencing run; The index image is normalized by a normalization function, where a normalized version of the index image from the current index sequencing cycle is (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A method comprising: 2. The normalization function is (i) the lower percentile of the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; and determining the top percentiles of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; In the normalized version of the index image, a first proportion of normalized intensity values below the lower percentile; a second percentage of normalized intensity values are above the upper percentile; a third proportion of normalized intensity values are between the lower and upper percentiles; Item 1. The artificial intelligence-based method according to item 1, wherein the 3. The nucleotides represented by the index images from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle are, in total, is cumulatively more diverse than the nucleotides represented by the index image alone from the current index sequencing cycle, Item 1. The artificial intelligence-based method according to item 1. 4. The artificial intelligence-based method of item 3, wherein at least one index image from the preceding index sequencing cycle and the subsequent index sequencing cycle exhibits one or more nucleotides in a detectable signal state. 5. The artificial intelligence-based method of item 3, wherein the nucleotides represented by the index image from the current index sequencing cycle are a low complexity pattern in which some of the four bases A, C, T, and G are represented at a frequency of less than 15%, 10%, or 5% of all nucleotides. 6. The artificial intelligence-based method of item 5, wherein the nucleotides represented by the index images from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle collectively cumulatively form a high complexity pattern in which each of the four bases A, C, T, and G is represented at a frequency of at least 20%, 25%, or 30% of all nucleotides. 7. Item 1. The artificial intelligence-based method of item 1, further comprising pre-processing the index images using a regularization function during training and inference of the neural network-based base caller. 8. preprocessing the index image using an augmentation function that generates an augmented version of the index image by multiplying intensity values of the index image by a scaling factor and adding an offset value to the result of the multiplication; generating index reads for the index sequence by processing the augmented version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; Item 1. The artificial intelligence-based method of item 1, further comprising: 9. Item 9. The artificial intelligence-based method of item 8, further comprising preprocessing the index images with the augmentation function only during training, and not during inference, of the neural network-based base code. 10. The index image, (i) intensity values of index images from one or more non-current index sequencing cycles; (ii) the intensity value of the index image from the current index sequencing cycle; and 2. The artificial intelligence-based method of claim 1, further comprising preprocessing using a normalization function to generate a normalized version of the index image from the current index sequencing cycle based on: 11. The artificial intelligence-based method of item 10, wherein the non-current index sequencing cycle comprises an initial index sequencing cycle of sequencing. 12. The artificial intelligence-based method of item 10, wherein the non-current index sequencing cycle comprises an intermediate index sequencing cycle of sequencing. 13. The artificial intelligence-based method of item 10, wherein the non-current index sequencing cycle comprises a terminal index sequencing cycle of sequencing. 14. The artificial intelligence-based method of item 13, wherein the non-current index sequencing cycle comprises a combination of an initial index sequencing cycle, an intermediate index sequencing cycle, and a terminal index sequencing cycle. 15. The artificial intelligence-based method of item 10, wherein at least one index image from a non-current index sequencing cycle exhibits one or more nucleotides in a detectable signal state. 16. An artificial intelligence-based method for base calling a sample in an index sequencing cycle of a sequencing run, comprising: The index images generated during the index sequencing cycle are normalized by a normalization function, which normalizes a normalized version of the index image from the current index sequencing cycle: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on For a given specimen being base-called in the current index sequencing cycle, The index image patch is extracted from normalized versions of the index image from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle. Extracting each normalized index image patch to represent intensity emissions of a particular analyte, adjacent analytes, and their surrounding background generated as a result of nucleotide incorporation in the corresponding index sequence of the particular analyte and several adjacent analytes during the current index sequencing cycle; convolving the normalized index image patch through a convolutional neural network to generate a convolutional representation; base calling a particular analyte in the current index sequencing cycle based on the convolutional representation; A method comprising: 17. An artificial intelligence-based method for base calling target sequences and index sequences, wherein the target sequences are derived from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run, the method comprising: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; preprocessing the target image with a first normalization function that generates a normalized version of the target image from a current target sequencing cycle based solely on intensity values of the target image; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; The index image is normalized by a second normalization function, the normalized version of the index image from the current index sequencing cycle being: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a second normalization function, generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A method comprising: 18. The first normalization function is the lower percentile of the intensity values of the target image; the top percentile of the intensity values of the target image, In the normalized version of the target image, a first proportion of normalized intensity values below the lower percentile; a second percentage of normalized intensity values are above the upper percentile; a third proportion of normalized intensity values are between the lower and upper percentiles; Item 18. The artificial intelligence-based method according to item 17, wherein the method calculates: 19. The second normalization function is (i) the lower percentile of the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; and determining the top percentiles of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; In the normalized version of the index image, a first proportion of normalized intensity values below the lower percentile; a second percentage of normalized intensity values are above the upper percentile; a third proportion of normalized intensity values are between the lower and upper percentiles; Item 18. The artificial intelligence-based method according to item 17, wherein the method calculates: Index read and regular read 20. An artificial intelligence-based method for base calling target sequences and index sequences, wherein the target sequences are derived from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run, the method comprising: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A method comprising: 21. The normalization function is (i) the lower percentile of the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle; and determining the top percentiles of (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle; In the normalized version of the target image, a first proportion of normalized intensity values below the lower percentile; a second percentage of normalized intensity values are above the upper percentile; a third proportion of normalized intensity values are between the lower and upper percentiles; 21. The artificial intelligence-based method according to item 20, wherein the method calculates: 22. The normalization function is (i) the lower percentile of the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; and determining the top percentiles of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; In the normalized version of the index image, a first proportion of normalized intensity values below the lower percentile; a second percentage of normalized intensity values are above the upper percentile; a third proportion of normalized intensity values are between the lower and upper percentiles; 21. The artificial intelligence-based method according to item 20, wherein the method calculates: twenty three. 21. The artificial intelligence-based method of claim 20, further comprising preprocessing the target image and index image using a normalization function during training and inference of the neural network-based base caller. twenty four. preprocessing the target image with an augmentation function that generates an augmented version of the target image by multiplying intensity values of the target image by a scaling factor and adding an offset value to the result of the multiplication; generating targeted reads for the target sequence by processing the augmented version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; 21. The artificial intelligence-based method of claim 20, further comprising: twenty five. preprocessing the index image using an augmentation function that generates an augmented version of the index image by multiplying intensity values of the index image by a scaling factor and adding an offset value to the result of the multiplication; generating index reads for the index sequence by processing the augmented version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; 21. The artificial intelligence-based method of claim 20, further comprising: 26. 21. The artificial intelligence-based method of claim 20, further comprising preprocessing the target image and index image using the augmentation function only during training, and not during inference of the neural network-based base code. 27. An artificial intelligence-based method for base calling sequences, comprising: accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles.
[0191] Other implementations of the above-described methods may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods. 28. An artificial intelligence-based method for base calling sequences, comprising: accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; generating targeted reads for the target sequence by processing the target image through a neural network-based base caller and generating base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A method comprising: 29. A system including one or more processors coupled to a memory, wherein computer instructions for base calling an index sequence are loaded into the memory, the instructions, when executed on the processor, perform: accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence during the sequencing run; The index image is normalized by a normalization function, where a normalized version of the index image from the current index sequencing cycle is (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A system that performs operations including: 30. The system according to item 29, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 31. A system comprising one or more processors coupled to a memory, wherein the memory is loaded with computer instructions for base calling analytes in an index sequencing cycle of a sequencing run, the instructions, when executed on the processor, perform: The index images generated during the index sequencing cycle are normalized by a normalization function, which normalizes a normalized version of the index image from the current index sequencing cycle: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on For a given specimen being base-called in the current index sequencing cycle, The index image patch is extracted from normalized versions of the index image from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle. Extracting each normalized index image patch to represent intensity emissions of a particular analyte, adjacent analytes, and their surrounding background generated as a result of nucleotide incorporation in the corresponding index sequence of the particular analyte and several adjacent analytes during the current index sequencing cycle; convolving the normalized index image patch through a convolutional neural network to generate a convolutional representation; base calling a particular analyte in the current index sequencing cycle based on the convolutional representation; A system that performs operations including: 32. The system according to item 31, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 33. A system including one or more processors coupled to a memory, wherein the memory is loaded with computer instructions for base calling target sequences and index sequences, the target sequences originating from a plurality of samples and combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run. In the system, the instructions, when executed on the processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; preprocessing the target image with a first normalization function that generates a normalized version of the target image from a current target sequencing cycle based solely on intensity values of the target image; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; The index image is normalized by a second normalization function, the normalized version of the index image from the current index sequencing cycle being: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a second normalization function, generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A system that performs operations including: 34. The system according to item 33, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 35. A system including one or more processors coupled to a memory, wherein the memory is loaded with computer instructions for base calling target sequences and index sequences, the target sequences originating from a plurality of samples and combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run. In the system, the instructions, when executed on the processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A system that performs operations including: 36. The system according to item 35, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 37. A system comprising one or more processors coupled to a memory, wherein computer instructions for base calling a sequence are loaded into the memory, the instructions, when executed on the processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A system that performs operations including: 38. The system according to item 37, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 39. A system comprising one or more processors coupled to a memory, wherein computer instructions for base calling a sequence are loaded into the memory, the instructions, when executed on the processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; generating targeted reads for the target sequence by processing the target image through a neural network-based base caller and generating base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A system that performs operations including: 40. The system according to item 39, implementing each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 41. A non-transitory computer-readable storage medium having computer program instructions for base calling an index sequence, the instructions, when executed on a processor, accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence during the sequencing run; The index image is normalized by a normalization function, where a normalized version of the index image from the current index sequencing cycle is (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A non-transitory computer-readable storage medium for performing a method comprising: 42. The non-transitory computer-readable storage medium according to item 41, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 43. A non-transitory computer-readable storage medium having computer program instructions provided thereon for base calling analytes in an index sequencing cycle of a sequencing run, the instructions, when executed on a processor, performing: The index images generated during the index sequencing cycle are normalized by a normalization function, which normalizes a normalized version of the index image from the current index sequencing cycle: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a normalization function, which is generated based on For a given specimen being base-called in the current index sequencing cycle, The index image patch is extracted from normalized versions of the index image from the current index sequencing cycle, the preceding index sequencing cycle, and the subsequent index sequencing cycle. Extracting each normalized index image patch to represent intensity emissions of a particular analyte, adjacent analytes, and their surrounding background generated as a result of nucleotide incorporation in the corresponding index sequence of the particular analyte and several adjacent analytes during the current index sequencing cycle; convolving the normalized index image patch through a convolutional neural network to generate a convolutional representation; base calling a particular analyte in the current index sequencing cycle based on the convolutional representation; A non-transitory computer-readable storage medium for performing a method comprising: 44. A non-transitory computer-readable storage medium according to item 43, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 45. A non-transitory computer-readable storage medium having computer program instructions provided thereon for base calling target sequences and index sequences, wherein the target sequences are from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run. In the non-transitory computer-readable storage medium, the instructions, when executed on a processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; preprocessing the target image with a first normalization function that generates a normalized version of the target image from a current target sequencing cycle based solely on intensity values of the target image; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; The index image is normalized by a second normalization function, the normalized version of the index image from the current index sequencing cycle being: (i) intensity values of the index image from one or more preceding index sequencing cycles; (ii) intensity values of the index image from one or more subsequent index sequencing cycles; and (iii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using a second normalization function, generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A non-transitory computer-readable storage medium for performing a method comprising: 46. The non-transitory computer-readable storage medium according to item 45, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 47. A non-transitory computer-readable storage medium having computer program instructions provided thereon for base calling target sequences and index sequences, wherein the target sequences are from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run. In the non-transitory computer-readable storage medium, the instructions, when executed on a processor, perform: accessing a target image generated for the target sequence during a target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read of the target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read of the index sequence bound to the target sequence; A non-transitory computer-readable storage medium for performing a method comprising: 48. The non-transitory computer-readable storage medium according to item 47, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 49. A non-transitory computer-readable storage medium having computer program instructions for base calling a sequence, the instructions, when executed on a processor, accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; pre-processing the target image with a normalization function that generates a normalized version of the target image from a current target sequencing cycle based on (i) intensity values of the target image from one or more preceding target sequencing cycles, (ii) intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) intensity values of the target image from the current target sequencing cycle; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; pre-processing the index image with a normalization function that generates a normalized version of the index image from a current index sequencing cycle based on (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity value of the index image from the current index sequencing cycle; generating targeted reads for the target sequence by processing the normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the normalized version of the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A non-transitory computer-readable storage medium for performing a method comprising: 50. A non-transitory computer-readable storage medium according to item 49, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. 51. A non-transitory computer-readable storage medium having computer program instructions for base calling a sequence, the instructions, when executed on a processor, accessing a target image generated for the target sequence during a target sequencing cycle of the sequencing run, the target image showing intensity radiation generated as a result of nucleotide incorporation into the target sequence; accessing an index image generated for the index sequence during an index sequencing cycle of the sequencing run, the index image showing intensity radiation generated as a result of nucleotide incorporation into the index sequence during the sequencing run; generating targeted reads for the target sequence by processing the target image through a neural network-based base caller and generating base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing the index image through a neural network-based base caller to generate base calls for each of the index sequencing cycles; A non-transitory computer-readable storage medium for performing a method comprising: 52. A non-transitory computer-readable storage medium according to item 51, which implements each of the items ultimately dependent on items 1, 16, 17, 20, and 27. [Explanation of symbols]
[0192] 102 Indexed Library 104 Pooling 106 Sequencing 108 Demultiplexing 110 Alignment 116 Output Files 202 Targeted Lead 204 Index Read 212 target primers 222 target sequence 224 Index Primer 232 Index Array 302 Percentile Calculation Section 322 Index image of the first image channel 324 Index image of the first image channel 326 Index image of the first image channel 332 Index image of the second image channel 334 Index image of the second image channel 336 Index image of the second image channel 344 Normalization 354 Image Normalization Unit 364 Normalized index image in the first image channel 374 Index image normalized in the second image channel 402 index image normalized in the first image channel 404 Index image normalized in the first image channel 406 Normalized index image in the first image channel 412 Normalized index image in the second image channel 414 Index image normalized in the second image channel 416 Normalized index image in the second image channel 424 Patch Extraction Process 426 Input image data 430 Neural Network Based Base Cola 432 base calls 502 Initial Index Sequencing Cycles 512 intermediate index sequencing cycles 522 Image Selection Section 532 End-Index Sequencing Cycles 602 index images 604 index images 612 index images 614 index images 632 normalized images 702 sequencing runs 712 index images 714 Target Image 722 Second normalization function 724 First normalization function 732 normalized index images 734 Normalized Target Image 742 Demultiplexing 802 index images 804 Target Image 812 Image Intensifier 822 Augmented Index Image 824 Enhanced Target Image 830 Neural Network Based Base Cola 832 Demultiplexing 3200 Computer System 3210 Storage Subsystem 3222 Memory Subsystem 3232 Main Random Access Memory (RAM) 3234 dedicated memory (ROM) 3236 File Storage Subsystem 3238 User Interface Input Devices 3255 Bus Subsystem 3272 Central Processing Unit (CPU) 3274 Network Interface Subsystem 3276 User Interface Output Device 3278 Deep Learning Processor
Claims
1. 1. A computer-implemented, artificial intelligence-based method for base calling index sequences, comprising: accessing an index image for an index sequence during an index sequencing cycle of a sequencing run, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence of a cluster during the sequencing run; The index image is subjected to a normalization function, wherein a normalized version of the index image from the current index sequencing cycle is calculating differential percentiles from (i) index image intensity values from one or more preceding index sequencing cycles, (ii) index image intensity values from one or more subsequent index sequencing cycles, and (iii) index image intensity values from the current index sequencing cycle; generating a normalized version of the index image consisting of intensity values normalized according to the different percentiles from (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; preprocessing using a normalization function, which is generated based on generating index reads for the index sequence by processing a normalized version of the index image through a neural network-based base caller trained on a target image to generate base calls for each of the index sequencing cycles; A computer-implemented artificial intelligence-based method, including:
2. The normalization function is (i) the lower percentile of the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; and determining the top percentiles of (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle; In the normalized version of the index image, a first percentage of normalized intensity values are below the lower percentile; a second percentage of the normalized intensity values are above the upper percentile; a third percentage of the normalized intensity values are between the lower percentile and the upper percentile; the normalization function scales the normalized intensity values within an upper and lower percentile range; 10. The computer-implemented artificial intelligence based method of claim 1, wherein the method calculates:
3. 3. The computer-implemented artificial intelligence-based method of claim 1, wherein the nucleotides indicated by the index images from the current index sequencing cycle, the one or more preceding index sequencing cycles, and the one or more subsequent index sequencing cycles are collectively cumulatively more diverse than the nucleotides indicated by the index image from the current index sequencing cycle alone.
4. 4. The computer-implemented artificial intelligence-based method of claim 1, wherein at least one index image among the index images from the one or more preceding and subsequent index sequencing cycles exhibits one or more nucleotides in a detectable signal state.
5. 5. The computer-implemented artificial intelligence-based method of claim 3 or 4, wherein the nucleotides represented by the index image from the current index sequencing cycle are a low complexity pattern in which some of the four bases A, C, T, and G are represented at a frequency of less than 15%, 10%, or 5% of all the nucleotides.
6. 6. The computer-implemented artificial intelligence-based method of claim 3, wherein the nucleotides indicated by the index images from the current index sequencing cycle, the one or more preceding index sequencing cycles, and the one or more subsequent index sequencing cycles collectively cumulatively form a high complexity pattern in which each of the four bases A, C, T, and G is represented at a frequency of at least 20%, 25%, or 30% of all the nucleotides.
7. 7. The computer-implemented artificial intelligence-based method of claim 1, further comprising pre-processing the index images using the normalization function during training and inference of the neural network-based base caller.
8. preprocessing the index image using an augmentation function that generates an augmented version of the index image by multiplying intensity values of the index image by a scaling factor and adding an offset value; generating index reads for the index sequence by processing an augmented version of the index image through the neural network-based base caller to generate base calls for each of the index sequencing cycles; preprocessing the index images using the augmentation function only during training, not during inference, of the neural network-based base classifier; 8. The computer-implemented artificial intelligence-based method of claim 1, further comprising:
9. 10. The computer-implemented artificial intelligence-based method of claim 8, further comprising preprocessing the index images using the augmentation function only during training, and not during inference, of the neural network-based base caller.
10. The index image is (i) index image intensity values from one or more non-current index sequencing cycles; (ii) the intensity value of the index image from the current index sequencing cycle; and preprocessing using the normalization function to generate the normalized version of the index image from the current index sequencing cycle based on Further comprising:
10. The computer-implemented artificial intelligence-based method of claim 1, wherein at least one index image from the non-current index sequencing cycle exhibits one or more nucleotides in a detectable signal state.
11. 11. The computer-implemented artificial intelligence-based method of claim 10, wherein the one or more non-current index sequencing cycles comprise one or more initial index sequencing cycles of the sequencing run.
12. 12. The computer-implemented artificial intelligence-based method of claim 10 or 11, wherein the one or more non-current index sequencing cycles comprise an intermediate index sequencing cycle of the sequencing run.
13. 13. The computer-implemented artificial intelligence-based method of claim 10, wherein the one or more non-current index sequencing cycles comprise a terminal index sequencing cycle of the sequencing run.
14. 14. The computer-implemented artificial intelligence based method of claim 13, wherein the one or more non-current index sequencing cycles comprise a combination of an initial index sequencing cycle, an intermediate index sequencing cycle, and a terminal index sequencing cycle.
15. 15. The computer-implemented artificial intelligence-based method of any one of claims 10 to 14, wherein at least one index image from the one or more non-current index sequencing cycles shows one or more nucleotides in a detectable signal state.
16. 1. A computer-implemented, artificial intelligence-based method for base calling a sample in an index sequencing cycle of a sequencing run, comprising: During the index sequencing cycle, the index image is normalized by a normalization function, such that a normalized version of the index image from the current index sequencing cycle is calculating differential percentiles from (i) index image intensity values from one or more preceding index sequencing cycles, (ii) index image intensity values from one or more subsequent index sequencing cycles, and (iii) index image intensity values from the current index sequencing cycle; generating a normalized version of the index image from (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, the normalized version comprising intensity values normalized according to the different percentiles; preprocessing using a normalization function, which is generated based on For the particular specimen being base called in the current index sequencing cycle, generating a normalized index image patch from normalized versions of the index image from the current index sequencing cycle, the one or more previous index sequencing cycles, and the one or more subsequent index sequencing cycles; extracting each normalized index image patch to represent intensity emissions of the specific analyte, the one or more adjacent analytes, and their surrounding background generated as a result of nucleotide incorporation in the corresponding index sequences of the specific analyte and several adjacent analytes during the current index sequencing cycle; convolving the normalized index image patch through a convolutional neural network trained on a target image to generate a convolved representation; base calling the particular analyte in the current index sequencing cycle based on the convoluted representation; A computer-implemented artificial intelligence-based method, including:
17. 1. A computer-implemented, artificial intelligence-based method for base calling target sequences and index sequences, wherein the target sequences originate from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run, the method comprising: accessing a target image for the target sequence during the target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into clusters of the target sequence; preprocessing the target image using a first normalization function that generates a normalized version of the target image from a current target sequencing cycle based on intensity values of the target image; generating targeted reads for the target sequence by processing a normalized version of the target image through a neural network-based base caller to generate base calls for each of the targeted sequencing cycles; accessing an index image generated for the index sequence during the index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; The index image is subjected to a second normalization function, wherein a normalized version of the index image from the current index sequencing cycle is calculating a differential percentile from (i) intensity values of an index image from one or more preceding index sequencing cycles, (ii) intensity values of an index image from one or more subsequent index sequencing cycles, and (iii) intensity values of an index image from the current index sequencing cycle; generating a normalized version of the index image from (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, the normalized version comprising intensity values normalized according to the different percentiles; preprocessing using a second normalization function generated based on generating index reads for the index sequence by processing a normalized version of the index image through the neural network-based base caller trained on a target image to generate base calls for each of the index sequencing cycles; classifying each target read for a target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read for an index sequence bound to the target sequence; A computer-implemented artificial intelligence-based method, including:
18. The first normalization function is: a bottom percentile of the intensity values of the target image; the upper percentile of the intensity values of the target image; In the normalized version of the target image, a first percentage of normalized intensity values are below the lower percentile; a second percentage of the normalized intensity values are above the upper percentile; a third percentage of the normalized intensity values fall between the lower percentile and the upper percentile.
20. The computer-implemented artificial intelligence based method of claim 17, wherein the method calculates:
19. 1. A computer-implemented, artificial intelligence-based method for base calling target sequences and index sequences, wherein the target sequences originate from a plurality of samples and are combined with index sequences to form target index sequences, each index sequence being uniquely associated with a respective sample of the plurality of samples, the target index sequences being pooled for sequencing during a sequencing run, the target sequences being sequenced during a target sequencing cycle of the sequencing run, and the index sequences being sequenced during an index sequencing cycle of the sequencing run, the method comprising: accessing a target image for the target sequence during the target sequencing cycle, the target image showing intensity radiation generated as a result of nucleotide incorporation into clusters of the target sequence; The target image is normalized by a normalization function, where a normalized version of the target image from the current target sequencing cycle is calculating differential percentiles from (i) intensity values of a target image from one or more preceding target sequencing cycles, (ii) intensity values of a target image from one or more subsequent target sequencing cycles, and (iii) intensity values of a target image from the current target sequencing cycle; generating a normalized version of the target image from (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle, the normalized version comprising intensity values normalized according to the different percentiles; preprocessing using a normalization function, which is generated based on accessing an index image generated for the index sequence during the index sequencing cycle, the index image showing intensity radiation generated as a result of incorporation of nucleotides into the index sequence; The index image is normalized by a normalization function, where a normalized version of the index image from the current index sequencing cycle is calculating differential percentiles from (i) index image intensity values from one or more preceding index sequencing cycles, (ii) index image intensity values from one or more subsequent index sequencing cycles, and (iii) index image intensity values from the current index sequencing cycle; generating a normalized version of the index image from (i) the intensity values of the index image from one or more preceding index sequencing cycles, (ii) the intensity values of the index image from one or more subsequent index sequencing cycles, and (iii) the intensity values of the index image from the current index sequencing cycle, the normalized version comprising intensity values normalized according to the different percentiles; preprocessing using a normalization function, which is generated based on generating targeted reads for the target sequence by processing a normalized version of the target image through a neural network-based base caller trained on the target image to generate base calls for each of the targeted sequencing cycles; generating index reads for the index sequence by processing a normalized version of the index image through the neural network-based base caller to generate base calls for each of the index sequencing cycles; classifying each target read for a target sequence as belonging to a particular sample among the plurality of samples based on a corresponding index read for an index sequence bound to the target sequence; A computer-implemented artificial intelligence-based method, including:
20. The normalization function is (i) the lower percentile of the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle; and determining the top percentiles of (i) the intensity values of the target image from one or more preceding target sequencing cycles, (ii) the intensity values of the target image from one or more subsequent target sequencing cycles, and (iii) the intensity values of the target image from the current target sequencing cycle; In the normalized version of the target image, a first percentage of normalized intensity values are below the lower percentile; a second percentage of the normalized intensity values are above the upper percentile; a third percentage of the normalized intensity values fall between the lower percentile and the upper percentile.
20. The computer-implemented artificial intelligence based method of claim 19, wherein the method calculates:
Citation Information
Patent Citations
Data processing system and methods
US20120020537A1
Data processing system and methods
US8965076B2
Phasing correction
WO2018129314A1
Single light source, two-optical channel sequencing
WO2018165099A1