Artificial intelligence-based sequence metadata generation

Deep neural networks are used to enhance nucleic acid sequencing by resolving overlapping clusters, improving throughput and quality of sequence data, addressing limitations in existing sequencing technologies.

JP7767012B2Active Publication Date: 2025-11-11ILLUMINA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2020572715
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-21
Filing Date
2020-03-21
Publication Date
2025-11-11
Estimated Expiration
2040-03-21

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing technologies face challenges in resolving data from closely spaced or spatially overlapping nucleic acid clusters, leading to reduced throughput and quality of nucleic acid sequence information, which affects applications in genomics, diagnostics, and other biological research.

Method used

Employing deep neural networks, specifically convolutional neural networks, to analyze digital images of nucleic acid clusters, using techniques like subpixel base calling, regression models, and classification models to improve the resolution and accuracy of cluster metadata generation, thereby enhancing the throughput and quality of nucleic acid sequencing.

Benefits of technology

The proposed method increases the throughput and quality of nucleic acid sequence data, improving genomics, diagnostics, and other biological research by accurately resolving cluster data from overlapping clusters, reducing computational resource requirements, and enhancing sequencing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767012000001
    Figure 0007767012000001
  • Figure 0007767012000002
    Figure 0007767012000002
  • Figure 0007767012000003
    Figure 0007767012000003
Patent Text Reader

Abstract

The disclosed technique includes using a neural network to determine analyte metadata by (i) processing input image data derived from a sequence of an image set through the neural network to generate alternative representations of the input image data, the input image data having an array of units depicting analytes and their surrounding background, (ii) processing the alternative representations through an output layer to generate an output value for each unit in the array, (iii) thresholding the output values ​​of the units to classify a first subset of the units as background units depicting the surrounding background, and (iv) locating peaks in the output values ​​of the units and classifying a second subset of the units as center units that include centers of the analytes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority application) This application claims priority to or the benefit of the following applications:

[0002] U.S. Provisional Patent Application No. 62 / 821,602, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-1 / IP-1693-PRV);

[0003] U.S. Provisional Patent Application No. 62 / 821,618, entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed March 21, 2019 (Attorney Docket No. ILLM1008-3 / IP-1741-PRV);

[0004] U.S. Provisional Patent Application No. 62 / 821,681, entitled "Artificial Intelligence-Based Base Calling," filed March 21, 2019 (Attorney Docket No. ILLM1008-4 / IP-1744-PRV);

[0005] U.S. Provisional Patent Application No. 62 / 821,724, entitled "Artificial Intelligence-Based Quality Scoring," filed March 21, 2019 (Attorney Docket No. ILLM1008-7 / IP-1747-PRV);

[0006] U.S. Provisional Patent Application No. 62 / 821,766, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV);

[0007] Dutch Patent Application No. 2023310 (Attorney Docket No. ILLM1008-11 / IP-1693-NL), entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed on June 14, 2019;

[0008] Dutch Patent Application No. 2023311 (Attorney Docket No. ILLM1008-12 / IP-1741-NL), entitled "Artificial Intelligence-Based Generation of Sequencing Metadata," filed on June 14, 2019;

[0009] Dutch Patent Application No. 2023312 entitled "Artificial Intelligence-Based Base Calling" filed on June 14, 2019 (Attorney Docket No. ILLM1008-13 / IP-1744-NL);

[0010] Dutch Patent Application No. 2023314 entitled "Artificial Intelligence-Based Quality Scoring", filed on June 14, 2019 (Attorney Docket No. ILLM1008-14 / IP-1747-NL), and

[0011] Dutch Patent Application No. 2023316 (Attorney Docket No. ILLM1008-15 / IP-1752-NL), entitled "Artificial Intelligence-Based Sequencing," filed on June 14, 2019.

[0012] U.S. Patent Application No. 16 / 825,987, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 20, 2020 (Attorney Docket No. ILLM1008-16 / IP-1693-US);

[0013] U.S. Patent Application No. 16 / 825,991, entitled "Training Data Generation for Artificial Intelligence-Based Sequencing," filed March 20, 2020 (Attorney Docket No. ILLM1008-17 / IP-1741-US);

[0014] U.S. Patent Application No. 16 / 826,126, entitled "Artificial Intelligence-Based Base Calling," filed March 20, 2020 (Attorney Docket No. ILLM1008-18 / IP-1744-US);

[0015] U.S. Patent Application No. 16 / 826,134, entitled "Artificial Intelligence-Based Quality Scoring," filed March 20, 2020 (Attorney Docket No. ILLM1008-19 / IP-1747-US);

[0016] U.S. Patent Application No. 16 / 826,168, entitled "Artificial Intelligence-Based Sequencing," filed March 21, 2020 (Attorney Docket No. ILLM1008-20 / IP-1752-PRV);

[0017] PCT Patent Application No. PCT_____, entitled "Training Data Generation for Artificial Intelligence Based Sequencing" (Attorney Docket No. ILLM 1008-21 / IP-1693-PCT)

[0018] PCT Patent Application No. PCT___________ (Attorney Docket No. ILLM1008-23 / IP-1744-PCT), entitled "Artificial Intelligence-Based Base Calling," filed concurrently herewith and subsequently published as PCT International Publication No. WO____________;

[0019] PCT Patent Application No. PCT__________ (Attorney Docket No. ILLM1008-24 / IP-1747-PCT), entitled "Artificial Intelligence-Based Quality Scoring," filed concurrently herewith and subsequently published as PCT International Publication No. WO____________; and

[0020] PCT Patent Application No. PCT___________, entitled "Artificial Intelligence-Based Sequencing," filed concurrently herewith, and subsequently published as PCT International Publication No. WO____________ (Attorney Docket No. ILLM1008-25 / IP-1752-PCT).

[0021] The priority application is incorporated herein by reference for all purposes as if fully set forth herein. (built-in)

[0022] The following are incorporated by reference for all purposes as if fully set forth herein:

[0023] U.S. Provisional Patent Application No. 62 / 849,091, entitled "Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing," filed May 16, 2019 (Attorney Docket No. ILLM1011-1 / IP-1750-PRV);

[0024] U.S. Provisional Patent Application No. 62 / 849,132, entitled "Base Calling Using Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-2 / IP-1750-PR2);

[0025] U.S. Provisional Patent Application No. 62 / 849,133, entitled "Base Calling Using Compact Convolutions," filed May 16, 2019 (Attorney Docket No. ILLM1011-3 / IP-1750-PR3);

[0026] U.S. Provisional Patent Application No. 62 / 979,384, entitled "Artificial Intelligence-Based Base Calling of Index Sequences," filed February 20, 2020 (Attorney Docket No. ILLM1015-1 / IP-1857-PRV);

[0027] U.S. Provisional Patent Application No. 62 / 979,414, entitled "Artificial Intelligence-Based Many-To-Many Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1016-1 / IP-1858-PRV);

[0028] U.S. Provisional Patent Application No. 62 / 979,385, entitled "Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller," filed February 20, 2020 (Attorney Docket No. ILLM1017-1 / IP-1859-PRV);

[0029] U.S. Provisional Patent Application No. 62 / 979,412, entitled "Multi-Cycle Cluster Based Real Time Analysis System," filed February 20, 2020 (Attorney Docket No. ILLM1020-1 / IP-1866-PRV);

[0030] U.S. Provisional Patent Application No. 62 / 979,411, entitled "Data Compression for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1029-1 / IP-1964-PRV);

[0031] U.S. Provisional Patent Application No. 62 / 979,399, entitled "Squeezing Layer for Artificial Intelligence-Based Base Calling," filed February 20, 2020 (Attorney Docket No. ILLM1030-1 / IP-1982-PRV);

[0032] Liu P, Hemani A, Paul K, Weis C, Jung M, Wehn ​​N. 3D-Stacked Many-Core Architecture for Biological Sequence Analysis Problems. Int J Parallel Prog. 2017, 45(6):1420-60,

[0033] Z. Wu, K. Hammad, R. Mittmann, S. Magierowski, E. Ghafar-Zadeh, and X. Zhong, "FPGA-Based DNA Basecalling Hardware Acceleration", in Proc. IEEE 61st Int. Midwest Symp. Circuits Syst., Aug. 2018, pp. 1098-1101,

[0034] Z. Wu, K. Hammad, E. Ghafar-Zadeh, and S. Magierowski, "FPGA-Accelerated 3rd Generation DNA Sequencing", in IEEE Transactions on Biomedical Circuits and Systems, Volume 14, Issue 1, Feb. 2020, pp. 65-74,

[0035] Prabhakar et al., "Plasticine: A Reconfigurable Architecture for Parallel Patterns", ISCA'17, June 24-28, 2017, Toronto, ON, Canada.

[0036] M.Lin,Q.Chen,and S.Yan、「Network in Network」、in Proc.of ICLR,2014、

[0037] L.Sifre、「Rigid-motion Scattering for Image Classification,Ph.D.thesis,2014、

[0038] L.Sifre and S.Mallat、「Rotation,Scaling and Deformation Invariant Scattering for Texture Discrimination」、in Proc.of CVPR,2013、

[0039] F.Chollet、「Xception:Deep Learning with Depthwise Separable Convolutions」、in Proc.of CVPR,2017、

[0040] X.Zhang,X.Zhou,M.Lin,and J.Sun、「ShuffleNet:An Extremely Efficient Convolutional Neural Network for Mobile Devices」、in arXiv:1707.01083,2017、

[0041] K.He,X.Zhang,S.Ren,and J.Sun、「Deep Residual Learning for Image Recognition」、in Proc.of CVPR,2016、

[0042] S.Xie,R.Girshick,P.Dollar,Z.Tu,and K.He、「Aggregated Residual Transformation For Deep NeuroNetworks」、Proc.of CVPR,2017、

[0043] A.G.Howard,M.Zhu,B.Chen,D.Kalenichenko,W.Wang,T.Weyand,M.Andreetto,and H.Adam、「Mobilenets:Efficient Convolutional Neural Networks for Mobile Vision Applications」、in arXiv:1704.04861,2017、

[0044] M.Sandler,A.Howard,M.Zhu,A.Zhmoginov,and L.Chen、「MobileNetV2:Inverted Residuals and Linear Bottlenecks」、in arXiv:1801.04381v3,2018、

[0045] Z.Qin,Z.Zhang,X.Chen and Y.Peng、「FD-MobileNet:Improved MobileNet with a Fast Downsampling Strategy」、in arXiv:1802.03750,2018、

[0046] Liang-Chieh Chen,George Papandreou,Florian Schroff,and Hartwig Adam.Rethinking atrous convolution for semantic image segmentation.CoRR、abs / 1706.05587,2017、

[0047] J.Huang,V.Rathod,C.Sun,M.Zhu,A.Korattikara,A.Fathi,I.Fischer,Z.Wojna,Y.Song,S.Guadarrama,et al.Speed / accuracy trade-offs for modern convolutional object detectors.arXiv preprint arXiv:1611.10012,2016、

[0048] S.Dieleman,H.Zen,K.Simonyan,O.Vinyals,A.Graves,N.Kalchbrenner,A.Senior,and K.Kavukcuoglu、「WAVENET:A GENERATIVE MODEL FOR RAW AUDIO」、arXiv:1609.03499,2016、

[0049] S.O.Arik,M.Chrzanowski,A.Coates,G.Diamos,A.Gibiansky,Y.Kang,X.Li,J.Miller,A.Ng,J.Raiman,S.Sengupta and M.Shoeybi、「DEEP VOICE:REAL-TIME NEURAL TEXT-TO-SPEECH」、arXiv:1702.07825,2017、

[0050] F.Yu and V.Koltun、「MULTI-SCALE CONTEXT AGGREGATION BY DILATED CONVOLUTIONS」、arXiv:1511.07122,2016、

[0051] K.He,X.Zhang,S.Ren,and J.Sun、「DEEP RESIDUAL LEARNING FOR IMAGE RECOGNITION」、arXiv:1512.03385,2015、

[0052] R.K.Srivastava,K.Greff,and J.Schmidhuber、「HIGHWAY NETWORKS」、arXiv:1505.00387,2015、

[0053] G.Huang,Z.Liu,L.van der Maaten and K.Q.Weinberger、「DENTILY CONNECTED CONVOLUTIONAL NETWORKS」、arXiv:1608.06993,2017、

[0054] C.Szegedy,W.Liu,Y.Jia,P.Sermanet,S.Reed,D.Anguelov,D.Erhan,V.Vanhoucke,and A.Rabinovich、「GOING DEEPER WITH CONVOLUTIONS」、arXiv:1409.4842,2014、

[0055] S.Ioffe and C.Szegedy、「BATCH NORMALIZATION:ACCELERATING DEEP NETWORK TRAINING BY REDUCING INTERNAL COVARIATE SHIFT」、arXiv:1502.03167,2015、

[0056] J.M.Wolterink,T.Leiner,M.A.Viergever,and 1.Isgum、「DILATED CONVOLUTIONAL NEURAL NETWORKS FOR CARDIOVASCULAR MR SEGMENTATION IN CONGENITAL HEART DISEASE」、arXiv:1704.03669,2017、

[0057] L.C.Piqueras、「AUTOREGRESSIVE MODEL BASED ON A DEEP CONVOLUTIONAL NEURAL NETWORK FOR AUDIO GENERATION」、Tampere University of Technology,2016、

[0058] J.Wu、「Introduction to Convolutional Neural Networks」、Nanjing University,2017、

[0059] 「Illumina CMOS Chip and One-Channel SBS Chemistry」、Illumina,Inc.2018,2 pages、

[0060] "skikit-image / peak.py at master", GitHub, 5 pages, [Retrieved 2018-11-16]. Internet<URL:https: / / github.com / scikit-image / scikit-image / blob / master / skimage / feature / peak.py#L25> Search from,

[0061] "3.3.9.11. Watershed and random walker for segmentation," Scipy lecture notes, 2 pages, [Retrieved 2018-11-13]. Internet<URL:http: / / scipy-lectures.org / packages / scikit-image / auto_examples / plot_segmentations.html> Search from,

[0062] Mordvintsev, Alexander and Revision, Abid K., "Image Segmentation with Watershed Algorithm", Revision 43532856, 2013, 6 pages [Retrieved 2018-11-13]. Internet <URL:https: / / opencv-python-tutroals.readthedocs.io / en / latest / py_tutorials / py_imgproc / py_watershed / py_watershed.html> Search from,

[0063] Mzur, "Watershed.py", 25 October 2017, 3 pages, [Retrieved 2018-11-13]. Internet<URL:https: / / github.com / mzur / watershed / blob / master / Watershed.py> Search from,

[0064] Thakur,Pratibha,et.al.「A Survey of Image Segmentation Techniques」、International Journal of Research in Computer Applications and Robotics,Vol.2,Issue.4,,April 2014,Pg.:158-165、

[0065] Long,Jonathan,et.al.、「Fully Convolutional Networks for Semantic Segmentation」、:IEEE Transactions on Pattern Analysis and Machine Intelligence,Vol 39,Issue 4,1 April 2017,10 pages、

[0066] Ronneberger,Olaf,et.al.、「U-net:Convolutional networks for biomedical image segmentation」.In International Conference on Medical image computing and computer-assisted intervention,18 May 2015,8 pages、

[0067] Xie,W.,et.al.、「Microscopy cell counting and detection with fully convolutional regression networks」,Computer methods in biomechanics and biomedical engineering:Imaging & Visualization,6(3),pp.283-292,2018、

[0068] Xie, Yuanpu, et al., "Beyond classification: structured regression for robust cell detection using convolutional neural network", International Conference on Medical Image Computing and Computer-Assisted Intervention. October 2015, 12 pages.

[0069] Snuverink, I.A.F., "Deep Learning for Pixelwise Classification of Hyperspectral Images", Master of Science Thesis, Delft University of Technology, 23 November 2017, 19 pages.

[0070] Shevchenko, A., "Keras weighted categorical_crossentropy", 1 page, [Searched on January 15, 2019]. Retrieved from the Internet <URL:https: / / gist.github.com / skeeet / cad06d584548fb45eece1d4e28cfa98b>.

[0071] van den Assem, D.C.F., "Predicting periodic And chaotic signals using Wavenets", Master of Science Thesis, Delft University Of Technology, 18 August 2017, Pages 3-38.

[0072] I.J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, "CONVOLUTIONAL NETWORKS", Deep Learning, MIT Press, 2016, and

[0073] J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, and G. Wang, “RECENT ADVANCES IN CONVOLUTIONAL NEURAL NETWORKS”, arXiv:1512.07108, 2017.

[0074] FIELD OF THE INVENTION The technology of this disclosure relates to artificial intelligence computers and digital data processing systems and corresponding data processing methods and products for emulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the disclosed technology relates to using deep neural networks, such as deep convolutional neural networks, to analyze data. [Background technology]

[0075] The subject matter described in this section should not be assumed to be prior art merely as a result of its mention in this section. Likewise, it should not be assumed that the problems mentioned in this section, or problems related to the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to the implementation of the claimed technology.

[0076] Deep neural networks are a class of artificial neural networks that use multiple nonlinear and complex transformation layers to continuously model high-level functions. Deep neural networks provide feedback via backpropagation, which propagates the difference between observed and predicted outputs to adjust parameters. Deep neural networks have evolved with the availability of large training datasets, the power of parallel and distributed computing, and advanced training algorithms. Deep neural networks have driven major advances in many domains, such as computer vision, speech recognition, and natural language processing.

[0077] Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks have been successful in image recognition, particularly in structures including convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to exploit the continuous information of input data with periodic connections between their constituent units, such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks have been proposed for limited situations, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.

[0078] The goal of training a deep neural network is to optimize the weight parameters in each layer, gradually combining simpler features into more complex ones so that a better hierarchical representation can be learned from the data. A single cycle of the optimization process consists of the following: First, given a training dataset, a forward pass sequentially computes the outputs in each layer and propagates a functional signal forward through the network. At the final output layer, an objective loss function measures the error between the inferred output and a given label. To minimize the training error, a backward pass backpropagates the error signal using a chain rule and computes gradients for all weights throughout the neural network. Finally, stochastic parameters are updated using an optimization algorithm based on stochastic gradient descent. While batch gradient descent updates parameters for the entire dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small set of data examples. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency of each parameter and the momentum of the gradient, respectively.

[0079] Another core element in deep neural network training is regularization, which refers to a strategy intended to avoid overfitting and thus achieve good generalization performance. For example, weight decay adds a penalty term to the objective loss function so that weight parameters converge to smaller absolute values. Dropout randomly removes hidden units from a neural network during training, which can be viewed as a collection of possible subnetworks. To improve dropout's capabilities, a new activation function, maxout, and a variant of dropout for recurrent neural networks called rnnDrop have been proposed. Furthermore, batch normalization provides a new regularization method via scalar feature normalization for each activation within a mini-batch, learning its mean and variance as parameters.

[0080] Given the multidimensional and high-dimensional nature of sequence data, deep neural networks hold considerable promise for bioinformatics research due to their broad applicability and enhanced predictive capabilities. Convolutional neural networks have been employed to solve sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight-sharing strategy that is particularly useful for studying DNA, as it can capture short sequence motifs that reproduce local patterns in DNA that are predicted to have significant biological functions. A notable feature of convolutional neural networks is the use of convolutional filters.

[0081] Unlike traditional classification approaches based on carefully designed, manually crafted features, convolutional filters perform adaptive feature learning, similar to the process of mapping raw input data to information representations of knowledge. In this sense, convolutional filters function as a series of motif scanners, as a set of such filters can recognize relevant patterns in the input and update themselves during the training procedure. Recurrent neural networks can capture long-range dependencies in continuous data of various lengths, such as protein or DNA sequences.

[0082] Therefore, an opportunity arises to use a coherent deep learning-based framework for template generation and base calling.

[0083] In the era of high-throughput technologies, accumulating the highest yield of interpretable data at the lowest cost per effort remains a significant challenge. Cluster-based methods of nucleic acid sequencing, such as those that utilize bridge amplification for cluster formation, have made valuable contributions to the goal of increasing nucleic acid sequencing throughput. These cluster-based methods rely on sequencing dense populations of nucleic acids immobilized on a solid support and typically involve the use of image analysis software to suppress the optical signal generated during the simultaneous sequencing of multiple clusters located at distinct locations on the solid support.

[0084] However, such solid-phase nucleic acid cluster-based sequencing technology faces considerable obstacles that limit the amount of throughput that can be achieved.For example, in cluster-based sequencing methods, determining the nucleic acid sequences of two or more clusters that are too close to each other physically to be spatially resolved, or that actually physically overlap on a solid support, can pose obstacles.For example, current image analysis software may require valuable time and computational resources to determine which of two overlapping clusters an optical signal originates from.As a result, various detection platforms inevitably compromise the quantity and / or quality of nucleic acid sequence information that can be obtained.

[0085] High-density nucleic acid aggregate-based genomics methods extend to other areas of genome analysis as well. For example, nucleic acid cluster-based genomics can be used in sequencing applications, diagnostics and screening, gene expression analysis, epigenetic analysis, genetic analysis of polymorphisms, etc. Each of these nucleic acid cluster-based genomics techniques is limited by the inability to resolve data generated from closely spaced or spatially overlapping nucleic acid clusters.

[0086] Clearly, there is a need to improve the quality and quantity of nucleic acid sequence data that can be obtained rapidly and cost-effectively for a variety of applications, including genomics (e.g., for genomic characterization of any and all animal, plant, microbial, or other biological species or populations), pharmacogenomics, transcriptomics, diagnostics, prognosis, biomedical risk assessment, clinical and research genetics, personalized medicine, drug efficacy and drug interaction assessment, veterinary medicine, agricultural, evolutionary, and biological research, aquatic culture, forestry, marine research, ecological and environmental management, and other purposes.

[0087] The disclosed technology provides neural network-based methods and systems that address these and similar needs, including increasing the level of throughput in high-throughput nucleic acid sequencing technologies, and offers other related advantages. [Brief explanation of the drawings]

[0088] [Figure 1] 1 illustrates one embodiment of a processing pipeline for determining cluster metadata using subpixel base calls. [Figure 2] 1 shows an embodiment of a flow cell containing clusters within its tiles. [Figure 3] An example of an Illumina GA-IIx flow cell with eight lanes is shown. [Figure 4] An image set of four channel chemical sequence images is depicted, i.e., the image set has four sequence images captured using four different wavelength bands (image / imaging channels) in the pixel domain. [Figure 5] 1 is an embodiment of dividing a sequence image into sub-pixels (or sub-pixel regions). [Figure 6] During subpixel base calling, preliminary center coordinates of clusters identified by the base caller are shown. [Figure 7]An example of merging sub-pixel base calls generated over multiple sequencing cycles to generate a so-called "cluster map" containing cluster metadata is shown. [Figure 8a] 1 shows an example of a cluster map generated by merging subpixel base calls. [Figure 8b] 1 illustrates one embodiment of subpixel base calling. [Figure 9] 10 illustrates another example of a cluster map that identifies cluster metadata. [Figure 10] We show how the center of mass (COM) of discontinuous regions in a cluster map is calculated. [Figure 11] 10 illustrates one embodiment of a calculation of a weighted attenuation coefficient based on the Euclidean distance from a subpixel of a discontinuous region to a COM of the discontinuous region. [Figure 12] 10 illustrates one implementation of an exemplary ground truth attenuation map derived from an exemplary cluster map generated by sub-pixel base calling. [Figure 13] 1 illustrates one embodiment of deriving a ternary map from a cluster map. [Figure 14] 1 illustrates one embodiment of deriving a binary map from a cluster map. [Figure 15] FIG. 1 is a block diagram illustrating one embodiment for generating training data used to train a neural network-based template generator and a neural network-based base caller. [Figure 16] 1 illustrates the properties of the disclosed training examples used to train the neural network-based template generator and the neural network-based base caller. [Figure 17] 1 illustrates one embodiment of processing input image data through the disclosed neural network-based template generator to generate an output value for each unit in the array. In one embodiment, the array is an attenuation map. In another embodiment, the array is a ternary map. In yet another embodiment, the array is a binary map. [Figure 18] FIG. 1 illustrates one embodiment of a post-processing technique applied to attenuation maps, ternary maps, or binary maps generated by a neural network-based template generator to derive cluster metadata including cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and / or cluster boundaries. [Figure 19] 1 illustrates one embodiment for extracting cluster intensities in the pixel domain. [Figure 20] 1 illustrates one embodiment for extracting cluster intensities in the sub-pixel domain. [Figure 21a] Three different implementations of a neural network-based template generator are presented. [Figure 21b] 15 illustrates one embodiment of input image data provided as input to the neural network-based template generator 1512. The input image data includes a series of image sets having sequence images generated during a specific number of initial sequence cycles of a sequencing run. [Figure 22] FIG. 21B illustrates one embodiment of extracting patches from the series of image sets of FIG. 21b to generate a series of "downsized" image sets that form the input image data. [Figure 23] 21b shows one embodiment of upsampling the series of image sets of FIG. 21b to generate a series of "upsampled" image sets that form the input image data. [Figure 24] 24 illustrates one embodiment of extracting patches from the series of upsampled image sets of FIG. 23 to generate a series of "upsampled and downsized" image sets that form the input image data. [Figure 25] 1 illustrates one embodiment of an overall exemplary process for generating ground truth data for training a neural network-based template generator. [Figure 26] 1 illustrates one embodiment of a regression model. [Figure 27]1 illustrates one embodiment of generating a ground truth attenuation map from a cluster map, which is used as ground truth data for training a regression model. [Figure 28] 1 is an embodiment of training a regression model using a backpropagation-based gradient update technique. [Figure 29] 1 is an embodiment of template generation by a regression model during inference. [Figure 30] 10 illustrates one embodiment of post-processing the attenuation map to identify cluster metadata. [Figure 31] 1 illustrates one implementation of a watershed segmentation technique that identifies non-overlapping groups of adjacent clusters / inter-cluster sub-pixels that characterize clusters. [Figure 32] 1 is a table showing an exemplary U-Net structure for a regression model. [Figure 33] We present a different approach to extract cluster intensities using cluster shape information identified in a template image. [Figure 34] 1 illustrates different approaches to base calling using the output of a regression model. [Figure 35] Figure 1 shows the difference in base calling performance when the RTA base caller uses ground truth center of mass (COM) positions as cluster centers, as opposed to using non-COM positions as cluster centers. The results show that using COM positions improves base calling. [Figure 36] On the left, we show an example attenuation map that generated the regression model. On the right, Figure 36 also shows an example ground truth attenuation map that the regression model approximates during training. [Figure 37] 1 illustrates one embodiment of a peak locator that identifies cluster centers in an attenuation map by detecting peaks. [Figure 38] Peaks detected by the peak locator in the attenuation map generated by the regression model are compared with peaks in the corresponding ground truth attenuation map. [Figure 39] Demonstrate the performance of your regression model using precision and recall statistics. [Figure 40] Compare the performance of the RTA base caller and the regression model for a library concentration of 20 pM (normal operation). [Figure 41] Compare the performance of the RTA base caller and the regression model for a library concentration of 30 pM (high density run). [Figure 42] The number of non-overlapping, well-aligned read pairs, i.e., the number of pairs of reads where neither read is aligned within a reasonable distance as detected by the regression model, is compared to those detected by RTA base calling. [Figure 43] On the right, the first attenuation map generated by the regression model is shown. On the left, Figure 43 shows the second attenuation map generated by the regression model. [Figure 44] Compare the performance of the RTA base caller and the regression model for 40 pM library concentration (high density run). [Figure 45] The first attenuation map generated by the regression model is shown on the left. On the right, Figure 45 shows the results of thresholding, peak location processing, and watershed segmentation techniques applied to the first attenuation map. [Figure 46] 1 illustrates one embodiment of a binary classification model. [Figure 47] 1 is an implementation of training a binary classification model using a backpropagation-based gradient update technique with softmax scoring. [Figure 48] 1 is another implementation of training a binary classification model using a backpropagation-based gradient update technique with sigmoid scores. [Figure 49] 1 illustrates another embodiment of input image data provided to a binary classification model and corresponding class labels used to train the binary classification model. [Figure 50] 1 is an embodiment of template generation with a binary classification model during inference. [Figure 51]1 illustrates one embodiment in which the binary map is subjected to peak detection to identify cluster centers. [Figure 52a] An example binary map generated by a binary classification model is shown on the left. Figure 52a also shows an example ground truth binary map on the right, to which the binary classification model is proximate during training. [Figure 52b] Use accuracy statistics to indicate the performance of a binary classification model. [Figure 53] 1 is a table illustrating an example structure of a binary classification model. [Figure 54] 1 illustrates one embodiment of a three-way classification model. [Figure 55] 1 is an implementation of training a ternary classification model using a backpropagation-based gradient update technique. [Figure 56] 10 illustrates another implementation of input image data fed to a three-way classification model and the corresponding class labels used to train the three-way classification model. [Figure 57] 1 is a table illustrating an exemplary structure of a three-way classification model. [Figure 58] 1 is an embodiment of template generation with a three-way classification model during inference. [Figure 59] 1 shows a ternary map generated by a ternary classification model. [Figure 60] The unit array generated by the three-way classification model 5400 is shown along with the output values ​​for each unit. [Figure 61] 1 illustrates one embodiment in which the ternary map is subjected to post-processing to identify cluster centers, cluster backgrounds, and cluster interiors. [Figure 62a] 1 shows an exemplary prediction of a three-way classification model. [Figure 62b] 10 shows another exemplary prediction of a three-way classification model. [Figure 62c] 10 shows yet another exemplary prediction of a three-way classification model. [Figure 63] FIG. 62b shows one embodiment for deriving cluster centers and cluster shapes from the output of the ternary classification model of FIG. 62a. [Figure 64]Compare the base calling performance of a binary classification model, a regression model, and an RTA base caller. [Figure 65] We compare the performance of the ternary classification model with that of the RTA-based caller under three conditions, five sequence metrics, and two driving densities. [Figure 66] We compare the performance of the regression model with that of the RTA-based caller under three conditions, five sequence metrics, and two driving densities considered in Figure 65 . [Figure 67] We focus on the penultimate layer of the neural network-based template generator. [Figure 68] Visualize what the penultimate layer of a neural network-based template generator has learned as a result of backpropagation-based gradient update training. The illustrated embodiment visualizes 24 out of 32 trained convolution filters in the penultimate layer shown in Figure 67. [Figure 69] Cluster center predictions from a binary classification model (in blue) are overlaid on RTA base calls (in pink). [Figure 70] Cluster center predictions produced by RTA-based color (in pink) are overlaid on a visualization of the trained convolutional filters in the penultimate layer of a binary classification model. [Figure 71] 1 illustrates one embodiment of training data used to train a neural network-based template generator. [Figure 72] 10 is an implementation of using beads for image registration based on cluster center prediction of a neural network-based template generator. [Figure 73] 10 illustrates one embodiment of cluster statistics for clusters identified by a neural network-based template generator. [Figure 74]We show how the ability of the neural network-based template generator to distinguish between adjacent clusters improves as the number of initial sequencing cycles for which the input image data is used increases from 5 to 7. [Figure 75] Figure 1 shows the difference in base calling performance when the RTA base caller uses ground truth center of mass (COM) positions as cluster centers as opposed to when non-COM positions are used as cluster centers. [Figure 76] We show the performance of the neural network-based template generator on additional detected clusters. [Figure 77] 1 shows different datasets used to train the neural network-based template generator. [Figure 78A] 1 illustrates one embodiment of a sequence system, the sequence system including a configurable processor. [Figure 78B] 1 illustrates one embodiment of a sequence system, the sequence system including a configurable processor. [Figure 79] FIG. 1 is a simplified block diagram of a system for analysis of sensor data from a sequencing system, such as base call sensor output. [Figure 80] FIG. 1 is a simplified diagram illustrating aspects of a base calling operation, including the functionality of a runtime program executed by a host processor. [Figure 81] FIG. 79 is a simplified diagram of a configuration of a configurable processor such as that shown in FIG. [Figure 82] 78B is a computer system that can be used by the sequencing system of FIG. 78A to implement the techniques disclosed herein. DETAILED DESCRIPTION OF THE INVENTION

[0089] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.

[0090] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings:

[0091] The following description is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0092] (introduction) Base calling from digital images is massively parallel and computationally intensive, which presents a number of technical challenges that were identified before implementing our novel technique.

[0093] The signal from the image set being evaluated becomes increasingly weak as the base classification progresses periodically, especially over longer and longer strands of the base. As the base classification extends over the length of the strand, the signal-to-noise ratio decreases and the reliability decreases. An updated estimate of reliability is projected as the estimated reliability of the change in base classification.

[0094] A digital image is captured from the amplified clusters of sample strands. The sample is amplified by replicating the strands using various physical structures and chemicals. During sequencing by synthesis, tags are chemically bound in cycles and stimulated to glow. A digital sensor collects photons from the tags, which are read out as pixels to generate an image.

[0095] Interpreting digital images to classify bases requires resolving positional uncertainty, which is hindered by limited image resolution. At resolutions higher than those collected during base calling, it is clear that imaged clusters have irregular shapes and uncertain center locations. Cluster positions are not mechanically controlled, and therefore cluster centers are not aligned with pixel centers. The pixel center can be an integer coordinate assigned to the pixel. In other embodiments, it can be the upper left corner of the pixel. In yet other embodiments, it can be the center of gravity or center of mass of the pixel. Amplification does not produce uniform cluster shapes. Therefore, the distribution of cluster signals in the digital image is a statistical distribution rather than a regular pattern. We determine this positional uncertainty.

[0096] One of the signal classes does not produce a detectable signal and can be classified into a specific location based on the "dark" signal. Therefore, a template is required for classification during the dark cycle. Template generation resolves the initial location uncertainty using multiple imaging cycles to avoid missing dark signals.

[0097] Tradeoffs in image sensor size, magnification, and stepper design lead to pixel sizes that are too large to manipulate cluster centers to coincide with sensor pixel centers. This disclosure uses pixel in two senses. A physical sensor pixel is an area of ​​a photosensor that reports detected photons. A logical pixel, simply called a pixel, is the data corresponding to at least one physical pixel and readout from the sensor pixel. Pixels can be subdivided or "upsampled" into subpixels (e.g., 4x4 subpixels). To account for the possibility that all photons hit one side of a physical pixel and not the other, values ​​can be assigned to the subpixels by interpolation, such as bilinear interpolation or area weighting. Interpolation or bilinear interpolation is also applied when a pixel is reframed by applying an affine transformation to the data from the physical pixel.

[0098] Larger physical pixels are more sensitive to weak signals than smaller pixels. Digital sensors improve over time, but the physical limitations of collector surface area are unavoidable. Considering design tradeoffs, legacy systems are designed to collect and analyze image data from a 3x3 patch of sensor pixels, with the center of that cluster located at the center pixel of the patch.

[0099] High-resolution sensors capture only a portion of the imaged medium at a time. The sensor steps over the imaged medium, covering the entire field of view. Thousands of digital images can be collected during one processing cycle.

[0100] The sensor and illumination design are combined to distinguish at least four illumination response values ​​used to classify the bass. Using a conventional RGB camera with a Bayer color filter array, the four sensor pixels are combined into a single RGB value. This would reduce the effective sensor resolution by four times. Alternatively, images can be collected at a single location using different illumination wavelengths and / or different filters rotated in position between the imaged medium and the sensor. The number of images required to distinguish between the four base classes varies between systems. Some systems use one image with four intensity levels for different classes of bass. Other systems use two images with different illumination wavelengths (e.g., red and green) and / or filters with a type of hologram to classify the bass. Systems can also use four images with different illumination wavelengths and / or filters tuned to specific bass classes.

[0101] Highly parallel processing of digital images requires the alignment of relatively short strands, on the order of 30-2000 base pairs, to sequences of much longer lengths, potentially millions or even billions in length. Because redundant samples are desirable on the imaged medium, portions of a sequence may be covered by multiple sample reads. Millions, or at least hundreds of thousands, of sample clusters may be imaged from a single imaged medium. Large-scale processing of such many clusters increases sequencing capacity while decreasing costs.

[0102] Sequencing capacity is increasing at a pace that replicates Moore's Law. While the first sequencing costs billions of dollars, 2018 services such as Illumina™ provide results for hundreds of dollars. As sequencing becomes mainstream and unit costs fall, less computing power is available for classification, which increases the challenge of near-real-time classification. With these technical challenges in mind, the inventors turn to the disclosed technology.

[0103] The disclosed techniques improve processing both during template generation to resolve position uncertainty and during base classification of clusters at resolved positions. Applying the disclosed techniques can use cheaper hardware to reduce machine costs. Near real-time analysis can be cost-effective and reduce the delay between image collection and base classification.

[0104] The disclosed technology can use upsampled images generated by interpolating sensor pixels to subpixels, then generate templates that resolve positional uncertainty. The resulting subpixels are submitted to a base caller for classification, which treats the subpixels as if they were the center of a cluster. Clusters are identified from groups of adjacent subpixels that repeatedly receive the same base classification. This aspect of the technology can leverage existing base calling techniques to identify cluster shapes and super-search for cluster centers at subpixel resolution.

[0105] Another aspect of the disclosed technology is creating ground truth pairs of images with reliably identified cluster centers and / or cluster shapes. Deep learning systems and other machine learning approaches require substantial training sets. Human-curated data is expensive to compile. The disclosed technology can be used to generate large sets of sensitively classified training data in non-standard modes of operation without the intervention or expense of a human curator. The training data correlates raw images with cluster centers and / or cluster shapes available from existing classifiers in non-standard modes of operation, such as CNN-based deep learning systems. A single training image can be rotated and reflected to generate additional, equally valid examples. Training examples can be focused on regions of a predetermined size within the entire image. The context evaluated during base calling determines the size of the example training region, rather than the size of the image or the entire imaged medium.

[0106] The disclosed technology can generate different types of maps that can be used as training data or as templates for base classification, correlating cluster centers and / or cluster shapes with digital images. First, subpixels can be classified as cluster centers, thereby localizing cluster centers within a physical sensor pixel. Second, cluster centers can be calculated as the centroid of the cluster shape. This location can be reported to a selected numerical precision. Third, cluster centers can be reported at surrounding subpixels in an attenuation map, either at subpixel or pixel resolution. The attenuation map reduces the weight given to photons detected within a region as the region's separation from the cluster center increases, attenuating signals from more distant locations. Fourth, binary or ternary classification can be applied to subpixels or pixels within a cluster of neighboring regions. In binary classification, regions are classified as belonging to the cluster center or as background. In ternary classification, a third class type is assigned to regions that include the cluster interior but are not the cluster center. Subpixel classification of cluster center locations can be substituted for real-valued cluster center coordinates within a larger optical pixel.

[0107] Alternative map styles can be generated initially as a ground truth dataset or can be generated using a neural network after training. For example, clusters can be depicted as discontinuous regions of adjacent subpixels with appropriate classifications. The intensities of the mapped clusters from the neural network can be post-processed with a peak detector filter to calculate cluster centers if the centers have not already been determined. Adjacent regions can be assigned to distinct clusters by applying a so-called watershed analysis. When generated by a neural network inference engine, the map can be used as a template for evaluating sequences of digital images and classifying bases across cycles of base calling.

[0108] (Neural network-based template generation) The first step in template generation is to identify cluster metadata, which identifies the spatial distribution of clusters, including their centers, shapes, sizes, backgrounds, and / or boundaries.

[0109] (Identifying cluster metadata) FIG. 1 shows one embodiment of a processing pipeline for identifying cluster metadata using subpixel base calls.

[0110] Figure 2 shows one embodiment of a flow cell containing clusters within its tiles. The flow cell is divided into lanes. The lanes are further divided into non-overlapping regions called "tiles." During the sequencing procedure, the populations on the tiles and their surrounding background are imaged.

[0111] Figure 3 shows an exemplary Illumina GA-IIx™ flow cell with eight lanes. Figure 3 also shows a magnification of one tile and its cluster and their surrounding background.

[0112] FIG. 4 depicts an image set of a four-channel chemical sequence image, i.e., the image set has four sequence images captured using four different wavelength bands (image / imaging channels) in the pixel domain. Each image in the image set covers a tile of a flow cell and shows the intensity emission of clusters on the tile and their surrounding background, captured for a particular image channel during a particular one of multiple sequencing cycles of a sequencing run performed on the flow cell. In one embodiment, each imaging channel corresponds to one of multiple filter wavelength bands. In another embodiment, each imaging channel corresponds to one of multiple imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter. The intensity emission of the clusters comprises a signal detected from the analyte that can be used to classify bases associated with the analyte. For example, the intensity emission can be a signal indicative of photons emitted by a tag chemically attached to the analyte during a cycle in which the tag is stimulated and can be detected by one or more digital sensors.

[0113] FIG. 5 illustrates one embodiment for dividing a sequence image into subpixels (or subpixel regions). In another embodiment shown, quarter (0.25) subpixels are used, resulting in each pixel in the sequence image being divided into 16 subpixels. Assuming the sequence image shown has a resolution of 20×20 pixels, i.e., 400 pixels, the division produces 6,400 subpixels. Each of the subpixels is treated by a base caller as a region center for subpixel base calling. In some embodiments, the base caller does not use neural network-based processing. In other embodiments, the base caller is a neural network-based base caller.

[0114] For a given sequencing cycle and a particular subpixel, the base caller is configured with logic to generate base calls for the particular subpixel for the given sequencing cycle by performing image processing steps to extract the subpixel's intensity data from the corresponding image set for the sequencing cycle. This is done for each subpixel and for each of multiple sequencing cycles. Experiments were also performed using a quarter subpixel division of an 1800 x 1800 pixel resolution tile image from an Illumina MiSeq sequencer. Subpixel base calls were performed for 50 sequencing cycles and 10 tile lanes.

[0115] Figure 6 shows the preliminary center coordinates of clusters identified by the base caller during subpixel base calling. Figure 6 also shows the "origin subpixel" or "center subpixel" that contains the preliminary center coordinates.

[0116] 7 shows an example of merging subpixel base calls generated over multiple sequencing cycles to generate a so-called "cluster map" that contains cluster metadata. In the illustrated implementation, the subpixel base calls are merged using a breadth-first search approach.

[0117] Figure 8a shows an example of a cluster map generated by merging subpixel base calls. Figure 8b shows an example of a subpixel base call. Figure 8b also shows an embodiment in which a cluster map is generated by analyzing the base call sequences for each subpixel generated from the subpixel base calls.

[0118] (Sequencing image) Cluster metadata determination involves analyzing image data generated by a sequencing device 102 (e.g., Illumina's iSeq, HiSeqX, HiSeq3000, HiSeq4000, HiSeq2500, NovaSeq 6000, NextSeq, NextSeqDx, MiSeq, and MiSeqDx). The following description outlines how image data is generated and what it describes, according to one embodiment.

[0119] Base calling is the process by which the raw signal from the sequencing instrument 102, i.e., intensity data extracted from the image, is decoded into DNA sequence and quality scores. In one embodiment, the Illumina platform employs cyclic reversible termination (CRT) chemistry for base calling. This process relies on growing emerging DNA strands complementary to template DNA strands bearing modified nucleotides, tracking the emission signal of each newly added nucleotide. The modified nucleotides have a 3' removable block that anchors the fluorophore signal of the nucleotide type.

[0120] Sequencing is performed in repeated cycles, each of which involves three steps: (a) extending the transnucleotide strand by adding modified nucleotides; (b) exciting the fluorophore using one or more lasers in the optical system 104 and imaging through different filters in the optical system 104 to generate a sequence image 108; and (c) cleaving the fluorophore and removing the 3' block in preparation for the next sequencing cycle. The incorporation and imaging cycle is repeated for a specified number of sequencing cycles, defining the read length of the entire population. Using this approach, each cycle interrogates a new position along the template strand.

[0121] The Illumina platform's power stems from its ability to simultaneously run and sense millions or even billions of clusters undergoing CRT reactions. The sequencing process begins in a flow cell 202, a small glass slide that holds input DNA fragments during the sequencing process. The flow cell 202 is connected to a high-throughput optical system 104, which includes a microscope imager, excitation laser, and fluorescence filters. The flow cell 202 contains multiple chambers, called lanes 204. The lanes 204 are physically separated from one another and may contain differently tagged sequencing libraries, distinguishable without sample cross-contamination. An imager 106 (e.g., a solid-state imager such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) sensor) takes snapshots at multiple locations along the lane 204 in a series of non-overlapping regions, called tiles 206.

[0122] For example, there are 100 tiles per lane on the Illumina Genome Analyzer II and 68 tiles per lane in the Illumina HiSeq2000. Tile 206 holds hundreds of thousands to millions of clusters. An image generated from a tile with clusters shown as bright spots is shown at 208. Cluster 302 contains approximately 1,000 identical copies of the template molecule, but the clusters vary in size and shape. Clusters are grown from the template molecule by bridge amplification of the input library before sequencing runs. The purpose of amplification and cluster growth is to increase the intensity of the emitted signal, since the imaging device 106 cannot reliably sense a single fluorophore. However, because the physical distance between DNA fragments within cluster 302 is small, the imaging device 106 perceives the cluster of fragments as a single spot 302.

[0123] The output of the sequencing operation is a sequence image 108 that shows the intensity emission of clusters on a tile in the pixel domain, each for a particular combination of lane, tile, sequencing cycle, and fluorophore (208A, 208C, 208T, 208G).

[0124] In one embodiment, the biosensor comprises an array of optical sensors. The optical sensors are configured to sense information from corresponding pixel regions (e.g., reaction sites / wells / nanocells) on the detection surface of the biosensor. Analytes disposed within a pixel region are said to be associated with the pixel region, i.e., the associated analyte. During a sequencing cycle, the optical sensors corresponding to the pixel regions are configured to detect / capture / sense luminescence / photons from the associated analyte and, accordingly, generate a pixel signal for each imaged channel. In one embodiment, each imaging channel corresponds to one of a plurality of filter wavelength bands. In another embodiment, each imaging channel corresponds to one of a plurality of imaging events during a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.

[0125] The pixel signals from the optical sensors are communicated (e.g., via a communications port) to a signal processor coupled to the biosensor. For each sequencing cycle and each imaging channel, the signal processor generates an image in which pixels depict / contain / show / represent / characterize the pixel signals obtained from the corresponding optical sensors. In this manner, a pixel in the image corresponds to (i) the optical sensor of the biosensor that generated the pixel signal represented by the pixel, (ii) the analyte of interest whose radiation was detected by the corresponding optical sensor and converted into a pixel signal, and (iii) the pixel area on the detection surface of the biosensor that holds the analyte of interest.

[0126] For example, consider a sequencing operation that uses two different imaging channels: a red channel and a green channel. Then, in each sequencing cycle, the signal processor generates a red image and a green image. In this way, for a series of k sequencing cycles of a sequencing run, a sequence having k pairs of red and green images is generated as output.

[0127] Pixels in the red and green images (i.e., different imaging channels) have a one-to-one correspondence within a sequencing cycle. This means that corresponding pixels in a pair of red and green images show intensity data of the same associated analyte in different imaging channels. Similarly, pixels across pairs of red and green images have a one-to-one correspondence between sequencing cycles. This means that corresponding pixels in different pairs of red and green images show intensity data of the same associated analyte for different acquisition events / time steps (sequencing cycles) of a sequencing run.

[0128] Corresponding pixels in the red and green images (i.e., different imaging channels) can be considered as pixels of an "image per cycle," representing intensity data in a first red channel and a second green channel. An image per cycle, whose pixels depict pixel signals of a subset of the pixel area, i.e., a region (tile) of the sensing surface of the biosensor, is called a "tile per cycle image." Patches extracted from the tile per cycle image are called "image per cycle patches." In one embodiment, patch extraction is performed by an input preparer.

[0129] The image data includes a series of per-cycle image patches generated for a series of k-sequence cycles of a sequencing run. Pixels in the per-cycle image patch include intensity data for associated analytes, the intensity data being acquired for one or more imaging channels (e.g., red and green channels) by corresponding optical sensors configured to detect emissions from the associated analytes. In one embodiment, based on a single target cluster, the per-cycle image patch is centered on a central pixel that includes intensity data for the target-associated analyte and non-central pixels, and the non-central pixels in the per-cycle image patch include intensity data for associated analytes adjacent to the target-associated analyte. In one embodiment, the image data is prepared by an input preparer.

[0130] (subpixel base call) The disclosed technology accesses a series of image sets generated during a sequencing run. The image sets include sequence images 108. Each successive image set is captured during each sequencing cycle of a sequencing run. The series of images (or sequence images) captures clusters on the tiles of the flow cell and their surrounding background.

[0131] In one embodiment, the sequencing run utilizes four channel chemistry and each image set has four images. In another embodiment, the sequencing run utilizes two channel chemistry and each image set has two images. In yet another embodiment, the sequencing run utilizes one channel chemistry and each image set has two images. In yet another embodiment, each image set has only one image.

[0132] The sequence image 108 in the pixel domain is first converted to the sub-pixel domain by a sub-pixel addresser 110 to generate a sequence image 112 in the sub-pixel domain. In one embodiment, each pixel in the sequence image 108 is divided into 16 sub-pixels 502. Thus, in one embodiment, the sub-pixels 502 are quarter sub-pixels. In another embodiment, the sub-pixels 502 are half sub-pixels. As a result, each of the sequence images 112 in the sub-pixel domain has multiple sub-pixels 502.

[0133] The subpixels are then separately fed as inputs to base caller 114 to obtain base calls from base caller 114 that classify each of the subpixels as one of four bases (A, C, T, and G). This generates base call sequences 116 for each of the subpixels across multiple sequencing cycles of a sequencing run. In one embodiment, subpixels 502 are identified to base caller 114 based on their integer or non-integer coordinates. By tracking the emission signals from subpixels 502 across a set of images generated during multiple sequencing cycles, base caller 114 recovers the underlying DNA sequence of each subpixel. An example of this is shown in FIG. 8b.

[0134] In other embodiments, the disclosed technology classifies each of the subpixels as one of five bases (A, C, T, G, and N) from the base caller 114. In such embodiments, the N base calls represent undetermined base calls, typically due to low levels of extracted intensity.

[0135] Some examples of base callers 114 include non-neural network-based Illumina offerings, such as Real Time Analysis (RTA), the Firecrest program in the Genome Analyzer Analysis Pipeline, the Integrated Primary Analysis and Reporting (Ipar) machine, and the Off-Line Basecaller (OLB). For example, the base caller 114 generates base call sequences by interpolating subpixel intensities, including at least one of nearest neighbor intensity extraction, Gaussian intensity extraction, intensity extraction based on an average 2x2 subpixel area, intensity extraction based on a brightest test of a 2x2 subpixel area, intensity extraction based on an average 3x3 subpixel area, bilinear intensity extraction, bicubic intensity extraction, and / or intensity extraction based on weighted area coverage. These techniques are described in detail in the Appendix entitled "Intensity Extraction Methods."

[0136] In other embodiments, the base caller 114 may be a neural network-based base caller, such as the neural network-based base caller 1514 disclosed herein.

[0137] The base call sequences 116 for each subpixel are then provided as input to a searcher 118, which searches for substantially matching base call sequences for consecutive subpixels. The base call sequences for consecutive subpixels are "substantially matching" when a predetermined portion of the base calls match a criterion for each ordinal position (e.g., 41 matches in >= 45 cycles, 4 mismatches in <= 45 cycles, 4 mismatches in <= 50 cycles, or 2 mismatches in <= 34 cycles).

[0138] The searcher 118 then generates a cluster map 802 that identifies clusters, such as 804a-d, of adjacent subpixels that share substantially matching base call sequences. This application uses "discontiguous," "disjoint," and "non-overlapping" interchangeably. The search involves calling subpixels that comprise part of a cluster and allowing them to link the called subpixels to adjacent subpixels that share substantially matching base call sequences. In some implementations, the searcher 118 requires that at least some of the discontiguous regions have a predetermined minimum number of subpixels (e.g., more than 4, 6, or 10 subpixels) to be treated as a cluster.

[0139] In some embodiments, the base caller 114 also identifies preliminary center coordinates of the cluster. The subpixel containing the preliminary center coordinate is referred to as the origin subpixel. Some example preliminary center coordinates (604a-c) identified by the base caller 114 and corresponding origin subpixels (606a-c) are shown in FIG. 6. However, as explained below, identification of the origin subpixel (preliminary center coordinate of the cluster) is not required. In some embodiments, the searcher 118 uses a breadth-first search, starting with the origin subpixels 606a-c and continuing through successive non-origin subpixels 702a-c, to identify substantially matching base call sequences of the subpixels. This is optional, as explained below.

[0140] (cluster map) FIG. 8a shows an example of a cluster map 802 generated by merging subpixel base calls. The cluster map identifies multiple discontinuous regions (shown in various colors in FIG. 8a). Each discontinuous region includes non-overlapping groups of contiguous subpixels (from the sequence images from which the cluster map is generated via subpixel base calls) that represent respective clusters on the tile. The regions between the discontinuous regions represent the background on the tile. Subpixels within the background regions are referred to as "background subpixels." Subpixels within the discontinuous regions are referred to as "cluster subpixels" or "intra-cluster subpixels." In this description, the origin subpixel is the subpixel at which the preliminary center cluster coordinates, determined by the RTA or another base caller, are located.

[0141] The origin subpixel contains the preliminary center cluster coordinate, meaning that the area covered by the origin subpixel contains a coordinate location that coincides with the preliminary center cluster coordinate location. Because the cluster map 802 is an image of logical subpixels, the origin subpixel is a subset of the subpixels in the cluster map.

[0142] The search to identify clusters having substantially matching base call sequences for subpixels does not need to begin with identifying an origin subpixel (a preliminary center coordinate for the cluster), because the search can be performed for all subpixels and can begin at any subpixel (e.g., the 0,0 subpixel or any random subpixel). Thus, because each subpixel is evaluated to determine whether it shares a substantially matching base call sequence with another contiguous subpixel, the search does not depend on the origin subpixel, and the search can begin at any subpixel.

[0143] Regardless of whether the origin subpixel is used, certain clusters that do not include the origin subpixel (the initial center coordinate of the cluster) are identified by the base caller 114. Some examples of clusters that are identified by merging subpixel base calls and do not include the origin subpixel are clusters 812a, 812b, 812c, 812d, and 812e in FIG. 8a. Thus, the disclosed techniques identify additional or extra clusters whose centers may not have been identified by the base caller 114. Therefore, the use of the base caller 114 to identify the origin subpixel (the initial center coordinate of the cluster) is optional and not required to search for substantially matching base call sequences of consecutive subpixels.

[0144] In one embodiment, the origin subpixels (initial center coordinates of the clusters) identified by base caller 114 are first used to identify a first set of clusters (by identifying substantially matching base call sequences of consecutive subpixels). Subpixels that are not part of the first set of clusters are then used to identify a second set of clusters (by identifying substantially matching base call sequences of consecutive subpixels). This enables the disclosed techniques to identify additional or extra clusters whose centers are not identified by base caller 114. Finally, subpixels that are not part of the first and second sets of clusters are identified as background subpixels.

[0145] Figure 8b shows an example of sub-pixel base calling, where each sequencing cycle has an image set with four different images (i.e., A, C, T, and G images) captured using four different wavelength bands (image / imaging channels) and four different fluorescent dyes (one for each base).

[0146] In this example, a pixel in an image is divided into 16 subpixels. The subpixels are then base called separately by the base caller 114 in each sequencing cycle. To base call a given subpixel in a particular sequencing cycle, the base caller 114 uses the intensity of the given subpixel in each of the four A, C, T, and G images. For example, the intensity of the image area covered by subpixel 1 in each of the four A, C, T, and G images in cycle 1 is used to base call subpixel 1 in cycle 1. For subpixel 1, these image areas include the upper-left 1 / 16 area of ​​the upper-left pixel in each of the four A, C, T, and G images in cycle 1. Similarly, the intensity of the image area covered by subpixel m in each of the four A, C, T, and G images in cycle n is used to determine subpixel m in cycle n. For subpixel m, these image areas include the lower-right 1 / 16 area of ​​the lower-right pixel in each of the four A, C, T, and G images in cycle 1.

[0147] This process generates base call sequences 116 for each subpixel over multiple sequencing cycles. Searcher 118 then evaluates pairs of consecutive subpixels to determine whether they have substantially matching base call sequences. If yes, the pair of subpixels is stored in cluster map 802 as belonging to the same cluster within a discontinuous region. If no, the pair of subpixels is stored in cluster map 802 as not belonging to the same discontinuous region. Thus, cluster map 802 identifies consecutive sets of subpixels for which the base calls for the subpixels substantially match over multiple cycles. Cluster map 802 thus uses information from the multiple clusters to provide multiple clusters, each of which has a high degree of confidence in providing sequence data for a single DNA strand.

[0148] The cluster metadata generator 122 then processes the cluster map 802 to determine cluster metadata, which includes determining the spatial distribution of the clusters, including their centers (810a), shapes, sizes, backgrounds, and / or boundaries (FIG. 9).

[0149] In some implementations, the cluster metadata generator 122 identifies subpixels in the cluster map 802 as background that do not belong to any of the disjoint regions and therefore do not contribute to any clusters. Such subpixels are referred to as background subpixels 806a-c.

[0150] In some embodiments, the cluster map 802 identifies cluster boundaries 808a-c between two consecutive subpixels where the base call sequences do not substantially match.

[0151] The cluster map is stored in memory (e.g., cluster map data store 120) for use as ground truth for training classifiers such as neural network-based template generator 1512 and neural network-based base caller 1514. Cluster metadata may also be stored in memory (e.g., cluster metadata data store 124).

[0152] FIG. 9 shows another example of a cluster map that identifies cluster metadata including the spatial distribution of clusters, along with cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and / or cluster boundaries.

[0153] (Center of mass (COM)) Figure 10 shows how the center of mass (COM) of a discontinuous region in a cluster map is calculated, which can be used as the "corrected" or "improved" center of the corresponding cluster in downstream processing.

[0154] In some implementations, for each cluster, a center of mass calculator 1004 determines the cluster's superlocation center coordinate 1006 by calculating the center of mass of the discontinuous regions of the cluster map as the average of the coordinates of each contiguous subpixel that forms the discontinuous region. The cluster's superlocation center coordinates are then stored for each cluster in memory for use as ground truth for training the classifier.

[0155] In some implementations, the subpixel classifier identifies, for each cluster, a center-of-mass subpixel 1008 within the discontinuous regions 804a-d of the cluster map 802 at the cluster's superposition center coordinates 1006.

[0156] In another alternative embodiment, the cluster map is upsampled using interpolation, and the upsampled cluster map is stored in memory for use as ground truth for training the classifier.

[0157] (Attenuation coefficient and attenuation map) FIG. 11 illustrates one implementation of a calculation of a weighted attenuation coefficient for a subpixel based on the Euclidean distance from the subpixel to the center of mass (COM) of the discontinuous region to which the subpixel belongs. In another illustrated implementation, the weighted attenuation coefficient is assigned its highest value to the subpixel containing the COM and decreases for subpixels further away from the COM. The weighted attenuation coefficient is used to derive a ground truth attenuation map 1204 from the cluster map generated from the subpixel base calls described above. The ground truth attenuation map 1204 includes an array of units, with each unit in the array assigned at least one output value. In some implementations, the units are subpixels, and each subpixel is assigned an output value based on the weighted attenuation coefficient. The ground truth attenuation map 1204 is then used as ground truth for training the disclosed neural network-based template generator 1512. In some implementations, information from the ground truth attenuation map 1204 is also used to prepare the input of the disclosed neural network-based base caller 1514.

[0158] 12 shows one implementation of an exemplary ground truth attenuation map 1204 derived from an exemplary cluster map generated by subpixel base calling as described above. In some implementations, in the cluster-by-cluster upsampled cluster map, for each cluster, a value is assigned to each contiguous subpixel within a discrete region based on an attenuation coefficient 1102 that is proportional to the distance 1106 of the contiguous subpixel from the center-of-mass subpixel 1104 within the discrete region to which the adjacent subpixel belongs.

[0159] 12 shows a ground truth attenuation map 1204. In one implementation, the subpixel values ​​are intensity values ​​normalized between zero and one. In another implementation, in the upsampled cluster map, all subpixels identified as background are assigned the same predetermined value. In some implementations, the predetermined value is a zero intensity value.

[0160] In some implementations, the ground truth attenuation map 1204 is generated by the ground truth attenuation map generator 1202 from the upsampled cluster map, which represents contiguous subpixels in discontinuous regions based on their assigned values. The ground truth attenuation map 1204 is stored in memory for use as ground truth for training the classifier. In one implementation, each subpixel in the ground truth attenuation map 1204 has a normalized value between zero and one.

[0161] (Three-class map) FIG. 13 illustrates one implementation of deriving a ground truth ternary map 1304 from a cluster map. The ground truth ternary map 1304 includes an array of units and assigns at least one output value to each unit in the array. By name, the ternary map implementation of the ground truth ternary map 1304 assigns three output values ​​to each unit in the array, such that for each unit, the first output value corresponds to the classification label or score of the background class, the second output value corresponds to the classification label or score of the cluster center class, and the third output value corresponds to the classification label or score of the cluster / cluster inner class. The ground truth ternary map 1304 is used as ground truth data for training the neural network-based template generator 1512. In some implementations, information from the ground truth ternary map 1304 is also used to prepare the input of the neural network-based base caller 1514.

[0162] 13 shows an exemplary ground truth ternary map 1304. In another implementation, in the upsampled cluster map, contiguous subpixels in discontinuous regions are classified by cluster by the ground truth ternary map generator 1302 as cluster interior subpixels that belong to the same cluster, center-of-mass subpixels as cluster center subpixels, and background subpixels as subpixels that do not belong to any cluster. In some implementations, the classifications are stored in the ground truth ternary map 1304. These classifications and the ground truth ternary map 1304 are stored in memory for use as ground truth for training a classifier.

[0163] In another alternative embodiment, for each cluster, the coordinates of the cluster interior subpixel, cluster center subpixel, and background subpixel are stored in memory for use as ground truth for training the classifier. The coordinates are then downscaled by the factor used to upsample the cluster map. The downscaled coordinates for each cluster are then stored in memory for use as ground truth for training the classifier.

[0164] In yet another embodiment, the ground truth ternary map generator 1302 uses the cluster map to generate ternary ground truth data 1304 from the upsampled cluster map. The ternary ground truth data 1304 labels background subpixels as belonging to a background class, cluster center subpixels as belonging to a cluster center class, and cluster interior subpixels as belonging to a cluster interior class. In some visualization embodiments, color coding can be used to depict and distinguish between different class labels. The ternary ground truth data 1304 is stored in memory for use as ground truth for training a classifier.

[0165] (Binary (2-class) map) FIG. 14 illustrates one embodiment of deriving a ground truth binary map 1404 from a cluster map. The binary map 1404 includes an array of units and assigns at least one output value to each unit in the array. By name, the binary map assigns two output values ​​to each unit in the array, such that for each unit, the first output value corresponds to the classification label or score of the cluster center class and the second output value corresponds to the classification label or score of a non-center class. The binary map is used as ground truth data for training a neural network-based template generator 1512. In some embodiments, information from the binary map is also used to prepare input for the neural network-based base caller 1514.

[0166] 14 shows a ground truth binary map 1404. The ground truth binary map generator 1402 uses the cluster map 120 to generate binary ground truth data 1404 from the upsampled cluster map. The binary ground truth data 1404 labels cluster center subpixels as belonging to a cluster center class and labels all other subpixels as belonging to a non-center class. The binary ground truth data 1404 is stored in memory for use as ground truth for training a classifier.

[0167] In some embodiments, the disclosed technology generates a cluster map 120 for multiple tiles of a flow cell, stores the cluster map in memory, and determines the spatial distribution of clusters within the tile based on the cluster map 120, including their shapes and sizes. The disclosed technology then classifies the subpixels for each cluster in the upsampled cluster map 120 of the clusters within the tile into cluster interior subpixels, cluster center subpixels, and background subpixels that belong to the same cluster. The disclosed technology then stores the classification in memory for use as ground truth for training a classifier, and for each cluster within the cluster, stores the coordinates of the cluster interior subpixels, cluster center subpixels, and background subpixels in memory for use as ground truth for training the classifier. The disclosed technology then downscales the coordinates by the factor used to upsample the cluster map and stores the downscaled coordinates for each cluster in memory for use as ground truth for training the classifier.

[0168] In some embodiments, the flow cell has at least one patterned surface with an array of wells that occupy clusters. In such embodiments, based on the determined shapes and sizes of the clusters, the disclosed techniques identify (1) which wells are substantially occupied by at least one group, (2) which wells are minimally occupied, and (3) which wells are co-occupied by multiple groups. This allows for the determination of metadata for each of multiple clusters that co-occupy the same well, i.e., the center, shape, and size of two or more clusters that share the same well.

[0169] In some embodiments, the solid support on which a sample is amplified into clusters comprises a patterned surface. A "patterned surface" refers to the arrangement of distinct regions within or on an exposed layer of a solid support. For example, one or more regions can be features where one or more amplification primers are present. The features can be separated by interstitial regions where no amplification primers are present. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repetitive sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random sequence of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. No. 8,778,849, U.S. Pat. No. 9,079,148, U.S. Pat. No. 8,778,848, and U.S. Patent Application Publication No. 2014 / 0243224, each of which is incorporated herein by reference.

[0170] In some embodiments, the solid support comprises an array of wells or depressions on its surface, which can be fabricated as commonly known in the art using a variety of techniques, including but not limited to photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the technique used will depend on the composition and shape of the array substrate.

[0171] The features within the patterned surface may be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM; see, e.g., U.S. Patent Application Publication Nos. 2013 / 184796, WO 2016 / 066586, and 2015-002813, each of which is incorporated by reference in its entirety). This process creates a gel pad used for sequencing, which can be stable over multiple cycles and across sequencing runs. Covalently attaching a polymer to the wells is useful for maintaining the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel need not be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477, which is incorporated herein by reference in its entirety), which is not covalently bonded to any portion of the structured substrate, can be used as the gel material.

[0172] In certain other embodiments, structured substrates can be created by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azido-SFA version, and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can then be attached to the gel material. A solution of target nucleic acids (e.g., a fragmented human genome) can then be contacted with the polished substrate such that individual target nucleic acids seed individual wells through interaction with primers bound to the gel material, but the target nucleic acids do not occupy the interstitial regions due to the inactivity or inactivity of the gel material. Amplification of the target nucleic acids will be confined to the wells because the absence or inactivity of gel in the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is conveniently manufacturable and scalable, utilizing micro- or nano-fabrication methods.

[0173] As used herein, the term "flow cell" refers to a chamber that includes a solid surface through which one or more fluidic reagents can flow. Examples of flow cells and related fluidic systems and detection platforms that can be easily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, U.S. Patent No. 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, 7,405,281, and 2008 / 0108082, each of which is incorporated herein by reference.

[0174] Throughout this disclosure, the terms "P5" and "P7" are used when referring to amplification primers. It will be understood that any suitable amplification primers can be used in the methods presented herein, and the use of P5 and P7 is exemplary only. The use of amplification primers such as P5 and P7 on flow cells is known in the art, as exemplified by the disclosures of International Publication Nos. WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957, the entire contents of which are incorporated herein by reference. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful in the methods presented herein for amplifying complementary sequences and sequences. Similarly, any suitable reverse amplification primer can be useful in the methods provided herein for complementary sequences and amplification of sequences, whether immobilized or in solution. Those skilled in the art will understand how to design and use suitable primer sequences for capture and amplification of nucleic acids provided herein.

[0175] In some embodiments, the flow cell has at least one non-patterned surface, and the clusters are scattered non-uniformly on the non-patterned surface.

[0176] In some embodiments, the density of the clusters is about 100,000 clusters / mm 2 ~approximately 1,000,000 clusters / mm 2 In another embodiment, the density of the clusters is in the range of about 1,000,000 clusters / mm 2 ~approximately 10,000,000 clusters / mm 2 The range is.

[0177] In one embodiment, the preliminary center coordinates of the clusters determined by the base caller are defined within a template image of the tile, hi some embodiments, the pixel resolution, image coordinate system, and measurement scale of the image coordinate system are the same as the template image and the image.

[0178] In another embodiment, the disclosed technology relates to determining metadata about clusters on tiles of a flow cell. First, the disclosed technology accesses (1) a set of images of the tile captured during a sequencing run and (2) preliminary center coordinates of the clusters determined by a base caller.

[0179] Then, for each image set, the disclosed technique obtains one of four bases: (1) an origin subpixel containing preliminary center coordinates and (2) a predetermined neighborhood of consecutive subpixels contiguous to each of the origin subpixels. This generates a base call sequence for each origin subpixel and each predetermined neighborhood of consecutive subpixels. The predetermined neighborhood of consecutive subpixels can be an m×n subpixel patch centered on a subpixel containing the origin subpixel. In one embodiment, the subpixel patch is 3×3 subpixels. In other embodiments, the image patch can be any size, such as 5×5, 15×15, or 20×20. In other embodiments, the predetermined neighborhood of consecutive subpixels can be an n-connected subpixel neighborhood centered on a subpixel containing the origin subpixel.

[0180] In one embodiment, the disclosed technique identifies as background those sub-pixels in the cluster map that do not belong to any of the disjoint regions.

[0181] The disclosed technique then generates a cluster map that identifies clusters as discontinuous regions of adjacent subpixels that (a) are contiguous to at least a portion of a corresponding one of the origin subpixels and (b) share a substantially matching base call sequence of one of the four bases with at least a portion of a corresponding one of the origin subpixels.

[0182] The disclosed technique then stores the cluster map in memory and determines the shape and size of the clusters based on the discontinuous regions in the cluster map. In other embodiments, the centers of the clusters are also determined.

[0183] (Generating training data for the template generator) FIG. 15 is a block diagram illustrating one embodiment for generating the training data used to train the neural network-based template generator 1512 and the neural network-based base caller 1514.

[0184] 16 illustrates characteristics of the disclosed training examples used to train the neural network-based template generator 1512 and the neural network-based base caller 1514. Each training example corresponds to a tile and is labeled with a corresponding ground truth data representation. In some implementations, the ground truth data representation is a ground truth mask or ground truth map that identifies ground truth cluster metadata in the form of a ground truth attenuation map 1204, a ground truth ternary map 1304, or a ground truth binary map 1404. In some implementations, multiple training examples correspond to the same tile.

[0185] In one embodiment, the disclosed technology relates to generating training data 1504 for neural network-based template generation and base calling. First, the disclosed technology accesses multiple images 108 of a flow cell 202 captured over multiple cycles of a sequencing run. The flow cell 202 has multiple tiles. In the multiple images 108, each tile has a series of image sets generated over multiple cycles. Each image in the sequence of image sets 108 shows the intensity emission of a particular cluster 302 of tiles and their surrounding background 304 at a particular cycle.

[0186] The training set constructor 1502 then constructs a training set 1504 having a plurality of training examples. As shown in FIG. 16 , each training example corresponds to a particular one of the tiles and includes image data from at least some of the image sets in the sequence of image sets 1602 for that particular one of the tiles. In one implementation, the image data includes images from at least some of the image sets in the sequence of image sets 1602 for that particular one of the tiles. For example, the images may have a resolution of 1800×1800. In other implementations, the images may have any resolution, such as 100×100, 3000×3000, 10000×10000, etc. In yet other implementations, the image data includes at least one image patch from each of the images. In one implementation, the image patch covers a portion of a particular one of the tiles. In one example, the image patch may have a resolution of 20×20. In other embodiments, the image patches can have any resolution, such as 50x50, 70x70, 90x90, 100x100, 3000x3000, 10000x10000, and the like.

[0187] In some implementations, the image data includes an upsampled representation of the image patch. The upsampled representation may have a resolution of, for example, 80 x 80. In other implementations, the upsampled representation may have any resolution, such as 50 x 50, 70 x 70, 90 x 90, 100 x 100, 3000 x 3000, 10000 x 10000, etc.

[0188] In some embodiments, the multiple training examples correspond to the same particular one of the tiles and each include a different image patch from a respective image of at least some of the image sets in the sequence of image sets 1602 for the same particular one of the tiles. In such implementations, at least some of the different image patches overlap one another.

[0189] The ground truth generator 1506 then generates at least one ground truth data representation for each of the training examples, which identifies the spatial distribution of the clusters and at least one of their surrounding background as represented by the image data, including at least one of the cluster shapes, cluster sizes, and / or cluster boundaries, and / or cluster centers.

[0190] In one implementation, the ground truth data representation identifies clusters as discontinuous regions of adjacent sub-pixels, with the centers of the clusters being the centers of the mass sub-pixels within corresponding ones of the discontinuous regions, and their surrounding background.

[0191] In one implementation, the ground truth data representation has an upsampled resolution of 80 x 80. In other implementations, the ground truth data representation can have any resolution, such as 50 x 50, 70 x 70, 90 x 90, 100 x 100, 3000 x 3000, 10000 x 10000, etc.

[0192] In one embodiment, the ground truth data representation identifies each sub-pixel as being either cluster-centered or non-centered, hi another embodiment, the ground truth data representation identifies each sub-pixel as being either cluster-interior, cluster-centered, or surrounding background.

[0193] In some implementations, the disclosed technology stores a training set 1504 and associated ground truth data 1508 in memory as training data 1504 for training a neural network-based template generator 1512 and a neural network-based base caller 1514. The training is handled by a trainer 1510.

[0194] In some embodiments, the disclosed techniques generate training data for a variety of flow cells, sequencing instruments, sequencing protocols, sequencing chemistries, sequencing reagents, and cluster densities.

[0195] (Neural network-based template generator) In inference or production embodiments, the disclosed techniques use peak detection and segmentation to determine cluster metadata. The disclosed techniques process input image data 1702 derived from a series of image sets 1602 via a neural network 1706 to generate alternative representations 1708 of the input image data 1702. For example, an image set may be for a particular sequence cycle and include four images, one for each image channel A, C, T, and G. Thus, for a sequence that runs for 50 sequence cycles, there will be 50 such image sets, for a total of 200 images. When arranged in time, the image sets, with four image patches per image set, form the series of image sets 1602. In some implementations, image patches of a specific size are extracted from each image in the 50 image sets to form four image patch sets per image patch set, which in one implementation is the input image data 1702. In other implementations, the input image data 1702 includes image patch sets having four image patches per image patch set for image patch sets with fewer than 50 sequencing cycles, i.e., fewer than 1, 2, 3, 15, or 20 sequencing cycles.

[0196] 17 illustrates one implementation of processing input image data 1702 through a neural network-based template generator 1512 to generate an output value for each unit in an array. In one implementation, the array is an attenuation map 1716. In another implementation, the array is a ternary map 1718. In yet another implementation, the array is a binary map 1720. Thus, the array may represent one or more characteristics of each of multiple locations represented in the input image data 1702.

[0197] Unlike training the template generator using the structure of the previous figure, including the ground truth attenuation map 1204, the ground truth ternary map 1304, and the ground truth binary map 1404, the attenuation map 1716, the ternary map 1718, and / or the binary map 1720 are generated by forward propagation of the trained neural network-based template generator 1512. The forward propagation can be during training or during inference. During training, backpropagation-based gradient updates cause the attenuation map 1716, the ternary map 1718, and the binary map 1720 (i.e., cumulatively the output 1714) to progressively match or approach the ground truth attenuation map 1204, the ground truth ternary map 1304, and the ground truth binary map 1404, respectively.

[0198] The size of the image array analyzed during inference, according to one embodiment, depends on the size of the input image data 1702 (e.g., the same or an upscaled or downscaled version). Each unit can represent a pixel, a subpixel, or a superpixel. The output value per unit of the array can characterize / represent / indicate an attenuation map 1716, a ternary map 1718, or a binary map 1720. In some implementations, the input image data 1702 is also an array of units at pixel, subpixel, or superpixel resolution. In another such implementation, the neural network-based template generator 1512 uses semantic segmentation techniques to generate output values ​​for each unit in the input array. Further details regarding the input image data 1702 can be found in Figures 21b, 22, 23, and 24 and their discussions.

[0199] In some embodiments, the neural network-based template generator 1512 is a fully convolutional network, such as that described in J. Long, E. Shelhamer, and T. Darrell, "Fully convolutional networks for semantic segmentation," CVPR, (2015), incorporated herein by reference. In other embodiments, the neural network-based template generator 1512 is a U-Net network with skip connections between the decoder and the encoder, such as that described in Ronneberger O, Fischer P, Brox T, "U-net: Convolutional networks for biomedical image segmentation," Med. Image Comput. Comput. Assist. Interv. (2015), available at http: / / link.springer.com / chapter / 10.1007 / 978-3-319-24574-4_28, incorporated herein by reference. The U-Net structure resembles an autoencoder with two main substructures: The system comprises: 1) an encoder that takes an input image and reduces its spatial resolution through multiple convolutional layers to generate an encoding; and 2) a decoder that takes the representation, encoding and increasing the spatial resolution, to generate a reconstructed image as output. U-Net introduces two innovations to this structure: first, the objective function is set to reconstruct a segmentation mask using a loss function; and second, the convolutional layers of the encoder are connected to corresponding layers of the same resolution in the decoder using skip connections. In yet a further embodiment, the neural network-based template generator 1512 is a deep fully convolutional segmentation neural network having an encoder subnetwork and a corresponding decoder network. In another such implementation, the encoder subnetwork includes a hierarchy of encoders, and the decoder subnetwork includes a hierarchy of decoders that map low-resolution encoder feature maps to full-input-resolution feature maps.Further details regarding segmentation networks can be found in the appendix entitled "Segmentation Networks."

[0200] In one embodiment, neural network-based template generator 1512 is a convolutional neural network. In another embodiment, neural network-based template generator 1512 is a recurrent neural network. In yet another embodiment, neural network-based template generator 1512 is a residual neural network having residual Box and residual connections. In a further embodiment, neural network-based template generator 1512 is a combination of a convolutional neural network and a recurrent neural network.

[0201] It will be appreciated that the neural network-based template generator 1512 (i.e., the neural network 1706 and / or the output layer 1710) can use various padding and striding configurations. It can use different output functions (e.g., classification or regression) and may or may not include one or more fully connected layers. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or asexual convolution, transposed convolution, depth-separable convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It can use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous SGD. It includes nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, Pejol connections, activation functions (e.g., nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), sigmoid, and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.

[0202] In some embodiments, each image in the sequence of image sets 1602 covers a tile and shows the intensity emission of a cluster on the tile and their surrounding background, captured for a particular imaging channel during a particular one of multiple sequencing cycles of a sequencing run performed on the flow cell. In one embodiment, input image data 1702 includes at least one image patch from each of the images in the sequence of image sets 1602. In another such embodiment, the image patch covers a portion of a tile. In one example, the image patch has a resolution of 20x20. In other cases, the resolution of the image patch may range from 20x20 to 10000x10000. In another embodiment, input image data 1702 includes upsampled subpixel resolution representations of image patches from each of the images in the sequence of image sets 1602. In one example, the upsampled subpixel representation has a resolution of 80x80. In other cases, the resolution of the upsampled subpixel representation may range from 80x80 to 10000x10000.

[0203] The input image data 1702 has an array of units 1704 that describe clusters and their surrounding background. For example, an image set may be for a particular sequence cycle and include four images, one for each image channel A, C, T, and G. Thus, for a sequence that runs for 50 sequence cycles, there will be 50 such image sets, for a total of 200 images. When arranged in time, the image sets, with four image patches per image set, form the sequence of image sets 1602. In some implementations, image patches of a particular size are extracted from each image in the 50 image sets to form 50 image patch sets, with four image patches per image patch set, which in one implementation is the input image data 1702. In other implementations, the input image data 1702 includes image patch sets with four image patches per image patch set for fewer than 50 sequencing cycles, i.e., fewer than 1, 2, 3, 15, or 20 sequencing cycles. An alternative representation is a feature map. The feature maps may be convolutional features or convolutional representations when the neural network is a convolutional neural network. The feature maps may be hidden state features or hidden state representations when the neural network is a recurrent neural network.

[0204] The disclosed technique then processes the alternative representation 1708 through an output layer 1710 to generate an output 1714 having an output value 1712 for each unit in the array 1704. The output layer may be a classification layer, such as a softmax or sigmoid, that generates a per-unit output value. In one implementation, the output layer is a ReLU layer or any other activation function layer that generates a per-unit output value.

[0205] In one embodiment, the units in the input image data 1702 are pixels, and therefore an output value 1712 for each pixel is generated at the output 1714. In another embodiment, the units in the input image data 1702 are sub-pixels, and therefore an output value 1712 for each sub-pixel is generated at the output 1714. In yet another embodiment, the units in the input image data 1702 are super-pixels, and therefore an output value 1712 for each super-pixel is generated at the output 1714. Deriving cluster metadata from attenuation maps, ternary maps and / or binary maps

[0206] 18 illustrates one embodiment of a post-processing technique applied to the attenuation map 1716, ternary map 1718, or binary map 1720 generated by the neural network-based template generator 1512 to derive cluster metadata including cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and / or cluster boundaries. In some embodiments, the post-processing technique is applied by a post-processor 1814, which further includes a threshold holder 1802, a peak locator 1806, and a divider 1810.

[0207] The input to the thresholder 1802 is an attenuation map 1716, a ternary map 1718, or a binary map 1720 generated by a template generator 1512, such as the disclosed neural network-based template generator. In one implementation, the thresholder 1802 applies a threshold to values ​​in the attenuation map, ternary map, or binary map to identify background units 1804 (i.e., subpixels that characterize non-cluster background) and non-background units. In other words, once the output 1714 is generated, the thresholder 1802 applies a threshold to the output values ​​of the units 1712 to classify or reclassify a first subset of the units 1712 as “background units” 1804 that depict the background surrounding the cluster and “non-background units” that represent units that may belong to the cluster. The threshold applied by the thresholder 1802 may be preset.

[0208] The input to the peak locator 1806 is also the attenuation map 1716, the ternary map 1718, or the binary map 1720 generated by the neural network-based template generator 1512. In one embodiment, the peak locator 1806 applies peak detection of values ​​in the attenuation map 1716 to the ternary map 1718 or the binary map 1720 to identify center units 1808 (i.e., center subpixels that characterize cluster centers). In other words, the peak locator 1806 processes the output values ​​of the units 1712 in the output 1714 and classifies a second subset of the units 1712 as "center units" 1808 that contain the centers of the clusters. In some embodiments, the centers of the clusters detected by the peak locator 1806 are also the centers of mass of the clusters. The center units 1808 are then provided to a divider 1810. Further details regarding the peak locator 1806 can be found in the appendix entitled "Peak Detection."

[0209] Thresholding and peak detection can occur in parallel or after one another, i.e., they are not dependent on each other.

[0210] The input to the divider 1810 is also the attenuation map 1716, ternary map 1718, or binary map 1720 generated by the neural network-based template generator 1512. Additional supplemental inputs to the divider 1810 include the thresholded units (background, non-background) 1804 identified by the thresholder 1802 and the center unit 1808 identified by the peak locator 1806. The divider 1810 uses the background, non-background 1804, and center unit 1808 to identify discontinuous regions 1812 (i.e., non-overlapping groups of adjacent clusters / intra-cluster sub-pixels that characterize clusters). In other words, the divider 1810 processes the output values ​​of the units 1712 in the output 1714 and uses the background and non-background units 1804, as well as the center unit 1808, to determine the shape of the cluster 1812 as a non-overlapping region of contiguous units separated by the background unit 1804 and centered on the center unit 1808. The output of the divider 1810 is cluster metadata 1812. The cluster metadata 1812 identifies cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and / or cluster boundaries.

[0211] In one embodiment, divider 1810 starts with central unit 1808 and, for each central unit, determines a group of consecutive units that represent the same cluster whose center of mass is contained in the central unit. In one embodiment, divider 1810 uses a so-called "watershed" segmentation technique to subdivide consecutive clusters into multiple adjacent clusters at valleys in intensity. Further details regarding watershed segmentation techniques and other segmentation techniques can be found in the Appendix entitled "Watershed Segmentation."

[0212] In one implementation, the output values ​​of unit 1712 in output 1714 are continuous values, such as those encoded in ground truth attenuation map 1204. In another implementation, the output values ​​are softmax scores, such as those encoded in ground truth ternary map 1304 and ground truth binary map 1404. In one implementation of ground truth attenuation map 1204, consecutive units in corresponding ones of the non-overlapping regions have output values ​​weighted according to the consecutive units' distance from the central unit in the non-overlapping region to which the adjacent units belong. In such an implementation, the central unit has the highest output value in each one of the non-overlapping regions. As described above, during training, backpropagation-based gradient updates cause attenuation map 1716, ternary map 1718, and binary map 1720 (i.e., cumulatively output 1714) to progressively match or approach ground truth attenuation map 1204, ground truth ternary map 1304, and ground truth binary map 1404, respectively.

[0213] (Pixel domain - intensity extraction from regular cluster shapes) The discussion now turns to how the cluster shapes determined by the disclosed techniques can be used to extract the intensities of the clusters. Because clusters typically have irregular shapes and contours, the disclosed techniques can be used to identify which sub-pixels contribute to the irregularly shaped, discontinuous regions that represent the cluster shape.

[0214] 19 illustrates one implementation of extracting cluster intensities in the pixel domain. A "template image" or "template" can refer to a data structure that includes or identifies cluster metadata 1812 derived from an attenuation map 1716, a ternary map 1718, and / or a binary map 1718. The cluster metadata 1812 identifies cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and / or cluster boundaries.

[0215] In some implementations, the template image is in the upsampled sub-pixel domain, distinguishing cluster boundaries at a fine level. However, the sequence image 108 containing cluster and background intensity data is typically in the pixel domain. Therefore, the disclosed technology proposes two approaches to extract intensities of irregularly shaped clusters from the optical pixel resolution sequence image using cluster shape information encoded in the upsampled sub-pixel resolution template image. In the first approach, shown in FIG. 19, non-overlapping groups of contiguous sub-pixels identified in the template image are located in the pixel resolution sequence image, and their intensities are extracted by interpolation. Further details regarding this intensity extraction technique can be found in FIG. 33 and its discussion.

[0216] In one implementation, when the non-overlapping regions have irregular contours and the units are sub-pixels, the cluster intensity 1912 of a given cluster is determined by intensity extractor 1902 as follows:

[0217] First, subpixel locator 1904 identifies the subpixels that contribute to the cluster intensity of a given cluster based on corresponding non-overlapping regions of adjacent subpixels that identify the shape of the given cluster.

[0218] The subpixel locator 1904 then locates the identified subpixels within one or more optical pixel resolution images 1918 generated for one or more imaging channels in the current sequencing cycle. In one embodiment, integer or non-integer coordinates (e.g., floating points) are located within the optical resolution image, the pixel resolution image, after downscaling based on a downscaling factor that matches the upsampling factor used to create the subpixel domain.

[0219] An interpolator and subpixel intensity combiner 1906 for the identified subpixels in the processed images then combines the interpolated intensities and normalizes the combined interpolated intensities to generate a cluster intensity per image for a given cluster in each of the images. The normalization is performed by normalizer 1908 and is based on a normalization factor. In one embodiment, the normalization factor is the number of identified subpixels. This is done to normalize / account for different cluster sizes and non-uniform illumination that clusters receive depending on their location on the flow cell.

[0220] Finally, a cross-channel sub-pixel intensity accumulator 1910 combines the per-image cluster intensities for each of the images to determine a cluster intensity 1912 for a given cluster in the current sequence cycle.

[0221] The given cluster is then base called based on the cluster strength 1912 in the current sequencing cycle by any one of the base calls discussed in this application to generate base call 1916.

[0222] However, in some implementations, when the cluster size is large enough, the outputs of the neural network-based base caller 1514, i.e., the attenuation map 1716, the ternary map 1718, and the binary map 1720, are in the optical pixel domain. Thus, in such implementations, the template image is also in the optical pixel domain.

[0223] (Subpixel domain - intensity extraction from regular cluster shapes) FIG. 20 illustrates a second approach to extracting cluster intensities in the subpixel domain. In this second approach, the sequence images are optically upsampled from pixel resolution to subpixel resolution. This results in a correspondence between the cluster shapes representing the subpixels in the template image and the cluster intensities representing the subpixels in the upsampled sequence images. Cluster intensities are then extracted based on the correspondence. Further details regarding this intensity extraction technique can be found in FIG. 33 and its discussion.

[0224] In one implementation, when the non-overlapping regions have irregular contours and the units are sub-pixels, the cluster intensity 2012 of a given cluster is determined by the intensity extractor 2002 as follows:

[0225] First, the subpixel locator 2004 identifies the subpixels that contribute to the cluster intensity of a given cluster based on corresponding non-overlapping regions of adjacent subpixels that define the shape of the given cluster.

[0226] The subpixel locator 2004 then locates the identified subpixels in one or more subpixel resolution images 2018 that are upsampled from the corresponding optical pixel resolution images 1918 generated for one or more imaging channels in the current sequencing cycle. The upsampling can be performed by nearest neighbor intensity extraction, Gaussian intensity extraction, intensity extraction based on averaging 2x2 subpixel areas, intensity extraction based on brightest test of 2x2 subpixel areas, average 3x3 subpixel areas, bilinear intensity extraction, dual intensity extraction, and / or intensity extraction based on weighted region coverage. These techniques are described in detail in the Appendix entitled "Intensity Extraction Methods." The template image can, in some embodiments, serve as a mask for intensity extraction.

[0227] A subpixel intensity combiner 2006 in each of the upsampled images then combines the intensities of the identified subpixels and normalizes the combined intensities to generate a cluster intensity per image for a given cluster in each of the upsampled images. The normalization is performed by a normalizer 2008 and is based on a normalization factor. In one embodiment, the normalization factor is the number of identified subpixels. This is done to normalize / account for different cluster sizes and non-uniform illumination that clusters receive depending on their location on the flow cell.

[0228] Finally, a cross-channel sub-pixel intensity accumulator 2010 combines the cluster intensities per image for each upsampled image to determine a cluster intensity 2012 for a given cluster in the current sequence cycle.

[0229] The given cluster is then called based on the cluster strength 2012 in the current sequencing cycle by any one of the base calls discussed in this application to generate a base call 2016.

[0230] (A type of neural network-based template generator) This discussion details three different implementations of the neural network-based template generator 1512, shown in Figure 21a, including (1) an attenuation map-based template generator 2600 (also called a regression model), (2) a binary map-based template generator 4600 (also called a binary classification model), and (3) a ternary map-based template generator 5400 (also called a ternary classification model).

[0231] In one embodiment, regression model 2600 is a fully convolutional network. In another embodiment, regression model 2600 is a U-Net network with skip connections between the decoder and encoder. In one implementation, binary classification model 4600 is a fully convolutional network. In another embodiment, binary classification model 4600 is a U-Net network with skip connections between the decoder and encoder. In one implementation, ternary classification model 5400 is a fully convolutional network. In another embodiment, ternary classification model 5400 is a U-Net network with skip connections between the decoder and encoder.

[0232] (Input image data) 21b illustrates one embodiment of input image data 1702 provided as input to the neural network-based template generator 1512. The input image data 1702 includes a series of image sets 2100 having sequence images 108 generated during a specific number of initial sequence cycles of a sequencing sequence (e.g., the first 2 to 7 sequencing cycles).

[0233] In some embodiments, the intensities of the sequence images 108 are corrected for background and / or aligned with each other using an affinity transform. In one embodiment, the sequencing run utilizes four channel chemistry and each image set has four images. In another embodiment, the sequencing run utilizes two channel chemistry and each image set has two images. In yet another embodiment, the sequencing run utilizes one channel chemistry and each image set has two images. In yet another embodiment, each image set has only one image. These and other different embodiments are described in Appendices 6 and 9.

[0234] Each image 2116 in the series of image set 2100 covers a tile 2104 of flow cell 2102 and shows the intensity emission of clusters 2106 on tile 2104 and their surrounding background captured for a particular image channel in a particular one of multiple sequencing cycles of a sequencing run. In one example, for cycle t1, the image set includes four images 2112A, 2112C, 2112T, 2112G, including one image for each base A, C, T, and G labeled with a corresponding fluorescent dye and imaged in a corresponding wavelength band (image / imaging channel).

[0235] For illustrative purposes, in image 2112G, Figure 21b shows cluster intensity emission as 2108 and background intensity emission as 2110. In another example, for cycle tn, the image set also includes four images 2114A, 2114C, 2114T, 2114G, including one image for each base A, C, T, and G labeled with a corresponding fluorescent dye and imaged in a corresponding wavelength band (image / imaging channel). Also for illustrative purposes, in image 2114A, Figure 21b shows cluster intensity emission as 2118, and in image 2114T, background intensity emission is shown as 2120.

[0236] The input image data 1702 is encoded using intensity channels (also called imaging channels). For each c-image acquired from the sequencer for a particular sequencing cycle, a separate imaging channel is used to encode its intensity signal data. For example, consider a sequencing using a two-channel chemistry that generates a red image and a green image in each sequencing cycle. In such a case, the input data 2632 includes (i) a first red imaging channel having w×h pixels that represents the intensity radiation of one or more clusters and their surrounding background captured in the red image, and (ii) a second green imaging channel having w×h pixels that represents the intensity radiation of one or more clusters and their surrounding background captured in the green image.

[0237] (non-image data) In another embodiment, the input data to the neural network-based template generator 1512 and the neural network-based base caller 1514 is based on pH changes induced by the release of hydrogen ions during molecular elongation, which are detected and converted to voltage changes proportional to the number of incorporated bases (e.g., in the case of Ion Torrent).

[0238] In yet another embodiment, input data is constructed from nanopore sensing, which uses a biosensor to measure the disruption of electrical current as an analyte passes through or near the nanopore's opening. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore, and a potential difference is applied across the membrane. Nucleotides present within the pore affect the pore's electrical resistance, so current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (due to its "squish" appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values ​​taken at a 4 kHz frequency (for example). With a DNA strand velocity of ~450 base pairs per second, this gives, on average, approximately nine raw observations per base. This signal is then processed to identify disruption of the pore signal corresponding to individual reads. The stretching of these raw signals, called bases, is the process of converting DAC values ​​into sequences of DNA bases. In some embodiments, the input data includes normalized or scaled DAC values.

[0239] In another embodiment, image data is not used as input to the neural network-based template generator 1512 or the neural network-based base caller 1514. Instead, the input to the neural network-based template generator 1512 and the neural network-based base caller 1514 is based on pH changes induced by the release of hydrogen ions during molecular elongation. The pH changes are detected and converted to voltage changes proportional to the number of bases incorporated (e.g., in the case of Ion Torrent).

[0240] In yet another embodiment, input to the neural network-based template generator 1512 and the neural network-based base caller 1514 is constructed from nanopore sensing, which uses a biosensor to measure the disruption of electrical current as an analyte passes through or near the opening of the nanopore. For example, Oxford Nanopore Technologies (ONT) sequencing is based on the following concept: a single strand of DNA (or RNA) is passed through a membrane via a nanopore, and a potential difference is applied across the membrane. Nucleotides present within the pore affect the electrical resistance of the pore, so that current measurements over time can indicate the sequence of DNA bases passing through the pore. This current signal (due to its "squish" appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values ​​taken at a 4 kHz frequency (for example). With a DNA strand velocity of ~450 base pairs per second, this gives, on average, approximately nine raw observations per base. This signal is then processed to identify breaks in the aperture signal that correspond to individual reads. Extension of these raw signals is a process that results in base calling and converting the DAC values ​​into sequences of DNA bases. In some embodiments, input data 2632 includes normalized or scaled DAC values.

[0241] (patch extraction) Figure 22 shows one embodiment of extracting patches from the series of image sets 2100 of Figure 21b to generate a series of "downsized" image sets that form the input image data 1702. In another embodiment shown, the sequence images 108 in the series of image sets 2100 are of size LxL (e.g., 2000x2000). In other embodiments, L is any number ranging from 1 to 10,000.

[0242] In one embodiment, patch extractor 2202 extracts patches from sequence images 108 in the series of image sets 2100 to generate a series of downsized image sets 2206, 2208, 2210, and 2212. Each image in the series of downsized image sets is a patch of size M×M (e.g., 20×20) extracted from a corresponding sequence-determined image in the series of image sets 2100. The size of the patch can be preset. In other alternative embodiments, M is any number in the range of 1 to 1000.

[0243] 22 shows four exemplary series of downsized image sets. A first exemplary series of downsized image sets 2206 is extracted from coordinates 0,0 to 20,20 in the sequence of images 108 in the series of image sets 2100. A second exemplary series of downsized image sets 2208 is extracted from coordinates 20,20 to 40,40 in the sequence of images 108 in the series of image sets 2100. A third exemplary series of downsized image sets 2210 is extracted from coordinates 40,40 to 60,60 in the sequence of images 108 in the series of image sets 2100. A fourth exemplary series of downsized image sets 2212 is extracted from coordinates 60,60 to 80,80 in the sequence of images 108 in the series of image sets 2100.

[0244] In some implementations, the series of downsized image sets forms the input image data 1702 that is provided as input to the neural network-based template generator 1512. Multiple series of downsized image sets can be provided simultaneously as an input batch, and a separate output can be generated for each series in the input batch.

[0245] (upsampling) FIG. 23 illustrates one implementation of upsampling the series of image sets 2100 of FIG. 21 b to generate a series of “upsampled” image sets 2300 that form the input image data 1702 .

[0246] In one embodiment, the upsampler 2302 upsamples the sequence images 108 in the set of images 2100 by an upsampling factor (eg, 4x) and a set of upsampled images 2300 .

[0247] In another illustrated embodiment, the sequence images 108 in the series of image sets 2100 are of size L×L (e.g., 2000×2000) and are upsampled by an upsampling factor of 4 to generate upsampled images of size U×U (e.g., 8000×8000) in the series of upsampled image sets 2300.

[0248] In one embodiment, the sequence images 108 in the set of sequential images 2100 are fed directly to the neural network-based template generator 1512, and upsampling is performed by an initial layer of the neural network-based template generator 1512. That is, the upsampler 2302 is part of the neural network-based template generator 1512 and operates as the first layer that upsamples the sequence images 108 in the set of sequential images 2100 and generates the set of sequential upsampled images 2300.

[0249] In some implementations, the series of upsampled image sets 2300 form the input image data 1702 that is provided as input to the neural network-based template generator 1512 .

[0250] FIG. 24 illustrates one implementation of extracting patches from the series of upsampled image sets 2300 of FIG. 23 to generate a series of “upsampled and downsized” image sets 2406, 2408, 2410, and 2412 that form the input image data 1702.

[0251] In one embodiment, patch extractor 2202 extracts patches from upsampled images in a series of upsampled image sets 2300 to generate a series of upsampled image sets 2406, 2408, 2410 and a downsized image set 2412. Each upsampled image in the series of upsampled image sets and the downsized image set is a patch of size M×M (e.g., 80×80) extracted from a corresponding upsampled image in the series of upsampled image sets 2300. The size of the patch can be preset. In other alternative embodiments, M is any number in the range of 1 to 1000.

[0252] 24 shows four example series of upsampled and downsized image sets. A first example series of upsampled and downsized image sets 2406 is extracted from coordinates 0,0 to 80,80 within the upsampled images in the series of upsampled image sets 2300. A second example series of upsampled and downsized image sets 2408 is extracted from coordinates 80,80 to 160,160 within the upsampled images in the series of upsampled image sets 2300. A third example series of upsampled and downsized image sets 2410 is extracted from coordinates 160,160 to 240,240 within the upsampled images in the series of upsampled image sets 2300. A fourth example series of upsampled and downsized image sets 2412 is extracted from coordinates 240,240 to 320,320 within the upsampled images in the series of upsampled image sets 2300.

[0253] In some implementations, a series of upsampled and downsized image sets form the input image data 1702 that is provided as input to the neural network-based template generator 1512. Multiple series of upsampled and downsized image sets can be provided simultaneously as an input batch, and a separate output can be generated for each series in the input batch.

[0254] (output) The three models are trained to produce different outputs. This is achieved by using different types of ground truth data representations as training labels. The regression model 2600 is trained to produce outputs that characterize / represent the so-called "attenuation map" 1716. The binary classification model 4600 is trained to produce outputs that characterize / represent the so-called "binary map" 1720. The ternary classification model 5400 is trained to produce outputs that characterize / represent the so-called "ternary map" 1718.

[0255] The output 1714 of each type of model includes a unit array 1712. The unit 1712 can be a pixel, subpixel, or superpixel. The output of each type of model includes an output value for each unit such that the output values ​​of the unit array together characterize / represent / represent an attenuation map 1716 in the case of a regression model 2600, a binary map 1720 in the case of a binary classification model 4600, or a ternary map 1718 in the case of a ternary classification model 5400. Details are as follows:

[0256] (Ground truth data generation) 25 shows one implementation of an overall exemplary process for generating ground truth data for training the neural network-based template generator 1512. For a regression model 2600, the ground truth data may be the attenuation map 1204. For a binary classification model 4600, the ground truth data may be the binary map 1404. For a ternary classification model 5400, the ground truth data may be the ternary map 1304. The ground truth data is generated from cluster metadata. The cluster metadata is generated by the cluster metadata generator 122. The ground truth data is generated by the ground truth data generator 1506.

[0257] In another illustrated embodiment, ground truth data is generated for tile A on lane A of flow cell A. The ground truth data is generated from sequence images 108 of tile A captured during a sequencing run. The sequence images 108 of tile A are in the pixel domain. In one example with a four-channel chemistry that generates four sequence images per sequencing cycle, 200 sequence images 108 are accessed for 50 sequencing cycles. Each of the 200 sequence images 108 shows the intensity emissions of clusters on tile A and their surrounding background captured in a particular image channel at a particular sequencing cycle.

[0258] The sub-pixel addresser 110 converts the sequence image 108 into the sub-pixel domain (eg, by dividing each pixel into multiple sub-pixels) and generates the sequence image 112 in the sub-pixel domain.

[0259] A base caller 114 (e.g., RTA) then processes the sequence image 112 in the subpixel domain and generates base calls for each subpixel and each of the 50 sequencing cycles, referred to herein as "subpixel base calls."

[0260] The subpixel base calls 116 are then merged to generate a base call sequence for each subpixel over 50 sequencing cycles. Each subpixel base call sequence has 50 base calls, i.e., one base call for each of the 50 sequencing cycles.

[0261] The searcher 118 evaluates the base call sequences of consecutive subpixels on a pairwise basis. The search involves evaluating each subpixel to determine which subpixels in the consecutive subpixels share substantially matching base call sequences. The base call sequences of consecutive subpixels "substantially match" when a predetermined portion of the base calls match a criterion per ordinal position (e.g., 41 matches in >= 45 cycles, 4 mismatches in <= 45 cycles, 4 mismatches in <= 50 cycles, or 2 mismatches in <= 34 cycles).

[0262] In some embodiments, the base caller 114 also identifies preliminary center coordinates of the cluster. The subpixel containing the preliminary center coordinate is referred to as the center or origin subpixel. Some example preliminary center coordinates (604a-c) identified by the base caller 114 and corresponding origin subpixels (606a-c) are shown in FIG. 6. However, as explained below, identification of the origin subpixel (preliminary center coordinate of the cluster) is not required. In some embodiments, the searcher 118 uses a breadth-first search, starting with the origin subpixels 606a-c and continuing through successive non-origin subpixels 702a-c, to identify substantially matching base call sequences of the subpixels. This is optional, as explained below.

[0263] The search for a substantially matching base call sequence of a subpixel does not require identification of an origin subpixel (initial center coordinate of a cluster) because the search can be performed for all subpixels and the search need not start at the origin subpixel, but instead at any subpixel (e.g., the 0,0 subpixel or any random subpixel). Thus, because each subpixel is evaluated to determine whether it shares a substantially matching base call sequence with another contiguous subpixel, the search can start at any subpixel, rather than relying on the origin subpixel.

[0264] Regardless of whether the origin subpixel is used, certain clusters that do not include the predicted origin subpixel (the initial center coordinate of the cluster) are identified by base caller 114. Some examples of clusters that are identified by merging subpixel base calls and do not include the origin subpixel are clusters 812a, 812b, 812c, 812d, and 812e in Figure 8a. Therefore, the use of base caller 114 to identify the origin subpixel (the initial center coordinate of the cluster) is optional and not required for the search for substantially matching base call sequences of subpixels.

[0265] Searcher 118: (1) identifies contiguous subpixels having substantially matching base call sequences as so-called "discontinuous regions," (2) further evaluates the base call sequences of these subpixels that do not belong to any of the non-connected regions already identified in (1) to obtain additional discontinuous regions, and (3) then identifies background subpixels as subpixels that do not belong to any of the discontinuous regions already identified in (1) and (2). Action (2) enables the disclosed techniques to identify additional or additional clusters whose centers are not identified by base caller 114.

[0266] The results of searcher 118 are encoded in a so-called "cluster map" for tile A and stored in cluster map data store 120. In the cluster map, each of the clusters on tile A is identified by respective discontinuous regions of adjacent sub-pixels, and background sub-pixels separate the isolated regions to identify the surrounding background on tile A.

[0267] Center of mass (COM) calculator 1004 determines the center of each of the clusters on Tile A by calculating the COM of each of the discontinuous regions as the average of the coordinates of each contiguous subpixel that forms the discontinuous region. The center of mass of the cluster is stored as COM data 2502.

[0268] Subpixel classifier 2504 uses cluster map and COM data 2502 to generate subpixel classifications 2506. Subpixel classification 2506 classifies (1) background subpixels, (2) COM subpixels (one COM subpixel for each discrete region, including the COM of each discrete region), and (3) cluster / intra-cluster subpixels that form each discrete region. That is, each subpixel in the cluster map is assigned one of three categories.

[0269] Based on the subpixel classification 2506 in some embodiments, (i) a ground truth attenuation map 1204 is generated by the ground truth attenuation map generator 1202, (ii) a ground truth binary map 1304 is generated by the ground truth binary map generator 1302, and (iii) a ground truth ternary map 1404 is generated by the ground truth ternary map generator 1402.

[0270] 1. (Regression model) Figure 26 shows one example of a regression model 2600. In another embodiment shown, the regression model 2600 is a fully convolutional network 2602 that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map low-resolution encoder feature maps to full-input resolution attenuation maps 1716. In another embodiment, the regression model 2600 is a U-Net network 2604 with skip connections between the decoder and encoder. Further details regarding segmentation networks can be found in the appendix titled "Segmentation Networks."

[0271] (Attenuation map) 27 illustrates one implementation of generating a ground truth attenuation map 1204 from a cluster map 2702. The ground truth attenuation map 1204 is used as ground truth data for training the regression model 2600. In the ground truth attenuation map 1204, the ground truth attenuation map generator 1202 assigns a weighted attenuation value to each neighboring subpixel based on a weighted attenuation coefficient. The weighted attenuation value is proportional to the Euclidean distance of the neighboring subpixel from the center of mass (COM) subpixel within the discontinuous region to which the neighboring subpixel belongs, such that the weighted attenuation value is highest (e.g., 1 or 100) for COM subpixels and decreases for subpixels further away from the COM subpixel. In some implementations, the weighted attenuation value is multiplied by a preset coefficient, such as 100.

[0272] Furthermore, the ground truth attenuation map generator 1202 assigns the same predetermined value (eg, the minimum background value) to all background sub-pixels.

[0273] The ground truth attenuation map 1204 represents contiguous subpixels in discontinuous regions and background subpixels based on assigned values. The ground truth attenuation map 1204 also stores the assigned values ​​in a unit array, with each unit in the array representing a corresponding subpixel in the input.

[0274] (training) FIG. 28 shows one implementation of training 2800 of a regression model 2600 using a back-propagation based gradient update technique to modify the parameters of the regression model 2600 until the attenuation map 1716 generated by the regression model 2600 as the training output during training 2800 progressively approaches or matches the ground truth attenuation map 1204 of the ground.

[0275] Training 2800 involves iteratively optimizing to minimize an error 2806 between the attenuation map 1716 and the ground truth attenuation map 1204 and updating the parameters of the regression model 2600 based on the error 2806. In one embodiment, the loss function is the mean squared error, which is minimized for each subpixel between the weighted attenuation values ​​of corresponding subpixels in the attenuation map 1716 and the ground truth attenuation map 1204.

[0276] Training 2800 includes hundreds, thousands, and / or millions of forward propagations 2808 and back propagations 2810, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by an annotator 2806. Training 2800 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.

[0277] (inference) 29 is an implementation of template generation by regression model 2600 during inference 2900, in which attenuation map 1716 is generated by regression model 2600 as an inference output during inference 2900. An example of attenuation map 1716 is disclosed in the Appendix entitled "Regression_Model_Ouput." The Appendix includes unit-weighted attenuation output values ​​2910 that together represent attenuation map 1716.

[0278] Inference 2900 includes hundreds, thousands, and / or millions of forward propagations 2904, including parallelogram techniques such as batching. Inference 2900 is performed on inference data 2908, which includes a series of upsampled and downsized image sets, as input image data 1702. Inference 2900 can be operated on by tester 2906.

[0279] (basin separation) 30 includes (i) thresholding the attenuation map 1716 to identify background sub-pixels that characterize cluster backgrounds, and (ii) peak detection to identify center sub-pixels that characterize cluster centers. Thresholding is performed by a threshold keeper 1802 that uses a local threshold binary to generate a binarized output. Peak detection is performed by a peak locator 1806 to identify cluster centers. Further details regarding peak locators can be found in the Appendix titled "Peak Detection."

[0280] 31 shows one implementation of the watershed segmentation technique, which takes as input the background subpixels and the center subpixels, each identified by thresholder 1802, and peak locator 1806 finds the intensity valleys between adjacent clusters and outputs non-overlapping groups of adjacent cluster / intra-cluster subpixels that characterize the clusters. Further details regarding the watershed segmentation technique can be found in the Appendix entitled "Watershed Segmentation."

[0281] In one implementation, the watershed divider 3102 takes as input (1) the attenuation map 1716, (2) the negated output values ​​1802, and (3) the cluster centers identified by the peak locator 1806 as input (1) minus the output values ​​2910. Based on the input, the watershed divider 3102 then generates an output 3104. In the output 3104, each cluster center is identified as a unique set / group of subpixels that belong to the cluster center (as long as the subpixels are "1" in the binary output, i.e., not background subpixels). Further, clusters are filtered based on containing at least four subpixels. The watershed divider 3102 can be part of the divider 1810, which in turn is part of the post-processor 1814.

[0282] (Network structure) 32 is a table showing an example U-Net structure for regression model 2600, including details of the layers, the dimensionality of the layer outputs, the magnitude of the model parameters, and the interconnections between layers for regression model 2600. Similar details are disclosed in the file entitled "Regression_Model_Example_Architecture" submitted as an appendix to this application.

[0283] (Cluster intensity extraction) Figure 33 shows a different approach to extracting cluster intensities using cluster shape information identified in a template image. As described above, the template image identifies cluster shape information at upsampled sub-pixel resolution. However, the cluster intensity information is in the sequence image 108, which is typically at optical resolution.

[0284] According to the first approach, the coordinates of the sub-pixels are located in the sequence image 108 and their respective intensities are extracted using bilinear interpolation and normalized based on the count of the sub-pixels contributing to the cluster.

[0285] The second approach uses a weighted area coverage technique to modulate the intensity of a pixel according to the number of subpixels that contribute to the pixel. Again, the modulated pixel intensity is normalized by a subpixel count parameter.

[0286] The third technique uses quadratic interpolation to upsample the sequence images to the subpixel domain, sums the intensities of the upsampled pixels that belong to a cluster, and normalizes the summed intensity based on the count of the upsampled pixels that belong to the cluster.

[0287] (Experimental results and discussion) Figure 34 shows different approaches to base calling using the output of the regression model 2600. In the first approach, cluster centers identified from the output of the neural network-based template generator 1512 in the template image are fed into a base caller (e.g., Illumina's Time Analysis software, referred to herein as "RTA base calling") for base calling.

[0288] In the second approach, instead of cluster centers, cluster intensities extracted from sequence images based on cluster shape information in the template image are fed to the RTA base caller for base calling.

[0289] Figure 35 shows the difference in base calling performance when RTA base calling uses ground truth center of mass (COM) positions as cluster centers, as opposed to using non-COM positions as cluster centers. The results show that using COMs improves base calling.

[0290] (Example of model output) Figure 36 shows, on the left, an exemplary attenuation map 1716 produced by the regression model 2600. Figure 36 also shows, on the right, an exemplary ground truth attenuation map 1204 that the regression model 2600 approximates during training.

[0291] Both the attenuation map 1716 and the ground truth attenuation map 1204 depict clusters as discontinuous regions of adjacent sub-pixels, with the cluster centers shown as the central sub-pixels at the center of mass of the corresponding regions of the discontinuous regions, and the clusters as their surrounding background.

[0292] Additionally, consecutive sub-pixels in corresponding ones of the discontinuous regions have values ​​weighted according to the distance of the consecutive sub-pixel from the central sub-pixel in the discontinuous region to which the adjacent sub-pixels belong. In one implementation, the central sub-pixel has the highest value in the corresponding one of the discontinuous regions. In one implementation, all background sub-pixels have the same minimum background value in the attenuation map.

[0293] 37 shows one embodiment of a peak locator 1806 that identifies cluster centers in an attenuation map by detecting peaks 3702. Further details regarding the peak locator can be found in the Appendix entitled "Peak Detection."

[0294] Figure 38 compares the peaks detected by the peak locator 1806 in the attenuation map 1716 generated by the regression model 2600 with the peaks in the corresponding ground truth attenuation map 1204. The red markers are the peaks predicted by the regression model 2600 as cluster centers, and the green markers are the ground truth centers of cluster agglomerations.

[0295] (Further experimental results and considerations) 39 shows the performance of regression model 2600 using precision and recalibration statistics. The precision and recalibration statistics demonstrate that regression model 2600 is successful in recovering all identified cluster centers.

[0296] Figure 40 compares the performance of regression model 2600 with the RTA base caller for a library concentration of 20 pM (normal operation). Running the RTA base caller, regression model 2600 identifies 34,323 (4.46%) clusters in a higher cluster density environment (i.e., 988,884 clusters).

[0297] Figure 40 also shows the results of other sequencing metrics such as the number of clusters that pass the chemistry filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads that aligned to the reference sequence ("% mismatch"), the number of reads that do not match the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases").

[0298] Figure 41 compares the performance of regression model 2600 with the RTA base caller for a 30 pM library concentration (high density run). Running with the RTA base caller, regression model 2600 identifies 34,323 (6.27%) more clusters in a much higher cluster density environment (i.e., 1,351,588 clusters).

[0299] Figure 41 also shows the results of other sequencing metrics such as the number of clusters that pass the chemistry filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads that aligned to the reference sequence ("% mismatch"), the number of reads that do not match the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases").

[0300] Figure 42 compares the number of two non-overlapping (unique or overlapping duplicate) correct read pairs, i.e., the number of paired reads where both reads are aligned within a reasonable distance, detected by regression model 2600 with those detected by the RTA-based color. The comparison is performed for both a normal run at 20 pM and a high-density run at 30 pM.

[0301] More importantly, Figure 42 shows that the disclosed neural network-based template generator can detect more clusters with fewer sequencing cycles input to template generation. With just four sequencing cycles, regression model 2600 identifies 11% more non-overlapping correct read pairs than the RTA base caller in a 20 pM regular run and 33% more correct read pairs than the RTA base caller in a 30 pM high-density run. With seven sequencing cycles, regression model 2600 identifies 4.5% more non-overlapping correct read pairs than the RTA base caller in a 20 pM regular run and 6.3% more correct read pairs than the RTA base caller in a 30 pM high-density run.

[0302] Figure 43 shows on the right the first attenuation map generated by regression model 2600. The first attenuation map identifies the clusters and their surrounding background imaged during normal operation at 20 pM, along with their spatial distribution showing cluster shape, cluster size, and cluster centers.

[0303] On the left, Figure 43 shows the second attenuation map generated by regression model 2600. The second attenuation map identifies the clusters and their surrounding background imaged during the 30 pM high-density run, along with their spatial distribution showing cluster shape, cluster size, and cluster centers.

[0304] Figure 44 compares the performance of regression model 2600 with the RTA base caller for a library concentration of 40 pM (high density run). Regression model 2600 produced 89,441,688 more aligned bases than the RTA base caller in a much higher cluster density environment (i.e., 1,509,395 clusters).

[0305] Figure 44 also shows the results of other sequencing metrics such as the number of clusters that pass the chemistry filter ("%PF" (pass filter)), the number of aligned reads ("% aligned"), the number of overlapping reads ("% duplicates"), all reads aligned to the reference sequence ("% mismatch"), the number of reads that mismatch the reference sequence for a quality score of 30 and the bases referred to above ("%Q30 bases"). (Further examples of model outputs)

[0306] Figure 45 shows, on the left, the first attenuation map generated by regression model 2600. The first attenuation map identifies the clusters and their surrounding background imaged during normal operation at 40 pM, along with their spatial distribution showing cluster shape, cluster size, and cluster centers.

[0307] In the top right, Figure 45 shows the results of thresholding and peak location applied to the first attenuation map to distinguish each cluster from each other and from the background and identify their respective cluster centers. In some embodiments, the intensity of each cluster is identified and a chassis filter (or pass filter) is specified that is applied to reduce the mismatch rate.

[0308] 2. (Binary classification model) Figure 46 shows one example of a binary classification model 4600. In another embodiment shown, the binary classification model 4600 is a deep fully convolutional segmentation neural network that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map low-resolution encoder feature maps to full input resolution binary maps 1720. In another embodiment, the binary classification model 4600 is a U-Net network with skip connections between the decoders and encoders. Further details regarding segmentation networks can be found in the appendix titled "Segmentation Networks."

[0309] (binary map) The final output layer of the binary classification model 4600 is a unit-wise classification layer that generates a classification label for each unit in the output array. In some implementations, the unit-wise partitioning layer is a subpixel-wise classification layer that generates a softmax classification score distribution for each subpixel in the binary map 1720 across two classes, i.e., the cluster center class and non-cluster class, and the classification label for a given subpixel are determined from the corresponding softmax classification score distribution.

[0310] In another alternative embodiment, the unit-wise classification layer is a sub-pixel-wise classification layer that generates a sigmoid classification score for each sub-pixel in the binary map 1720 such that the activation of the unit is interpreted as the probability that the unit belongs to a first class, and conversely, one minus one gives the probability of belonging to a second class.

[0311] The binary map 1720 represents each subpixel based on its predicted classification score. The binary map 1720 also stores the predicted classification scores in an array of units, with each unit in the array representing a corresponding subpixel in the input.

[0312] (training) FIG. 47 illustrates one implementation of training 4700 a binary classification model 4600 using a backpropagation-based gradient update technique to modify the parameters of the binary classification model 4600 until the binary map 1720 of the binary classification model 4600 progressively approaches or matches the ground truth binary map 1404.

[0313] In the illustrated implementation, the final output layer of the binary classification model 4600 is a softmax-based per-subpixel classification layer. In another implementation of softmax, the ground truth binary map generator 1402 assigns each ground truth subpixel either (i) a cluster center value pair (e.g., [1, 0]) or (ii) a non-center value pair (e.g., [0, 1]).

[0314] In a cluster center value pair [1,0], the first value [1] represents the cluster center class label and the second value [0] represents the non-center class label. In a non-center value pair [0,1], the first value [0] represents the cluster center class label and the second value [1] represents the non-center class label.

[0315] The ground truth binary map 1404 represents each subpixel based on an assigned value pair / value. The ground truth binary map 1404 also stores the assigned value pairs / values ​​in a unit array, with each unit in the array representing a corresponding subpixel in the input.

[0316] Training involves iteratively optimizing a loss function that minimizes the error 4706 (e.g., softmax error) between the binary map 1720 and the ground truth binary map 1404, and updating the parameters of the binary classification model 4600 based on the error 4706.

[0317] In one implementation, the loss function is a custom weighted binary cross-entropy loss, and the error 4706 is minimized for each subpixel between the predicted classification score (e.g., softmax score) and the labeled class score (e.g., softmax score), as shown in FIG. 47, and between the labeled class score (e.g., softmax score) of the corresponding subpixel in the binary map 1720 and the ground truth binary map 1404.

[0318] The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the appendix titled "Custom-Weighted Loss Function."

[0319] Training 4700 includes hundreds, thousands, and / or millions of forward propagations 4708 and back propagations 4710, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by an annotator 2806. Training 2800 can be operated on by a trainer 1510 using a stochastic gradient update algorithm such as Adam.

[0320] FIG. 48 is another embodiment of training 4800 a binary classification model 4600 in which the final output layer of the binary classification model 4600 is a sigmoid-based sub-pixel-wise classification layer.

[0321] In another sigmoid implementation, the ground truth binary map generator 1302 assigns each ground truth subpixel either (i) a cluster center value (e.g., [1]) or (ii) a non-center value (e.g., [0]). The COM subpixel is assigned the cluster center value pair / value, and all other subpixels are assigned the non-center value pair / value.

[0322] For cluster central values, values ​​between 0 and 1 (e.g., values ​​above 0.5) represent central class labels. For non-central values, values ​​below the threshold between 0 and 1 (e.g., values ​​below 0.5) represent non-central class labels.

[0323] The ground truth binary map 1404 represents each subpixel based on an assigned value pair / value. The ground truth binary map 1404 also stores the assigned value pairs / values ​​in a unit array, with each unit in the array representing a corresponding subpixel in the input.

[0324] Training involves iteratively optimizing a loss function that minimizes the error 4806 (e.g., a sigmoid error) between the binary map 1720 and the ground truth binary map 1404, and updating the parameters of the binary classification model 4600 based on the error 4806.

[0325] In one implementation, the loss function is a custom weighted binary cross-entropy loss, and the error 4806 is minimized for each subpixel between the predicted scores (e.g., sigmoid scores) of corresponding subpixels in the binary map 1720 and the ground truth binary map 1404, as shown in FIG. 48, and the labeled scores (e.g., sigmoid scores) of corresponding subpixels in the binary map 1720 and the ground truth binary map 1404, as shown in FIG.

[0326] The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the appendix titled "Custom-Weighted Loss Function."

[0327] Training 4800 includes hundreds, thousands, and / or millions of forward propagations 4808 and back propagations 4810, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 2800 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.

[0328] FIG. 49 shows another implementation of input image data 1702 provided to a binary classification model 4600 and the corresponding class labels 4904 used to train the binary classification model 4600.

[0329] In another illustrated embodiment, input image data 1702 is serially upsampled to comprise downsized image set 4902. Class labels 4904 include two classes: (1) "no cluster center" and (2) "cluster center," which are distinguished using different output values: (1) light green units / subpixels 4906 represent subpixels predicted by binary classification model 4600 to not contain cluster centers, and (2) dark green subpixels 4908 represent units / subpixels predicted by binary classification model 4600 to contain cluster centers.

[0330] (inference) Figure 50 is an implementation of template generation by a binary classification model 4600 during inference 5000 in which a binary map 1720 is generated by the binary classification model 4600 as an inference output during inference 5000. An example of a binary map 1720 includes binary classification scores 5010 per unit that together represent the binary map 1720. In a softmax application, the binary map 1720 has a first array 5002a of classification scores per unit for non-center classes and a second array 5002b of classification scores per unit for cluster center classes.

[0331] Inference 5000 includes hundreds, thousands, and / or millions of forward propagations 5004, including parallelogram techniques such as batching. Inference 5000 is performed on inference data 2908, which includes a series of upsampled and downsized image sets, as input image data 1702. Inference 5000 can be operated by tester 2906.

[0332] In some implementations, the binary map 1720 is subjected to the post-processing techniques described above, such as thresholding, peak detection, and / or watershed division, to generate cluster metadata.

[0333] (Peak detection) Figure 51 shows one embodiment of subjecting a binary map 1720 to peak detection to identify cluster centers. As described above, the binary map 1720 is an array of units that classify each subpixel based on a predicted classification score, with each unit in the array representing a corresponding subpixel in the input. The classification score can be a softmax score or a sigmoid score.

[0334] In softmax applications, the binary map 1720 includes two arrays: (1) a first array 5002a of classification scores per unit of the non-center classes, and (2) a second array 5002b of classification scores per unit of the cluster center classes, where each unit in both arrays represents a corresponding subpixel in the input.

[0335] To determine which sub-pixels in the input contain cluster centers and which do not, the peak locator 1806 applies peak detection on the units in the binary map 1720. Peak detection identifies units with classification scores (e.g., softmax / sigmoid scores) above a preset threshold. The identified units are inferred as cluster centers, and their corresponding sub-pixels in the input are determined to contain cluster centers and are stored as cluster center sub-pixels in the sub-pixel classification data store 5102. Further details regarding the peak locator 1806 can be found in the Appendix entitled "Peak Detection."

[0336] The remaining units in the input and their corresponding subpixels do not contain cluster centers and are stored as non-central subpixels in the subpixel classification data store 5102.

[0337] In some implementations, before applying peak detection, units with classification scores below a certain background threshold (e.g., 0.3) are set to zero. In some implementations, such units and their corresponding subpixels in the input are inferred to represent the background surrounding the cluster and are stored as background subpixels in subpixel classification data store 5102. In other implementations, such units are considered noise and can be ignored.

[0338] (Example of model output) Figure 52a shows an example binary map generated by binary classification model 4600 on the left. Figure 52a also shows an example ground truth binary map on the right that binary classification model 4600 approximates during training. The binary map has multiple subpixels and classifies each subpixel as either a cluster center or a non-center. Similarly, the ground truth binary map has multiple subpixels and classifies each subpixel as either a cluster center or a non-center. (Experimental results and discussion)

[0339] Figure 52b shows the performance of the binary classification model 4600 using recalibration and refinement statistics. By applying these statistics, the binary classification model 4600 performs the RTA base caller.

[0340] (Network structure) 53 is a table illustrating an example structure of a binary classification model 4600, along with details of the layers, the dimensionality of the layer outputs, the magnitude of the model parameters, and the interconnections between layers of the binary classification model 4600. Similar details are disclosed in the appendix entitled "Binary_Classification_Model_Example_Architecture."

[0341] 3. Three-way (three-class) classification model Figure 54 shows one implementation of a ternary classification model 5400. In another implementation shown, the ternary classification model 5400 is a deep fully convolutional segmentation neural network that processes input image data 1702 through an encoder sub-network and a corresponding decoder sub-network. The encoder sub-network includes a hierarchy of encoders. The decoder sub-network includes a hierarchy of decoders that map low-resolution encoder feature maps to full-input-resolution ternary maps 1718. In another embodiment, the ternary classification model 5400 is a U-Net network with skip connections between the decoders and encoders. Further details regarding segmentation networks can be found in the Appendix titled "Segmentation Networks." (Triple Map)

[0342] The final output layer of the ternary classification model 5400 is a unit-wise classification layer that generates a classification label for each unit in the output array. In some implementations, the unit-wise segmentation layer is a subpixel-wise classification layer that generates a softmax classification score distribution for each subpixel in the ternary map 1718 across three classes: background class, cluster center class, and cluster / intra-cluster class, and the classification label for a given subpixel is determined from the corresponding softmax classification score distribution.

[0343] The ternary map 1718 represents each subpixel based on its predicted classification score. The ternary map 1718 also stores the predicted classification scores in an array of units, with each unit in the array representing a corresponding subpixel in the input. (training)

[0344] FIG. 55 illustrates one implementation of training 5500 a ternary classification model 5400 using a backpropagation-based gradient update technique that modifies the parameters of the ternary classification model 5400 until the ternary map 1718 of the ternary classification model 5400 progressively approaches or matches the training ground truth ternary map 1304.

[0345] In the illustrated implementation, the final output layer of the ternary classification model 5400 is a softmax-based per-subpixel classification layer. In another implementation, the softmax ternary map generator 1402 assigns each ground truth triplet either (i) a background value triplet (e.g., [1, 0, 0]), (ii) a cluster center value triplet (e.g., [0, 1, 0]), or (iii) a cluster / cluster interior value triplet (e.g., [0, 0, 1]).

[0346] Background subpixels are assigned background value triplets, center of mass (COM) subpixels are assigned cluster center value triplets, and cluster / intra-cluster subpixels are assigned cluster / intra-cluster value triplets.

[0347] In the background value triplet [1,0,0], the first value [1] represents the background class label, the second value [0] represents the cluster center label, and the third value [0] represents the cluster / intra-cluster class label.

[0348] In the cluster center triplet [0, 1, 0], the first value [0] represents the background class label, the second value [1] represents the cluster center label, and the third value [0] represents the cluster / intra-cluster class label.

[0349] In the cluster / cluster inner value triplet [0, 0, 1], the first value [0] represents the background class label, the second value [0] represents the cluster center label, and the third value [1] represents the cluster / cluster inner class label.

[0350] The ground truth ternary map 1304 represents each subpixel based on an assigned value triplet. The ground truth ternary map 1304 also stores the assigned triplets in a unit array, with each unit in the array representing a corresponding subpixel in the input.

[0351] Training involves iteratively optimizing a loss function that minimizes the error 5506 (e.g., softmax error) between the ternary map 1718 and the ground truth ternary map 1304, and updating the parameters of the ternary classification model 5400 based on the error 5506.

[0352] In one implementation, the loss function is a custom weighted categorization cross-entropy loss, and the error 5506 is minimized for each subpixel between the predicted classification score (e.g., softmax score) and the labeled class score (e.g., softmax score), as shown in FIG. 54, and between the labeled class score (e.g., softmax score) of the corresponding subpixel in the ternary map 1718 and the ground truth ternary map 1304.

[0353] The custom-weighted loss function gives more weight to the COM subpixel each time it is misclassified, multiplied by the corresponding reward (or penalty) weight specified in the reward (or penalty) matrix. Further details regarding the custom-weighted loss function can be found in the appendix titled "Custom-Weighted Loss Function."

[0354] Training 5500 includes hundreds, thousands, and / or millions of forward propagations 5508 and back propagations 5510, including parallelogram techniques such as batching. Training data 1504 includes a series of upsampled and downsized image sets as input image data 1702. Training data 1504 is annotated with ground truth labels by annotator 2806. Training 5500 can be operated on by trainer 1510 using a stochastic gradient update algorithm such as Adam.

[0355] FIG. 56 shows one embodiment of input image data 1702 provided to a three-way classification model 5400 and the corresponding class labels used to train the three-way classification model 5400.

[0356] In another illustrated embodiment, input image data 1702 is serially upsampled to comprise downsized image set 5602. Class labels 5604 include three classes: (1) a "background class," (2) a "cluster center class," and (3) a "cluster interior class," which are distinguished using different output values. For example, some of these different output values ​​can be visually represented as follows: (1) gray units / subpixels 5606 represent subpixels predicted by ternary classification model 5400 to be background, (2) dark green units / subpixels 5608 represent subpixels predicted by ternary classification model 5400 to contain cluster centers, and (3) light green subpixels 5610 represent subpixels predicted by ternary classification model 5400 to contain the interiors of clusters.

[0357] (Network structure) 57 is a table illustrating an example structure of a ternary classification model 5400, along with details of the layers, the dimensionality of the layer outputs, the magnitude of the model parameters, and the interconnections between layers of the ternary classification model 5400. Similar details are disclosed in the appendix entitled "Ternary_Classification_Model_Example_Architecture."

[0358] (inference) Figure 58 illustrates an implementation of template generation by a ternary classification model 5400 during inference 5800 in which a ternary map 1718 is generated by the ternary classification model 5400 as an inference output during inference 5800. An example of the ternary map 1718 is disclosed in the appendix titled "Ternary_Classification_Model_Output." The appendix includes binary classification scores 5810 per unit that together represent the ternary map 1718. In a softmax application, the appendix has a first array 5802a of classification scores per unit for background classes, a second array 5802b of classification scores per unit for cluster center classes, and a third array 5802c of classification scores per unit for cluster / intra-cluster classes.

[0359] Inference 5800 includes hundreds, thousands, and / or millions of forward propagations 5804, including parallelogram techniques such as batching. Inference 5800 is performed on inference data 2908, which includes a series of upsampled and downsized image sets, as input image data 1702. Inference 5800 can be operated by tester 2906.

[0360] In some implementations, the ternary map 1718 is generated by the ternary classification model 5400 using post-processing techniques described above, such as thresholding, peak detection, and / or watershed segmentation.

[0361] FIG. 59 graphically illustrates the ternary map 1718 generated by the ternary classification model 5400 with the ternary softmax classification score distributions for three corresponding classes, namely, the background class 5906, the cluster center class 5902, and the cluster / cluster inner class 5904, respectively.

[0362] Figure 60 shows the unit array generated by the ternary classification model 5400 along with the output values ​​for each unit. As shown, each unit has three output values ​​for three corresponding classes: background class 5906, cluster center class 5902, and cluster / cluster inner class 5904. For each classification (column-wise), each unit is assigned the class with the highest output value, as indicated by the class in parentheses for each unit. In some implementations, the output values ​​6002, 6004, and 6006 are analyzed for each of the respective classes 5906, 5902, and 5904 (row-wise).

[0363] (Peak detection and watershed division) Figure 61 shows one embodiment in which the ternary map 1718 is subjected to post-processing to identify cluster centers, cluster backgrounds, and cluster interiors. As described above, the ternary map 1718 is an array of units that classify each subpixel based on a predicted classification score, with each unit in the array representing a corresponding subpixel in the input. The classification score may be a softmax score.

[0364] For softmax applications, the ternary map 1718 includes three arrays: (1) a first array 5802a of per-unit classification scores for the background class, (2) a second array 5802b of per-unit classification scores for the cluster center classes, and (3) a third array 5802c of per-unit classification scores for the cluster inner classes. In all three arrays, each unit represents a corresponding subpixel in the input.

[0365] To determine which subpixels in the input contain cluster centers that contain the interior of a cluster and include background, the peak locator 1806 applies peak detection to the softmax values ​​in the ternary map 1718 of the cluster center class 5802b. Peak detection identifies units that have a classification score (e.g., a softmax score) above a preset threshold. The identified units are inferred as cluster centers, and their corresponding subpixels in the input are determined to contain cluster centers and are stored as cluster center subpixels in the subpixel classification and segmentation data store 6102. Further details regarding the peak locator 1806 can be found in the Appendix titled "Peak Detection."

[0366] In some implementations, before applying peak detection, units with classification scores below a certain noise threshold (e.g., 0.3) are set to zero. Such units can be considered noise and can be ignored.

[0367] Also, those subpixels in the input that have a classification score for background class 5802a above a certain background threshold (e.g., 0.5 or greater) are inferred to indicate background surrounding the cluster and are stored as background subpixels in the subpixel classification and segmentation data store 6102.

[0368] A watershed segmentation algorithm operated by watershed segmentation 3102 is then used to determine the shape of the clusters. In some implementations, background units / subpixels are used as a mask by the watershed segmentation algorithm. The classification scores of units / subpixels inferred as cluster centers and cluster interiors are summed to generate so-called "cluster labels." The cluster centers are used as watershed markers for separation by intensity valleys by the watershed segmentation algorithm.

[0369] In one embodiment, the negatively polarized cluster labels are provided as an input image to a watershed divider 3102, which performs the segmentation and generates cluster shapes as discontinuous regions of adjacent cluster interior subpixels separated by background subpixels. Additionally, each discontinuous region includes a corresponding cluster center subpixel. In some embodiments, the corresponding cluster center subpixel is the center of the region to which it belongs. In other embodiments, the center of mass (COM) of the discontinuous region is calculated based on the underlying location coordinates and stored as the new center of the cluster.

[0370] The output of the watershed segmenter 3102 is stored in the subpixel classification and segmentation data store 6102. Further details regarding the watershed segmentation algorithm and other segmentation algorithms can be found in the appendix entitled "Watershed Segmentation."

[0371] Example outputs of the peak locator 1806 and the watershed divider 3102 are shown in FIGS. 62a, 62b, 63, and 64.

[0372] (Example of model output) Figure 62a shows an example prediction of the ternary classification model 5400. Figure 62a shows four maps, each with a unit arrangement. The first map 6202 (far left) shows the output value of each unit of the cluster center class 5802b. The second map 6204 shows the output value of each unit of the cluster / cluster inner class 5802c. The third map 6206 (far right) shows the output value of each unit of the background class 5802a. The fourth map 6208 (bottom) is a binary mask of the ground truth ternary map 6008, assigning each unit the class label with the highest output value.

[0373] Figure 62b shows another example prediction of the ternary classification model 5400. Figure 62b shows four maps, each with a unit arrangement. The first map 6212 (bottom) shows the output value of each unit of the cluster / cluster inner class. The second map 6214 shows the output value of each unit of the cluster center class. The third map 6216 (rightmost) shows the output value of each unit of the background class. The fourth map (top) 6210 is a ground truth ternary map that assigns each unit the class label with the highest output value.

[0374] Figure 62c shows yet another example prediction of the ternary classification model 5400. Figure 64 shows four maps, each with a unit arrangement. The first map 6220 (bottom) shows the output value of each unit of the cluster / cluster inner class. The second map 6222 shows the output value of each unit of the cluster center class. The third map 6224 (rightmost) shows the output value of each unit of the background class. The fourth map 6218 (top) is a ground truth ternary map that assigns each unit the class label with the highest output value.

[0375] Figure 63 shows one embodiment for deriving cluster centers and cluster shapes from the output of the ternary classification model 5400 of Figure 62a by subjecting the output to post-processing. The post-processing (e.g., peak locations, watershed divisions) produces cluster shape data and other metadata that are identified in a cluster map 6310.

[0376] (Experimental results and discussion) Figure 64 compares the performance of the binary classification model 4600, the regression model 2600, and the RTA base caller. Performance is evaluated using various sequencing metrics. One metric is the total number of clusters detected ("# clusters"), which can be measured by the number of unique cluster centers detected. Another metric is the number of detected clusters that pass the chicity filter ("%PF" (pass filter)). During cycles 1-25 of the sequencing run, the chicity filter removes the least-confident clusters from the image extraction results. A cluster "passes the filter" if no more than one base call has a chicity value of less than 0.6 in the first 25 cycles. The chicity filter is defined as the ratio of the brightest base intensity divided by the sum of the brightest test and the second brightest base intensity. This metric goes beyond the quantity of detected clusters and also their quality, i.e., how many of the detected clusters can be used for accurate base calling and downstream secondary and ternary analyses, such as variant calling and variant pathogenicity annotation.

[0377] Other metrics measuring how good a detected cluster is for downstream analysis include the number of aligned reads generated from the detected cluster ("% Aligned"), the number of duplicate reads generated from the detected cluster ("% Duplicate"), the number of reads generated from the detected cluster that mismatch the reference sequence for all reads aligned to the reference sequence ("Mismatched"), the number of reads generated from the detected cluster that are ignored for alignment because their portion does not sufficiently match the reference sequence on either side ("% Soft Clipped"), the number of bases called for the detected cluster that have a quality score of 30 and are above ("% Q30 Bases"), the number of paired reads generated from the detected cluster that are aligned within a reasonable distance ("Total Correct Read Pairs"), and the number of unique or overlapping correct read pairs generated from the detected cluster ("Non-Overlapping Correct Read Pairs").

[0378] As shown in FIG. 64, both the binary classification model 4600 and the regression model 2600 perform well with the RTA base caller in generating templates for the majority of metrics.

[0379] Figure 65 compares the performance of the three-way classification model 5400 with that of the RTA-based caller under three conditions, five sequencing metrics, and two run densities.

[0380] In the first scenario, referred to as "RTA," cluster centers are detected by the RTA base caller, intensity extraction from the clusters is performed by the RTA base caller, and the clusters are also base called using the RTA base caller. In the second scenario, referred to as "RTA IE," cluster centers are detected by the ternary classification model 5400, but intensity extraction from the clusters is performed by the RTA base caller, and the clusters are also base called using the RTA base caller. In the third scenario, referred to as "Self IE," cluster centers are detected by the ternary classification model 5400, and intensity extraction from the clusters is performed using the cluster shape-based intensity extraction technique disclosed herein (note that cluster shape information is generated by the ternary classification model 5400). However, the clusters are base called using the RTA base caller.

[0381] Performance is compared between the three-way classification model 5400 and the RTA base caller along the following five metrics: (1) total number of detected clusters (“# clusters”), (2) number of detected clusters passing the chemistry filter (“#PF”), (3) number of unique or overlapping suitable read pairs generated from the detected clusters (“#non-overlapping suitable read pairs”), (4) sequence reads generated from the detected clusters and the reference sequence after alignment (“Mismatch Rate”), and (5) percentage of mismatches between detected clusters with a quality score of 30 (“%Q30”).

[0382] We compare performance between the three-way classification model 5400 and the RTA base caller under three conditions, comparing five metrics for two types of sequencing runs: (1) a normal run with a 20 pM library concentration, and (2) a high-density run with a 30 pM library concentration.

[0383] As shown in FIG. 65, a three-way classification model 5400 implements the RTA base caller for all metrics.

[0384] FIG. 66 shows that under the same three conditions, five metrics, and two execution densities, regression model 2600 outperforms the RTA base caller for all metrics.

[0385] FIG. 67 focuses on the final layer 6702 of the neural network-based template generator 1512.

[0386] Figure 68 visualizes what the final layer 6702 of the neural network-based template generator 1512 has learned as a result of back-propagation-based gradient update training. The illustrated embodiment visualizes 24 out of 32 convolution filters of the final layer 6702 overlaid on the ground truth cluster shapes. As shown in Figure 68, the final layer 6702 has learned cluster metadata including the spatial distribution of clusters, such as cluster centers, cluster shapes, cluster sizes, cluster backgrounds, and cluster boundaries.

[0387] Figure 69 overlays the cluster center predictions of binary classification model 4600 (in blue) on the RTA-based color (in pink) predictions. Predictions are made for sequencing image data from an Illumina NextSeq sequencer.

[0388] Figure 70 overlays cluster center predictions made by RTA base calling (in pink) on a visualization of the trained convolutional filters of the final layer of binary classification model 4600. These convolutional filters are learned as a result of sequencing image data from an Illumina NextSeq sequencer.

[0389] Figure 71 shows one embodiment of training data used to train the neural network-based template generator 1512. In this alternative embodiment, the training data is obtained from a high-density flow cell that generates data using Storm probe images. In another embodiment, the training data is obtained from a high-density flow cell that generates data with fewer bridge amplification cycles.

[0390] FIG. 72 is an example of using beads for image registration based on cluster center prediction in a neural network-based template generator 1512.

[0391] Figure 73 shows one embodiment of cluster statistics for clusters identified by the neural network-based template generator 1512. The cluster statistics include the number of contributing subpixels and the cluster size based on GC content.

[0392] Figure 74 shows how the ability of the neural network-based template generator 1512 to distinguish between adjacent clusters improves as the number of initial sequencing cycles in which the input image data 1702 is used increases from 5 to 7. For five sequencing cycles, a single cluster is identified by a single, discontinuous region of contiguous subpixels. For seven sequencing cycles, the single cluster is split into two adjacent clusters, each with its own discontinuous region of adjacent subpixels.

[0393] Figure 75 shows the difference in base calling performance when RTA-based color uses ground truth mass (COM) positions as cluster centers as opposed to when non-COM positions are used as cluster centers.

[0394] FIG. 76 shows the performance of the neural network-based template generator 1512 on additional detected clusters.

[0395] FIG. 77 shows different data sets used to train the neural network-based template generator 1512.

[0396] (Sequencing System) 78A and 78B show one embodiment of a sequencing system 7800A. The sequencing system 7800A includes a configurable processor 7846. The configurable processor 7846 implements the base calling techniques disclosed herein. A sequencing system is also referred to as a "sequencer."

[0397] The sequencing system 7800A can obtain any information or data related to at least one of biological or chemical substances. In some embodiments, the sequencing system 7800A is a workstation, which can be similar to a benchtop device or desktop computer. For example, most (or all) of the systems and components for performing the desired reactions can be within a common housing 7802.

[0398] In certain embodiments, the sequencing system 7800A is a nucleic acid sequencing system configured for a variety of applications, including, but not limited to, de novo sequencing, resequencing of whole genomes or targeted genomic regions, and metagenomics. The sequencer may also be used for DNA or RNA analysis. In some embodiments, the sequencing system 7800A may also be configured to generate reaction sites within a biosensor. For example, the sequencing system 7800A may be configured to receive a sample and generate surface-bound clusters of clonovirus-amplified nucleic acids from the sample. Each cluster may constitute or be part of a reaction site within a biosensor.

[0399] The exemplary sequencing system 7800A may include a system receptacle or interface 7810 configured to interact with a biosensor 7812 to effect a desired reaction within the biosensor 7812. In the description that follows with respect to FIG. 78A , the biosensor 7812 is loaded into the system receptacle 7810. However, it is understood that a cartridge including the biosensor 7812 may be inserted into the system receptacle 7810, and that in some conditions the cartridge may be temporarily or permanently removed. As discussed above, the cartridge may include, among other things, fluid control and fluid storage components.

[0400] In certain embodiments, the sequencing system 7800A is configured to perform multiple parallel reactions within the biosensor 7812. The biosensor 7812 includes one or more reaction sites where desired reactions can occur. The reaction sites may be immobilized, for example, on a solid surface of the biosensor or on beads (or other movable substrates) located within corresponding reaction chambers of the biosensor. The reaction sites may include, for example, clusters of clonovirus-amplified nucleic acids. The biosensor 7812 may include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell attached thereto. The flow cell may include one or more flow paths that receive solutions from the sequencing system 7800A and direct the solutions toward the reaction sites. Optionally, the biosensor 7812 may be configured to engage a thermal element to transfer thermal energy into or out of the flow paths.

[0401] The sequencing system 7800A may include various components, assemblies, and systems (or subsystems) that interact with each other to perform a predetermined method or assay protocol for biological or chemical analysis. For example, the sequencing system 7800A includes a system controller 7806 that may be in communication with the various components, assemblies, and subsystems of the sequencing system 7800A and also includes a biosensor 7812. For example, in addition to the system vessel 7810, the sequencing system 7800A may also include a fluid control system 7808 that controls the flow of fluids in the fluidic network and biosensor 7812 of the sequencing system 7800A, a fluid reservoir system 7814 that holds any fluids (e.g., gases or liquids) that may be used by the bioassay system, a temperature control system 7804 that may regulate the temperature of the fluids in the fluidic network, the fluid reservoir system 7814, and / or the biosensor 7812, and an illumination system 7816 configured to illuminate the biosensor 7812. As described above, when a cartridge having a biosensor 7812 is loaded into the system container 7810, the cartridge may also include fluid control and fluid storage components.

[0402] The sequencing system 7800A may also include a user interface 7818 for interacting with a user. For example, the user interface 7818 may include a display 7820 for displaying or requesting information from the user and a user input device 7822 for receiving user input. In some embodiments, the display 7820 and the user input device 7822 are the same device. For example, the user interface 7818 may include a touch-sensitive display configured to detect the presence of an individual touch and identify the location of the touch on the display. However, other user input devices 7822, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc., may also be used. As described in more detail below, the sequencing system 7800A may communicate with various components, including a biosensor 7812 (e.g., in the form of a cartridge), to perform desired reactions. The sequencing system 7800A may also be configured to analyze data obtained from the biosensor to provide desired information to the user.

[0403] The system controller 7806 may comprise a microcontroller, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRAs), a logic circuit, or any other circuit or processor capable of performing the functions described herein. The above examples are merely illustrative and are not intended to limit the definition and / or meaning of the term system controller. In an exemplary embodiment, the system controller 7806 executes a set of instructions stored in one or more storage elements, memories, or modules to at least one of acquire and analyze detection data. The detection data may include multiple sequences of pixel signals, thereby allowing sequences of pixel signals from each of millions of sensors (or pixels) to be detected over many base call cycles. The storage elements may be in the form of information sources or physical memory elements within the sequencing system 7800A.

[0404] The instruction set may include various commands that instruct the sequencing system 7800A or biosensor 7812 to perform specific operations, such as the methods and processes of various embodiments described herein. The instruction set may be in the form of a software program, which may form part of a tangible, non-transitory computer-readable medium or media. As used herein, the terms "software" and "firmware" are used interchangeably and include any computer program stored in memory executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and thus not limiting of the types of memory that may be used to store a computer program.

[0405] The software may be in various forms, such as system software or application software. Furthermore, the software may be in the form of a collection of separate programs, or a program module or portion of a program module within a larger program. The software may also include modular programming in the form of object-oriented programming. After acquiring the detection data, the detection data may be processed automatically by the sequencing system 7800A in response to user input, or in response to a request made by another processing machine (e.g., a remote request via a communications link). In the illustrated embodiment, the system controller 7806 includes an analysis module 7844. In other embodiments, the system controller 7806 does not include the analysis module 7844, but instead has access to the analysis module 7844 (e.g., the analysis module 7844 may be separately hosted on the cloud).

[0406] The system controller 7806 may be connected to the biosensor 7812 and other components of the sequencing system 7800A via a communication link. The system controller 7806 may also be communicatively connected to an off-site system or server. The communication link may be a wire, a cord, or wireless. The system controller 7806 may receive user input or commands from a user interface 7818 and user input devices 7822.

[0407] The fluid control system 7808 includes a fluid network and is configured to direct the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with a biosensor 7812 and a fluid storage system 7814. For example, fluid may be selected from the fluid storage system 7814 and directed to the biosensor 7812 in a controlled manner, or fluid may be drawn from the biosensor 7812 and directed to, for example, a waste reservoir within the fluid storage system 7814. Although not shown, the fluid control system 7808 may include a flow sensor that detects the flow rate or pressure of the fluid within the fluid network. The sensor may be in communication with the system controller 7806.

[0408] The temperature control system 7804 is configured to regulate the temperature of fluids in different regions of the fluid network, the fluid reservoir system 7814, and / or the biosensor 7812. For example, the temperature control system 7804 may include a thermal circulator that interacts with the biosensor 7812 and controls the temperature of fluids flowing along reaction sites within the biosensor 7812. The temperature control system 7804 may also regulate the temperature of solid elements or components of the sequencing system 7800A or the biosensor 7812. Although not shown, the temperature control system 7804 may include sensors for detecting the temperature of the fluids or other components. The sensors may be in communication with the system controller 7806.

[0409] The fluid storage system 7814 is in fluid communication with the biosensor 7812 and may store various reaction components or reactants used to carry out a desired reaction. The fluid storage system 7814 may also store fluids for cleaning or rinsing the fluidic network and the biosensor 7812 and for diluting reactants. For example, the fluid storage system 7814 may include various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, etc. Additionally, the fluid storage system 7814 may also include a waste reservoir for receiving waste from the biosensor 7812. In embodiments including a cartridge, the cartridge may include one or more of a fluid storage system, a fluid control system, or a temperature control system. Accordingly, one or more of the components described herein for these systems may be contained within the cartridge housing. For example, the cartridge may have various reservoirs for storing samples, reagents, enzymes, other biomolecules, buffers, aqueous, and non-polar solutions, waste, etc. Thus, one or more of the fluid reservoir system, fluid control system, or temperature control system may be removably engaged with the bioassay system via a cartridge or other biosensor.

[0410] The illumination system 7816 may include a light source (e.g., one or more LEDs) and multiple optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. The optical components may be, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, etc. In embodiments using an illumination system, the illumination system 7816 may be configured to direct excitation light to the reaction sites. As an example, a fluorophore may be excited by a green wavelength of light, and therefore the wavelength of the excitation light may be approximately 532 nm. In one embodiment, the illumination system 7816 is configured to generate illumination parallel to a surface normal to the surface of the biosensor 7812. In another embodiment, the illumination system 7816 is configured to generate illumination that is off-angled relative to the surface normal to the surface of the biosensor 7812. In yet another embodiment, the illumination system 7816 is configured to generate illumination having multiple angles, including some parallel illumination and some off-angle illumination.

[0411] The system receptacle or interface 7810 is configured to engage the biosensor 7812 in at least one of mechanical, electrical, and fluidic manners. The system receptacle 7810 can hold the biosensor 7812 in a desired orientation to facilitate fluid flow through the biosensor 7812. The system receptacle 7810 can also include electrical contacts configured to engage the biosensor 7812 so that the sequencing system 7800A can communicate with and / or provide power to the biosensor 7812. Additionally, the system receptacle 7810 can include a fluid port (e.g., a nozzle) configured to engage the biosensor 7812. In some embodiments, the biosensor 7812 is removably coupled to the system receptacle 7810 both electrically and fluidically.

[0412] Additionally, the sequencing system 7800A may communicate remotely with other systems or networks, or with other bioassay systems 7800A. Detection data obtained by the bioassay system 7800A may be stored in a remote database.

[0413] FIG. 78B is a block diagram of a system controller 7806 that can be used in the system of FIG. 78A. In one implementation, the system controller 7806 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include algorithms (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithms for performing particular processes. The system controller 7806 is conceptually illustrated as a collection of modules, but may also be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 7806 may be implemented using an off-the-shelf PC with a single processor or multiple processors, with functional operations distributed among the processors. As a further option, the modules described below may be implemented using a hybrid configuration in which certain modular functions are implemented using dedicated hardware, while remaining modular functions are implemented using an off-the-shelf PC, etc. The modules may also be implemented as software modules within a processing unit.

[0414] During operation, the communication port 7850 may transmit information (e.g., commands) to the biosensor 7812 ( FIG. 78A ) and / or the subsystems 7808, 7814, 7804 ( FIG. 78A ). In some embodiments, the communication port 7850 may output multiple sequences of pixel signals. The communication link 7834 may receive user input from the user interface 7818 ( FIG. 78A ) and transmit data or information to the user interface 7818. Data from the biosensor 7812 or the subsystems 7808, 7814, 7804 may be processed in real time by the system controller 7806 during a bioassay session. Additionally or alternatively, data may be temporarily stored in system memory during a bioassay session and processed slower than real time or for offline operation.

[0415] As shown in FIG. 78B, the system controller 7806 may include multiple modules 7826-7848 in communication with a main control module 7824 along with a central processing unit (CPU) 7852. The main control module 7824 may be in communication with a user interface 7818 ( FIG. 78A ). While the modules 7826-7848 are shown in direct communication with the main control module 7824, the modules 7826-7848 may also be in direct communication with each other, the user interface 7818, and the biosensor 7812. The modules 7826-7848 may also be in communication with the main control module 7824 via other modules.

[0416] The plurality of modules 7826-7848 includes system modules 7828-7832, 7826 that communicate with subsystems 7808, 7814, 7804, and 7816, respectively. The fluid control module 7828 may communicate with the fluid control system 7808 to control valves and flow sensors in the fluid network to control the flow of one or more fluids through the fluid network. The fluid storage module 7830 may notify a user when fluid is low or when a waste reservoir is at or near capacity. The fluid storage module 7830 may also communicate with a temperature control module 7832 so that fluids can be stored at a desired temperature. The illumination module 7826 may communicate with the illumination system 7816 to illuminate the reaction sites at specified times during a protocol, such as after a desired reaction (e.g., a binding event) has occurred. In some embodiments, the illumination module 7826 may communicate with the illumination system 7816 to illuminate the reaction sites at a specified angle.

[0417] The plurality of modules 7826-7848 may also include an apparatus module 7836 that communicates with the biosensor 7812 and an identification module 7838 that determines identification information associated with the biosensor 7812. The apparatus module 7836 may, for example, communicate with the system receptacle 7810 to confirm that the biosensor has established electrical and fluidic connectivity with the sequencing system 7800A. The identification module 7838 may receive a signal that identifies the biosensor 7812. The identification module 7838 may use the identification information of the biosensor 7812 to provide other information to a user. For example, the identification module 7838 may determine and subsequently display the lot number, manufacturing date, or a recommended protocol to be run on the biosensor 7812.

[0418] The plurality of modules 7826-7848 also includes an analysis module 7844 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 7812. The analysis module 7844 includes memory (e.g., RAM or flash) for storing the detection / image data. The detection data can include multiple sequences of pixel signals, such that sequences of pixel signals from each of millions of sensors (or pixels) can be detected over many base call cycles. The signal data can be stored for subsequent analysis or transmitted to the user interface 7818 to display desired information to a user. In some embodiments, the signal data can be processed by a solid-state image sensor (e.g., a CMOS image sensor) before the analysis module 7844 receives the signal data.

[0419] Analysis module 7844 is configured to acquire image data from the photodetector in each of a plurality of sequencing cycles. The image data is derived from the luminescence signals detected by the photodetector and processes the image data for each of the plurality of sequencing cycles via neural network-based template generator 1512 and / or neural network-based base caller 1514 to generate base calls for at least some of the analytes in each of the plurality of sequencing cycles. The photodetector may be part of one or more overhead cameras (e.g., a CCD camera in an Illumina GAIIx that takes images of the clusters on biosensor 7812 from above) or may be part of biosensor 7812 itself (e.g., a CMOS image sensor in an Illumina iSeq that is below the clusters on biosensor 7812 and takes images of the clusters from the bottom).

[0420] The output of the photodetector is a sequence image showing the intensity emissions of each cluster and their surrounding background. The sequence images show the intensity emissions generated as a result of incorporating nucleotides into a sequence during sequencing. The intensity emissions are from the associated analytes and their surrounding background. The sequence images are stored in memory 7848.

[0421] Protocol modules 7840 and 7842 communicate with main control module 7824 to control the operation of subsystems 7808, 7814, and 7804 in implementing a predetermined assay protocol. Protocol modules 7840 and 7842 may include instruction sets for instructing sequencing system 7800A to perform specific operations according to a predetermined protocol. As shown, the protocol module may be a sequence synthesis (SBS) module 7840 configured to issue various commands to execute a sequence-by-sequence synthesis process. In SBS, the extension of nucleic acid primers along a nucleic acid template is monitored to determine the sequence of nucleotides in the template. The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme) or ligation (e.g., catalyzed by a ligase enzyme). In certain polymer-based SBS embodiments, fluorescently labeled nucleotides are added to primers (thereby extending the primers) in a template-dependent manner, such that detection of the order and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, to initiate the first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc. can be delivered into / through a flow cell containing an array of nucleic acid templates. The nucleic acid templates may be located at corresponding reaction sites. Primer extension can detect incorporated labeled nucleotides through an imaging event, and these reaction sites can be detected. During the imaging event, an illumination system 7816 can provide excitation light to the reaction sites. Optionally, the nucleotides can further include a reversible termination feature that terminates further primer extension once the nucleotide is added to the primer. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer so that continued extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, in another embodiment using a reversible termination, a command can be given to deliver a deblocking reagent to the flow cell (either before or after detection).One or more commands can be given to effect wash(s) between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing techniques are described, for example, in Bentley et al., Nature 456:53-59 (20078), WO 04 / 0178497, U.S. Patent No. 7,057,026, WO 91 / 066778, U.S. Patent No. 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, U.S. Patent No. 7,405,2781, and 20078 / 01470780782 (each of which is incorporated herein by reference).

[0422] In the nucleotide delivery step of the SBS cycle, any single type of nucleotide can be delivered at a time, or multiple different nucleotide types (e.g., A, C, T, and G) can be delivered. In nucleotide delivery configurations where only a single type of nucleotide is present at a time, different nucleotides do not need to have distinct labels because they can be distinguished based on the temporal separation inherent in individualized delivery. Thus, a sequencing method or device can use single-color detection. For example, the excitation source only needs to provide excitation at a single wavelength or a single wavelength range. In nucleotide delivery configurations where delivery results in multiple different nucleotides being present in the flow cell at a given time, the sites at which different nucleotide types incorporate can be distinguished based on the different fluorescent labels attached to each nucleotide type in the mixture. For example, four different nucleotides, each bearing one of four different fluorophores, can be used. In one embodiment, four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, fewer than four different excitation sources can be used, but optical filtering of the excitation radiation from a single source can be used to generate different excitation radiation ranges in the flow cell.

[0423] In some embodiments, fewer than four different colors can be detected in a mixture having four different nucleotides. For example, pairs of nucleotides can be detected at the same wavelength but can be distinguished based on differences in intensity for one member of the pair, or based on a change to one member of the pair (e.g., via chemical, photochemical, or physical modification) that results in the appearance or disappearance of a distinct signal compared to the signal detected for the other member of the pair. Exemplary devices and methods for distinguishing four different nucleotides using detection of fewer than four colors are described, for example, in U.S. Patent Nos. 61 / 5378,294 and 61 / 619,78778, which are incorporated herein by reference in their entireties. U.S. Patent Application No. 13 / 624,200, filed September 21, 2012, is incorporated by reference in its entirety.

[0424] The multiple protocol modules may also include a sample preparation (or generation) module 7842 configured to issue commands to the fluidic control system 7808 and the temperature control system 7804 to amplify the product in the biosensor 7812. For example, the biosensor 7812 may be coupled to a sequencing system 7800A. The amplification module 7842 can issue instructions to the fluidic control system 7808 to deliver the necessary amplification components to a reaction chamber in the biosensor 7812. In other embodiments, the reaction site may already contain some components for amplification, such as template DNA and / or primers. After delivering the amplification components to the reaction chamber, the amplification module 7842 can instruct the temperature control system 7804 to cycle through different temperature steps according to a known amplification protocol. In some embodiments, amplification and / or nucleotide incorporation is performed isothermally.

[0425] The SBS module 7840 can issue commands to perform bridge PCR, in which clusters of clonal amplicons are formed over localized regions within the flow cell channel. After generating amplicons via bridge PCR, the amplicons may be "linearized" to create single-stranded template DNA, and sstDNA and sequencing primers may be hybridized to universal sequences flanking the region of interest. For example, reversible terminator-based sequencing by synthesis methods can be used, as described above or as follows.

[0426] Each base calling or sequencing cycle can extend the sstDNA by a single base, which can be achieved, for example, by using a modified DNA polymerase and a mixture of four types of nucleotides. Different types of nucleotides can have unique fluorescent labels, and each nucleotide can further have a reversible terminator that allows only a single base to be incorporated in each cycle. After addition of a single base to the sstDNA, excitation light can enter the reaction site and fluorescent emission can be detected. After detection, the fluorescent label and terminator can be chemically cleaved from the sstDNA. Another similar base calling or sequencing cycle can be as follows: In such a sequencing protocol, the SBS module 7840 can instruct the fluid control system 7808 to direct the flow of reagent and enzyme solutions through the biosensor 7812. Exemplary reversible terminator-based SBS methods that can be utilized with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 017878901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 027814714709, WO 05 / 0657814, U.S. Patent Application Publication No. 2005 / 014700900, WO 06 / 078B199, and WO 07 / 01470251, each of which is incorporated by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in U.S. Pat. No. 7,541,444, U.S. Pat. No. 7,057,026, U.S. Pat. No. 7,414,14716, U.S. Pat. No. 7,427,673, U.S. Pat. No. 7,566,537, U.S. Pat. No. 7,592,435, and WO 07 / 1478353678, each of which is incorporated herein by reference in its entirety.

[0427] In some embodiments, the amplification and SBS modules may operate in a single assay protocol, eg, template nucleic acids are amplified and subsequently sequenced within the same cartridge.

[0428] The sequencing system 7800A may also allow the user to reconfigure the assay protocol. For example, the determination system 7800A may provide the user with options through the user interface 7818 to modify the determined protocol. For example, if it is determined that the biosensor 7812 is to be used for amplification, the sequencing system 7800A may request the temperature of the annealing cycle. Additionally, the sequencing system 7800A may issue a warning to the user if the user provides user input that is not generally accepted for the selected assay protocol.

[0429] In an embodiment, biosensor 7812 includes millions of sensors (or pixels), each of which generates a plurality of sequences of pixel signals over successive base call cycles. Analysis module 7844 detects the plurality of sequences of pixel signals and attributes them to corresponding sensors (or pixels) according to the row and / or column positions of the sensors on the array of sensors.

[0430] FIG. 79 is a simplified block diagram of a system for analyzing sensor data, such as base call sensor output, from a sequencing system 7800A. In the example of FIG. 79, the system includes a configurable processor 7846. The configurable processor 7846 can execute a base caller (e.g., neural network-based template generator 1512 and / or neural network-based base caller 1514) in coordination with a runtime program executed by a central processing unit (CPU) 7852 (i.e., a host processor). The sequencing system 7800A includes a biosensor 7812 and a flow cell. The flow cell may include one or more tiles in which clusters of genetic material are exposed to a series of analyte flows that are used to trigger reactions within the clusters to identify bases in the genetic material. A sensor detects reactions for each cycle of sequencing in each tile of the flow cell to provide tile data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a sequence of base calls for each group of genetic material sensed during the base calling operation.

[0431] The system in this example includes a CPU 7852 that executes a runtime program for coordinating base calling operations, and memory 7848B that stores sequences of arrays of tile data, base call reads generated by the base calling operations, and other information used in the base calling operations. Also in this figure, the system includes memory 7848A that stores configuration files (or files), such as FPGA bit files and model parameters for a neural network used to configure and reconfigure configurable processor 7846. Sequencing system 7800A can include programs for configuring the configurable processor, and in some embodiments, can include a reconfigurable processor that runs a neural network.

[0432] The sequencing system 7800A is coupled to a configurable processor 7846 by a bus 7902. The bus 7902 can be implemented using high-throughput technology, such as bus technology compatible with the PCIe (Peripheral Component Interconnect Express) standard currently maintained and developed by the PCI-SIG (PCI Special Interest Group) standard. Also, in this example, a memory 7848A is coupled to the configurable processor 7846 by a bus 7906. The memory 7848A can be on-board memory located on a circuit board with the configurable processor 7846. The memory 7848A is used for fast access by the configurable processor 7846 of working data used in base calling operations. The bus 7906 can also be implemented using high-throughput technology, such as bus technology compatible with the PCIe standard.

[0433] Configurable processors, including field programmable gate arrays (FPGAs), coarse-grained configurable reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than can be achieved using general-purpose processors running computer programs. Configuring a configurable processor involves compiling a functional description to generate a configuration file, sometimes referred to as a bitstream or bitfile, and distributing the configuration file to configurable elements on the processor. The configuration file configures the circuit to set dataflow patterns, including the use of distributed memory and other on-chip memory resources, lookup table contents, the operation of configurable logic blocks, and configurable execution units such as configurable interconnects and other elements of a configurable array. A configuration file is reconfigurable if it can be changed in the field by modifying a loaded configuration file. For example, the configuration file may be stored in a volatile SRAM element, a non-volatile read-write memory element, or distributed among an array of configurable elements on a configurable or reconfigurable processor. Various commercially available configurable processors are suitable for use in basecall operations as described herein.Examples include Google's Tensor Processing Unit (TPU)™, GX4 Rackmount Series™, GX9 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ (Snapdragon processors™), NVIDIA Volta™, NVIDIA's Drive PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, Arm DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, Xilinx Alveo™ U200, Xilinx Alveo™ U250, Xilinx Alveo™ U280, Intel / Altera Stratix™ GX2800, Intel / Altera Stratix™ GX2800, and Intel Stratix™ GX10M. In some embodiments, the host CPU may be implemented on the same integrated circuit as the configurable processor.

[0434] The embodiments described herein use a configurable processor 7846 to implement the neural network-based template generator 1512 and / or the neural network-based base caller 1514. The configuration file for the configurable processor 7846 can be implemented by specifying the logic functions to be performed using a high-level description language HDL or a register-transfer level RTL language specification. This specification can be compiled using resources designed for a selected configurable processor to generate the configuration file. The same or similar specifications can be compiled to generate designs for application-specific integrated circuits that may not be configurable processors.

[0435] Thus, alternatives to the configurable processor 7846 in all embodiments described herein include a configured processor including an application specific ASIC or dedicated integrated circuit or set of integrated circuits, or is a system-on-chip SOC device, or a graphics processing unit (GPU) processor or a coarse-grained reconfigurable architecture (CGRA) processor configured to perform neural network-based base call operations as described herein.

[0436] In general, the configurable and configured processors described herein that are configured to perform neural network execution are referred to herein as neural network processors.

[0437] Configurable processor 7846, in this example, is configured using a program executed by CPU 7852 or by a configuration file loaded by other source to configure an array of configurable elements 7916 (e.g., configuration logic blocks (CLBs), such as look-up tables (LUTs), flip-flops, arithmetic processing units (PMUs), and compute memory units (CMUs), configurable I / O blocks, programmable interconnects) to perform base calling functions. In this example, the configuration includes data flow logic 7908 coupled to buses 7902 and 7906, which performs the function of distributing data and control parameters among elements used in base calling operations.

[0438] Configurable processor 7846 is also configured with base calling execution logic 7908 to execute neural network-based template generator 1512 and / or neural network-based base caller 1514. Logic 7908 includes multi-cycle execution clusters (e.g., 7914), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters can be selected according to tradeoffs with the desired throughput of operation and available resources on configurable processor 7846.

[0439] The multi-cycle execution clusters are coupled to the data flow logic 7908 by data paths 7910 implemented using configurable interconnect and memory resources on the configurable processor 7846. The multi-cycle execution clusters are also coupled to the data flow logic 7908 by control paths 7912 implemented using configurable interconnect and memory resources, such as the configurable processor 7846. Provisions are provided to provide control signals indicative of available execution clusters, provide input units to the available execution clusters for execution of the neural network-based template generator 1512 and / or neural network-based base caller 1514, provide trained parameters for the neural network-based template generator 1512 and / or neural network-based base caller 1514, provide output patches of base call classification data, and other control data used in the execution of the neural network-based template generator 1512 and / or neural network-based base caller 1514.

[0440] The configurable processor 7846 is configured to execute an execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514 using the trained parameters to generate classification data for detection cycles of the base calling operation. The execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514 executes to generate classification data for subject detection cycles of the base calling operation. The execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514 operates in a sequence including a number N of arrays of tile data from each detection cycle of the N detection cycles, where the N detection cycles provide sensor data for different base calling operations for one base position per operation in the time sequence in the embodiments described herein. Optionally, some of the N detection cycles can exit the sequence as needed according to the particular neural network model being executed. The number N can be any number greater than 1. In some embodiments described herein, the N detection cycles represent a set of detection cycles for at least one detection cycle preceding the subject detection cycle and at least one detection cycle following the subject detection cycle. Embodiments described herein include those in which the number N is an integer greater than or equal to 5.

[0441] The data flow logic 7908 is configured to use an input unit for a given run that includes tile data for N arrays of spatially aligned patches to move the tile data and at least some of the trained parameters of the model parameters from memory 7848A to the configurable processor 7846 for execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514. The input unit can be moved by direct memory access operations in a single DMA operation, or in smaller units that move between available time slots in coordination with the execution of the deployed neural network.

[0442] The tile data of the sensing cycles described herein can include an array of sensor data having one or more features. For example, the sensor data can include two images analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The tile data can also include metadata about the images and sensors. For example, in a base calling embodiment, the tile data can include information about the alignment of the images with clusters, such as distance from center information indicating the distance of each pixel in the array of sensor data from the center of the group of genetic material on the tile.

[0443] During execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514, the tile data may also include data generated during execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514, referred to as intermediate data that is not recalculated but can be recalculated during execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514. For example, during execution of the neural network-based template generator 1512 and / or the neural network-based base caller 1514, the data flow logic 7908 may write the intermediate data to memory 7848A in place of sensor data for a given patch of the array of tile data. Such embodiments are described in more detail below.

[0444] As shown, a system for analyzing base calling sensor output is described that includes a memory (e.g., 7848A) accessible by a runtime program that stores tile data including sensor data for tiles from a detection cycle of a base calling operation. The system also includes a neural network processor, such as configurable processor 7846, having access to the memory. The neural network processor is configured to perform a neural network execution using trained parameters to generate classification data for the detection cycle. As described herein, the neural network execution operates on a sequence of N arrays of tile data from each of the N detection cycles comprising a subject cycle to generate classification data for the subject cycle. Data flow logic 908 is provided to move the tile data and trained parameters from the memory to the neural network processor for execution of the neural network, using input units including data for the N arrays of spatially aligned patches from each of the N detection cycles.

[0445] Also described is a system in which the neural network processor has access to a memory and includes a plurality of execution clusters, the execution clusters being configured to execute a neural network. Data flow logic 7908 accesses the memory and executes a cluster in the plurality of execution clusters to provide an input unit of tile data to an available execution cluster in the plurality of execution clusters, the input unit including a number N of spatially aligned patches of the array of tile data from each sensing cycle, and causing the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches of the subject sensing cycle, where N is greater than 1.

[0446] Figure 80 is a simplified diagram illustrating aspects of a base calling operation, including runtime program functionality executed by a host processor. In this diagram, image sensor output from a flow cell is provided on line 8000 to image processing thread 8001, which can perform processes on the image, such as aligning and positioning individual tiles within an array of sensor data and resampling the image, which can be used by a process to calculate a tile cluster mask for each tile within the flow cell, which can be used by a process to identify pixels within the array of sensor data that correspond to clusters of genetic material on the corresponding tile of the flow cell. The output of image processing thread 8001 is provided on line 8002 to dispatch logic 8010 within the CPU, which transfers the data to data cache 8004 (e.g., SSD storage) on high-speed bus 8003 or high-speed bus 8005, depending on the state of the base calling operation, to neural network processor hardware 8020, such as configurable processor 7846 of Figure 79. The processed and transformed image can be stored on data cache 8004 to detect previously used cycles. Hardware 8020 returns the classification data output by the neural network to dispatch logic 8080, which passes the information to data cache 8004 or on line 8011 to thread 8002, which uses the classification data to perform base calling and quality score calculations and can arrange the data in a standard format for base called reads. The output of thread 8002, which performs base calling and quality score calculations, is provided on line 8012 to thread 8003, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base call output to a specified destination for consumption by the customer.

[0447] In some embodiments, the host may include a thread (not shown) that performs final processing of the output of the hardware 8020 supporting the neural network. For example, the hardware 8020 may provide classification data output from the final layer of a multi-cluster neural network. The host processor may perform output activation functions, such as a softmax function, over the classification data to populate the data used by the base calling and quality scoring thread 8002. The host processor may also perform input operations (not shown), such as batch normalization of the tile data before input to the hardware 8020.

[0448] FIG. 81 is a simplified diagram of a configurable processor 7846 configuration such as that of FIG. 79. In FIG. 81, the configurable processor 7846 includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 8100 including data flow logic 7908 as described with reference to FIG. 79. The wrapper 8100 manages interfacing and coordination with the runtime program in the CPU via CPU communication link 8109 and manages communication with on-board DRAM 8102 (e.g., memory 7848A) via DRAM communication link 8110. The data flow logic 7908 in the wrapper 8100 provides patch data obtained by traversing an array of tile data on the on-board DRAM 8102 to clusters 8101 for a number N of cycles, and obtains and delivers process data 8115 from clusters 8101 to the on-board DRAM 8102. The wrapper 8100 also manages the transfer of data between the on-board DRAM 8102 and host memory for both input arrays of tile data and output patches of classification data. The wrapper forwards the patch data on line 8113 to the assigned cluster 8101. The wrapper provides trained parameters such as weights and biases on line 8112 to the cluster 8101 obtained from onboard DRAM 8102. The wrapper provides configuration and control data on line 8111 to the cluster 8101 provided by, or generated in response to, a runtime program on the host via CPU communication link 8109. The cluster can also provide status signals on line 8116 to the wrapper 8100 that are used in conjunction with control signals from the host to provide spatially aligned patch data and to run a multi-cycle neural network on the patch data using the resources of the cluster 8101.

[0449] As described above, multiple clusters may reside on a single configurable processor managed by a wrapper 8100 configured to run on corresponding ones of the multiple patches of tile data. Each cluster may be configured to provide classification data for base calls in a subject detection cycle using the tile data of multiple sensing cycles described herein.

[0450] In an example system, model data, including kernel data such as filter weights and biases, can be sent from the host CPU to the configurable processor, so that the model can be updated as a function of cycle number. Base calling operations can typically involve hundreds of sensing cycles. In some embodiments, base calling operations can include paired end reads. For example, model-trained parameters can be updated every 20 cycles (or other number of cycles) or according to an update pattern implemented in a particular system and neural network model. In some embodiments, where a sequence for a given string within a genetic cluster on a tile includes paired end reads that include a first portion extending from (or above) the first end of the string and a second portion extending above (or below) the second end of the string, trained parameters can be updated at the transition from the first portion to the second portion.

[0451] In some embodiments, image data for multiple cycles of sensor data for a tile can be sent from the CPU to the wrapper 8100. The wrapper 8100 can optionally perform some preprocessing and transformation of the sensor data and write that information to on-board DRAM 8102. The input tile data for each sensing cycle can include an array of sensor data containing 4000 x 3000 pixels / tile or more per tile, with two features representing the colors of two images of the tile and including one or two bytes per pixel. In an embodiment where the number N is three sensing cycles used in each implementation of the multi-cycle neural network, the array of tile data for each implementation of the multi-cycle neural network can consume several hundred megabytes per tile. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and tile.

[0452] In operation, if a multi-cycle cluster is available, the wrapper assigns the patch to the cluster. The wrapper fetches the next patch of tile data for the cross section of the tile and sends it to the assigned cluster along with the appropriate control and configuration information. The cluster can be configured with enough memory on the configurable processor to have enough memory to hold the patch of data, including the patch, from multiple cycles in some systems being processed in place, and in various embodiments is processed using a ping-pong buffer technique or a raster scan technique.

[0453] When an assigned cluster completes its operation of the neural network for the current patch and generates an output patch, it signals the wrapper. The wrapper either reads the output patch from the assigned cluster, or the assigned cluster pushes data to the wrapper. The wrapper then assembles the output patch for the processed tile in DRAM 8102. Once processing of the entire tile is complete and the output patch of data is transferred to DRAM, the wrapper sends the processed output array back to the host / CPU in a specific format. In some embodiments, the on-board DRAM 8102 is managed by memory management logic within the wrapper 8100. A runtime program can control the sequencing operations to complete analysis of all tile data arrays for every cycle executed in a continuous flow to provide real-time analysis.

[0454] (Technical Improvements and Terminology) Base calling involves incorporating or attaching a fluorescently labeled tag with the analyte. The analyte may be a nucleotide or oligonucleotide, and the tag may be a specific nucleotide type (A, C, T, or G). Excitation light is directed at the tagged analyte, causing the tag to emit a detectable fluorescent signal or intensity emission. The intensity emission indicates the photons emitted by the excitation tag chemically bound to the analyte.

[0455] Throughout this application, including the claims, when "images, image data, or image regions showing the intensity emissions of analytes and their surrounding background are used, they refer to the intensity emissions of tags attached to the analytes. Those skilled in the art will understand that the intensity emissions of attached tags represent or correspond to the intensity emissions of the analytes to which the tags are attached, and are therefore used interchangeably. Similarly, a characteristic of an analyte refers to a tag attached to the analyte, or a characteristic of the intensity emissions from the attached tag. For example, the center of the analyte refers to the center of the intensity emissions emitted by the tag attached to the analyte. In another example, the background around the analyte refers to the background around the intensity emissions emitted by the tag attached to the analyte.

[0456] The literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, trees, and web pages, are expressly incorporated by reference in their entirety. In the event that one or more of the incorporated literature and similar materials differs from or contradicts this application, including but not limited to defined terms, term usage, described techniques, etc., this application controls.

[0457] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid template or its complement, e.g., a nucleic acid sample, such as a DNA or RNA polynucleotide or other nucleic acid sample. Accordingly, certain implementations of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher rates of collection of DNA or RNA sequence data, greater efficiency in sequence data collection, and / or lower costs of obtaining such sequence data, compared to previously available methods.

[0458] The disclosed technology uses neural networks to identify the centers of solid-phase nucleic acid clusters and analyze optical signals generated during sequencing of such clusters, unambiguously distinguishing between adjacent, neighboring, or overlapping clusters and assigning sequencing signals to single, discrete source clusters. These and related embodiments thus enable the recovery of meaningful information, such as sequence data, from regions of high-density cluster arrays, where useful information may not have been previously obtained from such regions due to confounding effects of overlapping or closely spaced neighboring clusters, including the effects of overlapping signals (e.g., as used in nucleic acid sequencing).

[0459] As described in more detail below, in certain embodiments, compositions are provided that include a solid support immobilized with one or more nucleic acid clusters, as provided herein. Each cluster contains multiple immobilized nucleic acids of the same sequence and has a distinct center with a detectable central label as provided herein, which is distinguishable from the nucleic acids immobilized in the surrounding area within the cluster. Also described herein are methods for producing and using such clusters with distinct centers.

[0460] Embodiments of the present disclosure will find use in many situations where benefits derive from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, including high-throughput nucleic acid sequencing, development of image analysis algorithms for assigning optical or other signals to individual source clusters, and other applications where recognition of the center of immobilized nucleic acid clusters is desirable and beneficial.

[0461] In certain embodiments, the present invention contemplates methods related to high-throughput nucleic acid analysis, such as nucleic acid sequencing (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetics analysis, genomic methylation analysis, allele-specific primer extension (APSE), genetic diversity profiling, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequencing, and the like. Those skilled in the art will understand that a variety of different nucleic acids can be analyzed using the methods and compositions of the present invention.

[0462] Although implementations of the present invention are described in the context of nucleic acid sequencing, they are applicable in any field in which image data acquired at different times, spatial locations, or other temporal or physical aspects are analyzed. For example, the methods and systems described herein are useful in the fields of molecular biology and cell biology, where image data from microarrays, biological specimens, cells, organisms, etc., are acquired and analyzed at different times or perspectives. Images can be obtained using any number of techniques known in the art, including, but not limited to, fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomographic scanning, etc. As another example, the methods and systems described herein can be applied when image data acquired by surveillance, aerial, or satellite imaging techniques, etc., are acquired and analyzed at different times or perspectives. The methods and systems are particularly useful for analyzing images acquired within a field of view, in which observed specimens remain in the same location relative to each other within the field of view. However, specimens may have different characteristics in separate images; for example, specimens may appear different in separate images of the field of view. For example, the analyte may appear to be different in color for a given analyte detected in different images, may indicate a change in the intensity of the signal detected for a given analyte in different images, or even the appearance of a signal for a given analyte in one image and the disappearance of the signal for that analyte in another image.

[0463] The examples described herein may be used in a variety of biological or chemical processes and systems for academic or commercial analysis. More specifically, the examples described herein may be used in a variety of processes and systems in which it is desirable to detect an event, characteristic, quality, or property indicative of a specified reaction. For example, the examples described herein include optical detection devices, biosensors, and components thereof, as well as bioassay systems operating in conjunction with biosensors. In some embodiments, the devices, biosensors, and systems may include a flow cell and one or more optical sensors coupled (removably or fixedly) together in a substantially monolithic structure.

[0464] The devices, biosensors, and bioassay systems may be configured to perform multiple designated reactions that can be detected individually or collectively. The devices, biosensors, and bioassay systems may be configured to perform multiple cycles in which multiple designated reactions occur in parallel. For example, the devices, biosensors, and bioassay systems may be used to sequence high-density arrays of DNA features through repeated cycles of enzymatic manipulation and light or image detection / capture. Thus, the devices, biosensors, and bioassay systems (e.g., via one or more cartridges) may include one or more microfluidic channels that deliver reagents or other reaction components into the reaction solution, biosensors, and bioassay systems. In some examples, the reaction solution may be substantially acidic, such as comprising a pH of about 5 or less, or about 4 or less, or about 3 or less. In some other examples, the reaction solution may be substantially alkaline / basic, such as comprising a pH of about 8 or more, or about 9 or more, or about 10 or more. As used herein, the term "acidic" and grammatical variants thereof refer to pH values ​​less than about 7, and the terms "basic," "alkaline," and grammatical variants thereof refer to pH values ​​greater than about 7.

[0465] In some embodiments, the reaction sites are provided or spaced in a predetermined manner, such as a uniform or repeating pattern. In some other embodiments, the reaction sites are randomly distributed. Each of the reaction sites can be associated with one or more light guides and one or more light sensors that detect light from the associated reaction site. In some embodiments, the reaction sites are located within a reaction recess or chamber that can at least partially compartmentalize a designated reaction.

[0466] As used herein, a "designated reaction" includes a change in at least one chemical, electrical, physical, or optical property (or quality) of a chemical or biological substance of interest, such as an analyte of interest. In certain examples, the designated reaction is a positive binding event, such as the incorporation of a fluorescently labeled biomolecule with a fluorescently labeled biomolecule. More generally, the designated reaction may be a chemical conversion, chemical change, or chemical interaction. The designated reaction may also be a change in an electrical property. In certain examples, the designated reaction includes the incorporation of an analyte with a fluorescently labeled molecule. The analyte may be an oligonucleotide, and the fluorescently labeled molecule may be a nucleotide. The designated reaction may be detected when excitation light is directed at the oligonucleotide bearing the labeled nucleotide, causing the fluorophore to emit a detectable fluorescent signal. In alternative examples, the detected fluorescence is the result of chemiluminescence or bioluminescence. The specified reaction can also, for example, increase fluorescence (or Forster) resonance energy transfer (FRET) by bringing a donor fluorophore into close proximity with an acceptor fluorophore, decrease FRET by separating the donor and acceptor fluorophores, increase fluorescence by separating a quencher from a fluorophore, or decrease fluorescence by co-localizing a quencher and fluorophore.

[0467] As used herein, "reaction solution," "reaction component," or "reactant" includes any substance that can be used to obtain at least one specified reaction. For example, potential reaction components include, for example, reagents, enzymes, samples, other biomolecules, and buffers. A reaction component may be delivered to a reaction site in solution and / or immobilized at the reaction site. A reaction component may interact directly or indirectly with another substance, such as an analyte of interest, immobilized at the reaction site. As noted above, the reaction solution may be substantially acidic (i.e., comprising a relatively high degree of acidity) (e.g., including a pH of about 5 or less, a pH of about 4 or less), or a pH of about 3 or less, or substantially alkaline / basic (i.e., comprising a relatively high degree of alkaline / basicity) (e.g., including a pH of about 8 or more, a pH of about 9 or more, or a pH of about 10 or more).

[0468] As used herein, the term "reaction site" refers to a localized region where at least one specified reaction can occur. A reaction site may include a support surface of a reaction structure or substrate onto which a substance can be immobilized. For example, a reaction site may include the surface of a reaction structure (which may be disposed within a channel of a flow cell) having reaction components thereon, e.g., a colony of nucleic acids thereon. In some such examples, the nucleic acids in the colonies have the same sequence, e.g., are clonal copies of a single-stranded or double-stranded template. However, in some examples, a reaction site may contain only a single nucleic acid molecule, e.g., in single-stranded or double-stranded form.

[0469] The multiple reaction sites may be randomly distributed along the reaction structure or may be arranged in a predetermined manner (e.g., parallel in a matrix such as a microarray). A reaction site may also include a reaction chamber or recess that at least partially defines a spatial region or volume configured to compartmentalize a specified reaction. As used herein, the term "reaction chamber" or "reaction recess" includes a defined spatial region of a support structure (often in fluid communication with a flow path). A reaction recess may be at least partially isolated from the surrounding environment or spatial region. For example, multiple reaction recesses may be separated from each other by a shared wall, such as a detection surface. As a more specific example, a reaction recess may be a nanocell that includes a depression, well, groove, cavity, or depression defined by the inner surface of the detection surface, and may have an opening or aperture (i.e., an open side) to allow the nanocell to be in fluid communication with the flow path.

[0470] In some embodiments, the reaction recess of the reaction structure is sized and shaped relative to a solid (including a semi-solid) so that the solid can be fully or partially inserted therein. For example, the reaction recess may be sized and shaped to accommodate a capture bead. The capture bead may have clonovirus-amplified DNA or other material thereon. Alternatively, the reaction recess may be sized and shaped to receive a number of beads or solid substrates. As another example, the reaction recess may be filled with a porous gel or material configured to control diffusion or filter fluids or solutions that may flow into the reaction recess.

[0471] In some embodiments, an optical sensor (e.g., a photodiode) is associated with a corresponding reaction site. The optical sensor associated with a reaction site is configured to detect light emission from the associated reaction site via at least one light guide when a designated reaction occurs at the associated reaction site. In some cases, multiple optical sensors (e.g., several pixels of a light detection or camera device) may be associated with a single reaction site. In other cases, a single optical sensor (e.g., a single pixel) may be associated with a single reaction site or with a group of reaction sites. The optical sensors, reaction sites, and other features of the biosensor may be configured such that at least a portion of the light is directly detected by the optical sensor without being reflected.

[0472] As used herein, "biological or chemical substances" include biomolecules, subject samples, subject analytes, and other chemical compounds. Biological or chemical substances may be used to detect, identify, or analyze other chemical compounds, or to act as intermediaries for studying or analyzing other chemical compounds. In certain examples, biological or chemical substances include biomolecules. As used herein, "biomolecules" include at least one of biopolymers, nucleotides, nucleic acids, polynucleotides, oligonucleotides, proteins, enzymes, polypeptides, antibodies, antigens, ligands, receptors, polysaccharides, carbohydrates, polyphosphates, cells, tissues, organisms, or fragments thereof, or any other biologically active chemical compounds, such as analogs or mimetics of the foregoing species. In further examples, the biological or chemical substances or biomolecules detect the product of another reaction, such as an enzyme or reagent, e.g., the product of an enzyme or reagent, such as an enzyme or reagent used to detect pyrophosphate in a pyrosequencing reaction. Enzymes and reagents useful for pyrophosphate detection are described, for example, in U.S. Patent Publication No. 2005 / 0244870 A1, which is incorporated by reference in its entirety.

[0473] The biomolecules, samples, and biological substances or chemicals may be naturally occurring or synthetic and may be suspended in a solution or mixture within the reaction wells or regions. The biomolecules, samples, and biological substances or chemicals may also be bound to a solid phase or gel material. The biomolecules, samples, and biological substances or chemicals may also include pharmaceutical compositions. In some cases, the biomolecules, samples, and biological substances or chemicals of interest may be referred to as targets, probes, or analytes.

[0474] As used herein, "biosensor" includes devices including a reaction structure with multiple reaction sites configured to detect a specified reaction occurring at or near the reaction site. A biosensor may include a solid-state photodetector or "imaging" device (e.g., a CCD or CMOS photodetector) and, optionally, a flow cell attached thereto. The flow cell may include at least one flow path in fluid communication with the reaction site. As one specific example, the biosensor is configured to be fluidically and electrically coupled to a biological assay system. The bioassay system may deliver reaction solutions to the reaction sites according to a predetermined protocol (e.g., sequence number synthesis) and perform multiple imaging events. For example, the bioassay system may flow reaction solutions along the reaction sites. At least one of the reaction solutions may contain four types of nucleotides with the same or different fluorescent labels. The nucleotides may bind to corresponding oligonucleotides, etc., in the reaction sites. The bioassay system can then illuminate the reaction sites using an excitation light source (e.g., a solid-state light source such as a light-emitting diode (LED)). The excitation light may have a predetermined wavelength or wavelengths, including a range of wavelengths. Fluorescent labels excited by incident excitation light can provide an emission signal (e.g., of a different wavelength or wavelengths of light than the excitation light, and potentially different from each other) that can be detected by a photosensor.

[0475] As used herein, the term "immobilized," when used with reference to a biomolecule or biological substance or chemical, includes substantially attaching the biomolecule or biological substance or chemical to a surface, such as the detection surface of an optical detection device or a reaction structure. For example, a biomolecule or biological substance or chemical may be immobilized on the surface of a reaction structure using adsorption techniques, including non-covalent bonding (e.g., electrostatic forces, van der Waals, and hydrophobic interfacial dehydration), as well as covalent bonding techniques in which functional groups or linkers facilitate binding of the biomolecule to the surface. Immobilizing a biomolecule or biological substance or chemical on a surface may be based on the properties of the surface, the liquid medium carrying the biomolecule or biological substance or chemical, and the properties of the biomolecule or biological substance or chemical itself. In some cases, the surface may be functionalized (e.g., chemically or physically modified) to facilitate immobilization of a biomolecule (or biological substance or chemical) on the surface.

[0476] In some embodiments, nucleic acids can be immobilized on a reaction structure, such as the surface of a reaction well. In certain embodiments, the devices, biosensors, bioassay systems, and methods described herein can include the use of naturally occurring nucleotides and enzymes configured to interact with naturally occurring nucleotides. Naturally occurring nucleotides include, for example, ribonucleotides or deoxyribonucleotides. Naturally occurring nucleotides can be in monophosphate, diphosphate, or triphosphate form and can have a base selected from adenine (A), thymine (T), uracil (U), guanine (G), or cytosine (C). However, it will be understood that non-naturally occurring nucleotides, modified nucleotides, or analogs of the above nucleotides can be used.

[0477] As described above, biomolecules, biological substances, or chemicals may be immobilized at reaction sites within the reaction recesses of the reaction structure. Such biomolecules or biological substances may be physically held or immobilized within the reaction recesses by interference fitting, adhesion, covalent bonding, or entrapment. Examples of articles or solids that can be placed within the reaction recesses include polymer beads, pellets, agarose gel, powders, quantum dots, or other solids that can be compressed and / or held within the reaction chamber. In certain embodiments, the reaction recesses may be coated or filled with a hydrogel layer that can covalently bond to DNA oligonucleotides. In certain examples, nucleic acid superstructures such as DNA balls can be placed within the reaction recesses by, for example, attaching them to the inner surface of the reaction recesses or by immersing them in a liquid within the reaction recesses. DNA balls or other nucleic acid superstructures can be formed and then placed within the reaction recesses. Alternatively, DNA balls can be synthesized in situ in the reaction recesses. The substances immobilized within the reaction recesses can be in a solid, liquid, or gaseous state.

[0478] As used herein, the term "analyte" is intended to mean a point or region of a pattern that can be distinguished from other points or regions according to their relative position. An individual analyte can include one or more molecules of a particular type. For example, an analyte can include a single target nucleic acid molecule having a particular sequence, or an analyte can include several nucleic acid molecules having the same sequence (and / or its complementary sequence). Different molecules that are different analytes of a pattern can be differentiated from one another according to the location of the analyte within the pattern. Exemplary analytes include wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, ridges on a substrate, pads of gel material on a substrate, or channels in a substrate.

[0479] Any of a variety of target analytes to be detected, characterized, or identified can be used in the devices, systems, or methods described herein. Exemplary analytes include, but are not limited to, nucleic acids (e.g., DNA, RNA, or analogs thereof), proteins, polysaccharides, cells, antibodies, epitopes, receptors, ligands, enzymes (e.g., kinases, phosphatases, or polymerases), small molecule drug candidates, cells, viruses, organisms, etc.

[0480] The terms "analyte," "nucleic acid," "nucleic acid molecule," and "polynucleotide" are used interchangeably herein. In various embodiments, a nucleic acid may be used as a template (e.g., a nucleic acid template or a nucleic acid complement complementary to a nucleic acid template) as provided herein for certain types of nucleic acid analysis, including, but not limited to, nucleic acid amplification, nucleic acid expression analysis, and / or nucleic acid sequencing, or a suitable combination thereof. Nucleic acids in certain implementations include, for example, linear polymers of deoxyribonucleotides in 3'-5' phosphodiester chains, or deoxyribonucleic acid (DNA), such as single- and double-stranded DNA, genomic DNA, copy DNA or complementary DNA (cDNA), recombinant DNA, or any form of synthetic or modified DNA. In other embodiments, nucleic acids include, for example, linear polymers of ribonucleotides in 3'-5' phosphodiester or other linkages, such as ribonucleic acid (RNA), including single- and double-stranded RNA, messenger (mRNA), copy or complementary RNA (cRNA), or spliced ​​mRNA, ribosomal RNA, small nuclear RNA (snoRNA), microRNA (miRNA), small interfering RNA (sRNA), piRNA (piRNA), or any form of synthetic or modified RNA. Nucleic acids used in the compositions and methods of the invention may vary in length and may be intact or full-length molecules or fragments, or smaller portions of larger nucleic acid molecules. In certain embodiments, nucleic acids may bear one or more detectable labels, as described elsewhere herein.

[0481] The terms "specimen," "cluster," "nucleic acid cluster," "nucleic acid colony," and "DNA cluster" are used interchangeably and refer to multiple copies of a nucleic acid template and / or its complement attached to a solid support. Typically, in certain preferred embodiments, a nucleic acid cluster comprises multiple copies of a template nucleic acid and / or its complement attached to a solid support via their 5' ends. The copies of the nucleic acid strands that make up a nucleic acid cluster can be in single-stranded or double-stranded form. The copies of the nucleic acid template present within a cluster can have nucleotides at corresponding positions that differ from each other due to, for example, the presence of a labeled moiety. Corresponding positions can also include analog structures with different chemical structures but similar Watson-Crick base pairing properties, such as uracil and thymine.

[0482] Colonies of nucleic acids may also be referred to as "nucleic acid clusters." Nucleic acid colonies can optionally be generated by cluster amplification or bridge amplification techniques, as described in more detail elsewhere herein. Multiple repeats of a target sequence can be present in a single nucleic acid molecule, such as a disruptor generated using a rolling circle amplification procedure.

[0483] The nucleic acid clusters of the present invention can have different shapes, sizes, and densities depending on the conditions used. For example, the clusters can be substantially circular, polyhedral, donut-shaped, or ring-shaped. The diameter of the nucleic acid cluster can be designed to be about 0.2 μm to about 6 μm, about 0.3 μm to about 4 μm, about 0.4 μm to about 3 μm, about 0.5 μm to about 2 μm, about 0.75 μm to about 1.5 μm, or any intermediate diameter. In certain embodiments, the diameter of the nucleic acid cluster is about 0.5 μm, about 1 μm, about 1.5 μm, about 2 μm, about 2.5 μm, about 3 μm, about 4 μm, about 5 μm, or about 6 μm. The diameter of the nucleic acid cluster can be affected by numerous parameters, including, but not limited to, the number of amplification cycles performed in producing the cluster, the length of the nucleic acid template, or the density of primers attached to the surface on which the cluster is formed. The density of the nucleic acid cluster is typically less than 0.1 μm / mm 2 , 1 / mm 2 , 10 / mm 2 , 100 / mm 2 , 1,000 / mm 2 , 10,000 / mm 2 ~100,000 / mm 2 The present invention, in part, is directed to higher density nucleic acid clusters, e.g., 100,000 / mm 2 ~1,000,000 / mm 2 , and 1,000,000 / mm 2 ~10,000,000 / mm 2 Further plans are underway.

[0484] As used herein, an "analyte" is a specimen or region of interest within a field of view. When used in connection with a microarray device or other molecular analysis device, an analyte refers to a region occupied by similar or identical molecules. For example, an analyte can be an amplification oligonucleotide or any other group of polynucleotides or polypeptides having the same or similar sequence. In other embodiments, an analyte can be any element or group of elements that occupy a physical region on a sample. For example, an analyte can be a packet of land, a body of water, etc. When analytes are imaged, each analyte has some area. Thus, in many embodiments, an analyte is not simply a single pixel.

[0485] The distance between analytes can be described in any number of ways. In some embodiments, the distance between analytes can be described from the center of one analyte to the center of another analyte. In other embodiments, the distance can be described from the edge of one analyte to the edge of another analyte, or between the outermost identifiable points of each analyte. The edge of an analyte can be described as a theoretical or actual physical boundary on the chip, or some point within the boundary of the analyte. In other embodiments, the distance can be described with respect to a fixed point on the sample or an image of the sample.

[0486] Generally, some embodiments are described herein with respect to analytical methods. It will be understood that systems for performing the methods in an automated or semi-automated manner are also provided. Thus, the present disclosure provides a neural network-based template generation and base calling system, which can include a processor, a storage device, and a program for image analysis, the program including instructions for performing one or more of the methods described herein. Thus, the methods described herein can be performed, for example, on a computer having components described herein or known in the art.

[0487] The methods and systems described herein are useful for analyzing any of a variety of objects. Particularly useful objects are solid supports or solid surfaces with attached analytes. The methods and systems described herein offer advantages when used with objects having repeating patterns of analytes in the xy plane. One example is a microarray having a collection of cells, viruses, nucleic acids, proteins, antibodies, carbohydrates, small molecules (such as drug candidates), biologically active molecules, or other analytes of interest.

[0488] The number of applications of arrays containing analytes containing biological molecules such as nucleic acids and polypeptides has increased. Such microarrays typically contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) probes, which are specific for nucleotide sequences present in humans and other organisms. In certain applications, for example, individual DNA or RNA probes can be attached to individual analytes on the array. Test samples, such as those from known humans or organisms, can be exposed to the array so that target nucleic acids (e.g., gene fragments, mRNA, or amplicons) hybridize to complementary probes for each analyte in the array. The probes can be labeled through target-specific processes (e.g., due to labels present on the target nucleic acids or due to enzyme labels on the probes or targets present in hybridized form in the analyte). Analytes can then be examined by scanning specific light frequencies over the analytes to identify which target nucleic acids are present in the sample.

[0489] Biological microarrays can be used for gene sequencing and similar applications. Generally, gene sequencing involves determining the order of nucleotides in a length of target nucleic acid, such as a fragment of DNA or RNA. Relatively short sequences are typically sequenced in each analyte, and the resulting sequence information can be used in various bioinformatics methods to reliably determine the sequence of many widely varying lengths of genetic material from which the fragments are derived. Automated computer-based algorithms for signature fragment identification have been developed and have more recently been used in genome mapping, gene identification, and their functions. Microarrays are particularly useful for characterizing genome content because of the large number of variants present, which is an alternative to conducting numerous experiments for individual probes and targets. Microarrays are an ideal format for conducting such studies in a practical manner.

[0490] Any of a variety of analyte arrays (also called "microarrays") known in the art can be used in the methods or systems described herein. A typical array contains analytes, each having an individual probe or a population of probes. In the latter case, the population of probes in each analyte is typically homogeneous, having a single type of probe. For example, in the case of nucleic acid sequences, each analyte can have multiple nucleic acid molecules, each having a common sequence. However, in some embodiments, the population in each analyte of the array can be heterogeneous. Similarly, a protein sequence can have analytes having a single protein or a population of proteins, typically, but not necessarily, having the same amino acid sequence. Probes can be attached to the surface of the array, for example, by covalently linking the probe to the surface or through non-covalent interaction(s) between the probe and the surface. In some embodiments, probes, such as nucleic acid molecules, can be attached to the surface via a gel layer, as described, for example, in U.S. Patent Application No. 13 / 784,368 and U.S. Patent Application No. 2011 / 0059865(A1), each of which is incorporated herein by reference.

[0491] Exemplary arrays include, but are not limited to, BeadChip arrays available from Illumina, Inc. (San Diego, Calif.), or others, such as those described below, in which probes are attached to beads present on a surface (e.g., beads in wells on a surface). See U.S. Pat. Nos. 6,266,459, 6,355,431, 6,770,441, 6,859,570, or 7,622,294, or PCT Publication WO 00 / 63437, each of which is incorporated herein by reference. Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays or other microarrays synthesized according to a technique sometimes referred to as VLSIPS™ (Very Large Scale Immobilized Polymer Synthesis) technology. Spotted microarrays can also be used in methods or systems according to some embodiments of the present disclosure. An exemplary spotted microarray is the CodeLink™ Array available from Amersham Biosciences. Another useful microarray is one produced using inkjet printing methods, such as SurePrint™ Technology available from Agilent Technologies.

[0492] Other useful sequences include those used in nucleic acid sequencing applications.For example, the sequences that have amplicons of genome fragments (often referred to as clusters) are particularly useful, such as those described in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, or U.S. Patent No. 7,057,026, or U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.Another type of sequence that is useful for nucleic acid sequencing is the sequence of particles that are generated from emulsion PCR technology. Examples are described in Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), WO 05 / 010145, U.S. Patent Application Publication No. 2005 / 0130173, or U.S. Patent Application Publication No. 2005 / 0064460, each of which is incorporated herein by reference in its entirety.

[0493] Arrays used for nucleic acid sequencing often have a random spatial pattern of nucleic acid analytes. For example, the HiSeq or MiSeq sequencing platforms available from Illumina Inc. (San Diego, Calif.) utilize flow cells in which nucleic acid sequences are formed by random seeding followed by bridge amplification. However, patterned arrays can also be used for nucleic acid sequencing or other analytical applications. Examples of patterned arrays, their uses, and methods for their use are described in U.S. Patent Application Nos. 13 / 787,396, 13 / 783,043, 13 / 784,368, U.S. Patent Application Publication Nos. 2013 / 0116153 (A1), and 2012 / 0316086 (A1), each of which is incorporated herein by reference. Analytes in such patterned arrays can be used to capture single nucleic acid template molecules for subsequent formation of homogeneous colonies, for example, via bridge amplification. Such patterned arrays are particularly useful for nucleic acid sequencing applications.

[0494] The size of the analytes on an array (or other object used in the methods or systems herein) can be selected to suit a particular application. For example, in some embodiments, the analytes of the array can have a size that accommodates only a single nucleic acid molecule. A surface with multiple analytes in this size range is useful for constructing an array of molecules for detection with single molecule resolution. Analytes in this size range are also useful for use in arrays with analytes each comprising a colony of nucleic acid molecules. Thus, each analyte in an array can be approximately 1 mm 2 Below, approximately 500μm 2 Below, approximately 100μm 2 Below, approximately 10μm 2 Below, approximately 1μm 2 Below, about 500nm 2 Less than or equal to about 100 nm 2 Below, approximately 10nm 2 Below, approximately 5nm 2 Less than or equal to 1 nm 2Alternatively or additionally, the specimens in the array may have an area of ​​about 1 mm 2 More than approximately 500μm 2 More than approximately 100μm 2 or more, about 10μm 2 or more, approximately 1μm 2 or more, about 500nm 2 Over 100nm 2 or more, about 10nm 2 or more, about 5nm 2 More than or about 1n m 2 or greater. Indeed, the analyte can have a size within a range between upper and lower limits selected from those exemplified above. While several size ranges for surface analytes have been exemplified with respect to nucleic acids and nucleic acid scales, it will be understood that analytes in these size ranges can be used in applications that do not involve nucleic acids. It will further be understood that the size of the analyte need not necessarily be limited to the scale used in nucleic acid applications.

[0495] In embodiments involving an object having multiple analytes, such as an array of analytes, the analytes can be distinct, separated by a space between them. Arrays useful in the invention can have analytes separated by an edge-to-edge distance of at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, arrays can have analytes separated by an edge-to-edge distance of at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can apply to the average edge-to-edge spacing and edge-to-edge spacing of the analytes, as well as the minimum or maximum spacing.

[0496] In some embodiments, the analytes in the array need not be distinct; instead, adjacent analytes can abut one another. Whether the analytes are distinct or not, the size of the analytes and / or the pitch of the analytes can be varied to allow the array to have a desired density. For example, the average analyte pitch in a regular pattern can be at most 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, 0.5 μm, or less. Alternatively or additionally, the average analyte pitch in a regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more. These ranges can also apply to the maximum or minimum pitch of a regular pattern. For example, the maximum analyte pitch in the regular pattern can be 100 μm or less, 50 μm or less, 10 μm or less, 5 μm or less, 1 μm or less, 0.5 μm or less, and / or the minimum analyte pitch in the regular pattern can be at least 0.5 μm, 1 μm, 5 μm, 10 μm, 50 μm, 100 μm, or more.

[0497] The density of analytes in an array can also be understood in terms of the number of analytes present per unit area. For example, the average density of analytes for an array is at least about 1×10 3 specimens / mm 2 , 1×10 4 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 6 specimens / mm 2 , 1×10 7 specimens / mm 2 , 1×10 8 specimens / mm 2 , or 1 × 10 9 specimens / mm 2 Alternatively, or in addition, the average density of analytes on the array can be at most about 1 x 10 9 specimens / mm 2 , 1×10 8 specimens / mm 2 , 1×10 7 specimens / mm 2, 1×10 6 specimens / mm 2 , 1×10 5 specimens / mm 2 , 1×10 4 specimens / mm 2 , or 1 × 10 3 specimens / mm 2 It can be the following:

[0498] The above ranges can apply, for example, to all or part of a regular pattern that includes all or part of an array of analytes.

[0499] The analytes within the pattern can have any of a variety of shapes. For example, when viewed in a two-dimensional plane, such as on the surface of an array, the analytes may appear rounded, circular, oval, rectangular, square, symmetrical, asymmetrical, triangular, polygonal, etc. The analytes can be arranged in a regular repeating pattern, including, for example, a hexagonal or rectilinear pattern. The pattern can be selected to achieve a desired level of packing. For example, circular analytes are optimally packed in a hexagonal arrangement. Of course, other packaging configurations can also be used for circular analytes, and vice versa.

[0500] A pattern can be characterized in terms of the number of analytes present in a subset that forms the smallest geometric unit of the pattern. A subset can include, for example, at least about 2, 3, 4, 5, 6, 10, or more analytes. Depending on the size and density of the analytes, a geometric unit can be as small as 1 mm 2 , 500 μm 2 , 100 μm 2 , 50 μm 2 , 10 μm 2 , 1 μm 2 , 500nm 2 , 100 nm 2 , 50nm 2 , 10nm 2 Alternatively or additionally, the geometric unit may be 10 nm 2 , 50nm 2 , 100 nm 2 , 500nm 2 , 1 μm2 , 10 μm 2 , 50 μm 2 , 100 μm 2 , 500 μm 2 , 1mm 2 The properties of the specimens in the geometric unit, such as shape, size, pitch, etc., can be selected from those described herein more generally for specimens in an array or pattern.

[0501] Arrays with a regular pattern of analytes may be ordered with respect to the relative location of the analytes, but random with respect to one or more other characteristics of each analyte. For example, in the case of nucleic acid sequences, the nucleic acid analytes may be regular with respect to their relative location, but random with respect to sequence knowledge regarding the nucleic acid species present in any particular analyte. As a more specific example, a nucleic acid sequence formed by seeding a repeating pattern of analytes with template nucleic acids and amplifying the template in each analyte to form copies of the template in the analytes (e.g., via cluster amplification or bridge amplification) will have a regular pattern of nucleic acid analytes, but will be random with respect to the distribution of sequences of the nucleic acids across the array. Thus, detection of the presence of nucleic acid material on an array can result in a repeating pattern of analytes, whereas sequence-specific detection can result in a non-repeating distribution of signal across the array.

[0502] It will be understood that descriptions of pattern, order, randomness, etc. herein relate not only to analytes on an object, such as analytes on an array, but also to analytes in an image. Thus, the pattern, order, randomness, etc. can exist in any of a variety of formats used to store, manipulate, or communicate image data, including, but not limited to, computer-readable media or computer components such as a graphical user interface or other output device.

[0503] As used herein, the term "image" is intended to mean a representation of all or a portion of an object. The representation may be an optically detected reproduction. For example, an image may be obtained from fluorescence, luminescence, scattering, or absorption signals. The portion of the object present in the image may be the surface or other xy plane of the object. Typically, an image is a two-dimensional representation, but in some cases, information in an image can be derived from three or more dimensions. An image need not include optically detected signals. Non-optical signals may instead be present. An image may be provided in a computer-readable format or medium, such as one or more of those described elsewhere herein.

[0504] As used herein, "image" refers to a reproduction or representation of at least a portion of a sample or other object. In some embodiments, the reproduction is an optical reproduction, e.g., produced by a camera or other optical detector. The reproduction may be a non-optical reproduction, e.g., a representation of electrical signals obtained from an array of nanopore analytes or a representation of electrical signals obtained from an ion-sensitive CMOS detector. In certain embodiments, non-optical reproductions may be excluded from the methods or apparatus described herein. The image may have a resolution capable of distinguishing analytes present at any of a variety of intervals, including, for example, those spaced less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm apart.

[0505] As used herein, "acquisition," "capture," and like terms refer to any part of the process of acquiring an image file. In some embodiments, data acquisition can include generating an image of the specimen, looking for a signal in the specimen, directing a detection device to look for or generate an image of the signal, and providing instructions for further analysis or transformation of the image file, and instructions for any number of transformations or manipulations of the image file.

[0506] As used herein, the term "template" refers to a representation of the location or relationship between signals or analytes. Thus, in some embodiments, the template is a physical grid having a representation of signals corresponding to analytes in a sample. In some embodiments, the template can be a chart, table, text file, or other computer file indicating locations corresponding to analytes. In the embodiments presented herein, the template is generated to track the location of analytes across a set of images of the sample captured at different reference points. For example, the template can be x, y coordinates or a series of values ​​describing the direction and / or distance of one analyte relative to another analyte.

[0507] As used herein, the term "specimen" can refer to an object or region of an object from which an image is captured. For example, in an embodiment in which an image is taken from the surface of soil, a parcel of land may be the specimen. In other embodiments in which biomolecular analysis is performed within a flow cell, the flow cell may be divided into any number of subdivisions, each of which may be a specimen. For example, the flow cell may be divided into various channels or lanes, and each lane may be further divided into 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 140, 160, 180, 200, 400, 600, 800, 1000, or more distinct regions to be imaged. One example flow cell has eight lanes, each divided into 120 specimens or tiles. In other embodiments, samples may be generated in multiple tiles, or even the entire flow cell. Thus, each specimen image can represent a larger surface area than is imaged.

[0508] References to ranges and sequential lists of numbers set forth herein will be understood to include not only the numbers recited but all real numbers between the recited numbers.

[0509] As used herein, a "reference point" refers to any temporal or physical distinction between images. In another preferred embodiment, the reference point is a time point. In a more preferred embodiment, the reference point is a time point or cycle during the sequencing reaction. However, the term "reference point" can also include other aspects that distinguish or separate images, such as angle, rotation, time, or other aspects that can distinguish or separate the images.

[0510] As used herein, a "subset of images" refers to a group of images within a set. For example, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, or any number of images selected from a set of images. In certain other embodiments, a subset may include 1, 2, 3, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, or less, or any number of images selected from a set of images. In another preferred embodiment, the images are obtained from one or more sequencing cycles, with four images associated with each cycle. Thus, for example, a subset may be a group of 16 images acquired over four cycles.

[0511] Base refers to a nucleotide base or nucleotide, (adenine), C (cytosine), T (thymine), or G (guanine). This application uses "base(s)" and "nucleotide(s)" interchangeably.

[0512] The term "chromosome" refers to a gene carrier of the present invention in a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized numbering system for individual human genome chromosomes is used herein.

[0513] The term "site" refers to a unique location (e.g., chromosome ID, chromosomal location and orientation) on a reference genome. In some embodiments, a site may be a residue, sequence tag, or the location of a segment on a sequence. The term "locus" may be used to refer to a specific location of a nucleic acid sequence or polymorphism on a reference chromosome.

[0514] The term "sample" as used herein typically refers to a sample derived from a biological fluid, cell, tissue, organ, or organism containing the nucleic acid to be sequenced and / or phased, or a sample derived from a mixture of nucleic acids containing at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine needle biopsy samples (e.g., surgical biopsy, needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations thereof, or fractions or derivatives thereof. While samples are often collected from human subjects (e.g., patients), samples can be collected from any organism that has chromosomes, including, but not limited to, dogs, cats, horses, goats, sheep, cattle, pigs, etc. Samples can be used directly as obtained from a biological source or after pretreatment to modify the sample's characteristics. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, and the like.

[0515] The term "sequence" includes or refers to a chain of nucleotides linked together. The nucleotides can be based on DNA or RNA. It should be understood that a sequence may contain multiple subsequences. For example, a single sequence (e.g., a PCR amplicon) may have 350 nucleotides. A sample read may contain multiple subsequences within these 350 nucleotides. For example, a sample read may include first and second flanking subsequences, e.g., 20-50 nucleotides. The first and second flanking subsequences may be located on either side of a repeat segment with corresponding subsequences (e.g., 40-100 nucleotides). Each of the flanking subsequences may include (or a portion of) a primer subsequence (e.g., 10-30 nucleotides). For ease of reading, the term "subsequence" is referred to as "sequence," but it is understood that two sequences need not be distinct from each other on a common strand. To distinguish between the various sequences described herein, the sequences may be labeled with different labels (e.g., target sequence, primer sequence, flanking sequence, reference sequence, etc.). Other terms, such as "allele," may be given different labels to distinguish between similar entities. Applications use "read(s)" and "sequence read(s)" interchangeably.

[0516] The term "paired end sequencing" refers to a sequencing method that sequences both ends of a target fragment. Paired end sequencing can facilitate the detection of genome rearrangements and repeated segments, as well as gene fusions and novel transcripts. Methods for paired end sequencing are described in International Publication No. 07010252, International Application No. PCTGB2007 / 003798, and U.S. Patent Application Publication No. 2009 / 0088327, each of which is incorporated herein by reference. In one example, a series of operations may be performed as follows: (a) generating clusters of nucleic acids; (b) linearizing the nucleic acids; (c) hybridizing a first sequencing primer and performing repeated cycles of extension, scanning, and deblocking; (d) "flipping" the target nucleic acid on the flow cell surface by synthesizing a complementary copy; (e) linearizing the resynthesized strand; and (f) hybridizing a second sequencing primer and performing repeated cycles of extension, scanning, and deblocking. The inversion operation can deliver the reagents described above for a single cycle of bridge amplification.

[0517] The term "reference genome" or "reference sequence" refers to a specific known genome sequence, either partial or complete, of any organism that can be used to reference an identified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. A genome includes both genetic and non-coding sequences of DNA. A reference sequence may be larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 times larger, or at least about 10 times larger, or at least about 10 times larger. In one example, the reference genome sequence is of the full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference chromosome is a chromosome sequence from human genome version hg19. Such sequences are sometimes referred to as chromosomal reference sequences, although the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (strands, etc.) of any species. In various embodiments, a reference genome is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, a reference sequence may be taken from a specific individual. In other embodiments, "genome" also covers so-called "graph genomes," which use specific storage formats and representations of genome sequences. In one embodiment, a graph genome stores data in a linear file. In another embodiment, a graph genome refers to a representation in which alternative sequences (e.g., different copies of a chromosome with small differences) are stored as different paths in a graph.Further information regarding the implementation of graph genomes can be found at https: / / www.biorxiv.org / content / biorxiv / early / 2018 / 03 / 20 / 194530.full.pdf, the contents of which are incorporated herein by reference in their entirety.

[0518] The term "read" refers to a collection of sequence data describing a fragment of a nucleotide sample or reference. The term "read" can refer to a sample read and / or a reference read. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample or reference. A read may be symbolically represented by the base pair sequence (ATCG) of the sample or reference fragment. The read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing instrument or indirectly from stored sequence information about the sample. In some cases, the read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region that can be aligned and specifically assigned to, for example, a chromosome or genomic region or gene.

[0519] Next-generation sequencing methods include, for example, sequencing by synthesis (Illumina), pyrosequencing (454), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), and sequencing by ligation (SOLiD sequencing). Depending on the sequencing method, the length of each read can vary from about 30 bp to over 10,000 bp. For example, DNA sequencing using a SOLiD sequencer generates nucleic acid reads of about 50 bp. In another example, Ion Torrent sequencing generates nucleic acid reads of up to 400 bp, while 454 pyrosequencing generates nucleic acid reads of about 700 bp. In yet another example, single-molecule real-time sequencing can generate reads of 10,000 bp to 15,000 bp. Thus, in certain embodiments, nucleic acid sequence reads have lengths of 30-100 bp, 50-200 bp, or 50-400 bp.

[0520] The terms "sample read," "sample sequence," or "sample fragment" refer to sequence data relating to a genomic sequence of interest from a sample. For example, a sample read includes sequence data from a PCR amplicon having forward and reverse primer sequences. The sequence data can be obtained from any selected sequence procedure. A sample read can be, for example, a sequence-by-synthesis (SBS) reaction, a sequencing-ligation reaction, or any other suitable sequencing method in which it is desirable to determine the length and / or identity of repetitive elements. A sample read can be a consensus (e.g., average or weighted) sequence derived from multiple sample reads. In certain embodiments, providing a reference sequence includes identifying a locus of interest based on primer sequences of a PCR amplicon.

[0521] The term "raw fragment" refers to sequence data of a portion of a genome sequence of interest that at least partially overlaps a specified or secondary location of interest within a sample read or sample fragment. Non-limiting examples of product fragments include double-stitched fragments, simple stitched fragments, and simple unstitched fragments. The term "raw" is used to indicate that a raw fragment contains sequence data that has some relationship to the sequence data in a sample read, regardless of whether the raw fragment exhibits supporting variants that correspond to and authenticate or confirm potential variants in the sample read. The term "raw fragment" does not indicate that the fragment necessarily contains supporting variants that verify the variant call in the sample read. For example, when a sample read is determined by a variant calling application to exhibit a first variant, the variant calling application may determine that one or more raw fragments lack a corresponding type of "supporting" variant that would otherwise be expected to occur given the variants in the sample read.

[0522] The terms "mapping," "aligned," "aligning," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining whether the reference sequence contains the read sequence. If the reference sequence is read, the read may be mapped to the reference sequence, or in certain alternative embodiments, may be mapped to a specif...

Claims

1. 1. A computer-implemented method for determining cluster metadata from image data generated based on one or more clusters, the computer-implemented method comprising: receiving as input said image data derived from a sequence of images; each image in the sequence of images represents an imaged region and shows intensity radiation of the one or more clusters and a background surrounding the one or more clusters in a corresponding one of a plurality of sequencing cycles of a sequencing run; processing the image data through a neural network to generate a feature map comprising alternative representations of the image data, the feature map comprising at least one of convolutional features or hidden state features generated from the image data; The neural network is trained on cluster metadata determination tasks, including determining cluster backgrounds, cluster centers, and cluster shapes; processing the feature maps through an output layer to generate an output comprising a classification score for the image data; the output identifies a center of the one or more clusters as a central unit at the center of mass of the discrete region based on coordinates of adjacent units forming the discrete region; applying thresholding to the output classification scores to classify a first subset of portions of the imaged region as background portions that depict the surrounding background of the one or more clusters; applying peak detection to the output classification scores to find peak intensities in the image data and classifying a second subset of the portions of the imaged area as containing the centers of the one or more clusters; and applying a segmenter to the classification scores to determine a shape of the one or more clusters by determining non-overlapping, contiguous portions of the imaged region.

2. The output generated by the output layer of the neural network is whether the portion of the imaged area represents the background or a cluster; and whether a part represents the center of several consecutive image parts that each represent the same cluster; and the output identifies the one or more clusters whose intensity radiation is described by the image data as discontinuous regions of adjacent units, and a background around the one or more clusters as a background unit that does not belong to any of the discontinuous regions.

3. the neighboring units of the discontinuous region have intensity values ​​weighted according to the distance of the neighboring units from a central unit in the discontinuous region to which the neighboring units belong, and the output is a binary map classifying each part as a cluster or background; The computer-implemented method of claim 2 , wherein the output is a ternary map that classifies each part as a cluster, background, or center.

4. determining position coordinates of the centers of the one or more clusters based on the peak intensities found in the image data; downscaling the position coordinates by an upsampling factor used to create the image data; 4. The computer-implemented method of claim 2 or 3, further comprising: storing the downscaled position coordinates in a memory for use in base calling the one or more clusters.

5. classifying the adjacent units within the discontinuous region as intra-cluster units belonging to the same cluster; 5. The computer-implemented method of claim 4, further comprising: storing, for each cluster, the classified and downscaled position coordinates of the intra-cluster units in the memory for use in base calling the one or more clusters.

6. obtaining training data for training the neural network; the training data includes a plurality of training examples and corresponding ground truth data; Each training example includes training image data from a sequence of image sets associated with a plurality of image channels; each image in the sequence of the image set represents an imaging area corresponding to a tile of a flow cell and depicts the intensity emission of a cluster on the tile and the surrounding background of the cluster captured for a particular image channel from the plurality of image channels in a particular one of a plurality of sequencing cycles of a sequencing run performed on the flow cell; each ground truth data identifying one or more characteristics of a respective portion of the plurality of training examples; training the neural network using gradient descent training techniques and generating training outputs for the plurality of training examples that progressively match the corresponding ground truth data; iteratively optimizing a loss function that minimizes the error between the training outputs and the corresponding ground truth; The computer-implemented method of claim 2, further comprising: updating parameters of the neural network based on the error.

7. the one or more characteristics of the respective portions of the plurality of training examples identified by the corresponding ground truth data are: Clusters whose intensity emissions are shown by the training image data of corresponding training examples as discontinuous regions of adjacent units; the cluster center as a central unit at the center of mass of the discontinuous region; and the background surrounding the cluster as a background unit that does not belong to any of the discontinuous regions, wherein the one or more characteristics indicate whether a unit is central or non-central.

8. and upon error convergence after the last iteration, storing updated parameters of the neural network in memory for application to further neural network-based template generation and base calling.

7. The computer-implemented method of claim 6, wherein in the corresponding ground truth data, the neighboring units within the discontinuous region have intensity values ​​weighted according to the neighboring units' distance from a central unit within the discontinuous region to which the neighboring units belong, and wherein in the ground truth data, the central unit has the highest intensity value within the respective one of the discontinuous regions.

9. 7. The computer-implemented method of claim 6, wherein the loss function is a mean squared error, and the error is minimized for each unit between the normalized intensity values ​​of corresponding units in the training output and the corresponding ground truth data.

10. In the training data, a plurality of training examples each include a different portion of each image in a corresponding sequence of an image set of the same tile as the training image data; At least some of the different portions overlap each other, and in the corresponding ground truth data: all units classified as cluster centers are assigned the same first predetermined class score, including a first softmax score; 7. The computer-implemented method of claim 6, wherein all units classified as non-central are assigned the same second predetermined class score comprising a second softmax score.

11. The loss function is a custom weighted binary cross-entropy loss, and the error is minimized unit-by-unit between the predicted classification score of the corresponding unit in the training output and the class score of the corresponding unit in the training output and the corresponding ground truth data, and in the corresponding ground truth data: All units classified as background are assigned the same first predetermined class score, all units classified as cluster centers are assigned the same second predetermined class score, and all units classified as cluster interiors are assigned the same third predetermined class score; 7. The computer-implemented method of claim 6.

12. applying thresholding to the output values ​​of the training output to classify a first subset of units of the training image data as background units indicative of the background surrounding the one or more clusters; finding peak intensities in the image data by applying peak detection to the output values ​​of the training outputs and classifying a second subset of units of the training image data as corresponding to central units comprising centers of the clusters; applying a segmenter to the output values ​​of the units to determine the shape of the cluster as non-overlapping regions of consecutive units separated by the background units and centered on the central unit, the segmenter starting from the central unit and determining for each central unit successively adjacent units that represent the same cluster whose centers are contained in the central unit; The computer-implemented method of any one of claims 6 to 11, further comprising:

13. 13. The computer-implemented method of claim 12, wherein the non-overlapping regions have irregular contours; The cluster strength of a given cluster is Identifying units of the training image data that contribute to the cluster strength of the given cluster based on corresponding non-overlapping regions of adjacent units that identify the shape of the given cluster; locating the identified units within one or more optical pixel resolution images generated for one or more image channels in the current sequencing cycle; interpolating the intensities of the identified units in each of the images, combining the interpolated intensities, and normalizing the combined interpolated intensities to generate a per-image cluster intensity for the given cluster in each of the images; combining the per-image cluster intensities for each of the images to determine a cluster intensity for the given cluster in the current sequencing cycle; The computer-implemented method further comprising determining by:

14. 13. The computer-implemented method of claim 12, wherein the non-overlapping regions have irregular contours; The cluster strength of a given cluster is Identifying units of the training image data that contribute to the cluster strength of the given cluster based on corresponding non-overlapping regions of adjacent units that identify the shape of the given cluster; locating the identified units in one or more corresponding optically upsampled unit resolution images, the pixel resolution images generated for one or more image channels in the current sequencing cycle; combining the intensities of the identified units in each of the upsampled images and normalizing the combined intensities to generate a per-image cluster intensity for the given cluster in each of the upsampled images; combining the per-image cluster intensities for each of the upsampled images to determine a cluster intensity for the given cluster in the current sequencing cycle; The computer-implemented method further comprising determining by:

15. the normalization is based on a normalization factor; the normalization factor is the number of the specified units; 15. A computer-implemented method according to claim 13 or 14.

Citation Information

Patent Citations

  • Automated detection and repositioning of micro-objects in microfluidic devices

    WO2018102748A1