AI-based quality scoring
By applying neural network-based methods in nucleic acid sequencing technology, using deep neural networks to analyze sequencing image data, the problems of throughput limitation and low data quality in the prior art are solved, and efficient and low-cost high-quality nucleic acid sequencing are achieved.
Patent Information
- Application Number
- CN202080005431.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-20
- Filing Date
- 2020-03-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-05-07
AI Technical Summary
In the existing high-throughput nucleic acid sequencing technology, the throughput limitation and data quality are not high, making it difficult to obtain high-quality nucleic acid sequencing data quickly and cost-effectively.
Using neural network-based methods and systems, deep neural networks, especially convolutional neural networks (CNNs), sequencing image data is analyzed, base detection and sequence generation are realized, and the quality and quantity of sequencing data are improved.
It improves the throughput level of nucleic acid sequencing technology, enhances the quality and accuracy of sequencing data, reduces sequencing costs, and is suitable for various biological applications.
Smart Images

Figure CN112789680B_ABST
Abstract
Description
[0001] Priority application
[0002] This patent application claims priority or the benefit of the following patent applications:
[0003] U.S. Provisional Patent Application No. 62 / 821,602, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on March 21, 2019 (Attorney Docket No. ILLM 1008-1 / IP-1693-PRV);
[0004] U.S. Provisional Patent Application No. 62 / 821,618, entitled “Artificial Intelligence-Based Generation of Sequencing Metadata,” filed on March 21, 2019 (Attorney Docket No. ILLM 1008-3 / IP-1741-PRV);
[0005] U.S. Provisional Patent Application No. 62 / 821,681, entitled “Artificial Intelligence-Based Base Calling,” filed on March 21, 2019 (Attorney Docket No. ILLM 1008-4 / IP-1744-PRV);
[0006] U.S. Provisional Patent Application No. 62 / 821,724, entitled “Artificial Intelligence-Based Quality Scoring,” filed on March 21, 2019 (Attorney Docket No. ILLM 1008-7 / IP-1747-PRV);
[0007] U.S. Provisional Patent Application No. 62 / 821,766, entitled “Artificial Intelligence-Based Sequencing,” filed on March 21, 2019 (Attorney Docket No. ILLM1008-9 / IP-1752-PRV);
[0008] Netherlands Patent Application No. 2023310, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on June 14, 2019 (Attorney Docket No. ILLM1008-11 / IP-1693-NL);
[0009] Netherlands Patent Application No. 2023311, entitled “Artificial Intelligence-Based Generation of Sequencing Metadata,” filed on June 14, 2019 (Attorney Docket No. ILLM 1008-12 / IP-1741-NL);
[0010] Netherlands Patent Application No. 2023312, entitled “Artificial Intelligence-Based Base Calling,” filed on June 14, 2019 (Attorney Docket No. ILLM 1008-13 / IP-1744-NL);
[0011] Netherlands Patent Application No. 2023314, entitled “Artificial Intelligence-Based QualityScoring,” filed on June 14, 2019 (Attorney Docket No. ILLM 1008-14 / IP-1747-NL); and
[0012] Netherlands Patent Application No. 2023316, entitled “Artificial Intelligence-Based Sequencing,” filed on June 14, 2019 (Attorney Docket No. ILLM 1008-15 / IP-1752-NL).
[0013] U.S. non-provisional patent application No. 16 / 825,987, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-16 / IP-1693-US);
[0014] U.S. non-provisional patent application No. 16 / 825,991, entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-17 / IP-1741-US);
[0015] U.S. non-provisional patent application No. 16 / 826,126, entitled “Artificial Intelligence-Based Base Calling,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-18 / IP-1744-US);
[0016] U.S. non-provisional patent application No. 16 / 826,134, entitled “Artificial Intelligence-Based Quality Scoring,” filed on March 20, 2020 (Attorney Docket No. ILLM 1008-19 / IP-1747-US);
[0017] U.S. non-provisional patent application No. 16 / 826,168, entitled “Artificial Intelligence-Based Sequencing,” filed on March 21, 2020 (Attorney Docket No. ILLM 1008-20 / IP-1752-PRV);
[0018] concurrently filed PCT Patent Application No. PCT________ (Attorney Docket No. ILLM1008-21 / IP-1693-PCT) entitled “Training Data Generation for Artificial Intelligence-Based Sequencing,” which was subsequently published as PCT Publication No. WO____________;
[0019] concurrently filed PCT Patent Application No. PCT________ (Attorney Docket No. ILLM 1008-22 / IP-1741-PCT) entitled “Artificial Intelligence-Based Generation of Sequencing Metadata,” which subsequently published as PCT Publication No. WO____________;
[0020] concurrently filed PCT Patent Application No. PCT________ (Attorney Docket No. ILLM 1008-23 / IP-1744-PCT) entitled “Artificial Intelligence-Based Base Calling,” which subsequently published as PCT Publication No. WO____________; and
[0021] concurrently filed PCT Patent Application No. PCT________ (Attorney Docket No. ILLM 1008-25 / IP-1752-PCT) entitled “Artificial Intelligence-Based Sequencing,” which subsequently published as PCT Publication No. WO____________;
[0022] These priority applications are hereby incorporated by reference as if fully set forth herein for all purposes.
[0023] Literature Incorporation
[0024] The following documents are incorporated by reference as if fully set forth herein for all purposes:
[0025] U.S. Provisional Patent Application No. 62 / 849,091, “Systems and Devices for Characterization and Performance Analysis of Pixel-Based Sequencing,” filed on May 16, 2019 (Attorney Docket No. ILLM 1011-1 / IP-1750-PRV);
[0026] U.S. Provisional Patent Application No. 62 / 849,132, entitled “Base Calling Using Convolutions,” filed on May 16, 2019 (Attorney Docket No. ILLM 1011-2 / IP-1750-PR2);
[0027] U.S. Provisional Patent Application No. 62 / 849,133, entitled “Base Calling Using Compact Convolutions,” filed on May 16, 2019 (Attorney Docket No. ILLM 1011-3 / IP-1750-PR3);
[0028] U.S. Provisional Patent Application No. 62 / 979,384, entitled “Artificial Intelligence-Based Base Calling of Index Sequences,” filed on February 20, 2020 (Attorney Docket No. ILLM 1015-1 / IP-1857-PRV);
[0029] U.S. Provisional Patent Application No. 62 / 979,414, entitled “Artificial Intelligence-Based Many-To-ManyBase Calling,” filed on February 20, 2020 (Attorney Docket No. ILLM 1016-1 / IP-1858-PRV);
[0030] U.S. Provisional Patent Application No. 62 / 979,385, entitled “Knowledge Distillation-Based Compression of Artificial Intelligence-Based Base Caller,” filed on February 20, 2020 (Attorney Docket No. ILLM 1017-1 / IP-1859-PRV);
[0031] U.S. Provisional Patent Application No. 62 / 979,412, entitled “Multi-Cycle Cluster Based Real Time Analysis System,” filed on February 20, 2020 (Attorney Docket No. ILLM 1020-1 / IP-1866-PRV);
[0032] U.S. Provisional Patent Application No. 62 / 979,411, entitled “Data Compression for Artificial Intelligence-Based Base Calling,” filed on February 20, 2020 (Attorney Docket No. ILLM 1029-1 / IP-1964-PRV);
[0033] U.S. Provisional Patent Application No. 62 / 979,399, entitled “Squeezing Layer for Artificial Intelligence-Based Base Calling,” filed on February 20, 2020 (Attorney Docket No. ILLM 1030-1 / IP-1982-PRV);
[0034] Liu P, Hemani A, Paul K, Weis C, Jung M, Wehn N, “3D-Stacked Many-Core Architecture for Biological Sequence Analysis Problems,” Int J Parallel Prog, 2017, vol. 45(no. 6): 1420–1460.
[0035] Z. Wu, K. Hammad, R. Mittmann, S. Magierowski, E. Ghafar-Zadeh, and X. Zhong, “FPGA-Based DNA Basecalling Hardware Acceleration,” in Proc. IEEE 61st Int. Midwest Symp. Circuits Syst., August 2018, pp. 1098–1101.
[0036] Z. Wu, K. Hammad, E. Ghafar-Zadeh, and S. Magierowski, “FPGA-Accelerated 3rd Generation DNA Sequencing,” IEEE Transactions on Biomedical Circuits and Systems, Vol. 14, No. 1, February 2020, pp. 65–74.
[0037] Prabhakar et al., “Plasticine: A Reconfigurable Architecture for Parallel Patterns,” ISCA 17, June 24–28, 2017, Toronto, ON, Canada;
[0038] M.Lin, Q.Chen and S.Yan, "Network in Network", Proc.of ICLR, 2014;
[0039] L.Sifre, "Rigid-motion Scattering for Image Classification", Ph.D. thesis, 2014;
[0040] L.Sifre and S.Mallat, "Rotation, Scaling and Deformation InvariantScattering for Texture Discrimination", Proc.of CVPR, 2013;
[0041] F.Chollet, "Xception: Deep Learning with Depthwise SeparableConvolutions", Proc.of CVPR, 2017;
[0042] X. Zhang, X. Zhou, M. Lin and J. Sun, “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices,” arXiv:1707.01083, 2017;
[0043] K. He, X. Zhang, S. Ren and J. Sun, “Deep Residual Learning for Image Recognition”, Proc. of CVPR, 2016;
[0044] S. Xie, R. Girshick, P. Dollár, Z. Tu and K. He, “Aggregated Residual Transformations for Deep Neural Networks”, Proc. of CVPR, 2017;
[0045] A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto and H. Adam, “Mobilenets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, arXiv:1704.04861, 2017;
[0046] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov and L. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, arXiv:1801.04381v3, 2018;
[0047] Z. Qin, Z. Zhang, X. Chen and Y. Peng, “FD-MobileNet: Improved MobileNet with a Fast Downsampling Strategy”, arXiv:1802.03750, 2018;
[0048] Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam, "Rethinking atrous convolution for semantic image segmentation", CoRR, abs / 1706.05587, 2017;
[0049] J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, et al., "Speed / accuracy trade-offs for modern convolutional object detectors", arXiv preprint arXiv:1611.10012, 2016;
[0050] S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, "WAVENET: A GENERATIVE MODEL FOR RAW AUDIO", arXiv:1609.03499, 2016;
[0051] S. S. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, "DEEP VOICE: REAL-TIME NEURAL TEXT-TO-SPEECH", arXiv:1702.07825, 2017;
[0052] F. Yu and V. Koltun, "MULTI-SCALE CONTEXT AGGREGATION BY DILATED CONVOLUTIONS", arXiv:1511.07122, 2016;
[0053] K. He, X. Zhang, S. Ren, and J. Sun, "DEEP RESIDUAL LEARNING FOR IMAGE RECOGNITION", arXiv:1512.03385, 2015;
[0054] R. K. Srivastava, K. Greff, and J. Schmidhuber, "HIGHWAY NETWORKS", arXiv:1505.00387, 2015;
[0055] G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, "DENSELY CONNECTED CONVOLUTIONAL NETWORKS", arXiv:1608.06993, 2017;
[0056] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "GOING DEEPER WITH CONVOLUTIONS", arXiv:1409.4842, 2014;
[0057] S. Ioffe and C. Szegedy, "BATCH NORMALIZATION: ACCELERATING DEEP NETWORK TRAINING BY REDUCING INTERNAL COVARIATE SHIFT", arXiv:1502.03167, 2015;
[0058] J. M. Wolterink, T. Leiner, M. A. Viergever, and I. "DILATED CONVOLUTIONAL NEURAL NETWORKS FOR CARDIOVASCULAR MR SEGMENTATION IN CONGENITAL HEART DISEASE", arXiv:1704.03669, 2017;
[0059] L.C. Piqueras, “AUTOREGRESSIVE MODEL BASED ON A DEEP CONVOLUTIONAL NEURAL NETWORK FOR AUDIO GENERATION,” Tampere University of Technology, 2016;
[0060] J. Wu, “Introduction to Convolutional Neural Networks”, Nanjing University, 2017;
[0061] "Illumina CMOS Chip and One-Channel SBS Chemistry", Illumina, Inc, 2018, page 2;
[0062] “skikit-image / peak.py at master”, GitHub, page 5, [retrieved on November 16, 2018], retrieved from the internet <URL: https: / / github.com / scikit-image / scikit-image / blob / master / skimage / feature / peak.py#L25 >
[0063] “3.3.9.11. Watershed and random walker for segmentation”, Scipy lecturenotes, page 2, [retrieved on November 13, 2018], retrieved from the internet <URL: http: / / scipy- lectures.org / packages / scikit-image / auto_examples / plot_segmentations.html >
[0064] Mordvintsev, Alexander and Revision, Abid K., "Image Segmentation with Watershed Algorithm", Version 43532856, 2013, p. 6, [retrieved on November 13, 2018], retrieved from the internet <URL: https: / / opencv-python-tutroals.readthedocs.io / en / latest / py_ tutorials / py_imgproc / py_watershed / py_water shed.html >
[0065] Mzur, “Watershed.py,” October 25, 2017, p. 3, [retrieved November 13, 2018], retrieved from the internet<URL:https: / / github.com / mzur / watershed / blob / master / Watershed.py> ;
[0066] Thakur, Pratibha, et al., “A Survey of Image Segmentation Techniques,” International Journal of Research in Computer Applications and Robotics, Vol. 2, No. 4, April 2014, pp. 158–165.
[0067] Long, Jonathan, et al., “Fully Convolutional Networks for Semantic Segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, vol. 4, April 1, 2017, p. 10.
[0068] Ronneberger, Olaf, et al., “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, May 18, 2015, p. 8;
[0069] Xie, W. et al., “Microscopy cell counting and detection with fully convolutional regression networks,” Computer methods in biomechanics and biomedical engineering: Imaging & Visualization, Vol. 6 (Vol. 3), pp. 283-292, 2018.
[0070] Xie, Yuanpu et al., "Beyond classification: structured regression for robust cell detection using convolutional neural network", International Conference on Medical Image Computing and Computer-Assisted Intervention, October 12, 2015, page 12;
[0071] Snuverink, IAF, “Deep Learning for Pixelwise Classification of Hyperspectral Images,” Master of Science Thesis, Delft University of Technology, November 23, 2017, p. 19.
[0072] Shevchenko, A., "Keras weighted categorical_crossentropy", Page 1, [Retrieved January 15, 2019], Retrieved from the internet <URL: https: / / gist.github.com / skeeet / cad06d584548f b45eece1d4e28cfa98b >
[0073] van den Assem, DCF, “Predicting periodic and chaotic signals using Wavenets,” Master of Science thesis, Delft University of Technology, August 18, 2017, pp. 3–38;
[0074] I.J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Convolutional Networks,” in Deep Learning, MIT Press, 2016; and
[0075] J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, and G. Wang, “RECENT ADVANCES IN CONVOLUTIONAL NEURAL NETWORKS,” arXiv:1512.07108, 2017. Technical Field
[0076] The technology disclosed herein relates to artificial intelligence-type computers and digital data processing systems, as well as corresponding data processing methods and products for simulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems). The technology also includes systems for inferring uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the technology disclosed herein relates to the use of deep neural networks, such as deep convolutional neural networks, for analyzing data. Background Art
[0077] The subject matter discussed in this section should not be considered prior art simply because it is mentioned in this section. Similarly, problems mentioned in this section or associated with the subject matter provided as background technology should not be considered to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to specific implementations of the claimed technology.
[0078] A deep neural network is an artificial neural network that uses multiple layers of nonlinear and complex transformations to continuously model high-level features. Deep neural networks provide feedback via backpropagation, which uses the difference between observed and predicted outputs to adjust parameters. Deep neural networks have evolved with the availability of large training datasets, the power of parallel and distributed computing, and sophisticated training algorithms. They have driven significant advances in many fields, such as computer vision, speech recognition, and natural language processing.
[0079] Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks are particularly successful in image recognition with an architecture comprising convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to exploit the sequential information of input data and have cyclic connections between building blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks for limited contexts have been proposed, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.
[0080] The goal of training a deep neural network is to optimize the weight parameters in each layer, gradually combining simpler features into complex ones, so that the most appropriate layered representation can be learned from the data. A single cycle of the optimization process proceeds as follows. First, given a training dataset, a forward pass sequentially computes the output of each layer and propagates the function signal forward through the network. In the final output layer, the target loss function measures the error between the inferred output and the given label. To minimize the training error, a backward pass uses the chain rule to backpropagate the error signal and calculate the gradient with respect to all weights in the entire neural network. Finally, an optimization algorithm based on stochastic gradient descent is used to update the weight parameters. While batch gradient descent performs parameter updates for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small data example. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency and momentum of the gradient of each parameter, respectively.
[0081] Another core element in deep neural network training is regularization, which refers to strategies designed to avoid overfitting and thus achieve good generalization performance. For example, weight decay adds a penalty factor to the objective loss function so that the weight parameters converge to a small absolute value. Dropout randomly removes hidden units from a neural network during training and can be thought of as an ensemble of possible subnetworks. To enhance the capabilities of dropout, new activation functions, maximum output, and a dropout variant of recurrent neural networks (called rnnDrop) have been proposed. In addition, batch normalization provides a new regularization method by normalizing the scalar features of each activation within a mini-batch and learning each mean and variance as parameters.
[0082] Given the multidimensionality and high-dimensionality of sequence data, deep neural networks hold great promise in bioinformatics research due to their broad applicability and enhanced predictive power. Convolutional neural networks have been used to address sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight-sharing strategy that is particularly useful for studying DNA because they can capture sequence motifs, which are short, recurring local patterns in DNA that are hypothesized to have significant biological functions. The hallmark of convolutional neural networks is the use of convolutional filters.
[0083] Unlike traditional classification methods based on carefully designed and handcrafted features, convolutional filters perform adaptive learning of features, similar to the process of mapping raw input data into an information representation of knowledge. In this sense, convolutional filters act as a series of motif scanners, as a set of such filters is able to identify relevant patterns in the input and update themselves during the training process. Recurrent neural networks can capture long-range dependencies in sequence data of varying lengths, such as protein or DNA sequences.
[0084] Therefore, there is an opportunity to use a principled deep learning-based framework for template generation and base calling.
[0085] In the era of high-throughput technologies, accumulating the most interpretable data at the lowest cost per work remains a major challenge. Cluster-based nucleic acid sequencing methods, such as those that use bridge amplification to form clusters, have made important contributions to the goal of increasing nucleic acid sequencing throughput. These cluster-based methods rely on sequencing dense populations of nucleic acids immobilized on a solid support and typically involve using image analysis software to deconvolute the optical signals generated during the simultaneous sequencing of multiple clusters located at different locations on the solid support.
[0086] The present invention relates to the sequencing technology of solid-phase nucleic acid cluster.However, this type of sequencing technology based on solid-phase nucleic acid cluster is still faced with sizable obstacles, and these obstacles limit achievable flux.For example, in the sequencing method based on cluster, determining physically too close to each other and unable to spatially resolve or in fact may have an obstacle aspect the nucleic acid sequence of two or more clusters that physically overlap on solid support.For example, current image analysis software may need valuable time and computing resources to determine which cluster in two overlapping clusters has sent light signal.Therefore, for multiple detection platforms, the compromise about the quantity and / or quality of obtainable nucleic acid sequence information is inevitable.
[0087] Genomic approaches based on high-density nucleic acid clusters have also expanded into other areas of genomic analysis. For example, nucleic acid cluster-based genomics can be used for sequencing applications, diagnostics and screening, gene expression analysis, epigenetic analysis, and genetic analysis of polymorphisms. Each of these nucleic acid cluster-based genomic technologies is also limited when data generated from closely spaced or spatially overlapping nucleic acid clusters cannot be interpreted.
[0088] Clearly, there remains a need to increase the quality and quantity of nucleic acid sequencing data that can be rapidly and cost-effectively obtained for a wide variety of uses, including genomics (e.g., for genomic characterization of any and all animal, plant, microbial, or other biological species or populations), pharmacogenetics, transcriptomics, diagnostics, prognosis, biomedical risk assessment, clinical and research genetics, personalized medicine, drug efficacy and drug interaction assessment, veterinary medicine, agriculture, evolutionary and biodiversity studies, aquaculture, forestry, oceanography, ecological and environmental management, and other purposes.
[0089] The disclosed technology provides neural network-based methods and systems that address these and similar needs, including increasing throughput levels in high-throughput nucleic acid sequencing technologies, and provide other related advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] This patent or patent application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings are also available in pairs via the Supplemental Content tab.
[0091] In the drawings, like reference characters generally refer to like parts throughout the different views. Additionally, the drawings are not necessarily drawn to scale, emphasis instead being placed on illustrating the principles of the disclosed technology. In the following description, various implementations of the disclosed technology are described with reference to the following drawings, wherein:
[0092] Figure 1 The processing stages used by an RTA base caller for base calling are shown according to one specific implementation.
[0093] Figure 2 One specific implementation of base calling using the disclosed neural network-based base caller is illustrated.
[0094] Figure 3 It is a specific implementation of transforming the location / position information of cluster centers identified from the output of a neural network-based template generator from the sub-pixel domain to the pixel domain.
[0095] Figure 4 is a specific implementation that uses a cycle-specific transformation and an image channel-specific transformation to derive so-called "transformed cluster centers" from reference cluster centers.
[0096] Figure 5 An image patch is illustrated which is part of the input data fed to a neural network based base caller.
[0097] Figure 6Depicted is one specific implementation of determining distance values for a distance channel when a single target cluster is being base called by a neural network-based base caller.
[0098] Figure 7 A specific implementation of encoding the distance value between the calculated pixel and the target cluster pixel by pixel is shown.
[0099] Figure 8a Described is one specific implementation of determining distance values for a distance channel when multiple target clusters are simultaneously base called by a neural network-based base caller.
[0100] Figure 8b It is shown that for each target cluster, some nearest pixels are determined based on the distance from the pixel center to the nearest cluster center.
[0101] Figure 9 A specific implementation of pixel-by-pixel encoding of the calculated minimum distance value between a pixel and the nearest cluster in the cluster is shown.
[0102] Figure 10 One specific implementation of pixel-to-cluster classification / attribution / categorization using what is referred to herein as "cluster shape data" is illustrated.
[0103] Figure 11 A specific implementation of calculating distance values using cluster shape data is shown.
[0104] Figure 12 One specific implementation of pixel-by-pixel encoding of the calculated distance values between pixels and assigned clusters is shown.
[0105] Figure 13 One specific implementation of a specialized architecture for a neural network-based base caller that isolates the processing of data from different sequencing cycles is illustrated.
[0106] Figure 14 Depicts a specific implementation of isolated convolution.
[0107] Figure 15a Describes a specific implementation of combined convolution.
[0108] Figure 15b Depicts another implementation of combined convolution.
[0109] Figure 16 One specific implementation of the convolutional layers of a neural network-based base caller is shown, where each convolutional layer has a set of convolutional filters.
[0110] Figure 17 Depicted are two configurations of scaling channels that complement the image channels.
[0111] Figure 18a One specific implementation of input data for a single sequencing cycle that produces a red image and a green image is illustrated.
[0112] Figure 18b We illustrate one implementation of a range channel that provides an additive bias that is incorporated into the feature maps generated from the image channels.
[0113] Figure 19a 、 Figure 19b and Figure 19c Depicts a specific implementation of base calling for a single target cluster.
[0114] Figure 20 One specific implementation of simultaneous base calling for multiple target clusters is shown.
[0115] Figure 21 One specific implementation is shown of simultaneously performing base calling on multiple target clusters at multiple subsequent sequencing cycles, thereby simultaneously generating a base called sequence for each of the multiple target clusters.
[0116] Figure 22 A dimensionality diagram is illustrated for a single cluster base calling implementation.
[0117] Figure 23 A dimensional diagram illustrating a multiple cluster, single sequencing cycle base calling implementation.
[0118] Figure 24 A dimensional diagram illustrating a specific implementation of base calling for multiple clusters and multiple sequencing cycles is shown.
[0119] Figure 25a Depicted is an exemplary array input configuration for multi-cycle input data.
[0120] Figure 25b An exemplary stacked input configuration for multi-cycle input data is shown.
[0121] Figure 26a One implementation is depicted in which pixels of an image patch are reconstructed so that the center of a base-called target cluster is centered in a center pixel.
[0122] Figure 26b Another exemplary reconstructed / shifted image patch is depicted, where (i) the center of the central pixel coincides with the center of the target cluster, and (ii) the non-central pixels are equidistant from the center of the target cluster.
[0123] Figure 27 One specific implementation is shown for base calling a single target cluster at the current sequencing cycle using a standard convolutional neural network and reconstructed input.
[0124] Figure 28 One specific implementation of base calling for multiple target clusters at the current sequencing cycle using a standard convolutional neural network and aligned inputs is shown.
[0125] Figure 29 One specific implementation of base calling for multiple target clusters at multiple sequencing cycles using a standard convolutional neural network and aligned inputs is shown.
[0126] Figure 30 One specific implementation of training a neural network-based base caller is shown.
[0127] Figure 31a Depicted is a specific implementation of a hybrid neural network for use as a neural network-based base caller.
[0128] Figure 31b Shown is a specific implementation of 3D convolution used by the recurrent module of the hybrid neural network to produce the current hidden state representation.
[0129] Figure 32 One specific implementation of processing per-cycle input data for a single sequencing cycle in the series of t sequencing cycles to be base called by a cascade of convolutional layers of a convolutional module is shown.
[0130] Figure 33 Depicted is one specific implementation of mixing the input data for each cycle of a single sequencing cycle with the corresponding convolutional representation of the single sequencing cycle produced by the cascade of convolutional layers of a convolutional module.
[0131] Figure 34 One implementation of arranging the flattened mixed representations of subsequent sequencing cycles into a stack is shown.
[0132] Figure 35a Shows that Figure 34 The stack of undergoes recursive application of 3D convolution in the forward and backward directions and produces one specific implementation of a base call for each of the clusters at each sequencing cycle of the series of t sequencing cycles.
[0133] Figure 35b A specific implementation of processing a 3D input volume x(t) comprising a flattened set of hybrid representations is shown by applying the input gate, activation gate, forget gate, and output gate of a long short-term memory (LSTM) network using 3D convolution. The LSTM network is part of the recurrent module of the hybrid neural network.
[0134] Figure 36One specific implementation of balancing trinucleotides (3-mers) in training data for training a neural network-based base caller is shown.
[0135] Figure 37 The base calling accuracy of the RTA base caller was compared with the base calling accuracy of the neural network-based base caller.
[0136] Figure 38 The block-to-block generalization of the RTA base caller is compared with the block-to-block generalization of the neural network-based base caller on the same block.
[0137] Figure 39 The block-to-block generalization of the RTA base caller is compared with the block-to-block generalization of the neural network-based base caller on the same block and on different blocks.
[0138] Figure 40 The block-to-block generalization of the RTA base caller is also compared with that of a neural network-based base caller on different blocks.
[0139] Figure 41 Shown is how different sizes of image patches fed as input to a neural network-based base caller affect base calling accuracy.
[0140] Figure 42 、 Figure 43 、 Figure 44 and Figure 45 The neural network-based base caller is shown to generalize slot-to-slot on training data from Acinetobacter baumannii and Escherichia coli.
[0141] Figure 46 Shown above relative to Figure 42 、 Figure 43 、 Figure 44 and Figure 45 The error distribution of the slot-to-slot generalization.
[0142] Figure 47 will be Figure 46 Error distribution of the detected error sources attributed to low cluster intensities in the green channel.
[0143] Figure 48 The error distributions of the RTA base caller and the neural network-based base caller for two sequencing runs (Read 1 and Read 2) were compared.
[0144] Figure 49a Shown are the run-to-run generalization of the neural network-based base caller on four different instruments.
[0145] Figure 49b Shown is the run-to-run generalization of the neural network-based base caller across four different runs performed on the same instrument.
[0146] Figure 50 Shown are the genomic statistics of the training data used to train the neural network-based base caller.
[0147] Figure 51 Shown is the genomic context of the training data used to train a neural network-based base caller.
[0148] Figure 52 The base calling accuracy of the neural network-based base caller in base-called long reads (e.g., 2×250) is shown.
[0149] Figure 53 A specific implementation of how a neural network-based base caller focuses on a central cluster of pixels and their neighboring pixels on an image patch is shown.
[0150] Figure 54 Various hardware components and configurations for training and running a neural network-based base caller according to one implementation are shown. In other implementations, different hardware components and configurations are used.
[0151] Figure 55 Various sequencing tasks that can be performed using neural network-based base callers are shown.
[0152] Figure 56 It is a scatter plot visualized by t-distributed stochastic neighbor embedding (t-SNE) and depicts the base calling results of the neural network-based base caller.
[0153] Figure 57 One specific implementation of selecting base call confidence probabilities produced by a neural network-based base caller for use in quality scoring is illustrated.
[0154] Figure 58 A specific implementation of neural network-based quality scoring is shown.
[0155] Figures 59a to 59b Depicts one specific implementation of the correspondence between quality scores and base call confidence predictions made by a neural network-based base caller.
[0156] Figure 60 One specific implementation of inferring quality scores from base call confidence predictions made by a neural network-based base caller during inference is shown.
[0157] Figure 61One specific implementation of training a neural network-based quality scorer to process input data derived from sequencing images and directly generate quality indicators is shown.
[0158] Figure 62 One specific implementation is shown where quality indicators are generated directly during inference as output of a neural network-based quality scorer.
[0159] Figure 63A and Figure 63B A specific implementation of a sequencing system is described. The sequencing system includes a configurable processor.
[0160] Figure 63C is a simplified block diagram of a system for analyzing sensor data (such as base calling sensor output) from a sequencing system.
[0161] Figure 64A is a simplified diagram illustrating aspects of base calling operations, including the functionality of a runtime program executed by a host processor.
[0162] Figure 64B Is a configurable processor such as Figure 63C A simplified diagram depicting the configuration of a configurable processor.
[0163] Figure 65 Can be Figure 63A The sequencing system is a computer system used to implement the technology disclosed in this article.
[0164] Figure 66 Different implementations of data preprocessing are shown, which may include data normalization and data augmentation.
[0165] Figure 67 shows that when a neural network-based base caller is trained on bacterial data and tested on human data, Figure 66 Our data normalization technique (DeepRTA(Normalization)) and data augmentation technique (DeepRTA(Enhancement)) reduced the base calling error percentage, where bacterial data and human data shared the same assays (e.g., both contained intron data).
[0166] Figure 68 It is shown that when a neural network-based base caller is trained on non-exonic data (e.g., intronic data) and tested on exonic data, Figure 66 The data normalization technology (DeepRTA(Normalization)) and data enhancement technology (DeepRTA(Enhancement)) of the proposed method reduce the percentage of base calling errors. DETAILED DESCRIPTION
[0167] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed implementations will be apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Therefore, the disclosed technology is not intended to be limited to the specific implementations shown but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0168] Introduction
[0169] When classifying bases in a sequence of digital images, the neural network processes multiple image channels in the current cycle, as well as image channels in past and future cycles. Within a cluster, some chains may extend before or after the main process of synthesis, where out-of-phase markers are called prephasing or phasing. Given the empirically observed low rates of prephasing and postphasing, almost all noise in the signal generated by prephasing and postphasing can be processed by the neural network, which processes digital images in the current, past, and future cycles (only three cycles).
[0170] In the digital image pipeline within the current cycle, careful registration to align images within the cycle significantly contributes to accurate base classification. Combinations of wavelength and misaligned illumination sources, as well as other sources of error, produce small, correctable differences in the measured cluster center positions. A general affine transformation with translation, rotation, and scaling capabilities can be applied to precisely align cluster centers across the entire image patch. Affine transformations can be used to reconstruct image data and account for shifts in cluster centers.
[0171] Reconstructing image data means interpolating the image data, typically by applying an affine transformation. Reconstruction can place the center of the cluster of interest in the middle of the center pixel of a pixel patch. Alternatively, the reconstruction can align the image with a template to overcome jitter and other differences during image collection. Reconstruction involves adjusting the intensity values of all pixels in the pixel patch. Bilinear and bicubic interpolation methods and weighted area adjustment are alternative strategies.
[0172] In some implementations, cluster center coordinates can be fed to the neural network as an additional image channel.
[0173] Distance signals can also contribute to base classification. Several types of distance signals reflect the separation of regions from cluster centers. The strongest light signal is considered to coincide with the cluster center. Light signals along the periphery of a cluster sometimes include stray signals from nearby clusters. It has been observed that classification is more accurate when the contribution of the signal component decays according to its separation from the cluster center. The distance signals that play a role include a single cluster distance channel, multiple cluster distance channels, and multiple cluster shape-based distance channels. A single cluster distance channel is applied to a patch with a cluster center in the center pixel. Therefore, the distance of all regions in the patch is the distance to the cluster center in the center pixel. Pixels that do not belong to the same cluster as the center pixel can be marked as background rather than given a calculated distance. Multiple cluster distance channels precalculate the distance from each region to the nearest cluster center. This has the possibility of connecting a region to the wrong cluster center, but this possibility is low. Multiple cluster shape-based distance channels associate regions (sub-pixels or pixels) to the pixel center that produces the same base classification through adjacent regions. Through calculation, the possibility of measuring the distance to the wrong pixel is avoided. Multiple clusters and multiple cluster shape-based distance signal methods have the advantage of being subject to pre-computation and working with multiple clusters in an image.
[0174] Neural networks can use shape information to separate signal from noise to improve the signal-to-noise ratio. In the above discussion, several methods for region classification and providing distance channel information were identified. In any of these methods, regions can be marked as background and not as part of a cluster to define cluster edges. Neural networks can be trained to utilize the resulting information about irregular cluster shapes. Distance information and background classification can be used in combination or separately. As cluster density increases, separating signals from adjacent clusters becomes increasingly important.
[0175] One approach to increasing parallel processing scale is to increase cluster density on the imaging medium. Increasing density has the disadvantage of increasing background noise when reading clusters with adjacent neighbors. For example, using shape data rather than arbitrary patches (e.g., 3×3 pixel patches) can help maintain signal separation as cluster density increases.
[0176] In one aspect of the disclosed technology, base classification scores can also be used to predict quality. The disclosed technology includes correlating classification scores directly or through a prediction model with traditional Sanger or Phred quality Q scores. Scores such as Q20, Q30, or Q40 are calculated by Q=-10log 10P is logarithmically related to the probability of base sorting error. The correlation between class score and Q score can be performed using a multi-output neural network or multivariate regression analysis. During base sorting, the advantage of real-time calculation of quality score is that defective sequencing can be terminated early. Applicants have found that it is possible to occasionally (rarely) decide to terminate the run at one-eighth to one-quarter of the way through the entire analysis sequence. The decision to terminate can be made after 50 cycles or after 25 to 75 cycles. In the order-checking process that would originally have run 300 to 1000 cycles, early termination results in significant resource savings.
[0177] Specialized convolutional neural network (CNN) architectures can be used to classify bases within multiple cycles. One specialization involves isolation between digital image channels during processing of the initial layers. The convolution filter stack can be constructed to isolate processing between cycles, thereby preventing crosstalk between digital image sets from different cycles. The motivation for isolating processing between cycles is that images acquired at different cycles have residual registration errors and are therefore misaligned with respect to each other and have random translational offsets. This occurs because the motion accuracy of the sensor's motion stage is limited, and also because images acquired in different frequency channels have different optical paths and wavelengths.
[0178] The motivation for using image sets from subsequent cycles is that the contributions to pre-phasing and post-phasing the signal in a particular cycle are second-order contributions. Therefore, it can help the convolutional neural network to structurally isolate the lower-level convolutions of the digital image sets between image collection cycles.
[0179] The convolutional neural network architecture can also be specifically designed to process information about clustering. Templates of cluster centers and / or shapes provide additional information that the convolutional neural network combines with digital image data. Classification of cluster centers and distance data can be applied repeatedly across cycles.
[0180] The convolutional neural network can be configured to classify multiple clusters in an image field. When classifying multiple clusters, the distance channel of a pixel or sub-pixel can more compactly contain distance information relative to the nearest cluster center or the center of an adjacent cluster to which the pixel or sub-pixel belongs. Alternatively, a large distance vector can be provided for each pixel or sub-pixel, or at least for each pixel or sub-pixel that contains a cluster center, which gives complete distance information from the cluster center to all other pixels that are contextual to the given pixel.
[0181] Some combinations of template generation and base calling can use area-weighted variations instead of distance channels. We now discuss how to use the output of the template generator directly instead of the distance channel.
[0182] We discuss three considerations that influence the direct application of a template image to modify pixel values: whether the image set is processed in the pixel domain or the sub-pixel domain; how the area weights are computed in either domain; and, in the sub-pixel domain, the application of the template image as a mask to modify the interpolated intensity values.
[0183] Performing base classification in the pixel domain has the advantage of not requiring an increase in computation due to upsampling (e.g., 16x). In the pixel domain, even the top convolution layer can have sufficient cluster density to justify performing computations that would otherwise be lost, rather than adding logic to cancel unnecessary computations. We begin with an example using template image data directly in the pixel domain without a range channel.
[0184] In some implementations, classification is focused on specific clusters. In these cases, pixels on the periphery of a cluster may have different modified intensity values, depending on which adjacent cluster is the focus of the classification. A template image in the sub-pixel domain may indicate that overlapping pixels contribute to the intensity values of two different clusters. When two or more adjacent or contiguous clusters both overlap a pixel, we refer to an optical pixel as an "overlapping pixel"; both of these two or more adjacent or contiguous clusters contribute to the intensity reading from the optical pixel. Watershed analysis (named for the division of rainwater flows into different water bodies at ridge lines) can be applied to separate even adjacent clusters. When data is received for cluster-by-cluster classification, a template image can be used to modify the intensity data of overlapping pixels along the periphery of the cluster. The overlapping pixels may have different modified intensities, depending on which cluster is the focus of the classification.
[0185] The modified intensity of a pixel can be reduced based on the sub-pixel contribution of the overlapping pixel to the primary cluster (i.e., the cluster to which the pixel belongs or the cluster whose intensity emission is primarily depicted by the pixel) rather than the distant cluster (i.e., the non-primary cluster whose intensity emission is depicted by the pixel). Assume that 5 sub-pixels are part of the primary cluster and 2 sub-pixels are part of the distant cluster. Then, 7 sub-pixels contribute to the intensity of the primary cluster or the distant cluster. During focus on the primary cluster, in one embodiment, the intensity of the overlapping pixel is reduced by 7 / 16 because 7 of the 16 sub-pixels contribute to the intensity of the primary cluster or the distant cluster. In another embodiment, the intensity is reduced by 5 / 16 based on the area of the sub-pixels contributing to the primary cluster divided by the total number of sub-pixels. In a third embodiment, the intensity is reduced by 5 / 7 based on the area of the sub-pixels contributing to the primary cluster divided by the total area of the contributing sub-pixels. When the focus shifts to the distant cluster, the latter two calculations change, resulting in a fraction with a numerator of "2".
[0186] Of course, further reduction in intensity can be applied if the distance channel is considered together with the sub-pixel map of the cluster shape.
[0187] Once the pixel intensities of the clusters that are the focus of classification have been modified using the template image, the modified pixel values are convolved through the layers of a neural network-based classifier to produce modified images. These modified images are used to classify bases in subsequent sequencing cycles.
[0188] Alternatively, classification in the pixel domain can be performed simultaneously for all pixels or all clusters in an image block. In this case, only one modification of the pixel value can be applied to ensure reusability of intermediate calculations. Any of the scores given above can be used to modify the pixel intensity, depending on whether a smaller or larger intensity attenuation is required.
[0189] Once the pixel intensities of an image patch have been modified using the template image, the pixels and surrounding context are convolved through layers of a neural network-based classifier to produce a modified image. Performing convolution on the image patch allows for the reuse of intermediate computations between pixels that already share context. These modified images are then used to classify bases in subsequent sequencing cycles.
[0190] This description can be used in parallel to apply area weights in the sub-pixel domain. Parallel means that the weights can be calculated for individual sub-pixels. The weights can, but do not have to, be the same for different sub-pixel portions of an optical pixel. Repeating the above scenario with a main cluster of 5 and a distant cluster of 2 sub-pixels in overlapping pixels, respectively, the intensity distribution of the sub-pixels belonging to the main cluster can be 7 / 16, 5 / 16, or 5 / 7 of the pixel intensity. Similarly, if the range channel is considered together with the sub-pixel map of the cluster shape, further intensity reduction can be applied.
[0191] Once the pixel intensities of an image patch have been modified using the template image, the subpixels and surrounding context are convolved through layers of a neural network-based classifier to produce a modified image. Performing convolution on the image patch allows for the reuse of intermediate computations between subpixels that share context. These modified images are then used to classify bases in subsequent sequencing cycles.
[0192] Another alternative is to apply a template image as a binary mask in the sub-pixel domain to the image data interpolated into the sub-pixel domain. The template image can be arranged to require background pixels between clusters or to allow sub-pixels from different clusters to be adjacent. The template image can be used as a mask. If an interpolated pixel is classified as background in the template image, the mask determines whether the interpolated pixel retains the value assigned by the interpolation method or receives the background value (e.g., zero).
[0193] Similarly, once the pixel intensities of an image patch have been masked using a template image, the subpixels and surrounding context are convolved through layers of a neural network-based classifier to produce a modified image. Performing convolution on the image patch allows for the reuse of intermediate computations between subpixels that share context. These modified images are then used to classify bases in subsequent sequencing cycles.
[0194] Features of the disclosed techniques can be combined to classify any number of clusters within a shared context, thereby reusing intermediate computations. In one implementation, at optical pixel resolution, approximately ten percent of the pixels hold cluster centers to be classified. In conventional systems, 3×3 optical pixel groups are grouped as potential signal contributors to cluster centers for analysis, given the observation of irregularly shaped clusters. Even with a 3×3 filter away from the top convolutional layer, the cluster density may roll up into pixels at the cluster center based on light signals from substantially more than half of the optical pixels. Only at the supersampled resolution does the cluster center density at the top convolutional layer drop to less than one percent.
[0195] In some implementations, the shared context is substantial. For example, a 15×15 optical pixel context can contribute to accurate base classification. The equivalent context for a 4x upsampling would be 60×60 sub-pixels. This level of context helps the neural network recognize the effects of uneven lighting and background during imaging.
[0196] The disclosed technique uses small filters at lower convolutional layers to combine cluster boundaries in the template input with boundaries detected in the digital image input. The cluster boundaries help the neural network separate the signal from the background conditions and normalize the image processing relative to the background.
[0197] The disclosed technique essentially reuses intermediate computations. Assume that 20 to 25 cluster centers appear within a context area of 15×15 optical pixels. The first layer of convolution is then reused 20 to 25 times in a block-by-block rollup. The reuse factor decreases layer by layer until the penultimate layer, where the reuse factor at optical resolution first drops to less than 1x.
[0198] Block-by-block roll-up training and inference from multiple convolutional layers applies subsequent roll-ups to blocks of pixels or sub-pixel blocks. There is an overlap region around the perimeter of the block where data used during the roll-up of a first data block overlaps with and can be reused for a second rolled-up block. Within the block, in a central region surrounded by this overlap region are pixel values and intermediate computations that can be rolled up and reused. Using the overlap region, convolution results that gradually reduce the size of the context field (e.g., from 15×15 to 13×13 by applying a 3×3 filter) can be written to the same storage block that holds the convolved values, saving memory without compromising reuse of underlying computations within the block. For larger blocks, sharing intermediate computations in the overlap region requires fewer resources. For smaller blocks, it may be possible to compute multiple blocks simultaneously to share intermediate computations in the overlap region.
[0199] Larger filters and dilation reduce the number of convolutional layers after lower convolutional layers react to cluster boundaries in the template and / or digital image data, which speeds up computation without harming classification.
[0200] The input channels for the template data can be selected so that the template structure is consistent with the classification of multiple cluster centers in the digital image field. The two alternatives described above do not meet this consistency criterion: reconstruction in the entire context and distance mapping. Reconstruction places the center of only one cluster at the center of the optical pixel. For classification of multiple clusters, it is better to provide a center offset for pixels classified as holding a cluster center.
[0201] Unless each pixel has its own distance map across the entire context, it is difficult to perform distance mapping across the entire context area (if provided). Simpler distance maps provide useful consistency for classifying multiple clusters from a digital image input patch.
[0202] The neural network can learn from the classification in the template of pixels or sub-pixels at the cluster boundaries, so the distance channel can be replaced by a template that provides binary or ternary classification and is accompanied by a cluster center offset channel. When used, the distance map can give the distance of the pixel from the center of the cluster to which the pixel (or sub-pixel) belongs. Or the distance map can give the distance to the nearest cluster center. The distance map can encode the binary classification with the label value assigned to the background pixels, or it can be a separate channel for pixel classification. Combined with the cluster center offset, the distance map can encode the ternary classification. In some specific implementations, especially in specific implementations where the pixel classification is encoded with one or two bits, it may be desirable to use separate channels for pixel classification and distance, at least during development.
[0203] The disclosed techniques may include reducing computations to save some computational resources in upper layers. A cluster center offset channel or a ternary classification map may be used to identify the centers of pixel convolutions that do not contribute to the final classification of the pixel center. In many hardware / software implementations, performing lookups and skipped convolution rollups during inference may be more efficient than performing even nine multiplications and eight additions to apply a 3×3 filter. In custom hardware that routes computations for parallel execution, each pixel may be classified within the pipeline. The cluster center map may then be used after the final convolution to obtain results for only the pixels coinciding with the cluster center, since only these pixels require final classification. Similarly, in the optical pixel domain, at the currently observed cluster density, rollup computations will be obtained for approximately ten percent of the pixels. In a 4x upsampled domain, on some hardware, more layers may benefit from skipped convolutions since less than one percent of sub-pixel classifications will be obtained in the top layer.
[0204] Neural network-based base calling
[0205] Figure 1 The processing stages used by the RTA base caller for base calling are shown according to one specific implementation. Figure 1 Also shown are the processing stages used by the disclosed neural network-based base caller for base calling according to two specific implementations. Figure 1 As shown, the neural network-based base caller 218 can simplify the base calling process by eliminating many of the processing stages used by the RTA base caller. This simplification improves the accuracy and scale of base calling. In a first embodiment of the neural network-based base caller 218, the neural network-based base caller uses the position / location information of the cluster center identified from the output of the neural network-based template generator 1512 to perform base calling. In a second embodiment, the neural network-based base caller 218 does not use the position / location information of the cluster center to perform base calling. When a patterned flow cell is designed for cluster generation, a second embodiment is used. The patterned flow cell includes a nanowell that is precisely positioned relative to a known reference position and provides a pre-arranged cluster distribution on the patterned flow cell. In other embodiments, the neural network-based base caller 218 performs base calling on the clusters generated on the random flow cell.
[0206] We will now discuss neural network-based base calling, where a neural network is trained to map sequence images to base calls. This discussion will proceed as follows. First, we will describe the inputs to the neural network. Next, we will describe the structure and form of the neural network. Finally, we will describe the outputs of the neural network.
[0207] enter
[0208] Figure 2 One specific implementation of base calling using neural network 206 is illustrated.
[0209] Main input: image channel
[0210] The primary input to the neural network 206 is image data 202. The image data 202 originates from the sequencing image 108 generated by the sequencer 222 during a sequencing run. In one embodiment, the image data 202 comprises n×n image patches extracted from the sequencing image 222, where n is any number in the range of 1 to 10,000. The sequencing run generates m images per sequencing cycle for the corresponding m image channels, and an image patch is extracted from each of the m images to prepare image data for the particular sequencing cycle. In different embodiments, such as four-channel chemistry, two-channel chemistry, and single-channel chemistry, m is 4 or 2. In other embodiments, m is 1, 3, or greater than 4. In some embodiments, the image data 202 is in the optical pixel domain, and in other embodiments, in the upsampled sub-pixel domain.
[0211] Image data 202 includes data for multiple sequencing cycles (e.g., a current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles). In one implementation, image data 202 includes data for three sequencing cycles, such that the data for the pending base call for the current (time t) sequencing cycle is accompanied by (i) data for the left flank / context / previous / previous / previous (time t-1) sequencing cycle and (ii) data for the right flank / context / next / subsequent / after (time t+1) sequencing cycle. In other implementations, image data 202 includes data for a single sequencing cycle.
[0212] Image data 202 depicts the intensity emission of one or more clusters and their surrounding background. In one embodiment, when base calling is to be performed on a single target cluster, image patches are extracted from sequencing image 108 in such a way that each image patch contains the center of the target cluster in its center pixel, a concept referred to herein as "target cluster-centered patch extraction."
[0213] The image data 202 is encoded in the input data 204 using intensity channels (also referred to as image channels). For each of the m images obtained from the sequencer 222 for a particular sequencing cycle, its intensity data is encoded using a separate image channel. For example, consider a sequencing run using a dual-channel chemistry that produces a red image and a green image in each sequencing cycle. Then, the input data 204 includes (i) a first red image channel having n×n pixels, which depicts the intensity emission of the one or more clusters and their surrounding background captured in the red image; and (ii) a second green imaging channel having n×n pixels, which depicts the intensity emission of the one or more clusters and their surrounding background captured in the green image.
[0214] In one embodiment, the biosensor includes an array of photosensors. The photosensors are configured to sense information from corresponding pixel areas (e.g., reaction sites / wells / nanopores) on the detection surface of the biosensor. An analyte disposed in a pixel area is said to be associated with the pixel area, i.e., an associated analyte. In a sequencing cycle, the photosensors corresponding to the pixel areas are configured to detect / capture / sense the emission / photons from the associated analytes, and in response, generate pixel signals for each imaging channel. In one embodiment, each imaging channel corresponds to a filter wavelength band in a plurality of filter wavelength bands. In another embodiment, each imaging channel corresponds to an imaging event in a plurality of imaging events in a sequencing cycle. In yet another embodiment, each imaging channel corresponds to a combination of illumination with a specific laser and imaging through a specific optical filter.
[0215] The pixel signals from the photosensors are transmitted to a signal processor coupled to the biosensor (e.g., via a communication port). For each sequencing cycle and each imaging channel, the signal processor generates an image whose pixels each depict / contain / indicate / represent / characterize the pixel signals obtained from the corresponding photosensors. Thus, a pixel in the image corresponds to: (i) a photosensor of the biosensor that generates the pixel signal depicted by the pixel, (ii) an associated analyte whose emissions are detected by the corresponding photosensor and converted into a pixel signal, and (iii) a pixel region on the detection surface of the biosensor that holds the associated analyte.
[0216] For example, consider a sequencing run using two different imaging channels (i.e., a red channel and a green channel). Then, in each sequencing cycle, the signal processor generates a red image and a green image. Thus, for a series of k sequencing cycles in a sequencing run, a sequence with k pairs of red and green images is generated as output.
[0217] Pixels in the red and green images (i.e., different imaging channels) correspond one-to-one within a sequencing cycle. This means that corresponding pixels in a pair of red and green images depict intensity data for the same associated analyte, despite being in different imaging channels. Similarly, pixels on a pair of red and green images correspond one-to-one between sequencing cycles. This means that corresponding pixels in different pairs of red and green images depict intensity data for the same associated analyte, despite being acquired for different acquisition events / time steps (sequencing cycles) of a sequencing run.
[0218] Corresponding pixels in the red and green images (i.e., different imaging channels) can be considered as pixels of a "per-cycle image" expressing intensity data in the first red channel and the second green channel. A per-cycle image whose pixels depict the pixel signals of a subset of a pixel region (i.e., an area (patches) of the detection surface of the biosensor) is referred to as a "per-cycle patch image." Patches extracted from the per-cycle patch image are referred to as "per-cycle image patches." In one embodiment, patch extraction is performed by an input preparer.
[0219] The image data includes a sequence of image patches for each cycle generated for a series of k sequencing cycles of a sequencing run. The pixels in the image patch for each cycle contain intensity data for the associated analyte, and intensity data for one or more imaging channels (e.g., a red channel and a green channel) are obtained by corresponding light sensors that are configured to detect emission from the associated analyte. In one embodiment, when base calling is to be performed on a single target cluster, the image patch for each cycle is centered on a center pixel containing intensity data for the target associated analyte, and the non-center pixels in the image patch for each cycle contain intensity data for associated analytes adjacent to the target associated analyte. In one embodiment, the image data is prepared by an input preparer.
[0220] Non-image data
[0221] In another embodiment, the input data for the neural network-based base caller 218 and the neural network-based quality scorer 6102 is based on the pH change caused by the release of hydrogen ions during molecular extension. The pH change is detected and converted into a voltage change that is proportional to the number of bases incorporated (e.g., in the case of Ion Torrent).
[0222] In another specific embodiment, the input data for the neural network-based base detector 218 and the neural network-based quality scorer 6102 is created based on nanopore sensing, which uses a biosensor to measure the interruption of the current when the analyte passes through the nanopore or near its pore opening, while determining the type of base. For example, Oxford Nanopore Technology (ONT) sequencing is based on the following concept: single-stranded DNA (or RNA) is passed through a membrane via a nanopore and a voltage difference is applied across the membrane. The nucleotides present in the pore will affect the resistance of the pore, so the current measurement results over time can indicate the sequence of DNA bases passing through the pore. This current signal (called a "squiggle" due to its appearance when plotted) is the raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values acquired at a frequency of (for example) 4kHz. In the case of a DNA chain speed of approximately 450 base pairs / second, this gives an average of approximately nine raw observations per base. The signal is then processed to identify the interruption of the pore opening signal corresponding to each reading. The most efficient use of these raw signals is to perform base calling, the process of converting DAC values into DNA base sequences.In some implementations, the input data includes normalized or scaled DAC values.
[0223] Supplementary input: distance channel
[0224] The image data 202 is accompanied by supplementary distance data (also referred to as a distance channel). The distance channel provides an additive bias that is incorporated into the feature map generated from the image channel. This additive bias contributes to the accuracy of base calling because it is based on the distance from the pixel center to the cluster center, which is encoded pixel by pixel in the distance channel.
[0225] In a "single target cluster" base calling implementation, for each image channel (image patch) in input data 204, a supplementary distance channel identifies the distance of the center of its pixels from the center of the target cluster containing its center pixel and for which base calling is to be performed. The distance channel thus indicates the respective distances of the pixels of the image patch from the center pixel of the image patch.
[0226] In a "multi-cluster" base calling implementation, for each image channel (image patch) in the input data 204, a supplemental distance channel identifies the center-to-center distance of each pixel from the nearest one of the clusters, which is selected based on the center-to-center distance between the pixel and each of the clusters.
[0227] In a "multi-cluster shape-based" base calling implementation, for each image channel (image patch) in the input data 204, a supplemental distance channel identifies the center-to-center distance of each cluster pixel from an assigned cluster, which is selected based on classifying each cluster pixel into only one cluster.
[0228] Supplementary input: scaling channel
[0229] The image data 202 is accompanied by supplemental scaling data (also referred to as a scaling channel) that accounts for varying cluster sizes and non-uniform lighting conditions. The scaling channel also provides an additive bias that is incorporated into the feature map generated from the image channel. This additive bias contributes to the accuracy of base calling because it is based on the average intensity of the center cluster pixels encoded pixel by pixel in the scaling channel.
[0230] Supplementary input: cluster center coordinates
[0231] In some implementations, location / position information 216 (eg, xy coordinates) of cluster centers identified from the output of the neural network-based template generator 1512 is fed to the neural network 206 as a supplemental input.
[0232] Supplementary input: cluster affiliation information
[0233] In some implementations, the neural network 206 receives as a supplemental input cluster attribute information that classifies pixels or subpixels as background pixels or subpixels, cluster center pixels or subpixels, and cluster / cluster interior pixels or subpixels that depict / contribute to / belong to the same cluster. In other implementations, an attenuation map, a binary map, and / or a ternary map, or variations thereof, are fed as supplemental input to the neural network 206.
[0234] Preprocessing: Intensity Modification
[0235] In some implementations, the input data 204 does not include a range channel, and instead the neural network 206 receives as input modified image data that is modified according to the output (i.e., attenuation map, binary map, and / or ternary map) of the neural network-based template generator 1512. In such implementations, the intensity of the image data 202 is modified to account for the absence of the range channel.
[0236] In other implementations, the image data 202 is subjected to one or more lossless transform operations (e.g., convolution, deconvolution, Fourier transform), and the resulting modified image data is fed as input to the neural network 206.
[0237] Network structure and form
[0238] Neural network 206 is also referred to herein as "neural network-based base caller" 218. In one implementation, neural network-based base caller 218 is a multilayer perceptron. In another implementation, neural network-based base caller 218 is a feedforward neural network. In yet another implementation, neural network-based base caller 218 is a fully connected neural network. In another implementation, neural network-based base caller 218 is a fully convolutional neural network. In yet another implementation, neural network-based base caller 218 is a semantic segmentation neural network.
[0239] In one implementation, the neural network-based base caller 218 is a convolutional neural network (CNN) having multiple convolutional layers. In another implementation, the neural network-based base caller is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another implementation, the neural network-based base caller includes both a CNN and an RNN.
[0240] In other embodiments, the neural network-based base caller 218 may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or atrous convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It may use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross entropy / softmax loss, binary cross entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It may use any parallelism, efficiency, and compression scheme, such as TFRecords, compressed encoding (e.g., PNG), sharpening, parallel calling of map transformations, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous SGD. It may include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (for example, nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential lining unit (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (for example, maximum or average pooling), global average pooling layers and attention mechanisms.
[0241] A neural network-based base caller 218 processes the input data 204 and generates an alternative representation 208 of the input data 204. The alternative representation 208 is a convolutional representation in some implementations and a hidden representation in other implementations. The alternative representation 208 is then processed by an output layer 210 to generate an output 212. The output 212 is used to generate a base call, as described below.
[0242] Output
[0243] In one implementation, the neural network-based base caller 218 outputs base calls for a single target cluster for a particular sequencing cycle. In another implementation, the neural network-based base caller outputs base calls for each target cluster in a plurality of target clusters for a particular sequencing cycle. In yet another implementation, the neural network-based base caller outputs base calls for each target cluster in a plurality of target clusters for each sequencing cycle in a plurality of sequencing cycles, thereby generating a base call sequence for each target cluster.
[0244] Distance channel calculation
[0245] We now discuss how to obtain appropriate position / location information (eg, xy coordinates) of the cluster centers for use in calculating the distance values of the distance channel.
[0246] Reduce coordinates
[0247] Figure 3 It is a specific implementation of transforming the location / position information of the cluster centers identified from the output of the neural network-based template generator 1512 from the sub-pixel domain to the pixel domain.
[0248] The cluster center position / localization information is used for neural network-based base calling to at least: (i) construct input data by extracting an image patch from the sequencing image 108 that contains the center of the target cluster to be base called in its center pixel, (ii) construct a distance channel that identifies the distance of the center of the pixel of the image patch from the center of the target cluster containing its center pixel, and / or (iii) as a supplemental input 216 to the neural network-based base caller 218.
[0249] In some implementations, cluster center position / localization information is identified from the output of the neural network-based template generator 1512 at an upsampled sub-pixel resolution. However, in some implementations, the neural network-based base caller 218 operates on image data at optical pixel resolution. Therefore, in one implementation, the cluster center position / localization information is transformed into the pixel domain by scaling down the coordinates of the cluster centers by the same upsampling factor used to upsample the image data fed as input to the neural network-based template generator 1512.
[0250] For example, consider that the image patch data fed as input to the neural network-based template generator 1512 is obtained by upsampling the image 108 by the upsampling factor f in some initial sequencing cycles. Then, in one embodiment, the coordinates of the cluster centers 302 generated by the neural network-based template generator 1512 and the post-processor 1814 and stored in the template / template image 304 are divided by f (divider). These reduced cluster center coordinates are referred to herein as "reference cluster centers" 308 and are stored in the template / template image 304. In one embodiment, this reduction is performed by the reducer 306.
[0251] Coordinate transformation
[0252] Figure 4 is one specific implementation that uses a cycle-specific transform and an image channel-specific transform to derive the so-called "transformed cluster centers" 404 from the reference cluster centers 308. The motivation for doing so is first discussed.
[0253] Sequencing images acquired at different sequencing cycles are misaligned relative to each other and have random translational offsets. This occurs due to the limited motion accuracy of the sensor's motion stage and also because images acquired in different image / frequency channels have different optical paths and wavelengths. As a result, there is an offset between the reference cluster center and the position / location of the cluster center in the sequencing image. This offset varies between images captured at different sequencing cycles and within images captured at the same sequencing cycle in different image channels.
[0254] To account for this shift, cycle-specific and image channel-specific transformations are applied to the reference cluster centers to produce corresponding transformed cluster centers for the image patches of each sequencing cycle. The cycle-specific and image channel-specific transformations are obtained through an image registration process that uses image correlation to determine a full six-parameter affine transformation (e.g., translation, rotation, scaling, shearing, right reflection, left reflection) or a Procrustes transformation (e.g., translation, rotation, scaling, optionally extended to aspect ratio), further details of which can be found in Appendices 1, 2, 3, and 4.
[0255] For example, consider that the reference cluster centers for the four cluster centers are (x1, y1); (x2, y2); (x3, y3); (x4, y4), and the sequencing run uses a dual-channel chemistry where a red image and a green image are produced at each sequencing cycle. Then, for exemplary sequencing cycle 3, the cycle-specific and image-channel-specific transformations for the red image are The cycle-specific and image channel-specific transforms for the green image are
[0256] Similarly, for exemplary sequencing cycle 9, the cycle-specific and image channel-specific transformations for the red image are The cycle-specific and image channel-specific transforms for the green image are
[0257] Then, the transformed cluster centers of the red image of sequencing cycle 3 is by transforming Transformed cluster centers of the green image obtained by applying the transformation to the reference cluster centers (x1, y1); (x2, y2); (x3, y3); (x4, y4) and sequencing cycle 3 is by transforming Applied to the reference cluster centers (x1,y1); (x2,y2); (x3,y3); (x4,y4).
[0258] Similarly, the transformed cluster centers of the red image of sequencing cycle 9 is by transforming Transformed cluster centers of the green image at sequencing cycle 9, applied to the reference cluster centers (x1, y1); (x2, y2); (x3, y3); (x4, y4) is by transforming Applied to the reference cluster centers (x1,y1); (x2,y2); (x3,y3); (x4,y4).
[0259] In one implementation, these transformations are performed by transformer 402 .
[0260] The transformed cluster centers 404 are stored in the template / template image 304 and are respectively: (i) used for patch extraction from the corresponding sequencing image 108 (e.g., by a patch extractor 406), (ii) used in the distance formula (iii) used to calculate the distance channel for the corresponding image patch, and (iv) as a supplementary input to the neural network-based base caller 218 for the corresponding sequencing cycle that was base called. In other embodiments, different distance formulas can be used, such as distance squared, e^-distance, and e^-distance squared.
[0261] Image patch
[0262] Figure 5 An image patch 502 is shown, which is part of the input data fed to the neural network-based base caller 218. The input data includes a sequence of per-cycle image patch sets generated for a series of sequencing cycles of a sequencing run. Each per-cycle image patch set in the sequence has an image patch for a corresponding image channel in one or more image channels.
[0263] For example, consider that the sequencing run uses a dual-channel chemistry that produces a red image and a green image at each sequencing cycle, and that the input data includes data spanning a series of three sequencing cycles of the sequencing run: the current (time t) sequencing cycle at which a base call is to be made, the previous (time t-1) sequencing cycle, and the subsequent (time t+1) sequencing cycle.
[0264] The input data then includes the following sequence of image patch sets for each cycle: a current cycle image patch set having a current red image patch and a current green image patch extracted from the red sequencing image and the green sequencing image captured at the current sequencing cycle, respectively; a previous cycle image patch set having a previous red image patch extracted from the red sequencing image and the green sequencing image captured at the previous sequencing cycle, and a subsequent red image patch and a subsequent green image patch extracted from the red sequencing image and the green sequencing image captured at the subsequent sequencing cycle, respectively.
[0265] Each image patch may be of size n x n, where n may be any number in the range of 1 to 10,000. Each image patch may be in the optical pixel domain or in the upsampled sub-pixel domain. Figure 5 In the illustrated implementation, the extracted image patch 502 includes pixel intensity data for pixels covering / depicting the plurality of clusters 1-m and their surrounding background. Furthermore, in the illustrated implementation, the image patch 502 is extracted such that the center pixel of the image patch contains the center of the target cluster being base called.
[0266] exist Figure 5 In , pixel centers are depicted by black rectangles and have integer position / location coordinates, and cluster centers are depicted by purple circles and have floating point position / location coordinates.
[0267] Distance calculation for a single target cluster
[0268] Figure 6 One specific implementation of determining distance values 602 for a distance channel when a single target cluster is being base called by the neural network-based base caller 218 is depicted. The center of the target cluster is contained in the center pixel of an image patch fed as input to the neural network-based base caller 218. The distance value is calculated pixel by pixel such that for each pixel the distance between its center and the center of the target cluster is determined. Thus, a distance value is calculated for each pixel in each of the image patches that are part of the input data.
[0269] Figure 6 Three distance values d1, dc, and dn are shown for a particular image patch. In one embodiment, the distance value 602 is calculated using the following distance formula: The distance formula operates based on the transformed cluster centers 404. In other implementations, different distance formulas may be used, such as distance squared, e^-distance, and e^-distance squared.
[0270] In other implementations, when the image patch is at an upsampled sub-pixel resolution, the distance value 602 is calculated in the sub-pixel domain.
[0271] Therefore, in a single target cluster base calling implementation, the range channel is calculated only relative to the target cluster being base called.
[0272] Figure 7 A specific implementation of pixel-by-pixel encoding 702 of the distance value 602 calculated between the pixel and the target cluster is shown. In one specific implementation, the distance value 602 as part of the distance channel in the input data is supplemented as "pixel distance data" for each corresponding image channel (image patch). Returning to the example of the red image and the green image generated by the sequencing cycle, the input data includes a red distance channel and a green distance channel, and the red image channel and the green image channel are supplemented as pixel distance data.
[0273] In other implementations, when the image patch is at an upsampled sub-pixel resolution, the range channel is encoded on a sub-pixel basis.
[0274] Distance calculation for multiple target clusters
[0275] Figure 8aDepicted is one specific implementation of determining distance values 802 for a distance channel when multiple target clusters 1-m are simultaneously base called by the neural network-based base caller 218. The distance values are calculated pixel by pixel, such that for each pixel, the distance between its center and the corresponding center of each of the multiple clusters 1-m is determined, and the minimum distance value (in red) is assigned to the pixel.
[0276] Thus, the distance channel identifies the center-to-center distance of each pixel from the nearest cluster of the clusters, which is selected based on the center-to-center distance between the pixel and each of the clusters. In the illustrated implementation, Figure 8a The pixel center to cluster center distances for two pixels and four cluster centers are shown. Pixel 1 is closest to cluster 1, and pixel n is closest to cluster 3.
[0277] In one specific implementation, the distance value 802 is calculated using the following distance formula: The distance formula operates based on the transformed cluster centers 404. In other implementations, different distance formulas may be used, such as distance squared, e^-distance, and e^-distance squared.
[0278] In other implementations, when the image patch is at an upsampled sub-pixel resolution, the distance value 802 is calculated in the sub-pixel domain.
[0279] Therefore, in multiple cluster base calling implementations, the distance channel is calculated relative to the nearest cluster from the multiple clusters.
[0280] Figure 8b It is shown that for each target cluster 1-m, some nearest pixels are determined based on the distance 804 (dl, d2, d23, d29, d24, d32, dn, dl3, dl4, etc.) of the pixel center to the nearest cluster center.
[0281] Figure 9 One implementation is shown where the calculated minimum distance value between the pixel and the nearest cluster is encoded pixel by pixel 902. In other implementations, when the image patch is at an upsampled sub-pixel resolution, the distance channel is encoded on a sub-pixel basis.
[0282] Cluster shape-based distance calculation for multiple target clusters
[0283] Figure 10 One specific implementation of using pixel-to-cluster classification / attribution / categorization 1002 (referred to herein as "cluster shape data" or "cluster shape information") to determine cluster distance values 1102 for distance channels is illustrated when multiple target clusters 1-m are simultaneously base called by the neural network-based base caller 218. First, the following is a brief overview of how the cluster shape data is generated.
[0284] As described above, the output of the neural network-based template generator 1512 is used to classify pixels into: background pixels, center pixels, and cluster / cluster-internal pixels that depict / contribute to / belong to the same cluster. This pixel-to-cluster classification information is used to attribute each pixel to only one cluster, regardless of the distance between the pixel center and the cluster center, and is stored as cluster shape data.
[0285] exist Figure 10 In the specific implementation shown, background pixels are colored gray, pixels belonging to cluster 1 are colored yellow (Cluster 1 pixels), pixels belonging to cluster 2 are colored green (Cluster 2 pixels), pixels belonging to cluster 3 are colored red (Cluster 3 pixels), and pixels belonging to cluster m are colored blue (Cluster m pixels).
[0286] Figure 11 A specific implementation of using cluster shape data to calculate distance values 1102 is shown. First, the present invention explains why distance information calculated without considering cluster shape is prone to error. Then the present invention explains how cluster shape data overcomes this limitation.
[0287] In a "multiple clusters" base calling implementation that does not use cluster shape data ( Figures 8a to 8b and Figure 9 ), the center-to-center distance value of a pixel is calculated relative to the nearest cluster from multiple clusters. Now, consider the case where a pixel belonging to cluster A is farther away from the center of cluster A but closer to the center of cluster B. In this case, without cluster shape data, the pixel is assigned a distance value calculated relative to cluster B (to which the pixel does not belong) rather than a distance value calculated relative to cluster A (to which the pixel does belong).
[0288] The "multi-cluster shape-based" base calling implementation avoids this by using true pixel-to-cluster mappings (as defined in the original image data and produced by the neural network-based template generator 1512).
[0289] With respect to pixels 34 and 35, a clear difference between the two implementations can be seen. Figure 8b In , the distance values of pixels 34 and 35 are calculated relative to the nearest center of cluster 3, without considering the cluster shape data. However, in Figure 11 , based on the cluster shape data, the distance values 1102 of pixels 34 and 35 are calculated relative to cluster 2 (to which they actually belong).
[0290] exist Figure 11In
[0015] , cluster pixels depict cluster intensity, and background pixels depict background intensity. A cluster distance value identifies the center-to-center distance of each cluster pixel from an assigned cluster in the cluster, where the assigned cluster is selected based on classifying each cluster pixel into only one of the clusters. In some implementations, background pixels are assigned a predetermined background distance value, such as 0 or 0.1 or some other minimum value.
[0291] In one specific implementation, as described above, the cluster distance value 1102 is calculated using the following distance formula: The distance formula operates based on the transformed cluster centers 404. In other implementations, different distance formulas may be used, such as distance squared, e^-distance, and e^-distance squared.
[0292] In other implementations, when the image patch is at an upsampled sub-pixel resolution, cluster distance values 1102 are calculated in the sub-pixel domain, and cluster and background attribution 1002 occurs on a sub-pixel basis.
[0293] Thus, in a multiple cluster shape-based base calling implementation, a distance channel is computed relative to an assigned cluster from the multiple clusters, which is selected based on classifying each cluster pixel into only one of the clusters according to a true pixel-to-cluster mapping defined in the original image data.
[0294] Figure 12 One implementation is shown in which the distance value 1002 calculated between the pixel and the assigned cluster is encoded pixel by pixel. In other implementations, when the image patch is at an upsampled sub-pixel resolution, the distance channel is encoded on a sub-pixel basis.
[0295] Deep learning is a powerful machine learning technique that uses multi-layer neural networks. A particularly successful network architecture in the computer vision and image processing domains is the convolutional neural network (CNN), in which each layer performs a feed-forward convolution transformation from an input tensor (a multidimensional dense array similar to an image) to an output tensor of a different shape. CNNs are particularly well-suited to image-like inputs due to the spatial coherence of images and the advent of general-purpose graphics processing units (GPUs), which can quickly train arrays up to 3D or 4D. Exploiting these image-like properties leads to superior empirical performance compared to other learning methods such as support vector machines (SVMs) or multi-layer perceptrons (MLPs).
[0296] We introduce a specialized architecture that enhances the standard CNN to process image data and supplements it with both distance and scale data. More details are below.
[0297] Specialized architecture
[0298] Figure 13 One specific implementation of a specialized architecture for a neural network-based base caller 218 is shown for isolating the processing of data from different sequencing cycles. The motivation for using a specialized architecture is first described.
[0299] As described above, the neural network-based base caller 218 processes data for the current sequencing cycle, one or more previous sequencing cycles, and one or more subsequent sequencing cycles. The data from the additional sequencing cycles provides sequence-specific context. The neural network-based base caller 218 learns the sequence-specific context during training and performs base calls based on this sequence-specific context. In addition, the data from the previous and next sequencing cycles provide second-order contributions to the prephasing and phasing signals for the current sequencing cycle.
[0300] Spatial convolutional layer
[0301] However, as mentioned above, images captured at different sequencing cycles and in different image channels are misaligned with respect to each other and have residual registration errors. To account for this misalignment, the specialized architecture includes spatial convolutional layers that do not mix information between sequencing cycles and only mix information within a sequencing cycle.
[0302] The spatial convolution layer uses so-called "isolated convolutions," which achieve isolation by independently processing the data for each of multiple sequencing cycles via a sequence of "dedicated, unshared" convolutions. This isolated convolution convolves the data and resulting feature maps only for a given sequencing cycle (i.e., within the cycle), and not for any other sequencing cycles.
[0303] For example, consider that the input data includes (i) current data of the current (time t) sequencing cycle for which base detection is to be performed, (ii) previous data of the previous (time t-1) sequencing cycle, and (iii) subsequent data of the previous (time t+1) sequencing cycle. The specialized architecture then initiates three separate data processing pipelines (or convolution pipelines), namely the current data processing pipeline, the previous data processing pipeline, and the subsequent data processing pipeline. The current data processing pipeline receives the current data of the current (time t) sequencing cycle as input, and independently processes the current data through multiple spatial convolution layers to produce a so-called "current spatial convolution representation" as the output of the final spatial convolution layer. The previous data processing pipeline receives the previous data of the previous (time t-1) sequencing cycle as input, and independently processes the previous data through multiple spatial convolution layers to produce a so-called "previous spatial convolution representation" as the output of the final spatial convolution layer. The subsequent data processing pipeline receives subsequent data of a subsequent (time t+1) sequencing cycle as input, and independently processes the subsequent data through multiple spatial convolutional layers to produce a so-called "subsequent spatial convolutional representation" as the output of the final spatial convolutional layer.
[0304] In some implementations, the current processing pipeline, the previous processing pipeline, and the subsequent processing pipeline are executed synchronously.
[0305] In some implementations, the spatial convolutional layer is part of a spatial convolutional network (or subnetwork) within a specialized architecture.
[0306] Temporal Convolutional Layer
[0307] The neural network-based base caller 218 also includes a temporal convolutional layer that mixes information between sequencing cycles (i.e., inter-cycle). The temporal convolutional layer receives its input from the spatial convolutional network and operates on the spatial convolutional representation produced by the final spatial convolutional layer of the corresponding data processing pipeline.
[0308] The freedom of inter-cycle operability of temporal convolutional layers stems from the fact that the misalignment property present in the image data fed as input to the spatial convolutional network is cleaned from the spatial convolutional representation by the cascade of isolated convolutions performed by the sequence of spatial convolutional layers.
[0309] The temporal convolution layer uses so-called "combined convolutions" that convolve the input channels in subsequent inputs group by group on a sliding window basis. In one specific implementation, these subsequent inputs are the subsequent outputs produced by the previous spatial convolution layer or the previous temporal convolution layer.
[0310] In some implementations, the temporal convolutional layer is part of a temporal convolutional network (or subnetwork) within a specialized architecture. The temporal convolutional network receives its input from the spatial convolutional network. In one implementation, the first temporal convolutional layer of the temporal convolutional network combines the spatial convolutional representations between sequencing cycles on a group-by-group basis. In another implementation, subsequent temporal convolutional layers of the temporal convolutional network combine subsequent outputs of previous temporal convolutional layers.
[0311] The output of the final temporal convolutional layer is fed into an output layer that produces outputs that are used to base call one or more clusters at one or more sequencing cycles.
[0312] Below is a more detailed discussion of isolated and combined convolutions.
[0313] Isolate Convolution
[0314] During the forward pass, the specialized architecture processes information from multiple inputs in two stages. In the first stage, isolating convolutions are used to prevent information from mixing between inputs. In the second stage, combining convolutions are used to mix information between inputs. The results from the second stage are used to perform a single inference on the multiple inputs.
[0315] This differs from batch mode techniques where a convolutional layer processes multiple inputs in a batch simultaneously and performs a corresponding inference for each input in the batch. In contrast, a specialized architecture maps the multiple inputs to a single inference. The single inference can include more than one prediction, such as a classification score for each of the four bases (A, C, T, and G).
[0316] In one embodiment, the inputs are time-ordered such that each input is generated at a different time step and has multiple input channels. For example, the multiple inputs may include the following three inputs: a current input generated by a current sequencing cycle at time step (t), a previous input generated by a previous sequencing cycle at time step (t-1), and a subsequent input generated by a subsequent sequencing cycle at time step (t+1). In another embodiment, each input is derived from a current output, a previous output, and a subsequent output generated by one or more previous convolutional layers, respectively, and includes k feature maps.
[0317] In one implementation, each input may include the following five input channels: a red image channel (red), a red range channel (yellow), a green image channel (green), a green range channel (purple), and a scale channel (blue). In another implementation, each input may include k feature maps produced by a previous convolutional layer, and each feature map is considered an input channel.
[0318] Figure 14One implementation of isolated convolution is depicted. Isolated convolution processes multiple inputs by applying a convolution filter to each input once synchronously. With isolated convolution, the convolution filters combine input channels from the same input and do not combine input channels from different inputs. In one implementation, the same convolution filter is applied synchronously to each input. In another implementation, a different convolution filter is applied synchronously to each input. In some implementations, each spatial convolution layer includes a set of k convolution filters, where each convolution filter is applied synchronously to each input.
[0319] Combined convolution
[0320] Combinatorial convolution mixes information between different inputs by grouping their corresponding input channels and applying a convolution filter to each group. The grouping of these corresponding input channels and the application of the convolution filter occur on a sliding window basis. In this context, a window spans two or more subsequent input channels, representing, for example, the output of two subsequent sequencing cycles. Since the window is a sliding window, most of the input channels are used in two or more windows.
[0321] In some implementations, the different inputs originate from a sequence of outputs generated by a previous spatial convolutional layer or a previous temporal convolutional layer. In the output sequence, these different inputs are arranged as subsequent outputs and are therefore considered subsequent inputs by a subsequent temporal convolutional layer. Then, in the subsequent temporal convolutional layer, these combined convolutions apply convolution filters to corresponding groups of input channels in these subsequent inputs.
[0322] In one embodiment, the subsequent inputs are time-ordered such that the current input is generated by the current sequencing cycle at time step (t), the previous input is generated by the previous sequencing cycle at time step (t-1), and the subsequent input is generated by the subsequent sequencing cycle at time step (t+1). In another embodiment, each subsequent input is derived from the current output, the previous output, and the subsequent output produced by one or more previous convolutional layers, respectively, and includes k feature maps.
[0323] In one implementation, each input may include the following five input channels: a red image channel (red), a red range channel (yellow), a green image channel (green), a green range channel (purple), and a scale channel (blue). In another implementation, each input may include k feature maps produced by a previous convolutional layer, and each feature map is considered an input channel.
[0324] The depth B of the convolution filter depends on the number of subsequent inputs whose corresponding input channels are convolved by the convolution filter group by group on a sliding window basis. In other words, the depth B is equal to the number of subsequent inputs and the group size in each sliding window.
[0325] exist Figure 15a In , corresponding input channels from two subsequent inputs are combined in each sliding window, and thus B=2. Figure 15b , the corresponding input channels from three subsequent inputs are combined in each sliding window, and thus B=3.
[0326] In one implementation, the sliding windows share the same convolutional filters. In another implementation, a different convolutional filter is used for each sliding window. In some implementations, each temporal convolutional layer includes a set of k convolutional filters, where each convolutional filter is applied to subsequent inputs on a sliding window basis.
[0327] filter banks
[0328] Figure 16 One specific implementation of the convolutional layers of the neural network based base caller 218 is shown in FIG. 1 , where each convolutional layer has a set of convolutional filters. Figure 16 In Figure 1, five convolutional layers are shown, each of which has a set of 64 convolutional filters. In some implementations, each spatial convolutional layer has a set of k convolutional filters, where k can be any number such as 1, 2, 8, 64, 128, 256, etc. In some implementations, each temporal convolutional layer has a set of k convolutional filters, where k can be any number such as 1, 2, 8, 64, 128, 256, etc.
[0329] Now let's discuss the supplementary scaling channel and how to calculate it.
[0330] Zoom Channel
[0331] Figure 17 Two configurations of the scaling channel that supplements the image channel are described. The scaling channel is encoded pixel by pixel in the input data fed to the base caller 218 based on the neural network. Different cluster sizes and uneven lighting conditions cause a wide range of cluster intensities to be extracted. The additive bias provided by the scaling channel makes the cluster intensities similar between clusters. In other specific implementations, when the image patch is in the sub-pixel resolution of upsampling, the scaling channel is encoded sub-pixel by sub-pixel.
[0332] When base calling a single target cluster, the scaling channels assign the same scaling value to all pixels. When base calling multiple target clusters simultaneously, the scaling channels assign different scaling values to groups of pixels based on the cluster shape data.
[0333] Zoom channel 1710 has the same zoom value (s1) for all pixels. The zoom value (s1) is based on the average intensity of the center pixel containing the center of the target cluster. In one embodiment, the average intensity is calculated by averaging the intensity values of the center pixel observed during two or more previous sequencing cycles, which produced A and T base calls for the target cluster.
[0334] Based on cluster shape data, scaling channel 1708 has different scaling values (s1, s2, s3, sm) for the corresponding pixel groups belonging to corresponding clusters. Each pixel group includes a center cluster pixel, and this center cluster pixel comprises the center of corresponding cluster. The scaling value for specific pixel group is based on the average intensity of its center cluster pixel. In a specific implementation, average intensity is calculated by averaging the intensity values of the center cluster pixels observed during two or more previous sequencing cycles, and these two or more previous sequencing cycles produce A and T base detection of corresponding clusters.
[0335] In some implementations, background pixels are assigned a background scaling value (sb), which can be 0 or 0.1 or some other minimum value.
[0336] In one implementation, the scaled channels 1706 and their scaled values are determined by an intensity scaler 1704. The intensity scaler 1704 uses cluster intensity data 1702 from previous sequencing cycles to calculate an average intensity.
[0337] In other implementations, the supplemental scaling channel may be provided as an input in different ways (such as before or at the last layer of the neural network-based base caller 218, before or at one or more intermediate layers of the neural network-based base caller 218), and provided as a single value, rather than encoding the supplemental scaling channel pixel by pixel to match the image size.
[0338] Now let's discuss the input data fed to the neural network based base caller 218.
[0339] Input data: image channel, distance channel and scale channel
[0340] Figure 18a One specific implementation of input data 1800 for a single sequencing cycle that produces a red image and a green image is shown. The input data 1800 includes the following:
[0341] • Red intensity data 1802 (red) for pixels in the image patch extracted from the red image. The red intensity data 1802 is encoded in the red image channel.
[0342] • Red distance data 1804 (yellow) that complements the red intensity data 1802 pixel by pixel. The red distance data 1804 is encoded in the red distance channel.
[0343] • Green intensity data 1806 (green) for pixels in the image patch extracted from the green image. The green intensity data 1806 is encoded in the green image channel.
[0344] • Green distance data 1808 (purple) that complements the green intensity data 1806 pixel by pixel. The green distance data 1808 is encoded in the green distance channel.
[0345] • Scaled data 1810 (blue) that complements the red intensity data 1802 and the green intensity data 1806 pixel by pixel. The scaled data 1810 is encoded in the scale channel.
[0346] In other specific implementations, the input data may include fewer or greater numbers of image channels and supplemental distance channels. In one example, for a sequencing run using four-channel chemistry, the input data includes four image channels and four supplemental distance channels for each sequencing cycle.
[0347] We now discuss how the distance channel and the scale channel contribute to base calling accuracy.
[0348] Additive bias
[0349] Figure 18b Shown is a specific implementation of a distance channel that provides an additive bias that is incorporated into the feature map generated from the image channels. This additive bias contributes to the accuracy of base calling because it is based on the distance from the pixel center to the cluster center, which is encoded pixel by pixel in the distance channel.
[0350] On average, approximately 3×3 pixels comprise a cluster. The density at the center of a cluster is expected to be higher than at the edges, as clusters grow outward from a substantially central location. Peripheral cluster pixels may contain conflicting signals from nearby clusters. Therefore, the center cluster pixels are considered the region of maximum intensity and serve as a beacon to reliably identify clusters.
[0351] The pixels of the image patch depict the intensity emissions of multiple clusters (e.g., 10 to 200 clusters) and their surrounding background. Additional clusters contain information from a wider radius and contribute to base call prediction by distinguishing the underlying base whose intensity emission is depicted in the image patch. In other words, the intensity emissions from a group of clusters cumulatively create an intensity pattern that can be assigned to a discrete base (A, C, T, or G).
[0352] We observed that explicitly conveying the distance of each pixel from the cluster center in the supplementary distance channels to the convolutional filters resulted in higher base calling accuracy. These distance channels convey to the convolutional filters which pixels contain the cluster center and which pixels are further away from the cluster center. These convolutional filters use this information to assign sequencing signals to their appropriate source clusters by focusing on (a) the center cluster pixels, their neighboring pixels, and the feature maps derived from them, rather than (b) the peripheral cluster pixels, background pixels, and the feature maps derived from them. In one example of this focus, these distance channels provide a positive additive bias incorporated in the feature map generated by (a), but provide a negative additive bias incorporated in the feature map generated by (b).
[0353] These range channels have the same dimensionality as the image channels. This allows the convolutional filters to evaluate the image and range channels separately within the local receptive field and coherently combine these evaluations.
[0354] When base calling a single target cluster, these range channels identify only one center cluster pixel at the center of the image patch. When base calling multiple target clusters simultaneously, these range channels identify multiple center cluster pixels distributed across the image patch.
[0355] The "single cluster" distance channel is applied to an image patch whose center pixel contains the center of a single target cluster to be base called. The single cluster distance channel includes the center-to-center distance from each pixel in the image patch to the single target cluster. In this implementation, the image patch also includes additional clusters adjacent to the single target cluster, but base calling is not performed on these additional clusters.
[0356] The "Multiple Clusters" distance channel is applied to image patches that contain the centers of multiple target clusters to be base called in their corresponding center cluster pixels. This multiple cluster distance channel includes the center-to-center distance from each pixel in the image patch to the nearest cluster in the multiple target clusters. This has the potential to measure the center-to-center distance of the wrong cluster, but this probability is low.
[0357] The "multi-cluster shape-based" distance channel is applied to an image patch that contains the centers of multiple target clusters to be base called in its corresponding center cluster pixel, and for which pixel-to-cluster assignment information is known. The multi-cluster distance channel includes the center-to-center distance of each cluster pixel in the image patch to the cluster to which it belongs or is assigned in the multiple target clusters. Background pixels may be labeled as background rather than given a calculated distance.
[0358] Figure 18b Also shown is a specific implementation of a scale channel that provides an additive bias that is incorporated into the feature map generated from the image channel. This additive bias contributes to base calling accuracy because it is based on the average intensity of the center cluster pixels encoded pixel by pixel in the scale channel. The discussion of additive bias in the context of the range channel is similarly applicable to the scale channel.
[0359] Example of Additive Bias
[0360] Figure 18b An example is also shown of how additive biases are derived from the distance and scale channels and incorporated into the feature maps generated from the image channel.
[0361] exist Figure 18b In FIG, convolution filter i 1814 evaluates a local receptive field 1812 (magenta) across two image channels 1802 and 1806, two range channels 1804 and 1808, and a scale channel 1810. Because these range and scale channels are encoded separately, additive biasing occurs when the intermediate outputs 1816a-1816e of each of the channel-specific convolution kernels (or feature detectors) 1816a-1816e (plus bias 1816f) are accumulated channel-by-channel as the final output / feature map element 1820 of the local receptive field 1812. In this example, the additive biases provided by the two range channels 1804 and 1808 are intermediate outputs 1816b and 1816d, respectively. The additive bias provided by the scale channel 1810 is intermediate output 1816e.
[0362] This additive bias guides the feature map compilation process by placing greater emphasis on those features in the image channels that are considered more important and reliable for base calling (i.e., the pixel intensities of the center cluster pixel and its neighbors). During training, the weights of the convolutional kernels are updated by backpropagation of the gradients calculated from comparisons with the ground truth base calls to produce stronger activations for the center cluster pixel and its neighbors.
[0363] For example, considering that a pixel in a set of neighboring pixels covered by the local receptive field 1812 contains the cluster center, the distance channels 1804 and 1808 reflect the proximity of the pixel to the cluster center. Therefore, when the intensity intermediate outputs 1816a and 1816c are combined with the distance channel additive biases 1816b and 1816d at the channel-by-channel accumulation 1818, the result is a positive biased convolutional representation 1820 of the pixel.
[0364] In contrast, if the pixels covered by the local receptive field 1812 are not close to the cluster center, the distance channels 1804 and 1808 reflect their separation from the cluster center. Therefore, when the intensity intermediate outputs 1816a and 1816c are combined with the distance channel additive biases 1816b and 1816d at the channel-by-channel accumulation 1818, the result is a negatively biased convolutional representation 1820 of the pixel.
[0365] Similarly, the scaled channel additive bias 1816e derived from the scaled channel 1810 may positively or negatively bias the convolved representation 1820 of the pixel.
[0366] For clarity, Figure 18b A single convolution filter i 1814 is shown applied to input data 1800 for a single sequencing cycle. Those skilled in the art will appreciate that this discussion can be extended to multiple convolution filters (e.g., a filter bank with k filters, where k can be 8, 16, 32, 64, 128, 256, etc.), multiple convolution layers (e.g., multiple spatial and temporal convolution layers), and multiple sequencing cycles (e.g., t, t+1, t-1).
[0367] In other embodiments, these distance channels and scale channels are not encoded separately, but are applied directly to the image channels to generate modulated pixel multiplications because the distance channels and scale channels have the same dimensions as the image channels. In another embodiment, the weights of the convolution kernel are determined based on the distance channels and the image channels so as to detect the most important features in the image channels during the element-wise multiplication. In other embodiments, these distance channels and scale channels are not fed to the first layer, but are provided as auxiliary inputs to downstream layers and / or networks (e.g., to a fully connected network or classification layer). In yet another embodiment, these distance and scale channels are fed to the first layer and re-fed to the downstream layers and / or networks (e.g., via residual connections).
[0368] The above discussion is for 2D input data with k input channels. Those skilled in the art will appreciate the extensions to 3D input. In short, the volumetric input is a 4D tensor with dimensions k × l × w × h, where l is an additional dimension—the length. Each individual kernel is a 4D tensor, and by sweeping through the 4D tensor, a 3D tensor is obtained (the channel dimension is collapsed because it is not swept).
[0369] In other implementations, when the input data 1800 is at an upsampled sub-pixel resolution, these range channels and scale channels are encoded separately on a sub-pixel basis, and additive biasing occurs at the sub-pixel level.
[0370] Base calling using specialized architecture and input data
[0371] We now discuss how to use specialized architectures and input data for neural network-based base calling.
[0372] Single cluster base calling
[0373] Figure 19a 、 Figure 19b and Figure 19c Depicted is a specific implementation for base calling a single target cluster. The specialized architecture processes input data from three sequencing cycles (i.e., the current (time t) sequencing cycle for which base calling is to be performed, the previous (time t-1) sequencing cycle, and the subsequent (time t+1) sequencing cycle) and generates a base call for the single target cluster at the current (time t) sequencing cycle.
[0374] Figure 19a and Figure 19b The spatial convolutional layers are shown. Figure 19c Temporal convolutional layers are shown, along with some other non-convolutional layers. Figure 19a and Figure 19b In the figure, the vertical dashed lines define the spatial convolutional layer according to the feature map, and the horizontal dashed lines define the three convolutional pipelines corresponding to the three sequencing cycles.
[0375] For each sequencing cycle, the input data consists of a tensor of dimension n×n×m (e.g., Figure 18a ), where n represents the width and height of the square tensor, and m represents the number of input channels, so that the dimensions of the input data for the three loops are n×n×m×t.
[0376] Here, each tensor for each cycle contains the center of a single target cluster in the center pixel of its image channel. Each tensor also shows the intensity emission of the single target cluster, some neighboring clusters, and its surrounding background captured in each of the image channels at a particular sequencing cycle. Figure 19a, two exemplary image channels are depicted, namely a red image channel and a green image channel.
[0377] Each tensor for each loop also includes distance channels that complement the corresponding image channels (e.g., red and green distance channels). These distance channels identify the center-to-center distance of each pixel in the corresponding image channel to a single target cluster. Each tensor for each loop also includes a scaling channel that scales the intensity values in each of the image channels pixel by pixel.
[0378] This specialized architecture has five spatial convolutional layers and two temporal convolutional layers. Each spatial convolutional layer uses a Segregated convolution is applied using a set of k convolutional filters, where j represents the width and height of the square filter, and Denotes its depth. Each temporal convolutional layer applies combined convolution using a set of k convolutional filters of dimension j×j×α, where j denotes the width and height of the square filter and α denotes its depth.
[0379] The specialized architecture has a pre-classification layer (e.g., a flattening layer and a dense layer) and an output layer (e.g., a softmax classification layer). The pre-classification layer prepares the input for the output layer. The output layer generates a base call for a single target cluster at the current (time t) sequencing cycle.
[0380] The decreasing spatial dimension
[0381] Figure 19a 、 Figure 19b and Figure 19c Also shown are the resulting feature maps (convolutional representations or intermediate convolutional representations or convolutional features or activation maps) produced by the convolutional filters. Starting from the tensor for each cycle, the spatial dimensions of these resulting feature maps decrease by a constant step size from one convolutional layer to the next, a concept referred to herein as "continuously decreasing spatial dimensions." Figure 19a 、 Figure 19b and Figure 19c In , an exemplary constant step size of 2 is used for the decreasing spatial dimension.
[0382] The decreasing spatial dimensionality is represented by the following formula: "Current feature map spatial dimension = Previous feature map spatial dimension - Convolution filter spatial dimension + 1." This decreasing spatial dimensionality causes the convolution filter to gradually narrow its focus on the center cluster pixel and its neighboring pixels, generating a feature map with features that capture the local dependencies between the center cluster pixel and its neighboring pixels. This, in turn, facilitates accurate base calling of clusters whose centers are contained within the center cluster pixel.
[0383] The isolated convolutions of these five spatial convolutional layers prevent information mixing between the three sequencing cycles and maintain three separate convolutional pipelines.
[0384] The combined convolution of these two temporal convolutional layers mixes the information between the three sequencing cycles. The first temporal convolutional layer convolves the subsequent and current spatial convolutional representations generated by the final spatial convolutional layer for the subsequent and current sequencing cycles, respectively. This produces a first temporal output. The first temporal convolutional layer also convolves the current and previous spatial convolutional representations generated by the final spatial convolutional layer for the current and previous sequencing cycles, respectively. This produces a second temporal output. The second temporal convolutional layer convolves the first and second temporal outputs and produces a final temporal output.
[0385] In some implementations, the final temporal output is fed to a flattening layer to produce a flattened output. The flattened output is then fed to a dense layer to produce a dense output. The dense output is processed by the output layer to produce a base call for a single target cluster at the current (time t) sequencing cycle.
[0386] In some implementations, the output layer generates the likelihood (classification scores) of bases that are A, C, T, and G incorporated into a single target cluster at the current sequencing cycle, and classifies the bases as A, C, T, or G based on these probabilities (e.g., selecting the base with the greatest likelihood, such as Figure 19a In such implementations, these likelihoods are the exponentially normalized scores produced by the softmax classification layer and sum to 1.
[0387] In some implementations, the output layer exports an output pair for a single target cluster. The output pair identifies the class labels of the bases incorporated into the single target cluster at the current sequencing cycle as A, C, T, or G, and performs base calling for the single target cluster based on these class labels. In one implementation, class label 1,0 identifies an A base; class label 0,1 identifies a C base; class label 1,1 identifies a T base; and class label 0,0 identifies a G base. In another implementation, class label 1,1 identifies an A base; class label 0,1 identifies a C base; class label 0.5,0.5 identifies a T base; and class label 0,0 identifies a G base. In yet another implementation, class label 1,0 identifies an A base; class label 0,1 identifies a C base; class label 0.5,0.5 identifies a T base; and class label 0,0 identifies a G base. In yet another specific implementation, class label 1,2 identifies an A base; class label 0,1 identifies a C base; class label 1,1 identifies a T base; and class label 0,0 identifies a G base.
[0388] In some implementations, the output layer derives a class label for a single target cluster, where the class label identifies whether the base incorporated into the single target cluster at the current sequencing cycle is A, C, T, or G, and base calling is performed for the single target cluster based on the class label. In one implementation, a class label of 0.33 identifies an A base; a class label of 0.66 identifies a C base; a class label of 1 identifies a T base; and a class label of 0 identifies a G base. In another implementation, a class label of 0.50 identifies an A base; a class label of 0.75 identifies a C base; a class label of 1 identifies a T base; and a class label of 0.25 identifies a G base.
[0389] In some implementations, the output layer derives a single output value, compares the single output value to a range of class values corresponding to the bases A, C, T, and G, assigns the single output value to a particular class value range based on the comparison, and base calls the single target cluster based on the assignment. In one implementation, the single output value is derived using a sigmoid function, and the single output value is in the range of 0 to 1. In another implementation, a class value range of 0 to 0.25 represents an A base, a class value range of 0.25 to 0.50 represents a C base, a class value range of 0.50 to 0.75 represents a T base, and a class value range of 0.75 to 1 represents a G base.
[0390] Those skilled in the art will appreciate that in other specific implementations, the specialized architecture can process input data for fewer or more sequencing cycles, and may include fewer or more spatial convolutional layers and temporal convolutional layers. In addition, the dimensions of the input data, the tensors for each cycle in the input data, the convolution filters, the resulting feature maps, and the outputs may be different. In addition, the number of convolutional filters in the convolutional layer may be different. It can use various padding and stride configurations. It can use different classification functions (e.g., sigmoid or regression) and may or may not include a fully connected layer. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or hole convolution, transposed convolution, depth-separable convolution, point-by-point convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It can use one or more loss functions such as logistic regression / logarithmic loss function, multi-class cross entropy / softmax loss function, binary cross entropy loss function, mean squared error loss function, L1 loss function, L2 loss function, smooth L1 loss function and Huber loss function. It can use any parallelism, efficiency and compression scheme, such as TFRecords, compressed encoding (e.g., PNG), sharpening, parallel detection of map transformations, batching, prefetching, model parallelism, data parallelism and synchronous / asynchronous SGD. It can include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential lining units (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (e.g., max or average pooling), global average pooling layers and attention mechanisms.
[0391] Having described single cluster base calling, we now discuss multiple cluster base calling.
[0392] Multiple cluster base calling
[0393] Depending on the size of the input data and the cluster density on the flow cell, anywhere between one hundred thousand and three hundred thousand clusters are base called simultaneously on each input by the neural network-based base caller 218. Extending this to a data parallelism and / or model parallelism strategy implemented on a parallel processor, using batches or mini-batches of size ten results in base calling between one hundred and three million clusters on a per-batch or per-mini-batch basis.
[0394] Depending on the sequencing configuration (e.g., cluster density, number of blocks on the flow cell), a block includes 20,000 to 300,000 clusters. In another embodiment, Illumina's NovaSeq sequencer has up to 4 million clusters per block. Thus, a sequencing image of a block (block image) can depict intensity emissions from 20,000 to 300,000 clusters and their surrounding background. Thus, in one embodiment, using input data that includes the entire block image results in simultaneous base calling of 300,000 clusters per input. In another embodiment, using image patches of size 15×15 pixels in the input data results in simultaneous base calling of less than one hundred clusters per input. Those skilled in the art will appreciate that these numbers can vary depending on the sequencing configuration, parallelization strategy, details of the architecture (e.g., based on optimal architecture hyperparameters), and available compute.
[0395] Figure 20 A specific implementation of simultaneous base calling for multiple target clusters is shown. The input data has three tensors for the three sequencing cycles described above. For each tensor of each cycle (e.g., Figure 18a The input tensor 1800 in FIG. 1 depicts the intensity emission captured in each of the image channels at a specific sequencing cycle for multiple target clusters to be base called and their surrounding background. In other embodiments, some additional neighboring clusters that are not base called are also included for context.
[0396] In a multiple cluster base calling implementation, each tensor for each cycle includes distance channels that complement the corresponding image channels (e.g., a red distance channel and a green distance channel). These distance channels identify the center-to-center distance of each pixel in the corresponding image channel to the nearest cluster in the multiple target clusters.
[0397] In a multiple cluster shape-based base calling implementation, each tensor for each cycle includes distance channels that complement corresponding image channels (e.g., a red distance channel and a green distance channel). These distance channels identify the center-to-center distance of each cluster pixel in the corresponding image channel to the cluster to which it belongs or is attributed in the multiple target clusters.
[0398] Each tensor for each cycle also includes a scaling channel that scales the intensity values in each image channel pixel by pixel.
[0399] exist Figure 20 In the example, the spatial dimension of each tensor for each loop is greater than Figure 19a The spatial dimensions shown. That is, in Figure 19a In the specific implementation of single target cluster base calling in , the spatial dimension of each tensor for each cycle is 15×15, while in Figure 20In a specific implementation, the spatial dimension of each tensor for each cycle is 114 × 114. According to some implementations, having a larger amount of pixelated data depicting the intensity emissions of additional clusters improves the accuracy of predicting base calls for multiple clusters simultaneously.
[0400] Avoid redundant convolution
[0401] In addition, the image channels in each tensor for each cycle are obtained from the image patches from which the sequenced images are extracted. In some implementations, there are overlapping pixels between the extracted image patches that are spatially adjacent (e.g., left, right, top, and bottom adjacent). Therefore, in one implementation, these overlapping pixels are not subjected to redundant convolutions, and the results from the previous convolution are reused later when the overlapping pixels are part of the subsequent input.
[0402] For example, consider extracting a first image patch of size n×n pixels from a sequenced image, and also extracting a second image patch of size m×m pixels from the same sequenced image, such that the first image patch and the second image patch are spatially adjacent and share an overlapping region of o×o pixels. Further consider that o×o pixels are convolved as part of the first image patch to produce a first convolved representation stored in memory. Then, when the second image patch is convolved, the o×o pixels are no longer convolved, but instead the first convolved representation is retrieved from memory and reused. In some implementations, n=m. In other implementations, they are not equal.
[0403] The input data is then processed through the spatial and temporal convolutional layers of the specialized architecture to produce a final temporal output of dimension w × w × k. Here again, in the phenomenon of continuously reducing spatial dimensions, the spatial dimension is reduced to a constant step size of 2 at each convolutional layer. That is, starting from the n × n spatial dimensions of the input data, the w × w spatial dimensions of the final temporal output are derived.
[0404] Then, based on the spatial dimension w×w of the final temporal output, the output layer generates a base call for each cell in the w×w cell set. In one specific implementation, the output layer is a softmax layer that generates a four-way classification score for the four bases (A, C, T, and G) on a cell-by-cell basis. That is, a base call is assigned to each cell in the w×w cell set based on the maximum classification score in the corresponding softmax quadruple, such as Figure 20In some implementations, a w×w cell set is derived as a result of processing the final temporal output through a flatten layer and a dense layer to produce a flattened output and a dense output, respectively. In such implementations, the flattened output has w×w×k elements, and the dense output has w×w elements forming the w×w cell set.
[0405] Base calls for multiple target clusters are obtained by identifying which of the base-called cells in the w×w cell set coincide with or correspond to center cluster pixels (i.e., pixels in the input data containing the respective centers of the multiple target clusters). Base calls for a given target cluster are assigned to cells that coincide with or correspond to the pixel containing the center of the given target cluster. In other words, base calls for cells that do not coincide with or correspond to center cluster pixels are filtered out. This functionality is performed by a base call filtering layer, which in some implementations is part of a specialized architecture or implemented as a post-processing module in other implementations.
[0406] In other specific implementations, base calls for the multiple target clusters are obtained by identifying which groups of base-called cells in a w×w cell set cover the same cluster (i.e., identifying groups of pixels in the input data that depict the same cluster). Then, for each cluster and its corresponding pixel group, the average (softmax probability) of the classification scores for the four corresponding base classes (A, C, T, and G) is calculated across the pixels in the pixel group, and the base class with the highest average classification score is selected for base calling the cluster.
[0407] During training, in some implementations, ground truth comparisons and error calculations are performed only for those cells that coincide with or correspond to the center cluster pixels, so that their predicted base calls are evaluated against the correct base calls identified as ground truth labels.
[0408] Having described multiple cluster base calling, we now discuss multiple cluster and multiple cycle base calling.
[0409] Multiple cluster and multiple cycle base calling
[0410] Figure 21 One specific implementation is shown of simultaneously performing base calling on multiple target clusters at multiple subsequent sequencing cycles, thereby simultaneously generating a base called sequence for each of the multiple target clusters.
[0411] In the single and multiple base calling implementations discussed above, the base call at one sequencing cycle (the current (time t) sequencing cycle) is predicted using data for three sequencing cycles (the current (time t), the previous / left flank (time t-1), and the subsequent / right flank (time t+1) sequencing cycles), where the right and left flank sequencing cycles provide sequence-specific context for the base triplet motif as well as second-order contributions to the prephasing and phasing signals. This relationship is represented by the following formula: "Number of sequencing cycles included in the input data (t) = Number of sequencing cycles for which a base call was made (y) + Number of right and left flank sequencing cycles (x)."
[0412] exist Figure 21 , the input data includes t per-cycle tensors for t sequencing cycles, such that the input data has dimensions n×n×m×t, where n=114, m=5, and t=15. In other implementations, these dimensions are different. In t sequencing cycles, the tth sequencing cycle and the first sequencing cycle are used as the right-wing and left-wing contexts x, and base calls are performed for the y sequencing cycles in between. Therefore, y=13, x=2, and t=y+x. Each tensor for each cycle includes an image channel, a corresponding distance channel, and a scaling channel, such as Figure 18a The input tensor in 1800.
[0413] The input data, with t per-cycle tensors, is then processed through the specialized architecture's spatial and temporal convolutional layers to produce y final temporal outputs, where each final temporal output corresponds to a corresponding sequencing cycle among the y sequencing cycles that were base called. Each of the y final temporal outputs has dimensions w × w × k. Here again, under the phenomenon of ever-decreasing spatial dimensionality, the spatial dimensions are reduced with a constant step size of 2 at each convolutional layer. That is, starting from the n × n spatial dimensions of the input data, each of the y final temporal outputs is derived with dimensions w × w.
[0414] Each of the y final time outputs is then processed synchronously by an output layer. For each of the y final time outputs, the output layer generates a base call for each cell in the w×w cell set. In one embodiment, the output layer is a softmax layer that generates a four-way classification score for the four bases (A, C, T, and G) on a cell-by-cell basis. That is, a base call is assigned to each cell in the w×w cell set based on the maximum classification score in the corresponding softmax quadruple, such as Figure 20As shown. In some implementations, as the y final temporal outputs are processed by the flatten layer and the dense layer to produce corresponding flattened outputs and dense outputs, respectively, a wxw cell set is derived for each of these final temporal outputs. In such implementations, each flattened output has w×w×k elements, and each dense output has w×w elements forming a w×w cell set.
[0415] For each of the y sequencing cycles, base calls for the multiple target clusters are obtained by identifying which of the base-called cells in the corresponding w×w cell set coincide with or correspond to center cluster pixels (i.e., pixels containing the centers of the multiple target clusters in the input data). Base calls are assigned to cells that coincide with or correspond to the pixel containing the center of the given target cluster. In other words, base calls for cells that do not coincide with or correspond to center cluster pixels are filtered out. This functionality is performed by a base call filtering layer, which in some implementations is part of a specialized architecture or implemented as a post-processing module in other implementations.
[0416] During training, in some implementations, ground truth comparisons and error calculations are performed only for those cells that coincide with or correspond to the center cluster pixels, so that their predicted base calls are evaluated against the correct base calls identified as ground truth labels.
[0417] Based on each input, the result is a base call for each target cluster in the plurality of target clusters at each of y sequencing cycles, i.e., a base called sequence of length y for each target cluster in the plurality of target clusters. In other specific implementations, y is 20, 30, 50, 150, 300, etc. Those skilled in the art will appreciate that these numbers can vary depending on the sequencing configuration, parallelization strategy, details of the architecture (e.g., based on optimal architecture hyperparameters), and available compute.
[0418] End-to-end dimension diagram
[0419] The following discussion uses dimensionality diagrams to illustrate the underlying data dimensionality changes involved in generating base calls from image data, and different specific implementations of the dimensionality of the data operators that implement the data dimensionality changes.
[0420] exist Figure 22 、 Figure 23 and Figure 24 In , rectangles represent data operators, such as spatial convolution layers, temporal convolution layers, and softmax classification layers, and rounded rectangles represent data produced by the data operators (e.g., feature maps).
[0421] Figure 22A dimensionality diagram 2200 for a specific implementation of single cluster base calling is shown. Note that the "cycle dimension" of the input is three, and continues to be the "cycle dimension" of the resulting feature map until the first temporal convolutional layer. The cycle dimension of three represents three sequencing cycles, and its continuity means that the feature maps of the three sequencing cycles are generated and convolved separately, and no features are mixed between the three sequencing cycles. The isolated convolution pipeline is implemented by the depthwise isolated convolution filters of the spatial convolutional layers. Note that the "depth dimension" of the depthwise isolated convolution filters of these spatial convolutional layers is one. This enables these depthwise isolated convolution filters to convolve only the data and resulting feature map of a given sequencing cycle (i.e., within the cycle), and prevents them from convolving the data and resulting feature map of any other sequencing cycle.
[0422] In contrast, it is important to note that the depthwise combined convolutional filters of the temporal convolutional layer have a depth dimension of 2. This enables these depthwise combined convolutional filters to convolve the resulting feature maps of multiple sequencing cycles group by group and mix features between sequencing cycles.
[0423] Also notice that the "spatial dimension" keeps decreasing in constant steps of 2.
[0424] In addition, the vector with four elements is exponentially normalized by the softmax layer to generate classification scores (i.e., confidence scores, probabilities, likelihoods, softmax function scores) for the four bases (A, C, T, and G). The base with the highest (maximum) softmax function score is assigned to the single target cluster called by the base at the current sequencing cycle.
[0425] Those skilled in the art will appreciate that in other implementations, the dimensions shown may vary depending on the sequencing configuration, parallelization strategy, details of the architecture (e.g., based on optimal architecture hyperparameters), and available compute.
[0426] Figure 23 A dimensional diagram 2300 for a multiple cluster, single sequencing cycle base calling implementation is shown. The above discussion regarding cycle, depth, and spatial dimensions relative to single cluster base calling applies to this implementation.
[0427] Here, the softmax layer operates independently on each of the 10,000 cells and produces a quadruple of softmax function scores for each of the 10,000 cells. The quadruple corresponds to the four bases (A, C, T, and G). In some implementations, these 10,000 cells are derived from the conversion of 64,0000 flattened cells to 10,000 dense cells.
[0428] Then, according to the softmax function score quadruple of each unit in these 10,000 units, the base with the highest softmax function score in each quadruple is assigned to the corresponding unit in these 10,000 units.
[0429] Then, the 2,500 cells of these 10,000 cells that correspond to the 2,500 center cluster pixels containing the respective centers of the 2,500 target clusters that are being simultaneously base-called at the current sequencing cycle are selected. The bases assigned to these 2,500 cells are then assigned to the corresponding target clusters in the 2,500 target clusters.
[0430] Those skilled in the art will appreciate that in other implementations, the dimensions shown may vary depending on the sequencing configuration, parallelization strategy, details of the architecture (e.g., based on optimal architecture hyperparameters), and available compute.
[0431] Figure 24 A dimensional diagram 2400 for a multiple cluster, multiple sequencing cycle base calling implementation is shown. The above discussion regarding cycle, depth, and spatial dimensions relative to single cluster base calling applies to this implementation.
[0432] Furthermore, the above discussion regarding softmax-based base call classification relative to multiple cluster base calls also applies here. However, here, softmax-based base call classification for the 2,500 target clusters occurs simultaneously for each of the thirteen sequencing cycle bases that have been base called, thereby generating thirteen base calls for each of the 2,500 target clusters simultaneously.
[0433] Those skilled in the art will appreciate that in other implementations, the dimensions shown may vary depending on the sequencing configuration, parallelization strategy, details of the architecture (e.g., based on optimal architecture hyperparameters), and available compute.
[0434] Array input and stack input
[0435] We now discuss two configurations in which multi-cycle input data can be arranged into a neural network-based base caller. The first configuration is referred to as "array input," and the second configuration is referred to as "stacked input." Figure 25a shown in, and in the above with respect to Figures 19a to 24This array input encodes the input for each sequencing cycle in separate columns / blocks, since the image patches in the input for each cycle are misaligned relative to each other due to residual registration errors. A specialized architecture is used with this array input to isolate the processing of each separate column / block in a separate column / block. Additionally, a distance channel is computed using the transformed cluster centers to account for misalignment between image patches within a cycle and across cycles.
[0436] In contrast, Figure 25b The stacked input shown encodes inputs from different sequencing cycles in a single column / block. In one implementation, this eliminates the need to use this specialized architecture because the image patches in the stacked input are aligned to each other through affine transformation and intensity interpolation, which eliminates inter-cycle and intra-cycle residual registration errors. In some implementations, the stacked input has a common scaling channel for all inputs.
[0437] In another embodiment, intensity interpolation is used to reconstruct or shift the image patches so that the center of the center pixel of each image patch coincides with the center of the single target cluster being base called. This eliminates the need to use a supplementary distance channel because all non-center pixels are equidistant from the center of the single target cluster. The stacked input without the distance channel is referred to herein as the "reconstructed input" and is Figure 27 Shown in.
[0438] However, for base calling implementations involving multiple clusters, such reconstruction may not be feasible because there are image patches containing multiple center cluster pixels that have been base called. Stacked inputs that do not have a range channel and for which no reconstruction is performed are referred to herein as "aligned inputs" and are Figure 28 and Figure 29 Aligned inputs can be used when computing the range channel is not desirable (eg, due to computational limitations) and reconstruction is not feasible.
[0439] The following sections discuss various base calling implementations that do not use this specialized architecture and these supplementary distance channels, but instead use standard convolutional layers and filters.
[0440] Reconstructed input: aligned image patches without range channel
[0441] Figure 26a One embodiment of reconstructing 2600a the pixels of image patch 2602 so that the center of the base-called target cluster is centered in the center pixel is depicted. The center of the target cluster (purple) falls within the center pixel of image patch 2602 but is offset from the center of the center pixel (red), as shown in FIG. Figure 26a shown.
[0442] To eliminate the offset, the reconstructor 2604 compensates for the reconstruction by interpolating the intensities of the pixels in the shifted image patch 2602 and produces a reconstructed / shifted image patch 2606. In the shifted image patch 2606, the center of the center pixel coincides with the center of the target cluster. In addition, the non-center pixels are equidistant from the center of the target cluster. Interpolation can be performed by the following methods: nearest neighbor intensity extraction, Gaussian-based intensity extraction, intensity extraction based on the average value of a 2×2 sub-pixel area, intensity extraction based on the brightest point in a 2×2 sub-pixel area, intensity extraction based on the average value of a 3×3 sub-pixel area, bilinear intensity extraction, bicubic intensity extraction, and / or intensity extraction based on weighted area cover. These techniques are described in detail in the Appendix entitled "Intensity Extraction Methods."
[0443] Figure 26b Another exemplary reconstructed / shifted image patch 2600b is depicted in which (i) the center of the center pixel coincides with the center of the target cluster, and (ii) the non-center pixels are equidistant from the center of the target cluster. These two factors eliminate the need to provide a supplementary distance channel because all non-center pixels have the same proximity to the center of the target cluster.
[0444] Figure 27 A specific implementation of base calling a single target cluster at the current sequencing cycle using a standard convolutional neural network and reconstructed input is shown. In the illustrated implementation, the reconstructed input includes: the current image patch set for the current (t) sequencing cycle being base called, the previous image patch set for the previous (t-1) sequencing cycle, and the subsequent image patch set for the subsequent (t+1) sequencing cycle. Each image patch set has an image patch for a corresponding image channel from one or more image channels. Figure 27 Two image channels are depicted: a red channel and a green channel. Each image patch contains pixel intensity data covering the pixels of the base-called target cluster, some neighboring clusters, and the surrounding background. The reconstruction input also includes a common scaling channel.
[0445] Because the image patch is reconstructed or shifted to be centered on the center of the target cluster, the input to this reconstruction does not include any distance channel, as described above with respect to Figures 26a to 26b As described above. In addition, these image patches are aligned to each other to remove inter-cycle and intra-cycle residual registration errors. In one specific implementation, this is done using affine transformations and intensity interpolation, further details of which can be found in Appendices 1, 2, 3, and 4. These factors eliminate the need to use specialized architectures, and instead a standard convolutional neural network can be used with the reconstructed input.
[0446] In the illustrated implementation, the standard convolutional neural network 2700 includes seven standard convolutional layers using standard convolutional filters. This means that there is no isolated convolution pipeline to prevent data mixing between sequencing cycles (because the data is aligned and can be mixed). In some implementations, the phenomenon of decreasing spatial dimensionality is used to teach the standard convolutional filters to pay more attention to the center of the central cluster and its neighboring pixels rather than other pixels.
[0447] The reconstructed input is then processed through these standard convolutional layers to produce the final convolutional representation. Based on this final convolutional representation, the base calls for the target cluster at the current sequencing cycle are obtained in a similar manner using flattening layers, dense layers, and classification layers, as described above with respect to Figure 19c As stated.
[0448] In some implementations, the process is iterated over multiple sequencing cycles to generate base called sequences for the target cluster.
[0449] In other specific implementations, the process is iterated over multiple sequencing cycles for a plurality of target clusters to generate a base called sequence for each of the plurality of target clusters.
[0450] Aligned input: aligned image patches without range channel and without reconstruction
[0451] Figure 28 One implementation shows base calling for multiple target clusters at the current sequencing cycle using a standard convolutional neural network and aligned input. Here, reconstruction is not feasible because the image patch contains multiple center cluster pixels that were base called. Therefore, the image patch in the aligned input is not reconstructed. Furthermore, according to one implementation, a supplementary distance channel is not included due to computational considerations.
[0452] The aligned inputs are then processed through standard convolutional layers to produce a final convolutional representation. Based on this final convolutional representation, a flattening layer (optional), a dense layer (optional), a classification layer, and a base calling filter layer are used in a similar manner to obtain base calls for each target cluster at the current sequencing cycle, as described above with respect to Figure 20 As stated.
[0453] Figure 29 A specific implementation of base calling for multiple target clusters at multiple sequencing cycles using a standard convolutional neural network and aligned inputs is shown. The aligned inputs are processed through standard convolutional layers to produce a final convolutional representation for each of the y sequencing cycles that are base called. Based on these y final convolutional representations, a flattening layer (optional), a dense layer (optional), a classification layer, and a base calling filter layer are used in a similar manner to obtain a base call for each of the target clusters for each of the y sequencing cycles that are base called, as described above with respect to Figure 21 As stated.
[0454] Those skilled in the art will appreciate that in other specific implementations, a standard convolutional neural network can process reconstructed inputs for fewer or more sequencing cycles and can include fewer or more standard convolutional layers. In addition, the dimensions of the reconstructed input, the tensors for each cycle in the reconstructed input, the convolution filters, the resulting feature maps, and the outputs may be different. In addition, the number of convolutional filters in the convolutional layer may be different. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or dilated convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It can use one or more loss functions, such as logistic regression / logarithmic loss function, multi-class cross entropy / softmax loss function, binary cross entropy loss function, mean squared error loss function, L1 loss function, L2 loss function, smoothed L1 loss function, and Huber loss function. It can use any parallelism, efficiency and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel detection of map transformations, batching, prefetching, model parallelism, data parallelism and synchronous / asynchronous SGD. It can include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential lining units (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (e.g., max or average pooling), global average pooling layers and attention mechanisms.
[0455] train
[0456] Figure 30 One specific implementation of training 3000 a neural network-based base caller 218 is shown. Utilizing both specialized and standard architectures, the neural network-based base caller 218 is trained using a backpropagation-based gradient update technique that compares predicted base calls 3004 to correct base calls 3008 and calculates errors 3006 based on the comparisons. Errors 3006 are then used to calculate gradients, which are applied to the weights and parameters of the neural network-based base caller 218 during backpropagation 3010. Training 3000 is performed by a trainer 1510 using a stochastic gradient update algorithm, such as ADAM.
[0457] Trainer 1510 trains neural network-based base caller 218 using training data 3002 (derived from sequenced images 108) over thousands and millions of iterations of forward propagation 3012 to generate predicted base calls 3004 and backward propagation 3010 to update weights and parameters based on errors 3006. Additional details regarding training 3000 can be found in the Appendix entitled "Deep Learning Tools."
[0458] CNN — RNN-based base caller
[0459] Hybrid Neural Network
[0460] Figure 31a One embodiment of a hybrid neural network 3100a for use as a neural network-based base caller 218 is depicted. The hybrid neural network 3100a includes at least one convolutional module 3104 (or convolutional neural network (CNN)) and at least one recurrent module 3108 (or recurrent neural network (RNN)). The recurrent module 3108 uses and / or receives input from the convolutional module 3104.
[0461] The convolution module 3104 processes the input data 3102 through one or more convolutional layers and produces a convolution output 3106. In one embodiment, the input data 3102 includes only image channels or image data as the primary input, as described above in the section entitled "Input." The image data fed to the hybrid neural network 3100a can be the same as the image data 202 described above.
[0462] In another embodiment, in addition to image channels or image data, the input data 3102 also includes supplemental channels, such as a distance channel, a scale channel, cluster center coordinates and / or cluster affiliation information, as described above in the section entitled "Input".
[0463] The image data (i.e., input data 3102) depicts the intensity emission of one or more clusters and their surrounding background. The convolution module 3104 processes the image data of a series of sequencing cycles of the sequencing run through a convolutional layer and produces one or more convolution representations of the image data (i.e., convolution output 3106).
[0464] The series of sequencing cycles may include image data for t sequencing cycles to be base called, where t is any number between 1 and 1000. When t is between 15 and 21, we observe accurate base call results.
[0465] The recurrence module 3110 convolves the convolution output 3106 and generates a recurrence output 3110. Specifically, the recurrence module 3110 generates a current hidden state representation (i.e., the recurrence output 3110) based on convolving the convolution representation and the previous hidden state representation.
[0466] In one embodiment, the recursive module 3110 applies a three-dimensional (3D) convolution to the convolution representation and the previous hidden state representation and produces a current hidden state representation, which is mathematically expressed as:
[0467] h t =W1 3DCONV V t +W2 3DCONV h t-1 ,in
[0468] h t represents the current hidden state representation generated at the current time step t,
[0469] V t represents a set of convolutional representations that form the input volume at the current sliding window at the current time step t,
[0470] W1 3DCONV Indicates that it is applied to V t The weights of the first 3D convolution filter,
[0471] h t-1 represents the previous hidden state representation produced at the previous time step t-1, and
[0472] W2 3DCONV Indicates that it is applied to h t-1 The weights of the second 3D convolution filter.
[0473] In some implementations, since the weights are shared, W1 3DCONV and W2 3DCONV are the same.
[0474] Output module 3112 then generates base calls 3114 based on recursive output 3110. In some implementations, output module 3112 includes one or more fully connected layers and a classification layer (e.g., softmax). In such implementations, the current hidden state representation is processed by the fully connected layers, and the output of the fully connected layers is processed by the classification layer to generate base calls 3114.
[0475] Base calling 3114 includes base calling for at least one of the clusters and for at least one of the sequencing cycles. In some implementations, base calling 3114 includes base calling for each of the clusters and for each of the sequencing cycles. Thus, for example, when input data 3102 includes image data for twenty-five clusters and fifteen sequencing cycles, base calling 3102 includes base call sequences for fifteen base calls for each of the twenty-five clusters.
[0476] 3D convolution
[0477] Figure 31b One specific implementation of a 3D convolution 3100b used by the recurrent module 3110 of the hybrid neural network 3100b to produce a representation of the current hidden state is shown.
[0478] 3D convolution is a mathematical operation where each voxel present in the input volume is multiplied by the voxel in the equivalent position of the convolution kernel. Finally, the sum of the results is added to the output volume. Figure 31b , a representation of a 3D convolution operation can be observed where the highlighted voxels 3116a in the input 3116 are multiplied by their corresponding voxels in the kernel 3118. After these calculations, their sum 3120a is added to the output 3120.
[0479] Since the coordinates of the input volume are given by (x,y,z) and the convolution kernel has dimensions (P,Q,R), the 3D convolution operation can be mathematically defined as:
[0480] in
[0481] O is the result of convolution,
[0482] I is the input volume,
[0483] K is the convolution kernel, and
[0484] (p,q,r) are the coordinates of K.
[0485] For greater simplicity, the bias term is omitted in the above formula.
[0486] In addition to extracting spatial information from matrices like 2D convolutions, 3D convolutions also extract information that exists between consecutive matrices. This allows them to map both the spatial information of a 3D object and the temporal information of a set of sequential images.
[0487] Convolutional module
[0488] Figure 32One specific implementation of processing input data 3202 for each cycle of a single sequencing cycle in the series of t sequencing cycles to be base called is shown through a cascade 3200 of convolutional layers of a convolution module 3104 .
[0489] The convolution module 3104 processes each input data for each cycle separately through a cascade of convolutional layers 3200. A sequence of input data for each cycle is generated for a series of t sequencing cycles to be base called for a sequencing run, where t is any number between 1 and 1000. Thus, for example, when the series includes fifteen sequencing cycles, the sequence of input data for each cycle includes fifteen different input data for each cycle.
[0490] In one embodiment, each input data for each cycle includes only image channels (e.g., a red channel and a green channel) or image data (e.g., image data 202 described above). The image channels or image data depict the intensity emission of one or more clusters and their surrounding background captured at the corresponding sequencing cycle in the series. In another embodiment, in addition to the image channels or image data, each input data for each cycle also includes supplementary channels, such as a distance channel and a zoom channel (e.g., input data 1800 described above).
[0491] In the illustrated embodiment, the input data 3202 for each cycle includes two image channels, a red channel and a green channel, for a single sequencing cycle in the series of t sequencing cycles to be base called. Each image channel is encoded in an image patch of size 15×15. The convolution module 3104 includes five convolutional layers. Each convolutional layer has a set of twenty-five convolutional filters of size 3×3. In addition, the convolution filters use so-called SAME padding, which preserves the height and width of the input image or tensor. With SAME padding, padding is added to the input features so that the output feature map has the same size as the input features. In contrast, so-called VALID padding means there is no padding.
[0492] The first convolutional layer 3204 processes the input data 3202 for each cycle and produces a first convolutional representation 3206 of size 15×15×25. The second convolutional layer 3208 processes the first convolutional representation 3206 and produces a second convolutional representation 3210 of size 15×15×25. The third convolutional layer 3212 processes the second convolutional representation 3210 and produces a third convolutional representation 3214 of size 15×15×25. The fourth convolutional layer 3216 processes the third convolutional representation 3214 and produces a fourth convolutional representation 3218 of size 15×15×25. The fifth convolutional layer 3220 processes the fourth convolutional representation 3218 and produces a fifth convolutional representation 3222 of size 15×15×25. Note that SAME padding preserves the spatial dimensions of the resulting convolutional representations (e.g., 15×15). In some implementations, the number of convolution filters in a convolutional layer is a power of 2, such as 2, 4, 16, 32, 64, 128, 256, 512, and 1024.
[0493] As convolutions become deeper, information can be lost. To account for this, in some implementations, we use skip connections to (1) reintroduce the original input data for each cycle and (2) combine low-level spatial features extracted by earlier convolutional layers with high-level spatial features extracted by later convolutional layers. We observe that this improves base calling accuracy.
[0494] Figure 33 One embodiment of mixing 3300 the input data 3202 for each cycle of a single sequencing cycle with its corresponding convolutional representations 3206, 3210, 3214, 3218, and 3222 generated by the concatenation 3200 of the convolutional layers of the convolution module 3104 is depicted. The convolutional representations 3206, 3210, 3214, 3218, and 3222 are concatenated to form a sequence of convolutional representations 3304, which is then concatenated with the input data 3202 for each cycle to generate the mixed representation 3306. In other embodiments, summation is used instead of concatenation. In addition, mixing 3300 is performed by a mixer 3302.
[0495] Then, the flattener 3308 flattens the mixed representation 3306 and produces a flattened mixed representation 3310 for each cycle. In some implementations, the flattened mixed representation 3310 is a high-dimensional vector or a two-dimensional (2D) array that shares at least one dimension size (e.g., 15×1905, i.e., the same row-by-row dimensions) with the input data 3202 and the convolution representations 3206, 3210, 3214, 3218, and 3222 for each cycle. This introduces symmetry in the data, thereby facilitating feature extraction in downstream 3D convolutions.
[0496] Figure 32 and Figure 33Processing of per-cycle image data 3202 for a single sequencing cycle in the series of t sequencing cycles to be base called is shown. Convolution module 3104 processes the corresponding per-cycle image data for each of the t sequencing cycles individually and produces a corresponding per-cycle flattened mixture representation for each of the t sequencing cycles.
[0497] stacking
[0498] Figure 34 One embodiment of the invention is shown in which the flattened mixed representations of subsequent sequencing cycles are arranged into a stack 3400. In the illustrated embodiment, fifteen flattened mixed representations 3204a to 3204o of fifteen sequencing cycles are stacked in stack 3400. Stack 3400 is a 3D input volume that makes features from the spatial and temporal dimensions (i.e., multiple sequencing cycles) available in the same receptive field of a 3D convolutional filter. The stack is operated on by stacker 3402. In other embodiments, stack 3400 can be a tensor of any dimension (e.g., 1D, 2D, 4D, 5D, etc.).
[0499] Recursive modules
[0500] We use recursive processing to capture long-term dependencies in sequencing data, and specifically, consider second-order contributions from pre-phasing and phasing across-cycle sequencing images. Because time steps are used, recursive processing is used to analyze sequential data. The current hidden state representation at the current time step is a function of (i) the previous hidden state representation from the previous time step and (ii) the current input at the current time step.
[0501] The recursive module 3108 subjects the stack 3400 to a recursive application of 3D convolution in both the forward and backward directions (i.e., recursive process 3500), and produces a base call for each of the clusters at each of the t sequencing cycles in the series. The 3D convolution is used to extract spatiotemporal features from a subset of the flattened mixture representation in the stack 3400 on a sliding window basis. Each sliding window (w) corresponds to a respective sequencing cycle, and Figure 35a In some implementations, w is parameterized to 1, 2, 3, 5, 7, 9, 15, 21, etc., depending on the total number of sequencing cycles that are simultaneously base called. In one implementation, w is a fraction of the total number of sequencing cycles that are simultaneously base called.
[0502] Thus, for example, considering that each sliding window contains three subsequent flattened mixture representations from stack 3400, the stack includes fifteen flattened mixture representations 3204a through 3204o. Then, the first three flattened mixture representations 3204a through 3204c in the first sliding window correspond to the first sequencing cycle, the next three flattened mixture representations 3204b through 3204d in the second sliding window correspond to the second sequencing cycle, and so on. In some implementations, padding is used to encode a sufficient number of flattened mixture representations in the final sliding window corresponding to the final sequencing cycle, starting with the final flattened mixture representation 3204o.
[0503] At each time step, the recursive module 3108 accepts (1) the current input x(t) and (2) the previous hidden state representation h(t-1), and computes the current hidden state representation h(t). The current input x(t) includes only the subset of the flattened mixture representations from the stack 3400 that fall within the current sliding window ((w), orange). Thus, each current input x(t) at each time step is a 3D volume of multiple flattened mixture representations (e.g., 1, 2, 3, 5, 7, 9, 15, or 21 flattened mixture representations, depending on w). For example, when (i) the single flattened mixture representation is a two-dimensional (2D) representation of dimension 15×1905 and (ii) w is 7, then each current input x(t) at each time step is a 3D volume of dimension 15×1905×7.
[0504] The recursive module 3108 converts the first 3D convolution (W1 3DCONV ) is applied to the current input x(t), and the second 3D convolution (W2 3DCONV ) is applied to the previous hidden state representation h(t-1) to produce the current hidden state representation h(t). In some implementations, because the weights are shared, W1 3DCONV and W2 3DCONV are the same.
[0505] Gating Processing
[0506] In one implementation, the recurrent module 3108 processes the current input x(t) and the previous hidden state representation h(t-1) through a gated network, such as a long short-term memory (LSTM) network or a gated recurrent unit (GRU) network. For example, in an LSTM implementation, the current input x(t) along with the previous hidden state representation h(t-1) are processed through each of the four gates of the LSTM unit: input gate, activation gate, forget gate, and output gate. Figure 35b, which illustrates one implementation of processing 3500b a current input x(t) and a previous hidden state representation h(t-1) by an LSTM unit that applies a 3D convolution to the current input x(t) and the previous hidden state representation h(t-1) and produces the current hidden state representation h(t) as an output. In this implementation, the weights of the input gate, activation gate, forget gate, and output gate are subjected to a 3D convolution.
[0507] In some implementations, the gated unit (LSTM or GRU) does not use nonlinear / squeezing functions such as hyperbolic tangent function and sigmoid function.
[0508] In one embodiment, the current input x(t), the previous hidden state representation h(t-1), and the current hidden state representation h(t) are all 3D volumes with the same dimension and are processed or generated as 3D volumes through the input gate, the activation gate, the forget gate, and the output gate.
[0509] In one implementation, the 3D convolution of the recursive module 3108 uses a set of twenty-five convolution filters of size 3×3 with SAME padding. In some implementations, the size of the convolution filters is 5×5. In some implementations, the number of convolution filters used by the recursive module 3108 is factored into powers of 2, such as 2, 4, 16, 32, 64, 128, 256, 512, and 1024.
[0510] Bidirectional processing
[0511] The recursive module 3108 first processes the stack 3400 from the beginning to the end (top-down) on a sliding window basis and produces a sequence (vector) representing the current hidden state for the forward traversal.
[0512] The recursive module 3108 then processes the stack 3400 from the end to the beginning (bottom-up) on a sliding window basis and produces a sequence (vector) of current hidden state representations for backward / reverse traversal
[0513] In some implementations, for both directions, at each time step, this processing uses the gates of an LSTM or GRU. For example, at each time step, the forward current input x(t) is processed through the input gate, activation gate, forget gate, and output gate of the LSTM unit to produce the forward current hidden state representation And the backward current input x(t) is processed through the input gate, activation gate, forget gate and output gate of another LSTM unit to produce the backward current hidden state representation
[0514] Then, for each time step / sliding window / sequencing cycle, the recursive module 3108 combines (concatenates or sums or averages) the corresponding forward and backward current hidden state representations and produces a combined hidden state representation
[0515] The combined hidden representation is then processed through one or more fully connected networks to produce a dense representation. The dense representation is then processed through a softmax layer to produce the probability that each base incorporated into the cluster at a given sequencing cycle is A, C, T, and G. Based on this probability, the base is classified as A, C, T, or G. This is done synchronously or sequentially for each sequencing cycle (or each time step / sliding window) in the series of t sequencing cycles.
[0516] Those skilled in the art will appreciate that in other specific implementations, the hybrid architecture can process input data for fewer or more sequencing cycles and can include fewer or more convolutional layers and recursive layers. In addition, the dimensions of the input data, current and previous hidden representations, convolution filters, resulting feature maps, and outputs can be different. In addition, the number of convolution filters in the convolutional layer can be different. It can use various padding and stride configurations. It can use different classification functions (e.g., sigmoid or regression) and may or may not include a fully connected layer. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, expansion or hole convolution, transposed convolution, depth-separable convolution, point-by-point convolution, 1×1 convolution, grouped convolution, flat convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It can use one or more loss functions such as logistic regression / logarithmic loss function, multi-class cross entropy / softmax loss function, binary cross entropy loss function, mean squared error loss function, L1 loss function, L2 loss function, smooth L1 loss function and Huber loss function. It can use any parallelism, efficiency and compression scheme, such as TFRecords, compressed encoding (e.g., PNG), sharpening, parallel detection of map transformations, batching, prefetching, model parallelism, data parallelism and synchronous / asynchronous SGD. It can include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential lining units (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (e.g., max or average pooling), global average pooling layers and attention mechanisms.
[0517] Experimental results and observations
[0518] Figure 36 A specific implementation of trinucleotides (3-mers) in the training data for training a neural network-based base caller 218 is shown. Balance results in very little learning about the statistics of the genome in the training data, which then improves generalization. Heat map 3602 shows the balanced 3-mers of a first organism called "Acinetobacter baumannii" in the training data. Heat map 3604 shows the balanced 3-mers of a second organism called "Escherichia coli" in the training data.
[0519] Figure 37 The base calling accuracy of the RTA base caller is compared with the base calling accuracy of the neural network based base caller 218. Figure 37 As shown, the RTA base caller has a higher error percentage in both sequencing runs (Read: 1 and Read: 2). In other words, the neural network-based base caller 218 outperforms the RTA base caller in both sequencing runs.
[0520] Figure 38 The block-to-block generalization of the RTA base caller was compared to the block-to-block generalization of the neural network-based base caller 218 on the same block. That is, inference (testing) was performed using the neural network-based base caller 218 on the same block of data that was used for training.
[0521] Figure 39 The block-to-block generalization of the RTA base caller was compared to the block-to-block generalization of the neural network-based base caller 218 on the same block and on different blocks. That is, the neural network-based base caller 218 was trained on data from the cluster on the first block, but performed inference on data from the cluster on the second block. In a same-block implementation, the neural network-based base caller 218 was trained on data from the cluster on block five and tested on data from the cluster on block five. In a different-block implementation, the neural network-based base caller 218 was trained on data from the cluster on block ten and tested on data from the cluster on block five.
[0522] Figure 40 The block-to-block generalization of the RTA base caller is also compared to the block-to-block generalization of the neural network-based base caller 218 on a different block. In the specific implementation of the different blocks, the neural network-based base caller 218 is trained on data from the cluster on block ten and tested once on data from the cluster on block five, and then trained on data from the cluster on block twenty and tested on data from the cluster on block five.
[0523] Figure 41 Figure 2 shows how different sizes of image patches fed as input to the neural network-based base caller 218 affect base calling accuracy. In two sequencing runs (Read: 1 and Read: 2), the percentage of error decreases as the patch size increases from 3×3 to 11×11. That is, the neural network-based base caller 218 produces more accurate base calls with larger image patches. In some implementations, base calling accuracy and computational efficiency are balanced by using image patches no larger than 100×100 pixels. In other implementations, image patches as large as 3000×3000 pixels (and larger) are used.
[0524] Figure 42 、 Figure 43 、 Figure 44 and Figure 45 The neural network-based base caller 218 is shown performing slot-to-slot generalization on training data from Acinetobacter baumannii and Escherichia coli.
[0525] Go to Figure 43 In one embodiment, a neural network-based base detector 218 is based on the E. coli data training of the cluster on the first groove of the circulation cell, and based on the Acinetobacter baumannii data test of the cluster on both the first groove and the second groove of the circulation cell. In another embodiment, a neural network-based base detector 218 is based on the Acinetobacter baumannii data training of the cluster on the first groove, and based on the Acinetobacter baumannii data test of the cluster on both the first groove and the second groove. In yet another embodiment, a neural network-based base detector 218 is based on the E. coli data training of the cluster on the second groove, and based on the Acinetobacter baumannii data test of the cluster on both the first groove and the second groove. In yet another embodiment, a neural network-based base detector 218 is based on the Acinetobacter baumannii data training of the cluster on the second groove, and based on the Acinetobacter baumannii data test of the cluster on both the first groove and the second groove.
[0526] In one embodiment, the base caller 218 based on the neural network is trained based on the E. coli data of the cluster on the first groove of the circulation cell, and is tested based on the E. coli data of the cluster on both the first groove and the second groove of the circulation cell. In another embodiment, the base caller 218 based on the neural network is trained based on the Acinetobacter baumannii data of the cluster on the first groove, and is tested based on the E. coli data of the cluster on both the first groove and the second groove. In yet another embodiment, the base caller 218 based on the neural network is trained based on the E. coli data of the cluster on the second groove, and is tested based on the E. coli data of the cluster on the first groove. In yet another embodiment, the base caller 218 based on the neural network is trained based on the Acinetobacter baumannii data of the cluster on the second groove, and is tested based on the E. coli data of the cluster on both the first groove and the second groove.
[0527] exist Figure 43 , base calling accuracy (measured by percent error) is shown for each of these implementations for two sequencing runs (eg, Read: 1 and Read: 2).
[0528] Go to Figure 44 In one embodiment, a neural network-based base detector 218 is based on the Escherichia coli data training of the cluster on the first groove of the circulation cell, and based on the Acinetobacter baumannii data test of the cluster on the first groove. In another embodiment, a neural network-based base detector 218 is based on the Acinetobacter baumannii data training of the cluster on the first groove, and based on the Acinetobacter baumannii data test of the cluster on the first groove. In another embodiment, a neural network-based base detector 218 is based on the Escherichia coli data training of the cluster on the second groove, and based on the Acinetobacter baumannii data test of the cluster on the first groove. In another embodiment, a neural network-based base detector 218 is based on the Acinetobacter baumannii data training of the cluster on the second groove, and based on the Acinetobacter baumannii data test of the cluster on the first groove.
[0529] In one embodiment, the base caller 218 based on the neural network is trained based on the E. coli data of the cluster on the first groove of the circulation cell, and is tested based on the E. coli data of the cluster on the first groove. In another embodiment, the base caller 218 based on the neural network is trained based on the Acinetobacter baumannii data of the cluster on the first groove, and is tested based on the E. coli data of the cluster on the first groove. In another embodiment, the base caller 218 based on the neural network is trained based on the E. coli data of the cluster on the second groove, and is tested based on the E. coli data of the cluster on the first groove. In another embodiment, the base caller 218 based on the neural network is trained based on the Acinetobacter baumannii data of the cluster on the second groove, and is tested based on the E. coli data of the cluster on the first groove.
[0530] exist Figure 44 In , base calling accuracy (measured by percent error) is shown for each of these implementations for two sequencing runs (e.g., Read: 1 and Read: 2). Figure 43 and Figure 44 In comparison, it can be seen that the implementations covered by the latter lead to a 50% to 80% reduction in error.
[0531] Go to Figure 45 In one embodiment, the base detector 218 based on a neural network is trained based on the Escherichia coli data of the cluster on the first groove of the circulation cell, and is tested based on the Acinetobacter baumannii data of the cluster on the second groove. In another embodiment, the base detector 218 based on a neural network is trained based on the Acinetobacter baumannii data of the cluster on the first groove, and is tested based on the Acinetobacter baumannii data of the cluster on the second groove. In another embodiment, the base detector 218 based on a neural network is trained based on the Escherichia coli data of the cluster on the second groove, and is tested based on the Acinetobacter baumannii data of the cluster on the first groove. In the second first groove. In another embodiment, the base detector 218 based on a neural network is trained based on the Acinetobacter baumannii data of the cluster on the second groove, and is tested based on the Acinetobacter baumannii data of the cluster on the second groove.
[0532] In one embodiment, the base detector 218 based on the neural network is based on the E. coli data training of the cluster on the first groove of the circulation cell, and based on the E. coli data test of the cluster on the second groove. In another embodiment, the base detector 218 based on the neural network is based on the Acinetobacter baumannii data training of the cluster on the first groove, and based on the E. coli data test of the cluster on the second groove. In another embodiment, the base detector 218 based on the neural network is based on the E. coli data training of the cluster on the second groove, and based on the E. coli data test of the cluster on the second groove. In another embodiment, the base detector 218 based on the neural network is based on the Acinetobacter baumannii data training of the cluster on the second groove, and based on the E. coli data test of the cluster on the second groove.
[0533] exist Figure 45 In , base calling accuracy (measured by percent error) is shown for each of these implementations for two sequencing runs (e.g., Read: 1 and Read: 2). Figure 43 and Figure 45 In comparison, it can be seen that the implementations covered by the latter lead to a 50% to 80% reduction in error.
[0534] Figure 46 Shown above relative to Figure 42 、 Figure 43 、 Figure 44 and Figure 45 The error distribution generalized from channel to channel. In one embodiment, the error distribution detects errors in base calling A and T bases in the green channel.
[0535] Figure 47 will be Figure 46 Error distribution of the detected error sources attributed to low cluster intensities in the green channel.
[0536] Figure 48 The error distributions of the RTA base caller and the neural network-based base caller 218 for two sequencing runs (Read 1 and Read 2) were compared. This comparison confirmed the superior base calling accuracy of the neural network-based base caller 218.
[0537] Figure 49a Shown is the run-to-run generalization of the neural network-based base caller 218 on four different instruments.
[0538] Figure 49b The run-to-run generalization of the neural network-based base caller 218 across four different runs performed on the same instrument is shown.
[0539] Figure 50The genomic statistics of the training data used to train the neural network-based base caller 218 are shown.
[0540] Figure 51 The genomic context of the training data used to train the neural network-based base caller 218 is shown.
[0541] Figure 52 The base calling accuracy of the neural network-based base caller 218 in base called long reads (eg, 2×250) is shown.
[0542] Figure 53 One specific implementation of how the neural network-based base caller 218 focuses on a center cluster pixel and its neighboring pixels on an image patch is shown.
[0543] Figure 54 Various hardware components and configurations are shown according to one implementation for training and running the neural network-based base caller 218. In other implementations, different hardware components and configurations are used.
[0544] Figure 55 Shown are various sequencing tasks that can be performed using the neural network-based base caller 218. Some examples include quality scoring (QScoring) and variant classification. Figure 55 Also listed are some exemplary sequencing instruments for which the neural network-based base caller 218 performs base calling.
[0545] Figure 56 is a scatter plot 5600 visualized by t-distributed stochastic neighbor embedding (t-SNE) and depicts the base calling results of the neural network based base caller 218. The scatter plot 5600 shows that the base calling results are clustered into 64 (4 3 ) groups, where each group primarily corresponds to a specific input 3-mer (trinucleotide repeat pattern). This is the case because the neural network-based base caller 218 processes input data for at least three sequencing cycles and learns sequence-specific motifs to generate current base calls based on previous base calls and subsequent base calls.
[0546] Quality Rating
[0547] Quality scoring refers to the process of assigning a quality score to each base call. Quality scores are defined according to the Phred framework, which transforms the values of predicted features of a sequencing trajectory into probabilities based on a quality table. The quality table is obtained by training on a calibration dataset and is updated as the characteristics of the sequencing platform change. This probabilistic interpretation of quality scores allows for the fair integration of diverse sequence reads in downstream analyses such as variant calling and sequence assembly. Therefore, an effective model for defining quality scores is essential for any base caller.
[0548] Let's first describe what a quality score is. A quality score is a measure of the probability of sequencing errors in base calls. A high quality score indicates a more reliable base call, with a lower probability of error. For example, if a base has a quality score of Q30, the probability of incorrectly calling that base is 0.001. This also indicates a base call accuracy of 99.9%.
[0549] The following table shows the relationship between the base calling quality score and its corresponding error probability, base calling accuracy, and base calling error rate:
[0550]
[0551] We now describe how quality scores are generated. During a sequencing run, a quality score is assigned to each base call in each cluster on each block for each sequencing cycle. The Illumina quality score for each base call is calculated in a two-step process. For each base call, multiple quality predictor values are calculated. Quality predictor values are observable properties of the cluster from which the base call was extracted. These properties include attributes such as intensity distribution and signal-to-noise ratio and measure various aspects of base call reliability. They have been empirically determined to correlate with the quality of the base call.
[0552] A quality model (also called a quality table or Q-table) lists combinations of quality predictor values and associates them with corresponding quality scores; this relationship is determined through a calibration process using empirical data. To estimate a new quality score, a quality predictor value is calculated for each new base call and compared to the values in the precalibrated quality table.
[0553] We now describe how to calibrate a quality table. Calibration is a process in which a statistical quality table is derived from empirical data consisting of a variety of well-characterized human and nonhuman samples sequenced on multiple instruments. Using a modified version of the Phred algorithm, the quality table is developed and refined using properties of the raw signal and error rates determined by aligning reads to an appropriate reference sequence.
[0554] We now describe why quality tables change from time to time. Quality tables provide quality scores for runs generated based on a specific instrument configuration and chemistry version. When important characteristics of the sequencing platform change, such as new hardware, software, or chemistry versions, the quality model requires recalibration. For example, improvements in sequencing chemistry require recalibration of the quality table to accurately score the new data, which consumes significant processing time and computing resources.
[0555] Neural network-based quality scoring
[0556] We disclose a neural network-based technique for quality scoring that does not use quality predictor values or quality tables. Instead, it infers quality scores based on confidence in the predictors of a well-calibrated neural network. In the context of neural networks, "calibration" refers to the consistency or correlation between subjective predictions and empirical long-term frequencies. This is a deterministic notion of normality: if a neural network claims that a particular label is correct 90% of the time, then during the evaluation period, 90% of all labels that meet a 90% probability of being correct should be correct. Note that calibration is an orthogonal issue to accuracy: a neural network's predictions can be accurate but miscalibrated.
[0557] The disclosed neural networks are well calibrated because they are trained on a large-scale training set with a variety of sequencing characteristics that fully simulate the base calling domain of a real-world sequencing run. Specifically, sequencing images obtained from a variety of sequencing platforms, sequencing instruments, sequencing protocols, sequencing chemistries, sequencing reagents, cluster densities, and flow cells are used as training examples to train the neural network. In other specific implementations, different base calling and quality scoring models are used for different sequencing platforms, sequencing instruments, sequencing protocols, sequencing chemistries, sequencing reagents, cluster densities, and / or flow cells, respectively.
[0558] For each of the four base call classes (A, C, T, and G), a large number of sequencing images were used as training examples, which identified the intensity patterns representing the corresponding base call class under a wide range of sequencing conditions. This, in turn, eliminated the need to expand the classification capabilities of the neural network to new classes not present in the training. Furthermore, each training example was accurately labeled with the corresponding ground truth value based on alignment of the reads to the appropriate reference sequence. The result is a well-calibrated neural network whose confidence in its predictions can be interpreted as a measure of the certainty of the quality score, expressed mathematically below.
[0559] Let Y = {A, C, T, G} denote the class label set of base calling classes A, C, T, and G and let X denote the input space. Let N θ(y|x) represents the probability distribution predicted by one of the disclosed neural networks for the input x∈X and let θ represent the parameters of the neural network. i The training instance x i , neural network prediction label The prediction gets a correct score (if Then c i =1, otherwise 0), and the confidence score
[0560] Neural Network N θ (y|x) is well calibrated on the data distribution D because for all (x i ,y i )∈D and α, r i =α has a probability of c i = 1. For example, given 100 predictions in a sample from D, and each prediction has a confidence level of 0.8, 80 predictions are made by the neural network N θ (y|x) is correctly classified. More formally, P θ,D (r,c) represents the neural network N θ The distribution of r and c values of the prediction of (y|x) for D is expressed as Among them I α represents a small non-zero interval around α.
[0561] Because fully calibrated neural networks are trained based on a variety of training sets, unlike quality predictor values or quality tables, they are not specific to instrument configurations and chemistry versions. This has two advantages. First, fully calibrated neural networks eliminate the need to derive different quality tables from separate calibration processes for different sequencing instrument types. Second, for the same sequencing instrument, they eliminate the need for recalibration when the sequencing instrument's characteristics change. More details are below.
[0562] Inferring quality scores based on Softmax confidence probabilities
[0563] The first fully calibrated neural network is the neural network-based base caller 218, which processes input data from the sequencing image 108 and generates base call confidence probabilities for the bases A, C, T, and G. The base call confidence probabilities can also be considered as likelihoods or classification scores. In one embodiment, the neural network-based base caller 218 uses a softmax function to generate the base call confidence probabilities as softmax function scores.
[0564] Quality scores are inferred from the base call confidence probabilities generated by the softmax function of the neural network-based base caller 218 because the softmax function scores are calibrated (i.e., they represent the likelihood that the ground truth is correct) and therefore correspond naturally to quality scores.
[0565] We demonstrate the correspondence between base call confidence probabilities and quality scores by selecting a set of base call confidence probabilities generated by the neural network-based base caller 218 during training and determining their base call error rates (or base call accuracy rates).
[0566] Thus, for example, we select a base call confidence probability of "0.90" generated by the neural network-based base caller 218. When the neural network-based base caller 218 makes a base call prediction with a softmax function score of 0.90, we use a large number of examples (e.g., in the range of 10,000 to 1,000,000). These large number of examples can be obtained from a validation set or a test set. Then, based on a comparison with the corresponding ground truth base calls associated with the respective examples in the large number of examples, we determine how many of the examples in the large number of examples had correct base call predictions.
[0567] We observe that in ninety percent of these many instances, the base call was correctly predicted, with ten percent of the base calls being false. This means that for a softmax score of 0.90, the base call error rate is 10% and the base call accuracy is 90%, which in turn corresponds to a quality score of Q10 (see table above). Similarly, for other softmax scores such as 0.99, 0.999, 0.9999, 0.99999, and 0.999999, we observe a correspondence with quality scores of Q20, Q30, Q40, Q50, and Q60, respectively. This is in Figure 59a In other specific implementations, we observe a correspondence between the softmax function score and the quality score (such as Q9, Q11, Q12, Q23, Q25, Q29, Q37 and Q39).
[0568] We also observe a correspondence with the quality score of the group. For example, a softmax function score of 0.80 corresponds to the quality score Q06 of the group, a softmax function score of 0.95 corresponds to the quality score Q15 of the group, a softmax function score of 0.993 corresponds to the quality score Q22 of the group, a softmax function score of 0.997 corresponds to the quality score Q27 of the group, a softmax function score of 0.9991 corresponds to the quality score Q33 of the group, a softmax function score of 0.9995 corresponds to the quality score Q37 of the group, and a softmax function score of 0.9999 corresponds to the quality score Q40 of the group. This is in Figure 59b Shown in.
[0569] The sample sizes used herein are large to avoid the small sample problem and may be, for example, in the range of 10,000 to 1,000,000. In some implementations, the sample size of instances for determining the base call error rate (or base call accuracy) is selected based on the softmax function score being evaluated. For example, for a softmax function score of 0.99, the sample includes 100 instances, for a softmax function score of 0.999, the sample includes 1,000 instances, for a softmax function score of 0.9999, the sample includes 10,000 instances, for a softmax function score of 0.99999, the sample includes 100,000 instances, and for a softmax function score of 0.999999, the sample includes 1,000,000 instances.
[0570] Regarding softmax, it is the output activation function used for multi-class classification. Formally, training a so-called softmax classifier is a regression to class probabilities, not to a true classifier, because it does not return classes, but rather confidence predictions about the likelihood of each class. The softmax function takes the class values and converts them into probabilities that sum to 1. The softmax function compresses any k-dimensional vector of real values into a k-dimensional vector of real values in the range 0 to 1. Therefore, using the softmax function ensures that the output is a valid, exponentially normalized probability mass function (non-negative and summing to 1).
[0571] Taking into account is a vector The i-th element of
[0572] in
[0573] is a vector of length n, where n is the number of classes in the classification. The values of these elements are between 0 and 1 and sum to 1, so that they represent a valid probability distribution.
[0574] The exemplary softmax activation function 5706 is Figure 57 Softmax 5706 is applied to three classes as follows: Note that the three outputs always sum to 1. Therefore, they define a discrete probability mass function.
[0575] When used for classification, Give the probability of belonging to class i.
[0576]
[0577] The name "softmax" can be somewhat confusing. This function is more closely related to the argmax function than to the max function. The term "soft" comes from the fact that the softmax function is continuous and differentiable. The result of the argmax function is represented as a one-hot vector, which is not continuous or differentiable. Therefore, the softmax function provides a "softened" version of the argmax. It might be better to call the softmax function "softargmax", but the current name is a deeply ingrained convention.
[0578] Figure 57 One embodiment of selecting 5700 a base call confidence probability 3004 of a neural network-based base caller 218 for quality scoring is shown. The base call confidence probability 3004 of the neural network-based base caller 218 can be a classification score (e.g., a softmax function score or a sigmoid function score) or a regression score. In one embodiment, the base call confidence probability 3004 is generated during training 3000.
[0579] In some implementations, the selection 5700 is done based on quantization performed by a quantizer 5702 that accesses the base call confidence probabilities 3004 and generates a quantized classification score 5704. The quantized classification score 5704 can be any real number. In one implementation, based on the definition The selection formula is used to select the quantitative classification score 5704. In another embodiment, based on the definition of A selection formula is used to select the quantitative classification score 5704.
[0580] Figure 58One specific implementation of a neural network-based quality score 5800 is shown. For each of the quantified classification scores 5704, a base call error rate 5808 and / or a base call accuracy rate 5810 is determined by comparing its base call prediction 3004 with the corresponding ground truth base call 3008 (e.g., on batches with different sample sizes). This comparison is performed by a comparator 5802, which in turn includes a base call error rate determiner 5804 and a base call accuracy rate determiner 5806.
[0581] Then, to establish a correspondence between the quantitative classification score 5704 and the quality score, a fit determiner 5812 determines a fit between the quantitative classification score 5704 and its base call error rate 5808 (and / or its base call accuracy rate 5810). In one embodiment, the fit determiner 5812 is a regression model.
[0582] Based on the fit, the quality score is associated with the quantitative classification score 5704 via correlator 5814.
[0583] Figures 59a to 59b One specific implementation of the correspondence 5900 between quality scores and base call confidence predictions made by the neural network-based base caller 218 is depicted. The base call confidence probabilities of the neural network-based base caller 218 can be classification scores (e.g., softmax function scores or sigmoid function scores) or regression scores. Figure 59a This is a quality score corresponding scheme 5900a for quality score. Figure 59b This is a quality score correspondence scheme 5900a for the quality score of the group.
[0584] infer
[0585] Figure 60 One specific implementation of inferring a quality score based on base call confidence predictions made by the neural network-based base caller 218 during inference 6000 is shown. The base call confidence probabilities of the neural network-based base caller 218 can be classification scores (e.g., softmax function scores or sigmoid function scores) or regression scores.
[0586] During inference 6000, predicted base calls 6006 are assigned the quality scores 6008 to which their base call confidence probabilities (i.e., highest softmax function scores (red)) most closely correspond. In some implementations, quality score correspondences 5900 are made by looking up quality score correspondences 5900a-5900b and operated by a quality score inferor 6012.
[0587] In some implementations, the purity filter 6010 terminates base calling for a given cluster when the quality score 6008 assigned to its called bases, or the average quality score over subsequent base calling cycles, falls below a preset threshold.
[0588] Inference 6000 includes hundreds, thousands, and / or millions of iterations of forward propagation 6014, including parallelization techniques such as batch processing. Inference 6000 is performed on inference data 6002 including input data (wherein image channels are derived from sequenced images 108 and / or supplementary channels (e.g., distance channels, scale channels)). Inference 6000 is operated by a tester 6004.
[0589] Directly predict base call quality
[0590] The second well-calibrated neural network is a neural network-based quality scorer 6102 that processes input data derived from sequencing image 108 and directly generates quality indicators.
[0591] In one implementation, the neural network-based quality scorer 6102 is a multilayer perceptron (MLP). In another implementation, the neural network-based quality scorer 6102 is a feedforward neural network. In yet another implementation, the neural network-based quality scorer 6102 is a fully connected neural network. In another implementation, the neural network-based quality scorer 6102 is a fully convolutional neural network. In yet another implementation, the neural network-based quality scorer 6102 is a semantic segmentation neural network.
[0592] In one implementation, the neural network-based quality scorer 6102 is a convolutional neural network (CNN) having multiple convolutional layers. In another implementation, the neural network-based base caller is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another implementation, the neural network-based base caller includes both a CNN and an RNN.
[0593] In other embodiments, the neural network-based quality scorer 6102 may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or dilated convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1×1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It may use one or more loss functions such as logistic regression / logarithmic loss, multi-class cross entropy / softmax loss, binary cross entropy loss, mean squared error loss, L1 loss, L2 loss, smooth L1 loss, and Huber loss. It may use any parallelism, efficiency, and compression scheme, such as TFRecords, compressed encoding (e.g., PNG), sharpening, parallel detection of map transformations, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous SGD. It may include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (for example, nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential lining unit (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (for example, maximum or average pooling), global average pooling layers and attention mechanisms.
[0594] In some implementations, the neural network-based quality scorer 6102 has the same architecture as the neural network-based base caller 218.
[0595] The input data may include image channels and / or supplementary channels (e.g., distance channels, scaling channels) derived from the sequencing image 108. The neural network-based quality scorer 6102 processes the input data and generates an alternative representation of the input data. The alternative representation is a convolutional representation in some implementations and a hidden representation in other implementations. The alternative representation is then processed by an output layer to generate an output. The output is used to generate a quality indicator.
[0596] In one implementation, the same input data is fed to the neural network-based base caller 218 and the neural network-based quality scorer 6102 to produce (i) base calls from the neural network-based base caller 218 and (ii) corresponding quality indicators from the neural network-based quality scorer 6102. In some implementations, the neural network-based base caller 218 and the neural network-based quality scorer 6102 are jointly trained using end-to-end back-propagation.
[0597] In one implementation, the neural network-based quality scorer 6102 outputs a quality indicator for a single target cluster for a particular sequencing cycle. In another implementation, the neural network-based quality scorer outputs a quality indicator for each target cluster in a plurality of target clusters for a particular sequencing cycle. In yet another implementation, the neural network-based quality scorer outputs a quality indicator for each target cluster in a plurality of target clusters for each sequencing cycle in a plurality of sequencing cycles, thereby generating a quality indicator sequence for each target cluster.
[0598] In one implementation, the neural network-based quality scorer 6102 is a convolutional neural network trained on training examples that include data from sequencing images 108 labeled with ground truth base call quality values. The neural network-based quality scorer 6102 is trained using a backpropagation-based gradient update technique that progressively matches the convolutional neural network's 6102 base call quality predictions 6104 to the ground truth base call quality values 6108. In some implementations, if a base is an incorrect base call, the base is labeled as 0; otherwise, the base is labeled as 1. Thus, the output corresponds to a probability of error. In one implementation, this eliminates the need to use sequence context as an input feature.
[0599] The input module of the convolutional neural network 6102 feeds data from the sequencing images 108 captured at one or more sequencing cycles to the convolutional neural network 6102 to determine the quality of one or more bases called for one or more clusters.
[0600] The output module of the convolutional neural network 6102 converts the analysis performed by the convolutional neural network 6102 into an output 6202 that identifies the quality of the one or more bases called for the one or more clusters.
[0601] In one embodiment, the output module further comprises a softmax classification layer that generates the probability that the quality state is high quality, medium quality (optional, as shown by the dotted line), and low quality. In another embodiment, the output module further comprises a softmax classification layer that generates the probability that the quality state is high quality and low quality. Those skilled in the art will appreciate that other classes of quality scores that can be differently and distinguishably bucketed can be used. The softmax classification layer generates the probability that the quality is assigned a plurality of quality scores. Based on the probability, the quality is assigned a quality score from one of the plurality of quality scores. The quality score is logarithmically based on the probability of base call errors. The plurality of quality scores include Q6, Q10, Q15, Q20, Q22, Q27, Q30, Q33, Q37, Q40, and Q50. In another embodiment, the output module further comprises a regression layer that generates a continuous value for the recognition quality.
[0602] In some implementations, the neural network-based quality scorer 6102 further includes a supplemental input module that supplements the data from the sequencing image 108 with quality predictor values for the called bases, and feeds the quality predictor values along with the data from the sequencing image to the convolutional neural network 6102.
[0603] In some implementations, quality predictor values include online overlap, purity, phasing, start5, hexamer score, motif accumulation, endiness, near homopolymer, intensity decay, penultimate purity (chastity), signal overlap with background (SOWB) and / or shifted purity G adjustment. In other implementations, quality predictor values include peak height, peak width, peak position, relative peak position, peak height ratio, peak spacing ratio and / or peak correspondence. Additional details on quality predictor values can be found in U.S. Patent Publication Nos. 2018 / 0274023 and 2012 / 0020537, which are incorporated by reference as if fully set forth herein.
[0604] train
[0605] Figure 61One embodiment of training 6100 a neural network-based quality scorer 6102 to process input data derived from sequenced images 108 and directly generate quality indicators is shown. The neural network-based quality scorer 6102 is trained using a back-propagation-based gradient update technique that compares a predicted quality indicator 6104 with an accurate quality indicator 6108 and calculates an error 6106 based on the comparison. The error 6106 is then used to calculate a gradient that is applied to the weights and parameters of the neural network-based quality scorer 6102 during back-propagation 6110. Training 6100 is performed by a trainer 1510 using a stochastic gradient update algorithm such as ADAM.
[0606] Trainer 1510 trains a neural network-based quality scorer 6102 using training data 6112 (derived from sequenced images 108) over thousands and millions of iterations of forward propagation 6116 to generate predicted quality indicators and backward propagation 6110 to update weights and parameters based on errors 6106. In some implementations, training data 6112 is supplemented with quality predictor values 6114. Additional details regarding training 6100 can be found in the Appendix entitled "Deep Learning Tools."
[0607] infer
[0608] Figure 62 One embodiment of the invention is shown in which a quality indicator is generated directly during inference 6200 as an output of a neural network-based quality scorer 6102. Inference 6200 includes hundreds, thousands, and / or millions of iterations of forward propagation 6208, including parallelization techniques such as batch processing. Inference 6200 is performed on inference data 6204, including input data (wherein image channels are derived from sequenced images 108 and / or supplementary channels (e.g., distance channels, scale channels)). In some embodiments, inference data 6204 is supplemented with quality predictor values 6206. Inference 6200 is operated by a tester 6210.
[0609] Data preprocessing
[0610] In some implementations, the disclosed technology uses a preprocessing technique that is applied to pixels in the image data 202 and produces preprocessed image data 202p. In such implementations, the preprocessed image data 202p is provided as input to the neural network-based base caller 218 instead of the image data 202. Data preprocessing is performed by a data preprocessor 6602, which in turn can include a data normalizer 6632 and a data enhancer 6634.
[0611] Figure 66Different implementations of data preprocessing are shown, which may include data normalization and data augmentation.
[0612] Data normalization
[0613] In one embodiment, data normalization is applied to the pixels in the image data 202 on a patch-by-patch basis. This includes normalizing the intensity values of the pixels in the image patch so that the fifth percentile of the pixel intensity histogram of the resulting normalized image patch is zero and the ninety-fifth percentile is one. That is, in the normalized image patch, (i) 5% of the pixels have an intensity value less than zero, and (ii) another 5% of the pixels have an intensity value greater than one. The corresponding image patches of the image data 202 can be normalized individually, or the image data 202 can be normalized all at once. The result is a normalized image patch 6616, which is an example of preprocessed image data 202p. Data normalization is performed by a data normalizer 6632.
[0614] Data augmentation
[0615] In one specific implementation, data augmentation is applied to the intensity values of pixels in the image data 202. This includes (i) multiplying the intensity values of all pixels in the image data 202 by the same scaling factor, and (ii) adding the same offset value to the scaled intensity values of all pixels in the image data 202. For a single pixel, this can be represented by the following formula:
[0616] Enhanced pixel intensity (API) = aX + b
[0617] Where A is the scaling factor, X is the original pixel intensity, b is the offset value, and aX is the scaled pixel intensity
[0618] The result is an enhanced image patch 6626, which is also an example of pre-processed image data 202p. Data augmentation is performed by a data enhancer 6634.
[0619] Figure 67 shows that when the neural network based base caller 218 is trained on bacterial data and tested on human data, Figure 66 Our data normalization technique (DeepRTA(Normalization)) and data augmentation technique (DeepRTA(Enhancement)) reduced the base calling error percentage, where bacterial data and human data shared the same assays (e.g., both contained intron data).
[0620] Figure 68 It is shown that when the neural network based base caller 218 is trained on non-exonic data (e.g., intronic data) and tested on exonic data, Figure 66The data normalization technology (DeepRTA(Normalization)) and data enhancement technology (DeepRTA(Enhancement)) of the proposed method reduce the percentage of base calling errors.
[0621] In other words, Figure 66 The data normalization technique and data augmentation technique allow the neural network-based base caller 218 to better generalize to data not seen in training, thereby reducing overfitting.
[0622] In one implementation, data augmentation is applied during both training and inference. In another implementation, data augmentation is applied only during training. In yet another implementation, data augmentation is applied only during inference.
[0623] Sequencing system
[0624] Figure 63A and Figure 63B A specific implementation of a sequencing system 6300A is depicted. The sequencing system 6300A includes a configurable processor 6346. The configurable processor 6346 implements the base calling technology disclosed herein. The sequencing system is also referred to as a "sequencer."
[0625] The sequencing system 6300A can be operated to obtain any information or data related to at least one of a biological substance or a chemical substance. In some implementations, the sequencing system 6300A is a workstation that can be similar to a desktop device or desktop computer. For example, most (or all) of the systems and components used to perform the desired reactions can be located within a common housing 6302.
[0626] In certain embodiments, the sequencing system 6300A is a nucleic acid sequencing system configured for various applications, including but not limited to de novo sequencing, resequencing of whole genomes or target genome regions, and metagenomics. The sequencer can also be used for DNA or RNA analysis. In some embodiments, the sequencing system 6300A can also be configured to generate reaction sites in a biosensor. For example, the sequencing system 6300A can be configured to receive a sample and generate surface-attached clusters of clonally amplified nucleic acids derived from the sample. Each cluster can constitute or be part of a reaction site in a biosensor.
[0627] The exemplary sequencing system 6300A may include a system receptacle or interface 6310 configured to interact with a biosensor 6312 to perform a desired reaction within the biosensor 6312. Figure 63AIn the description, the biosensor 6312 is loaded into the system receptacle 6310. However, it should be understood that a cartridge including the biosensor 6312 can be inserted into the system receptacle 6310, and in some cases, the cartridge can be temporarily or permanently removed. As described above, the cartridge can also include, among other things, a fluid control component and a fluid storage component.
[0628] In a particular embodiment, the sequencing system 6300A is configured to perform a large number of parallel reactions within the biosensor 6312. The biosensor 6312 includes one or more reaction sites where the desired reaction can occur. The reaction sites can be, for example, fixed to a solid surface of the biosensor or fixed to beads (or other removable substrates) located within a corresponding reaction chamber of the biosensor. The reaction sites can include, for example, clusters of clonally amplified nucleic acids. The biosensor 6312 can include a solid-state imaging device (e.g., a CCD or CMOS imager) and a flow cell mounted thereon. The flow cell can include one or more flow channels that receive solution from the sequencing system 6300A and direct the solution to the reaction sites. Optionally, the biosensor 6312 can be configured to engage a thermal element for transferring thermal energy into or out of the flow channel.
[0629] The sequencing system 6300A can include various components, assemblies, and systems (or subsystems) that interact with each other to execute a predetermined method or assay protocol for a biological or chemical analysis. For example, the sequencing system 6300A includes a system controller 6306 that can communicate with the various components, assemblies, and subsystems of the sequencing system 6300A and the biosensor 6312. For example, in addition to the system receptacle 6310, the sequencing system 6300A can also include a fluid control system 6308 to control the flow of fluids throughout the fluid network of the sequencing system 6300A and the biosensor 6312; a fluid storage system 6314 configured to store all fluids (e.g., gases or liquids) that can be used by the bioassay system; a temperature control system 6304 that can regulate the temperature of fluids in the fluid network, the fluid storage system 6314, and / or the biosensor 6312; and an illumination system 6316 configured to illuminate the biosensor 6312. As described above, if a cartridge having a biosensor 6312 is loaded into the system receptacle 6310, the cartridge may also include a fluid control component and a fluid storage component.
[0630] As also shown in the figure, the sequencing system 6300A may include a user interface 6318 for interacting with the user. For example, the user interface 6318 may include a display 6320 for displaying or requesting information from the user and a user input device 6322 for receiving user input. In some specific implementations, the display 6320 and the user input device 6322 are the same device. For example, the user interface 6318 may include a touch-sensitive display that is configured to detect the presence of individual touches and also identify the position of the touch on the display. However, other user input devices 6322 may be used, such as a mouse, touchpad, keyboard, keypad, handheld scanner, voice recognition system, motion recognition system, etc. As will be discussed in more detail below, the sequencing system 6300A can communicate with various components including a biosensor 6312 (e.g., in the form of a cartridge) to perform the desired reaction. The sequencing system 6300A may also be configured to analyze the data obtained from the biosensor to provide the desired information to the user.
[0631] The system controller 6306 may include any processor-based or microprocessor-based system, including the use of a microcontroller, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a coarse-grained reconfigurable architecture (CGRA), a logic circuit, and any other circuit or processor capable of performing the functions described herein. The above examples are merely exemplary and are therefore not intended to limit the definition and / or meaning of the term system controller in any way. In an exemplary embodiment, the system controller 6306 executes an instruction set stored in one or more storage elements, memories, or modules to obtain at least one of the detection data and analyze the detection data. The detection data may include multiple pixel signal sequences so that the pixel signal sequence of each sensor (or pixel) from millions of sensors (or pixels) can be detected within many base detection cycles. The storage element may be in the form of an information source or physical memory element within the sequencing system 6300A.
[0632] The instruction set may include various commands that instruct the sequencing system 6300A or the biosensor 6312 to perform specific operations (such as the methods and processes of the various specific implementations described herein). The instruction set may be in the form of a software program that may form part of one or more tangible non-transitory computer-readable media. As used herein, the terms "software" and "firmware" are interchangeable and include any computer program stored in a memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The above memory types are exemplary only and do not limit the memory types that can be used to store computer programs.
[0633] The software can be in various forms, such as system software or application software. In addition, the software can be in the form of a set of independent programs, or in the form of a program module or a part of a program module in a larger program. The software can also include modular programming in the form of object-oriented programming. After obtaining the detection data, the detection data can be automatically processed by the sequencing system 6300A, processed in response to user input, or processed in response to the request (for example, by a remote request of a communication link) proposed by another processing machine. In the specific implementation of the example, the system controller 6306 includes an analysis module 6344. In other specific implementations, the system controller 6306 does not include the analysis module 6344, but an accessible analysis module 6344 (for example, the analysis module 6344 can be hosted separately on the cloud).
[0634] The system controller 6306 can be connected to the biosensor 6312 and other components of the sequencing system 6300A via a communication link. The system controller 6306 can also be communicatively connected to an off-site system or server. The communication link can be hardwired, wired, or wireless. The system controller 6306 can receive user input or commands from the user interface 6318 and the user input device 6322.
[0635] The fluid control system 6308 includes a fluid network and is configured to direct and regulate the flow of one or more fluids through the fluid network. The fluid network may be in fluid communication with the biosensor 6312 and the fluid storage system 6314. For example, a selected fluid may be drawn from the fluid storage system 6314 and directed to the biosensor 6312 in a controlled manner, or a fluid may be drawn from the biosensor 6312 and directed toward, for example, a waste reservoir within the fluid storage system 6314. Although not shown, the fluid control system 6308 may include a flow sensor that detects the flow rate or pressure of the fluid within the fluid network. The sensor may communicate with the system controller 6306.
[0636] The temperature control system 6304 is configured to regulate the temperature of fluids at various locations within the fluid network, fluid storage system 6314, and / or biosensor 6312. For example, the temperature control system 6304 may include a thermal cycler that interfaces with the biosensor 6312 and controls the temperature of fluids flowing along reaction sites within the biosensor 6312. The temperature control system 6304 may also regulate the temperature of solid elements or components of the sequencing system 6300A or the biosensor 6312. Although not shown, the temperature control system 6304 may include sensors for detecting the temperature of the fluids or other components. The sensors may communicate with the system controller 6306.
[0637] The fluid storage system 6314 is in fluid communication with the biosensor 6312 and can store various reaction components or reactants for performing the desired reaction therein. The fluid storage system 6314 can also store fluids for washing or cleaning the fluid network and the biosensor 6312 and for diluting the reactants. For example, the fluid storage system 6314 can include various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, and non-polar solutions. In addition, the fluid storage system 6314 can also include a waste reservoir for receiving waste from the biosensor 6312. In a specific implementation including a cartridge, the cartridge can include one or more of a fluid storage system, a fluid control system, or a temperature control system. Therefore, one or more components related to those systems described herein can be accommodated within the cartridge housing. For example, the cartridge can have various reservoirs to store samples, reagents, enzymes, other biomolecules, buffer solutions, aqueous solutions, and non-polar solutions, waste, etc. Therefore, one or more of the fluid storage system, fluid control system, or temperature control system can be removably engaged with the bioassay system via the cartridge or other biosensors.
[0638] The illumination system 6316 may include a light source (e.g., one or more LEDs) and a plurality of optical components for illuminating the biosensor. Examples of light sources may include lasers, arc lamps, LEDs, or laser diodes. The optical components may include, for example, reflectors, dichroic mirrors, beam splitters, collimators, lenses, filters, wedges, prisms, mirrors, detectors, and the like. In an implementation using an illumination system, the illumination system 6316 may be configured to direct excitation light to the reaction site. As an example, a fluorophore may be excited by green wavelength light, and thus the wavelength of the excitation light may be approximately 532 nm. In one implementation, the illumination system 6316 is configured to generate illumination parallel to the surface normal of the surface of the biosensor 6312. In another implementation, the illumination system 6316 is configured to generate illumination at an angle relative to the surface normal of the surface of the biosensor 6312. In yet another implementation, the illumination system 6316 is configured to generate illumination at multiple angles, including some parallel illumination and some angled illumination.
[0639] The system socket or interface 6310 is configured to engage the biosensor 6312 in at least one of a mechanical, electrical, and fluidic manner. The system socket 6310 can hold the biosensor 6312 in a desired orientation to facilitate fluid flow through the biosensor 6312. The system socket 6310 can also include electrical contacts configured to engage the biosensor 6312 so that the sequencing system 6300A can communicate with the biosensor 6312 and / or provide power to the biosensor 6312. Additionally, the system socket 6310 can include a fluid port (e.g., a nozzle) configured to engage the biosensor 6312. In some implementations, the biosensor 6312 is removably coupled to the system socket 6310 mechanically, electrically, and fluidically.
[0640] Additionally, the sequencing system 6300A can communicate remotely with other systems or networks or with other bioassay systems 6300A. Detection data obtained by the bioassay system 6300A can be stored in a remote database.
[0641] Figure 63B It is available in Figure 63A 6 is a block diagram of a system controller 6306 used in a system of FIG. In one embodiment, the system controller 6306 includes one or more processors or modules that can communicate with each other. Each of the processors or modules may include an algorithm (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or a sub-algorithm for performing a specific process. The system controller 6306 is conceptually shown as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, the system controller 6306 may be implemented using an off-the-shelf PC with a single processor or multiple processors, where functional operations are distributed between the processors. As a further option, the modules described below may be implemented using a hybrid configuration, where certain modular functions are performed using dedicated hardware, while the remaining modular functions are performed using an off-the-shelf PC, etc. The modules may also be implemented as software modules within a processing unit.
[0642] During operation, the communication port 6350 can send data to the biosensor 6312 ( Figure 63A ) and / or subsystems 6308, 6314, 6304 ( Figure 63A ) transmits information (e.g., commands) or receives information (e.g., data) from it. In a specific implementation, the communication port 6350 can output multiple pixel signal sequences. The communication link 6334 can be connected to the user interface 6318 ( Figure 63A) receives user input and transmits data or information to the user interface 6318. Data from the biosensor 6312 or subsystems 6308, 6314, 6304 can be processed in real time during a biometric session by the system controller 6306. Additionally or alternatively, data can be temporarily stored in system memory during a biometric session and processed at a slower rate than real time or offline operation.
[0643] like Figure 63B As shown, the system controller 6306 may include a plurality of modules 6326-6348 that communicate with a main control module 6324 and a central processing unit (CPU) 6352. The main control module 6324 may communicate with a user interface 6318 ( Figure 63A Although modules 6326-6348 are shown as communicating directly with the main control module 6324, modules 6326-6348 may also communicate directly with each other, the user interface 6318, and the biosensor 6312. In addition, modules 6326-6348 may communicate with the main control module 6324 through other modules.
[0644] A plurality of modules 6326-6348 include system modules 6328-6332, 6326 that communicate with subsystems 6308, 6314, 6304, and 6316, respectively. A fluid control module 6328 can communicate with a fluid control system 6308 to control valves and flow sensors of a fluid network, thereby controlling the flow of one or more fluids through the fluid network. A fluid storage module 6330 can notify a user when the fluid volume is low or when the waste reservoir is at or near capacity. The fluid storage module 6330 can also communicate with a temperature control module 6332 so that the fluid can be stored at a desired temperature. An illumination module 6326 can communicate with an illumination system 6316 to illuminate a reaction site at a specified time during the protocol, such as after a desired reaction (e.g., binding event) has occurred. In some implementations, an illumination module 6326 can communicate with an illumination system 6316 to illuminate a reaction site at a specified angle.
[0645] The plurality of modules 6326-6348 may also include a device module 6336 that communicates with the biosensor 6312 and an identification module 6338 that determines identification information associated with the biosensor 6312. The device module 6336 may, for example, communicate with the system receptacle 6310 to confirm that the biosensor has established electrical and fluidic connection with the sequencing system 6300A. The identification module 6338 may receive a signal identifying the biosensor 6312. The identification module 6338 may use the identity of the biosensor 6312 to provide additional information to the user. For example, the identification module 6338 may determine and subsequently display a batch number, manufacturing date, or a recommended protocol for use with the biosensor 6312.
[0646] The plurality of modules 6326-6348 also includes an analysis module 6344 (also referred to as a signal processing module or signal processor) that receives and analyzes signal data (e.g., image data) from the biosensor 6312. The analysis module 6344 includes a memory (e.g., RAM or flash memory) for storing detection / image data. The detection data may include a plurality of pixel signal sequences such that a pixel signal sequence from each of the millions of sensors (or pixels) may be detected within many base calling cycles. The signal data may be stored for subsequent analysis or may be transmitted to the user interface 6318 to display desired information to the user. In some implementations, the signal data may be processed by a solid-state imager (e.g., a CMOS image sensor) before the analysis module 6344 receives the signal data.
[0647] The analysis module 6344 is configured to obtain image data from the photodetector at each sequencing cycle of the plurality of sequencing cycles. The image data is derived from an emission signal detected by the photodetector, and the image data of each sequencing cycle of the plurality of sequencing cycles is processed by a neural network-based quality scorer 6102 and / or a neural network-based base caller 218, and base calls are generated for at least some of the analytes at each sequencing cycle of the plurality of sequencing cycles. The photodetector can be part of one or more overlooking cameras (e.g., a CCD camera of Illumina's GAIIx captures an image of the clusters on the biosensor 6312 from the top), or can be part of the biosensor 6312 itself (e.g., a CMOS image sensor of Illumina's iSeq is located below the clusters on the biosensor 6312 and captures an image of the clusters from the bottom).
[0648] The output of the photodetector is a sequencing image, each of which depicts the intensity emission of a cluster and its surrounding background. The sequencing image depicts the intensity emission resulting from the incorporation of nucleotides into the sequence during sequencing. The intensity emission comes from the associated analyte and its surrounding background. The sequencing image is stored in memory 6348.
[0649] Protocol modules 6340 and 6342 communicate with the main control module 6324 to control the operation of subsystems 6308, 6314 and 6304 when performing a predetermined assay scheme. Protocol modules 6340 and 6342 may include an instruction set for instructing the sequencing system 6300A to perform a specific operation according to a predetermined scheme. As shown in the figure, the scheme module can be a sequencing by synthesis (SBS) module 6340, which is configured to issue various commands for performing a sequencing by synthesis process. In SBS, the extension of nucleic acid primers along the nucleic acid template is monitored to determine the sequence of nucleotides in the template. The basic chemical process can be polymerization (e.g., catalyzed by a polymerase) or connection (e.g., catalyzed by a ligase). In a specific polymerase-based SBS implementation, fluorescently labeled nucleotides are added to primers (thereby extending the primers) in a template-dependent manner so that the detection of the order and type of nucleotides added to the primers can be used to determine the sequence of the template. For example, in order to start the first SBS cycle, a command can be issued to deliver one or more labeled nucleotides, DNA polymerases, etc. to / by a flow cell containing a nucleic acid template array. The nucleic acid template can be located at a corresponding reaction site. Those reaction sites in which primer extension causes the labeled nucleotides to be incorporated can be detected by imaging events. During imaging events, illumination system 6316 can provide excitation light to the reaction site. Optionally, the nucleotide can further include a reversible termination attribute, and once the nucleotide is added to the primer, the reversible termination attribute terminates further primer extension. For example, a nucleotide analog with a reversible terminator portion can be added to the primer so that subsequent extension does not occur until a deblocking agent is delivered to remove the portion. Therefore, for the specific implementation using reversible termination, a command can be issued to deliver the deblocking agent to the flow cell (before or after detection occurs). One or more commands can be issued to realize the washing between each delivery step. The cycle can then be repeated n times, so that the primer is extended by n nucleotides, thereby detecting a sequence of length n. Exemplary sequencing technologies are described in, e.g., Bentley et al., Nature 456:53-59 (20063); WO 04 / 0163497, US 7,057,026, WO 91 / 066763, WO 07 / 123744, US 7,329,492, US 7,211,414, US 7,315,019, US 7,405,2631, and US 20063 / 01470630632, each of which is incorporated herein by reference.
[0650] For the nucleotide delivery step of SBS circulation, a single type of nucleotide can be delivered once, or a plurality of different nucleotide types (e.g., A, C, T, and G together) can be delivered. For a nucleotide delivery configuration in which only a single type of nucleotide is present once, different nucleotides do not need to have different labels, because they can be distinguished based on the time intervals inherent in individualized delivery. Therefore, sequencing methods or devices can use monochromatic detection. For example, an excitation source only needs to provide excitation in a single wavelength or a single wavelength range. For a nucleotide delivery configuration in which delivery causes a plurality of different nucleotides to be present in the flow cell simultaneously, the sites of incorporation of different nucleotide types can be distinguished based on the different fluorescent labels attached to the corresponding nucleotide types in the mixture. For example, four different nucleotides can be used, each nucleotide having one of four different fluorophores. In a specific implementation, four different fluorophores can be distinguished using excitation in four different regions of the spectrum. For example, four different excitation radiation sources can be used. Alternatively, less than four different excitation sources can be used, but the optical filtering of the excitation radiation from a single source can be used to produce excitation radiation of different ranges at the flow cell.
[0651] In some specific implementations, can detect in the mixture with four kinds of different nucleotides and be less than four different colors.For example, nucleotide pair can detect under the same wavelength, but based on the intensity difference of a member in the pair relative to another member, or based on the change (for example, by chemical modification, photochemical modification or physical modification) of a member causing to be compared with the signal of another member of the pair detected that obvious signal appears or disappears and distinguishes.Use the detection being less than four colors to distinguish the exemplary device and method of four kinds of different nucleotides and be described in for example following patent: U.S. patent application serial number 61 / 5363,294 and 61 / 619,63763, these patent applications are incorporated herein by reference in their entirety.The U.S. application 13 / 624,200 that submitted on September 21, 2012 is also incorporated by reference in its entirety.
[0652] A plurality of scheme modules may also include a sample preparation (or generation) module 6342, which is configured to issue commands to the fluid control system 6308 and the temperature control system 6304 for amplifying the product in the biosensor 6312. For example, the biosensor 6312 may be coupled to the sequencing system 6300A. The amplification module 6342 may issue instructions to the fluid control system 6308 to deliver necessary amplification components to the reaction chamber in the biosensor 6312. In other specific implementations, the reaction site may have included some components for amplification, such as template DNA and / or primers. After the amplification components are delivered to the reaction chamber, the amplification module 6342 may instruct the temperature control system 6304 to cycle through different temperature stages according to known amplification schemes. In some specific implementations, amplification and / or nucleotide incorporation is carried out isothermally.
[0653] The SBS module 6340 can issue a command to perform bridge PCR, in which clusters of clonal amplicons are formed on localized regions within the channels of the flow cell. After amplicons are generated by bridge PCR, the amplicons can be "linearized" to prepare single-stranded template DNA or sstDNA, and sequencing primers can be hybridized to universal sequences flanking the region of interest. For example, a reversible terminator-based sequencing-by-synthesis method can be used as described above or as follows.
[0654] Each base detection or sequencing cycle can extend sstDNA by a single base, which can be accomplished, for example, by using a mixture of modified DNA polymerase and four types of nucleotides. Different types of nucleotides can have unique fluorescent labels, and each nucleotide can also have a reversible terminator that only allows single-base incorporation to occur in each cycle. After a single base is added to the sstDNA, excitation light can be incident on the reaction site and fluorescent emission can be detected. After detection, the fluorescent label and terminator can be chemically cut from the sstDNA. Next can be another similar base detection or sequencing cycle. In this sequencing scheme, the SBS module 6340 can instruct the fluid control system 6308 to guide reagents and enzyme solutions to flow through the biosensor 6312. Exemplary reversible terminator-based SBS methods that can be used with the devices and methods described herein are described in U.S. Patent Application Publication No. 2007 / 0166705 Al, U.S. Patent Application Publication No. 2006 / 016363901 Al, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439 Al, U.S. Patent Application Publication No. 2006 / 026314714709 Al, PCT Publication No. WO 05 / 0656314, U.S. Patent Application Publication No. 2005 / 014700900 Al, PCT Publication No. WO 06 / 063B199, and PCT Publication No. WO 07 / 01470251, each of which is incorporated herein by reference in its entirety. Exemplary reagents for reversible terminator-based SBS are described in: US 7,541,444, US 7,057,026, US 7,414,14716, US 7,427,673, US 7,566,537, US 7,592,435, and WO 07 / 1463353663, each of which is incorporated herein by reference in its entirety.
[0655] In some implementations, the amplification module and the SBS module can operate in a single assay protocol, where, for example, a template nucleic acid is amplified and then sequenced within the same cartridge.
[0656] Sequencing system 6300A may also allow the user to reconfigure the assay protocol. For example, sequencing system 6300A may provide the user with an option to modify the selected protocol via user interface 6318. For example, if biosensor 6312 is determined to be used for amplification, sequencing system 6300A may request a temperature for the annealing cycle. Furthermore, sequencing system 6300A may issue a warning to the user if the user has provided user input that is generally unacceptable for the selected assay protocol.
[0657] In a specific implementation, biosensor 6312 includes millions of sensors (or pixels), each of which generates multiple pixel signal sequences during subsequent base calling cycles. Analysis module 6344 detects the multiple pixel signal sequences based on the row-by-row and / or column-by-column positions of the sensors on the sensor array and attributes them to the corresponding sensors (or pixels).
[0658] Figure 63C is a simplified block diagram of a system for analyzing sensor data (such as base call sensor output) from sequencing system 6300A. Figure 63C In an example of , the system includes a configurable processor 6346. The configurable processor 6346 can execute a base caller (e.g., a neural network-based quality scorer 6102 and / or a neural network-based base caller 218) in coordination with a runtime program executed by a central processing unit (CPU) 6352 (i.e., a host processor). The sequencing system 6300A includes a biosensor 6312 and a circulation cell. The circulation cell may include one or more blocks in which clusters of genetic material are exposed to a sequence of an analyte stream that is used to cause a reaction in the cluster to identify the bases in the genetic material. The sensor senses the reaction for each cycle of the sequence in each block of the circulation cell to provide block data. Genetic sequencing is a data-intensive operation that converts base call sensor data into a base call sequence for each cluster of genetic material sensed during the base call operation.
[0659] The system in this example includes a CPU 6352 that executes a runtime program to coordinate base calling operations, a memory 6348B for storing the sequence of the block data array, the base called reads generated by the base calling operations, and other information used in the base calling operations. In addition, in this illustration, the system includes a memory 6348A to store configuration files (or files) such as FPGA bit files and model parameters for the neural network used to configure and reconfigure the configurable processor 6346 and execute the neural network. The sequencing system 6300A may include a program for configuring the configurable processor, and in some embodiments, the reconfigurable processor, to execute the neural network.
[0660] The sequencing system 6300A is coupled to the configurable processor 6346 via a bus 6389. The bus 6389 can be implemented using high-throughput technology, such as, in one example, a bus technology compatible with the PCIe standard (Peripheral Component Interconnect Express) currently maintained and developed by the PCI-SIG (PCI Special Interest Group). Also in this example, the memory 6348A is coupled to the configurable processor 6346 via a bus 6393. The memory 6348A can be an on-board memory provided on a circuit board having the configurable processor 6346. The memory 6348A is used by the configurable processor 6346 to access working data used in base calling operations at high speed. The bus 6393 can also be implemented using high-throughput technology, such as a bus technology compatible with the PCIe standard.
[0661] Configurable processors, including field programmable gate arrays (FPGAs), coarse-grained reconfigurable arrays (CGRAs), and other configurable and reconfigurable devices, can be configured to implement various functions more efficiently or faster than would be possible using a general-purpose processor executing a computer program. Configuring a configurable processor involves compiling a functional description to produce a configuration file, sometimes called a bitstream or bitfile, and distributing the configuration file to the configurable elements on the processor. The configuration file defines the logical functions to be performed by the configurable processor by configuring the circuit to set data flow patterns, the use of distributed memory and other on-chip memory resources, the contents of lookup tables, the operation of configurable logic blocks and configurable execution units (such as multiply-accumulate units, configurable interconnects, and other elements of the configurable array). A configurable processor is reconfigurable if the configuration file can be changed in the field by changing a loaded configuration file. For example, the configuration file can be stored in volatile SRAM elements, non-volatile read-write memory elements, and combinations thereof, distributed across an array of configurable elements on a configurable or reconfigurable processor. A variety of commercially available configurable processors are suitable for base calling operations as described herein. Examples include Google's Tensor Processing Unit (TPU) TM , rack solutions (such as GX4 Rackmount Series TM 、GX9 Rackmount Series TM ), NVIDIA DGX-1 TM , Microsoft's Stratix V FPGA TM , Graphcore’s Intelligent Processor Unit (IPU) TM 、Qualcomm's Snapdragon processors TM Zeroth Platform TM、NVIDIA's Volta TM 、NVIDIA's DRIVE PX TM , NVIDIA's JETSON TX1 / TX2MODULE TM 、Intel's Nirvana TM 、Movidius VPU TM ,Fujitsu DPI TM ARM's DynamicIQ TM 、IBM TrueNorth TM , with Testa V100s TM Lambda GPU server, Xilinx Alveo TM U200, Xilinx Alveo TM U250, Xilinx Alveo TM U280, Intel / Altera Stratix TM GX2800、Intel / AlteraStratix TM GX2800 and Intel Stratix T M GX10M. In some examples, the host CPU can be implemented on the same integrated circuit as the configurable processor.
[0662] The embodiments described herein use a configurable processor 6346 to implement a neural network-based quality scorer 6102 and / or a neural network-based base caller 218. A configuration file for the configurable processor 6346 can be implemented by specifying the logic functions to be performed using a high-level description language (HDL) or a register transfer level (RTL) language specification. The specification can be compiled using resources designed for the selected configurable processor to generate the configuration file. To generate a design for an application-specific integrated circuit (ASIC) that may not be a configurable processor, the same or similar specification can be compiled.
[0663] Thus, in all embodiments described herein, alternatives to the configurable processor 6346 include a configured processor comprising a dedicated ASIC or application specific integrated circuit or group of integrated circuits, or a system on a chip (SOC) device, or a graphics processing unit (GPU) processor or a coarse-grained reconfigurable architecture (CGRA) processor, configured to perform neural network-based base calling operations as described herein.
[0664] In general, configurable processors and processors configured as described herein that are configured to perform operations of a neural network are referred to herein as neural network processors.
[0665] In this example, the configurable processor 6346 is configured by a configuration file loaded by a program executed by the CPU 6352, or by other sources that configure an array of configurable elements 6391 (e.g., configuration logic blocks (CLBs), such as lookup tables (LUTs), flip-flops, computational processing units (PMUs) and computational memory units (CMUs), configurable I / O blocks, programmable interconnects) on the configurable processor to perform base calling functions. In this example, the configuration includes data flow logic 6397, which is coupled to buses 6389 and 6393 and performs the function of distributing data and control parameters between the elements used in the base calling operation.
[0666] In addition, the configurable processor 6346 is configured with base calling execution logic 6397 to execute the neural network-based quality scorer 6102 and / or the neural network-based base caller 218. The logic 6397 includes multi-cycle execution clusters (e.g., 6379), which in this example include execution cluster 1 through execution cluster X. The number of multi-cycle execution clusters can be selected based on a trade-off between the desired throughput of the operations involved and the available resources on the configurable processor 6346.
[0667] The multi-cycle execution clusters are coupled to the dataflow logic 6397 via a dataflow path 6399 implemented using configurable interconnect and memory resources on the configurable processor 6346. In addition, the multi-cycle execution clusters are coupled to the dataflow logic 6397 via a control path 6395 implemented using, for example, configurable interconnect and memory resources on the configurable processor 6346, which provides control signals indicating the availability of the execution clusters, readiness to provide input elements for execution of operations of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218, readiness to provide trained parameters for the neural network-based quality scorer 6102 and / or the neural network-based base caller 218, readiness to provide output patches of base call classification data, and other control data for executing the neural network-based quality scorer 6102 and / or the neural network-based base caller 218.
[0668] The configurable processor 6346 is configured to use trained parameters to perform the operation of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218 to generate classification data for the sensing cycle of the base calling operation. The neural network-based quality scorer 6102 and / or the neural network-based base caller 218 are performed to generate classification data for the subject sensing cycle of the base calling operation. The operation of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218 operates on a sequence (a digital N array of block data including corresponding sensing cycles from N sensing cycles), wherein the N sensing cycles provide sensor data for different base calling operations for one base position of each operation in the time series in the example described herein. Optionally, if desired, some of the N sensing cycles may be out of order depending on the specific neural network model being executed. The number N can be any number greater than 1. In some examples described herein, a sensing cycle in the N sensing cycles represents a set of sensing cycles that is at least one sensing cycle before a subject sensing cycle and at least one sensing cycle after a subject cycle in a time sequence. Examples are described herein where the number N is an integer equal to or greater than five.
[0669] The data flow logic 6397 is configured to move tile data and at least some trained parameters of the model parameters from the memory 6348A to the configurable processor 6346 for operation of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218 using input units for a given operation, the input units including tile data for spatially aligned patches of the N arrays. The input units may be moved via direct memory access operations in one DMA operation or in smaller units that are moved during available time slots in coordination with the execution of the deployed neural network.
[0670] Block data for sensing cycles as described herein may include a sensor data array having one or more features. For example, the sensor data may include two images that are analyzed to identify one of four bases at a base position in a genetic sequence of DNA, RNA, or other genetic material. The block data may also include metadata about the images and the sensor. For example, in an embodiment of a base calling operation, the block data may include information about the alignment of the image with the cluster, such as distance-to-center information indicating how far each pixel in the sensor data array is from the center of the cluster of genetic material on the block.
[0671] During execution of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218, as described below, the chunk data may also include data generated during execution of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218, referred to as intermediate data, which may be reused rather than recalculated during execution of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218. For example, during execution of the neural network-based quality scorer 6102 and / or the neural network-based base caller 218, the data flow logic 6397 may write the intermediate data to the memory 6348A in place of the sensor data for a given patch of the chunk data array. Embodiments similar to this are described in more detail below.
[0672] As shown, a system for analyzing base call sensor output is described, the system including a memory (e.g., 6348A) accessible by a runtime program that stores block data, the block data including sensor data for blocks of sensing cycles from a base call operation. In addition, the system includes a neural network processor, such as a configurable processor 6346 that can access the memory. The neural network processor is configured to perform operation of a neural network using trained parameters to generate classification data for the sensing cycles. As described herein, the operation of the neural network operates on a sequence of N arrays of block data from corresponding sensing cycles of N sensing cycles (including a subject cycle) to generate classification data for the subject cycle. Data flow logic 908 is provided to move the block data and trained parameters from the memory to the neural network processor using input units (including data from spatially aligned patches of N arrays of corresponding sensing cycles of the N sensing cycles) for operation of the neural network.
[0673] Additionally, a system is described in which a neural network processor has access to memory and includes a plurality of execution clusters, wherein an execution cluster of the plurality of execution clusters is configured to execute a neural network. Data flow logic 6397 can access the memory and an execution cluster of the plurality of execution clusters to provide an input unit of tile data to an available execution cluster of the plurality of execution clusters, the input unit including a number N spatially aligned patches from a tile data array of a corresponding sensing cycle (including a subject sensing cycle), and cause the execution cluster to apply the N spatially aligned patches to the neural network to generate an output patch of classification data for the spatially aligned patches for the subject sensing cycle, where N is greater than 1.
[0674] Figure 64Ais a simplified diagram illustrating various aspects of a base calling operation, which includes the functionality of a runtime program executed by a host processor. In the diagram, the output of the image sensor from the circulation cell is provided on line 6400 to an image processing thread 6401, which may perform processing on the image, such as alignment and arrangement in the sensor data array of the individual blocks and resampling of the image, and may be used by a process of calculating a block cluster mask for each block in the circulation cell, which identifies pixels in the sensor data array that correspond to clusters of genetic material on the corresponding block of the circulation cell. Depending on the state of the base calling operation, the output of the image processing thread 6401 is provided on line 6402 to scheduling logic 6410 in the CPU, which routes the block data array on a high speed bus 6403 to a data cache 6404 (e.g., an SSD storage device), or on a high speed bus 6405 to neural network processor hardware 6420, such as Figure 63C 6404 for use in previously used sensing cycles. Hardware 6420 returns the classified data output by the neural network to dispatch logic 6464, which passes the information to data cache 6404 or, on line 6411, to thread 6402, which performs base calling and quality score calculations using the classified data, and may arrange the data for base called reads in a standard format. The output of thread 6402, which performs base calling and quality score calculations, is provided on line 6412 to thread 6403, which aggregates the base called reads, performs other operations such as data compression, and writes the resulting base called output to a designated destination for client utilization.
[0675] In some embodiments, the host may include a thread (not shown) that performs final processing of the output of hardware 6420 to support the neural network. For example, hardware 6420 may provide output of classification data from the final layer of a multi-cluster neural network. The host processor may perform an output activation function, such as a softmax function, on the classification data to configure the data for use by base calling and quality scoring thread 6402. Additionally, the host processor may perform input operations (not shown), such as batch normalization of the block data before input to hardware 6420.
[0676] Figure 64B The processor 6346 can be configured such as Figure 63C A simplified diagram of the configuration of a configurable processor. Figure 64B In the embodiment, the configurable processor 6346 includes an FPGA with multiple high-speed PCIe interfaces. The FPGA is configured with a wrapper 6490, which includes a reference Figure 63CThe data flow logic 6397 described above is shown in Figure 6. The wrapper 6490 manages the interface and coordination with the runtime program in the CPU via the CPU communication link 6477 and manages communication with the onboard DRAM 6499 (e.g., memory 6348A) via the DRAM communication link 6497. The data flow logic 6397 in the wrapper 6490 provides patch data retrieved by traversing the tile data array on the onboard DRAM 6499 for N cycles to the cluster 6485 and retrieves process data 6487 from the cluster 6485 for delivery back to the onboard DRAM 6499. The wrapper 6490 also manages data transfer between the onboard DRAM 6499 and host memory for both the input array of tile data and the output patch of classification data. The wrapper transmits the patch data on line 6483 to the assigned cluster 6485. The wrapper provides trained parameters such as weights and biases to the cluster 6485 retrieved from the onboard DRAM 6499 on line 6481. The wrapper provides configuration and control data to cluster 6485 on line 6479, which the cluster provides from or generates in response to the runtime program on the host via CPU communication link 6477. The cluster may also provide status signals to wrapper 6490 on line 6489, which are used in conjunction with control signals from the host to manage the traversal of the tile data array to provide spatially aligned patch data and to execute multi-recurrent neural networks on the patch data using the resources of cluster 6485.
[0677] As described above, multiple clusters may exist on a single configurable processor managed by the wrapper 6490, the multiple clusters being configured to execute on corresponding patches of the multiple patches of block data. Each cluster may be configured to provide classification data for base calls in a sensing cycle of a subject using the block data of the multiple sensing cycles described herein.
[0678] In an example of a system, model data (including kernel data, such as filter weights and biases) can be sent from the host CPU to a configurable processor so that the model can be updated based on the number of cycles. As a representative example, a base calling operation may include approximately hundreds of sensing cycles. In some embodiments, the base calling operation may include double-ended reads. For example, model training parameters can be updated every 20 cycles (or other number of cycles), or according to an update pattern implemented for a particular system and neural network model. In some embodiments including double-ended reads, where the sequence of a given string in a genetic cluster on a block includes a first portion extending downward (or upward) along the string from a first end and a second portion extending upward (or downward) along the string from a second end, the trained parameters can be updated in the transition from the first portion to the second portion.
[0679] In some examples, image data for multiple cycles of sensory data for a tile can be sent from the CPU to a wrapper 6490. The wrapper 6490 can optionally perform some preprocessing and conversion on the sensory data and write the information to the onboard DRAM 6499. The input tile data for each sensing cycle can include an array of sensor data, comprising approximately 4000×3000 pixels or more per tile per sensing cycle, with two features representing the colors of the two images for the tile, and one or two bytes per pixel per feature. For embodiments where the number N is three sensing cycles to be used in each run of the multi-cycle neural network, the tile data array for each run of the multi-cycle neural network can consume on the order of hundreds of megabytes per tile. In some embodiments of the system, the tile data also includes an array of DFC data stored once per tile, or other types of metadata about the sensor data and the tile.
[0680] In operation, when a multi-cycle cluster is available, the encapsulator assigns patches to the cluster. The encapsulator retrieves the next patch of data for the chunk during a traversal of the chunk and sends it to the assigned cluster along with appropriate control and configuration information. The cluster can be configured with sufficient memory on the configurable processor to hold data patches being processed in-place, including patches from multiple cycles in some systems, as well as data patches to be processed when the current patch is finished processing using ping-pong buffering or raster scanning techniques in various embodiments.
[0681] When the assigned cluster completes its execution of the neural network for the current patch and produces an output patch, it will signal the packager. The packager will read the output patch from the assigned cluster, or alternatively, the assigned cluster pushes the data to the packager. The packager will then assemble the output patch for the processed block in DRAM 6499. When processing of the entire block is complete and the output patch of data has been transferred to DRAM, the packager sends the processed output array of the block back to the host / CPU in the specified format. In some embodiments, the onboard DRAM 6499 is managed by memory management logic in the packager 6490. The runtime program can control the sequencing operation to complete the analysis of all arrays of block data for all cycles in the run in a continuous stream, thereby providing real-time analysis.
[0682] Technical improvements and terminology
[0683] Base calling involves binding or attaching a fluorescently labeled tag to the analyte. The analyte can be a nucleotide or oligonucleotide, and the tag can be used for a specific nucleotide type (A, C, T, or G). Excitation light is directed toward the tagged analyte, and the tag emits a detectable fluorescent signal or intensity emission. The intensity emission indicates the photons emitted by the excited tag chemically attached to the analyte.
[0684] Throughout this application, including the claims, when phrases such as or similar to "an image, image data, or image region depicting the intensity emissions of an analyte and its surrounding background" are used, they refer to the intensity emissions of a label attached to the analyte. One skilled in the art will appreciate that the intensity emissions of an attached label represent or are equivalent to the intensity emissions of the analyte to which the label is attached and are therefore used interchangeably. Similarly, properties of an analyte refer to properties of a label attached to the analyte or properties of the intensity emissions from the attached label. For example, the center of an analyte refers to the center of the intensity emissions emitted by a label attached to the analyte. In another example, the surrounding background of an analyte refers to the surrounding background of the intensity emissions emitted by a label attached to the analyte.
[0685] All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and web pages, regardless of the format of these literature and similar materials, are expressly incorporated by reference in their entirety. If one or more of the incorporated literature and similar materials differs from or contradicts this application, including but not limited to defined terms, term usage, described techniques, etc., this application shall prevail.
[0686] The disclosed technology uses neural networks to improve the quality and quantity of nucleic acid sequence information that can be obtained from a nucleic acid sample (such as a nucleic acid template or its complementary sequence, e.g., a DNA or RNA polynucleotide or other nucleic acid sample). Thus, certain embodiments of the disclosed technology provide higher throughput polynucleotide sequencing, e.g., higher DNA or RNA sequence data collection rates, higher sequence data collection efficiencies, and / or lower costs for obtaining such sequence data, relative to previously available methods.
[0687] The disclosed technology uses a neural network to identify the centers of solid-phase nucleic acid clusters and analyzes the optical signals generated during sequencing of such clusters to unambiguously distinguish between adjacent, contiguous, or overlapping clusters in order to assign sequencing signals to a single discrete source cluster. Thus, these and related embodiments allow for the retrieval of meaningful information, such as sequence data, from regions of high-density cluster arrays where usable information was previously unavailable due to the confounding effects of overlapping or very closely spaced adjacent clusters, including overlapping signals emanating therefrom (e.g., as used in nucleic acid sequencing).
[0688] As described in more detail below, in some specific implementation, provide the composition that comprises solid carrier, this solid carrier has one or more nucleic acid clusters that are fixed thereto as provided herein.Each cluster comprises the immobilized nucleic acid of a plurality of identical sequences and has identifiable center, this identifiable center has the detectable center mark that provided herein, can distinguish the immobilized nucleic acid in the surrounding area in identifiable center and cluster by this detectable center mark.This paper has also described the method for making and using this type of cluster with identifiable center.
[0689] The disclosed embodiments will find use in many situations where advantage is gained from the ability to identify, determine, annotate, record, or otherwise assign the location of a substantially central location within a cluster, such as high-throughput nucleic acid sequencing, the development of image analysis algorithms for assigning optical or other signals to discrete source clusters, and other applications where identifying the center of an immobilized nucleic acid cluster is desirable and beneficial.
[0690] In certain embodiments, the present invention contemplates methods involving high-throughput nucleic acid analysis, such as nucleic acid sequence determination (e.g., "sequencing"). Exemplary high-throughput nucleic acid analyses include, but are not limited to, de novo sequencing, resequencing, whole genome sequencing, gene expression analysis, gene expression monitoring, epigenetic analysis, genomic methylation analysis, allele-specific primer extension (APSE), genetic diversity analysis, whole genome polymorphism discovery and analysis, single nucleotide polymorphism analysis, hybridization-based sequence determination methods, and the like. One skilled in the art will appreciate that a wide variety of nucleic acids can be analyzed using the methods and compositions of the present invention.
[0691] Although the specific implementation of the present invention is described with respect to nucleic acid sequencing, they are applicable to any field of image data collected at different time points, spatial positions or other times or physical perspectives. For example, the methods and systems described herein can be used in the field of molecule and cell biology, wherein image data from microarrays, biological specimens, cells, organisms, etc. are collected at different time points or perspectives and analyzed. Images can be obtained using any number of techniques known in the art, including but not limited to fluorescence microscopy, optical microscopy, confocal microscopy, optical imaging, magnetic resonance imaging, tomography, etc. For another example, the methods and systems described herein can be applied, wherein image data obtained by monitoring, airborne or satellite imaging technology, etc. are collected at different time points or perspectives and analyzed. The method and system are particularly useful for analyzing images obtained for a visual field, wherein the analyte being observed remains in the same position relative to each other in the visual field. However, the analyte may have different features in a separate image, for example, the analyte may look different in a separate image of the visual field. For example, an analyte may appear different in terms of its color detected in different images, a change in signal intensity for a given analyte detected in different images, or even the appearance of a signal for a given analyte detected in one image and the disappearance of a signal for that analyte detected in another image.
[0692] Examples described herein can be used for various biological or chemical processes and systems for academic or commercial analysis. More specifically, examples described herein can be used for various processes and systems for expecting to detect events, attributes, qualities or features indicating a given reaction. For example, examples described herein include light detection equipment, biosensors and components thereof, and bioassay systems operating together with biosensors. In some examples, equipment, biosensors and systems can include a flow cell and one or more light sensors that are coupled together with a substantially integrated structure (removably or fixedly).
[0693] These devices, biosensors and bioassay systems can be configured to perform multiple specified reactions that can be detected individually or together. These devices, biosensors and bioassay systems can be configured to perform multiple cycles, wherein the multiple specified reactions occur synchronously. For example, these devices, biosensors and bioassay systems can be used for sequencing the dense array of DNA features through the iterative cycle of enzyme manipulation and light or image detection / collection. Therefore, these devices, biosensors and bioassay systems (for example, via one or more boxes) may include one or more microfluidic channels that deliver the reagents or other reaction components in the reaction solution to the reaction site of these devices, biosensors and bioassay systems. In some examples, the reaction solution may be substantially acidic, such as having a pH of less than or equal to about 5, or less than or equal to about 4, or less than or equal to about 3. In some other examples, the reaction solution may be substantially alkaline / alkaline, such as having a pH of greater than or equal to about 8, or greater than or equal to about 9, or greater than or equal to about 10. As used herein, the term "acidity" and grammatical variations thereof refer to a pH value of less than about 7, and the terms "alkalinity," "alkaline," and grammatical variations thereof refer to a pH value of greater than about 7.
[0694] In some examples, the reaction sites are provided or spaced in a predetermined manner, such as in a uniform or repeating pattern. In some other examples, the reaction sites are randomly distributed. Each reaction site can be associated with one or more light guides and one or more light sensors that detect light from the associated reaction site. In some examples, the reaction sites are located in reaction recesses or reaction chambers, which can at least partially isolate the designated reactions therein.
[0695] As used herein, a "specified reaction" includes a change in at least one of the chemical, electrical, physical, or optical properties (or mass) of a chemical or biological substance of interest (e.g., an analyte of interest). In a specific example, the specified reaction is a positive binding event, for example, the binding of a fluorescently labeled biomolecule to the analyte of interest. More generally, the specified reaction can be a chemical transformation, a chemical change, or a chemical interaction. The specified reaction can also be a change in an electrical property. In a specific example, the specified reaction includes binding a fluorescently labeled molecule to the analyte. The analyte can be an oligonucleotide, and the fluorescently labeled molecule can be a nucleotide. The specified reaction can be detected when excitation light is directed to the oligonucleotide with the labeled nucleotide and the fluorophore emits a detectable fluorescent signal. In an alternative example, the detected fluorescence is the result of chemiluminescence or bioluminescence. The specified reaction can also be, for example, to increase fluorescence (or ) FRET, reducing FRET by separating the donor and acceptor fluorophores, increasing fluorescence by separating the quencher from the fluorophore, or decreasing fluorescence by colocalizing the quencher and fluorophore.
[0696] As used herein, "reaction solution", "reaction component" or "reactant" include any material that can be used to obtain at least one specified reaction. For example, possible reaction component includes, for example, reagent, enzyme, sample, other biomolecules and buffer. Reaction component can be delivered to the reaction site in the solution and / or be fixed at the reaction site. Reaction component can directly or indirectly interact with another substance, such as the analyte of interest that is fixed at the reaction site. As mentioned above, the reaction solution can be substantially acidic (that is, including relatively high acidity) (for example, having a pH less than or equal to about 5, a pH less than or equal to about 4, or a pH less than or equal to about 3) or substantially alkaline / alkaline (that is, including relatively high alkalinity / alkalinity) (for example, having a pH greater than or equal to about 8, a pH greater than or equal to about 9, or a pH greater than or equal to about 10).
[0697] As used herein, the term "reaction site" is a local area where at least one specified reaction can occur. The reaction site can include a reaction structure or a support surface of a substrate on which a substance can be fixed. For example, the reaction site can include a surface (which can be located in a channel of a flow cell) of a reaction structure with a reactive component (such as a nucleic acid colony thereon). In some such examples, the nucleic acid in the colony has the same sequence, for example, a cloned copy of a single-stranded or double-stranded template. However, in some examples, the reaction site can only include a single nucleic acid molecule, such as a single-stranded or double-stranded form.
[0698] A plurality of reaction sites can be randomly distributed along the reaction structure or arranged in a predetermined manner (for example, arranged side by side in a matrix, such as in a microarray). The reaction site may also include a reaction chamber or reaction groove, which at least partially defines a spatial region or volume configured to separate a specified reaction. As used herein, the term "reaction chamber" or "reaction groove" includes a defined spatial region (which is generally in communication with a flow channel fluid) of a support structure. The reaction groove may be at least partially separated from the surrounding environment of other or spatial regions. For example, a plurality of reaction grooves may be separated from each other by a common wall such as a detection surface. As a more specific example, the reaction groove can be a nanopore comprising a dent, pit, hole, groove, cavity or depression defined by the inner surface of the detection surface, and has an opening or pore (that is, open) so that the nanopore can be in communication with the flow channel fluid.
[0699] In some examples, the size and shape of the reaction groove of the reaction structure are set relative to the solid (including semi-solid) so that the solid can be fully or partially inserted therein. For example, the size and shape of the reaction groove can be set to accommodate capture beads. The capture beads can have cloned amplified DNA or other substances thereon. Alternatively, the size and shape of the reaction groove can be set to accommodate an approximate number of beads or solid substrates. For another example, the reaction groove can be filled with a porous gel or substance that is configured to control diffusion or filter the fluid or solution that can flow into the reaction groove.
[0700] In some examples, a light sensor (e.g., a photodiode) is associated with a corresponding reaction site. The light sensor associated with the reaction site is configured to detect light emission from the associated reaction site via at least one light guide when a specified reaction has occurred at the associated reaction site. In some cases, multiple light sensors (e.g., several pixels of a light detection or camera device) may be associated with a single reaction site. In other cases, a single light sensor (e.g., a single pixel) may be associated with a single reaction site or with a group of reaction sites. Other features of the light sensors, reaction sites, and biosensors may be configured so that at least some of the light is detected directly by the light sensor without being reflected.
[0701] As used herein, "biological or chemical substance" includes biomolecules, samples of interest, analytes of interest, and other compounds. Biological or chemical substances can be used to detect, identify, or analyze other compounds, or as intermediates for studying or analyzing other compounds. In a specific example, biological or chemical substances include biomolecules. As used herein, "biomolecules" include at least one of biopolymers, nucleosides, nucleic acids, polynucleotides, oligonucleotides, proteins, enzymes, polypeptides, antibodies, antigens, ligands, receptors, polysaccharides, carbohydrates, polyphosphates, cells, tissues, organisms, or their fragments, or analogs or mimetics of any other biologically active compounds such as the aforementioned substances. In another example, biological or chemical substances or biomolecules include enzymes or reagents for detecting the product of another reaction in a coupled reaction, such as enzymes or reagents, such as enzymes or reagents for detecting pyrophosphate in a pyrophosphate sequencing reaction. Enzymes and reagents that can be used for pyrophosphate detection are described in, for example, U.S. Patent Publication 2005 / 0244870A1, which is incorporated herein by reference in its entirety.
[0702] Biomolecules, samples, and biological or chemical substances can be naturally occurring or synthetic and can be suspended in a solution or mixture within a reaction groove or region. Biomolecules, samples, and biological or chemical substances can also be bound to a solid phase or gel material. Biomolecules, samples, and biological or chemical substances can also include pharmaceutical compositions. In some cases, the biomolecules, samples, and biological or chemical substances of interest can be referred to as targets, probes, or analytes.
[0703] As used herein, "biosensor" includes a device having a reaction structure with multiple reaction sites, which is configured to detect a specified reaction occurring at or near the reaction site. The biosensor may include a solid-state light...
Claims
1. A computer-implemented method of determining a quality score for a base call, the method comprising: processing input data for clusters through a neural network-based base caller and generating an alternative representation of the input data, wherein the input data for clusters is image data depicting intensity emissions of one or more clusters and their surrounding background, and wherein the alternative representation comprises a convolutional representation or a hidden representation; processing the alternative representation through a softmax function of an output layer to produce an output comprising a classification score, wherein the classification score in the output identifies the likelihood of bases being A, C, T, and G being incorporated into a particular one of the clusters; base calling the cluster based on the output; and determining a quality score for the called base from the likelihood identified by the output by applying a quantification scheme, wherein the quantization scheme is calibrated for training of a base caller based on the neural network, comprising quantifying classification scores of bases called by the neural network-based base caller in response to processing training data during the training; selecting a set of quantized classification scores; for each quantized classification score in the set, determining a base call error rate by comparing a predicted base call having the quantized classification score to a corresponding baseline true base call; determining a fit between the quantified classification score and its base calling error rate; and The quality score is associated with the quantized classification score based on the fit. 2 . The computer-implemented method of claim 1 , wherein the fitting exhibits a corresponding relationship between the quantified classification score and the quality score.
3. The computer-implemented method of any one of claims 1 to 2, wherein the set of quantized classification scores comprises a subset of the classification scores for predicted base calls generated by the neural network-based base caller in response to processing the training data during the training, and Wherein the classification score is a real number.
4. The computer-implemented method of claim 1 , wherein the set of quantized classification scores comprises all of the classification scores for predicted base calls generated by the neural network-based base caller in response to processing the training data during the training, and Wherein the classification score is a real number.
5. The computer-implemented method of claim 3, wherein the classification scores are exponentially normalized softmax scores that sum to 1 and are produced by a softmax output layer of the neural network-based base caller.
6. The computer-implemented method of claim 5, wherein the set of quantized classification scores is based on a set of quantized classification scores defined as And the selection formula of the softmax score applied to the exponential normalization is selected.
7. The computer-implemented method of claim 5, wherein the set of quantized classification scores is based on a set of quantized classification scores defined as And the selection formula of the softmax score applied to the exponential normalization is selected.
8. A computer-implemented method of assigning quality scores to base calls, the method comprising: processing input data for clusters through a neural network-based base caller and generating an alternative representation of the input data, wherein the input data for clusters is image data depicting intensity emissions of one or more clusters and their surrounding background, and wherein the alternative representation comprises a convolutional representation or a hidden representation; processing the alternative representation through a softmax function of an output layer to produce an output comprising a classification score, wherein the classification score in the output identifies the likelihood of bases being A, C, T, and G being incorporated into a particular one of the clusters; base calling the cluster based on the output; and determining a quality score for the called base from the likelihood identified by the output by applying a quantification scheme, wherein the quantization scheme is calibrated for training of a base caller based on the neural network, comprising quantifying classification scores of bases called by the neural network-based base caller in response to processing training data during the training; selecting a set of quantized classification scores; for each quantized classification score in the set, determining a base call error rate by comparing a predicted base call having the quantized classification score to a corresponding baseline true base call; determining a fit between the quantified classification score and its base calling error rate; and correlating the quality score with the quantized classification score based on the fit to generate a correlation; Based on the correlation, the quality scores are assigned to bases called by the neural network-based base caller during inference.
9. The computer-implemented method of claim 8, further comprising: assigning the quality scores based on applying a quality score correspondence scheme to the bases called by the neural network-based base caller during the inferring; and wherein the quality score mapping scheme maps a range of classification scores generated by the neural network-based base caller in response to processing inference data during the inference to corresponding quantized classification scores in the set.
10. The computer-implemented method of any one of claims 8 to 9, further comprising: During the inference, base calling is stopped for analytes with a quality score below a set threshold for the current base calling cycle.
11. The computer-implemented method of claim 8, further comprising: During the inference, base calling is stopped for analytes whose average quality scores are below a set threshold after subsequent base calling cycles.
12. The computer-implemented method of claim 8, wherein a sample size used to compare the predicted base calls to the corresponding benchmark true base calls is specific to each quantized classification score.
13. The computer-implemented method of claim 8, wherein the fit is determined using a regression model.
14. The computer-implemented method of claim 8, further comprising: For each quantified classification score, the base call accuracy is determined by comparing its predicted base call with the corresponding ground truth base call; as well as The fit between the quantified classification score and its base calling accuracy is determined.
15. The computer-implemented method of claim 8, wherein the corresponding benchmark true base calls are derived from well-characterized human and non-human samples sequenced on a variety of sequencing instruments, sequencing chemistries, and sequencing protocols.
Citation Information
Patent Citations
Emulsion compositions
US20050064460A1
Methods of amplifying and sequencing nucleic acids
US20050130173A1
Nucleic acid sequencing using microsphere arrays
US20050244870A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Modified nucleotides
US20070166705A1