Machine learning model for detecting air bubbles within a nucleotide sample slide for sequencing
By using a machine learning model to detect air bubbles in a nucleic acid sequencing system, the accuracy and platform adaptability issues of air bubble detection in existing technologies are resolved, thereby improving sequencing efficiency and data quality.
Patent Information
- Application Number
- CN202280021725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-02
- Filing Date
- 2022-03-23
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-03-23
AI Technical Summary
Existing nucleic acid sequencing systems cannot effectively detect the presence of air bubbles, resulting in low base detection accuracy. This necessitates inefficient resequencing and reanalysis, and the detection methods are limited to specific hardware, making them unsuitable for use in dry sequencing platforms.
Using machine learning models based on base detection and quality index data, air bubbles are detected in nucleotide sample slides. Convolutional neural networks and other models are used for bubble classification and identification. Cross-platform application requires no additional hardware.
It improves the accuracy and flexibility of bubble detection, reduces the need for resequencing, improves sequencing efficiency and data quality, and is suitable for various sequencing platforms.
Smart Images

Figure CN117043867B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 170,072, filed April 2, 2021, the entire contents of which are incorporated herein by reference. Background Technology
[0003] In recent years, biotechnology companies and research institutions have improved the hardware and software platforms used for sequencing and analyzing nucleotides. For example, some existing nucleic acid sequencing systems determine the individual nucleobases of nucleic acid sequences using conventional Sanger sequencing. In contrast, some existing systems determine such nucleobase sequences by performing sequencing-by-synthesis (SBS). By using SBS, existing systems can monitor the parallel synthesis of thousands, tens of thousands, or more nucleic acid polymers to detect more accurate bases from a larger base detection dataset and capture additional sequencing information. In some cases, existing systems synthesize oligonucleotides in monoclonal colonies within the wells of a nucleotide sample slide, such as a flow cell. For example, after a camera captures an image of a fluorescent tag (which illuminates the color from the nucleobases incorporated into such oligonucleotides), some existing systems send the image data to a device with sequencing data analysis software to analyze the base detection in the image data and determine the nucleobase sequence of the nucleic acid polymer (e.g., the gene coding region of the nucleic acid polymer).
[0004] Despite these advances in sequencing, existing nucleic acid sequencing systems exhibit several technical drawbacks, such as suppressing the accuracy of base detection and error detection, requiring inefficient resequencing and reanalysis of nucleotide samples, and limiting error detection to specific hardware on the sequencing equipment. In fact, existing systems often perform inaccurate base detection or capture unreliable image data because fluids and gases flowing through the sequencing equipment or slide can introduce irregularities beneath the image data. For example, air bubbles (e.g., air or oil bubbles) in the nucleotide sample slide can interfere with, introduce noise within, or otherwise cause data quality problems in the data features from such image data (used for base detection). Such bubbles not only distort the data features of base detection but also inhibit or slow down run quality or yield. Despite the problems caused by bubbles, existing nucleic acid sequencing systems and existing sequencing data analysis software often lack effective means of detecting bubbles.
[0005] Existing nucleic acid sequencing systems often inefficiently resequence and reanalyze nucleotide samples due in part to errors or other sequencing errors caused by bubbles. In particular, existing systems and software often perform or consume additional processing, computation, storage resources, and time to generate quality data to correct data affected by bubble interference. To illustrate, a sequencing run can suffer from many problem types, such as failed sequencing reactions, contamination, poor sample loading, or the presence of bubbles. Because existing systems often cannot identify the presence of bubbles or distinguish bubble interference from other errors, such systems often require users to repeat sequencing runs before successfully identifying problems.
[0006] While basic mechanical methods for detecting bubbles have been developed or contemplated, such detection methods are inefficient and can be limited to particular platform types. For example, existing nucleic acid sequencing systems often require additional information about a sequencing run to identify the presence of bubbles or other sources of sequencing errors. More specifically, conventional nucleic acid sequencing systems that flow fluid through a tubing to a cartridge often require additional hardware to capture data indicative of the presence of bubbles. For example, existing systems often require additional tubing cameras, tubing detectors, or other types of sensors. In some cases, such systems use ultrasonic or capacitive sensing detectors to identify bubbles passing through the tubing. But such local hardware on a sequencing device is limited to wet platforms with tubing, and requires additional processing, storage, and analysis resources to implement such bubble detection methods.
[0007] In addition to the inefficiencies of existing mechanisms for detecting bubbles in wet sequencing platforms, some such bubble detection methods are limited to particular hardware on a sequencing device. As noted above, some conventional nucleic acid sequencing systems attempt to detect bubbles by utilizing hardware-based bubble detectors. Even if some conventional nucleic acid sequencing systems can include sensors in the tubing or other components to detect bubbles, such detection hardware is not only expensive but also not feasible in dry sequencing platforms. For example, dry sequencing platforms often perform fluidic operations on disposable consumables that lack tubing to pool fluid into the consumable. Such dry sequencing platforms cannot utilize dedicated bubble detection sensors, or such sensors are impractical due to the need for expensive sequencing devices or extensive redesign of consumable nucleotide sample slides. SUMMARY
[0008] The present disclosure describes one or more embodiments of systems, methods, and non-transitory computer-readable storage media that provide benefits and / or address one or more of the above-described problems in the art. For example, the disclosed systems use machine learning models to accurately and efficiently detect when bubbles are affecting a nucleic acid sequencing run based on data captured (or derived from) during base callouts during such sequencing runs. To illustrate, the disclosed systems can receive, from a sequencing platform during a sequencing cycle, data identifying nucleobase callouts and data identifying quality metrics for such nucleobase callouts. Based on particular nucleobase callouts and threshold markings for the quality metrics, a machine learning model can detect the presence of bubbles in a nucleotide sample slide. By using callout data and quality metrics, the disclosed systems can use readily available sequencing data in a platform-agnostic approach to detect bubbles using uniquely trained machine learning models.
[0009] In some cases, the disclosed systems use machine learning models that are trained to identify bubbles within particular portions or cells (e.g., blocks) of a nucleotide sample slide (e.g., flow cell) during a sequencing cycle. In addition to simply detecting the presence of bubbles, in some examples, the disclosed systems can also classify different detected bubbles, such as oil bubbles, air bubbles, or ghost bubbles, or other outputs identified during sequencing, such as block registration failures and dropped blocks.
[0010] Additional features and advantages of the one or more embodiments of the present disclosure will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of such exemplary embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0011] Various embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:
[0012] Figure 1 An environment in which a bubble detection system according to one or more embodiments of the present disclosure can operate is illustrated.
[0013] Figure 2 An overview diagram of a bubble detection system that detects the presence of bubbles according to one or more embodiments of the present disclosure is illustrated.
[0014] Figure 3 An overview diagram of a bubble detection system operating on single-channel, dual-channel, and four-channel sequencing data according to one or more embodiments of the present disclosure is illustrated.
[0015] Figures 4A-4C An example chart illustrating data features corresponding to different error classifications according to one or more embodiments of the present disclosure is illustrated.
[0016] Figure 5 An exemplary bubble detection machine learning model is shown in accordance with one or more embodiments of the present disclosure.
[0017] Figures 6A-6C A bubble detection system training a bubble detection machine learning model is shown in accordance with one or more embodiments, along with an exemplary spatial image having bubbles within a flow cell.
[0018] Figure 7 A series of actions for detecting the presence of bubbles is shown in accordance with one or more embodiments of the present disclosure.
[0019] Figure 8 A block diagram of an exemplary computing device is shown in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] The present disclosure describes one or more embodiments of a bubble detection system that utilizes a machine learning model to detect the presence of bubbles within a nucleotide sample slide based on data captured (or derived from) during a nucleic acid sequencing run. In some embodiments, for example, the bubble detection system accesses or receives base call data for nucleobase calls during a sequencing cycle and quality data identifying quality metrics that estimate error in such nucleobase calls during the sequencing cycle. Such call data and quality data can be specific to a nucleotide sample slide, such as a flow cell or a portion of a slide. The bubble detection system determines, from the call data and quality data, a subpopulation of nucleobase calls (e.g., a subpopulation of adenine and guanine base calls) and a subpopulation of nucleotide calls that satisfy a threshold quality value. Based on these subpopulations of data as input, the bubble detection system utilizes a machine learning model to detect the presence of bubbles within the nucleotide sample slide. In some such embodiments, such bubble detection machine learning model classifies the type of bubble detected.
[0021] As just noted, in some embodiments, the bubble detection system receives call data that includes nucleobase calls for a nucleic acid polymer sequencing cycle. Typically, the bubble detection system receives call data identifying nucleobase calls for each sequencing cycle. The bubble detection system can receive call data organized or packaged according to various types of data. For example, the bubble detection system can receive call data organized on single channel data, dual channel data, or quadruple channel data. In any case, the bubble detection system can receive and utilize call data from various types of sequencing platforms.
[0022] As further noted above, the bubble detection system also receives quality data including a quality indicator that estimates an error in the called nucleobases of the cycle. In some embodiments, the quality indicator indicates a base calling accuracy for the nucleotide sample slide. For example, the quality indicator can include a value that indicates a probability of an incorrect base call. In one or more embodiments, the quality indicator includes a quality score (or Q-score) that indicates a probability of an incorrect base call for a portion of the nucleotide sample slide is 1 in 100 for a Q20 score, 1 in 1,000 for a Q30 score, 1 in 10,000 for a Q40 score, and so on. But the bubble detection system flexibly receives any number of quality indicators as part of determining the presence of a bubble.
[0023] In some embodiments, based on the called data, the bubble detection system determines a subset of the called nucleobases that correspond to at least one nucleobase. For example, in certain implementations, the bubble detection system determines a proportion of adenine calls, thymine calls, cytosine calls, or guanine calls. In one example, the bubble detection system determines a proportion or percentage of base calls that include an adenine call in each cycle, and a proportion or percentage of base calls that include a thymine call in each cycle. Thus, in certain implementations, the bubble detection system determines a percentage (or other subset) of called nucleobases that correspond to adenine and a percentage (or other subset) of called nucleobases that correspond to guanine within a particular portion of the nucleotide sample slide.
[0024] Based on the quality data, in certain cases, the bubble detection system can also determine a subset of the called nucleobases that meet a threshold quality indicator of the quality indicator. In some embodiments, the bubble detection system determines a threshold quality indicator. For example, the bubble detection system can determine that the threshold quality indicator for base calls in a cycle is equal to Q30, and that corresponds to a 99.9% accuracy or a 1 in 1,000 chance that a given base call is incorrect. The bubble detection system also determines a proportion or percentage of base calls that meet the determined threshold quality indicator. Specifically, the bubble detection system compares the quality indicators from the received quality data to the threshold quality indicator. Thus, in certain implementations, the bubble detection system determines a percentage (or other subset) of called nucleobases that meet the threshold quality indicator within a particular portion of the nucleotide sample slide.
[0025] After determining the relevant subsets of nucleobase calls, in some cases, the bubble detection system generates an input matrix for the bubble detection machine learning model that includes the first subset of nucleobase calls corresponding to at least one nucleobase and the second subset of nucleobase calls that meet the threshold quality metric. More specifically, in one example, the bubble detection system uses the subset of adenine calls, the subset of guanine calls, and the subset of nucleobase calls that meet the threshold quality metric (e.g., for each cycle within the total number of sequencing cycles) to compose the input matrix. The bubble detection system can accommodate various input sizes by adjusting the input matrix based on the number of sequencing cycles. For example, in one embodiment, the input matrix includes three one-dimensional input channels of length N, where the three input channels include the subset of adenine calls, the subset of guanine calls, and the second subset of nucleobase calls that meet the threshold quality metric, and N is equal to the number of sequencing cycles.
[0026] Regardless of the input form, the bubble detection system can use the bubble detection machine learning model to detect the presence of bubbles within the nucleotide sample slide based on the subsets of call data and quality data. To detect the presence of such bubbles, the bubble detection system can utilize various types of machine learning models. For example, in some embodiments, the bubble detection system utilizes a neural network, such as a convolutional neural network (CNN), to detect bubbles. In other embodiments, the bubble detection system utilizes other types of machine learning models to detect bubbles. For example, in some implementations, the bubble detection system implements a support vector machine (SVM) or an adaptive boosting machine learning model.
[0027] As described above, the bubble detection system provides several technical benefits and technical improvements over conventional nucleic acid sequencing systems and corresponding sequencing data analysis software. Specifically, the bubble detection system can improve the accuracy with which existing nucleic acid sequencing systems or corresponding software detect the presence of bubbles that interfere with sequencing. The disclosed bubble detection system introduces the first machine learning model in its class that detects bubbles within a nucleotide sample slide that are not matched by existing technology or the art. As described above, existing systems either do not directly detect bubbles that interfere with sequencing or use mechanical sensors to detect bubbles that are limited to specific platforms. Unlike such existing systems, the disclosed bubble detection system utilizes a machine learning model that is trained to accurately detect bubbles within a nucleotide sample slide based on a unique analysis of available data (i.e., call data that identifies nucleobase calls and quality data that identifies quality metrics for such nucleobase calls). By relying on call data and quality data, the bubble detection system can accurately detect the presence of bubbles within a nucleotide sample slide (and sometimes identify the type of bubble) using a trained bubble detection machine learning model. Unlike conventional and mechanical bubble detection methods, the bubble detection system can apply its machine learning model across various sequencing platforms by using readily available call data and quality data.
[0028] In addition to a new and accurate bubble detection method, in some embodiments, the bubble detection system can accurately detect the presence of bubbles within particular portions of a nucleotide sample slide (e.g., within a zone of a flow cell or within a group of zones of a flow cell) and corresponding call data that is affected by the bubble. More specifically, in certain instances, the bubble detection system utilizes a bubble detection machine learning model that passes call data and quality data specific to a portion of a slide to automatically detect portions of a nucleotide sample slide that are affected by a bubble. By specifying which portions of a nucleotide sample slide have been affected, the bubble detection system can cull inaccurate data and improve the accuracy and overall quality of sequencing data. To illustrate, in some implementations, the bubble detection system removes reads from portions of a nucleotide sample slide from call data or reduces quality metrics for reads or nucleobase calls that correspond to particular portions of a nucleotide sample slide that are affected by a bubble. In some instances, the bubble detection system removes nucleobase calls or reduces quality metrics when a detected bubble is equal to or exceeds a size threshold or when data characteristics of a nucleobase call differ from a standard by a particular threshold.
[0029] In addition to improving accuracy, the bubble detection system improves the efficiency with which conventional nucleic acid sequencing systems and corresponding sequencing data analysis software determine nucleobase sequences of nucleic acid polymers. By identifying when bubbles affect or otherwise interfere with nucleotide sample slides, the bubble detection system eliminates the need to exclude certain errors and thus run and re-run multiple sequencing cycles to achieve high quality data. In some such cases, the bubble detection system identifies specific portions of nucleotide sample slides affected by bubbles to specifically identify which corresponding portions of data are corrupted or interfered with by bubbles. Moreover, the bubble detection system can also improve the efficiency of sequencing by classifying particular types of bubbles (e.g., oil, air, or ghosting) or other particular error types used for correction (e.g., block registration failure or dropped blocks). Thus, the bubble detection system improves the efficiency of nucleic acid polymer sequencing by identifying and minimizing the number of data or cycles of portions of nucleotide sample slides that need to be discarded or reevaluated to accurately determine nucleic acid polymer sequences.
[0030] In addition to reducing resequencing attempts or identifying particular bubble-affected data, in some embodiments, the bubble detection system improves efficiency relative to conventional nucleic acid sequencing systems and corresponding sequencing data analysis software by reducing resources typically needed to identify bubbles within sequencing runs. As previously described, the bubble detection system utilizes a bubble detection machine learning model to detect bubbles within sequencing runs. In at least one embodiment, the bubble detection system utilizes a lightweight CNN to identify the presence of bubbles. Thus, in some embodiments, the bubble detection system more efficiently utilizes a lightweight machine learning model from a computational standpoint to analyze available call data and quality data from various sequencing platforms rather than requiring additional hardware (e.g., a tubing sensor) to be used on sequencing equipment or using a heavyweight neural network to process additional information. Thus, in such cases, the bubble detection system produces a low data footprint compared to using image or other sensor data to detect bubbles.
[0031] Regardless of improved efficiency, the bubble detection system also improves the flexibility with which nucleic acid sequencing systems and corresponding sequencing data analysis software detect bubbles therewith. As discussed above, in some embodiments, the bubble detection system is platform-agnostic and does not have additional tubing sensors like those on some fluid-based sequencing devices. Specifically, the bubble detection system flexibly utilizes base call and quality data that is readily accessible from many sequencing platforms. In at least one embodiment, the bubble detection system utilizes a CNN with adaptive max pooling layers that enable the bubble detection system to more flexibly analyze variable input sizes. Thus, the bubble detection system can be implemented and utilized by existing sequencing platforms without requiring additional hardware. Moreover, in some embodiments, the bubble detection system flexibly applies utilizing various configurable circuitry, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0032] As indicated by the above discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the bubble detection system. Additional details regarding the meaning of such terms are now provided. For example, as used herein, the term “nucleotide sample slide” refers to a plate or slide that includes oligonucleotides for sequencing nucleotide fragments of a sample. In some embodiments, the nucleotide sample slide includes a slide that contains a fluidic channel through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, the nucleotide sample slide includes a flow cell that includes small fluidic channels and short oligonucleotides that are complementary to the adapter sequences.
[0033] As used herein, the term “call data” refers to image data or other digital information that indicates individual nucleobases or a sequence of nucleobases of a nucleic acid polymer. Specifically, call data can include intensity values from a camera taken image of a nucleotide sample slide (e.g., color or light intensity values of individual clusters) or other data that indicates individual nucleobases or a sequence of nucleobases of a nucleic acid polymer. In addition to or instead of intensity values, call data can include chromatographic peaks or current changes that indicate individual nucleobases in a sequence. Additionally, in some embodiments, call data includes individual nucleobase calls that identify individual nucleobases (e.g., A, T, C, or G). For example, call data can include data of nucleobase calls in a sequence of a nucleic acid polymer, a number of nucleobase calls that correspond to a particular base (e.g., adenine, cytosine, thymine, or guanine). In some embodiments, call data includes information from sequencing devices that utilize sequencing by synthesis (SBS).
[0034] As used herein, the term "nucleobase call" refers to the assignment or determination of a particular nucleobase to add or incorporate into an oligonucleotide for a sequencing cycle. Specifically, a nucleobase call indicates the assignment or determination of a nucleotide type that has been incorporated into an oligonucleotide on a nucleotide sample slide. In some cases, a nucleobase call includes the assignment or determination of a nucleobase to an intensity value that is produced by a nucleotide of an oligonucleotide added to a nanopore of a nucleotide sample slide. Alternatively, a nucleobase call includes the assignment or determination of a nucleobase to a chromatographic peak or current change that is produced by a nucleotide passing through a nanopore of a nucleotide sample slide. By using nucleobase calls, a sequencing system determines the sequence of a nucleic acid polymer. For example, a single nucleobase call can include an adenine call, a cytosine call, a guanine call, or a thymine call.
[0035] As further used herein, the term "sequencing cycle" or simply "cycle" refers to an iteration of adding or incorporating a nucleobase to an oligonucleotide or an iteration of adding or incorporating nucleobases to an oligonucleotide in parallel. Specifically, a cycle can include an iteration of acquiring and analyzing one or more images having data indicative of an individual nucleobase being added or incorporated to an oligonucleotide or added or incorporated to an oligonucleotide in parallel. Thus, a cycle can be repeated as part of sequencing a nucleic acid polymer. For example, in one or more embodiments, each sequencing cycle involves a single read in which a DNA or RNA strand is read in a single direction or a paired-end read in which a DNA or RNA strand is read from both ends. Further, in certain cases, each sequencing cycle involves a camera taking an image of a nucleotide sample slide or portions of a nucleotide sample slide to generate image data for determining a particular nucleobase added or incorporated into a particular oligonucleotide. After the image capture phase, the sequencing system can remove certain fluorescent labels from the incorporated nucleobases and perform another sequencing cycle until the nucleic acid polymer has been completely sequenced. In one or more embodiments, "cycle" refers to a sequencing cycle within a sequencing-by-synthesis (SBS) run.
[0036] As used herein, the term "nucleic acid polymer" refers to a macromolecule composed of nucleic acid units. Specifically, a nucleic acid polymer can include a macromolecule composed of different nitrogen-containing heterocyclic bases in a sequence. For example, a nucleic acid polymer can include a fragment or molecule of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acid or chimeric or hybrid forms of nucleic acid described below. More specifically, in some cases, a nucleic acid polymer is a nucleic acid polymer found in a sample prepared or isolated by a kit and received by a sequencing device.
[0037] As used herein, the term "quality data" refers to information indicative of the accuracy or quality of a nucleobase call of a sequencing cycle. In particular, quality data is generally indicative of the accuracy of one or more base calls within a sequencing cycle. For example, quality data can include one or more quality indicators.
[0038] As used herein, the term "quality indicator" refers to a particular score or other measure indicative of the accuracy of a nucleobase call of a sequencing cycle. In particular, a quality indicator includes a value indicative of the likelihood that one or more predicted nucleobase calls contain an error. For example, in certain embodiments, a quality indicator can include a Q-score predictive of the error probability of any given base call within a sequencing cycle.
[0039] As used herein, the term "bubble" refers to a spherical or spheroid body or other container that encloses a gas, liquid, or other material. In particular, a bubble refers to a spherical body that can enter a nucleotide sample slide and can affect the quality of data of a sequencing cycle. For example, a bubble can include an air bubble or an oil bubble that occurs within a nucleotide sample slide.
[0040] Additional details regarding the bubble detection system will now be provided in connection with the illustrative drawings depicting example embodiments and implementations of the bubble detection system. For example, Figure 1 A schematic diagram of a system environment (or "environment") 100 in which a bubble detection system 106 operates in accordance with one or more embodiments is shown. As shown, the environment 100 includes one or more server devices 102 connected to a user client device 108 and a sequencing device 114 via a network 112. While Figure 1 Embodiments of the bubble detection system 106 are shown, but alternative embodiments and configurations are possible.
[0041] As Figure 1 As shown in the middle, the server devices 102, the user client device 108, and the sequencing device 114 are connected via the network 112. Accordingly, each component of the environment 100 can communicate via the network 112. The network 112 includes any suitable network over which computing devices can communicate. Example networks are discussed in greater detail below in connection with Figure 8 Example networks are discussed in greater detail below in connection with
[0042] As Figure 1As shown in FIG. 1, sequencing device 114 includes a device for sequencing nucleic acid polymers. In some embodiments, sequencing device 114 analyzes nucleic acid fragments extracted from a sample to generate data on sequencing device 114 directly or indirectly using the computer-implemented methods and systems described herein. More specifically, sequencing device 114 receives and analyzes nucleic acid fragments extracted from a sample within a nucleotide sample slide. In one or more embodiments, sequencing device 114 utilizes SBS to sequence nucleic acid polymers. In addition to or in lieu of communicating across network 112, in some embodiments, sequencing device 114 bypasses network 112 and communicates directly with user client device 108.
[0043] As Figure 1 As further shown, server device 102 can generate, receive, analyze, store, receive, and transmit electronic data, such as data for determining nucleobase calls or sequencing nucleic acid polymers. As Figure 1 As shown in FIG. 1, server device 102 can receive data from sequencing device 114. For example, server device 102 can collect and / or receive sequencing data, including call data, quality data, and other data related to sequencing nucleic acid polymers. Server device 102 can also communicate with user client device 108. In particular, server device 102 can send nucleobase sequences, error data, and other information to user client device 108.
[0044] In some embodiments, server device 102 includes a distributed server, where server device 102 includes many server devices distributed across network 112 and located in different physical locations. Server device 102 can include a content server, an application server, a communication server, a network hosting server, or another type of server.
[0045] As Figure 1 As further shown in FIG. 1, server device 102 can include sequencing system 104. Generally, sequencing system 104 analyzes sequencing data received from sequencing device 114 to determine nucleobase sequences of nucleic acid polymers. For example, sequencing system 104 can receive raw data from sequencing device 114 and determine nucleobase sequences of nucleic acid fragments. In some embodiments, sequencing system 104 determines sequences of nucleobases in DNA and / or RNA fragments. In addition to processing and determining sequences of nucleic acid polymers, sequencing system 104 also analyzes sequencing data to detect irregularities in sequencing cycles. In particular, sequencing system 104 can use bubble detection system 106 to detect bubbles within sequencing cycles and send corresponding notifications to user client device 108.
[0046] As just described, and as Figure 1As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting).
[0047] As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting). Figure 1 As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting).
[0048] Figure 1 The user client device 108 shown in FIG. 1 can include various types of client devices. For example, in some embodiments, the user client device 108 includes a non-mobile device, such as a desktop computer or a server, or other type of client device. In yet other embodiments, the user client device 108 includes a mobile device, such as a laptop, a tablet, a mobile phone, or a smart phone. Additional details regarding the user client device 108 are discussed below with respect to FIG. 2. Figure 8
[0049] As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting). Figure 1 As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting).
[0050] As shown in FIG. 1, the bubble detection system 106 analyzes data from the sequencing device 114 to detect the presence of a bubble within a nucleotide sample slide associated with the sequencing device 114. More specifically, in some embodiments, the bubble detection system 106 receives call data and quality data from the sequencing device 114. Based on the call data and the quality data, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. Based on the first subset of nucleobase calls and the second subset of nucleobase calls, the bubble detection system 106 implements a bubble detection machine learning model to detect the presence of a bubble. Accordingly, the bubble detection system 106 can include one or more machine learning models (e.g., neural networks, SVMs, adaptive boosting). Figure 1 As further shown, the bubble detection system 106 can be located on the user client device 108 as part of the sequencing application 110. As shown, in some embodiments, the bubble detection system 106 is implemented on (e.g., located entirely or partially on) the user client device 108. Additionally or alternatively, in some embodiments, the bubble detection system 106 is implemented on (e.g., located entirely or partially on) the sequencing device 114. In yet other embodiments, the bubble detection system 106 is implemented by one or more other components of the environment 100. Specifically, the bubble detection system 106 can be implemented across the server device 102, the network 112, the user client device 108, and the sequencing device 114 in a variety of different ways.
[0051] Although Figure 1 While the components of the environment 100 are shown as communicating via the network 112, in certain embodiments, the components of the environment 100 can also communicate directly with one another, bypassing the network. For example, and as previously noted, the user client device 108 can communicate directly with the sequencing device 114. Additionally, the user client device 108 can communicate directly with the bubble detection system 106. Also, the bubble detection system 106 can access one or more databases housed on or accessed by the server device 102, or elsewhere in the environment 100.
[0052] As indicated above, the bubble detection system 106 can detect the presence of a bubble within a nucleotide sample slide. For example, Figure 2 A bubble detection system 106 is shown performing a series of actions 200 to detect the presence of a bubble within a nucleotide sample slide, in accordance with one or more embodiments. As part of the series of actions 200, the bubble detection system 106 performs an action 202 of receiving call data, an action 204 of receiving quality data, an action 206 of determining first and second subsets of nucleobase calls, and an action 208 of detecting the presence of a bubble.
[0053] As Figure 2 As shown in the series of actions 200, the action 202 of receiving call data is performed. Specifically, when the action 202 is performed, the bubble detection system 106 receives call data that includes or indicates nucleobase calls for a sequencing cycle of a nucleic acid polymer. In some cases, the bubble detection system 106 accesses call data (e.g., imaging data from the sequencing device 114) from a sequencing device that indicates nucleobase calls for each sequencing cycle. For example, as shown in the series of actions 200, the bubble detection system 106 receives call data 210 that includes or indicates nucleobase calls for a sequencing cycle of a nucleic acid polymer. Figure 2As shown in FIG. 2, the bubble detection system 106 receives image data for each cycle, which includes intensity values indicative of adenine (A) calls, thymine (T) calls, cytosine (C) calls, or guanine (G) calls for each sequencing cycle and portion of the nucleotide sample slide. In some embodiments, the call data is also indicative of the total number or percentage of a particular nucleobase called within a particular cycle. Although Figure 2 While the call data is depicted as image data having colors indicative of intensity values, the bubble detection system 106 can receive call data in any suitable format, such as call data as part of a binary base call (BCL) sequence file or an InterOp metrics file.
[0054] In addition to, or in lieu of, receiving image data when performing the action 202, in certain implementations, the bubble detection system 106 receives call data including individual nucleobase calls across sequencing cycles of the nucleic acid polymer. For example, in some cases, the call data includes explicit information or textual indicators of A, T, C, or G calls for a particular cycle and portion of the nucleotide sample slide. As noted above, the call data can also include the total number or percentage of a particular nucleobase called within a particular cycle.
[0055] As Figure 2 As further shown in FIG. 2, the series of actions 200 includes the bubble detection system 106 performing an action 204 of receiving quality data. As indicated above, the quality data includes quality indicators estimating errors in the nucleobase calls for a cycle. In particular, the bubble detection system 106 receives quality data from the sequencing device indicative of a probability of an erroneous nucleobase call for each cycle. For example, as shown in FIG. 2, the quality data includes a quality indicator corresponding to the total number of called bases for each cycle. Although Figure 2 As further shown in FIG. 2, the series of actions 200 includes the bubble detection system 106 performing an action 204 of receiving quality data. As indicated above, the quality data includes quality indicators estimating errors in the nucleobase calls for a cycle. In particular, the bubble detection system 106 receives quality data from the sequencing device indicative of a probability of an erroneous nucleobase call for each cycle. For example, as shown in FIG. 2, the quality data includes a quality indicator corresponding to the total number of called bases for each cycle. Although Figure 2 While the quality data is depicted as a distribution of total base calls associated with a particular quality indicator, the bubble detection system 106 can receive quality data in any suitable format, such as quality indicators within a BCL file or an InterOp metrics file. In one or more embodiments, the quality data includes quality indicators as described in more detail below.
[0056] As further indicated above, in some embodiments, the quality indicator includes a quality score associated with a probability of incorrect nucleobase call or accuracy of base call. For example, in one or more embodiments, the quality indicator includes a Phred quality score based on the Phred algorithm or a modified Phred algorithm developed by Illumina, Inc. In some embodiments, the bubble detection system 106 determines or uses a Phred score as a quality indicator as described in Method and System for Determining the Accuracy of DNA Base Identifications (U.S. Patent No. 8,392,126, filed September 23, 2009), the contents of which are hereby incorporated by reference in their entirety. A Phred quality score Q10 is equal to a probability of 1 in 10 incorrect nucleobase calls, which means that 1 in 10 nucleobase sequencing reads can contain an error. The following table includes additional Phred quality scores and their equivalent probabilities of incorrect nucleobase call and accuracy of nucleobase call.
[0057]
[0058] Additional details regarding Phred quality scores are provided by Ewing B, Green P. Base-calling of Automated Sequencer Traces Using Phred. II. Error Probabilities. Genome Res., 1998 Mar.; 8(3): 186-194. PMID: 9521922, the entirety of which is incorporated herein by reference.
[0059] As Figure 2 Further shown in the series of actions 200 is an action 206 of determining a first subset of nucleobase calls and a second subset of nucleobase calls. Specifically, when performing the action 206, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator of the quality indicator. In some embodiments, the first subset and the second subset include a proportion or percentage of all nucleobase calls for a given cycle and a particular portion (e.g., a block) of the nucleotide sample slide. The following paragraphs provide additional details regarding the first subset and the second subset.
[0060] As Figure 2 Further shown in the series of actions 200 is an action 206 of determining a first subset of nucleobase calls and a second subset of nucleobase calls. Specifically, when performing the action 206, the bubble detection system 106 determines a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator of the quality indicator. In some embodiments, the first subset and the second subset include a proportion or percentage of all nucleobase calls for a given cycle and a particular portion (e.g., a block) of the nucleotide sample slide. The following paragraphs provide additional details regarding the first subset and the second subset. Figure 2As shown in FIG. 2, the bubble detection system 106 determines a first subset of adenine calls and a second subset of guanine calls for each cycle. In one or more embodiments, the first subset includes a percentage value indicative of a portion of all nucleobase calls corresponding to a particular nucleobase. While Figure 2 While FIG. 2 illustrates the bubble detection system 106 determining a first subset corresponding to at least one nucleobase 210 by determining a percentage of adenine calls and a percentage of guanine calls, the bubble detection system 106 can also determine a first subset including any combination of adenine calls, thymine calls, cytosine calls, and guanine calls.
[0061] As shown in FIG. 2, the bubble detection system 106 determines a first subset of adenine calls and a second subset of guanine calls for each cycle. In one or more embodiments, the first subset includes a percentage value indicative of a portion of all nucleobase calls corresponding to a particular nucleobase. While Figure 2 As further shown in FIG. 2, the bubble detection system 106 also determines a second subset meeting a threshold quality indicator 212. The bubble detection system 106 identifies the threshold quality indicator and determines a subset of nucleobase calls meeting the threshold quality indicator. In some implementations, the bubble detection system 106 determines a threshold quality indicator including a percentage or proportion of nucleobase calls meeting or exceeding a benchmark threshold quality indicator. To illustrate, in one or more embodiments, the bubble detection system 106 determines the threshold quality indicator to equal a Phred quality score Q30. For each cycle, the bubble detection system 106 determines a percentage (or other subset) of nucleobase calls meeting or exceeding the Q30 quality indicator.
[0062] After performing the act 206 of determining the first and second subsets of nucleobase calls, the bubble detection system 106 performs the act 208 of detecting the presence of a bubble. Specifically, when performing the act 208, the bubble detection system 106 detects the presence of a bubble within the nucleotide sample slide by utilizing a bubble detection machine learning model based on the first subset of nucleobase calls and the second subset of nucleobase calls. As shown in FIG. 2, for example, the bubble detection system 106 utilizes the bubble detection machine learning model 216 to analyze the input matrix 214 and generate the output 218. Figure 2 As shown in FIG. 2, for example, the bubble detection system 106 utilizes the bubble detection machine learning model 216 to analyze the input matrix 214 and generate the output 218.
[0063] In addition to the series of acts 200, in some cases, the bubble detection system 106 also provides an alert to a computing device indicative of the presence of a bubble. Specifically, the bubble detection system 106 provides a notification or alert for display via a computing device associated with a user. Additionally or alternatively, the bubble detection system 106 provides an alert to a sequencing device. In any case, the bubble detection system 106 can include an error classification within the alert indicative of a type of bubble or error. Further, the alert can include additional information including a portion of the nucleotide sample slide and / or a sequencing cycle in which the bubble occurred.
[0064] Additionally, in some embodiments, the bubble detection system 106 determines one or more corrective actions based on detecting the presence of a bubble. To illustrate, in some embodiments, the bubble detection system 106 reduces a quality indicator for particular reads in a cycle, a particular cycle, or a particular portion of the nucleotide sample slide based on detecting the presence of a bubble. In some cases, for example, the bubble detection system 106 can identify nucleobase calls in a cycle for which to reduce a quality indicator by identifying unique molecular identifiers (UMIs) for corresponding reads. Additionally or alternatively, the bubble detection system 106 can excise affected calls from call data based on identifying particular reads in a cycle, a particular cycle, or a particular portion of the nucleotide sample slide affected by a bubble. In some cases, based on determining a persistence of a bubble, the bubble detection system 106 can include a suggested action to address the bubble within an alert. For example, based on determining that a number of detected oil bubbles satisfies a threshold, the bubble detection system 106 provides an alert including a suggested action to check an oil leak of a component of the sequencing device or to reload the nucleotide sample slide.
[0065] As previously described, in some embodiments, the bubble detection system 106 identifies a particular portion of the nucleotide sample slide affected by a bubble. In one example, a portion of the nucleotide sample slide includes a tile of a flow cell. Accordingly, in one or more embodiments, the bubble detection system 106 performs a series of actions 200 for a particular portion of the nucleotide sample slide. Thus, in certain embodiments, the bubble detection system 106 receives call data and quality data across cycles for a single portion of the nucleotide sample slide. Accordingly, the bubble detection system 106 can identify a particular portion of the nucleotide sample slide affected by a bubble.
[0066] As Figure 2 As further shown in FIG. 2, the bubble detection system 106 utilizes an input matrix 214 as an input into a bubble detection machine learning model 216. In one or more embodiments, the input matrix 214 includes data for a first subset of nucleobase calls (e.g., a subset of adenine calls and a subset of guanine calls) corresponding to at least one nucleobase and a second subset of nucleobase calls meeting a threshold quality indicator. As described below with respect to FIG. 3, the size of the input matrix 214 can vary based on a number of sequencing cycles. Figure 5
[0067] As Figure 2 Further shown, the bubble detection system 106 implements a bubble detection machine learning model 216. The bubble detection machine learning model 216 extracts features from the input matrix 214 to identify the presence of bubbles within the nucleotide sample slide. The bubble detection machine learning model 216 can include various types of machine learning models. In some embodiments, the bubble detection machine learning model 216 includes a neural network, such as a CNN, or a different type of machine learning model, such as an SVM or a self-adaptive boosting machine learning model. Figure 5 The corresponding discussion further describes an exemplary CNN in accordance with one or more embodiments.
[0068] After passing the input matrix 214 through the bubble detection machine learning model 216, the bubble detection system 106 generates an output 218 with the bubble detection machine learning model 216. In some embodiments, the output 218 includes (i) an indication of bubbles within the nucleotide sample slide and (ii) an error classification. As Figure 2 As shown in FIG. 3, for example, the output 218 includes a potential error classification that includes an oil bubble, an air bubble, and a drop out. In additional embodiments, the output 218 includes an additional error classification of a ghost bubble. Figures 4A-4C The corresponding paragraphs further describe the error classifications generated by the bubble detection system 106 in accordance with one or more embodiments.
[0069] Figure 2 A general overview of the bubble detection system 106 that determines the presence of bubbles within a nucleotide sample slide in accordance with one or more embodiments is provided. As described above, the bubble detection system 106 can flexibly determine the presence of bubbles based on various types of detection data. Figure 3 Different types of detection data that the bubble detection system 106 can use to determine the presence of bubbles within a nucleotide sample slide are shown. Generally, Figure 3 Single channel data 302, dual channel data 304, and four channel data 306 obtained as part of an SBS cycle are shown. The following paragraphs further describe each of these types of data.
[0070] As Figure 3 As shown in FIG. 3, for example, the output 218 includes a potential error classification that includes an oil bubble, an air bubble, and a drop out. In additional embodiments, the output 218 includes an additional error classification of a ghost bubble. Figure 3As shown, single-channel data includes a dual-image synthesis 312 of a portion 310a of a nucleotide sample slide 308a for a given cycle of nucleic acid polymer sequencing. In some embodiments, the dual-image synthesis 312 comprises a combination of two images, each captured using the same detection channel, the same dye, or the same fluorescent label captured at different times. Unlike four-channel SBS chemistry, where the sequencer uses different fluorescent dyes or labels for each nucleobase, single-channel SBS chemistry uses one fluorescent dye, two chemical steps, and two imaging steps (producing two images) per sequencing cycle. In single-channel chemistry, for example, adenine has a removable label and is labeled only in the first image 318. Cytosine has a linker group that can bind the label and is labeled only in the second image 320. Thymine has a persistent fluorescent label and is therefore labeled in both the first image 318 and the second image 320. Guanine is unlabeled and therefore does not fluoresce in either image. The bubble detection system 106 determines nucleobase detection based on analysis of the different emission patterns of each base across the two images.
[0071] In one or more embodiments, the bubble detection system 106 obtains single-channel data based on intensity information. In such embodiments, instead of capturing two images, the sequencing system 104 captures a single image and associates different intensity values with different nuclei. Specifically, three or more nuclei are bound to a fluorescent dye or label at different intensities. The bubble detection system 106 can associate an intensity range with a specific nuclei, or associate the absence of the dye or label with a specific nuclei. Thus, the bubble detection system 106 uses a single channel to determine nuclei detection based on intensity data.
[0072] like Figure 3 As further shown, in some cases, the bubble detection system 106 receives detection data in the form of dual-channel data 304. Specifically, the dual-channel data 304 includes a dual-image synthesis 314 of a portion 310b of the nucleotide sample slide 308b. Specifically, the dual-image synthesis 314 includes two images, each captured using a detection channel specific to two different dyes or different fluorescent labels. The dual-channel SBS determines the detection of all four nucleosides relative to the four-channel SBS chemically simplified nucleotide detection by using two fluorescent dyes and the dual-image synthesis 314. For example, in one embodiment, the camera of the sequencing device captures images using red and green filter bands. Thymine nucleosides are labeled with green fluorescent groups, cytosine with red fluorescent groups, and adenine with both red and green fluorescent groups. Guanine is permanently black. The bubble detection system 106 determines nucleobase detection by using two filter channels to process the dual-image synthesis 314 and by determining which nucleosides are incorporated into each cluster within the portion 310b of the nucleotide sample slide 308b.
[0073] As further noted above, in some embodiments, the bubble detection system 106 receives the call data in the form of four-channel data 306. Specifically, the four-channel data 306 includes a four-image composite 316 of a portion 310c of a nucleotide sample slide 308c. Specifically, the four-image composite 316 includes four images each captured using a detection channel specific to one of four different dyes or fluorescent labels. The four-channel SBS cycle begins with a chemical step in which all four different labeled bases are added to the nucleotide sample slide. An imaging cycle begins and includes capturing the four-image composite 316 using four different filter channels or wavelength bands. The bubble detection system 106 processes the four-image composite 316 to determine which nucleobases were incorporated at each cluster position across the nucleotide sample slide.
[0074] The bubble detection system 106 determines a subset of nucleobase calls based on the call data. Specifically, the bubble detection system 106 stores, processes, and analyzes the single-channel data 302, the dual-channel data 304, and / or the four-channel data 306 to determine the base calls for each sequencing cycle. More specifically, the bubble detection system 106 identifies nucleobases by analyzing the different emission patterns for each nucleobase across the captured images. Upon completion of a sequencing cycle, the bubble detection system 106 determines a total number of nucleobase calls. The bubble detection system also determines a subset of individual nucleobase calls by comparing the number of calls for a particular nucleobase to the total number of nucleobase calls for the cycle. In one example, the bubble detection system 106 determines that 310 adenine calls out of 1000 total base calls for a given cycle. Based on this determination, the bubble detection system 106 determines that the subset of adenine calls (%A calls) is equal to 0.31.
[0075] As previously noted, in some embodiments, as part of detecting the presence of bubbles within a nucleotide sample slide, the bubble detection system 106 utilizes a bubble detection machine learning model to generate an error classification based on a subset of adenine calls, a subset of guanine calls, and a subset of nucleobase calls that meet threshold quality indicators for a nucleic acid polymer sequencing cycle. For example, in certain embodiments, the bubble detection system 106 generates an error classification that identifies errors caused by air bubbles, oil bubbles, ghost bubbles, or dropouts. Each error classification corresponds to different data features of indicators from the call data and the quality data.
[0076] The bubble detection system 106 can detect bubbles or classify errors corresponding to Figures 4A-4C Such errors are classified according to various data features depicted in FIGS. 1-3. According to one or more embodiments, Figure 4A 、 Figure 4B and Figure 4C An exemplary chart is shown that illustrates the progression of input data depicted as data features during a cycle within a sequencing run. Specifically,Figure 4A A data chart illustrating exemplary data features corresponding to a nucleotide sample slide with no air bubbles is shown. Figure 4B Exemplary data features corresponding to air bubbles, ghost bubbles, and oil bubbles are shown in accordance with one or more embodiments. Figure 4C Exemplary data features corresponding to suspect air bubbles, dropouts, and dropouts occurring within a single cycle are shown in accordance with one or more embodiments. While Figures 4A-4C A chart depicting data (including a subset of adenine calls, a subset of guanine calls, and a subset of nucleobase calls meeting a threshold metric) in an input bubble detection machine learning model, the bubble detection system 106 does not input the chart itself into such a model.
[0077] As an overview, Figures 4A-4C The charts in FIGS. 1-3 share some common features. For example, Figures 4A-4C Exemplary charts 412a-412g with data features corresponding to various error classifications are shown. Metrics illustrated by the shown charts 412a-412g include error percentages 404a-404g, adenine call percentages 406a-406g, guanine call percentages 408a-408g, and Q30 meeting percentages 410a-410g. More specifically, the charts 412a-412g indicate the progression of the metrics during sequencing cycles within a sequencing run. The error percentages 404a-404g indicate the predicted percentage of error in nucleobase calls in each cycle. The adenine call percentages 406a-406g indicate the percentage (or subset) of all nucleobase calls in each cycle that include adenine calls. Similarly, the guanine call percentages 408a-408g indicate the percentage (or subset) of all nucleobase calls in each cycle that include guanine calls. The Q30 meeting percentages 410a-410g indicate the percentage of nucleobase calls in each cycle that meet the Q30 threshold quality metric. In one or more other embodiments, the bubble detection system 106 extracts features from other metrics to identify and classify errors.
[0078] As described above, Figure 4A Chart 412a associated with no air bubbles is shown. Specifically, chart 412a displays data features for a nucleotide sample slide that does not contain air bubbles. Generally, no air bubbles corresponds to data features with relatively stable metrics. For example, the error percentage 404a, the adenine call percentage 406a, the guanine call percentage 408a, and the Q30 meeting percentage 410a remain relatively stable over the sequencing cycles. Chart 412a provides a baseline for comparing charts corresponding to different errors. Based on the data corresponding to chart 412a, the bubble detection system 106 does not detect the presence of air bubbles.
[0079] In contrast,Figure 4B A chart 412b with data characteristics indicative of air bubbles, a chart 412c with data characteristics indicative of ghost bubbles, and a chart 412d with data characteristics indicative of oil bubbles are shown. For example, the chart 412b includes indicators in the data characteristics reflecting nucleobase calls for a nucleotide sample slide containing air bubbles. Generally, air bubbles are produced by air entering the fluidic lines and channels within the nucleotide sample slide. When air bubbles occur and are captured during the imaging phase of the sequencing cycle, the air bubbles adversely affect the data quality of the sequencing reads. For example, during the imaging phase, air bubbles can blur portions of the image or reduce chemical efficiency. More specifically, air bubbles can enter the nucleotide sample slide from the pad of the nucleotide sample slide and de-aerate during imaging.
[0080] As indicated by the chart 412b, the air bubbles caused a sharp spike in both the error percentage 404b and the guanine call percentage 408b, while also causing a drop in the adenine call percentage 406b and the percentage of Q30s 410b. As further shown in the chart 412b, the air bubbles occurred between the 60thand 80thsequencing cycles. Based on the data corresponding to the data characteristics shown in the chart 412b, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as air bubbles. Figure 4B As further shown in the chart 412b, the air bubbles occurred between the 60thand 80thsequencing cycles. Based on the data corresponding to the data characteristics shown in the chart 412b, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as air bubbles.
[0081] As further shown in the chart 412c, the ghost bubbles occurred at some time after the 80thsequencing cycle. The ghost bubbles caused a rapid increase in the error percentage 404c and the error percentage 404c remained elevated for the remaining sequencing cycles. Additionally, the percentage of Q30s 410c reflected the error percentage 404c and dropped at the same sequencing cycles. As further shown in the chart 412c, the adenine call percentage 406c and the guanine call percentage 408c remained relatively similar to the control. Based on the data corresponding to the data characteristics shown in the chart 412c, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as ghost bubbles. Figure 4B As further shown in the chart 412c, the ghost bubbles occurred at some time after the 80thsequencing cycle. The ghost bubbles caused a rapid increase in the error percentage 404c and the error percentage 404c remained elevated for the remaining sequencing cycles. Additionally, the percentage of Q30s 410c reflected the error percentage 404c and dropped at the same sequencing cycles. As further shown in the chart 412c, the adenine call percentage 406c and the guanine call percentage 408c remained relatively similar to the control. Based on the data corresponding to the data characteristics shown in the chart 412c, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as ghost bubbles.
[0082] As further shown in the chart 412c, the ghost bubbles occurred at some time after the 80thsequencing cycle. The ghost bubbles caused a rapid increase in the error percentage 404c and the error percentage 404c remained elevated for the remaining sequencing cycles. Additionally, the percentage of Q30s 410c reflected the error percentage 404c and dropped at the same sequencing cycles. As further shown in the chart 412c, the adenine call percentage 406c and the guanine call percentage 408c remained relatively similar to the control. Based on the data corresponding to the data characteristics shown in the chart 412c, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as ghost bubbles.
[0083] As further shown in the chart 412c, the ghost bubbles occurred at some time after the 80thsequencing cycle. The ghost bubbles caused a rapid increase in the error percentage 404c and the error percentage 404c remained elevated for the remaining sequencing cycles. Additionally, the percentage of Q30s 410c reflected the error percentage 404c and dropped at the same sequencing cycles. As further shown in the chart 412c, the adenine call percentage 406c and the guanine call percentage 408c remained relatively similar to the control. Based on the data corresponding to the data characteristics shown in the chart 412c, the bubble detection system 106 would detect the presence of the bubbles and classify the bubbles as ghost bubbles. Figure 4BAs depicted in FIG. 4I, chart 412d illustrates metrics for a nucleotide sample slide containing air bubbles. Generally, oil bubbles occur when oil from components of the sequencing device enters the nucleotide sample slide. Similar to air bubbles, oil bubbles negatively impact data quality by affecting images captured during the imaging phase of sequencing cycles. More specifically, oil bubbles absorb dye or labels and emit fluorescence, causing the sequencing device to capture excess fluorescence. For example, and as shown in chart 412d, an oil bubble captured between the 20th and 40th sequencing cycle causes spikes in error percentage 404d and adenine call percentage 406d. Chart 412d also illustrates a smaller drop in guanine call percentage 408d and a more significant drop in percentage of Q30 410d. Based on the data corresponding to the data features shown in chart 412d, bubble detection system 106 would detect the presence of a bubble and classify the bubble as an oil bubble.
[0084] As indicated above, Figure 4C Exemplary charts corresponding to additional error classifications are shown. Specifically, Figure 4C Charts 412e, 412f, and 412g corresponding to a suspect bubble, a drop-out, and a drop-out within a single cycle, respectively, are shown.
[0085] As Figure 4C As shown in FIG. 4I, for example, chart 412e illustrates metrics for a nucleotide sample slide with a suspect bubble. Generally, a suspect bubble can indicate no bubble, one of the previously described bubbles (e.g., air bubble, ghost bubble, oil bubble), or another type of error. Specifically, while certain bubble classifications (e.g., air bubble, ghost bubble, and oil bubble) are associated with distinct data features, such data features can also include some variation. Additionally, other errors besides bubbles can affect the quality of the data. Thus, in some embodiments, bubble detection system 106 generates a classification of “no bubble” based on a subset of nucleobase calls corresponding to the data features in chart 412e. Alternatively, in certain implementations, bubble detection system 106 generates a classification of “unknown bubble type” or “unknown error type” based on a subset of nucleobase calls corresponding to the data features in chart 412e. In one or more embodiments, a suspect bubble classification corresponds to data features that are slightly different from the typical data features of a particular bubble classification or no bubble data features (e.g., as shown in FIG. 4I). Figure 4A
[0086] To illustrate, chart 412e shows a spike in error percentage 404e and a corresponding dip in percentage passing Q30 410e. But adenine call percentage 406e and guanine call percentage 408e of chart 412e remain relatively unaffected. In one or more embodiments, bubble detection system 106 determines a classification of a suspect bubble based on features of an input matrix that resemble features of air, oil, or ghost bubbles but differ from features of air, oil, or ghost bubbles beyond a threshold. Based on data corresponding to the data features shown in chart 412e, bubble detection system 106 would detect the presence of a bubble but not classify the bubble.
[0087] Figure 4C Charts 412f and 412g are also shown corresponding to a sample slide with a missing nucleotide. Generally, missing refers to when a camera does not capture or captures a limited amount of image data of a portion (e.g., a zone within a flow cell) or a cluster within a portion of a nucleotide sample slide. Such missing differs from and is not referring to image data with a dark signal or intensity value that indicates a nucleotide that lacks a particular fluorescent label or a nucleotide with a label that has not been irradiated by light of a particular wavelength. Missing can occur in various stages of a sequencing cycle. As shown in chart 412f, missing can occur during a cluster or portion registration phase of SBS sequencing. Additionally, and as shown in chart 412g, missing can occur during a single cycle.
[0088] As noted above, chart 412f shows the impact of missing that occurs during cluster or portion registration. Generally, a cluster refers to a group of nucleic acid fragments or clonal fragments from a sample. Specifically, a cluster represents thousands of copies of the same DNA or RNA fragment. For example, in one or more embodiments, a cluster is immobilized in a portion of a nucleotide sample slide. In some embodiments, clusters can be spaced evenly using a patterned nucleotide sample slide.
[0089] During cluster and portion registration, sequencing system 104 records the location of clusters and portions for imaging. In some embodiments, sequencing system 104 also records intensity values during cluster and portion registration. Generally, missing that occurs during cluster registration results in sequencing system 104 being unable to register a particular cluster for the duration of a sequencing cycle. As shown in chart 412f, missing that occurs during portion or cluster registration causes a longer lasting impact. Specifically, error percentage 404f indicates a sharp increase around the 120th sequencing cycle and percentage passing Q30 410f indicates a corresponding dip. Based on data corresponding to the data features shown in chart 412f, bubble detection system 106 would detect a missing event during registration.
[0090] Dropouts occurring during cluster and partial registration can have various causes. For example, dropouts during cluster registration can indicate the presence of a bubble covering an entire portion of a nucleotide sample slide. Additionally, dropouts during cluster registration can indicate other types of irregularities. For example, dropouts can indicate an error in software or hardware functionality. In one example, dropouts indicate a failure of a direct memory access (DMA) transfer between a sequencing device and a user client device or a server device. Additionally, dropouts can indicate a hardware failure in a sensor or camera that causes excision of data related to a particular nucleotide sample slide portion or cluster. For example, a sensor within a sequencing device can be out of focus.
[0091] As Figure 4C As further shown in chart 412g, bubble detection system 106 can detect dropouts occurring during sequencing cycles. Specifically, during a given cycle, a sequencing device can erroneously omit data for a cluster or portion of a nucleotide sample slide. For example, a sequencing device can experience a mechanical error that causes a sensor to drop an entire cluster or portion of a nucleotide sample slide during a cycle. In another example, a sequencing device experiences a real-time analysis (RTA) error that causes dropouts during a sequencing run. As shown in chart 412g, dropouts in a single sequencing cycle can manifest as a significant drop in percentage of Q30 410g and a smaller corresponding drop in error percentage 404g. Furthermore, both adenine call percentage 406g and guanine call percentage 408g have data gaps corresponding to cycles affected by dropouts. Based on the data corresponding to the data features shown in chart 412f, bubble detection system 106 would detect a dropout event during a single cycle.
[0092] Figures 4B-4C Exemplary charts showing data features of various error classifications are shown. In some embodiments, bubble detection system 106 utilizes a bubble detection machine learning model to extract features from an input matrix and determine the presence of a bubble and a corresponding classification of the bubble. As previously described, a bubble detection machine learning model can include a neural network. Figure 5 An exemplary configuration of a bubble detection neural network is shown, in accordance with one or more embodiments. Specifically, Figure 5 A bubble detection neural network 500 including a feature extraction layer 502, a classification layer 504, and an adaptive max pooling layer 508 is shown. As shown, bubble detection neural network 500 includes a trained neural network applied by bubble detection system 106 to an input matrix 510. Bubble detection system 106 further generates an output classification 506 by utilizing bubble detection neural network 500.
[0093] As Figure 5As shown in FIG. 5, the bubble detection neural network 500 includes a trained neural network. Specifically, in one or more embodiments, the bubble detection system 106 trains the bubble detection neural network 500 with a training dataset. In one embodiment, the bubble detection system 106 accesses a training dataset that includes ground truth classifications for training input matrices. Figure 6A and the corresponding discussion provide additional description regarding how the bubble detection system 106 trains the bubble detection neural network 500 in accordance with one or more embodiments.
[0094] As Figure 5 As further shown in FIG. 5, the bubble detection system 106 applies the bubble detection neural network 500 to the input matrix 510 after training. As shown in Figure 5 As shown in FIG. 5, for each portion of the nucleotide sample slide (e.g., a block of the flow cell), the input matrix 510 includes three one-dimensional input channels of length N, where N equals the number of SBS cycles in the run. In some embodiments, the three one-dimensional input channels include a subset of adenine calls that meet a threshold quality metric (e.g., %Q30), a subset of guanine calls, and a subset of nucleobase calls. The size of the input matrix 510 is variable and can thus accommodate a wide range of sequencing run lengths.
[0095] In addition to training machine learning models to detect and classify bubbles, in certain implementations, the bubble detection system 106 trains such models to distinguish bubbles introduced during a particular sequencing chemistry step or phase. Bubbles occurring at different SBS or Sanger chemistry steps or phases can elicit unique data signatures. For example, by using training data corresponding to such unique data signatures (specific to the chemistry step or phase at which a bubble enters or interferes with the nucleotide sample slide), the bubble detection system 106 can train a bubble detection machine learning model to detect and distinguish bubbles introduced during a particular SBS chemistry step or phase. In some embodiments, for example, the bubble detection system 106 distinguishes bubbles introduced during a sequencing step (e.g., incorporation or deblocking) or during an imaging step (e.g., scanning mixing of reagents in the flow cell).
[0096] As indicated above, and as Figure 5As shown in FIG. 5, in some embodiments, the bubble detection neural network 500 includes a light-weight CNN. The bubble detection neural network 500 can include a CNN with lower network layers (e.g., convolution and deconvolution layers) and higher neural network layers (e.g., fully connected layers). In alternative embodiments, the bubble detection neural network 500 employs a different neural network architecture. Further, in some implementations, the bubble detection neural network 500 does not use a down-sampling method, such as implementing a max-pooling layer to compress dimensions after a convolution operation. In such implementations, the bubble detection system 106 excludes the max-pooling layer to maintain representation size, particularly for short sequencing runs (e.g., N = 36).
[0097] As shown in FIG. 5, in some embodiments, the bubble detection neural network 500 includes a light-weight CNN. The bubble detection neural network 500 can include a CNN with lower network layers (e.g., convolution and deconvolution layers) and higher neural network layers (e.g., fully connected layers). In alternative embodiments, the bubble detection neural network 500 employs a different neural network architecture. Further, in some implementations, the bubble detection neural network 500 does not use a down-sampling method, such as implementing a max-pooling layer to compress dimensions after a convolution operation. In such implementations, the bubble detection system 106 excludes the max-pooling layer to maintain representation size, particularly for short sequencing runs (e.g., N = 36). Figure 5 As further shown in FIG. 5, the bubble detection neural network 500 includes an adaptive max-pooling layer 508. In some implementations, the adaptive max-pooling layer 508 is positioned between the feature extraction layer 502 and the classification layer 504 of the bubble detection neural network 500. By implementing the adaptive max-pooling layer 508, the bubble detection system 106 specifies a representation size and spatially collapses features for input into the classification layer 504. The implementation of the adaptive max-pooling layer 508 improves the efficiency of the bubble detection neural network 500. In alternative CNNs as shown in FIG. 6, in some cases, the bubble detection neural network 500 does not include the adaptive max-pooling layer 508. Figure 5 As further shown in FIG. 5, the bubble detection neural network 500 includes an adaptive max-pooling layer 508. In some implementations, the adaptive max-pooling layer 508 is positioned between the feature extraction layer 502 and the classification layer 504 of the bubble detection neural network 500. By implementing the adaptive max-pooling layer 508, the bubble detection system 106 specifies a representation size and spatially collapses features for input into the classification layer 504. The implementation of the adaptive max-pooling layer 508 improves the efficiency of the bubble detection neural network 500. In alternative CNNs as shown in FIG. 6, in some cases, the bubble detection neural network 500 does not include the adaptive max-pooling layer 508.
[0098] In some embodiments, by using the adaptive max-pooling layer 508, the bubble detection neural network 500 becomes translationally invariant. More specifically, a translationally invariant network produces the same output regardless of certain changes in the input. In one example, a translationally invariant version of the bubble detection neural network 500 simply indicates the presence and classification of bubbles within a portion of the nucleotide sample slide, but does not indicate the particular cycle in which the bubble occurred. By removing or adjusting parameters of the adaptive max-pooling layer 508, the bubble detection system 106 can specify additional classifications to include in the output. For example, the bubble detection neural network 500 can generate an indication of the particular cycle in which a bubble occurred in addition to the error classification.
[0099] As indicated above, Figure 5 The classification layer 504 is shown as part of the bubble detection neural network 500. As shown here, the classification layer 504 includes a fully connected neural network that classifies the features extracted by the feature extraction layer 502. In one or more implementations, the classification layer 504 can generate a multi-class output and indicate multiple error classifications for a single portion of the nucleotide sample slide. For example, the classification layer 504 can generate a classification of both oil bubbles and air bubbles for a single portion.
[0100] As indicated above, Figure 5As further shown, the bubble detection neural network 500 includes an output classification 506. In some embodiments, the bubble detection neural network 500 outputs a corresponding confidence or probability score. Based on determining that the confidence or probability score for a particular classification meets a confidence threshold, the bubble detection system 106 determines the particular classification of bubble, air bubble, or dropout for the input matrix 510. In other words, the bubble detection system 106 detects a bubble or dropout event and classifies the same as a bubble, air bubble, or dropout based on a confidence score meeting a particular threshold. While Figure 5 The output classification 506 can include any number of additional classifications, shown for bubble, air bubble, and dropout classifications. For example, the output classification 506 can include a ghost bubble classification, a registration dropout classification, an imaging dropout classification, a suspicious bubble classification, and other error classifications.
[0101] Figure 5 The bubble detection neural network 500 in FIG. 5 illustrates an exemplary configuration of a CNN according to one or more embodiments. In other embodiments, the bubble detection system 106 utilizes machine learning models having various other configurations. Alternatively, the bubble detection system 106 can utilize neural networks having different configurations to identify particular cycles affected by bubbles. For example, in certain embodiments, the bubble detection system 106 incorporates an attention layer into the CNN to generate a classification indicating a particular location (e.g., cluster, portion) on the nucleotide sample slide affected by a bubble. The bubble detection system 106 can also implement other types of deep neural networks. For example, the bubble detection system 106 can implement a long short-term memory (LSTM) network or other type of recurrent neural network. Further, in additional embodiments, the bubble detection system 106 utilizes different types of machine learning models as the bubble detection neural network 500. In some examples, the bubble detection system 106 utilizes an SVM or an AdaBoost machine learning model.
[0102] In some embodiments, the bubble detection system 106 uses nucleobase call data corresponding to spatial images (or reconstructed spatial images) to detect the presence of bubbles within a portion of a nucleotide sample slide. For example, and as previously described, the bubble detection system 106 can use spatial images of a portion (e.g., tile) or sub-portion (e.g., sub-tile) of a nucleotide sample slide to train an image machine learning model to detect or classify bubbles. In some embodiments, for example, the bubble detection system 106 identifies ground truth classification labels of nucleobase call data (e.g., from a BCL or BAM file) corresponding to spatial image data having a correct detected presence or absence of a bubble to train a bubble detection machine learning model (e.g., the bubble detection neural network 500).
[0103] As just suggested, Figure 5Overall, a bubble detection system 106 that trains an image machine learning model and a bubble detection machine learning model using nucleobase call data corresponding to spatial images is shown in accordance with one or more embodiments. Specifically, Figures 6A-6C The bubble detection system 106 is shown training the image machine learning model using spatial images of portions of nucleotide sample slides, generating ground truth classification labels for such spatial images and corresponding nucleobase call data, and utilizing the nucleobase call data and ground truth classification labels to further train the bubble detection machine learning model. Figure 6A Exemplary spatial images generated by the bubble detection system 106 are shown in accordance with one or more embodiments. Figure 6B An exemplary sequencing run image depicting a portion of a nucleotide sample slide is shown in accordance with one or more embodiments.
[0104] As noted above, in some implementations, the bubble detection system 106 utilizes the image machine learning model 608 to detect or classify bubbles based on spatial images (or reconstructed spatial images) of portions or sub-portions of nucleotide sample slides. To illustrate, Figure 6C The bubble detection system 106 is shown training the image machine learning model 608 using the spatial images 606a-606n, and identifying nucleobase call data 602a-602n and ground truth classification labels 604a-604n corresponding to the spatial images 606a-606n. The bubble detection system 106 then uses the nucleobase call data 602a-602n and ground truth classification labels 604a-604n to train the bubble detection machine learning model 622. Although Figure 6A The bubble detection system 106 is shown training the image machine learning model 608, such training or use of the image machine learning model 608 is optional, and is representative of one or more embodiments. Indeed, in some embodiments, the bubble detection system 106 uses some or all of the nucleobase call data 602a-602n and ground truth classification labels 604a-604n to train the bubble detection machine learning model 622 without training or using the image machine learning model 608. Thus, Figure 6A A dashed line is included around the image machine learning model 608, as well as the corresponding outputs and determinations, to indicate that such training and use is optional.
[0105] For simplicity, the present disclosure describes an initial training iteration, followed by a summary of subsequent training iterations as Figure 6A depicted in FIG. 6B. By way of overview, in FIG. 6B, the bubble detection system 106 is shown training the image machine learning model 608 using the spatial images 606a-606n, and identifying nucleobase call data 602a-602n and ground truth classification labels 604a-604n corresponding to the spatial images 606a-606n. The bubble detection system 106 then uses the nucleobase call data 602a-602n and ground truth classification labels 604a-604n to train the bubble detection machine learning model 622. Although Figure 6AIn the depicted initial training iteration, the bubble detection system 106 utilizes the nucleobase call data 602a to generate or reconstruct spatial images 606a. The bubble detection system 106 utilizes the spatial images 606a as input for the image machine learning model 608 to subsequently generate bubble classifications 610a.
[0106] As just Figure 6A indicated and as shown, the bubble detection system 106 utilizes the nucleobase call data 602a-602n to generate the spatial images 606a-606n. In one or more embodiments, the nucleobase call data 602a-602n includes nucleobase calls and quality indicators corresponding to portions or sub-portions within a nucleotide sample slide for a given sequencing cycle. In some cases, the bubble detection system 106 accesses the nucleobase call data 602a-602n from a BCL sequence file or a BAM (*.bam) file. Some such nucleobase call data can, for example, include a pattern of nucleobase calls (e.g., a circular pattern of A calls or G calls) that indicates the presence of a bubble within a tile or sub-tile of a nucleotide sample slide.
[0107] As Figure 6A further shown in the above, in one or more embodiments, the bubble detection system 106 generates or reconstructs the spatial images 606a-606n based on the nucleobase call data 602a-602n. Generally, the bubble detection system 106 incorporates the nucleobase calls into a spatial pattern by generating a spatial representation of the nucleobase calls from a BCL or BAM file arranged according to the locations of the clusters on the nucleotide sample slide. In one example, the bubble detection system 106 color-encodes the spatial images 606a-606n by associating nucleobases with particular colors. For example, the bubble detection system 106 can associate A calls with yellow, G calls with blue, C calls with red, and T calls with green. The bubble detection system 106 Figure 6A An exemplary spatial image is shown in accordance with one or more embodiments.
[0108] In one or more embodiments, the bubble detection system 106 reduces the size of the spatial images 606a-606n prior to inputting them into the image machine learning model 608. In at least one example, the bubble detection system 106 down-samples the spatial images 606a-606n. For example, the bubble detection system 106 processes the spatial images 606a-606n to remove high-frequency information and retain low-frequency information for input. Thus, in some cases, the bubble detection system 106 can apply the image machine learning model 608 to a low-frequency version of the spatial images 606a-606n to improve efficiency.
[0109] For example, after inputting spatial image 606a as part of the initial training iteration, bubble detection system 106 executes image machine learning model 608. As mentioned above, image machine learning model 608 can be a neural network, such as CNN. In some cases, to name a few examples, image machine learning model 608 takes the form of a dense convolutional network (DenseNet) or a residual neural network (ResNet).
[0110] like Figure 6B As further shown, upon receiving input data for the initial training iteration, the image machine learning model 608 determines a bubble classification 610a. Additionally, the image machine learning model 608 predicts the location of detected bubbles within a portion or sub-portion of the nucleotide sample slide based on a spatial pattern within the input data. For example, the image machine learning model 608 generates a bubble classification 610a that includes markers indicating the presence and location of bubbles within a portion of the nucleotide sample slide. Typically, bubbles are associated with circular spatial patterns within the nucleotide detection data 602a or spatial image 606a. Therefore, in some embodiments, the bubble classification 610a includes a bubble classification along with the location of the bubbles. For example, the bubble classification 610a may indicate a predicted portion or sub-portion of the nucleotide sample slide containing bubbles or portions of bubbles. Similarly, the bubble classification 610a may indicate a predicted portion or sub-portion of the nucleotide sample slide that does not contain bubbles or portions of bubbles.
[0111] like Figure 6A As further shown, the bubble detection system 106 uses a loss function 612 to compare bubble classification 610a with a baseline truth classification marker 604a. In some embodiments, the baseline truth classification marker 604a includes a baseline truth bubble classification and bubble location corresponding to the nucleobase detection data 602a. For example, the baseline truth classification marker 604a may indicate (i) a specific portion or sub-portion of a nucleotide sample slide containing a bubble or a portion of a bubble, and (ii) a specific portion or sub-portion of a nucleotide sample slide that does not contain a bubble or a portion of a bubble.
[0112] Depending on the form of the image machine learning model 608, the bubble detection system 106 can use a variety of loss functions for the loss function 612. In certain embodiments, the bubble detection system 106 uses a cross-entropy loss function (e.g., for a CNN). For example, the bubble detection system 106 can use a pixel-level cross-entropy loss function for a DenseNet or a ResNet or some other suitable loss function (e.g., pixel-level LI or L2, perceptual loss at feature part level). Regardless of the form of the loss function 612, the bubble detection system 106 determines losses 614a-614n from the loss function 612 based on a comparison of the bubble classifications 610a with the ground truth classification labels 604a. Indeed, in certain implementations, the losses 614a-614n can include individual losses for particular portions (e.g., tiles or sub-tiles) of the nucleotide sample slide.
[0113] Based on the determined losses 614a-614n from the loss function 612, the bubble detection system 106 then adjusts the parameters of the image machine learning model 608. By adjusting the parameters, the bubble detection system 106 improves the accuracy with which the image machine learning model 608 determines the presence and location of bubbles based on the spatial image through multiple training iterations. Indeed, as further shown in Figure 6A the bubble detection system 106 performs a subsequent training iteration. As suggested Figure 6A In some embodiments, the bubble detection system 106 iteratively inputs the spatial image 606b-606n into the image machine learning model 608 to generate bubble classifications 610b-610n, iteratively compares the bubble classifications 610b-610n with the ground truth classification labels 604b-604n to determine losses 614b-614n, and iteratively adjusts the parameters of the image machine learning model 608, as suggested. In some cases, the bubble detection system 106 performs training iterations until the parameters (e.g., values or weights) of the image machine learning model 608 do not significantly change across training iterations or otherwise meet a convergence criterion.
[0114] As noted above, in some embodiments, the bubble detection system 106 utilizes the image machine learning model 608 as part of identifying a training data set for a bubble detection machine learning model. Additionally or alternatively, in some embodiments, the bubble detection system 106 utilizes the image machine learning model 608 as a bubble detection machine learning model. In yet additional embodiments, the bubble detection system 106 utilizes the image machine learning model 608 in addition to a bubble detection machine learning model 622 to improve the accuracy of the generated classifications. In one example, the bubble detection system 106 utilizes the image machine learning model 608 to remove false positives generated by the bubble detection machine learning model 622.
[0115] As just described, in certain implementations, the bubble detection system 106 utilizes the image machine learning model 608 to identify or generate a training data set 620 for the bubble detection machine learning model. For example, in some cases, as part of the training data set 620, the bubble detection system 106 identifies a nucleotide base call from the nucleotide base calls 602a-602n for which the image machine learning model 608 correctly detected the presence (or absence) of a bubble within the portion (e.g., tile or sub-tile) of the nucleotide sample slide depicted by the corresponding spatial image. Having identified such a nucleotide base call from the BCL or BAM file used to train the data set 620, the bubble detection system 106 likewise identifies a corresponding ground truth classification label from the ground truth classification labels 604a-604n that correctly indicates the presence (or absence) of a bubble for the training data set 620. In some cases, the ground truth classification label is modified to correctly indicate the presence (or absence) of a bubble within the portion of the nucleotide sample slide - for the corresponding nucleotide base call selected for inclusion within the training data set 620. As Figure 6A As shown in FIG. 6B, the bubble detection system 106 selects a combination of (i) a nucleotide base call, (ii) a corresponding quality indicator, and (iii) a corresponding ground truth classification label of a spatial image to include within the training data set 620 for which the image machine learning model 608 produced a correct detection of the presence or absence of a bubble.
[0116] In an alternative to using the image machine learning model 608 to identify the training data set 620, in some implementations, as part of the training data set 620, the bubble detection system 106 identifies a nucleotide base call from the nucleotide base calls 602a-602n for which a researcher correctly detected the presence (or absence) of a bubble within the portion (e.g., tile or sub-tile) of the nucleotide sample slide depicted by the corresponding spatial image. In other words, in some implementations, the bubble detection system 106 uses the spatial images 606a-606n identified by a human with expertise (rather than the image machine learning model 608) to select a nucleotide base call from the nucleotide base calls 602a-602n for inclusion within the training data set 620. In some such cases, the bubble detection system 106 uses a nucleotide base call from a BCL or BAM file corresponding to such a spatial image having a portion containing a bubble (or no bubble) identified by the human. As Figure 6AAs shown in FIG. 6, the bubble detection system 106 alternatively selects a combination of (i) a nucleobase call, (ii) a corresponding quality indicator, and (iii) a corresponding ground truth classification label of a spatial image in which a technician or researcher correctly detected the presence or absence of a bubble to include within the training dataset 620.
[0117] Regardless of how the training dataset 620 is selected, as shown in FIG. 6, the bubble detection system 106 uses the training dataset 620 to train the bubble detection machine learning model 622. For example, in an initial training iteration, the bubble detection system 106 inputs an input matrix including a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator from the training dataset 620. Alternatively, the bubble detection system 106 inputs nucleobase calls arranged according to clusters within a portion of a nucleotide sample slide and corresponding quality indicators from the training dataset 620. Figure 6A As further shown in FIG. 6, the bubble detection system 106 utilizes the training dataset 620 to train the bubble detection machine learning model 622 (e.g., the bubble detection neural network 500 shown in FIG. 5). As indicated above, in some cases, the bubble detection system 106 utilizes a training input matrix from the training dataset 620 that includes a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. More specifically, the bubble detection system 106 generates a training input matrix that includes a subset (e.g., percentage) of adenine calls, a subset of guanine calls, and a subset of nucleobase calls that meet a threshold quality indicator (e.g., Q30) from the training dataset 620. In such embodiments, the bubble detection machine learning model 622 is trained to generate an error classification (e.g., air bubble, oil bubble, etc.). Figure 6A As further shown in FIG. 6, the bubble detection system 106 utilizes the training dataset 620 to train the bubble detection machine learning model 622 (e.g., the bubble detection neural network 500 shown in FIG. 5). As indicated above, in some cases, the bubble detection system 106 utilizes a training input matrix from the training dataset 620 that includes a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator. More specifically, the bubble detection system 106 generates a training input matrix that includes a subset (e.g., percentage) of adenine calls, a subset of guanine calls, and a subset of nucleobase calls that meet a threshold quality indicator (e.g., Q30) from the training dataset 620. In such embodiments, the bubble detection machine learning model 622 is trained to generate an error classification (e.g., air bubble, oil bubble, etc.).
[0118] In an alternative to inputting such subsets of nucleobase calls from the training dataset 620, in some embodiments, the bubble detection system 106 inputs nucleobase calls arranged according to clusters within a portion of a nucleotide sample slide and corresponding quality indicators into the bubble detection machine learning model 622. By using nucleobase calls arranged according to clusters as input to the bubble detection machine learning model 622, the bubble detection system 106 can identify patterns of nucleobase calls that indicate the presence or absence of a bubble. For example, such nucleobase calls can reflect a pattern of nucleobase calls (e.g., a circular pattern of A calls or a circular pattern of G calls) that indicates the presence of a bubble within a portion (e.g., tile or sub-tile) of a nucleotide sample slide.
[0119] Regardless of the form of the training dataset 620, as shown in FIG. 6, the bubble detection system 106 uses the training dataset 620 to train the bubble detection machine learning model 622. For example, in an initial training iteration, the bubble detection system 106 inputs an input matrix including a first subset of nucleobase calls corresponding to at least one nucleobase and a second subset of nucleobase calls that meet a threshold quality indicator from the training dataset 620. Alternatively, the bubble detection system 106 inputs nucleobase calls arranged according to clusters within a portion of a nucleotide sample slide and corresponding quality indicators from the training dataset 620. Figure 5
[0120] Based on the input data, the bubble detection machine learning model 622 determines a predicted classification label 624 indicative of the presence or absence of a bubble. In some cases, the predicted classification label 624 is indicative of the presence or absence of a particular type of particulate (e.g., air bubble, oil bubble) and a particular portion of the nucleotide sample slide. For example, the predicted classification label 624 can be indicative of the presence or absence of a bubble within a tile or sub-tile of a flow cell. As indicated above, in one or more embodiments, the bubble detection system 106 determines a confidence score corresponding to the individual classification from the predicted classification label 624. Accordingly, the bubble detection system 106 can determine the predicted classification label 624 based on the generated confidence score.
[0121] As further shown in Figure 6A the bubble detection system 106 uses a loss function 626 to compare the predicted classification label 624 to a corresponding ground truth classification label from the training data set 620. In some implementations, the ground truth classification label from the training data set 620 includes a ground truth bubble classification and bubble location corresponding to the input nucleotide base call data and quality indicators. Similar to the training process described above, for example, the ground truth classification label can be indicative of (i) a particular portion or sub-portion of the nucleotide sample slide that contains a bubble or a portion of a bubble and (ii) a particular portion or sub-portion of the nucleotide sample slide that does not contain a bubble or a portion of a bubble.
[0122] Depending on the form of the bubble detection machine learning model 622, the bubble detection system 106 can use a variety of loss functions for the loss function 626. In certain embodiments, the bubble detection system 106 uses a cross-entropy loss function (e.g., for a CNN). But any suitable loss function can be used as the loss function 626. Regardless of the form of the loss function 626, the bubble detection system 106 determines a loss 628a from the loss function 626 based on the comparison of the predicted classification label 624 to the corresponding ground truth classification label from the training data set 620. Indeed, in certain implementations, the loss 628a can include individual losses for particular portions (e.g., tiles or sub-tiles) of the nucleotide sample slide.
[0123] Based on the determined loss 628a from the loss function 626, the bubble detection system 106 then adjusts the parameters of the bubble detection machine learning model 622. By adjusting the parameters, the bubble detection system 106 improves the accuracy with which the bubble detection machine learning model 622 determines the presence and location of bubbles in multiple training iterations. Indeed, as further shown in Figure 6A the bubble detection system 106 performs a subsequent training iteration. As further shown in Figure 6AAs suggested, in some embodiments, the bubble detection system 106 iteratively inputs data derived from the nucleobase calls and the quality metrics from the training data set 620 into the bubble detection machine learning model 622 to generate predicted classification labels, iteratively compares the predicted classification labels to the corresponding ground truth classification labels from the training data set 620 to determine a loss 628a-628n, and iteratively adjusts the parameters of the bubble detection machine learning model 622. In some cases, the bubble detection system 106 performs training iterations until the parameters (e.g., values or weights) of the bubble detection machine learning model 622 do not significantly change across training iterations or otherwise meet a convergence criterion.
[0124] In addition to generating predicted classification labels, in some implementations, the bubble detection system 106 trains the bubble detection machine learning model 622 to infer the size of a bubble. Specifically, the bubble detection machine learning model 622 can extract features from the nucleobase calls of the training data set 620 to predict the size of an identified bubble. To illustrate, the bubble detection system 106 can train the bubble detection machine learning model 622 to determine a predicted diameter of a bubble based on spatial data derived from the nucleobase calls and the quality metrics. Alternatively, the bubble detection system 106 trains the bubble detection machine learning model 622 to determine the size of a bubble based on the intensity of a sharp peak or a drop in percentage of nucleobase calls or percentage of Q30. Thus, the bubble detection system 106 can train the bubble detection machine learning model 622 to generate a predicted bubble size based on an analysis of the input data.
[0125] As previously described, in some embodiments, the bubble detection system 106 reduces the quality metric (e.g., Q-score) of a given read, cycle, portion, or sub-portion of a nucleotide sample slide based on determining the presence of a bubble. In some embodiments, the bubble detection system 106 reduces the quality metric based on the size or diameter of a detected bubble. For example, the bubble detection system 106 generates a predicted diameter of a detected bubble using the bubble detection machine learning model 622 and correlates larger diameter sizes to larger reductions in the quality metric. Moreover, in some embodiments, the bubble detection system 106 determines a threshold bubble diameter value below which the bubble detection system 106 does not change the quality metric. Specifically, the bubble detection system 106 can determine that smaller bubbles have an insignificant impact on read quality.
[0126] As previously described, the bubble detection system 106 can identify or generate a spatial image that includes a spatial pattern corresponding to the nucleobase calls. Figure 6B An exemplary spatial image is shown in accordance with one or more embodiments. Specifically, Figure 6BA spatial image 636 is shown, comprising a block 640 with a spatial pattern 638. As shown, the bubble detection system 106 constructs the spatial image 636 using nucleobase detection 642. Alternatively, the bubble detection system 106 receives the spatial image 636 as a spatial image for a technician or researcher to identify bubbles within the block 640.
[0127] As previously described, in some embodiments, the bubble detection system 106 can analyze the shape of a spatial pattern identified within the spatial image 636 to determine the presence or absence of bubbles or other artifacts. For example, as Figure 6B As indicated, the bubble detection machine learning model 622 can detect circular patterns detected by G as representing bubbles. In practice, in some embodiments, the bubble detection system 106 associates circular spatial patterns detected by specific nucleobases (e.g., A or G detection) with bubbles, and associates non-circular or alternative spatial patterns with other types of artifacts. For the latter type of artifact, for example, the bubble detection system 106 can associate alternative spatial patterns with artifacts such as low-occupancy regions or amplicon regions.
[0128] To help visualize real-world examples of air bubbles within a nucleotide sample slide, this disclosure includes Figure 6C . Specifically, Figure 6C A sequencing run image 650 depicts a portion of a flow pool 658, including blocks 656a to 656c. (See image 650.) Figure 6C As shown, sequencing run image 650 depicts black circular regions corresponding to bubbles 654a to 654c that pass through or exist within different blocks. For example, Figure 6C Bubble 654b is shown spanning blocks 656a and 656b, while bubble 654c is contained within block 656c.
[0129] Figure 6C An exemplary sequencing run image is shown, illustrating the appearance of bubbles in the flow cell. As previously mentioned, accessing, storing, and processing image data is computationally expensive and often impractical. Therefore, in some embodiments, the bubble detection system 106 does not access the sequencing run image 650, but instead accesses and processes nucleobase detection data and quality metrics (from various file types) to confirm the presence or absence of bubbles, as described above.
[0130] Figures 1-6B The corresponding text and examples provide numerous different methods, systems, devices, and non-transitory computer-readable media for the bubble detection system 106. In addition to the foregoing, flowcharts (such as those illustrating actions to achieve specific results) may also be provided. Figure 7The one or more embodiments are described with reference to flowcharts whose operations are illustrated in FIGS. 1-7. Additionally, the actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar actions.
[0131] Figure 7 A flowchart illustrating a series of actions 700 for detecting the presence of air bubbles within a nucleotide sample slide is shown. While Figure 7 Actions in accordance with one embodiment are shown. Alternative embodiments can omit, add to, reorder, and / or modify Figure 7 any of the actions shown in FIGS. 1-7. Figure 7 The actions of FIG. 7 can be performed as part of a method. Alternatively, a non-transitory computer-readable medium can include instructions that, when executed by one or more processors, cause a computing device to perform the actions of FIG. 7. In some embodiments, a system can perform the actions of FIG. 7. Figure 7 Figure 7
[0132] In one or more embodiments, the series of actions 700 is implemented on one or more computing devices, such as the computing device shown in FIG. 1. Additionally, in some embodiments, the series of actions 700 is implemented in a digital environment for nucleic acid polymer sequencing. For example, the series of actions 700 is implemented on a computing device having a memory that includes an air bubble detection machine learning model. In some embodiments, the memory also stores training data that includes a ground truth classification and a training input matrix. Figure 8 As shown in FIG. 7, the series of actions 700 includes an action 702 of receiving call data. Specifically, this action 702 includes receiving call data for a nucleotide sample slide, the call data including a nucleobase call for a cycle of nucleic acid polymer sequencing. In some embodiments, this action 702 also includes receiving the call data including the nucleobase call based on: single channel intensity data including a single image of each portion of the nucleotide sample slide for a given cycle of the nucleic acid polymer sequencing; dual channel data including two images of each portion of the nucleotide sample slide for the given cycle of the nucleic acid polymer sequencing; or quad channel data including four images of each portion of the nucleotide sample slide for the given cycle of the nucleic acid polymer sequencing.
[0133] Figure 7 The series of actions 700 shown in FIG. 7 includes an action 704 of receiving quality data. Specifically, this action 704 includes receiving quality data for the nucleotide sample slide, the quality data including a quality indicator that estimates an error in the nucleobase call for the cycle.
[0134] Figure 7
[0135] The series of acts 700 includes an act 706 of determining a first subset of nucleobase detections and a second subset of nucleobase detections. Specifically, the act 706 includes determining, from the nucleobase detections of the cycle, a first subset of the nucleobase detections corresponding to at least one nucleobase and a second subset of the nucleobase detections meeting a threshold quality indicator of the quality indicators. In some embodiments, the act 706 further includes determining the first subset of the nucleobase detections corresponding to the at least one nucleobase by determining at least one of a subset of adenine detections, a subset of thymine detections, a subset of cytosine detections, or a subset of guanine detections of the cycle of the nucleic acid polymer sequencing.
[0136] As further shown in Figure 7 The series of acts 700 includes an act 708 of detecting, with a bubble detection neural network, a presence of a bubble. Specifically, the act 708 includes detecting, with a bubble detection machine learning model based on the first subset of nucleobase detections and the second subset of nucleobase detections, a presence of a bubble within the nucleotide sample slide. Additionally, in one or more embodiments, the bubble detection neural network includes at least one of a support vector machine or an adaptive boosting machine learning model.
[0137] In some implementations, the act 708 further includes detecting, with the bubble detection machine learning model, the presence of the bubble by extracting features from an input matrix with a layer of the bubble detection machine learning model, the input matrix including the subset of adenine detections, the subset of guanine detections, and the second subset of nucleobase detections meeting the threshold quality indicator of the cycle of the nucleic acid polymer sequencing. Moreover, in one or more embodiments, the act 708 includes detecting the presence of a bubble by detecting at least one of an air bubble, an oil bubble, or a ghost bubble within the nucleotide sample slide. Additionally, in some embodiments, the bubble detection machine learning model includes a convolutional neural network including a feature extraction layer, a classification layer, and an adaptive max pooling layer between the feature extraction layer and the classification layer.
[0138] In one or more embodiments, the act 708 further includes additional acts of detecting the presence of the bubble by: generating, with the bubble detection machine learning model, a probability that a portion of the nucleotide sample slide contains the bubble; and determining that the probability meets a threshold indicative of the presence of the bubble.
[0139] In some embodiments, the series of acts 700 includes additional acts of receiving detection data and quality data for a portion of a nucleotide sample slide and detecting a presence of a bubble within a portion of a nucleotide sample slide. More specifically, in some embodiments, the additional acts further include detecting the presence of a bubble within a portion of a nucleotide sample slide by detecting a bubble within a zone of a flow cell.
[0140] Additionally, in some embodiments, the series of acts 700 further includes the additional act of determining the presence of a bubble during one or more of the cycles of the sequencing of the nucleic acid polymer.
[0141] Further, in one or more embodiments, the series of acts 700 further includes the act of providing an alert for display on the computing device, the alert indicating the presence of a bubble within the nucleotide sample slide.
[0142] Additionally, in some embodiments, the series of acts 700 includes the additional act of determining the presence of a bubble during one of the cycles of the sequencing of the nucleic acid polymer.
[0143] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which the nucleic acid is attached to a fixed position in an array such that its relative position does not change and in which the array is repeatedly imaged. Embodiments in which images are obtained in different color channels (e.g., in accord with different labels used to distinguish one type of nucleotide base from another) are particularly suitable. In some embodiments, the process of determining the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) technologies.
[0144] SBS technologies generally include the enzymatic extension of a nascent nucleic acid strand by repeated addition of nucleotides against a template strand. In traditional SBS methods, a single nucleotide monomer can be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0145] The SBS technologies described below can utilize single-end sequencing or paired-end sequencing. In single-end sequencing, a sequencing device reads a fragment from one end to the other to generate a sequence of base pairs. In contrast, during paired-end sequencing, a sequencing device begins with one read, completes the read of a particular read length in the same direction, and begins another read from the opposite end of the fragment.
[0146] SBS can utilize nucleotide monomers with a terminator moiety or nucleotide monomers lacking any terminator moiety. Methods that use nucleotide monomers lacking terminators include, for example, pyrophosphate sequencing and sequencing using gamma-phosphate labeled nucleotides, as described in further detail below. In methods that use nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is typically variable, and the number depends on the template sequence and the manner of nucleotide delivery. For SBS techniques that utilize nucleotide monomers with a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used, as in the case of traditional Sanger sequencing with dideoxynucleotides, or the terminator can be reversible, as in the case of sequencing methods developed by Solexa (now Illumina, Inc.).
[0147] SBS techniques can utilize nucleotide monomers with a label moiety or nucleotide monomers lacking a label moiety. Thus, incorporation events can be detected based on the properties of the label, such as fluorescence of the label; properties of the nucleotide monomer, such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; and the like. In embodiments in which two or more different nucleotides are present in the sequencing reagents, the different nucleotides can be distinguishable from one another, or alternatively, two or more different labels can be indistinguishable under the detection technology used. For example, different nucleotides present in the sequencing reagents can have different labels, and they can be distinguished using appropriate optics, as exemplified by sequencing methods developed by Solexa (now Illumina, Inc.).
[0148] A preferred embodiment includes pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as specific nucleotides are incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) “Real-time DNA sequencing using detection of pyrophosphate release.” Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) “Pyrosequencing sheds light on DNA sequencing.” Genome Res., 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) “A sequencing method based on real-time pyrophosphate.” Science 281(5375), 363; U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; and U.S. Patent No. 6,274,320, the disclosures of which are incorporated by reference herein in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to ATP by an adenosine triphosphate (ATP) sulfurylase, and the level of ATP produced is detected by the photons produced by a luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal produced as nucleotides are incorporated at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C, or G). The images obtained after each nucleotide type is added will differ in terms of which features of the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature will remain unchanged in the images. The images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after the array is treated with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for a sequencing-by- inversion method.
[0149] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides that comprise, for example, cleavable or photobleavable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method is commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, the disclosure of each of which is incorporated herein by reference. The availability of fluorescently labeled terminators, where termination can be reversible and the fluorescent label can be cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0150] Preferably, in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured after the labels are incorporated into the arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has a label that is spectrally distinct. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid features that have incorporated a particular type of nucleotide. Different features will be present or absent in different images due to the different sequence content of each feature. However, the relative positions of the features will remain unchanged in the images. Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the labels can be removed and the reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and before the subsequent cycle can provide the advantage of reducing background signal and cross-talk between cycles. Examples of labels and removal methods that can be used are set forth below.
[0151] In certain embodiments, some or all of the nucleotide monomers can include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore can include a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15: 1767-1776 (2005), which is incorporated herein by reference). Other methods have decoupled terminator chemistry from cleavage of fluorescent labels (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005), which is incorporated by reference in its entirety). Ruparel et al. describe the development of reversible terminators that use a small 3' allyl group to block extension, but which can be easily deblocked by a short treatment with a palladium catalyst. The fluorophore is attached to the base via a photo-cleavable linker that can be easily cleaved by exposure to long wavelength UV light for 30 seconds. Thus, disulfide reduction or photo-cleavage can be used as a cleavable linker. Another approach to reversible termination is to use natural termination that occurs after placing a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as a highly efficient terminator by steric and / or electrostatic blockage. The presence of one incorporation event prevents further incorporation unless the dye is removed. Cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent No. 7,427,673 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated by reference herein in their entireties.
[0152] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated by reference herein in their entireties.
[0153] Some embodiments can use fewer than four different labels to use detection of four different nucleotides. For example, SBS can be performed with the methods and systems described in the incorporated material of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity of one member of the pair relative to the other, or based on a change (e.g., by chemical modification, photochemical modification, or physical modification) of one member of the pair that results in a clear signal appearance or disappearance compared to the signal of the other member of the pair that is detected. As a second example, three of the four different nucleotide types are capable of being detected under certain conditions, while the fourth nucleotide type lacks a label that is detectable under those conditions or that is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). The first three nucleotide types can be determined to be incorporated into the nucleic acid based on the presence of their respective signals, and the fourth nucleotide type can be determined to be incorporated into the nucleic acid based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type can include labels that are detected in two different channels, while the other nucleotide types are detected in no more than one channel. The three exemplary configurations described above are not considered to be mutually exclusive, and can be used in various combinations. An exemplary embodiment that combines all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP with a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP with a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first channel and the second channel (e.g., dTTP with at least one label that is detected in both channels when excited by the first excitation wavelength and / or the second excitation wavelength), and a fourth nucleotide type that lacks a label that is detected in either channel (e.g., dGTP without a label).
[0154] Further, as described in the incorporated material of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called single-dye sequencing method, a first nucleotide type is labeled, but the label is removed after the first image is generated, and only a second nucleotide type is labeled after the first image is generated. A third nucleotide type retains its label in both the first image and the second image, and a fourth nucleotide type remains unlabeled in both images.
[0155] Some embodiments can utilize edge-sequencing-by-ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and determine incorporation of such oligonucleotides. The oligonucleotides typically have different labels related to the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show nucleic acid features that have incorporated a particular type of label. Due to the different sequence content of each feature, different features will be present or absent in different images, but the relative positions of the features will remain unchanged in the images. Images obtained by connection-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent No. 6,969,488, U.S. Patent No. 6,172,218, and U.S. Patent No. 6,306,597, the disclosures of which are incorporated by reference herein in their entireties.
[0156] Some embodiments can utilize nanopore sequencing (Deamer, D.W. and Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope", Nat. Mater., 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, a target nucleic acid is passed through a nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as a-hemolysin. As the target nucleic acid is passed through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductivity of the pore. (U.S. Patent No. 7,001,792; Soni, G.V. and Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K., "Nanopore-based single-molecule DNA analysis.", Nanomed., 2, 459-481 (2007); Cockroft, S.L., Cgu, J., Amorin, M. and Ghadiri, M.R., "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.", J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. In particular, data can be processed as images, in accordance with the exemplary processing of optical images and other images described herein.
[0157] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interactions between a polymerase carrying a fluorophore and a gamma-phosphate labeled nucleotide, as described, for example, in U.S. Patent No. 7,329,492 and U.S. Patent No. 7,211,414 (each of which is incorporated herein by reference), or nucleotide incorporation can be detected with zero-mode waveguides, as described, for example, in U.S. Patent No. 7,315,019 (which is incorporated herein by reference), and nucleotide incorporation can be detected using fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). Illumination can be limited to a volume on the order of femtoliters around surface-tethered polymerases, such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.” Science 299, 682-686 (2003); Lundquist, P. M. et al., “Parallel confocal detection of single molecules in real time.” Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures.” Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein in their entireties by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0158] Some SBS embodiments include detecting protons released upon nucleotide incorporation extension products. For example, sequencing based on detection of released protons can use electrical detectors and related technology commercially available from Ion Torrent, Inc. (Guilford, CT, which is a subsidiary of Life Technologies) or sequencing methods and systems described in US 2009 / 0026082 Al, US 2009 / 0127589 Al, US 2010 / 0137143 Al, or US 2010 / 0282617 Al, each of which is incorporated herein by reference. The methods set forth herein using kinetic exclusion to amplify target nucleic acids can be readily applied to substrates for detecting protons. More specifically, the methods set forth herein can be used to generate clonal populations of amplicons for detecting protons.
[0159] The SBS methods described above can advantageously be performed in a variety of formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle, or binding to a polymerase or other molecule attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature), or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail below.
[0160] The methods described herein can use an array having features at any of a variety of densities, including, for example, at least about 10 features / cm 2 , 100 features / cm 2 , 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 or higher.
[0161] The methods set forth herein have the advantage that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Accordingly, the present disclosure provides integrated systems that are capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Thus, the integrated systems of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, including components such as pumps, valves, reservoirs, fluidic lines, etc. The flow cell can be configured for and / or used in detecting target nucleic acids in the integrated system. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and U.S. Serial No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for the flow cell, one or more fluidic components of the integrated system can be used for the amplification methods and the detection methods. By way of nucleic acid sequencing embodiments, one or more fluidic components of the integrated system can be used for the amplification methods set forth herein and for delivering sequencing reagents in sequencing methods such as those exemplified above. Alternatively, the integrated system can include separate fluidic systems to perform the amplification methods and to perform the detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and also determining nucleic acid sequences include, but are not limited to, the MiSeq® platform (Illumina, Inc., San Diego, CA) and the apparatus described in U.S. Serial No. 13 / 273,666, which is incorporated herein by reference. TM The MiSeq® platform (Illumina, Inc., San Diego, CA) and the apparatus described in U.S. Serial No. 13 / 273,666, which is incorporated herein by reference.
[0162] The sequencing systems described above sequence nucleic acid polymers present in a sample received by the sequencing apparatus. As defined herein, "sample" and its derivatives are used in their broadest sense to include any specimen, culture, etc. suspected of containing a target. In some embodiments, the sample includes nucleic acids in DNA, RNA, PNA, LNA, chimeric or hybridized form. The sample can include any biological, clinical, surgical, agricultural, atmospheric, or aquatic based specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh-frozen or formalin-fixed paraffin-embedded nucleic acid specimens. It is also contemplated that the source of the sample can be: a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, a nucleic acid sample (matched thereto) from a single individual (such as a tumor sample and a normal tissue sample), or a sample from a single source containing two different forms of genetic material (such as maternal DNA and fetal DNA obtained from a maternal subject), or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of the nucleic acid material can include nucleic acids obtained from a neonate, for example, nucleic acids typically used for neonatal screening.
[0163] The nucleic acid sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another embodiment, the low molecular weight material includes enzymatically fragmented or mechanically fragmented DNA. The sample can comprise cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resections, and other clinically or laboratory obtained samples. In some embodiments, the sample can be an epidemiological sample, an agricultural sample, a forensic sample, or a pathogenic sample. In some embodiments, the sample can include nucleic acid molecules obtained from animals, such as human or mammalian sources. In another embodiment, the sample can include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecules can be an archived or extinct sample or species.
[0164] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples having low quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, the forensic sample can include nucleic acid obtained from a crime scene, nucleic acid obtained from a missing persons DNA database, nucleic acid obtained from a laboratory associated with a forensic investigation, or include a forensic sample obtained by law enforcement, one or more military services, or any such personnel. The nucleic acid sample can be a purified sample or a lysate containing crude DNA, for example, derived from a buccal swab, paper, fabric, or other substrate that can be impregnated with saliva, blood, or other bodily fluid. Thus, in some embodiments, the nucleic acid sample can comprise a small amount of DNA (such as genomic DNA), or a fragmented portion of DNA. In some embodiments, the target sequence can be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence can be obtained from a victim's hair, skin, tissue sample, autopsy, or remains. In some embodiments, the nucleic acid comprising one or more target sequences can be obtained from a deceased animal or human. In some embodiments, the target sequence can include nucleic acid obtained from non-human DNA, such as microbial, plant, or insect DNA. In some embodiments, the target sequence or amplified target sequence is directed toward the purpose of human identity identification. In some embodiments, the disclosure generally relates to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure generally relates to methods of human identity identification using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identity identification sample containing at least one target sequence can be amplified using any one or more target-specific primers disclosed herein or using the primer criteria outlined herein.
[0165] The components of the bubble detection system 106 can include software, hardware, or both. For example, the components of the bubble detection system 106 can include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., the user client device 108). When executed by the one or more processors, the computer-executable instructions of the bubble detection system 106 can cause the computing device to perform the bubble detection methods described herein. Alternatively, the components of the bubble detection system 106 can include hardware, such as a specialized processing device to perform certain functions or groups of functions. Additionally or alternatively, the components of the bubble detection system 106 can include a combination of computer-executable instructions and hardware.
[0166] Additionally, components of the bubble detection system 106 performing the functions described herein with respect to the bubble detection system 106 can be implemented, for example, as part of a standalone application, as a module of an application, as a plug-in of an application, as a library function or function that can be called by other applications, and / or as a cloud computing model. Thus, components of the bubble detection system 106 can be implemented as part of a standalone application on a personal computing device or mobile device. Additionally or alternatively, components of the bubble detection system 106 can be implemented in any application that provides sequencing services, including but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TmSight” are registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0167] As discussed in greater detail below, embodiments of the present disclosure can include or utilize a special-purpose or general-purpose computer that includes computer hardware, such as, for example, one or more processors and system memory. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer- executable instructions and / or data structures. In particular, one or more of the processes described herein can be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium (e.g., memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0168] A computer-readable medium can be any available medium or means that can be accessed by a general purpose or special purpose computing system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Accordingly, embodiments of the present disclosure can include at least two distinct kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0169] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0170] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmission media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0171] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received by way of network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Accordingly, it should be understood that non-transitory computer-readable storage media (devices) can be included in computing system components that also (or even primarily) utilize transmission media.
[0172] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer- executable instructions are executed on a general purpose computer to transform the general purpose computer into a special purpose computer that performs elements of the present disclosure. Computer- executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0173] Those skilled in the art will appreciate that the disclosure can be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0174] Embodiments of the disclosure can also be implemented in a cloud computing environment. In this description and the following claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, a consumer can request and be granted near-instantaneous access to or sharing of cloud computing resources. A cloud computing model can be implemented to provide availability of the shared pool resources and services to a consuming entity which can include an entity with on-going computing processing needs which is implemented with or without associated deployed software, such as an entity implementing a mission critical business application function. Through this model, even advanced communications functions, which can include a combination of software functioning, resource allocation, load balancing, replication, management, and other cloud computing functions, can be obtained by a consuming entity as needed, and then released when no longer needed, alleviating the necessity of a consumer to purchase, install, and manage the underlying cloud computing system hardware and software.
[0175] A cloud computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and the like. A cloud computing model can also exhibit various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). A cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and the like. In this description and the following claims, a "cloud computing environment" is an environment in which cloud computing is employed.
[0176] Figure 8 A block diagram illustrating a computing device 800, which can be configured to perform one or more of the processes described above, is shown. One will appreciate that one or more computing devices, such as computing device 800, can implement bubble detection system 106 and sequencing system 104. As shown, computing device 800 can include a processor 802, a memory 804, a storage device 806, an I / O interface 808, and a communication interface 810, which can be communicatively coupled by way of a communication infrastructure 812. In certain embodiments, computing device 800 can include fewer or more components than those shown in FIG. 8. The components of computing device 800 shown in FIG. 8 are depicted as being Figure 8 on a single computing device, but one will appreciate that the components of computing device 800 can be distributed across multiple computing devices. Figure 8 The following paragraphs describe in more detail the components of computing device 800 shown in FIG. 8. Figure 8 The following paragraphs describe in more detail the components of computing device 800 shown in FIG. 8.
[0177] In one or more embodiments, the processor 802 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, the processor 802 can fetch, decode, and execute, to name but a few, instructions for dynamically modifying a workflow. The processor 802 can retrieve (or fetch) the instructions from, for example, an internal register, an internal cache, the memory 804, or the storage device 806, and decode and execute them. The memory 804 can be a volatile or non-volatile memory for storing data, metadata, and program instructions to be executed by the processor. The storage device 806 includes a storage device for storing data or instructions for performing the methods described herein, such as a hard disk, a flash drive, or other digital storage device.
[0178] The I / O interface 808 allows a user to provide input to, receive output from, and otherwise transfer data to and from the computing device 800. The I / O interface 808 can include a mouse, a keypad, or a keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other well-known I / O devices, or combinations of such I / O interfaces. The I / O interface 808 can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (for example, a display screen), one or more output drivers (for example, display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 808 is configured to provide graphical data to a display for presentation to a user. The graphical data can represent one or more graphical user interfaces and / or any other graphical content serving a particular implementation.
[0179] The communication interface 810 can include hardware, software, or both. In any case, the communication interface 810 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 800 and one or more other computing devices or networks. As an example and not by way of limitation, the communication interface 810 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network.
[0180] Additionally, the communication interface 810 can facilitate communication with various types of wired or wireless networks. The communication interface 810 can also facilitate communication using various communication protocols. The communication infrastructure 812 can also include hardware, software, or both providing communication facilities used by the components of the computing device 800 to each other and to other devices, including other computing devices coupled to the network(s) 814. For example, the communication interface 810 can enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein using one or more networks and / or protocols. To illustrate, a sequencing process can allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0181] In the foregoing specification, the disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the disclosure are described with reference to details discussed, and the accompanying drawings illustrate the various embodiments. The description above and the drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the disclosure.
[0182] The disclosure can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described implementations are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein can be performed with fewer or additional steps / actions, or in different orders. Additionally, the steps / actions described herein can be repeated or performed in parallel with each other or with different instances of the same or similar steps / actions. Accordingly, the scope of the application is indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1. A system comprising: At least one processor; and A non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: The detection data is received from a nucleotide sample slide, and the detection data includes the detection of cyclic nucleobases from nucleic acid polymer sequencing. Quality data, including quality indicators, are received for the nucleotide sample slide. The quality index estimates the error in the nucleobase detection of the cycle; From the nucleobase detections in the cycle, determine a first subset of nucleobase detections corresponding to at least one nucleobase and a second subset of nucleobase detections that meet the threshold quality index of the quality index; as well as The presence of bubbles in the nucleotide sample slide is detected using a bubble detection machine learning model based on the first subset and the second subset of nucleobases detected.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: The detection data and the quality data received from the portion of the nucleotide sample slide; and The presence of the air bubbles is detected within the portion of the nucleotide sample slide.
3. The system of claim 2, further comprising instructions that, when executed by at least one processor, cause the system to detect the presence of the bubbles in the portion of the nucleotide sample slide by detecting the bubbles in a block of the flow cell.
4. The system of claim 1, further comprising instructions that, when executed by at least one processor, cause the system to determine the first subset of nucleobase detection corresponding to the at least one nucleobase by determining at least one of adenine detection subset, thymine detection subset, cytosine detection subset, or guanine detection subset of the cycle of nucleic acid polymer sequencing.
5. The system of claim 4, further comprising instructions that, when executed by at least one processor, cause the system to detect the presence of the bubble by extracting features from an input matrix using layers of the bubble detection machine learning model, the input matrix comprising a subset of adenine detection, a subset of guanine detection, and a second subset of nucleobase detection conforming to the threshold quality metric of the cycle of the nucleic acid polymer sequencing.
6. The system of claim 1, further comprising instructions that, when executed by at least one processor, cause the system to detect the presence of the bubble by detecting at least one of an air bubble, an oil bubble, or a ghost bubble within the nucleotide sample slide.
7. The system according to claim 1, wherein the bubble detection machine learning model comprises a convolutional neural network, the convolutional neural network comprising a feature extraction layer, a classification layer, and an adaptive max pooling layer between the feature extraction layer and the classification layer.
8. The system of claim 1, further comprising instructions that, when executed by at least one processor, cause the system to detect the presence of the bubble by: The probability that a portion of the nucleotide sample slide contains the air bubble is generated using the bubble detection machine learning model; and The probability is determined to meet the threshold indicating the presence of the bubble.
9. The system of claim 1, further comprising instructions that, when executed by at least one processor, cause the system to receive the detection data including the nucleobase detection based on: Single-channel data, which includes a single image of each portion of the nucleotide sample slide for a given cycle of sequencing the nucleic acid polymer; Dual-channel data, the dual-channel data comprising two images of each portion of the nucleotide sample slide for a given cycle of sequencing the nucleic acid polymer; or Four-channel data, comprising four images of each portion of the nucleotide sample slide for a given cycle of sequencing the nucleic acid polymer.
10. The system of claim 1, further comprising instructions that, when executed by at least one processor, cause the system to determine the presence of the bubble during one or more cycles of the nucleic acid polymer sequencing cycle.
11. A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause a computing device to: The detection data is received from a nucleotide sample slide, and the detection data includes the detection of cyclic nucleobases from nucleic acid polymer sequencing. For the nucleotide sample slide, quality data including quality indicators are received, which estimate the error in the detection of the nucleobases in the cycle; From the nucleobase detections in the cycle, determine a first subset of nucleobase detections corresponding to at least one nucleobase and a second subset of nucleobase detections that meet the threshold quality index of the quality index; as well as The presence of bubbles in the nucleotide sample slide is detected using a bubble detection machine learning model based on the first subset and the second subset of nucleobases detected.
12. The non-transitory computer-readable medium of claim 11, wherein the bubble detection machine learning model comprises at least one of a support vector machine or an adaptive augmentation machine learning model.
13. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by at least one processor, cause the computing device to provide an alarm for display on the computing device indicating the presence of the bubble within the nucleotide sample slide based on the detection of the bubble's presence.
14. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by the at least one processor, cause the computing device to: The detection data and the quality data received from the portion of the nucleotide sample slide; and The presence of the air bubbles is detected within the portion of the nucleotide sample slide.
15. The non-transitory computer-readable medium of claim 14, further comprising instructions that, when executed by at least one processor, cause the computing device to detect the presence of the bubble in the portion of the nucleotide sample slide by detecting the bubble in a block of the flow cell.
16. The non-transitory computer-readable medium of claim 11, further comprising instructions that, when executed by at least one processor, cause the computing device to determine the presence of the bubble during one cycle of the nucleic acid polymer sequencing cycle.
17. A computer-implemented method, the method comprising: The detection data is received from a nucleotide sample slide, and the detection data includes the detection of cyclic nucleobases from nucleic acid polymer sequencing. For the nucleotide sample slide, quality data including quality indicators are received, which estimate the error in the nucleotide detection during the cycle: From the nucleobase detections in the cycle, determine a first subset of nucleobase detections corresponding to at least one nucleobase and a second subset of nucleobase detections that meet the threshold quality index of the quality index; as well as The presence of bubbles in the nucleotide sample slide is detected using a bubble detection machine learning model based on the first subset and the second subset of nucleobases detected.
18. The computer-implemented method of claim 17, wherein determining the first subset of nucleobase detection corresponding to the at least one nucleobase comprises determining at least one of adenine detection subset, thymine detection subset, cytosine detection subset, or guanine detection subset of the cycle of nucleic acid polymer sequencing.
19. The computer-implemented method of claim 17, further comprising modifying a quality index for nucleobase detection based on detecting the presence of the bubble using the bubble detection machine learning model.
20. The computer-implemented method of claim 17, wherein detecting the presence of the bubble comprises detecting at least one of an air bubble, an oil bubble, or a ghost bubble within the nucleotide sample slide.
Citation Information
Patent Citations
Method of nucleic acid amplification
US20050100900A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1