To facilitate the secure execution of external workflows for genome sequencing diagnostics.
The diagnostic workflow system addresses inflexibility, security, and inefficiency by using a container orchestration engine to securely execute external workflows, ensuring data integrity and regulatory compliance while optimizing resource allocation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2022-09-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing diagnostic systems are inflexible, insecure, and inefficient in performing genomic diagnostics, often requiring platform-specific applications and exposing sequencing data to vulnerabilities, leading to computational inefficiencies and failure to meet regulatory standards.
A diagnostic workflow system utilizing a container orchestration engine to execute external sequencing diagnostic workflows, isolating workflow containers to maintain data security and efficiently allocate computing resources, allowing flexible execution of third-party applications while adhering to regulatory standards.
The system enhances flexibility, security, and computational efficiency by isolating workflow containers, preventing data exposure, and optimizing resource use, thus meeting regulatory standards and improving diagnostic performance.
Smart Images

Figure 0007857410000001 
Figure 0007857410000002 
Figure 0007857410000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit and priority of U.S. Patent Application No. 17 / 935476, filed on September 26, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 293587, filed on December 23, 2021. The foregoing applications are hereby incorporated by reference in their entirety.
Background Art
[0002] In recent years, biotechnology companies and computer science institutions have been improving the hardware and software for generating diagnoses about the nucleotide sequences of genomic samples. In particular, some existing diagnostic platforms generate nucleotide - base calls from nucleotide reads of a sample nucleotide sequence and / or perform diagnoses on nucleotide - base calls for various purposes. For example, existing diagnostic systems implement diagnostic applications (e.g., cancer screening assays) for screening nucleotide sequences for cancer by detecting specific gene markers within the nucleotide - base calls of a sample sequence. Some existing diagnostic systems similarly perform other diagnoses, such as genetic tests for other genetic conditions (or the tendency to develop genetic conditions) or for determining other genetic traits.
[0003] Despite these recent advances, existing diagnostic systems continue to exhibit numerous shortcomings or disadvantages. For example, many conventional diagnostic systems are rigidly tied to a specific set of internal diagnostic applications, limiting their scope and usefulness. In fact, many conventional systems can only perform genomic diagnostics using applications specifically designed for and installed on the system. Therefore, when a bioinformatics practitioner requires a specific diagnostic analysis of nucleotide sequences, existing systems may be unable to provide the necessary analytical data if a platform-specific diagnostic application for the analysis has not yet been written and / or installed within the system.
[0004] Apart from their lack of flexibility, some conventional diagnostic systems exhibit coding or network vulnerabilities that compromise or expose private or confidential genetic data. More specifically, in the case of existing systems that attempt to facilitate the integration of external diagnostic applications, these systems often compromise the security of sequencing data (and other information) in exchange for the flexibility that allows external workflows (including internal and / or external applications) to access sequencing data to perform diagnostics. In fact, some conventional systems are vulnerable to harmful external diagnostic workflows that maliciously or unintentionally damage or corrupt sequencing data related to nucleotide-based calls and / or sequencing data related to other diagnostic applications. As a result, many of these existing diagnostic systems fail to meet one or more diagnostic criteria required by various regulatory bodies (e.g., in vitro diagnostic criteria set by the U.S. Food and Drug Administration).
[0005] Additionally, some conventional diagnostic systems are inefficient. More specifically, existing systems often consume computing resources inefficiently when performing genomic diagnostic workflows. For example, some existing systems lack internal programming or other considerations for computing resource management, instead flooding available processors and memory with numerous (simultaneous) requests and instructions for the many processes involved in generating diagnostic data (often resulting in backlogs and slowdowns). Consequently, such existing systems are slow to produce the requested diagnostic results for the diagnostic workflow, or are unable to produce results completely, instead resulting in computational errors. [Overview of the project]
[0006] This disclosure describes a method, a non-temporary computer-readable medium, and embodiments of a system that can flexibly, securely, and efficiently facilitate the execution of external workflows for the diagnostic analysis of nucleotide sequencing data. For example, the disclosed system may utilize a container orchestration engine to enable an external system (e.g., a third-party system) to generate and implement a workflow for analyzing sequencing data (e.g., sequencing data generated by a sequencing instrument and / or variant analysis model). In some cases, the disclosed system further utilizes the container orchestration engine to generate sequencing data, such as nucleotide base calls, in order to implement a diagnostic workflow for analyzing the sequencing data. For example, the disclosed system may utilize the container orchestration engine to identify workflow containers that compartmentally define the individual functions of the workflow. In some such cases, the disclosed system may require the container orchestration engine to perform an external sequencing diagnostic workflow outside of the workflow for the variant analysis model and to implement the external sequencing diagnostic workflow in the workflow container. When executing individual workflow containers, the disclosed system can also isolate the workflow containers to prevent access to or corruption of sequencing data or other workflow data. [Brief explanation of the drawing]
[0007] For a detailed explanation, please refer to the diagrams briefly described below. [Figure 1] A block diagram of a sequencing system including a diagnostic workflow system, according to one or more embodiments, is shown. [Figure 2] An overview of performing an external sequencing diagnostic workflow using one or more embodiments is illustrated. [Figure 3]An exemplary flow for performing an external sequencing diagnostic workflow on sequencing data using a container orchestration engine, according to one or more embodiments, is illustrated. [Figure 4] The diagram illustrates an exemplary depiction of security permissions that a diagnostic workflow system assigns to a workflow container of an external sequencing diagnostic workflow, according to one or more embodiments. [Figure 5] This diagram illustrates the use of a container orchestration engine to schedule the execution of workflow containers in one or more embodiments. [Figure 6] An exemplary flow for determining whether a diagnostic application is compatible with the execution mode of a container orchestration engine is illustrated, based on one or more embodiments. [Figure 7] An illustrative architectural diagram of a diagnostic workflow system related to an overall sequencing environment, according to one or more embodiments, is shown. [Figure 8] A flowchart illustrating a series of actions for performing an external sequencing diagnostic workflow using one or more embodiments is provided. [Figure 9] An exemplary block diagram of a computing device for implementing one or more embodiments of the present disclosure is shown. [Figure 10] A block diagram illustrating an exemplary optical system for image-based genome sequencing, according to one or more embodiments, is shown. [Figure 11] An exemplary imaging apparatus for image-based genome sequencing, according to one or more embodiments, is illustrated. [Figure 12] Exemplary diagrams illustrating image-based genome sequencing using one or more embodiments are provided. [Modes for carrying out the invention]
[0008] This disclosure describes embodiments of a diagnostic workflow system that facilitate the execution of external workflows for the diagnostic analysis of nucleotide sequencing data. In particular, the diagnostic workflow system can identify or generate sequencing data, such as nucleotide reads and nucleotide base calls, for sample nucleotide sequences. In addition, the diagnostic workflow system can perform different diagnoses on sequencing data from either an internal workflow or an external workflow (including sequencing diagnostic applications) to analyze the sequencing data. In some cases, the diagnostic workflow system implements a diagnostic workflow designed to analyze sequencing data to identify genetic markers for a particular disease or genetic trait of a sample. In certain embodiments, the diagnostic workflow is requested and / or implemented internally (e.g., from a third-party device outside the diagnostic workflow system's server). In other embodiments, the diagnostic workflow system requires the implementation of an external sequencing diagnostic workflow, which includes an application supplied from outside the diagnostic workflow system (e.g., from an external system or another system) for applications and / or native variant analysis models adapted to the diagnostic workflow system. Nevertheless, when implementing an external sequencing diagnostic workflow, the diagnostic workflow system can maintain data security and integrity by utilizing a container orchestration engine to isolate or silo individual workflow containers for executing individual workflow functions without exposing sequencing data.
[0009] As described above, in certain implementations, a diagnostic workflow system requires that diagnostic analysis be performed on genomic sequences using an external sequencing diagnostic workflow. For example, a diagnostic workflow system requires that a server implement an external sequencing diagnostic workflow developed by an external entity, rather than being part of a set of internal workflows for variant analysis models. In further embodiments, in some cases, a diagnostic workflow system receives requests from client devices associated with an external system to process sequencing data and perform diagnostic analysis (e.g., on disease or genetic conditions). To facilitate external system access to internally generated sequencing data, a diagnostic workflow system may utilize a container orchestration engine designed to isolate and implement individual workflow containers. Specifically, a diagnostic workflow system may apply a container orchestration engine that allows a server or other computing device to access the diagnostic workflow system's services and sequencing data for tasks of the external sequencing diagnostic workflow, with granular control over how the sequencing data is used in each individual process.
[0010] As described, the diagnostic workflow system can run variant analysis models and utilize a container orchestration engine housed on a server device located on a local network along with the genome sequencing processor. Within a closed environment of the server device, including the container orchestration engine and / or variant analysis models, as well as a local connection to the genome sequencing processor, the diagnostic workflow system can securely perform sequencing diagnoses and execute analytical protocols, even allowing external applications to access diagnostic results without the risk of corruption or exposure of sequencing data. For example, the diagnostic workflow system can provide access permissions to sequencing data (or other data) on a basis specifically tailored to individual workflow containers (in relation to different types of data) to perform diagnoses, including external sequencing diagnostic workflows from external applications. Depending on the execution mode of the container orchestration engine, the diagnostic workflow system can identify compatible diagnostic applications that satisfy security and analytical standards set by regulatory bodies for in vitro diagnostics (IVD) (e.g., standards set by the U.S. Food and Drug Administration or any other agency). The diagnostic workflow system can also identify compatible diagnostic applications for investigational use only (IUO) and research use only (RUO) analyses.
[0011] In some embodiments, a diagnostic workflow system schedules or allocates computing resources to execute diagnostic workflows, such as external sequencing diagnostic workflows. For example, the diagnostic workflow system utilizes a container orchestration engine to allocate computing resources, such as processing power and memory, to execute the processes of individual workflow containers. More specifically, the diagnostic workflow system may communicate with genome sequencing processors, such as field programmable gate arrays (FPGAs) or central processing units (CPUs), which are designed and programmed to perform functions for the diagnostic workflows. In some cases, the genome sequencing processor can execute a limited number of processes at a time (e.g., one). As a result, the diagnostic workflow system utilizes the resource allocation capabilities of the container orchestration engine to schedule the execution of workflow containers for external sequencing diagnostic workflows according to the available computing resources (e.g., executing them sequentially as each process completes for each container).
[0012] As suggested above, embodiments of the diagnostic workflow system offer several advantages, benefits, and / or improvements over existing diagnostic systems. For example, in some embodiments, the diagnostic workflow system is more flexible than many existing diagnostic systems. While many existing systems are limited to implementing system-specific internal diagnostic applications and workflows to perform diagnostics on genome sequencing data, the diagnostic workflow system can facilitate the execution of external (e.g., third-party) sequencing diagnostic applications and workflows. Specifically, the diagnostic workflow system can utilize a container orchestration engine to enable the reception of external sequencing diagnostic workflows (e.g., via upload) and execute the external sequencing diagnostic workflow to perform diagnostic analysis on nucleotide base calls (or other sequencing data) of sample nucleotide sequences (as directed by an external, locally unnetworked system). In some embodiments, the diagnostic workflow system can even flexibly execute external sequencing diagnostic workflows during (or after) the sequencing process to generate nucleotide base calls of sample nucleotide sequences (unlike conventional systems).
[0013] In addition to increased flexibility, in some embodiments, diagnostic workflow systems offer improved data security compared to conventional diagnostic systems. While some conventional systems may attempt to facilitate diagnostic analysis from external applications, these conventional systems risk exposure and corruption of sequencing data or code for diagnostic applications when implementing external (or non-native) workflows for variant analysis models or other internal applications used for diagnosis. In fact, existing diagnostic systems lack mechanisms or models for securely running third-party or other external applications on computing devices that run software applications configured for genetic diagnosis.
[0014] In contrast, the diagnostic workflow system maintains the security and integrity of sequencing data even when implementing external sequencing diagnostic workflows from sources outside the diagnostic workflow system's secure local network. Unlike existing diagnostic systems that may expose the genetic information of a sample, the diagnostic workflow system, for example, utilizes a container orchestration engine that isolates individual workflow containers that define the processes of the external sequencing diagnostic workflow, preventing even undesirable or accidental exposure of sensitive sequencing data or other genetic information. By running the workflow container for the external sequencing diagnostic workflow on its own and configuring targeted security permissions for the external sequencing diagnostic workflow, the diagnostic workflow system encodes the secure execution of the external sequencing diagnostic workflow without accessing internal metrics for sequencing run software or variant analysis models. In fact, unlike existing diagnostic systems, the diagnostic workflow system goes beyond standard encryption by protecting the external sequencing diagnostic workflow in a separate container to prevent the external sequencing diagnostic workflow from corrupting the code or internal metrics of the application for sequencing run or variant analysis, and by granting targeted security permissions (e.g., read-only permissions) to such containers. As a result of its improved security and analysis protocols, the diagnostic workflow system can satisfy various standards for IVD (unlike some conventional systems).
[0015] Furthermore, in some embodiments, the diagnostic workflow system improves computational efficiency compared to existing diagnostic systems by organizing the workflow to run using the limited resources of the FPGA or CPU. For example, instead of flooding available processors and other computing resources with excessive processing demands for the diagnostic workflow, as in some existing systems, the diagnostic workflow system can utilize a container orchestration engine to intelligently allocate computing resources for running individual workflow containers. Specifically, the diagnostic workflow system can utilize a container orchestration engine to specify the iterative or sequential processing (coordinated with or after processing of sequencing data) of each workflow container via a genome sequencing processing device such as an FPGA or CPU. Thus, the diagnostic workflow system can avoid the backlog and / or slowdowns that commonly occur in conventional systems when running a diagnostic workflow on sequencing data.
[0016] As suggested by the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of a diagnostic workflow system. Further details regarding the meanings of these terms used in the present disclosure are provided below. As used in the present disclosure, for example, the term "sample nucleotide sequence" or "sample sequence" refers to the sequence of nucleotides isolated or extracted from a sample organism (or a copy of such an isolated or extracted sequence). In particular, a sample nucleotide sequence is isolated or extracted from a sample organism and includes segments of a nucleic acid polymer composed of nitrogenous heterocyclic bases. For example, a sample nucleotide sequence can include segments of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acids or chimeric or hybrid forms of nucleic acids described below. More specifically, in some cases, a sample nucleotide sequence is that found in a sample prepared or isolated by a kit and received by a sequencing device.
[0017] As further used herein, the term "sequencing data" refers to data or information regarding the nucleotide sequences of one or more genomic samples. For example, sequencing data can include data generated by a sequencing device and / or a variant analysis model. In some cases, sequencing data includes nucleotide reads, nucleotide base calls, and / or sequencing metrics related to a sample nucleotide sequence. In one or more embodiments, sequencing data is generated using the unique methods and processes of a genomic analysis platform that includes a variant analysis model implemented by a genomic sequence processing device and is specific to a particular nucleotide sequence.
[0018] In certain embodiments, sequencing data is stored or located on one or more workflow data sources. The term “workflow data source” refers to a network storage location or repository for storing one or more types of sequencing data. For example, a workflow data source may include locations specific to a particular type of sequencing data, such as an input directory, an output directory, or an application directory. Within an input directory, sequencing data may include data from sample nucleotide sequences written or generated by a variant analysis model (e.g., nucleotide-based calling). Within an output directory, sequencing data may include output information (e.g., diagnostic results) written or generated by a container orchestration engine that performs the workflow. Within an application directory, sequencing data may include definitions of diagnostic applications and / or diagnostic workflows (e.g., external sequencing diagnostic workflows) that include instructions for inputs, outputs, and other execution information for performing the diagnostics.
[0019] In relation to this, the term “nucleotide-based call” (or sometimes simply “call”) refers to the determination or prediction of a specific nucleotide base (or nucleotide base pair) for genomic coordinates or oligonucleotides of a sample genome during a sequencing cycle. In particular, a nucleotide-based call may indicate (i) the determination or prediction of the type of nucleotide base embedded within oligonucleotides on a nucleotide sample slide (e.g., a read-based nucleotide-based call), or (ii) the determination or prediction of the type of nucleotide base present in genomic coordinates or regions within the sample genome, including variant calls or non-variant calls in digital output files. In some cases, for nucleotide reads, a nucleotide-based call includes the determination or prediction of a nucleotide base based on intensity values resulting from fluorescently tagged nucleotides attached to oligonucleotides on a nucleotide sample slide (e.g., in a well of a flow cell). Alternatively, a nucleotide-based call includes the determination or prediction of a nucleotide base to chromatogram peaks or current changes resulting from nucleotides passing through the nanopores of a nucleotide sample slide. In contrast, a nucleotide-based call may also include an initial or final prediction of a nucleotide base in genomic coordinates of the sample genome for a variant call file or other base call output file, based on nucleotide reads corresponding to genomic coordinates. Therefore, nucleotide base calls can include base calls corresponding to genomic coordinates and a reference genome, such as indicators of variants or non-variants at specific locations corresponding to the reference genome. In fact, nucleotide base calls can refer to variant calls that include base calls that are part of single nucleotide polymorphisms (SNPs), insertions or deletions (indels), or structural variants, but are not limited to these. By using nucleotide base calls, sequencing systems determine the sequence of nucleic acid polymers.For example, a single nucleotide base call can include an adenine call, a cytosine call, a guanine call, or a thymine call (abbreviated as A, C, G, T) for DNA, or a uracil call (instead of a thymine call) (abbreviated as U) for RNA.
[0020] As used herein, the term "sequencing metric" refers to a quantitative measurement or score indicating the degree to which an individual nucleotide base call (or sequence of nucleotide base calls) is aligned, compared, or quantified with respect to genomic coordinates or genomic regions of a reference genome, with respect to nucleotide base calls from a nucleotide read, or with respect to an external genomic sequencing or genomic structure. For example, a sequencing metric can include a quantitative measurement or score indicating (i) the degree to which an individual nucleotide base call aligns, maps, or covers genomic coordinates or a reference base of a reference genome, (ii) the degree to which a nucleotide base call compares to a reference or alternative nucleotide read with respect to mapping, mismatches, base call quality, or other raw sequencing metrics, or (iii) the degree to which a genomic coordinate or region corresponding to a nucleotide base call demonstrates mapping potential, repetitive base call content, DNA structure, or other generalized metrics.
[0021] In related terms, as used herein, the term "nucleotide read" (or sometimes simply "read") refers to the inferred sequence of one or more nucleotide bases (or nucleotide base pairs) from all or part of a sample nucleotide sequence. In particular, a nucleotide read includes the sequence of determined or predicted nucleotide base calls for a nucleotide fragment (or group of monoclonal nucleotide fragments) from a sequencing library corresponding to a genomic sample. For example, a diagnostic workflow system determines a nucleotide read by generating nucleotide base calls for nucleotide bases that have passed through nanopores of a nucleotide sample slide, determined via fluorescence tagging, or determined from wells in a flow cell.
[0022] As described herein, in some embodiments, the diagnostic workflow system utilizes a variant analysis model to generate nucleotide base calls for genomic coordinates. As used herein, the term “variant analysis model” refers to a model comprising an algorithm or set of algorithms for analyzing data (e.g., base call data) about a sample nucleotide sequence. In some cases, the variant analysis model is a probabilistic model that generates sequencing data from nucleotide reads of a sample nucleotide sequence, including nucleotide base calls (e.g., variant calls) and associated metrics (e.g., base call quality metrics). For example, in some cases, the variant analysis model refers to a Bayesian probabilistic model that generates variant calls based on nucleotide reads of a sample nucleotide sequence. Such a model may include a model for secondary analysis performed by a server running variant call software to align the nucleotide reads of the sample with respect to a reference genome, determine the genetic variant of the sample based on the nucleotide reads aligned with respect to the reference genome, and determine one or more of the quality metrics, allele frequency metrics, or other sequencing metrics. Variant analysis models may also include multiple components, including, but are not limited to, different software applications or components for mapping and alignment, sorting, duplicate marking, read pile-up depth calculation, and variant calling. In some cases, the variant analysis model refers to the ILLUMINA DRAGEN model for the variant calling function and the mapping and alignment function.
[0023] As described, in some embodiments, a diagnostic workflow system utilizes a container orchestration engine to organize the execution of an external sequencing diagnostic workflow. As used herein, the term “container orchestration engine” refers to an application for automating the deployment, scaling, and management of containerized software services and applications. For example, a container orchestration engine may include a software application having a microservices architecture that runs individual workflow containers as part of a diagnostic workflow (e.g., an external sequencing diagnostic workflow). The container orchestration engine can treat each container separately to perform distinct functions (e.g., containerized tasks) that are partitioned and can be added to or removed from the workflow.
[0024] Relatedly, the term “workflow container” (or sometimes simply “container”) refers to a unit of software that packages code (and all its dependencies) for portable deployment. For example, a workflow container contains a partitioned or containerized, operable, portable, executable code body that performs a specific function or task. In some cases, a workflow container is executable to perform a function or task (e.g., a process or thread) to produce a specific output (e.g., a final output or intermediate output to feed into another container) from a fragment of sequencing data. Diagnostic workflow systems can treat workflow containers separately and isolate some containers from others to allow and / or prevent access to sequencing data (e.g., within one or more workflow data sources) in a specific coordinated manner. In some cases, a container refers to a NEXTFLOW container and / or a KUBERNETES container.
[0025] Additionally, as used herein, the term “sequencing diagnostic workflow” refers to a set of tasks or functions organized and coordinated together to generate a set of diagnostic outputs from sequencing data for a sample nucleotide sequence. For example, a sequencing diagnostic workflow may include any number of tasks providing any number of functions, which may include secondary analysis, tertiary analysis, custom QC logic, reporting, or other desired functionality. In some cases, multiple entities may develop a single workflow that can be deployed locally (e.g., on an edge server) or in the cloud.
[0026] An “external sequencing diagnostic workflow” refers to a sequencing diagnostic workflow that is outside of a genome analysis software platform or outside of one or more applications from a larger genome analysis software platform. For example, an external sequencing diagnostic workflow includes sequencing diagnostic workflows that are not native to (or not integrated as part of) the software (or another software application or set of software applications) for variant analysis models used for diagnosis and that follow standardized genetic diagnostic protocols. In some cases, a diagnostic workflow system receives or identifies external sequencing diagnostic workflows from third-party systems that do not share a local network with the server device hosting the diagnostic workflow system (e.g., with the container orchestration engine and / or variant analysis models). An external sequencing diagnostic workflow can refer to a newly generated workflow or a modified version of an existing workflow defined by an external device operated by an external entity (e.g., an entity not part of the diagnostic workflow system). In some cases, an external sequencing diagnostic workflow can refer to add-on software for a diagnostic workflow system that is adaptable to plug into a variant analysis model for performing secondary and / or tertiary analysis of nucleotide base calls.
[0027] An external sequencing diagnostic workflow can be part of or defined by a diagnostic application (e.g., an external diagnostic application). A “diagnostic application” can refer to a package of custom workflow content, deployable by a diagnostic workflow system, including workflow definitions, containerized tasks, custom user interfaces, custom microservices, and / or reference data. A diagnostic application may have a specific structure or file type (e.g., a tape archive or “TAR” file) that defines the workflow and other application data. In some embodiments, the diagnostic application is a single deployable unit that can be installed on a system completely disconnected from the internet. The application package may be signed to ensure authenticity and validity. The package may be uploaded to a server via a UI portal (e.g., using a browser). In some cases, the external sequencing diagnostic workflow is part of an external diagnostic application generated by (and received from) an external (e.g., third-party) system.
[0028] The following paragraphs describe the diagnostic workflow system with respect to exemplary embodiments and exemplary diagrams illustrating implementation configurations. For example, Figure 1 illustrates a schematic diagram of a system environment (or "environment") 100 in which the diagnostic workflow system 106 operates according to one or more implementation configurations. As illustrated, the environment 100 includes one or more server devices 102 connected via a network 112 to a client device 108 and an array determination device 114. While Figure 1 shows one embodiment of the diagnostic workflow system 106, alternative embodiments and configurations are described below.
[0029] As shown in Figure 1, the server device 102, the client device 108, and the sequence determination device 114 can communicate with each other via the network 112. The network 112 includes any suitable network on which the computing devices can communicate. An exemplary network is discussed in additional detail below with respect to Figure 9.
[0030] As shown in Figure 1, the sequencing apparatus 114 includes an apparatus for sequencing nucleic acid polymers. In some embodiments, the sequencing apparatus 114 analyzes nucleic acid segments or oligonucleotides extracted from a sample to generate nucleotide reads or other data, either directly or indirectly on the sequencing apparatus 114 using computer implementation methods and systems (as described herein). More specifically, the sequencing apparatus 114 receives and analyzes nucleic acid sequences extracted from a sample in a nucleotide sample slide (e.g., a flow cell). In one or more embodiments, the sequencing apparatus 114 sequences nucleic acid polymers into nucleotide reads using SBS. In addition to, or as an alternative to, communicating via the network 112, in some embodiments, the sequencing apparatus 114 bypasses the network 112 and communicates directly with server equipment 102 and / or client equipment 108. In fact, in some embodiments, the sequencing device 114 and the server device 102 share a local network (e.g., hosted on the same or different servers), as indicated by the dashed boxes, while the client device 108 does not share the local network and instead communicates via network 112.
[0031] As further shown in Figure 1, the server device 102 can generate, receive, analyze, store, and transmit digital data such as data for determining nucleotide base calls, sequencing nucleic acid polymers, and / or performing diagnostics on nucleotide sequences. As shown in Figure 1, the sequencing device 114 can transmit call data from the sequencing device 114 (and the server device 102 can receive call data). The server device 102 can also communicate with the client device 108. In particular, the server device 102 can transmit data to the client device 108 that includes variant call files, or other information indicating nucleotide base calls, sequencing metrics, error data, diagnostic information, or other metrics associated with nucleotide base calls.
[0032] In some embodiments, the service device 102 includes a local server device located at or near the same physical location as the sequencing device 114. In fact, in some embodiments, the server device 102 and the sequencing device 114 are integrated into the same computing device, as indicated by the dotted lines around the server device 102 and the sequencing device 114.
[0033] Rather than being located locally with the sequencing device 114, in some embodiments, the server device 102 comprises a collection of distributed servers, and the server device 102 comprises several server devices distributed across the network 112 and located in the same or different physical locations. As suggested, in some cases, the server device 102 houses the genome analysis platform 104. The server device 102 may also include content servers, application servers, communication servers, web hosting servers, or other types of servers.
[0034] As further shown in Figure 1, the server device 102 may include a genome analysis platform 104 for generating and analyzing sequencing data. Generally, the genome analysis platform 104 includes a variant analysis model 107 that generates and / or analyzes call data, such as sequencing metrics received from the sequencing device 114, to determine the nucleotide-based sequences of nucleic acid polymers. For example, the variant analysis model 107 may receive raw data from the sequencing device 114 and determine the nucleotide-based sequences for nucleic acid segments. In some embodiments, the variant analysis model 107 determines the nucleotide-based sequences in DNA and / or RNA segments or oligonucleotides. In addition to processing and determining sequences for nucleic acid polymers, the variant analysis model 107 also generates a variant call file showing one or more nucleotide-based calls and / or variant calls for one or more genomic coordinates. Furthermore, in some embodiments, the genome analysis platform 104 includes a diagnostic workflow system 106.
[0035] As described above and as illustrated in Figure 1, the diagnostic workflow system 106 analyzes sequencing data, such as call data and / or sequencing metrics (e.g., from the sequencing instrument 114 or variant analysis model 107), to generate or determine various diagnoses. For example, the diagnostic workflow system 106 performs a diagnostic analysis of sequencing data for a sample nucleotide sequence. In some cases, the diagnostic workflow system 106 performs a diagnostic analysis to diagnose or determine a tendency toward one or more diseases or genetic conditions. In fact, the diagnostic workflow system 106 performs a diagnosis on the sequencing data and runs an external sequencing diagnostic workflow (such as one received from the client instrument 108) to determine the likelihood of revealing a genetic condition or trait.
[0036] As further illustrated in Figure 1, the client device 108 can generate, store, receive, and transmit digital data. In particular, the client device 108 can receive sequencing metrics from the sequencing device 114. Furthermore, the client device 108 can communicate with the server device 102 to receive variant call files containing nucleotide-based calls and / or other metrics such as call quality, genotype index, and genotype quality. Thus, the client device 108 can present or display information about nucleotide-based calls within a graphical user interface to the user associated with the client device 108. In addition, the client device 108 can generate and provide (e.g., upload via network 112) an external sequencing diagnostic workflow containing one or more workflow containers for performing diagnostic analysis of sequencing data. In fact, the client device 108 can receive user interactions via the graphical user interface to select and sequence or organize workflow containers for the external sequencing diagnostic workflow. The client device 108 can also receive diagnostic results from the server device 102 and display the diagnostic results within the graphical user interface.
[0037] The client device 108 illustrated in Figure 1 can include various types of client devices. For example, in some embodiments, the client device 108 includes non-mobile devices such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 108 includes mobile devices such as laptops, tablets, mobile phones, or smartphones. Further details regarding the client device 108 are discussed below with reference to Figure 9.
[0038] As further illustrated in Figure 1, the client device 108 includes a sequencing application 110. The sequencing application 110 may be a web application or a native application (e.g., a mobile application, a desktop application) stored and executed on the client device 108. The sequencing application 110 may include commands to cause the client device 108 to receive data from the diagnostic workflow system 106 and to present data from a variant call file for display on the client device 108. Furthermore, the sequencing application 110 can instruct the client device 108 to visualize the external sequencing diagnostic workflow and / or the workflow containers arranged within the external sequencing diagnostic workflow, as well as to display the diagnostic results received from the server device when the external sequencing diagnostic workflow is executed.
[0039] As further illustrated in Figure 1, the diagnostic workflow system 106 may reside on the client device 108 or on the sequencing device 114 as part of the sequencing application 110. Therefore, in some embodiments, the diagnostic workflow system 106 is implemented on the client device 108 (e.g., entirely or partially). In yet other embodiments, the diagnostic workflow system 106 is implemented by one or more other components of the environment 100, such as the sequencing device 114. In particular, the diagnostic workflow system 106 can be implemented in various different ways across the server device 102, the network 112, the client device 108, and the sequencing device 114. For example, the diagnostic workflow system 106 can be downloaded from the server device 102 to the client device 108 and / or the sequencing device 114, and all or part of the functionality of the diagnostic workflow system 106 is performed on each device within the environment 100.
[0040] As further illustrated in Figure 1, the environment 100 includes a database 116. The database 116 can store information such as external sequencing diagnostic workflows, diagnostic results, variant call files, sample nucleotide sequences, and sequencing data such as nucleotide reads, nucleotide base calls, variant calls, and sequencing metrics. In some embodiments, a server device 102, a client device 108, and / or a sequencing device 114 communicate with the database 116 (e.g., via a network 112) to store and / or access information such as external sequencing diagnostic workflows, diagnostic results, variant call files, sample nucleotide sequences, and sequencing data such as nucleotide reads, nucleotide base calls, variant calls, and sequencing metrics. In some cases, the database 116 also stores one or more models, such as variant analysis models 107.
[0041] Figure 1 illustrates the components of environment 100 communicating via network 112, but in certain implementations, the components of environment 100 can also communicate directly with each other, bypassing network 112. For example, as mentioned above, in some implementations, the client device 108 can communicate directly with the sequence determination device 114. Additionally, in some embodiments, the client device 108 communicates directly with the diagnostic workflow system 106. Furthermore, the diagnostic workflow system 106 can access one or more databases housed in or accessed by a server device 102 or other location within environment 100.
[0042] As described, in certain described embodiments, the diagnostic workflow system 106 performs an external sequencing diagnostic workflow. In particular, the diagnostic workflow system 106 performs an external sequencing diagnostic workflow generated by an external system or an external device operated by an external system, such as a third-party system or third-party device (e.g., client device 108), in order to perform a diagnostic analysis of sequencing data. Figure 2 illustrates an exemplary overview of performing an external sequencing diagnostic workflow according to one or more embodiments. Figure 2 provides a general overview of performing an external sequencing diagnostic workflow, and additional details regarding various operations are provided thereafter by reference to subsequent figures.
[0043] As illustrated in Figure 2, the diagnostic workflow system 106 performs operation 202 to receive an external sequencing diagnostic workflow. More specifically, the diagnostic workflow system 106 receives an external sequencing diagnostic workflow generated by an external system. In some embodiments, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow directly from an external device, such as a client device 108. In other embodiments, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow from a different server (e.g., a web hosting server) that communicates with the client device 108 by hosting a website on which the client device 108 uploads the external sequencing diagnostic workflow.
[0044] In fact, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow from a web hosting server that receives the external sequencing diagnostic workflow as an upload (as part of the application). In some embodiments, the web hosting server is part of the diagnostic workflow system 106, while in other embodiments, the web hosting server is external to the diagnostic workflow system 106. Thus, in some cases, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow, while in other cases, the external sequencing diagnostic workflow is already downloaded or otherwise included as part of the diagnostic workflow system 106.
[0045] As further illustrated in Figure 2, the diagnostic workflow system 106 performs operation 204 to identify sequencing metrics. More specifically, the diagnostic workflow system 106 identifies sequencing data, such as nucleotide base calls and / or sequencing metrics, associated with one or more sample nucleotide sequences. In some cases, the diagnostic workflow system 106 generates sequencing data using a variant analysis model (e.g., variant analysis model 107), while in other cases, the diagnostic workflow system 106 receives sequencing data from another server hosting the variant analysis model. In certain embodiments, the diagnostic workflow system 106 identifies sequencing data specific to a particular sample.
[0046] Additionally, the diagnostic workflow system 106 performs operation 206 to execute the external sequencing diagnostic workflow. Specifically, the diagnostic workflow system 106 executes the external sequencing diagnostic workflow by executing a set of workflow containers that constitute the external sequencing diagnostic workflow. For example, the diagnostic workflow system 106 executes an external sequencing diagnostic workflow generated by an external system and / or uploaded by an external device (e.g., client device 108). In some cases, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow from a server that directly receives the external sequencing diagnostic workflow via upload and relays the external sequencing diagnostic workflow to a server of the diagnostic workflow system 106 (e.g., server device 102).
[0047] In fact, the diagnostic workflow system 106 executes an external sequencing diagnostic workflow generated by an external system operated by an external entity. More specifically, the diagnostic workflow system 106 facilitates or enables the generation and / or sequencing of an external sequencing diagnostic workflow via an external device (e.g., a client device 108). For example, the external device can sequence or organize individual workflow containers to perform discretized functions or tasks as part of performing a specific diagnostic test on a sample nucleotide sequence. In some embodiments, the workflow containers are predefined or pre-generated by the diagnostic workflow system 106 and stored for use (by the external system) in sequencing the external sequencing diagnostic workflow. In other embodiments, the diagnostic workflow system 106 facilitates the generation of entirely new workflow containers by the external system.
[0048] In some cases, the external sequencing diagnostic workflow is part of an external application. For example, the diagnostic workflow system 106 can execute an external sequencing diagnostic workflow defined by an application. The application may include a workflow that includes a custom user interface, reference data, and various composition containers, as defined by the external system. The diagnostic workflow system 106 can implement an external application to execute the external sequencing diagnostic workflow and provide diagnostic results via a web user interface (e.g., a custom interface defined by the application).
[0049] In certain embodiments, the external application is a modified version of an internal application that has been pre-generated and stored within the diagnostic workflow system 106 (e.g., pre-authorized to conform to one or more standardized genetic diagnostic protocols such as IVD, or to conform to other analytical protocols such as IUO or RUO). The diagnostic workflow system 106 can receive or identify modified versions of internal applications that have been modified or generated by the external system. In fact, the external system may modify the internal application to change, add, and / or remove one or more workflow containers within the workflow. For example, if a stored internal application includes a workflow that performs a diagnosis against a specific nucleotide base call to identify a particular genetic condition (e.g., a TSO500 screening assay designed to detect a specific cancer marker), the external system may modify the application (thus generating the external application) to analyze a different nucleotide base call and / or process other sequencing data in a slightly different way.
[0050] To execute an external sequencing diagnostic workflow, the diagnostic workflow system 106 requires the container orchestration engine to implement an external sequencing diagnostic workflow for performing diagnostic analysis on nucleotide base calls for the sample nucleotide sequence. More specifically, the diagnostic workflow system 106 requires the container orchestration engine to implement an external sequencing diagnostic workflow according to its sequence as defined in the external diagnostic application. For example, the diagnostic workflow system 106 utilizes the container orchestration engine to perform diagnostic analysis of the sample nucleotide sequence by executing individual workflow containers sequenced within the external sequencing diagnostic workflow.
[0051] As shown, in some embodiments, the diagnostic workflow system 106 performs operation 206 to execute an external sequencing diagnostic workflow after identifying sequencing data (as performed by the variant analysis model 107) or after sequencing is complete. For example, the diagnostic workflow system 106 executes the external sequencing diagnostic workflow after the variant analysis performed by the variant analysis model 107 has identified or generated nucleotide base calls for the sample nucleotide sequence. For example, in some embodiments, the diagnostic workflow system 106 detects the completion of the sequencing analysis by the variant analysis model 107, which automatically triggers the execution of the external sequencing diagnostic workflow.
[0052] In other embodiments, the diagnostic workflow system 106 performs operation 206 to execute the external sequencing diagnostic workflow during variant analysis or while the variant analysis model 107 is generating nucleotide base calls. More specifically, the diagnostic workflow system 106 also executes the external sequencing diagnostic workflow via the container orchestration engine (e.g., simultaneously or in parallel) while communicating with or utilizing the variant analysis model 107. For example, the diagnostic workflow system 106 generates specific nucleotide base calls required by a particular workflow container of the external sequencing diagnostic workflow, and as the variant analysis model 107 is generating other nucleotide base calls, the diagnostic workflow system 106 executes those executable parts of the external sequencing diagnostic workflow based on the nucleotide base calls available at that time. As sequencing progresses, the diagnostic workflow system 106 continues to execute the external sequencing diagnostic workflow until a complete diagnostic analysis is completed based on the complete set of nucleotide base calls from the variant analysis model 107.
[0053] As further illustrated in Figure 2, the diagnostic workflow system 106 performs operation 208 to provide diagnostic workflow results. In particular, the diagnostic workflow system 106 provides diagnostic workflow results generated as a result or product of an external sequencing diagnostic workflow. In some cases, the diagnostic workflow system 106 provides the diagnostic workflow results to a web hosting server for relaying to an external system operated by an external entity. In other cases, the diagnostic workflow system 106 provides the diagnostic workflow results directly to an external system (e.g., via a client device 108). For example, the diagnostic workflow system 106 provides the diagnostic workflow results for display via the client device 108 within a graphical user interface.
[0054] As described above, in certain described embodiments, the diagnostic workflow system 106 executes an external sequencing diagnostic workflow generated by an external system (one or more server hosts) of the diagnostic workflow system 106. In particular, the diagnostic workflow system 106 facilitates the generation and execution of custom sequencing diagnostic workflows for external systems to leverage their own sequencing data, such as nucleotide-based calls, generated by an internal variant analysis model (e.g., variant analysis model 107) for their own diagnostic purposes, without risking exposure to or corruption of sequencing data. Figure 3 illustrates an exemplary flow for executing an external sequencing diagnostic workflow on internal sequencing data in a secure environment, according to one or more embodiments.
[0055] As illustrated in Figure 3, the diagnostic workflow system 106 utilizes the container orchestration engine 318 to execute the external sequencing diagnostic workflow 310. Specifically, the diagnostic workflow system 106 executes the external sequencing diagnostic workflow 310 to perform a diagnostic analysis on the nucleotide base call 306 generated by the sequencing instrument 302. In fact, as shown, the sequencing instrument 302 analyzes the sample nucleotide sequence to generate a nucleotide base call 306 for the sample sequence (e.g., via synthetic sequencing, i.e., "SBS"). The sequencing instrument 302 then provides the nucleotide base call 306 to the variant analysis model 314 (the host instrument), which then generates a variant call 316 from the nucleotide base call 306.
[0056] In fact, the variant analysis model 314 generates the variant call 316 from the nucleotide base call 306 and / or other sequencing data such as sequencing metrics. For example, in some cases, the sequencing device 302 generates the nucleotide base call 306 and sequencing data such as sequencing metrics from the sample nucleotide sequence. The diagnostic workflow system 106 can access or otherwise utilize the sequencing data (including the nucleotide base call 306) as a basis for performing tertiary analysis or additional diagnostics (e.g., via the container orchestration engine 318).
[0057] As shown, the diagnostic workflow system 106 utilizes the container orchestration engine 318 to perform additional diagnostics or tertiary analyses on the variant call 316 (based on the nucleotide-based call 306 and other sequencing data in the corresponding genomic coordinates). For example, the diagnostic workflow system 106 utilizes the container orchestration engine 318 to execute sequencing diagnostic workflows, such as the external sequencing diagnostic workflow 310. In one or more embodiments, the container orchestration engine 318 and the variant analysis model 314 reside on a shared network 312, such as a local area network, to keep the variant call 316 and other sequencing data within a closed environment (for example, for additional data security). In fact, the diagnostic workflow system 106 communicates with both the variant analysis model 314 and the container orchestration engine 318 via the shared network 312, with the variant analysis model 314 residing on one server device and the container orchestration engine 318 residing on different server devices within the shared network 312.
[0058] In certain implementations, the diagnostic workflow system 106 implements an external sequencing diagnostic workflow 310 that is generated or modified by an external system (e.g., a third-party system). In some cases, the external system refers to a system located outside the shared network 312, or a system operated separately by an external (e.g., third-party) entity away from the genome analysis platform 104. In these, or other, the client device 304 (e.g., client device 108) is associated with the external system (e.g., a part thereof). In some embodiments, the client device 304 generates or sequences the external sequencing diagnostic workflow 310 as part of an external application 308.
[0059] Along these lines, the external application 308 includes a definition of the external sequencing diagnostic workflow 310, along with metadata such as workflow resource data (e.g., genome), Docker images containing software dependencies (e.g., variant analysis model 314), the name and version of the external application 308, how the external application 308 interfaces with the sequencing instrument 302 (e.g., compatible index adapter kit and library preparation kit), and how the diagnostic analysis can be configured and what parameters can be specified (e.g., by the client instrument 304). The client instrument 304 either sequences a workflow container for the external sequencing diagnostic workflow 310 from scratch to generate the external sequencing diagnostic workflow 310 for performing a specific diagnosis, or modifies an existing sequence of a workflow container (e.g., as part of a workflow or application already installed on the genome analysis platform 104).
[0060] For example, the diagnostic workflow system 106 provides access via a web interface to one or more internal sequencing diagnostic workflows (or applications) that conform to one or more standardized genetic diagnostic protocols or other analysis protocols. The diagnostic workflow system 106 further facilitates changes or modifications to sequencing diagnostic workflows by adding, deleting, or modifying the definitions of one or more workflow containers within the workflow (thus generating external sequencing diagnostic workflows). In addition, the diagnostic workflow system 106 facilitates the generation of entirely new external sequencing diagnostic workflows by sequencing available workflow containers and / or by generating or defining new workflow containers.
[0061] As shown, in one or more embodiments, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow 310 from a client device 304 (for example, as part of an external application 308). In other embodiments, the diagnostic workflow system 106 receives the external sequencing diagnostic workflow 310 from another server communicating with the client device 304. The diagnostic workflow system 106 further imports the external sequencing diagnostic workflow 310 into the container orchestration engine 318 to execute the external sequencing diagnostic workflow 310 and generate the diagnostic workflow result 326. In fact, the diagnostic workflow system 106 processes the variant call 316 (or nucleotide-based call 306 and / or other sequencing data) according to the workflow container defined in the external sequencing diagnostic workflow 310 to execute the external sequencing diagnostic workflow 310 internally without further communication to or from an external system.
[0062] In some embodiments, the diagnostic workflow system 106 executes an external sequencing diagnostic workflow 310 to perform a diagnostic analysis on the nucleotide base call 306 (and other sequencing data) and / or the variant call 316. In some cases, the diagnostic workflow system 106 receives instructions to perform a sequencing run (e.g., via a user interface presented on the client device 304). In response to these instructions, the diagnostic workflow system 106 utilizes the sequencing device 302 to generate the nucleotide base call 306 from the sample nucleotide sequence and further utilizes the variant analysis model 314 to generate the variant call 316 from the nucleotide base call 306. In some embodiments, the diagnostic workflow system 106 receives the nucleotide base call 306 (and other sequencing data) from the sequencing device 302 operating on a separate server. In these, or other embodiments, the diagnostic workflow system 106 receives the variant call 316 from the variant analysis model 314 operating on a separate server. The diagnostic workflow system 106 further performs a diagnostic analysis of the external sequencing diagnostic workflow 310 either during or after the sequencing run.
[0063] As shown, the diagnostic workflow system 106 executes the external sequencing diagnostic workflow 310 via the workflow execution service 320 of the container orchestration engine 318. In particular, the diagnostic workflow system 106 utilizes a specific workflow container of the container orchestration engine 318, referred to as the workflow execution service 320, which triggers the execution of the external sequencing diagnostic workflow 310. For example, the workflow execution service 320 triggers execution by detecting that sequencing is complete (e.g., the sequencing instrument 302 has completed generating the nucleotide base call 306, and / or the variant analysis model 314 has completed generating the variant call 316 from the nucleotide base call 306), or by receiving post-defined parameters for the execution of the external sequencing diagnostic workflow 310 (e.g., via a user interface presented on the client instrument 304).
[0064] In some cases, the post defines parameters for implementing the external sequencing diagnostic workflow 310, such as inputs, outputs, memory allocations, and / or required versions of a variant analysis model 314 compatible with the external sequencing diagnostic workflow 310. In some cases, the post includes the location of the workflow definition file (e.g., within the application directory), the location of the workflow resource file (e.g., within the application directory), the location of the sequencing input directory (e.g., the execution folder), the location of the output directory, and any user configuration parameters for the workflow. In fact, based on the post, the diagnostic workflow system 106 can provide a number of functions for running the external sequencing diagnostic workflow 310, including (i) containerized task execution, (ii) task orchestration, (iii) management of inputs and outputs for individual workflow containers of tasks and the entire workflow, (iv) variant analysis model (DRAGEN) acceleration, (v) automated workflow execution (e.g., sequencing to diagnostics without user interaction), (vi) system resource management (e.g., scheduling tasks for workflow containers based on CPU, RAM, storage, and FPGA resources), and (vii) workflow status.
[0065] The diagnostic workflow system 106 provides additional security in preventing exposure or corruption of sequencing data, such as nucleotide-based calls 306 and / or variant calls 316. In fact, the diagnostic workflow system 106 isolates or silos the individual workflow containers of the external sequencing diagnostic workflow 310 to limit the workflow data sources they can access. As shown, the diagnostic workflow system 106 grants only limited access to specific workflow containers of sequencing data and does not grant access to other workflow containers. For example, the container orchestration engine 318 utilizes the workflow main container 322 as an overall process for organizing the execution of additional workflow containers, such as workflow container 324, which constitutes the majority of the functionality for the external sequencing diagnostic workflow 310. As depicted, the diagnostic workflow system 106 grants limited access to the workflow main container 322 and does not grant access to sequencing data to workflow container 324 (and other task-specific workflow containers). By partitioning the workflow container in this way and controlling data access, while still performing detailed diagnostics on sequencing data, the diagnostic workflow system 106 provides data security and protocol compliance for various standardized genetic diagnostic protocols such as IVD, or other standardized analytical (e.g., non-diagnostic) protocols such as IUO and RUO. Further details regarding data-specific isolation of the workflow container are provided below with reference to Figure 4.
[0066] In another embodiment, the diagnostic workflow system 106 facilitates an integrated runtime between the cloud and the edge. More specifically, the diagnostic workflow system 106 can enable an external system to initiate execution via the variant analysis model 314 and provide an external sequencing diagnostic workflow for implementation by the container orchestration engine 318. In some cases, the diagnostic workflow system 106 can execute the external sequencing diagnostic workflow (e.g., external sequencing diagnostic workflow 310) using edge server devices or distributed cloud servers at nearly the same runtime.
[0067] In addition, the diagnostic workflow system 106 provides portability, for example, the ability to deploy components of an external sequencing diagnostic workflow or the entire solution on different servers, as well as sequencing devices (e.g., sequencing device 302). In fact, in some cases, the container orchestration engine 318 orchestrates the execution of different workflow containers across different components such as sequencing device 302, variant analysis model 314, and / or various server devices. For example, some workflow containers of the external sequencing diagnostic workflow 310 are executed by a first server device, while other workflow containers are executed by another server device or sequencing device 302. Thus, the diagnostic workflow system 106 facilitates the unified activation of workflows between cloud servers and edge servers. Therefore, once a workflow container performs a task and writes output, the diagnostic workflow system 106 can leverage the container orchestration engine 318 to access the output to run another workflow container elsewhere in the environment (e.g., write once and run everywhere).
[0068] As described above, in certain embodiments, the diagnostic workflow system 106 isolates individual workflow containers of the external sequencing diagnostic workflow to protect sequencing data. Specifically, the diagnostic workflow system 106 determines which workflow containers of the external sequencing diagnostic workflow are permitted to access which workflow data sources for read and / or write permissions. Figure 4 illustrates an exemplary depiction of permissions that the diagnostic workflow system 106 assigns to a particular workflow container in one or more embodiments.
[0069] As illustrated in Figure 4, the diagnostic workflow system 106 isolates individual workflow containers within a set of workflow containers 402 for an external sequencing diagnostic workflow (e.g., external sequencing diagnostic workflow 310). For example, the diagnostic workflow system 106 assigns specific permissions to each of the workflow containers 402 to read from and / or write to a particular workflow data source 414. Thus, each granular task performed as part of the external sequencing diagnostic workflow can only access data in the specified or permitted workflow data sources.
[0070] In fact, the diagnostic workflow system 106 can specify workflow data sources 414 that store different types of workflow data, and can activate access to one source while preventing access to another source for a given workflow container. In some cases, the diagnostic workflow system 106 mounts the workflow data source 414 as read-only for one or more workflow containers, while mounting it as read-write for other workflow containers. By selectively specifying data access for each of the workflow containers 402, the diagnostic workflow system 106 reduces exposure to sequencing data. As a result, the diagnostic workflow system 106 can prevent data corruption or other harmful effects that may result from otherwise exposing sequencing data to potentially harmful sequencing diagnostic workflows.
[0071] As shown, the diagnostic workflow system 106 can implement the variant analysis model 404 (e.g., variant analysis model 314) as a workflow container (or group of workflow containers). In fact, the variant analysis model 404 does not necessarily have to be part of a container orchestration engine (e.g., container orchestration engine 318), but the genome analysis platform 104 housing the container orchestration engine 318 and the variant analysis model 404 can nevertheless utilize workflow containers for individual tasks. As shown in Figure 4, the diagnostic workflow system 106 allows the variant analysis model 404 to both read and write to each of the workflow data sources 414, namely the input directory 416, the output directory 418, and the application directory 420.
[0072] Generally, the input directory 416 stores run folders, sample sheets, sample mappings, and other sequencing-related files, including sequencing data such as nucleotide base calls 306, variant calls 316 (e.g., variant call files or fields), and sequencing metrics generated by the sequencing instrument 302 and / or variant analysis model 404. The output directory 418 stores information such as diagnostic workflow results and intermediate outputs generated by the workflow container and input into other workflow containers. In addition, the application directory 420 stores application-specific files such as workflow definitions, including input specifications, output specifications, memory allocations, compatibility metrics (e.g., indicating compatible versions of variant analysis model 404), and definitions of various workflow containers included as part of an external sequencing diagnostic workflow.
[0073] As illustrated in Figure 4, the diagnostic workflow system 106 isolates the application programming interface (API) 406. Specifically, the diagnostic workflow system 106 grants read-only permissions to the workflow execution service API 406 for each of the workflow data sources 414, including the input directory 416, the output directory 418, and the application directory 420. In some cases, the workflow execution service API 406 provides access to and can invoke various workflow containers. For example, the workflow execution service API 406 can initiate the execution of an external sequencing diagnostic workflow based on a post provided by a client device (e.g., client device 304).
[0074] In addition, the diagnostic workflow system 106 isolates the workflow execution service worker 408 (e.g., workflow execution service 320). Specifically, the diagnostic workflow system 106 allows the workflow execution service worker 408 to read only from both the input directory 416 and the application directory 420. The diagnostic workflow system 106 further allows the workflow execution service worker 408 to read from and write to the output directory 418. In fact, the workflow execution service worker 408 organizes an implementation form of an additional workflow container for the external sequencing diagnostic workflow and further writes the diagnostic workflow results to the output directory 418.
[0075] Furthermore, the diagnostic workflow system 106 isolates the workflow main container 410. Specifically, similar to the workflow execution service worker 408, the diagnostic workflow system 106 allows the workflow main container 410 to read only from the input directory 416 and the application directory 420. The diagnostic workflow system 106 also allows the workflow main container 410 to read from and write to the output directory 418.
[0076] As further illustrated in Figure 4, the diagnostic workflow system 106 isolates the workflow worker container 412 (e.g., workflow container 324). Specifically, similar to the workflow execution service worker 408 and the workflow main container 410, the diagnostic workflow system 106 allows the workflow worker container 412 to read only from the input directory 416 and the application directory 420. The diagnostic workflow system 106 further allows the workflow worker container 412 to read from and write to the output directory 418.
[0077] As described above, in certain embodiments, the diagnostic workflow system 106 utilizes a container orchestration engine to organize the execution of workflow containers (for example, for individual tasks). In particular, the diagnostic workflow system 106 utilizes a container orchestration engine to schedule the computing resources available for performing tasks in the workflow containers. Figure 5 illustrates the use of a container orchestration engine to schedule the execution of workflow containers in one or more embodiments.
[0078] As illustrated in Figure 5, the diagnostic workflow system 106 utilizes a container orchestration engine 502 (e.g., container orchestration engine 318) to execute workflow containers for sequencing diagnostic workflows (e.g., external sequencing diagnostic workflow 310). In particular, the container orchestration engine 502 orchestrates the performance of workflow containers A 504, B 506, and C 508 using genome sequencing processing devices (e.g., FPGAs or CPUs). The depiction in Figure 5 is an example, and in some cases, the container orchestration engine 502 orchestrates the performance of different workflow containers across different servers or devices with different available resources, some of which may utilize CPUs and others which may utilize FPGAs (or other computing processors).
[0079] In scheduling the performance of the workflow container, the container orchestration engine 502 determines the available computing resources (e.g., FPGA such as DRAGEN FPGA, or CPU) of the genome sequencing processor 510. The available resources can fluctuate over time as the genome sequencing processor 510 performs different tasks from either the container orchestration engine 502 or another source. Therefore, the diagnostic workflow system 106 utilizes the container orchestration engine 502 to implement adaptive and elastic resource allocation and scale with available resources such as CPU, RAM, FPGA resources, variant analysis model (DRAGEN) accelerator, and several available worker nodes.
[0080] As shown, the container orchestration engine 502 communicates with the genome sequencing processor 510 to execute a container A execution 512 that runs the workflow container A 504. In fact, based on the determination that the genome sequencing processor 510 has available computing resources such as available worker nodes, available RAM, available processing power, and / or available FPGA processing power, the container orchestration engine 502 provides the workflow container A 504. The genome sequencing processor 510 then executes the container A execution 512. In some embodiments, the container A execution 512 generates an intermediate container output that the diagnostic workflow system 106 stores in a network location (e.g., input directory 416, output directory 418, or application directory 420) for use as input for subsequent workflow containers.
[0081] As further illustrated in Figure 5, the container orchestration engine 502 provides workflow container B 506 and workflow container C 508 for execution by the genome sequencing processor 510. In fact, the container orchestration engine 502 determines or detects when execution of container A 512 is complete and then provides workflow container B 506 when the genome sequencing processor 510 is available. The genome sequencing processor 510 then executes execution of container B 514.
[0082] In some cases, the genome sequencing processor 510 is an FPGA that executes processes sequentially (for example, it cannot run multiple workflow containers simultaneously). For example, the genome sequencing processor 510 implements a variant analysis model (e.g., variant analysis model 404) to generate nucleotide-based calls and other sequencing data by performing operations sequentially. Therefore, in some cases, the genome sequencing processor 510 is busy performing operations and cannot perform another operation until the previous operation is complete. Accordingly, the container orchestration engine 502 schedules the resources of the genome sequencing processor 510 to perform tasks of workflow containers based on availability.
[0083] In some embodiments, the genome sequencing processor 510 performs other tasks unrelated to the sequencing diagnostic workflow (or for a different sequencing diagnostic workflow). Therefore, the container orchestration engine 502 organizes the execution of workflow containers 504-508 based on the availability of the genome sequencing processor 510. As shown, the genome sequencing processor 510 performs other process executions 516 (unrelated to the sequencing diagnostic workflow) after container B execution 514 and before container C execution. In fact, once the other process executions 516 are complete, the container orchestration engine 502 determines that the genome sequencing processor 510 is available and provides workflow container C 508 to the genome sequencing processor 510 for execution. The genome sequencing processor 510 then performs container C execution 518.
[0084] The container orchestration engine 502 continuously orchestrates the performance of workflow containers for a sequencing diagnostic workflow (e.g., an external sequencing diagnostic workflow) until it is complete. For example, the container orchestration engine 502 utilizes the genome sequencing processor 510 and / or other processor devices to execute various workflow containers (in the appropriate order defined by the workflow) until all workflow containers have been executed. The container orchestration engine 502 further generates diagnostic workflow results from the overall workflow. The diagnostic workflow system 106 can provide the results for display within a user interface on a client device (e.g., client device 304). By scheduling resource allocations to execute workflow containers, the diagnostic workflow system 106 (via the container orchestration engine 502) can efficiently utilize computing resources by reducing downtime and preventing duplicate processes (especially for FPGAs or other series-oriented devices) to reduce crashes and slowdowns.
[0085] As described above, in certain embodiments, the diagnostic workflow system 106 determines the execution mode of the container orchestration engine. In particular, the diagnostic workflow system 106 determines the execution mode based on a standardized genetic diagnostic protocol that the container orchestration engine is intended to satisfy when executing a particular workflow (e.g., an external sequencing diagnostic workflow). Figure 6 illustrates an exemplary flow for determining or identifying diagnostic applications (and / or incompatible diagnostic applications) that are compatible with the execution mode of the container orchestration engine, according to one or more embodiments.
[0086] As illustrated in Figure 6, the diagnostic workflow system 106 utilizes the container orchestration engine 602 to execute a diagnostic workflow that includes various workflow containers, such as the workflow container 608. In fact, as described, the container orchestration engine 602 utilizes the workflow execution service 604 and the workflow main container 606 to trigger and facilitate the execution of the workflow. In some embodiments, the external sequencing diagnostic workflow (or its parent application) specifies a particular execution mode 610 for the container orchestration engine 602. For example, the workflow may represent a standardized genetic diagnostic protocol such as IVD, or another standardized analysis protocol such as IUO or RUO that the container orchestration engine 602 must adhere to when executing the workflow.
[0087] Based on detecting or determining the execution mode 610 for the container orchestration engine 602, the diagnostic workflow system 106 identifies a set of compatible diagnostic applications 612. For example, the diagnostic workflow system 106 identifies diagnostic applications or workflows approved by a specific regulatory body (e.g., the FDA) to satisfy a standardized genetic diagnostic protocol such as IVD. The diagnostic workflow system 106 further excludes or prevents access to other diagnostic applications that do not satisfy a standardized genetic diagnostic protocol (e.g., applications that satisfy lower bars such as IUO or RUO but not IVD). In another embodiment, the diagnostic workflow system 106 identifies analytical applications that satisfy (at least) lower bars such as IUO or RUO (including applications that satisfy higher bars, but preventing access to other applications or workflows that do not). Thus, the diagnostic workflow system 106 ensures that only compatible diagnostic applications 612 that are compatible with the execution mode 610 are available for a particular sequencing run. As shown, the diagnostic workflow system 106 determines that applications A 614 and C 618 are compatible with execution mode 610 of the container orchestration engine 602, while application B 616 is not compatible with execution mode 610.
[0088] The diagnostic workflow system 106 further grants or permits access to compatible diagnostic applications 612. For example, the diagnostic workflow system 106 grants access to a client device (e.g., client device 304) to select one or more applications or workflows and apply them to a sequencing run performed in execution mode 610 (e.g., by presenting them via a user interface). In some cases, the diagnostic workflow system 106 grants access to the container orchestration engine 602 to execute all or part of one or more compatible diagnostic applications 612 (e.g., specific workflow containers within them). For example, as part of the execution of an external sequencing diagnostic workflow, the diagnostic workflow system 106 may only access applications that conform to standardized genetic diagnostic protocols. In some cases, the diagnostic workflow system 106 grants access to individual workflow containers within a compatible diagnostic application 612 for application to or adaptation to an external sequencing diagnostic workflow implemented by the container orchestration engine 602.
[0089] As described above, in certain embodiments, the diagnostic workflow system 106 utilizes containers and pods to execute external workflows associated with nucleotide reads and base calls of sample sequences. In particular, the diagnostic workflow system 106 can analyze sequencing data via the diagnostic workflow to identify genetic markers or genetic traits exhibited within a genomic sample. Figure 7 illustrates illustrative diagrams of the components, applications, devices, and containers of a system architecture (e.g., installed on a local server device) associated with implementing an external diagnostic workflow according to one or more embodiments.
[0090] As illustrated in Figure 7, the diagnostic workflow system 106 communicates with various components or systems and utilizes a sequencing device (e.g., sequencing device 114) to perform sequencing operations used and / or commanded by a version of the variant analysis model on a local server (e.g., variant analysis model 107 on server device 102). For example, the diagnostic workflow system 106 communicates with a cloud-based interface for BaseSpace Sequencing Hub ("BSSH") or Research Only ("RUO") and a laboratory information management system ("lab information management system, LIMS") to generate base calls and other sequencing data for the nucleotide bases of a genomic sample.
[0091] Based on information from BSSH RUO and / or LIMS, the diagnostic workflow system 106 performs a real-time analysis ("real-time analysis, RTA") of the sample. More specifically, the diagnostic workflow system 106 performs an RTA to determine base calls, variant calls, and / or various metrics from the nucleotide base of the genomic sample according to the sequencing plan. Based on the RTA, the diagnostic workflow system 106 generates a binary base call ("binary base call, BCL") file containing the raw data generated and output by one or more sequencing runs (e.g., via RTA). In fact, the BCL file may show base calls, variant calls, and / or other sequencing information for interpretation by variant analysis models and / or some other systems.
[0092] To organize or plan RTA sequencing runs, the diagnostic workflow system 106 provides control software (including, for example, a user interface) for planning or scheduling sequencing runs for specific samples. In fact, the diagnostic workflow system 106 provides control software and a user interface for planning one or more sequencing runs, for example, to test a genomic sample for specific genetic markers according to planning parameters. For example, the control software allows the user to specify parameters for a sequencing run and / or test for specific markers. As shown, the diagnostic workflow system 106 can integrate control software for a sequencing instrument with a user interface web portal (including a standalone web browser and control software integration) to interface with the sequencing instrument for planning sequencing runs.
[0093] In some cases, the diagnostic workflow system 106 facilitates local planning for sequencing runs, and the planning software (e.g., control software) is hosted on a local server device, such as a local edge server. In these, or other, cases, the diagnostic workflow system 106 facilitates cloud planning for sequencing runs, and the planning software (e.g., control software) is hosted on a cloud server rather than a local server. Similarly, the execution of a variant analysis model can be local or cloud-based, depending on whether the server hosting the variant analysis model is a local server (e.g., server device 102). Thus, (i) the diagnostic workflow system 106 can be executed locally on the sequencing device 114 or on a local server device located nearby, or remotely on a cloud-based server device, in combination with (ii) a variant analysis model 107 that is executed locally on the sequencing device 114 or on a local server device located nearby, or remotely on a cloud-based server device.
[0094] As further illustrated in Figure 7, the system architecture 700 of the diagnostic workflow system 106 includes or communicates with containers or systems associated with one or more core services. In fact, as shown, the diagnostic workflow system 106 includes services of the system architecture 700. To manage or organize the various services of the system architecture 700, the system architecture 700 includes a container orchestration engine 701 (e.g., K3S or Kubernetes) to manage and implement various pods and containers associated with performing genomic analysis via the diagnostic workflow as described herein. As described, the diagnostic workflow system 106 utilizes the container orchestration engine 701 to organize or coordinate the diagnostic workflow to analyze genomic sequences for base calls, variant calls (e.g., as indicated by applications of third-party systems). The container orchestration engine 701 also includes pods and containers to perform other functions, including user management, application management, run management, variant analysis model management, instrument management, data copying, and audit logging.
[0095] For example, the system architecture 700 includes a user management service 702 which includes one or more user management pods or containers. The user management service 702 performs various processes or functions to provide the entire single sign-on (SSO) experience system. Specifically, the user management service 702 may include one or more containers or pods which include, or access to, user information for a third-party system in order to determine a diagnostic workflow (e.g., from one of the third-party systems) for analyzing a genome sequence, including user settings or preferences for performing a diagnostic workflow. Based on the determination of the diagnostic workflow and / or user settings, the user management service 702 may communicate with other services of the system architecture 700 to initiate the execution of the diagnostic workflow and, accordingly, analyze the genome sequence.
[0096] In addition, the system architecture 700 includes or utilizes an application management service 704 that communicates with the container orchestration engine 701. For example, the application management service 704 manages the installation of application packages for diagnostic workflows. In some cases, the application management service 704 further includes a resource manager. The resource manager can access or utilize genome analyzer resources as specified by the application specification and / or as part of the diagnostic workflow. More specifically, the resource manager identifies resource labels for accessing specified resources, such as FPGAs or CPUs, as schedulable resources for access via the container orchestration engine. In fact, in some cases, the application management service 704 includes (or receives from a third-party system) an application specification indicating an FPGA or CPU or some other genome analyzer for running the diagnostic workflow of a genome analysis application (or a particular workflow pod), and therefore the resource manager accesses or communicates with the specified instrument (or other resource) to facilitate the execution of the genome analysis application (or a particular workflow pod).
[0097] As further shown, the system architecture 700 includes or utilizes a run management and orchestration service 706. More specifically, the run management and orchestration service 706 includes one or more containers or pods for facilitating and executing genomic analysis through diagnostic workflows such as sequencing runs, primary analyses, secondary analyses, or tertiary analyses. In fact, the run management and orchestration service 706 includes computer code or instructions for executing sequencing runs (and / or further analyses) according to the installed version of the variant analysis model. For example, the run management and orchestration service 706 communicates with the workflow engine 714 to execute custom diagnostic workflows for applications such as applications associated with third-party systems (e.g., oncology assay applications such as the TSO500 application, QC application, or another application). The run management and orchestration service 706 further includes code for communicating with the data copy service 712 to copy input and output sequencing data (e.g., from BCL files generated by the sequencing instrument) for performing genome analysis and / or storing in a database such as local network attached storage ("network attached storage, NAS"), server message block ("server message block, SMB"), or common internet file system ("common internet file system, CIFS").
[0098] In addition, the system architecture 700 includes a variant analysis model management service 708. Specifically, the variant analysis model management service 708 includes one or more containers or pods for managing variant analysis models (e.g., variant analysis model 107) for performing genomic analysis. For example, the variant analysis model management service 708 implements a specific diagnostic workflow using a variant analysis model to detect a genetic marker for a particular state within a sample genome sequence. In addition, the variant analysis model management service 708 manages model peripherals such as licensing, self-testing, and version authentication for variant analysis models.
[0099] As further illustrated in Figure 7, the system architecture 700 includes an instrument management service 710. In one or more embodiments, the instrument management service 710 includes one or more containers or pods for pairing and monitoring instruments used as part of a sequencing workflow and / or a post-sequencing genome analysis workflow. For example, the instrument management service 710 manages sequencing instruments and / or variant analysis model instruments to pair compatible instruments with (or vice versa) indicated versions of variant analysis models. The system architecture 700 further includes an audit logging service 716 for monitoring and logging the performance of instruments, variant analysis model components, and / or containers within the application workflow. For example, the audit logging service 716 detects and logs errors or other audit information associated with the system architecture 700.
[0100] As suggested above in Figure 1 and elsewhere, using the system architecture 700 illustrated in Figure 7, the diagnostic workflow system 106 can be deployed locally on an edge server (e.g., local server device 102) or in the cloud, such as on a cloud-based server hosting Illumina Connected Analytics ("ICA") and / or on a cloud-based server from Amazon Web Services ("AWS"). For example, the diagnostic workflow system 106 can run locally on a local server device or remotely on a cloud-based server device as part of planning software that plans resources based on user input for sequencing runs or other assays, and the variant analysis model 107 can similarly run locally on a local server device or remotely on a cloud-based server device to analyze BCL data and determine variant calls or other metrics.
[0101] Referring now to Figure 8, this figure illustrates an exemplary flowchart of a set of operations in which an external sequencing diagnostic workflow is performed using a container orchestration engine, according to one or more embodiments. Figure 8 illustrates the operations according to one embodiment, but alternative embodiments may omit, add, rearrange, and / or modify any of the operations shown in Figure 8. The operations in Figure 8 may be performed as part of a method. Alternatively, a non-temporary computer-readable storage medium, when executed by one or more processors, may contain instructions that cause a computing device to perform the operations depicted in Figure 8. In a further embodiment, the system comprises at least one processor and a non-temporary computer-readable medium, when executed by one or more processors, containing instructions that cause the system to perform the operations in Figure 8.
[0102] As shown in Figure 8, the sequence of operations 800 includes an operation 802 for identifying nucleotide base calls. In particular, operation 802 may include identifying nucleotide base calls generated by a variant analysis model for a sample nucleotide sequence.
[0103] In addition, the set of operations 800 includes operation 804, which requires the container orchestration engine to implement an external sequencing diagnostic workflow. Specifically, operation 804 may include requiring the container orchestration engine associated with the variant analysis model to implement an external sequencing diagnostic workflow for performing a diagnostic analysis on nucleotide base calls for sample nucleotide sequences. For example, operation 804 may involve implementing an external sequencing diagnostic workflow identified from an external application that is separate from the container orchestration engine and the variant analysis model.
[0104] As further illustrated in Figure 8, the sequence of operations 800 includes an operation 806 for identifying one or more workflow containers. In particular, operation 806 may include identifying one or more workflow containers associated with each function of the external sequencing diagnostic workflow.
[0105] Furthermore, the set of operations 800 includes operation 808, which executes an external sequencing diagnostic workflow using one or more workflow containers. In particular, operation 808 may include executing the external sequencing diagnostic workflow by utilizing a container orchestration engine to implement one or more workflow containers. For example, operation 808 may involve executing the external sequencing diagnostic workflow after variant analysis by a variant analysis model has generated nucleotide base calls for the sample nucleotide sequence, or during variant analysis by a variant analysis model. In some cases, operation 808 involves selectively preventing one or more workflow containers from accessing the sequencing data of the sample nucleotide sequence while the external sequencing diagnostic workflow is being executed. In these, or other, cases, operation 808 involves utilizing a container orchestration engine located on a server device hosting the variant analysis model. In fact, operation 808 may involve executing an external sequencing diagnostic workflow generated by an external system separate from the container orchestration engine and the variant analysis model.
[0106] In some embodiments, the sequence of operations 800 includes operations that determine a diagnostic execution mode corresponding to a standardized genetic diagnostic protocol. The sequence of operations 800 may also include operations that grant the client device access only to diagnostic applications compatible with the diagnostic execution mode. In certain cases, the sequence of operations 800 includes operations that receive an external sequencing diagnostic workflow generated by an external device operated by an external entity.
[0107] The sequence of actions 800 may include actions that control access to different workflow data sources for one or more workflow containers in order to prevent access to sequencing data. For example, the sequence of actions 800 may include actions that prevent one of the workflow containers from accessing the sequencing data of sample nucleotide sequences. Preventing a workflow container from accessing the sequencing data of sample nucleotide sequences may involve preventing the workflow container from accessing one or more workflow data sources, including input directories, output directories, and application directories.
[0108] The sequence of operations 800 may also include operations to receive label metrics that define versions of variant analysis models and memory allocations to be used to execute the external sequencing diagnostic workflow. Additionally, the sequence of operations 800 may include operations to specify multiple workflow data sources that store different types of workflow data. Furthermore, the sequence of operations 800 may include operations to activate access to a first workflow data source among multiple workflow data sources for one workflow container from among one or more workflow containers, while preventing access to other workflow data sources among multiple workflow data sources. The sequence of operations 800 may include operations to mount multiple workflow data sources as read-only for one or more workflow containers. In some embodiments, the sequence of operations 800 includes operations to trigger the execution of the external sequencing diagnostic workflow by receiving post-defined parameters for implementing the external sequencing diagnostic workflow via the container orchestration engine.
[0109] In some embodiments, the sequence of operations 800 includes operations that satisfy one or more standardized genetic diagnostic protocols while also running an external sequencing diagnostic workflow, by encoding a workflow execution application that grants the external sequencing diagnostic workflow read-only access to sequencing data associated with nucleotide base calls of sample nucleotide sequences during the execution of the external sequencing diagnostic workflow.
[0110] The methods described herein can be used in conjunction with various nucleic acid sequencing techniques. Particularly applicable techniques involve attaching nucleic acids to fixed positions within an array so that their relative positions do not change, and the array being repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of a target nucleic acid may be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.
[0111] SBS technology generally involves the enzymatic elongation of a nascent nucleic acid chain by the repeated addition of nucleotides to a template chain. In conventional SBS methods, a single nucleotide monomer may be delivered to the target nucleotide in the presence of polymerase at each delivery. However, the method described herein allows for the delivery of two or more types of nucleotide monomers to the target nucleic acid in the presence of polymerase during delivery.
[0112] SBS can utilize nucleotide monomers with a terminator moiety, or nucleotide monomers lacking any terminator moiety. Methods utilizing nucleotide monomers lacking a terminator include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in further detail below. In methods using nucleotide monomers without a terminator, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques utilizing nucleotide monomers with a terminator moiety, the terminator may be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or it may be reversible, as in sequencing methods developed by Solexa (now Illumina).
[0113] SBS technology can use nucleotide monomers having a labeling moiety or nucleotide monomers lacking a labeling moiety. Therefore, integration events can be detected based on the characteristics of the label, such as fluorescence of the label; the characteristics of the nucleotide monomer, such as molecular weight or charge; and by-products of nucleotide integration, such as pyrophosphate release. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or alternatively, two or more different labels may be distinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent may have different labels, and they can be distinguished using a suitable optical system, as exemplified by the sequencing method developed by Solexa (now Illumina).
[0114] A preferred embodiment is pyrosequencing (pyrosequencing) technology. Pyrosequencing detects the release of inorganic pyrophosphate (Ppi) when specific nucleotides are incorporated into the nascent DNA chain (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing," Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate," Science 281(5375),363, U.S. Patent Nos. 6,210,891, 6,258,568 and 6,274,320 (the entire disclosure thereof is incorporated herein by reference). In pyrosequencing, released Ppi can be detected by immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of ATP produced is detected via photons produced by luciferase. The nucleic acids to be sequenced can be bound to features in an array, and the array can be imaged to capture the chemiluminescent signal produced by incorporating the nucleotides into the array features. After processing the array with a specific nucleotide type (e.g., A, T, C, or G), an image can be obtained. The images obtained after the addition of each nucleotide type differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative position of each feature remains unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after processing an array with each different nucleotide type can be processed in the same manner as images obtained from different detection channels for a reversible terminator-based sequencing method, as illustrated herein.
[0115] In another exemplary type of SBS, cyclic sequencing is achieved by stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, such as those described in International Publication No. 04 / 018497 and U.S. Patent No. 7,057,026, whose disclosure is incorporated herein by reference. This technique has been commercialized by Solexa (now Illumina Inc.) and is also described in International Publication No. 91 / 06678 and International Publication No. 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators with both ends reversible and the fluorescent label cleaved facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-operated to efficiently incorporate and extend these modified nucleotides.
[0116] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be captured after the incorporation of the label into the arrayed nucleic acid features. In certain embodiments, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally different label. Four images can then be obtained by using a selective detection channel for each of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image shows a nucleic acid feature incorporating a particular type of nucleotide. Because the sequence content of each feature part is different, different feature parts may or may not be present in different images. However, the relative positions of the features remain unchanged within the images. Images obtained from such a reversible terminator-SBS method can be stored, processed, and analyzed as described herein. Following the image acquisition step, the label can be removed, and the reversible terminator portion can be removed for subsequent nucleotide addition and detection cycles. Removing the label after detection in a specific cycle and before subsequent cycles has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are listed below.
[0117] In certain embodiments, some or all of the nucleotide monomers may include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore (fluor) may be a fluorophore (fluor) attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other methods separate the chemistry of the terminator from the cleavage of the fluorescent label (this entire method is incorporated herein by reference, Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005)). Ruparel et al. describe the development of a reversible terminator that uses a small amount of 3' allyl group to block elongation but can be easily unblocked by short-term treatment with a palladium catalyst. The fluorophore was attached to the group via a photocleavable linker that can be readily cleaved by 30 seconds of exposure to long-wavelength UV light. Therefore, either disulfide reduction or photocleavage can be used as a cleavable linker. Another method to reversible termination is the use of a natural termination followed by the placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator via steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore (fluor) and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patents 7,427,673 and 7,057,026, and these disclosures are incorporated herein by reference in their entirety.
[0118] Additional exemplary SBS systems and methods that can be used in conjunction with the methods and systems described herein are described in U.S. Patent Publication No. 2007 / 0166705, U.S. Patent Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Publication No. 2006 / 0240439, U.S. Patent Publication No. 2006 / 0281109, International Publication No. 05 / 065814, U.S. Patent Publication No. 2005 / 0100900, International Publication No. 06 / 064199, International Publication No. 07 / 010,251, U.S. Patent Publication No. 2012 / 0270305, and U.S. Patent Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entirety.
[0119] Several embodiments can utilize the detection of four different nucleotides using fewer than four different labels. For example, SBS can be carried out using the method and system described in the incorporated material, U.S. Patent Application Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types can be detected at the same wavelength but can be distinguished based on a difference in intensity for one member of the pair, or based on a change in one member of the pair (e.g., through chemical modification, photochemical modification, or physical modification) that causes a noticeable signal to appear or disappear compared to the signal detected for the other members of the pair. As a second example, three of the four different nucleotide types can be detected under specific conditions, while a fourth nucleotide type has no detectable label under those conditions or is minimally detectable under those conditions (e.g., minimal detection by background fluorescence). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid can be determined based on the absence of any signal or minimal detection. As a third example, one nucleotide type may include a label that is detected by two different channels, while other nucleotide types are detected by one or fewer channels. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method using a first nucleotide type detected in a first channel (e.g., dATP with a label detectable in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detectable in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detectable in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type that is not detected in any channel or has minimally detected labels (e.g., unlabeled dGTP).
[0120] Furthermore, as described in the incorporated document, U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such a so-called one-dye sequencing method, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.
[0121] Several embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligases to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. Oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence into which the oligonucleotide hybridizes. As with other SBS methods, an image can be obtained after processing an array of nucleic acid features with a labeled sequencing reagent. Each image shows a nucleic acid feature with a specific type of label incorporated. Because the sequence content of each feature region is different, different images may or may not contain different feature regions, but the relative positions of the feature regions remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used in conjunction with the methods and systems described herein are described in U.S. Patents 6,969,488, 6,172,218, and 6,306,597, and these disclosures are incorporated herein by reference in their entirety.
[0122] Several embodiments can utilize nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis." Acc. Chem. Res. 35: 817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. Golovchenko, "DNA molecules and composition in solid nanopore microscopy." Nat. Mater. 2: 611-615 (2003). These disclosures are incorporated herein by reference in their entirety). In such embodiments, the target nucleic acid passes through the nanopore. The nanopore may be a synthetic pore such as α-hemolysin or a biological membrane protein. As target nucleic acids pass through nanopores, each base pair can be identified by measuring the variation in the electrical conductance of the pores. (U.S. Patent No. 7,001,792, “Advances toward ultrafast DNA sequencing using solid-phase nanopores,” Clin. Chem. 53, 1996-2001 (2007), “Nanopore-based single-molecule DNA analysis,” Nanomed. 2, 459-481 (2007), “Single-molecule nanopore devices detect DNA polymerase activity at single-nucleotide resolution,” J. Am Chem. Soc. 130, 818-820 (2008). These disclosures are incorporated herein by reference in their entirety.) Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as an image in accordance with the exemplary processing of optical and other images described herein.
[0123] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected, for example, via fluorescence resonance energy transfer (FRET) interaction between a fluorophore-containing polymerase and a γ-phosphate-labeled nucleotide, as described in U.S. Patent Nos. 7,329,492 and 7,211,414, each incorporated herein by reference; or nucleotide incorporation can be detected using a zero-mode waveguide, as described in U.S. Patent No. 7,315,019, each incorporated herein by reference, and fluorescent nucleotide analogs and manipulated polymerases, as described in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0109082, each incorporated herein by reference. Illumination can be limited to a zeptolite-scale volume around the surface-tethered polymerase so that the incorporation of fluorescently labeled nucleotides can be observed with low background (see Levene, MJ et al., “Zero-mode waveguide for single-molecule analysis at high concentrations,” Science, 299, 682-686 (2003); Lundquist, PM et al., “Parallel confocal detection of single molecules in real time,” Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for target immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures,” Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008); these disclosures are incorporated herein by reference in their entirety). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0124] Some SBS embodiments include the detection of protons released during the incorporation of nucleotides into the extension product. For example, sequencing based on the detection of released protons may use electrodetectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Publications 2009 / 0026082(A1), 2009 / 0127589(A1), 2010 / 0137143(A1), or 2010 / 0282617(A1), each of which is incorporated herein by reference. The methods herein for amplifying target nucleic acids using dynamic exclusion can be readily applied to substrates used for proton detection. More specifically, the methods herein can be used to generate a clonal population of amplicons used for proton detection.
[0125] The SBS method described above can be advantageously implemented in a multiplex format so that multiple different target nucleic acids are manipulated simultaneously. In certain embodiments, different target nucleic acids can be processed on a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a multiplex format. In embodiments using surface-bound target nucleic acids, the target nucleic acids may be in array form. In array form, the target nucleic acids can typically be bound to the surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent bonding, binding to beads or other particles, or binding to polymerase or other molecules bound to the surface. The array may contain a single copy of the target nucleic acid at each site (also referred to as a feature), or multiple copies having the same sequence may be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, which are described in more detail below.
[0126] The method described herein can use arrays having any of the following densities of feature parts: for example, at least about 10 feature parts / cm², 100 feature parts / cm², 500 feature parts / cm², 1,000 feature parts / cm², 5,000 feature parts / cm², 10,000 feature parts / cm², 50,000 feature parts / cm², 100,000 feature parts / cm², 1,000,000 feature parts / cm², 5,000,000 feature parts / cm², or more.
[0127] An advantage of the methods described herein is that they provide the rapid and efficient parallel detection of multiple target nucleic acids. Therefore, this disclosure provides an integrated system that allows nucleic acids to be prepared and detected using techniques known in the art, such as those exemplified above. Accordingly, the integrated system of this disclosure may include a fluid component capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, and the system may include components such as pumps, valves, reservoirs, and fluid lines. A flow cell may constitute and / or be used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent No. 2010 / 0111768(A1) and U.S. Patent Application No. 13 / 273666, each of which is incorporated herein by reference. As exemplified with respect to flow cells, one or more fluid components of the integrated system may be used in amplification and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more fluid components of the integrated system may be used for the delivery of sequencing reagents in the amplification method described herein and in sequencing methods such as those exemplified above. Alternatively, the integrated system may include separate fluid systems for carrying out amplification and detection methods. Examples of integrated sequencing systems capable of producing amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq® platform (Illumina Inc., San Diego, CA) and the apparatus described in U.S. Patent Application No. 13 / 273,666, incorporated herein by reference.
[0128] The sequencing system described above sequences nucleic acid polymers present in a sample received by a sequencing device. As defined herein, “sample” and its derivatives are used in the broadest sense and include any sample, culture, etc., suspected to contain a target. In some embodiments, the sample includes DNA, RNA, PNA, LNA, chimeric or hybrid nucleic acids. The sample may include any biological, clinical, surgical, agricultural, air, or water sample containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, e.g., genomic DNA, fresh-frozen or formalin-fixed paraffin-embedded nucleic acid sample. The sample may also originate from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples from a single individual such as tumor and normal tissue samples (fit), or a sample from a single source containing two different forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of the nucleic acid material may include nucleic acids obtained from newborns, such as those typically used in newborn screening.
[0129] Nucleic acid samples may include high molecular weight substances such as genomic DNA (gDNA). Samples may include low molecular weight substances such as nucleic acid molecules obtained from FFPE or stored DNA samples. In another embodiment, the low molecular weight substance includes enzymatically or mechanically fragmented DNA. Samples may include cell-free circulating DNA. In some embodiments, samples may include nucleic acid molecules obtained from biopsies, tumors, scrapes, swabs, blood, mucus, urine, plasma, semen, hair, laser-captured microscopy, surgical excisions, and other clinical or laboratory-obtained samples. In some embodiments, samples may be epidemiological, agricultural, forensic, or pathogenic samples. In some embodiments, samples may include nucleic acid molecules obtained from animals such as humans or mammalian sources. In another embodiment, samples may include nucleic acid molecules obtained from non-mammalian sources such as plants, bacteria, viruses, or fungi. In some embodiments, the source of nucleic acid molecules may be stored or extinct samples or species.
[0130] Furthermore, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA, from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or forensic samples obtained by law enforcement agencies, one or more military services or members of such services. The nucleic acid sample may be crude DNA containing a purified sample or lysate derived from, for example, an oral swab, paper, cloth, or other substrate that can be impregnated with saliva, blood, or other bodily fluids. Thus, in some embodiments, the nucleic acid sample may include small amounts of DNA or fragmented portions of DNA, such as genomic DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from the victim's hair, skin, tissue sample, autopsy, or corpse. In some embodiments, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence may include nucleic acids obtained from non-human DNA, such as microbial, plant, or entomological DNA. In some embodiments, the target sequence or amplified target sequence is intended for human identification. In some embodiments, this disclosure generally relates to a method for identifying features of forensic specimens. In some embodiments, this disclosure generally relates to a human identification method using one or more target-specific primers disclosed herein, or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic specimen or human identification specimen containing at least one target sequence may be amplified using one or more of the target-specific primers disclosed herein, or using the primer criteria outlined herein.
[0131] The components of the diagnostic workflow system 106 may include software, hardware, or both. For example, the components of the diagnostic workflow system 106 may include one or more instructions stored on a computer-readable storage medium and executable by the processor of one or more computing devices (e.g., client device 108). When executed by one or more processors, the computer-executable instructions of the diagnostic workflow system 106 can cause the computing device to perform the bubble detection method described herein. Alternatively, the components of the diagnostic workflow system 106 may include hardware such as dedicated processing units for performing a particular function or set of functions. Additionally, or alternatively, the components of the diagnostic workflow system 106 may include a combination of computer-executable instructions and hardware.
[0132] Furthermore, the components of the diagnostic workflow system 106 that perform the functions described herein may be implemented, for example, as part of a standalone application, as a module of an application, as a plug-in to an application, as a library function that can be called by other applications, and / or as a cloud computing model. Thus, the components of the diagnostic workflow system 106 may be implemented as part of a standalone application on a personal computing device or a mobile device. Additionally or alternatively, the components of the diagnostic workflow system 106 may be implemented in any application that provides sequencing determination services, including, but not limited to, Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumina," "BaseSpace," "DRAGEN," and "TruSight" are registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0133] Embodiments of the present disclosure may include, or utilize, a dedicated or general-purpose computer, including, for example, one or more processors and system memory, among other computer hardware, as will be discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be embodied in a non-temporary computer-readable medium and at least partially implemented as instructions executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-temporary computer-readable medium (e.g., memory), executes those instructions, and thereby performs one or more processes, including one or more of the processes described herein.
[0134] A computer-readable medium can be any available medium that can be accessed by a general-purpose computer system or a dedicated computer system. A computer-readable medium that stores computer-executable instructions is a non-temporary computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, embodiments of the present disclosure may include, but are not limited to, two distinctly different types of computer-readable mediums: a non-temporary computer-readable storage medium (device) and a transmission medium.
[0135] Non-temporary computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSDs) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store desired program code means in the form of computer-executable instructions or data structures and can be accessed by a general-purpose or dedicated computer.
[0136] A “network” is defined as one or more data links that enable the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred to or provided to a computer via a network or another communication connection (either hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately recognizes the connection as a transmission medium. A transmission medium can be used to carry desired program code means in the form of computer-executable instructions or data structures and may include networks and / or data links that can be accessed by general-purpose or dedicated computers. The above combinations should also be included within the scope of computer-readable media.
[0137] Furthermore, upon reaching various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-temporary computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link may be buffered in RAM within a network interface module (e.g., NIC) and then ultimately transferred to computer system RAM and / or a less volatile computer storage medium (device) within the computer system. Therefore, it should be understood that non-temporary computer-readable storage media (devices) can be included in computer system components that also (or more primarily) utilize the transmission medium.
[0138] Computer executable instructions include instructions and data that, when executed by a processor, cause a general-purpose computer, a dedicated computer, or a dedicated processing unit to perform a certain function or set of functions. In some embodiments, computer executable instructions are executed on a general-purpose computer and transform the general-purpose computer into a dedicated computer implementing the elements of the Disclosure. Computer executable instructions may be, for example, binary, intermediate format instructions such as assembly language, or even source code. While the subject matter is described in language specific to structural features and / or methodological behavior, it should be understood that the subject matter as defined in the appended claims is not necessarily limited to the described features or behaviors described above. Rather, the described features and behaviors are disclosed as exemplary forms that implement the claims.
[0139] Those skilled in the art will understand that the disclosure can be implemented in network computing environments having many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure can also be implemented in distributed system environments where both local and remote computer systems linked over a network (by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) perform tasks. In a distributed system environment, program modules can reside in both local and remote memory storage devices.
[0140] Embodiments of this disclosure can also be implemented in a cloud computing environment. In this specification, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing may be used in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly configured via virtualization, exposed with low administrative effort or service provider interaction, and then scaled accordingly.
[0141] A cloud computing model can consist of various characteristics such as on-demand self-service, wide-area network access, resource pooling, rapid resilience, and measured service. A cloud computing model can also expose various service models, such as Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). A cloud computing model can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, and hybrid cloud. In this specification and in the claims, “cloud computing environment” refers to an environment in which cloud computing is employed.
[0142] Figure 9 illustrates a block diagram of a computing device 900 that may be configured to perform one or more of the processes described above. It will be understood that one or more computing devices, such as computing device 900, may implement the diagnostic workflow system 106 and the genome analysis platform 104. As shown in Figure 9, computing device 900 may comprise a processor 902, memory 904, storage device 906, I / O interface 908, and communication interface 910, which may be communicatively coupled by a communication infrastructure 912. In certain embodiments, computing device 900 may include fewer or more components than those shown in Figure 9. The following paragraphs describe in further detail the components of computing device 900 shown in Figure 9.
[0143] In one or more embodiments, the processor 902 includes hardware for executing instructions, such as instructions that constitute a computer program. Not limited to, as an example, to execute instructions for dynamically modifying a workflow, the processor 902 may retrieve (or fetch) instructions from internal registers, internal cache, memory 904, or storage device 906, decode them, and execute them. Memory 904 may be volatile or non-volatile memory used to store data, metadata, and programs for execution by the processor. Storage device 906 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for carrying out the methods described herein.
[0144] The I / O interface 908 enables a user to provide input to the computing device 900, receive output from it, transfer data to it, and receive data from it. The I / O interface 908 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of such I / O interfaces. The I / O interface 908 may include, but is not limited to, one or more devices for presenting output to the user, including a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 908 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be useful in a particular implementation.
[0145] The communication interface 910 may include hardware, software, or both. In any case, the communication interface 910 may provide one or more interfaces for communication (e.g., packet-based communication) between the computing device 900 and one or more other computing devices or networks. In examples, but not limited to, the communication interface 910 may include a network interface controller (NIC) or network adapter for communication with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communication with wireless networks such as Wi-Fi.
[0146] Additionally, the communication interface 910 can facilitate communication with various types of wired or wireless networks. The communication interface 910 can also facilitate communication using various communication protocols. The communication infrastructure 912 may also include hardware, software, or both that connect the components of the computing device 900 to one another. For example, the communication interface 910 uses one or more networks and / or protocols to enable multiple computing devices connected by a particular infrastructure to communicate with each other to carry out one or more aspects of the processes described herein. For example, a sequencing process can enable multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0147] In the aforementioned specification, the disclosure was described with reference to certain exemplary embodiments. Various embodiments and aspects of the disclosure are described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings are illustrative of the disclosure and should not be construed as limiting the disclosure. Numerous specific details are described in order to provide a complete understanding of the various embodiments of the disclosure.
[0148] This disclosure may be embodied in other specific forms without departing from its spirit or essential features. The embodiments described herein should be considered in all respects to be illustrative and not limiting. For example, the methods described herein may be carried out using fewer or more steps / operations, or the steps / operations may be carried out in a different order. In addition, the steps / operations described herein may be repeated or carried out in parallel with each other, or in parallel with different occurrences of the same or similar steps / operations. Accordingly, the scope of this application is indicated by the appended claims rather than by the foregoing description. All changes included in the meaning of the claims and the scope of equivalents are encompassed within those scopes.
[0149] As used herein, the term “object” includes all objects suitable for imaging, viewing, analyzing, inspecting, or profiling using the optical systems described herein. As mere examples, an object may include a semiconductor wafer or chip, a recordable medium, a sample, a flow cell, a particle, a slide, or a microarray. Generally, an object includes one or more surfaces and / or one or more interfaces whose profile the user may wish to image, view, analyze, inspect, and / or determine. An object may have surfaces or interfaces with relief features such as wells, pits, ridges, bumps, or beads.
[0150] As indicated above in the description of “Sample,” the sample may be imaged or scanned for subsequent analysis. In certain embodiments, the sample may include the biological or chemical substance of interest and, optionally, an optical substrate supporting the biological or chemical substance. Thus, the sample may or may not include an optical substrate. As used herein, the term “biological or chemical substance” is not intended to be limiting and may include a variety of biological or chemical substances suitable for imaged or examined with the optical systems described herein. For example, biological or chemical substances include biomolecules such as nucleosides, nucleic acids, polynucleotides, oligonucleotides, proteins, enzymes, polypeptides, antibodies, antigens, ligands, receptors, polysaccharides, carbohydrates, polyphosphates, nanopores, organelles, lipid layers, cells, tissues, organisms, and biologically active compounds such as analogs or mimetic compounds of the aforementioned species.
[0151] Biological or chemical substances may be supported by optical substrates. As used herein, the term “optical substrate” is not intended to be limiting and may include a variety of materials that support a biological or chemical substance and allow the biological or chemical substance to be observed, imaged, and examined. For example, an optical substrate may include a transparent material that reflects part of the incident light and refracts part of the incident light. Alternatively, an optical substrate may be, for example, a mirror that completely reflects incident light so that light does not pass through the optical substrate. Typically, an optical substrate has a flat surface. However, an optical substrate may have a surface with relief features such as wells, pits, ridges, bumps, and beads.
[0152] In exemplary embodiments, the optical substrate is a flow cell having a channel through which nucleic acids are sequenced. However, in alternative embodiments, the optical substrate may comprise one or more slides, planar chips (such as those used in microarrays), or microparticles. Where the optical substrate comprises multiple microparticles supporting biological or chemical material, the microparticles may be held by another optical substrate, such as a slide or grooved plate. In certain embodiments, the optical substrate comprises a diffraction grating-based coding optical identification element similar to or identical to that described in the pending U.S. Patent Application No. 10 / 661,234, filed September 12, 2003, entitled “Diffraction Grating Based Optical Identification Element,” which is incorporated herein by reference in its entirety and discussed in more detail below. The bead cells or plates for holding the optical identification elements are described in U.S. Patent Application No. 10 / 661,836, pending, filed September 12, 2003, entitled "Method and Apparatus for Aligning Microbeads in Order to Interrogate the Same," and U.S. Patent No. 7,164,533, issued January 16, 2007, entitled "Hybrid Random Bead / Chip Based Microarray," as well as U.S. Patent Application No. 60 / 609,583, filed September 13, 2004, entitled "Improved Method and Apparatus for Aligning Microbeads in Order to Interrogate the Same," and "Method and Apparatus for Aligning Microbeads in Order to Interrogate the These may be similar to or identical to those described in U.S. Patent Application No. 60 / 1010,910, titled “Same,” each of which is incorporated herein by reference in whole.
[0153] As used herein, the terms “optical component” or “focusing component” include various elements that affect the transmission of light. Optical components may include, for example, reflectors, polarizers, beam splitters, collimators, lenses, filters, wedges, prisms, and mirrors.
[0154] As examples, the optical systems described herein may be constructed to include various components and assemblies as described in PCT application PCT / US07 / 07991, filed March 30, 2007, entitled "System and Devices for Sequence by Synthesis Analysis," and / or as various components and assemblies as described in PCT application PCT / US2008 / 077850, filed September 26, 2008, entitled "Fluorescence Excitation and Detection System and Method," and the complete subject matter of both is incorporated herein by reference in whole. In certain embodiments, the optical system may include various components and assemblies as described in U.S. Patent No. 7,329,860, and the complete subject matter of both is incorporated herein by reference in whole. The optical system may also include various components and assemblies as described in U.S. Patent Application No. 12 / 638,770, filed December 15, 2009, and the complete subject matter of both is incorporated herein by reference in whole.
[0155] In certain embodiments, the methods and optical systems described herein may be used to sequence nucleic acids. For example, the synthetic sequencing (SBS) protocol is particularly applicable. In SBS, multiple fluorescently labeled nucleotides are used to sequence high-density clusters (possibly millions of clusters) of amplified DNA present on the surface of an optical substrate (e.g., a surface that at least partially defines the channels in the flow cell). The flow cell may contain a nucleic acid sample for sequencing, in which the flow cell is placed in a suitable flow cell holder. The sample for sequencing may take the form of single nucleic acid molecules separated from each other so as to be individually degradable, amplified populations of nucleic acid molecules in the form of clusters or other features, or beads attached to one or more molecules of nucleic acid. The nucleic acid may be prepared to contain oligonucleotide primers adjacent to an unknown target sequence. To initiate a first SBS sequencing cycle, one or more different labeled nucleotides and DNA polymerase, etc., may be flowed into / through the flow cell by a fluid flow subsystem (not shown). A single type of nucleotide can be added at once, or the nucleotides used in the sequencing procedure can be specially designed to have reversible termination properties, thus allowing each cycle of the sequencing reaction to occur simultaneously in the presence of several types of labeled nucleotides (e.g., A, C, T, G). The nucleotides may contain detectable labeling moieties, such as fluorophores. When four nucleotides are mixed together, the polymerase can select and incorporate the correct bases, and each sequence is extended by a single base. One or more lasers can excite nucleic acids and induce fluorescence. The fluorescence emitted from nucleic acids is based on the fluorophores of the incorporated bases, and different fluorophores may emit light at different wavelengths.Exemplary sequencing methods are described, for example, in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.
[0156] Other sequencing techniques applicable to the use of the methods and systems described herein include pyrosequencing, nanopore sequencing, and ligation sequencing. Particularly useful exemplary pyrosequencing techniques and samples are described in U.S. Patents 6,210,891, 6,258,568, 6,274,320, and Ronaghi, Genome Research 11:3-11 (2001), each of which is incorporated herein by reference. Similarly useful exemplary nanopore techniques and samples are described in Deamer et al., Acc. Res. 35:817-825 (2002), Li et al., Nat. Mater. 2:611-615 (2003), Soni et al., Clin Chem. 53:1996-2001 (2007), Healy et al., Nanomed. 2:459-481 (2007), and Cockroft et al., J. am. Chem. Soc. 130:818-820, and U.S. Patent No. 7,001,792, each of which is incorporated herein by reference. These systems can use any of a variety of samples, including substrates having beads produced by emulsion PCR, substrates having zero-mode waveguides, substrates having biological nanopores within a lipid bilayer, solid substrates having synthetic nanopores, and others known in the art. Such samples are described in relation to various sequencing techniques in the references cited above, and further in U.S. Patent Publication Nos. 2005 / 0042648, 2005 / 0079510, 2005 / 0130173, and International Publication No. 05 / 010145, each of which is incorporated herein by reference.
[0157] In other embodiments, the optical systems described herein may be used for the detection of samples including a microarray. A microarray may include a collection of different probe molecules attached to one or more substrates so that different probe molecules can be distinguished from one another according to their relative positions. The array may include a collection of different lobe molecules or probe molecules, each located at different addressable positions on the substrate. Alternatively, the microarray may include separate optical substrates, such as beads, each having a different collection of probe molecules, which can be identified according to the position of the optical substrate on the surface to which the substrate is attached, or according to the position of the substrate in a liquid. Exemplary arrays with separate substrates located on the surface include, but are not limited to, the Sentrix® Array or Sentrix® BeadChip Array available from Illumina®, Inc. (San Diego, CA), or others containing beads in wells, as described in U.S. Patent Nos. 6,266,459, 6,355,431, 6,770,441, and 6,859,570, and International Publication No. 00 / 63437 (each of which is incorporated herein by reference). Other arrays having particles on the surface include those described in U.S. Patent Application Publication No. 2005 / 0227252, International Publication No. 05 / 033681, and International Publication No. 04 / 024328, each of which is incorporated herein by reference.
[0158] For example, any of the various microarrays known in the art, including those described herein, can be used in embodiments of the present invention. A typical microarray contains regions, sometimes called features, each having a population of probes. The populations of probes in each region are typically homogeneous, each having a single type of probe, but in some embodiments, the populations can each be heterogeneous. The regions or features of the array are typically distinct and separated from one another by space. The size of the features and / or the spacing between regions can be varied so that the array can be high-density, medium-density, or low-density. High-density arrays are characterized by having regions separated by less than about 15 μm. Medium-density arrays have regions separated by about 15 to 30 μm, while low-density arrays have regions separated by more than 30 μm. Arrays useful in the present invention may have regions separated by less than 100 μm, 50 μm, 10 μm, 5 μm, 1 μm, or 0.5 μm. The apparatus or method of the embodiments of the present invention can be used to image an array with sufficient resolution to distinguish parts at the above density or density range.
[0159] Further examples of commercially available microarrays that can be used include, for example, Affymetrix® GeneChip® microarrays, or, for example, U.S. Patents No. 5,324,633, 5,744,305, 5,451,683, 5,482,867, 5,491,074, 5,624,711, 5,795,716, 5,831,070, 5,856,101, 5,858,659, 5,874,219, and 5,968,74 This includes other microarrays synthesized in accordance with the techniques sometimes referred to as VLSIPS® (Very Large Scale Immobilized Polymer Synthesis) techniques, as described in No. 0, No. 5,974,164, No. 5,981,185, No. 5,981,956, No. 6,025,601, No. 6,033,860, No. 6,090,555, No. 6,136,269, No. 6,022,963, No. 6,083,697, No. 6,291,183, No. 6,309,831, No. 6,416,949, No. 6,428,752, and No. 6,482,591 (each of which is incorporated herein by reference). Spotted microarrays can also be used in methods or systems according to embodiments of the present invention. An exemplary spotted microarray is the CodeLink® Array, available from Amersham Biosciences. Another useful microarray is manufactured using inkjet printing methods such as SurePrint® Technology, available from Agilent Technologies.
[0160] The systems and methods described herein can be used to detect the presence of specific target molecules in a sample contacted with a microarray. This can be determined, for example, based on the binding of a labeled target analyte to a specific probe of the microarray, or due to target-dependent modification of a specific probe to incorporate, remove, or alter the label at the probe position. Any one of the various assays can be used to identify or characterize targets using microarrays, for example, those described in U.S. Patent Application Publications 2003 / 0108867, 2003 / 0108900, 2003 / 0170684, 2003 / 0207295, or 2005 / 0181394 (each of which is incorporated herein by reference).
[0161] Exemplary labels that can be detected according to embodiments of the present invention include, for example, chromophores, luminescent phores, fluorophores, optically encoded nanoparticles, diffraction grating-encoded particles, and Ru(bpy) when present on a microarray. 32+This includes, but is not limited to, electrochemiluminescent labels or portions that can be detected based on optical properties. Useful fluorophores include, for example, fluorescent lanthanide complexes including those of Europium and Terbium, fluorescein, rhodamine, tetramethylrhodamine, eosin, erythrosine, coumarin, methyl-coumarin, pyrene, malacite green, Cy3, Cy5, stilbene, Lucifer Yellow, Cascade Blue (trademark), Texas Red, Alexa dye, phycoerythrin, body pea, and others known in the art, such as those described in Haugland, Molecular Probes Handbook, (Eugene, Oreg.) 6th Edition, The Synthegen catalog (Houston, Tex.), Lakowicz, Principles of Fluorescence Spectroscopy, 2nd Ed., Plenum Press New York (1999), or International Publication 98 / 59066 (each of which is incorporated herein by reference).
[0162] In certain embodiments, the optical system can be configured for, for example, time delay integration (TDI) in a line scanning embodiment, as described in U.S. Patent No. 7,329,860, the complete subject matter of which is incorporated herein by reference. For example, the optical assembly may have a 0.75 NA lens and a focusing accuracy of + / - 125 to 500 nm. The resolution can be 50 to 100 nm. The system may be capable of obtaining 1,000 to 10,000 unfiltered measurements per second.
[0163] While embodiments are illustrated with respect to the detection of samples containing biological or chemical substances supported by an optical substrate, it will be understood that other samples may be analyzed, examined, or imaged by embodiments described herein. Other exemplary samples include, but are not limited to, biological specimens such as cells or tissues, and electronic chips such as those used in computer processors. Some examples of applications include microscopy, satellite scanners, high-resolution reprography, fluorescence imaging, nucleic acid analysis and sequencing, DNA sequencing, synthetic sequencing, microarray imaging, and imaging of holographically encoded particles.
[0164] In other embodiments, the optical system may be configured to inspect an object in order to determine specific characteristics or structures of the object. For example, the optical system may be used to inspect the surface of an object (e.g., a semiconductor chip, a silicon wafer) in order to determine whether there are any deviations or defects on the surface.
[0165] Figure 10 illustrates a block diagram of an optical system 1000 formed according to one embodiment. As a mere example, the optical system 1000 may be a sampler imaging device for imaging a sample of interest for analysis. In other embodiments, the optical system 1000 may be a surface shape measuring device for determining the surface profile (e.g., topography) of an object. Furthermore, various other types of optical systems may use the mechanisms and systems described herein. In the illustrated embodiment, the optical system 1000 includes an optical assembly 1006, an object holder 1002 for supporting an object 1010 near the focal plane FP of the optical assembly 1006, and a stage controller 1015 configured to move the object holder 1002 laterally (along the X and / or Y axes extending within this page) or vertically / height along the Z axis. The optical system 1000 may also include a system controller or computing system 1020 operably coupled to the optical assembly 1006, the stage controller 1015, and / or the object holder 1002.
[0166] In certain embodiments, the optical system 1000 is a sample imaging device configured to image a sample. Although not shown, the sample imaging device may include other subsystems or devices for carrying out various assay protocols. As a mere example, the sample may include a flow cell having a flow channel. The sample imaging device may include a fluid control system, which includes a liquid reservoir fluidically coupled to the flow channel through a fluid network. The sample imaging device may also include a temperature control system, which may have a heater / cooler configured to regulate the temperature of the sample and / or the fluid flowing through the sample. The temperature control system may include sensors for detecting the temperature of the fluid.
[0167] As shown, the optical assembly 1006 is configured to direct input light to object 1010 and receive output light and direct it to one or more detectors. The output light may be input light that has been reflected and / or refracted by object 1010, and / or the output light may be light emitted from object 1010. To direct the input light, the optical assembly 1006 may include at least one reference light source 1012 and at least one excitation light source 1014 that direct light, such as a light beam having a predetermined wavelength, through one or more optical components of the optical assembly 1006. The optical assembly 1006 may include various optical components, including a conjugate lens 1018, for directing input light to object 1010 and output light to detectors.
[0168] In exemplary embodiments, the absolute wavelength light source 1012 may be used by a distance measuring system or focus control system (or focusing mechanism) of the optical system 1000, and the excitation light source 1014 may be used to excite biological or chemical substances in the object 1010 if the object 1010 contains a biological or chemical sample. The excitation light source 1014 may be arranged to illuminate the bottom surface of the object 1010 in TIRF imaging, etc., or to illuminate the top surface of the object 1010 in epifluorescence imaging, etc. As shown in Figure 10, the conjugate lens 1018 directs the input light to a focal region 1022 located in the focal plane FP. The lens 1018 has an optical axis 1024 and is measured along the optical axis 1024 to be positioned at a working distance WD1 from the object 1010. The stage controller 1015 may move the object 1010 in the Z direction to adjust the working distance WD1, for example, so that a portion of the object 1010 is located in the focal region 1022.
[0169] To determine whether object 1010 is in focus (i.e., sufficiently within the focal region 1022 or focal plane FP), the optical assembly 1006 is configured to direct at least one pair of light beams toward the focal region 1022 where object 1010 is approximately located. Object 1010 reflects the light beams. More specifically, the outer surface of object 1010 or an interface within object 1010 reflects the light beams. The reflected light beams then return to lens 1018 and propagate through it. As shown, each light beam has an optical path that includes a portion that has not yet been reflected by object 1010 and a portion that has been reflected by object 1010. The portions of the optical path before reflection are designated as incident light beams 1030A and 1032A and are indicated by arrows pointing toward object 1010. The portions of the optical path reflected by object 1010 are designated as reflected light beams 1030B and 1032B and are indicated by arrows pointing toward object 1010. For illustrative purposes, the light beams 1030A, 1030B, 1032A, and 1032B are shown having different optical paths within the lens 1018 and near the object 1010. However, in the exemplary embodiment, the light beams 1030A and 1032B are configured to propagate in opposite directions and have the same or substantially overlapping optical paths within the lens 1018 and near the object 1010, and the light beams 1030B and 1032A are configured to propagate in opposite directions and have the same or substantially overlapping optical paths within the lens 1018 and near the object 1010.
[0170] In the embodiment shown in Figure 10, the light beams 1030A, 1030B, 1032A, and 1032B pass through the same lens used for imaging. In an alternative embodiment, a light beam used for distance measurement or focusing may pass through a different lens not used for imaging. In this alternative embodiment, lens 1018 is dedicated to passing beams 1030A, 1030B, 1032A, and 1032B for distance measurement or focusing, and a separate lens (not shown) is used to image object 1010. Similarly, it will be understood that the systems and methods described herein for focusing and distance measurement may result from using a common objective lens shared with the imaging optical system, or alternatively, the objective lenses exemplified herein may be dedicated to focusing or distance measurement.
[0171] The reflected light beams 1030B and 1032B propagate through the lens 1018 and may optionally be further directed by other optical components of the optical assembly 1006. As shown, the reflected light beams 1030B and 1032B are detected by at least one focus detector 1044. In the illustrated embodiment, both reflected light beams 1030B and 1032B are detected by a single focus detector 1044. The reflected light beams may be used to determine the relative separation RS1. For example, the relative separation RS1 may be determined by the distance (i.e., separation distance) that separates the beam spots from the reflected light beams 1030B and 1032B incident on the focus detector 1044. The relative separation RS1 may be used to determine the degree of focus of the optical system 1000 with respect to the object 1010. However, in an alternative embodiment, each reflected light beam 1030B and 1032B may be detected by a separate corresponding focus detector 1044, and the relative separation RS1 may be determined based on the position of the beam spot on the corresponding focus detector 1044.
[0172] If object 1010 is not within a sufficient depth of field, the computing system 1020 may operate the stage controller 1015 to move object holder 1002 to a desired position. Alternatively to or in addition to moving object holder 1002, the optical assembly 1006 may be moved in the Z direction and / or along the XY plane.
[0173] For example, if object 1010 is located above the focal plane FP (or focal region 1022), object 1010 can be moved relative to the focal plane FP by a distance ΔZ1, and if object 1010 is located below the focal plane FP (or focal region 1022), object 1010 can be moved relative to the focal plane FP by a distance ΔZ2. In some embodiments, the optical system 1000 can replace lens 1018 with another lens 1018 or other optical component to move the focal region 1022 of the optical assembly 1006.
[0174] The embodiments described above and in Figure 10 are presented relating to a system for controlling focus or determining the degree of focus. The system is also useful for determining the working distance WD1 between an object 1010 and a lens 1018. In such embodiments, a focus detector 1044 can function as a working distance detector, and the distance at which a beam spot on the working distance detector is separated can be used to determine the working distance between the object 1010 and the lens 1018. For the sake of ease of description, various embodiments of the system and method are illustrated herein relating to controlling focus or determining the degree of focus. It will be understood that the system and method can also be used to determine the working distance between an object and a lens. Similarly, the system and method can also be used to determine the surface profile of an object.
[0175] In an exemplary embodiment, during operation, the excitation light source 1014 directs input light (not shown) towards an object 1010 to excite a fluorescently labeled biological or chemical substance. The label of the biological or chemical substance provides an optical signal 1040 (also called emission) having a predetermined wavelength. The optical signal 1040 is received by the lens 1018 and then directed to at least one object detector 1042 by other optical components of the optical assembly 1006. Although the illustrated embodiment shows only one object detector 1042, the object detector 1042 may comprise multiple detectors. For example, the object detector 1042 may include a first detector configured to detect one or more wavelengths of light and a second detector configured to detect one or more different wavelengths of light. The optical assembly 1006 may include lens / filter assemblies that direct different optical signals along different optical paths to the corresponding object detectors. Such optical systems are described in further detail in PCT application PCT / US 07 / 07991, filed on 30 March 2007, entitled "System and Devices for Sequence by Synthesis Analysis," and PCT application PCT / US2008 / 077850, filed on 26 September 2008, entitled "Fluorescence Excitation and Detection System and Method," the complete subject matter of both of which is incorporated herein by reference in whole.
[0176] The object detector 1042 communicates object data relating to the detected optical signal 1040 to the computing system 1020. The computing system 1020 may then record, process, analyze, and / or communicate the data to other users or computing systems, including remote computing systems, via a communication line (e.g., the Internet). In an embodiment, the object data may include imaging data that is processed to generate an image of the object 1010. The image may then be analyzed by a user of the computing system and / or the optical system 1000. In other embodiments, the object data may include not only emission from biological or chemical substances, but also light that has been reflected and / or refracted by an optical substrate or other component. For example, the optical signal 1040 may include light reflected by encoding particles such as the holographically encoded optical identification elements described above.
[0177] In some embodiments, a single detector may provide both of the functions described above with respect to the object detector 1042 and the focus detector 1044. For example, a single detector may detect reflected light beams 1030B and 1032B, as well as the optical signal 1040.
[0178] The optical system 1000 may include a user interface 1025 that interacts with the user through the computing system 1020. For example, the user interface 1025 may include a display (not shown) that shows and requests information from the user, and a user input device (not shown) for receiving user input.
[0179] The computing system 1020 may, among other things, include an object analysis module 1050 and a focus control module 1052. The focus control module 1052 is configured to receive focus data obtained by the focus detector 1044. The focus data may include signals representing beam spots incident on the focus detector 1044. The data may be processed to determine relative separation (e.g., separation distance between beam spots). The degree of focus of the optical system 1000 with respect to object 1010 may then be determined based on the relative separation. In certain embodiments, the working distance WD1 between object 1010 and lens 1018 may be determined. Similarly, the object analysis module 1050 may receive object data acquired by the object detector 1042. The object analysis module may process or analyze the object data to generate an image of the object.
[0180] Furthermore, the controller system 1020 may include any processor-based or microprocessor-based system, including systems using microcontrollers, reduced instruction set computers (RISC), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), logic circuits, and any other circuits or processors capable of performing the functions described herein. The above embodiments are illustrative and are therefore not intended to limit the definition and / or meaning of the term system controller. In exemplary implementations, the computing system 1020 executes a set of instructions stored in one or more storage elements, memories, or modules for at least one to acquire and analyze object data. The storage elements may be in the form of information sources or physical memory elements within the optical system 1000.
[0181] The set of instructions may include a variety of commands that instruct the optical system 1000 to perform a specific protocol. For example, the set of instructions may include a variety of commands for performing an assay to image an object 1010 or for determining the surface profile of an object 1010. The set of instructions may also be in the form of a software program. As used herein, the terms “software” and “firmware” are interchangeable and include any computer program stored in memory executed by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The memory types described above are illustrative and are therefore not limited to the types of memory available for storing computer programs.
[0182] As described above, the excitation light source 1014 generates excitation light directed toward object 1010. The excitation light source 1014 may generate one or more laser beams at one or more predetermined excitation wavelengths. The light may move in a raster pattern across parts of object 1010, such as groups of columns and rows of object 1010. Alternatively, the excitation light may illuminate one or more entire areas of object 1010 at once and stop continuously through the areas in a “step-and-shoot” scanning pattern. Line scanning can also be used, for example, as described in U.S. Patent No. 7,329,860, the complete subject matter of which is incorporated herein by reference in its entirety. Object 1010 generates an optical signal 1040, which may include light emitted in response to illumination of a label within object 1010, and / or light reflected or refracted by the optical substrate of object 1010. Alternatively, the optical signal 1040 can be generated without illumination, based entirely on the luminescence properties of the substance within the object 1010 (e.g., radioactive or chemiluminescent components within the object).
[0183] The object detector 1042 and the focus detector 1044 may be, for example, a photodiode or a camera. In some embodiments of this specification, the detectors 1042 and 1044 may comprise a camera having a 1-megapixel CCD-based optical imaging system, such as a 1002×1004 CCD camera having 8 gm pixels, which can image an area of 0.4×0.4 mm per tile at 20x magnification and using excitation light having an optional laser spot size of 0.5×0.5 mm (e.g., a square spot, or a circle or elliptical spot with a diameter of 0.5 mm). The camera may optionally have more or fewer than 1 million pixels, for example, a 4-megapixel camera may be used. In many embodiments, the camera's readout speed is desirable to be as fast as possible, for example, the transfer speed may be 10 MHz or more, for example, 20 or 30 MHz. More pixels generally mean that a larger area of the surface, and therefore more sequencing reactions or other optically detectable events, can be imaged simultaneously for a single exposure. In certain embodiments, a CCD camera / TIRF laser may collect approximately 6,400 images to examine 1,600 tiles (as described herein, images are optionally taken in four different colors per cycle using a combination of filters, dichroic and detectors). For a 1-megapixel CCD, a particular image may optionally contain approximately 5,000 to 50,000 randomly spaced unique nucleic acid clusters (i.e., images on the flow cell surface). With an imaging rate of 2 seconds per tile for four colors and a density of 25,000 clusters per tile, the system herein can optionally quantify approximately 45 million features per hour. With faster imaging rates and higher cluster densities, the imaging rate can be improved. For example, with a 20MHz camera readout speed and decomposed clusters of 20 pixels each, the readout can be 1 million clusters per second.The detector can be configured, for example, for time-delay integral (TDI) in a line scanning embodiment, as described in U.S. Patent No. 7,329,860, the complete subject matter of which is incorporated herein by reference. Other useful detectors include, but are not limited to, optical quadrant photodiode detectors, such as those available from Pacific Silicon Sensor (Westlake Village, Calif.) that have a 2x2 array of individual photodiode active regions fabricated on a single chip, or position-sensitive detectors, such as those available from Hamamatsu Photonics, KK (Hamamatsu City, Japan) that have a monolithic PIN photodiode with uniform resistance in one or two dimensions.
[0184] Figure 11 is a perspective view of a sample imaging device 1100 formed according to one embodiment. As shown, the sample imaging device 1100 includes an imaging device base 1102 supporting a stage 1104 having a sample holder 1106 thereon. The sample holder 1106 is configured to support one or more optical substrates 1108 during an imaging session. The optical substrates 1108 are illustrated as flow cells in Figure 11. However, other samples may be used.
[0185] The sample imaging device 1100 also includes a housing 1110 (illustrated by dashed lines) and a support column 1112 that supports the housing 1110. The housing 1110 can enclose at least a portion of the optical assembly 1114. The optical assembly 1114 may include a focus assembly 1116 and a sample detection assembly 1130. For example, the focus assembly 1116 may include an autofocus line scanning camera that receives a reflected light beam to determine the degree of focus of the sample imaging device 1100. The sample imaging device 1100 may also include a filter wheel 1122 and an alignment mirror 1124 that directs light toward the sample detector 1132, shown as the K4 camera in Figure 11.
[0186] Figure 12 illustrates an implementation of a sequencing system 1210 configured to process molecular samples that can be sequenced to determine their components, component order, and overall structure. The system includes an instrument 1212 that receives and processes biological samples. A sample source 1214 provides a sample 1216, often containing tissue samples. The sample source may include individuals or subjects such as humans, animals, microorganisms, plants, or other donors (including environmental samples), or any other subject containing organic molecules of the subject to be sequenced. Of course, the system can be used with samples other than those taken from living organisms, including synthesized molecules. Often, molecules include DNA, RNA, or other base-paired molecules, and their sequences may define genes and variants with specific functions of the final subject.
[0187] Sample 1216 is introduced into a sample / library preparation system 1218. This system can isolate, disrupt, and otherwise prepare the sample for analysis. The resulting library contains the target molecule at a length that facilitates sequencing. The resulting library is then supplied to instrument 1212, where sequencing is performed. In practice, the library, which may also be called a template, is combined with reagents in an automated or semi-automated process and then introduced into a flow cell before sequencing.
[0188] In the implementation shown in Figure 12, the instrument includes a flow cell or array 1220 that receives a sample library. The flow cell includes one or more fluid channels that enable sequencing chemical phenomena to occur, including the attachment of molecules from the library and amplification at locations or sites that can be detected during the sequencing operation. For example, the flow cell / array 1220 may include a sequencing template immobilized on one or more surfaces at locations or sites. The “flow cell” may include patterned arrays such as microarrays and nanoarrays. In practice, the locations or sites may be arranged on one or more surfaces of a support in a regular repeating pattern, a complex non-repeating pattern, or a random arrangement. To enable sequencing chemical phenomena to occur, the flow cell also allows for the introduction of substances, such as various reagents, buffers, and other reaction media used for reactions, flushing, etc. The substances flow through the flow cell and may come into contact with the target molecules at individual sites.
[0189] In this apparatus, the flow cell 1220 is mounted on a movable stage 1222 that can move in one or more directions, as indicated by reference no. 1224 in this implementation. The flow cell 1220 may be provided in the form of a removable and replaceable cartridge that can interface with ports on the movable stage 1222 or other components of the system to allow, for example, reagents and other fluids to be delivered to or from the flow cell 1220. The stage is associated with an optical detection system 1226 that can direct radiation or light 1228 onto the flow cell during sequencing. The optical detection system may employ various methods, such as fluorescence microscopy, for the detection of analytes located at sites in the flow cell. In a non-limiting embodiment, the optical detection system 1226 may employ confocal line scanning to generate progressive pixelated image data that can be analyzed to locate individual sites in the flow cell and to determine the type of nucleotide immediately attached to or bound to each site. Other imaging techniques may also be appropriately employed, such as techniques in which one or more radiating points are scanned along the sample, or techniques employing "step-and-shoot" imaging techniques. The optical detection system 1226 and the stage 1222 may cooperate to maintain the flow cell and detection system in a static relationship while acquiring a regional image, or, as described above, the flow cell may be scanned in any appropriate mode (e.g., point scanning, line scanning, "step-and-shoot" scanning).
[0190] While many different techniques can be used for imaging, or more generally, for detecting molecules at a site, the currently envisioned implementation can utilize confocal optical imaging at wavelengths that cause excitation of a fluorescent tag. The tag is excited by its absorption spectrum and returns a fluorescent signal by its emission spectrum. The optical detection system 1226 is configured to process pixelated image data at a resolution that allows for analysis of the signal emission site and to capture such signals for processing and storing the resulting image data (or data derived therefrom).
[0191] In sequencing operations, the cycle operation or process is implemented in an automated or semi-automated manner, where the reaction is facilitated, for example, using a single nucleotide or oligonucleotide, followed by flushing, imaging, and deblocking in preparation for subsequent cycles. A sample library prepared for sequencing and immobilized on a flow cell can undergo numerous such cycles before all useful information is extracted from the library. The optical detection system 1226 can generate image data from scanning the flow cell (and its sites) during each cycle of the sequencing operation by using an electronic detection circuit (e.g., a camera or imaging electronic circuit or chip). The obtained image data can then be analyzed to identify the location of individual sites in the image data, and molecules present at those sites can be analyzed and characterized, for example, by referring to a specific color or wavelength of light detected at a particular location (a characteristic emission spectrum of a particular fluorescent tag), as indicated by a group or cluster of pixels in the image data at that location. In DNA or RNA sequencing applications, for example, four common nucleotides may be represented by identifiable fluorescence emission spectra (wavelengths or wavelength ranges of light). Each emission spectrum can then be assigned a value corresponding to its nucleotide. Based on this analysis and by tracking the cyclic values determined for each site, individual nucleotides and their order can be determined for each site. These sequences can then be further processed to assemble longer segments, including genes, chromosomes, etc. As used in this disclosure, the terms “automated” and “semi-automated” mean that once the operation is initiated, or once a process involving the operation is initiated, the operation is performed by system programming or configuration with little or no human interaction.
[0192] In the illustrated configuration, reagent 1230 is drawn into or aspirated into the flow cell via valve 1232. The valve may access the reagent from the receiver or container where it is stored, via a pipette or sipper (not shown in Figure 12), etc. Valve 1232 may allow for reagent selection based on a defined sequence of operations to be performed. The valve may also receive commands to direct the reagent through channel 1234 to the flow cell 1220. An outlet or effluent channel 1236 directs the used reagent from the flow cell. In the illustrated configuration, pump 1238 moves the reagent through the system. The pump may also perform other useful functions, such as measuring reagent or other fluids through the system, or aspirating air or other fluids. An additional valve 1240 downstream of pump 1238 allows for proper directing of the used reagent to a disposal container or receiver 1242.
[0193] The device further includes various circuits that assist in commanding the operation of various system components, monitoring their operation through feedback from sensors, collecting image data, and processing the image data at least partially. In the implementation shown in Figure 12, the control / monitoring system 1244 includes a control system 1246 and a data acquisition and analysis system 1248. Both systems would include one or more processors (e.g., digital processing circuits such as microprocessors, multicore processors, FPGAs, or any other suitable processing circuits) and associated memory circuits 1250 (e.g., solid memory devices, dynamic memory devices, on and / or off-board memory devices, etc.) that can store machine-executable instructions for, for example, controlling one or more computers, processors, or other similar logic devices in order to provide specific functionality. Application-specific or general-purpose computers may constitute at least partially the control system and the data acquisition and analysis system. The control system may include (e.g., programmed) circuits configured to process commands for, for example, fluid dynamics, optics, stage control, and any other useful functions of the device. The data acquisition and analysis system 1248 interfaces with the optical detection system to command the movement of the optical detection system or the stage or both, the emission of light for periodic detection, the reception and processing of return signals, etc. The instrument may also include various interfaces, such as those shown in reference no. 1252, including an operator interface that allows control and monitoring of the instrument, movement of samples, initiation of automated or semi-automated sequencing operations, generation of reports, etc. Finally, in the implementation configuration of Figure 12, an external network or system 1254 may be coupled to and cooperate with the instrument for, for example, analysis, control, monitoring, maintenance, and other operations.
[0194] Figure 12 illustrates a single flow cell and fluid path, as well as a single optical detection system 1226; however, it should be noted that some instruments may accommodate two or more flow cells and fluid paths. For example, in the currently envisioned implementation, two such arrangements are provided to improve sequencing and throughput. In practice, any number of flow cells and fluid paths may be provided. These may utilize the same or different reagent receptacles, disposal receptacles, control systems, image analysis systems, etc. If provided, multiple fluid systems may be controlled individually or in a coordinated manner.
Claims
1. It is a system, A field-programmable gate array (FPGA) locally housed on a shared network server having a container orchestration engine and a variant analysis model, When executed by the FPGA, the system will Variant analysis is performed using the variant analysis model described above in order to generate nucleotide base calls for the sample nucleotide sequence, The diagnostic analysis is performed on the nucleotide base call to detect the genetic state of the sample nucleotide sequence by utilizing the FPGA to access the nucleotide base call in the shared network server and to sequentially execute processes for the diagnostic workflow container as scheduled by the container orchestration engine, wherein the diagnostic workflow container constitutes and executes an external sequencing diagnostic workflow from an external server without accessing the nucleotide base call in the shared network server. To prevent the exposure of the nucleotide base call to the external server, one or more workflow containers associated with each function of the external sequencing diagnostic workflow are isolated using targeted security authorizations. A system comprising: executing the external sequencing diagnostic workflow by utilizing the container orchestration engine to allocate computing resources of the FPGA to process one or more sequentially scheduled workflow containers; and a non-temporary computer-readable medium containing instructions to perform the above.
2. The system according to claim 1, further comprising instructions, when executed by the FPGA, causing the system to implement the external sequencing diagnostic workflow identified from an external application hosted on an external server separate from the shared network server of the variant analysis model and the container orchestration engine.
3. When executed by the FPGA, the system will Determining the diagnostic execution mode corresponding to the standardized genetic diagnostic protocol, The system according to claim 1 or 2, further comprising an instruction to a client device to grant it access only to diagnostic applications compatible with the diagnostic execution mode.
4. The system according to claim 1, further comprising instructions, when executed by the FPGA, causing the system to prevent the external sequencing diagnostic workflow from accessing the sequencing data associated with the nucleotide base call for the sample nucleotide sequence by applying read-only permissions to the workflow container used by the external sequencing diagnostic workflow.
5. The system according to claim 1, further comprising instructions, when executed by the FPGA, causing the system to selectively prevent one or more workflow containers from accessing sequencing data associated with the nucleotide base call for the sample nucleotide sequence while the external sequencing diagnostic workflow is being executed.
6. The system according to claim 1, further comprising, when executed by the FPGA, instructions causing the system to isolate the one or more workflow containers by utilizing the targeted security permissions in order to individually specify the sequence determination data access permissions for the one or more workflow containers.
7. The system according to claim 1, further comprising, when executed by the FPGA, instructions causing the system to receive the external sequencing diagnostic workflow generated by an external device operated by an external entity.
8. A computer implementation method, To generate nucleotide base calls for the sample nucleotide sequence, variant analysis is performed using a variant analysis model, and The diagnostic analysis is performed on the nucleotide base call to detect the genetic state of the sample nucleotide sequence by utilizing an FPGA locally housed on the shared network server having the container orchestration engine and the variant analysis model, in order to access the nucleotide base call in the shared network server and to sequentially execute processes for the diagnostic workflow container as scheduled by the container orchestration engine, wherein the diagnostic workflow container constitutes and executes an external sequencing diagnostic workflow from an external server without accessing the nucleotide base call in the shared network server. To prevent the exposure of the nucleotide base call to the external server, one or more workflow containers associated with each function of the external sequencing diagnostic workflow are isolated using targeted security authorization, A computer implementation method comprising executing the external sequencing diagnostic workflow by utilizing the container orchestration engine to allocate computing resources of the FPGA to process one or more sequentially scheduled workflow containers.
9. The computer implementation method according to claim 8, wherein the isolation of one or more workflow containers includes controlling access to different workflow data sources for one or more workflow containers in order to prevent access to sequencing data by the external sequencing diagnostic workflow.
10. The computer implementation method according to claim 8 or 9, further comprising receiving a label index defining the version of the variant analysis model and a memory allocation used to perform the external sequencing diagnostic workflow.
11. Specifying multiple workflow data sources that store different types of workflow data, The computer implementation method according to claim 8, further comprising activating access to a first workflow data source among the plurality of workflow data sources for one workflow container from among the one or more workflow containers, while preventing access to other workflow data sources among the plurality of workflow data sources.
12. The computer implementation method according to claim 8, further comprising mounting a plurality of workflow data sources as read-only for one or more workflow containers.
13. The computer implementation method according to claim 8, further comprising accessing the sequencing data associated with the nucleotide base call for the sample nucleotide sequence and preventing the external sequencing diagnostic workflow from accessing the sequencing data associated with the nucleotide base call for the sample nucleotide sequence by applying read-only permissions to the workflow container used by the external sequencing diagnostic workflow.
14. The computer implementation method according to claim 8, further comprising performing the external sequencing diagnostic workflow during variant analysis using the variant analysis model.
15. A non-temporary computer-readable medium, when executed by an FPGA, to a computing device, To generate nucleotide base calls for the sample nucleotide sequence, variant analysis is performed using a variant analysis model, and The diagnostic analysis is performed on the nucleotide base call to detect the genetic state of the sample nucleotide sequence by utilizing the FPGA locally housed on the shared network server, which has the container orchestration engine and the variant analysis model, in order to access the nucleotide base call in the shared network server and to sequentially execute processes for the diagnostic workflow container as scheduled by the container orchestration engine, wherein the diagnostic workflow container constitutes and executes an external sequencing diagnostic workflow from an external server without accessing the nucleotide base call in the shared network server. To prevent the exposure of the nucleotide base call to the external server, one or more workflow containers associated with each function of the external sequencing diagnostic workflow are isolated using targeted security authorizations. A non-temporary computer-readable medium including instructions to perform the external sequencing diagnostic workflow by utilizing the container orchestration engine to allocate computing resources of the FPGA to process one or more sequentially scheduled workflow containers.
16. The non-temporary computer-readable medium according to claim 15, further comprising instructions, when executed by the FPGA, causing the computing device to isolate one or more workflow containers that satisfy one or more standardized genetic diagnostic protocols while also executing the external sequencing diagnostic workflow, by encoding a workflow execution application that grants the external sequencing diagnostic workflow read-only access to sequencing data associated with the nucleotide base call of the sample nucleotide sequence during the execution of the external sequencing diagnostic workflow.
17. The non-temporary computer-readable medium according to claim 15 or 16, further comprising instructions, when executed by the FPGA, causing the computing device to execute the external sequencing diagnostic workflow generated by an external system on a server separate from the shared network server of the container orchestration engine and the variant analysis model.
18. The non-temporary computer-readable medium according to claim 15, further comprising instructions, when executed by the FPGA, causing the computing device to isolate the one or more workflow containers by preventing one of the workflow containers from accessing the sequencing data of the sample nucleotide sequence.
19. The non-temporary computer-readable medium according to claim 18, further comprising instructions, when executed by the FPGA, causing the computing device to prevent the workflow container from accessing the sequence determination data by preventing the workflow container from accessing one or more workflow data sources, including an input directory, an output directory, and an application directory.
20. The non-temporary computer-readable medium according to claim 15, further comprising instructions that, when executed by the FPGA, cause the computing device to trigger the execution of the external sequencing diagnostic workflow by receiving post-defined parameters for implementing the external sequencing diagnostic workflow via the container orchestration engine.