Method for controlling sequencing devices
A modular software platform for sequencing devices addresses assay availability and cost issues by enabling cloud-based delivery and local server integration, facilitating efficient and cost-effective sequencing assays for research and diagnostics.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LIFE TECHNOLOGIES CORP
- Filing Date
- 2020-07-31
- Publication Date
- 2026-05-12
AI Technical Summary
The use of sequencing technologies is limited by assay availability, sequencing execution time, and cost, particularly in high-quality sequencing processes, which are historically expensive and hinder widespread application in biology and medicine.
A modular software platform for sequencing devices that supports rapid assay development and deployment through cloud-based delivery of assay definition files, enabling multiple functional modes for research and diagnostic applications, and integrating with local server systems for data analysis and workflow management.
Facilitates efficient and cost-effective implementation of sequencing assays in laboratories, supporting both research and diagnostic uses with rapid adoption and streamlined assay configurations, enhancing laboratory molecular pathology testing capabilities.
Smart Images

Figure 0007857215000002 
Figure 0007857215000003 
Figure 0007857215000004
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 62 / 889,109, filed Aug. 20, 2019, and U.S. Provisional Application No. 62 / 704,806, filed May 29, 2020. The entire contents of the foregoing applications are incorporated herein by reference.
[0002] The present disclosure relates to the control of a sequencing device for next - generation sequencing (NGS), including digital delivery of modular software components that include assay workflows from cloud - based computing and storage system resources.
Background Art
[0003] Biological and medical academic research is increasingly focusing on nucleic acid sequencing to enhance biological research and medicine. For example, biologists and zoologists are focusing on sequencing to study animal movement, species evolution, and the origin of traits. The medical community uses sequencing to study the origin of diseases, drug sensitivity, and the origin of infections. Therefore, sequencing has broad applicability in substantially all aspects of biology, treatment, diagnosis, forensics, and academic research.
[0004] Nevertheless, the use of sequencing can be limited by assay availability, sequencing execution time, preparation time, and cost. Additionally, high - quality sequencing has historically been an expensive process, thus limiting its practice.
Summary of the Invention
[0005] Laboratory molecular pathology testing can be enhanced by a wide selection of assays developed for sequencing devices, such as NGS sequencing devices. Server systems configured to control sequencing devices can implement a modular software platform supporting rapid expansion of molecular testing menus, enabling rapid laboratory adoption of assays. Assay content and corresponding workflows can be delivered as modular software components from cloud-based computing and storage system resources (e.g., Thermo Fisher Cloud, Thermo Fisher Scientific, Waltham, Massachusetts). Assay configurations and corresponding workflows are delivered to the user's server system as modular software components in an assay definition file (ADF). The assay definition file supports backward compatibility of workflow software modules and separation of workflow software modules from the server system's platform software.
[0006] The server system and modular software components can be configured to control multiple functional modes, including Research Use Only (RUO) or Assay Development (AD) mode, and In Vitro Diagnostic (IVD) or Dx mode. The RUO or AD mode supports assay development and digital distribution for research applications and third-party development of assays (RUO and AD are used interchangeably). The IVD or Dx mode supports digital distribution of molecular diagnostic assays that meet local requirements for diagnostic applications (IVD and Dx are used interchangeably). Multiple functional modes allow the same NGS sequencing device to be used for both RUO and IVD assays.
[0007] According to an exemplary embodiment, a method is provided which includes: receiving an assay definition file from a server of a cloud computing and storage system, wherein the assay definition file includes code modules for configuring the assay; storing the code modules in the memory of the local server system; receiving sequencing data from a sequencing device, wherein the local server system is generating the sequencing data during sequencing execution for the assay; and applying an analysis pipeline for the assay to the sequencing data, wherein the analysis pipeline includes analysis steps performed by the processor of the local server system according to the code modules from the assay definition file to generate assay analysis results.
[0008] According to an exemplary embodiment, a local server system is provided, comprising: a memory; and a processor, which, when executed by the processor, causes the local server system to perform a method comprising: receiving an assay definition file from a server of a cloud computing and storage system, wherein the assay definition file includes code modules for structuring an assay; storing the code modules in the memory of the local server system; receiving sequencing data from a sequencing device, wherein the sequencing data is generated by the sequencing device during sequencing execution for an assay; and applying an analysis pipeline for an assay to the sequencing data, wherein the analysis pipeline includes analysis steps performed by the processor of the local server system according to the code modules from the assay definition file to generate assay analysis results. [Brief explanation of the drawing]
[0009] Novel features are specifically revealed in the attached claims. A better understanding of the features and advantages will be obtained by referring to the following detailed description illustrating exemplary embodiments and attached drawings.
[0010] [Figure 1] A schematic diagram of the server system components according to the embodiment is shown. [Figure 2] This is a block diagram of the analysis pipeline according to the embodiment. [Figure 3] This is a schematic diagram illustrating the generation of an assay definition file according to the embodiment. [Figure 4] This is a schematic diagram of an example of an assay definition file package. [Figure 5] This is a diagram illustrating an example of a sequencing device. [Figure 6] Figure 5 shows an example of the instrument deck of a sequencing determination device. [Figure 7] This diagram illustrates an example of a sequencing machine workflow. [Figure 8] This is a diagram of an example of a sequencing determination chip. [Figure 9] This is a block diagram illustrating an example of processing sequence determination data from multiple lanes of a sequence determination chip. [Modes for carrying out the invention]
[0011] As used herein, DNA (deoxyribonucleic acid) may be referred to as a chain of nucleotides consisting of four types of nucleotides: A (adenine), T (thymine), C (cytosine), and G (guanine), while RNA (ribonucleic acid) contains four types of nucleotides: A, U (uracil), G, and C. Certain pairs of nucleotides bind specifically to each other in a complementary manner (called complementary base pairs). That is, adenine (A) pairs with thymine (T) (however, in the case of RNA, adenine (A) pairs with uracil (U)), and cytosine (C) pairs with guanine (G). When a first nucleic acid chain binds to a second nucleic acid chain consisting of nucleotides complementary to the nucleotides of the first chain, the two chains join to form a double helix. In various embodiments, “nucleic acid sequencing data,” “nucleic acid sequencing information,” “nucleic acid sequence,” “genome sequence,” “gene sequence,” or “fragment sequence,” “nucleic acid sequence read,” or “nucleic acid sequencing read” refers to any information or data indicating the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine / uracil) in a DNA or RNA molecule (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, fragment, etc.). It should be understood that this instruction intends sequence information to be obtained using all types of available technologies, platforms, or techniques, including but not limited to capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection systems, and digital signature-based systems.
[0012] "Polynucleotide," "nucleic acid," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogues) linked by nucleoside bonds. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides are usually in size ranging from a few monomer units, e.g., 3-4, to several hundred monomer units. Whenever a polynucleotide, such as an oligonucleotide, is represented by a letter sequence such as "ATGCCTG," it will be understood that, unless otherwise indicated, the nucleotides are in a 5'-3' order from left to right, with "A" representing deoxyadenosine, "C" representing deoxycytidine, "G" representing deoxyguanosine, and "T" representing thymidine. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide containing a base, as is standard in the art.
[0013] The terms “variant,” “genomic variant,” or “genome variant” refer to a single or group of sequences (in DNA or RNA) that have been altered due to mutation, recombination / crossover, or genetic drift, so as to be referenced for a particular species or subpopulation within a particular species. Examples of types of genomic variants include, but are not limited to, single nucleotide polymorphisms (SNPs), copy number variations (CNVs), insertions / deletions (indels), single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), and inversions.
[0014] In various embodiments, genomic variants can be detected using nucleic acid sequencing systems and / or analysis of sequencing data. A sequencing workflow may begin with the test sample being sheared or digested into hundreds, thousands, or millions of small fragments sequenced by a nucleic acid sequencing instrument to provide hundreds, thousands, or millions of sequence reads, such as nucleic acid sequence reads. Each read can then be mapped to a reference or target genome, and in the case of mate-pair fragments, the reads can be paired, thereby enabling querying of repeating regions of the genome. The results of mapping and pairing can be used as input for various independent or integrated genomic variant (e.g., SNP, CNV, Indel, inversion, etc.) analysis tools.
[0015] The phrase "sample genome" can refer to the whole or partial genome of an organism.
[0016] As used herein, “targeting panel” refers to a set of target-specific primers designed for the selective amplification of a target gene sequence in a sample. In some embodiments, following the selective amplification of at least one target sequence, the workflow further comprises nucleic acid sequencing of the amplified target sequence.
[0017] As used herein, “target sequence” or “target gene sequence” and its derivatives refer to any single- or double-stranded nucleic acid sequence that can be amplified or synthesized in accordance with this disclosure, including any nucleic acid sequence suspected or expected to be present in a sample. In some embodiments, the target sequence includes at least a portion of a particular nucleotide sequence to be amplified or synthesized, or its complement, which exists in a double-stranded form and is present before the addition of a target-specific primer or an attached adapter. The target sequence may include nucleic acids that can be hybridized to primers useful in the amplification or synthesis reaction before extension by polymerase. In some embodiments, the term refers to a nucleic acid sequence in which the sequence identity, ordering, or location of nucleotides is determined by one or more of the methods of this disclosure.
[0018] As used herein, “target-specific primer” and its derivatives refer to single-stranded or double-stranded polynucleotides, typically oligonucleotides, that contain at least one sequence that is at least 50% complementary, typically at least 75% complementary, or at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98%, or at least 99% complementary, or identical to at least a portion of a nucleic acid molecule containing a target sequence. In such cases, the target-specific primer and the target sequence are described as “corresponding” to each other. In some embodiments, a target-specific primer can hybridize to at least a portion of its corresponding target sequence (or a complement of the target sequence), and such hybridization can optionally be carried out under standard hybridization conditions or strict hybridization conditions. In some embodiments, a target-specific primer cannot hybridize to the target sequence or its complement, but can hybridize to a portion of the nucleic acid chain containing the target sequence or its complement. In some embodiments, forward target-specific primers and reverse target-specific primers define a pair of target-specific primers that can be used to amplify a target sequence via template-dependent primer extension. Typically, each primer in a pair of target-specific primers contains at least one sequence that is substantially complementary to at least a portion of a nucleic acid molecule containing the corresponding target sequence, but less than 50% complementary to at least one other target sequence in the sample. In some embodiments, amplification can be performed using multiple pairs of target-specific primers in a single amplification reaction, where each primer pair comprises a forward target-specific primer and a reverse target-specific primer, each containing at least one sequence that is substantially complementary to or substantially identical to the corresponding target sequence in the sample, and each primer pair has a different corresponding target sequence. In various embodiments, the target nucleic acid produced by the amplification of multiple target-specific sequences from a population of nucleic acid molecules can be sequenced.In some embodiments, amplification may include hybridizing one or more target-specific primer pairs to a target sequence, extending a first primer of the primer pair, denaturing the extended first primer product from a population of nucleic acid molecules, hybridizing a second primer of the primer pair to the extended first primer product, extending the second primer to form a double-stranded product, and digesting the target-specific primer pair away from the double-stranded product to generate a plurality of amplified target sequences. In some embodiments, the amplified target sequences may be ligated to one or more adapters. In some embodiments, the adapters may include one or more nucleotide barcode or tagging sequences. In some embodiments, once ligated to an adapter, the amplified target sequences may undergo a nick translation reaction and / or further amplification to generate a library of adapter-ligated amplified target sequences.
[0019] In various embodiments, a method for performing multiplex PCR amplification includes contacting a population of target sequences with a plurality of target-specific primer pairs having forward and reverse primers to form a plurality of template / primer duplexes; adding a mixture of DNA polymerase and dNTPs to the plurality of template / primer duplexes for a sufficient amount of time and temperature to extend one or both of the forward and reverse primers of each target-specific primer pair via template-dependent synthesis, thereby generating a plurality of extended primer products / template duplexes; denaturing the extended primer products / template duplexes; annealing complementary primers from the target-specific primer pairs to the extended primer products; and extending the annealed primers in the presence of DNA polymerase and dNTPs to form a plurality of target-specific double-stranded nucleic acid molecules.
[0020] As used herein, the term "templating" refers to a process of generating two or more, or a plurality or population of substantially identical polynucleotides that can be used as templates in nucleic acid analysis methods, including, for example, nucleic acid sequencing such as synthetic sequencing, or a process of generating a substantially monoclonal population of nucleic acids. The polynucleotides generated by the templating process are typically referred to as nucleic acid templates.
[0021] In some embodiments, the nucleic acid sequencing instrument can be interfaced with a server system for control of various components of the sequencing instrument and processing of data output from sequencing runs on the sequencing instrument. Server system software can include web applications, databases, and analysis pipelines and can support connections from the sequencing instrument (FIG. 5). Server system software can provide the following major functions and application program interfaces (APIs). 1. APIs for user authentication and reagent tracking perform information execution and tracking / logging. Supported instruments can include sequencing instruments and extraction instruments. 2. APIs for a LIMS (Laboratory Information Management System) for sample, library, and run planning perform execution and draw out the execution status of the plan. 3. Support the management of sample and run data. 4. Support the execution of assay configurations and analysis pipelines for data analysis and report generation. 5. Interface with a software update server for software updates and maintenance. 6. Support configuration for connection to annotation and reporting systems such as Thermo Fisher Scientific's Ion Reporter deployed in a cloud-based system or a local system, and establish a secure authenticated connection with the cloud-based system for transferring mapped or unmapped BAM files. 7. Support configuration for connecting to resource systems in cloud computing environments such as Thermo Fisher Cloud, establish secure and authenticated connections with cloud resource systems to download software and system content, and transmit telemetry data.
[0022] Figure 1 shows a schematic diagram of the server system components. In some embodiments, the basic software architecture may include a web interface, a remote monitoring agent, a database, an API to instruments, an analysis pipeline, containerization of the analysis pipeline (e.g., using Docker), connectivity to an annotation and reporting system (e.g., Ion Reporter from Thermo Fisher Scientific), and a cloud-based support and resource system (e.g., Thermo Fisher Cloud). The cloud-based support and resource system, or cloud-based resource system, may be implemented in a cloud computing and storage system. The cloud-based support and resource system stores content including assay definition files. A server in the cloud computing and storage system may download content such as assay definition files to a local server system. The cloud-based support and resource system may receive telemetry data from the local server system. The server system, local server system, and user server system are used interchangeably as described herein.
[0023] In some embodiments, the user interface (UI) may be implemented via web application software. The UI may provide a sample management page. The sample management UI page allows the user to enter sample information into the system. Sample information includes a unique sample identifier (ID), sample name, and sample preparation reagent tracking information. Validation logic is built into the sample management flow, locking the sample preparation steps into a predefined assay workflow. The UI may provide an assay management page. The assay management UI page allows the user to view and create assays. The assay locks the workflow into predefined parameters for each step of the process. Validation logic may be built to ensure assay configuration. The UI may provide a run plan and monitoring page. The run plan and monitoring UI page allows the user to plan runs and monitor ongoing runs. The UI may provide an output data page. The output data UI page allows the user to view analysis results along with quality control (QC) metric evaluations, logs, and audit trails of the generated results. The UI may provide a configuration page. The configuration UI page allows the user to view and configure the system.
[0024] In some embodiments, the application programming interface (API) may be provided through the Java® platform. For example, the Java platform may include a Tomcat server, which can be used to build Web ARchive (WAR) files for web-based applications.
[0025] Code modules for various steps in an analysis pipeline can be referred to as actors in the context of the Kepler workflow engine. For example, a code module for an analysis step may be implemented by the binary code of a Java program contained in an actor jar. The Kepler workflow engine defines the processing components of a workflow as "actors" and chains together steps for execution by the algorithm or processor of the analysis pipeline (https: / / kepler-project.org). For example, the Kepler workflow engine can be used to construct the workflow of the analysis pipeline in Figure 1.
[0026] A server system may include one or more databases. For example, a server system may include a relational database for storing sample data, execution data, and system / user configuration. The relational database may include two separate databases: an assay development database and a Dx database. The assay development database may store sample data, execution data, and system / user configuration for the RUO or assay development mode of the operation. The Dx database may store sample data, execution data, and system / user configuration for the IVD or Dx mode of the operation.
[0027] The server system may include an annotation database, AnnotationDB, for storing annotation source data. For example, the annotation database may be implemented as a NoSQL or non-relational database, such as MongoDB. Each annotation source may be stored as a JSON (JavaScript Object Notation) string containing metadata indicating the source name and version. Each annotation source may contain a list of annotations keyed by an annotation ID. The server system may also include a variome database, VariomeDB, for storing variant information. For example, the variome database may be implemented as a NoSQL or non-relational database, such as MongoDB. VariomeDB may store a collection of variant call results on a particular sample. For example, records in JSON format may contain metadata to identify the sample.
[0028] For example, the AnnotationDB database may store one or more of the following annotation sources: 1. RefGene model: hg19_refgene_63, version 63 2. RefGene function regular transcript score: hg19_refgeneScores_4, version 4 3. dbSNP:dbsnp_138, version 138 4. Canonical RefSeq transcript: hg19_refgene_63, version 63 5.5000Exomes:hg_esp6500_1, Version 1 6. ClinVar:clinvar_1, version 1 7. DGV: dgv_20130723, version 20130723 8. OMIM:omim_03022014, version 03022014
[0029] Other annotation sources may be included. Other versions of the annotation sources mentioned above may be included. Annotation sources may provide publicly available annotation information or proprietary annotation information.
[0030] For each call in the Variome database, each annotation source may be queried for annotations that match the variant, and matching annotations may be stored in the Variome database as key-value pairs along with the variant. The annotated variant may be contained in a results file for the user, for example, an annotated VCF file. A VCF file is a tab-delimited text file used to store variants of gene sequences. In some embodiments, the annotation method for use in this teaching may include one or more features described in U.S. Patent Application Publication No. 2016 / 0026753, published on January 28, 2016, which is incorporated in whole herein.
[0031] In some embodiments, the server system may include an analysis pipeline to process sequencing data generated during sequencing runs for assays performed by a sequencing instrument. The sequencing instrument transfers sequencing data files and experiment log files to the server system memory, for example, raw .dat files, processed .dat files that produce blockwise 1.wells files, and thumbnail data. The analysis pipeline accesses the data files from memory and initiates data analysis for execution.
[0032] In some embodiments, Docker containers and Docker images can be used to package analysis pipelines and operating system-specific binaries. Docker is a tool used to create, deploy, and run applications by using containers. Containers allow applications with all the necessary parts, including libraries and other dependencies, to be bundled as a single package. This allows application software to use the same Linux® kernel as the host system. Docker image files can be packaged together with the libraries and binaries required by the analysis pipeline code. Docker can be used to adapt applications or algorithms to new or different versions of the operating system (OS) and to create Docker images of applications that are compatible with the OS version.
[0033] In some embodiments, the server system may include a crawler service for data transfer from the sequencing instrument to the analysis pipeline. The crawler is an event-based service that can be developed using the Java NIO watcher API (Application Programming Interface). NIO (Non-Blocking I / O) is a collection of Java programming language APIs that provide functionality for intensive input / output (I / O) operations. The crawler may monitor an FTP directory configured for the sequencing instrument to transfer execution data from the sequencing instrument to the analysis pipeline.
[0034] Figure 2 is a block diagram of an analysis pipeline according to an embodiment. The sequencing instrument generates raw data files (DAT or .dat files) during the sequencing run for the assay. Signal processing may be applied to the raw data to generate embedded signal measurement data for files such as 1.wells files, which are transferred to an FTP location on the server along with log information of the run. The signal processing step may derive background signals corresponding to the wells. The background signals may be subtracted from the measurement signals for the corresponding wells. The remaining signals may be fitted by an incorporation signal model to estimate incorporation in each nucleotide flow for each well. The output from the above signal processing is the signal measurements per well and per flow, which may be stored in files such as 1.wells files.
[0035] In some embodiments, the base calling step may perform phase estimation, normalization, identify the best partial sequence fit, and perform base calling using a solver algorithm. The base sequences of the sequence reads are stored in an unmapped BAM file. The base calling step may generate the total number of reads, the total number of bases, and the average read length as QC measurements to indicate base calling quality. Base calling may be performed by analyzing any preferred signal characteristics (e.g., signal amplitude or intensity). Signal processing and base calling for use in this teaching may include one or more features described in U.S. Patent Application Publication No. 2013 / 0090860, published April 11, 2013, U.S. Patent Application Publication No. 2014 / 0051584, published February 20, 2014, and U.S. Patent Application Publication No. 2012 / 0109598, published May 3, 2012, which are each incorporated herein by reference in their entirety.
[0036] Once the nucleotide sequence for sequence reading is determined, the sequence reading can be provided to an alignment step, for example, in an unmapped BAM file. The alignment step maps the sequence reading to a reference genome to determine the aligned sequence reading and associated mapping quality parameters. The alignment step may generate a percentage of mappable reads as a QC index indicating the quality of the alignment. The alignment results can be stored in a mapped BAM file. A method for aligning sequence readings for use in this instruction may include one or more features described in U.S. Patent Application Publication No. 2012 / 0197623, published on 2 August 2012, which is incorporated herein by reference in its entirety.
[0037] The structure of the BAM file format is described in the "Sequence Alignment / Map Format Specification" (https: / / github.com / samtools / hts-specs) dated September 12, 2014. As described herein, "BAM file" refers to a file compatible with the BAM format. As described herein, an "unmapped" BAM file refers to a BAM file that does not contain aligned sequence reading information and mapping quality parameters, and a "mapped" BAM file refers to a BAM file that contains aligned sequence reading information and mapping quality parameters.
[0038] In some embodiments, the variant calling step may include the detection of single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), polynucleotide polymorphisms (MNPs), and complex block substitution events. In various embodiments, the variant caller may be configured to communicate the variants called against the sample genome as *.vcf, *.gff, or *.hdf data files. The called variant information can be communicated using any file format, as long as the called variant information can be analyzed and / or extracted for analysis. The variant detection method for use in this teaching may include one or more features described in U.S. Patent Publication No. 2013 / 0345066, published December 26, 2013; U.S. Patent Publication No. 2014 / 0296080, published October 2, 2014; U.S. Patent Publication No. 2014 / 0052381, published February 20, 2014; and U.S. Patent No. 9,953,130, issued April 24, 2018, which are each incorporated herein by reference in their entirety. In some embodiments, the variant calling step may be applied to molecularly tagged nucleic acid sequence data. The variant detection method for molecularly tagged nucleic acid sequence data may include one or more features described in U.S. Patent Publication No. 2018 / 0336316, published November 22, 2018, which is each incorporated herein by reference in its entirety.
[0039] In some embodiments, the analysis pipeline may include a fusion analysis pipeline for fusion detection. A fusion detection method may include one or more features described in U.S. Patent Application Publication No. 2016 / 0019340, published on January 21, 2016, which is incorporated herein by whole reference. In some embodiments, the fusion analysis pipeline may be applied to molecularly tagged nucleic acid sequence data. A fusion detection method for molecularly tagged nucleic acid sequence data may include one or more features described in U.S. Patent Application Publication No. 2019 / 0087539, published on March 21, 2019, which are each incorporated herein by whole reference.
[0040] In some embodiments, the analysis pipeline may include a copy number variant analysis pipeline for detecting copy number variation. Methods for detecting copy number variation may include one or more features described in U.S. Patent Application Publication No. 2014 / 0256571, published September 11, 2014, U.S. Patent Application Publication No. 2012 / 0046877, published February 23, 2012, and U.S. Patent Application Publication No. 2016 / 0103957, published April 14, 2016, which are each incorporated herein by reference in their entirety.
[0041] In some embodiments, the server system software may support an encapsulated assay configuration that includes the assay name, assay type, panel, hotspot file (if any), reference name, control name (if any), quality control (QC) threshold, assay description (if any), data analysis parameters and values, instrument execution script name, and other configurations that define the assay. The entire set of information is referred to as the assay definition. The assay configuration and corresponding workflow may be delivered to the user as modular software components in an assay definition file (ADF). The server system software may import assay definition files containing assay configurations. The import process may be initiated by importing a zip file containing encrypted Debian files, which triggers the installation process. The user interface may provide a page for the user to select an ADF for import. An application store in a cloud-based support and resource system may store ADFs supporting various assays, panels, and workflows that are available for user selection for download to the user's local server system.
[0042] An Assay Definition File (ADF) is an encapsulated file that defines the configuration for a molecular test or assay, including the assay name, technology platform configuration (e.g., next-generation sequencing (NGS), chip type, chemistry type), workflow steps (sample preparation, instrument script, analysis, reporting), analytical algorithm, regulatory labels (e.g., for scientific research use only (RUO), for in vitro diagnostics (IVD), for Central European in vitro diagnostics (CE-IVD), for internal use only (IUO), etc.), targeted markers (panel), reference genome version, consumables, control values, QC thresholds, and reported genes and variants. The ADF provides a modular approach to building assay functionality for local sequencing instruments. Assay software may be provided by the ADF separately from the sequencing instrument's platform software.
[0043] The advantages of using ADF for assay configuration include: • Encapsulation of assay workflows and analyses • Install with a single click • The modular structure of the software, enabled by a Docker implementation that allows for separation from platform software, eliminates the need for re-validation after software updates for assay configuration. • Multi-layer encryption for secure delivery • Streamlined support for assay configurations for original equipment manufacturing (OEM) • Streamlined and customized reporting • Support for local regulatory requirements • The plug-and-play format supports technology-independent workflows. • Enables rapid expansion of molecular testing menus and adoption of assays in the laboratory.
[0044] In some embodiments, the assay definition file (ADF) may include one or more software code modules for the following steps: 1) library preparation, 2) templating, 3) sequencing, 4) analysis, 5) variant interpretation, and 6) report generation. For the library preparation and templating workflow step (Figure 7), the ADF may include scripts for library preparation, templating, and enrichment of the templated beads. For the sequencing and analysis workflow step, the ADF may include algorithm binary code and Docker image packages of parameters for the analysis pipeline described with respect to Figure 2. For the variant interpretation workflow step, the ADF may include a list of annotation sources that can be used for variant analysis and annotation. For the report generation workflow step, the ADF may include report templates and image files for use when generating the report.
[0045] An ADF may include instrument scripts for controlling workflow steps on a sequencing instrument. For example, the script may include parameters that control the amount of pipetting and robotic control. Instrument scripts can be customized for specific assays.
[0046] For example, for the sequencing and analysis steps, the ADF may include a Docker image of the end-to-end analysis pipeline. The Docker image may include OS-specific libraries and binaries for the algorithms of each step in the analysis pipeline. The algorithm binaries may include steps in the analysis pipeline, such as those described with respect to Figures 2 and 9, including signal processing, base calling, alignment, and variant calling. In another example, the ADF Debian file may package specific code modules for a particular assay, such as code modules for signal processing, base calling, and RNACounts.
[0047] The ADF may include scripts for configuring the reagent kit. These scripts support the calculation of consumables required for sequencing, as further described below with respect to Table 1. The configuration scripts included in the ADF may include one or more of the following: • Barcode set and chip Library kits and consumables, including functions for associating sample control configurations (e.g., sample inline control) with their QC parameters. Template kits and consumables, including a function to associate internal control values with QC parameters. • Sequence determination kit including a function to correlate internal control values with QC parameters.
[0048] An ADF may contain one or more reference genome files. Examples of reference genomes include hg19 and GRCH38. Reference genome files may be packaged in the main ADF along with workflow information. Alternatively, reference genome files may be packaged in a separate ADF that supplements the main ADF.
[0049] The ADF may include code modules for the fusion panel and fusion target region panel workflows. The ADF may also include fusion target region reference files and hotspot files for analysis.
[0050] An ADF can include assay parameters at various points in a workflow that can be configured by the user. Configurable parameters can be displayed in the user interface for user adjustment. New parameters can be added at any actor level. Configurable parameters can be passed to the analysis pipeline. Input formats for configurable assay parameters can include one or more of the following: single-string text, Boolean, multi-line text, floating-point numbers, radio buttons, dropdowns, and file uploads. For example, file uploads can use file formats such as .properties and .json.
[0051] An ADF may include QC parameters used for quality control and assay performance thresholds at various points in the workflow. For example, types of QC parameters may include execution QC parameters, sample QC parameters, internal control QC parameters, and assay-specific QC parameters. QC parameters may be defined by one or more of the following: data type (e.g., integer, floating-point), lower limit, upper limit, and default value.
[0052] The ADF may include designated data tab columns for result presentation selected from a database for a given assay. The selected data tab columns support the configuration of the user interface display of the results and the columns to be included in the PDF report for the assay. The ADF may include image files for result presentation for a given assay. The ADF may include support for multiple languages for the PDF report. The ADF may include a download file list for any files generated by the analysis pipeline for a given assay. The file list for samples or runs may be displayed in the user interface. The ADF may include a gene list. The gene list may be used to display a list of known genes for a given cancer type in the user interface and the PDF report.
[0053] An ADF may contain a set of plugins to be used for a given assay. The ADF may specify a set of plugins and their versions. If the ADF does not specify plugin versions, the latest versions of the plugins installed on the server system may be used for the given assay.
[0054] ADF may include new workflow templates to support the creation of custom assays. These new workflow templates may include a set of assay chevron steps. Parameters for these steps may be displayed.
[0055] An ADF may include a list of annotation sources and sets to support the construction of a new annotation set. An ADF may include a filter chain applied to variants detected by the analytical pipeline of a given assay. An ADF may include a rule set for variant annotation.
[0056] ADFs can be configured to support a variety of assays. Examples include, but are not limited to, oncology-related assays (e.g., Thermo Fisher Scientific's Oncomine assay), immuno-oncology-related assays (e.g., T-cell receptor (TCR), microsatellite instability (MSI), tumor mutational load (TML)), infectious disease-related assays (e.g., microbiome), reproductive health-related assays, and exome-related assays. ADFs can also be configured for custom assays.
[0057] Figure 3 is a schematic diagram illustrating the generation of an assay definition file according to an embodiment. The assay definition may be generated by build.sh, debscripts, and makedeb.sh, which start a database population and copy the assay information files to form a Debian file. The contents of the assay definition may include assay parameters, BED files (browser extensible data files that define chromosome locations or regions), panel files, gene lists, hotspot files (typically BED or VCF files that define regions in genes containing variants), and seed data containing acceptable reagents. The contents of the assay definition may include localized versions of the assay name, description, and reporting messages to support the display of assay information in different languages. The assay definition file may support the packaging of new analysis pipelines. The ADF may include optional post-processing scripts that can be executed for variant calls, fusion calls, and CNV calls, depending on the assay type. The ADF may include optional Docker container images for updates to binaries for specific analysis pipelines. The Docker container images may be packaged with the ADF to ensure that platform changes, such as the operating system or third-party libraries, do not have a strong impact on assay results or system functionality.
[0058] Debian files can be serialized to prevent unauthorized modification. The serialized assay definition can be further encrypted using the Advanced Encryption Standard (AES), a symmetric key algorithm. Text files containing assay metadata can also be encrypted using AES and the same encryption key. The encrypted assay definition file can be compressed into a zip file along with the encrypted metadata file. Other encryption methods can also be applied to the serialized assay definition information. For example, the metadata may include one or more of the following: • Analysis pipeline version, • Reference genome path for the reference genome file location, • The unique name of the assay, which is the internal name of the assay used to check for unique occurrences in the system. Docker image name used for launching the analysis and installing assay-dependent file references. • The names of dependency packages required to start the analysis pipeline.
[0059] Figure 4 is a schematic diagram of an example of assay definition file packaging. The assay definition file compressed in ZIP format 40 may include a serialized and encrypted assay definition Debian package 41, a serialized and encrypted metadata text file 42, and a serialized and encrypted optional Docker image Debian package 43. The server system may decrypt both the metadata text file 42 and the assay definition serialization file 41 before installing the assay definition Debian file.
[0060] The server system and modular software components may be configured to control multiple functional modes, including RUO or AD mode, and IVD or Dx mode. Referring to Figure 1, the Tomcat Server may be configured to include Web ARchive (WAR) files for RUO mode and WAR files for IVD mode. The server system may be configured to include a RUO variome database for variants detected by the RUO assay and an IVD variome database for variants detected by the IVD assay. The server system may be configured to include separate analysis pipelines and associated Kepler workflow engines for RUO mode and IVD mode. RUO Docker image files for the RUO assay may be configured as separate files from IVD Docker image files for the IVD assay. The relational database may be configured to have separate databases: an assay development (AD) database for RUO mode and a Dx database for IVD mode. Initially, a server system supporting only RUO mode may be configured to support both RUO and IVD modes through a software update.
[0061] ADFs can be generated separately for RUO-mode assays and IVD-mode assays. RUO-mode ADFs may include assay definitions for assays used in academic research. RUO-mode ADFs may be developed by third parties. IVD-mode ADFs may include assay definitions for assays compliant with local regulatory requirements for diagnostic use.
[0062] Figure 5 includes a diagram of an exemplary instrument 500 incorporating a three-axis pipetting robot. In one example, instrument 500 may be a sequencing device incorporating a sample preparation platform. For example, instrument 500 may include an upper and a lower section. The upper section may include a door 506 for accessing a deck 510 where samples, reagent containers, and other consumables are placed. The lower section may include a cabinet for storing additional reagent solutions and other parts of instrument 500. Furthermore, the instrument may include a user interface such as a touchscreen display 508.
[0063] In certain examples, the instrument 500 may be a sequencing instrument (a sequencing instrument, sequencing device, and sequencing apparatus used interchangeably). In some embodiments, the sequencing instrument includes an upper section, a display screen, and a lower section. In some embodiments, the upper section may include a deck supporting components of the sequencing instrument and consumables, including a template section, a sequencing tip, and reagent strip tubes and carriers. In some embodiments, the lower section may house reagent bottles containing reagents used for sequencing, and a waste container.
[0064] In some embodiments, a camera mounted on the cabinet of the upper section of the instrument is directed towards the deck to monitor which items are arranged as preparations for arrangement determination execution. The camera may acquire images at time intervals. For example, images may be acquired at intervals of 3-4 seconds or any preferred interval. The processor analyzes the images to detect the completion of a task by the user. The processor may provide feedback and instructions for the next task in the preparation via a display screen. The display screen may show graphic representations of instrument components and consumables to indicate instructions for the user.
[0065] An exemplary instrument deck 510 is shown in Figure 6 as instrument deck 600. Instrument deck 600 is housed in the upper section of the instrument within the field of view of one or more cameras. The sample preparation deck may include multiple locations configured to receive reagent strips, supplies, sequencing tips, and other consumables. As used herein, consumables are components used by the instrument that are replaced periodically when they are used. For example, consumables include flow cells and associated sensors, among other disposable components that are not part of the instrument's permanent components, such as reagent and solution strips or containers, pipette tips, microwell arrays, and other disposable components.
[0066] In this example, the instrument deck system 600 includes a pipetting robot 602 that accesses various reagent strips and containers, pipette tips, microwell arrays, and other consumables to implement the tests. Furthermore, the system may include a mechanism 604 for carrying out the tests. An exemplary mechanism 604 may include a mechanical conveyor or slide and fluid system.
[0067] In the example, the instrument deck 600 includes trays 606 or 608 for receiving solution or reagent strips of a specific configuration. In the example of a sequencing instrument, tray 606 can be used for library and template solutions in appropriately configured strips, and tray 608 can receive library and template reagents in an appropriate configuration.
[0068] Furthermore, the instrument can be configured to receive sequencing chips containing microwell arrays 610 and 612 at specific locations on the deck. For example, a sample can be supplied in the array of microwells of sequencing chip 612. In another example, the system can be configured to receive additional reagents 614 in different strip configurations. In yet another example, the reagent solution can be supplied in array 616. In a further example, a container array 620 can be supplied in combination with instrumentation such as a thermocycler. Furthermore, the system may include other instruments such as a centrifuge, which may be supplied with consumables such as tubing. Additionally, a tray can be provided to receive pipetting tips 622.
[0069] The proper provisioning of consumables at each of these locations can be monitored by a vision system including one or more cameras. The deck may be provided with one or more cameras for tracking the provisioning and protection of reagents and other consumables. When a reagent that should be used to carry out a plan is insufficient, or when reagent consumables are found to be used up, the user can be prompted through the user interface.
[0070] Figure 7 shows the workflow of a sequencing instrument. The top-level steps include library preparation, template creation, and sequencing.
[0071] The components of a sequencing device may include a sequencing chip (interchangeably a microchip, chip, or sensor device) containing a microwell array that fluidly communicates with a sensor array, and a flow cell having multiple lanes. Figure 8 shows an example of a sequencing chip 700 having four lanes 701, 702, 703, and 704. Each lane is accessed individually by its respective fluid inlet 710 and fluid outlet 712. Alternatively, the sensor device 700 may include fewer than four lanes or more than four lanes. For example, the sensor device 700 may include 1 to 10 lanes, such as 2 to 8 lanes or 4 to 6 lanes. The lanes can be fluidically separated from each other. Therefore, the lanes can be used in parallel or simultaneously at different times, depending on the nature of the execution plan.
[0072] It is advantageous to optimize the use of sequencing chip lanes for multiple assays. A given lane can accommodate two or more samples. In some embodiments, server system software may be provided for optimizing chip usage by applying one or more of the following rules. The maximum number of assays that can be included in a single planned run is equal to the number of available chip lanes. This rule applies to both new and used chips. The maximum number of assays possible in a single plan run may be adjusted depending on the number of lanes required by the assay. Rules for determining the number of lanes may include the following: ■ One assay per lane ■If the minimum number of reads per sample in the assay exceeds the lane capacity, calculate the number of lanes required, i.e., (minimum reads / lane capacity). For example, 2,000,000 / 1,300,000 = 1.54 lanes, round up to 2, so the assay requires 2 lanes. • The combined pool size of selected assays cannot exceed 8. ○ Combined pool size = Total (pool size of each assay) Regarding the AmpliSeq panel (Thermo Fisher Scientific), the pool size of the AmpliSeq assay = total (number of DNA pools, number of RNA pools). Regarding the AmpliSeq HD panel (Thermo Fisher Scientific), the pool size of the AmpliSeq HD assay = the number of TNA pools. The following rules may apply to PCR profiles. The number of individual PCR profiles (thermocycling) in a single plan execution cannot exceed two. For DNA and Fusion assays, DNA samples and Fusion samples must be assigned to separate zones. This rule limits the number of PCR profiles supported in a single design run. TNA, DNA, and Fusions assays can be performed in a single plan. In this case, if the PCR profiles for TNA and RNA are the same, TNA and RNA can be in the same zone. DNA may be in a different zone. ■ A PCR profile is defined per assay. ■ A PCR profile is an assay attribute that is stored in the database when an assay is stored. ■Regarding the factory-shipped assay, the PCR profile is pre-seeded. ■Regarding custom assays, users can edit the PCR profile during assay creation, which is explained in detail in the user story during assay creation. • Assays in a single planned run may have the same or different analytical pipeline versions. • Assays in a single design run can be of the same or different application types (DNA only, RNA only, DNA + RNA, etc.). • The number of flows for all assays in a single plan does not need to be the same. The maximum number of flows is used for execution. The analysis pipeline should only analyze data for the number of flows configured in the assay. Setting flow limit parameters corresponding to the assay may limit signal processing to the number of flows configured in the assay. • Assays in a single plan run may have different template sizes.
[0073] In some embodiments, the software may be configured to display a warning message if the chip type or capacity does not match the ongoing plan. In the following exemplary scenario, a confirmation dialog with a warning message may be displayed to the user. The user's confirmation selection is maintained, and the remaining verification may occur based on the user's selection to consider new chips or on-deck chips. If the selected assay chip type does not match the one on the deck, a confirmation dialog will appear with the warning message "The chip type on the deck does not match the selected assay," and you will be asked whether to click [Yes] or [Cancel] to consider a new chip, or to use the deck chip. If the number of selected assays exceeds the available lane capacity, a confirmation message will appear stating, "The chips on the deck have only N lanes and can process only N assays," and asking if you want to consider adding new chips. • If one assay is selected but the number of reads per sample exceeds the available lane capacity, a confirmation message will appear stating, "The selected assay exceeds the available lane capacity of the chip, and therefore the minimum number of reads per sample cannot be achieved." Should I consider a new chip? If [Yes], switch to new chip verification. If [No] is selected, the user can proceed with the selected option. Clicking [Next] will require the software to assign lanes for each selected assay. The lane assignment rules may be as follows: One assay per lane If the minimum number of reads per sample in an assay exceeds the lane capacity, calculate the required number of lanes, i.e., (minimum reads / lane capacity). For example, 2,000,000 / 1,300,000 = 1.54 lanes, round up to 2, and then the assay requires 2 lanes.
[0074] The chip lane allocation rules may include the following: • Number of lanes allocated to the assay = upper limit ((number of selected samples + control value) x minimum reading per sample / reading per lane) • If multiple lanes are assigned to an assay, the assigned lanes must be consecutive. After the final number of lanes required for the assay is determined on the sample page, the software must readjust the lane assignments to ensure consecutive lane assignments.
[0075] Figure 9 is an example block diagram for processing sequencing data from multiple lanes of a sequencing chip. Preprocessing may prepare analyses corresponding to each chip lane according to the assay assigned to the lane. For example, server software may create data structures such as pipeline folder structures for assays corresponding to individual lanes and folder structures for each sample in each lane. As described with respect to Figure 2, signal processing, e.g., signal measurements resulting from the 1.wells file, may be input to the parallel process block 810. The base calling step 820 may be applied to multiple signal measurements corresponding to each lane to determine the base sequences of multiple sequence reads for the lane. In step 830, the sequence reads per sample per lane are provided to the alignment step 840. The sequence reads may be provided to the alignment step, for example, in an unmapped BAM file per sample per lane. The alignment step 830 maps the sequence reads to a reference genome. The mapped reads per sample per lane may be stored in mapped BAM files corresponding to the sample and lane. The variant calling step 850 may be applied to the mapped reads corresponding to the sample and lane, according to the type of assay. The base calling step 820, alignment step 840, and variant calling step 850 are described with respect to Figure 2. The Kepler workflow engine may be applied to control one or more of the processing flows from the steps in Figure 9. Once the variant calling step 850 is completed for the samples and lanes, the results may be prepared for reporting in step 860. For example, the results may be used to input data into a unique PDF file and generate image files for a particular assay. In step 870, the results may be displayed to the user or provided in a PDF file.
[0076] In some embodiments, the server software may calculate the consumables required for sequence determination. Table 1 lists examples of consumable calculations. Table 1
[0077] According to an exemplary embodiment, a method is provided comprising: receiving an assay definition file in a local server system from a server of a cloud computing and storage system, wherein the assay definition file includes code modules for configuring the assay; storing the code modules in the memory of the local server system; receiving sequencing data in the local server system from a sequencing device, wherein the sequencing data is generated by the sequencing device during sequencing execution for the assay; and applying an analysis pipeline for the assay to the sequencing data, wherein the analysis pipeline includes analysis steps performed by the processor of the local server system according to code modules from the assay definition file to generate assay analysis results. The code modules for the analysis pipeline may include a code module for a base calling step, the base calling step generating sequence reads. The code modules for the analysis pipeline may include a code module for an alignment step, the alignment step generating aligned sequence reads. The code modules for the analysis pipeline may include a code module for a variant calling step, the variant calling step applied to the aligned sequence reads to generate variant calling results. The method may further include storing the variant calling results in a variome database of the local server system. This method may further include displaying assay analysis results, the display including image files for presenting results for the assay. The assay definition file may include image files for presenting results for the assay. The assay definition file may include a reference genome file. The assay definition file may include a list of annotation sources. The analysis pipeline may be applied in parallel to sequencing data corresponding to multiple lanes of a sequencing chip installed in the sequencing device. Each lane of the multiple lanes may correspond to its own assay, and the step of applying the analysis pipeline applies the analysis steps of each assay to the sequencing data for that lane.The method may further include displaying a page to the user on the local server system's user interface for selecting assay definition files to import from a cloud computing and storage system to the local server system. The method may further include multiple assay definition files, which include research-only (RUO) mode assay definition files and in vitro diagnostic (IVD) mode assay definition files.
[0078] According to an exemplary embodiment, a local server system is provided, comprising: a memory; and a processor, which, when executed by the processor, causes the local server system to perform a method comprising: receiving an assay definition file from a server of a cloud computing and storage system, wherein the assay definition file includes code modules for structuring an assay; storing the code modules in the memory of the local server system; receiving sequencing data from a sequencing device, wherein the sequencing data is generated by the sequencing device during sequencing execution for the assay; and applying an analysis pipeline for the assay to the sequencing data, wherein the analysis pipeline includes analysis steps performed by the processor of the local server system according to the code modules from the assay definition file to generate assay analysis results. The code modules for the analysis pipeline may include a code module for a base calling step, the base calling step generating sequence reads. The code modules for the analysis pipeline may include a code module for an alignment step, the alignment step generating aligned sequence reads. The code module for the analysis pipeline includes a code module for the variant calling step, which is applied to aligned sequence reads to generate variant calling results. The server system may further include a variome database for storing the variant calling results. The method may further include displaying the assay analysis results, the display including image files for result presentation for the assay. The assay definition file may include image files for result presentation for the assay. The assay definition file may include a reference genome file. The assay definition file may include a list of annotation sources.The analysis pipeline can be applied in parallel to sequencing data corresponding to multiple lanes of a sequencing chip installed in the sequencing device. Each lane of the multiple lanes may correspond to a respective assay, and the step of applying the analysis pipeline applies the analysis steps of each assay to the sequencing data for that lane. The method may further include displaying a page to the user in the user interface of the local server system for selecting assay definition files to import from a cloud computing and storage system to the local server system. The local server system may further include multiple assay definition files, which include research-use-only (RUO) mode assay definition files and in vitro diagnostic (IVD) mode assay definition files. The local server system may further include a first database and a second database, the first database storing information for the research-use-only (RUO) mode of the operation and the second database storing information for the in vitro diagnostic (IVD) mode of the operation.
[0079] According to various exemplary embodiments, one or more features of any one or more of the teachings and / or exemplary embodiments described above may be performed or implemented using appropriately configured and / or programmed hardware and / or software elements. The determination of whether an embodiment is implemented using hardware and / or software elements may be based on any factors, such as desired computing speed, output level, heat tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0080] Examples of hardware elements may include processors, microprocessors, input and / or output (I / O) devices (or peripherals) communicatively connected via local interface circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Local interfaces may include, for example, one or more buses or other wired or wireless connections, controllers, buffers (caches), drivers, repeaters, and receivers to enable proper communication between hardware components. A processor is a hardware device for executing software, particularly software stored in memory. A processor can be any custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with a computer, a semiconductor-based microprocessor (e.g., in the form of a microchip or chipset), a macroprocessor, or generally any device for executing software instructions. A processor can also represent a distributed processing architecture. I / O devices can include input devices such as keyboards, mice, scanners, microphones, touchscreens, interfaces for various medical devices and / or laboratory equipment, barcode readers, styluses, laser readers, and radio frequency device readers. Furthermore, I / O devices can also include output devices such as printers, barcode printers, and displays. Finally, I / O devices can also include devices that communicate as both inputs and outputs, such as modulators / demodulators (modems; for accessing another device, system, or network), radio frequency (RF) transceivers or other transceivers, telephone interfaces, bridges, routers, and so on.
[0081] Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, arithmetic codes, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Software in memory may include one or more separate programs that may contain an ordered list of executable instructions for implementing logical functions. Software in memory may include a system for identifying data flows in accordance with this teaching, as well as any suitable custom-made or commercial operating system (O / S) that can control the execution of other computer programs such as the system and provide scheduling, input / output control, file and data management, memory management, communication control, etc.
[0082] According to various exemplary embodiments, one or more features of any one or more of the above teachings and / or exemplary embodiments may be performed or implemented using a appropriately configured and / or programmed non-temporary machine-readable medium or article that, when performed by a machine, can store instructions or sets of instructions that can cause the machine to perform the methods and / or operations according to the exemplary embodiments. Such a machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, scientific instrument or experimental instrument, and may be implemented using any suitable combination of hardware and / or software. Machine-readable media or articles may include, for example, any suitable type of memory unit, memory device, memory article, memory medium, storage device, storage article, storage medium, and / or storage unit, such as memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disks, floppy disks, read-only compact disks (CD-ROMs), recordable compact disks (CD-Rs), rewritable compact disks (CD-RWs), optical disks, magnetic media, magneto-optical media, removable memory cards or disks, various types of digital multi-purpose disks (DVDs), tapes, cassettes, and any medium suitable for use in a computer. Memory may include any one or combination of volatile memory elements (e.g., random access memory (RAM, e.g., DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, EPROM, EEROM, flash memory, hard drives, tapes, CD-ROMs, etc.). Furthermore, memory may incorporate electrical, magnetic, optical, and / or other types of storage media. Memory can have a distributed architecture where various components are located geographically separated from each other, but are still accessed by the processor.Instructions may include source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, etc., implemented using any suitable type of code, such as any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0083] According to various exemplary embodiments, one or more features of any one or more of the teachings and / or exemplary embodiments described above may be performed or implemented using, at least in part, a distributed, clustered, remote, or cloud computing and storage system. In some embodiments, one or more users may access computers or servers of the cloud computing and storage system via an intranet and / or the internet. In some embodiments, users may remotely access servers of the cloud computing and storage system via a web client.
[0084] According to various exemplary embodiments, one or more features of any one or more of the teachings and / or exemplary embodiments described above may be done or implemented using any other entity including a source program, an executable program (objective code), a script, or a set of instructions to be performed. When it is a source program, the program may be translated through a compiler, assembler, interpreter, etc., which may or may not be contained in memory, in order to connect with the OS and function properly. Instructions may be written using (a) an object-oriented programming language having classes of data and methods, or (b) a procedural programming language having routines, subroutines, and / or functions, which may include, for example, C, C++, R, Pascal, Basic, Fortran, Cobol, Perl, Python, Java, and Ada.
[0085] According to various exemplary embodiments, one or more of the above exemplary embodiments may include transmitting, displaying, storing, printing, or outputting any information, signals, data, and / or intermediate or final results generated, accessed, or used by such exemplary embodiments to a user interface device, computer-readable storage medium, local computer system, or remote computer system. Such transmitted, displayed, stored, printed, or outputted information may take the form of, for example, searchable and / or filterable lists of executions and reports, images, tables, charts, graphs, spreadsheets, correlations, arrays, and combinations thereof.
[0086] While preferred embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided merely as examples. Numerous variations, modifications, and substitutions will be conceivable to those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in carrying out the present invention. The following claims define the scope of the present invention, and methods and structures within the scope of these claims, as well as their equivalents, are intended to be covered thereby.
Claims
1. The local server system receives an assay definition file from a server in a cloud computing and storage system, wherein the assay definition file includes code modules for constructing the assay, and the code modules constitute steps in the workflow for the assay, including library preparation, templating, sequencing, and analysis pipelines. The code module is stored in the memory of the local server system, The local server system receives sequencing data from a sequencing device, the sequencing data being generated by the sequencing device during sequencing execution according to an instrument script from the assay definition file for the assay, and the receiving of such data. Applying an analytical pipeline for the assay to the sequencing data, wherein the analytical pipeline includes analytical steps performed by the processor of the local server system according to the code module from the assay definition file to generate assay analysis results, method.
2. The method according to claim 1, wherein the code module for the analysis pipeline includes a code module for a base calling step, and the base calling step generates a sequence read.
3. The method according to claim 2, wherein the code module for the analysis pipeline includes a code module for an alignment step, and the alignment step generates an aligned sequence read.
4. The method according to claim 3, wherein the code module for the analysis pipeline includes a code module for a variant calling step, the variant calling step being applied to the aligned sequence reads to generate a variant calling result.
5. The method according to claim 4, further comprising storing the variant calling result in the variome database of the local server system.
6. The method according to claim 1, further comprising displaying the assay analysis results, wherein the display includes an image file for presenting the results for the assay.
7. The method according to claim 6, wherein the assay definition file includes the image file for presenting the results for the assay.
8. The method according to claim 1, wherein the assay definition file includes a reference genome file.
9. The method according to claim 1, wherein the assay definition file includes a list of annotation sources.
10. The method according to claim 1, wherein the analysis pipeline is applied in parallel to the sequencing data corresponding to multiple lanes of a sequencing chip installed in the sequencing device.
11. The method according to claim 10, wherein each of the plurality of lanes corresponds to a respective assay, and the step of applying the analysis pipeline is to apply the analysis step of each assay to the sequencing data for the lane.
12. The method according to claim 1, further comprising displaying a page to the user on the user interface of the local server system for selecting the assay definition file to import from the cloud computing and storage system to the local server system.
13. The method according to claim 1, further comprising a plurality of assay definition files, wherein the plurality of assay definition files include a research use only (RUO) mode assay definition file and an in vitro diagnostic (IVD) mode assay definition file.
14. It is a local server system, Memory and A processor, which, when executed by the processor, on the local server system, The local server system receives an assay definition file from a server of the cloud computing and storage system, wherein the assay definition file includes code modules for configuring the assay, and the code modules constitute steps in the workflow for the assay, including library preparation, templating, sequencing, and analysis pipelines. The code module is stored in the memory of the local server system, The local server system receives sequencing data from a sequencing device, the sequencing data being generated by the sequencing device during sequencing execution according to an instrument script from the assay definition file for the assay, and the receiving of such data. Applying an analytical pipeline for the assay to the sequencing data, wherein the analytical pipeline includes analytical steps performed by the processor of the local server system according to the code module from the assay definition file to generate assay analysis results, Configured to execute instructions that cause the method to be performed, Equipped with a processor, Local server system.
15. The local server system according to claim 14, wherein the code module for the analysis pipeline includes a code module for a base calling step, and the base calling step generates a sequence read.
16. The local server system according to claim 15, wherein the code module for the analysis pipeline includes a code module for an alignment step, and the alignment step generates an aligned sequence read.
17. The local server system according to claim 16, wherein the code module for the analysis pipeline includes a code module for a variant invocation step, the variant invocation step is applied to the aligned sequence reads to generate a variant invocation result.
18. The local server system according to claim 17, further comprising a variome database for storing the results of the aforementioned variant calling.
19. The local server system according to claim 14, further comprising a plurality of assay definition files, wherein the plurality of assay definition files include a research use-only (RUO) mode assay definition file and an in vitro diagnostic (IVD) mode assay definition file.
20. The local server system according to claim 14, further comprising a first database and a second database, the first database storing information for a research use-only (RUO) mode of the operation and the second database storing information for an in vitro diagnostic (IVD) mode of the operation.