Genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis

By employing integrated circuits with hardwired digital logic circuits for bioinformatics protocols, the challenges of labor-intensive and error-prone DNA/RNA sequencing are addressed, facilitating faster and more accurate genomic data processing for personalized medicine.

JP2025157232APending Publication Date: 2025-10-15EDICO GENOME CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025102586
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-01-11
Filing Date
2025-06-18
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Current bioinformatics systems for DNA/RNA sequencing are labor-intensive, time-consuming, and prone to errors, with high computational demands and costs, limiting the widespread adoption of personalized medicine.

Method used

Implementing bioinformatics protocols on integrated circuits with a combination of software and hardware processing platforms, utilizing hardwired digital logic circuits and interconnected processing engines to perform genetic analysis tasks more efficiently and accurately.

Benefits of technology

Accelerates DNA/RNA sequencing and analysis, reducing computational intensity and errors, enabling faster and more accurate genomic data processing for personalized medicine applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025157232000001_ABST
    Figure 2025157232000001_ABST
Patent Text Reader

Abstract

To provide a system, method, and apparatus for executing a genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis.SOLUTION: A system 1 includes an integrated circuit formed of hardwired digital logic circuits that are interconnected by physical electrical interconnects. One of the physical electrical interconnects forms an input to the integrated circuit connected with an electronic data source for receiving reads of genomic data. The hardwired digital logic circuits are arranged as a set of processing engines, each processing engine being formed of a subset of the hardwired digital logic circuits to perform processing in a sequence analysis pipeline on the reads of genomic data. Each subset is formed in a wired configuration to perform the processing in the sequence analysis pipeline.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 277,445, filed January 11, 2016, the contents of which are incorporated herein by reference in their entirety.

[0002] The subject matter described herein relates to bioinformatics, and more particularly to systems, apparatus, and methods for implementing bioinformatics protocols, such as performing one or more functions for analyzing genomic data on an integrated circuit, such as a hardware processing platform. [Background technology]

[0003] The goal of medical researchers and practitioners is to improve the safety, quality, and effectiveness of medical care for each individual patient. Personalized medicine is directed to achieving these goals at the individual level. For example, "genomics" and / or "bioinformatics" are fields of research that aim to improve the safety, quality, and effectiveness of preventive and therapeutic treatments at the individualized individual level. Thus, by utilizing genomics and / or bioinformatics techniques, an individual's genetic makeup, e.g., their genes, can be determined, and this knowledge can be used in developing preventive and / or therapeutic plans, including drug therapies, that are personalized for the individual, thereby allowing medicines to be tailored to meet each person's individual needs.

[0004] The desire to provide personalized care to individuals is transforming the healthcare system. This transformation of the healthcare system is likely to be fueled by groundbreaking innovations at the intersection of medical science and information technology, such as those represented by the fields of genomics and bioinformatics. Genomics and bioinformatics are therefore critical foundations on which this future will be built. Science has evolved dramatically since the first human genome was completely sequenced in 2000 at a total cost of over $1 billion. Today, we are achieving high-resolution sequencing at a cost of less than $1,000 per genome, making it economically feasible for the first time to move beyond the laboratory and be widely adopted in healthcare. Genomic data can therefore be a critical input into screening, therapeutic and / or preventative drug discovery, and / or disease treatment.

[0005] More specifically, genomics and bioinformatics are fields related to the application of information technology and computer science to the field of molecular biology. Specifically, bioinformatics techniques can be applied to process and analyze genomic data to improve the safety, quality, and effectiveness of healthcare at a personalized level by determining qualitative and quantitative information about various genomic data from individuals and the like that can be used by various practitioners in developing preventative and therapeutic methods to prevent or at least ameliorate disease states.

[0006] With a focus on advancing personalized medicine, bioinformatics can facilitate preventative rather than symptomatic personalized medicine, giving patients the opportunity to become more involved in their own health. Typically, this can be achieved through two guiding principles. First, federal initiatives can provide support for research addressing these individual modalities of disease and disease prevention, with the ultimate goal of tailoring diagnostic and preventive care to each person's unique genetic characteristics. Additionally, medical data can be aggregated to create a "network of networks" to help researchers establish patterns and identify genetic "definitions" for existing diseases.

[0007] The advantage of using bioinformatics techniques in such cases is that qualitative and / or quantitative analysis of molecular biological data can be performed much faster, and often more accurately, on wider sample sets, facilitating the emergence of personalized medicine systems.

[0008] Thus, in various instances, the molecular data to be processed in bioinformatics-based platforms typically relates to genomic data, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) data. For example, a well-known method for generating DNA and / or RNA data is DNA / RNA sequencing. DNA / RNA sequencing can be performed manually, such as in a laboratory, or by an automated sequencer, such as at a core sequencing facility, for the purpose of determining the genetic makeup of an individual's genetic material, e.g., a sample of DNA and / or RNA. The person's genetic information can then be used to determine variances from a reference, such as a reference sequence, haplotype, or theoretical haplotype. Such variant information can then be further processed and used to determine or predict the occurrence of a disease state in the individual.

[0009] For example, manual or automated DNA / RNA sequencing can be used to determine the sequence of nucleotide bases in DNA / RNA samples, such as samples obtained from test subjects. Using a variety of different bioinformatics techniques, these sequences can then be stitched together to generate the test subject's genome sequence. This sequence can then be compared with a reference genome sequence to determine how the test subject's genome sequence has changed from the reference genome sequence. This process involves determining variants in sampled sequences, which is the central task of bioinformatics methods.

[0010] For example, a central challenge in DNA sequencing is constructing a full-length genomic sequence, e.g., a chromosomal sequence, from a sample of genetic material that can be compared to a reference genomic sequence, such as to determine variants in the sampled full-length genomic sequence. Specifically, methods utilized in sequencing protocols do not produce full-length chromosomal sequences of the sample DNA.

[0011] Rather, fragments of sequence, typically 100-1000 nucleotides in length, are generated without any indication of where they align within the genome. Thus, to generate the genome structure of the full-length chromosomes, these fragments of DNA sequence must be mapped, aligned, merged, and / or compared to a reference genome sequence. Through such processing, variants of the sample genome sequence from the reference genome sequence can be determined.

[0012] However, because the human genome consists of approximately 3.1 billion base pairs, and because each fragment of sequence is typically only 100 to 500 nucleotides in length, the time and effort required to construct such a full-length genome sequence and determine the variants therein is enormous, often requiring the application of several different algorithms over an extended period of time using several different computer resources.

[0013] In certain cases, thousands to millions of fragments of DNA sequence are generated, aligned, and merged to construct a genome sequence roughly equivalent in length to a chromosome, a step in this process may include comparing a DNA fragment to a reference sequence to determine where the fragment aligns in the genome.

[0014] Constructing chromosome-length sequence and determining the variant of sampled sequence involves many such steps.Therefore, a wide variety of methods have been developed for carrying out these steps.For example, there are commonly used software implementations for carrying out one or a series of such steps in bioinformatics system.However, the common characteristics of such software-based bioinformatics method and system are that they are labor-intensive, take a long time to run on general-purpose processor, and are prone to error. Summary of the Invention [Problem to be solved by the invention]

[0015] Therefore, bioinformatics systems that can execute the algorithms implemented by such software in a less labor-intensive / processing-power-intensive manner and with a higher percentage accuracy would be useful. However, even as we approach the "$1,000 genome," the cost of analyzing, storing, and sharing this raw digital data far exceeds the cost of generating it. This data analysis bottleneck is a major obstacle standing between the ever-increasing amount of raw data and the real medical insights we hope to gain from it.

[0016] Thus, presented herein are systems, apparatus, and methods for implementing genomics and / or bioinformatics protocols, such as for performing one or more functions for analyzing genomic data, via software implementation and / or on integrated circuits, such as hardware processing platforms. For example, as described herein below, in various implementations, a combination of software-implementable and / or hardware accelerator solutions, such as those including an integrated circuit and software for interacting with the integrated circuit, may be utilized in performing such bioinformatics-related tasks, where the integrated circuit may be formed of one or more hardwired digital logic circuits, which may be interconnected by multiple physical electrical interconnects that may be arranged as a set of processing engines, each of which may be configured to perform one or more steps in a bioinformatics genetic analysis protocol. An advantage of this configuration is that bioinformatics-related tasks may be performed faster than software-only configurations, such as those typically involved in performing such tasks. However, such hardware accelerator technology is not currently typically utilized in the fields of genomics and / or bioinformatics. [Means for solving the problem]

[0017] The present disclosure relates to performing tasks, such as in bioinformatics protocols. In various instances, multiple tasks are performed, and in some instances, these tasks are performed in a manner that forms a pipeline, with each task and / or its substantial completion serving as a building block for each subsequent task until a desired end result is achieved. Thus, in various embodiments, the present disclosure is directed to performing one or more methods on one or more devices, the devices being optimized for performing those methods. In some embodiments, the one or more methods and / or one or more devices are organized into one or more systems.

[0018] For example, in some embodiments, the present disclosure is directed to systems, devices, and methods for implementing genomics and / or bioinformatics protocols, such as for performing one or more functions for producing and / or analyzing genetic data, that utilize, in various cases, innovative software and / or integrated circuits, such as those implemented in a combination of software and / or hardware processing platforms. For example, in one embodiment, a genomics and / or bioinformatics system is provided. The system may be involved in performing various bioanalytical result-producing and / or analytical functions that are optimized to run faster and / or with greater accuracy. Methods for performing these functions may be implemented in software or hardware solutions. Thus, in some cases, methods are presented that involve data production and / or acquisition and / or analysis that may include the execution of one or more algorithms, where the algorithms are optimized according to the manner in which they are implemented, e.g., software, hardware, or a combination of both. Specifically, if an algorithm is to be implemented in a software solution, the algorithm and / or its accompanying processes may be optimized for execution by that medium to run faster and / or with greater accuracy. Similarly, where the functions of an algorithm are to be implemented in hardware solutions, the hardware is designed to perform those functions and / or their accompanying processes in a manner optimized for execution by that medium so as to perform faster and / or with greater accuracy. Furthermore, where the functions involve a combination of software and / or hardware solutions, those functions and their accompanying processes are designed and configured to work seamlessly together to maintain or improve accuracy while achieving previously unattainable speeds.

[0019] Thus, in one aspect, presented herein are systems, devices, and methods for implementing bioinformatics protocols, such as for performing one or more functions for generating and / or analyzing genetic data, e.g., via one or more developed and / or optimized algorithms and / or on one or more optimized integrated circuits, such as one or more hardware processing platforms. Thus, in one instance, a method is provided for implementing one or more algorithms for performing one or more steps for generating and / or analyzing genomic data in a genomics and / or bioinformatics protocol. In another instance, a method is provided for performing the functions of one or more algorithms for performing one or more steps for analyzing genomic data in a bioinformatics protocol, the functions being at least partially implemented on an integrated circuit, such as one formed of one or more hardwired digital logic circuits. In such instances, the hardwired digital logic circuits may be interconnected, such as by one or more physical electrical interconnects, and may be arranged to function as one or more processing engines. In various instances, a plurality of hardwired digital logic circuits are provided, configured as a set of processing engines, each capable of performing one or more steps in a bioinformatics genetic analysis protocol, such as a bioinformatics processing pipeline.

[0020] More specifically, in one instance, a system for generating genetic sequence data is provided, including, for example, devices and methods for sequencing nucleic acids and / or for executing a sequence analysis pipeline on genetic sequence data. The system may include one or more of an electronic data source, a memory, and / or an integrated circuit, such as those associated with a DNA / RNA sequencing device such as those described herein. For example, in one embodiment, an electronic data source is included, which may be configured to generate and / or provide one or more digital signals, such as digital signals representing one or more reads of genetic data, where each read of genomic data includes a sequence of nucleotides. Additionally, the memory may be configured to store one or more genetic reference sequences and / or may be further configured to store an index, such as an index of annotated splice site data, of the one or more genetic reference sequences.

[0021] Still further, devices and / or methods for generating genetic sequence data are provided. For example, approaches to DNA / RNA analysis, such as for genetic diagnosis and / or sequencing, involving one or more of nucleic acid hybridization, detection, and / or sequencing reactions are provided. In various cases, the approaches may include hybridization and / or detection devices and / or procedures for performing one or more of the following steps. Specifically, for genetic analysis, a subject's RNA or DNA sample to be analyzed may be isolated and immobilized, for example, directly and / or indirectly, on a substrate, such as a substrate including a chemically sensitive one-dimensional (1-D) and / or two-dimensional (2D) reaction layer, e.g., a graphene reaction layer, and / or a three-dimensional (3D) reaction layer, and a probe of a known or to-be-detected genetic sequence, e.g., a disease marker, may be washed across the substrate, or vice versa. In various cases, one or more of the subject's RNA or DNA sample and / or probes may be labeled.

[0022] In other cases, such as those in which the substrate includes a 1D or 2D, e.g., graphene, reactive layer, and / or other chemically sensitive reactive layer, labels or probes, such as chemical or radioactive labels, may not be required and / or may not be included. In either case, if the disease marker is present, a binding event, e.g., hybridization, occurs, and as provided herein, the hybridization event is detectable, e.g., via a labeled analyte or probe and / or via an appropriately configured reactive layer, thereby detecting the presence of the disease marker. If the disease marker is not present, there is no reaction and therefore no detection. Of course, in some cases, the absence of a binding event may be an indicative event. Thus, the system may be configured such that a hybridization event may or may not be detected, indicating the presence or absence of a disease marker in a subject's sample.

[0023] Similarly, for DNA and / or RNA sequencing, an unknown nucleic acid sequence, e.g., a single-stranded sequence of a subject's DNA or RNA, whose nucleotide identity is to be determined, is first isolated, amplified, and immobilized on a substrate, which may include a 1D, 2D, e.g., graphene layer, 3D, or other configured reaction layer, as described herein. Next, a known nucleic acid, e.g., a nucleotide base that may be labeled with an identifiable tag, is contacted with the unknown nucleic acid sequence in the presence of a polymerase. As noted, labeled reactants do not need to be included if the reaction event occurs near an appropriately configured reaction layer, e.g., a reaction layer including graphene.

[0024] Thus, when hybridization occurs, nucleic acid binds to complementary bases in unknown sequence, for example, in the sample DNA or RNA being sequenced, and is immobilized on the surface of the substrate, such as near the reaction layer.Then, the binding event can be detected, for example, optically, electrically, and / or through a suitable detectable reaction that occurs in the reaction layer.Then, these steps are repeated until the entire DNA or RNA sample is completely sequenced.Usually, these steps are carried out by a next-generation sequencer as known in the art, or they can be carried out according to the device and method described herein, so that thousands to millions of sequencing reactions can be carried out and / or processed simultaneously, and the resulting digital data can be analyzed in multiple bioinformatics processing pipelines, etc., together with the innovative sequencing device and process disclosed herein.

[0025] For example, in one aspect, such as that relating to the innovative sequencing device presented herein, an appropriately configured sequencing platform can be provided as a field effect transistor (FFT) including a chemical reaction layer, such as one for use in performing hybridization and / or sequencing reactions. Specifically, such a field effect transistor (FFT) can be fabricated on a primary structure, such as a wafer, e.g., a silicon wafer. In various cases, this primary structure can include one or more additional structures, such as an insulator material layer, e.g., in a stacked configuration. For example, the insulator material can be included on top of the silicon wafer primary structure and can be an inorganic material, such as silicon oxide, e.g., silicon dioxide, or silicon nitride, or an organic material, such as polyimide, BCB, or other similar materials.

[0026] The primary structure and / or insulator layer may include additional structures including one or more conductive sources and / or conductive drains, spaced apart from one another by a space and embedded in the primary structure and / or insulator material layer, and / or may be flush with the top and / or bottom surfaces of the insulator to form top and / or bottom gates. In various cases, the structure, e.g., a silicon wafer structure, may further include or be otherwise associated with an integrated circuit, such as a processor, e.g., a microprocessor, for processing generated data, e.g., data derived by a sensor, e.g., data derived as a result of a sequencing reaction, near the gate region. Thus, multiple structures may be configured as or otherwise include integrated circuits and / or may exist as an ASIC, a structured ASIC, or an FPGA.

[0027] In particular, these structures may be configured as complementary metal oxide semiconductors (CMOS), which may be configured as chemically sensitive FET sensors including one or more reactive regions, such as conductive source, conductive drain, and / or gate regions, which may themselves include micro- or nanochannel, chamber, and / or well configurations, and the sensors may be adapted to communicate with a processor. For example, a FET may include a CMOS configuration having or otherwise associated with an integrated circuit fabricated on a silicon wafer further including an insulator layer having embedded therein a conductive source and a conductive drain, which may be made of a metal such as damascene copper. In various instances, CMOS and related structures may include a surface, e.g., a top surface, which may include channels and / or chambers for forming reaction wells, where the reaction well surfaces are configured to extend from a conductive source to a conductive drain and may be adapted to receive various reagents that aid in carrying out biochemical reactions, such as DNA or RNA hybridization and / or sequencing reactions.

[0028] In some cases, the surface and / or channel and / or chamber may comprise one-dimensional transistor material, two-dimensional transistor material, three-dimensional transistor material, etc. In various cases, one-dimensional (1D) transistor material may be included, and the 1D material may be comprised of carbon nanotubes or semiconductor nanowires, which may be formed as sheets or channels in various cases, and / or may comprise nanopores in various cases, although in many cases nanopores are not included or required. In various cases, two-dimensional (2D) transistor material may be included, and the 2D material may comprise graphene layers, silicene, molybdenum disulfide, black phosphorus, and / or metal dichalcogenides. Three-dimensional (3D) configurations may also exist. In various cases, the surface and / or channel may comprise a dielectric layer. Additionally, in various cases, a reactive layer, e.g., an oxide layer, may be provided on the surface and / or within the channel and / or chamber, such as by being laminated or otherwise deposited on the 1D layer, the 2D layer, e.g., graphene, or the 3D layer. Such an oxide layer may be a silicon oxide, such as aluminum oxide or silicon dioxide. In various cases, a passivation layer may be provided on the surface and / or within the channel and / or chamber, such as by being laminated or otherwise deposited on the 1D layer, the 2D layer, e.g., graphene, or the 3D layer, and / or on the associated reactive layer on the surface and / or channel and / or chamber.

[0029] In certain cases, the primary and / or secondary and / or tertiary structures may be fabricated or otherwise configured to include chamber or well structures within and / or on the surface, e.g., in a manner that forms a reaction region. For example, the well structures may be disposed on a portion of the surface, e.g., the outer surface, of the primary and / or secondary and / or tertiary structures. In some cases, the well structures may be configured as macrochambers or nanochambers, may be formed on top of or otherwise include at least a portion of a 1D material, a 2D material, e.g., graphene, and / or a 3D material, and / or may additionally include a reaction layer, e.g., an oxide and / or a passivation layer. In various cases, the chamber and / or well structures may define openings, such as openings that allow access to the interior of the chamber, e.g., allowing direct contact with the 1D, e.g., carbon nanotube or nanowire, 2D, e.g., graphene, or 3D surface and / or channel and / or chamber. In certain cases, the chambers and / or wells may be dimensioned to be microchambers or nanochambers.

[0030] Thus, a further aspect of the present disclosure is a biosensor for performing nucleic acid sequencing reactions, etc. The biosensor includes a CMOS structure, which may be configured as a chemically sensitive FET sensor and may include a metal containing a source and drain, for example, a damascene copper source and / or drain, which further includes a surface, such as a reaction region including a 1D or 2D stacked surface, e.g., a graphene stacked surface, or a 3D surface, extending from the source to the drain. Specifically, the reaction region may include or be configured as a well or chamber structure that may be located on a portion of the outer surface of a 1D or 2D stacked well. In such cases, the well structure may be configured to define an opening that allows direct contact of nanotubes, nanowires, and / or graphene with the surface of the well or chamber. In various cases, an oxide layer and / or a passivation layer may be provided within or on the chamber surface. Thus, in some cases, a chemically sensitive transistor, such as a field effect transistor (FFT), may be provided that includes one or more nanowells or microwells for carrying out the sequencing reaction.

[0031] In some embodiments, the chemically sensitive field-effect transistor may include multiple wells and may be configured as an array, e.g., a sensor array. One or more such arrays may be utilized to detect the presence and / or changes in concentration of various analyte types in a wide variety of chemical and / or biological processes, including DNA and / or RNA hybridization and / or DNA or RNA sequencing reactions. For example, the devices described herein and / or systems including the devices may be utilized in methods for analyzing biological or chemical materials, such as whole genome sequencing and / or analysis, genotyping, microarray analysis, panel analysis, exome analysis, microbiome analysis, and / or clinical analyses, such as cancer analysis, NIPT analysis, and / or UCS analysis.

[0032] Thus, in certain embodiments, graphene FET (gFET) arrays may be utilized to facilitate DNA and / or RNA sequencing and processing techniques, such as in genetic analysis pipelines, as described herein. For example, CMOS FETs, e.g., graphene FET (gFET) arrays, may be configured to include reaction wells containing reaction layers adapted to detect changes in hydrogen ion concentration (pH), other analyte concentration changes, and / or binding events associated with chemical processes, such as those associated with DNA or RNA synthesis, such as within a gated reaction chamber or well of a gFET-based sensor. Such chemically sensitive field-effect transistors may include or be adapted to be associated with one or more integrated circuits, and / or may be adapted to increase the measurement sensitivity and / or accuracy of the sensor and / or associated array, such as by including one or more surfaces within a reaction chamber or well having at least one surface layered with 1D and / or 2D and / or 3D materials, dielectric or reaction layers, passivation layers, etc.

[0033] Accordingly, aspects of the present disclosure may include one or more integrated circuits, which may be formed of one or more sets of hardwired digital logic circuits, such as when the sets of hardwired digital logic circuits are interconnected by a plurality of physical electrical interconnects, and which may be adapted to participate in DNA or RNA hybridization and / or sequencing reactions, e.g., performing and / or detecting primary processing, and / or which may be further adapted to process the results of the primary processing, such as in one or more secondary and / or tertiary processing steps. In such cases, the integrated circuit may include inputs, such as via one or more of the plurality of physical electrical interconnects, to interface with an electronic data generating source, such as a sequencing CMOS FET and / or next generation sequencer of the present disclosure, which electronic data generating source is configured to generate such data, e.g., in the form of a plurality of sequenced segments, e.g., reads, of genomic data. In particular cases, the one or more integrated circuits may include a set of hardwired digital logic circuits configured to execute secondary and / or tertiary processing analysis pipelines on generated reads of genomic data, and thus may be connected to an electronic data generating source, such as through one or more of the associated interconnects.

[0034] In such cases, the hardwired digital logic of the integrated circuit and / or associated interconnects may be configured to be capable of receiving one or more reads of genomic data, e.g., from an electronic data source. In certain cases, one or more of the hardwired digital logic circuits may be configured as a set of processing engines, such as where each processing engine is formed of a subset of the hardwired digital logic circuits, and configured to perform one or more steps in a sequencing and / or analysis pipeline, e.g., on multiple reads of genomic data. In such cases, each subset of the hardwired digital logic may, in some cases, be hardwired to perform one or more steps in the sequencing and / or analysis pipeline. However, as indicated above, one or more of the steps in the sequencing and / or analysis pipeline may be configured to be implemented in software, such as where the software and / or hardware are adapted to operate in an optimized manner with respect to each other.

[0035] Thus, in various instances, a plurality of hardwired digital logic circuits are provided, the hardwired digital logic circuits arranged as a set of processing engines, one or more of which may include one or more of a sequencing module, a mapping module, an alignment module, a sorting module, a variant calling module, and / or one or more tertiary processing modules, as described herein. For example, in various embodiments, one or more of the processing engines may include a mapping module, which may be in a wired configuration and may be further configured to communicate with or otherwise be associated with memory on the device, e.g., via a suitably configured interconnect, to access indexes including one or more of the gene reference sequences, one or more reads of the generated sequencing data, and / or splice site indexes (e.g., in the case of RNA sequencing), and to utilize those indexes and reads to perform one or more mapping operations.

[0036] In particular, a suitably configured processing engine may include, or may otherwise be adapted as, a mapping module for performing one or more mapping operations, including, for example, accessing indices of one or more genetic reference sequences from a memory, such as by one or more of a plurality of physical electrical interconnects, to map a plurality of reads to one or more segments of the one or more genetic reference sequences. Additionally, in various embodiments, one or more of the processing engines may include an alignment module, which may be wired and configured to access one or more genetic reference sequences from a memory, such as by one or more of a plurality of physical electrical interconnects, to align a plurality of reads to one or more segments of the one or more genetic reference sequences.

[0037] Further, in various embodiments, one or more of the processing engines may include a sorting module, which may be wired and configured to access one or more aligned reads from a memory, such as via one or more of a plurality of physical electrical interconnects, to sort each aligned read, such as according to one or more locations in one or more genetic reference sequences. In such cases, one or more of the plurality of physical electrical interconnects may include an output from an integrated circuit, such as for communicating result data from the mapping module and / or alignment module and / or sorting module. Furthermore, in certain embodiments, as indicated above, one or more of the processing engines may be configured to interact with various software-implemented processing functions, such as via one or more interconnects, e.g., multiple physical electrical interconnects, for performing one or more steps in an analytical pipeline, including performing one or more RNA and / or DNA sequencing and / or variant calling protocols.

[0038] In various cases, one or more integrated circuits may include a master controller to establish a wired configuration for each subset of hardwired digital logic circuits to perform one or more of the mapping, alignment, and / or sorting functions, which may be configured as one or more steps in a sequence analysis pipeline and / or may include performing one or more aspects of the sequencing and / or variant calling functions. Furthermore, in various embodiments, one or more integrated circuits disclosed herein may not be non-volatile, such as when they are configured as a field programmable gate array (FPGA) with hardwired digital logic circuits, where the wired configuration may be established at the time of manufacturing the integrated circuit. In various other embodiments, the integrated circuit may be configured as an application-specific integrated circuit (ASIC) with hardwired digital logic circuits. In various other embodiments, the integrated circuit may be configured as a structured application-specific integrated circuit (structured ASIC) with hardwired digital logic circuits.

[0039] In some cases, one or more integrated circuits, e.g., CMOS FET sequencing and / or biosensors, and / or one or more associated memories, may be housed on an expansion card such as a peripheral component interconnect (PCI) card; for example, in various embodiments, the integrated circuits of the present disclosure may be chips with PCIe cards. In various cases, the integrated circuits and / or chips may be components within a sequencer, such as an automated sequencer utilizing FET sensors and / or NGS, and / or in other embodiments, the integrated circuits and / or expansion cards may be accessible via the internet, e.g., via the cloud. Furthermore, in some cases, the memory may be volatile random access memory (RAM) or DRAM.

[0040]

[0006] Thus, in one aspect, an apparatus is provided for performing one or more steps of a sequence analysis pipeline on genetic data, such as genetic data including one or more of a genetic reference sequence, an index of one or more genetic reference sequences, an index of one or more splice sites, e.g., an index or table of annotated splice sites, and / or a plurality of reads of genetic data, e.g., DNA or RNA. In various cases, the apparatus may include an integrated circuit, which may include one or more, e.g., a set, of hardwired digital logic circuits, which may be interconnected by one or more physical electrical interconnects, etc. In some cases, one or more of the plurality of physical electrical interconnects may include an input, such as for receiving a plurality of reads of genomic data, such as from a sequencing device as disclosed herein. Additionally, the set of hardwired digital logic circuits may be further configured to access one or more gene reference sequences and / or annotation splice site indexes via one of a plurality of physical electrical interconnects, and to map multiple reads of DNA and / or RNA to one or more segments of the one or more gene reference sequences according to the one or more indexes, etc.

[0041] In various embodiments, the index may include one or more hash tables, such as a primary hash table and / or a secondary hash table and / or a splice site table. For example, a primary hash table may be included, and in such cases, a set of hardwired digital logic circuits may be configured to do one or more of the following: extract one or more seeds of genetic data from a plurality of reads of the genetic data; perform a primary hash function on the one or more seeds to generate a lookup address for each of the one or more seeds of genetic data; access the primary hash table using the lookup address to provide a location within one or more genetic reference sequences for each of the one or more seeds of genetic data. In various cases, the one or more seeds of genetic data may have a fixed number of nucleotides.

[0042] Further, in various embodiments, the index may include a secondary hash table, e.g., the set of hardwired digital logic circuits configured for at least one of: extending at least one of the one or more seeds with additional neighboring nucleotides to produce at least one extended seed of the genetic data; performing a hash function, e.g., a secondary hash function, on the at least one extended seed of the genetic data to generate a second lookup address for the at least one extended seed; and accessing the secondary hash table, e.g., using the second lookup address, to provide a location within the one or more genetic reference sequences for each of the at least one extended seed of the genetic data. In various instances, the secondary hash function may be performed by the set of hardwired digital logic circuits, such as when the primary hash table returns an extension record that instructs the set of hardwired digital logic circuits to extend at least one of the one or more seeds with additional neighboring nucleotides. In some cases, the extension record may specify the number of additional neighboring nucleotides by which at least one or more seeds are extended and / or the manner in which the seeds are extended, e.g., extending each end of the seed equally by an even number "x" of nucleotides.

[0043] Furthermore, as is known, DNA encodes genes. However, for a gene to be expressed, its genetic code must be transcribed and translated into a protein. Specifically, a gene can be transcribed in the cell nucleus by the enzyme RNA polymerase into messenger RNA (mRNA) transcripts or other types of RNA (e.g., transfer RNA). The intermediate RNA transcript is a single-stranded copy of the gene, except that thymine (T) bases in DNA are transcribed into uracil (U) bases in RNA. However, immediately after this copy is produced, the copy's sequence contains both various intron copies and exon copies, where the various intron copies usually need to be excised, for example, by a spliceosome, leaving only the exon copies to be joined at "splice sites" (which are not immediately apparent after ligation) to form the codon region. The spliced ​​mRNA containing the coding region is then transported from the cell nucleus to ribosomes, which decode the mRNA into proteins, with each group of three RNA nucleotides forming a codon that codes for one amino acid. During the decoding process, strings of amino acids are strung together, linked together and glycosylated, to form proteins that make up the body's cells, tissues, and organs. In this way, genes in DNA act as the initial instructions for making proteins.

[0044] Thus, because DNA contains both coding regions, e.g., exons, and non-coding regions, e.g., introns, mapping and / or aligning and / or sorting RNA back to gene precursors in genomic DNA can be complicated. Specifically, each gene often exists on one strand of a double-stranded DNA duplex as a series of exons (coding segments) separated by introns (non-coding segments). Some genes have only a single exon, but most have several exons (separated by introns), and some have hundreds or thousands of exons. Exons are generally several hundred nucleotides long but can be as short as a single nucleotide or tens or hundreds of thousands of nucleotides long. Introns are usually several thousand nucleotides long, with some exceeding one million nucleotides. Thus, when mapping, aligning, and / or sorting from RNA, for example, spliced ​​mRNA, portions of the spliced ​​mRNA may come from different regions of DNA that may be separated from each other by one or two, or even a million or more nucleotides, which greatly complicates RNA processing.

[0045] However, certain aspects of the present disclosure overcome these challenges through methods described herein, thus enabling fast and accurate transcriptome-wide RNA sequencing, mapping, alignment, and / or sorting. More specifically, when RNA processing is involved, the index may include one or more tables, such as a hash table or other index, which includes or is otherwise associated with a table that allows for rapid lookup of various known or determined splice sites used by biological systems when transcribing RNA from DNA, as described in detail below. In such cases, the RNA-aware mapper / aligner may be configured to process such splice sites and consider RNA sequence reads that correspond to segments of transcribed and spliced ​​RNA, such as when the read spans one or more splice sites, which means that with respect to the DNA-oriented reference genome, a first portion of the read should come from and map to a first exon, a second portion of the read should map to a second exon, etc. Thus, the index may include or otherwise be associated with one or more splice site tables, and a set of hardwired digital logic circuits may be configured to utilize said splice site data to perform one or more of: determine and / or extract one or more seeds of genetic data, e.g., RNA data, from a plurality of reads of the genetic RNA data; perform a function, e.g., a hash function, on the one or more seeds of the genetic RNA data, etc., to generate a lookup address for each of the one or more seeds; and access the hash table using the lookup address to provide a location within one or more genetic reference sequences for each of the one or more seeds of the genetic RNA data.

[0046] Additionally, in one aspect, an apparatus is provided for performing one or more steps of a sequence analysis pipeline on genetic sequence data, e.g., either DNA or RNA, where the genetic sequence data includes one or more of one or more genetic reference sequences, which may include both exons and introns, one or more genetic reference sequence indices and / or annotated splice site indices, and multiple reads of genomic data. In various cases, the apparatus may include an integrated circuit, which may include one or more, e.g., a set, of hardwired digital logic circuits, which may be interconnected by one or more physical electrical interconnects, etc. In some cases, one or more of the multiple physical electrical interconnects may include an input, e.g., for receiving multiple reads of genomic data, which may have been pre-processed as described herein to be mapped. Additionally, the set of hardwired digital logic circuits may be further configured to access one or more genetic reference sequences via one of a plurality of physical electrical interconnects to receive position information, e.g., from a mapper, that designates one or more segments of the one or more reference sequences, and to align multiple reads to one or more segments of the one or more genetic reference sequences.

[0047] Thus, in various cases, the wired configuration of the set of hardwired digital logic circuits further includes a wavefront processor configured to align multiple reads of DNA or RNA genetic data to one or more segments of one or more genetic reference sequences, and may be formed with the wired configuration of the set of hardwired digital logic circuits. In some embodiments, the wavefront processor may be configured to process an array of cells of an alignment matrix, such as a matrix defined by a subset of the set of hardwired digital logic circuits. For example, in some cases, the alignment matrix may define a first axis, e.g., representing one of the multiple reads, and a second axis, e.g., representing one or more segments of the one or more genetic reference sequences. In such cases, the wavefront processor may be configured to generate a wavefront pattern of cells extending across the array of cells from the first axis to the second axis, and may further be configured to generate a score for each cell, e.g., in the wavefront pattern of cells, which may represent a degree of match between one of the multiple reads and one of the segments of the one or more genetic reference sequences.

[0048] In such cases, the wavefront processor may be further configured to manipulate the wavefront pattern of the cell across the alignment matrix such that the highest score may be at the center of the wavefront pattern of the cell. Additionally, in various embodiments, the wavefront processor may be further configured to backtrace one or more, e.g., all, locations in the scored wavefront pattern of the cell through previous locations in the alignment matrix, follow one or more, e.g., all, of the backtraced paths until convergence is generated, and generate a CIGAR sequence from this convergence based on the backtracing.

[0049] In some embodiments, the wired configuration of the set of hardwired digital logic circuits for aligning multiple reads to one or more segments of one or more genetic reference sequences may include a wired configuration for implementing the Burrows-Wheeler algorithm as described above, e.g., for mapping prior to alignment, and / or for implementing a Smith-Waterman and / or Needleman-Wunsch scoring algorithm. In such cases, the Smith-Waterman and / or Needleman-Wunsch scoring algorithm may be configured to implement scoring parameters that are affected by base quality scores. Further, in some embodiments, the Smith-Waterman scoring algorithm may be an affine Smith-Waterman scoring algorithm.

[0050] In certain embodiments, the device may include an integrated circuit, which may include one or more, e.g., sets, of hardwired digital logic circuits, which may be interconnected by one or more physical electrical interconnects, etc. In some of these cases, one or more of the plurality of physical electrical interconnects may include an input, such as for receiving a plurality of reads of genomic data as described herein to be mapped and / or aligned, which reads may have been pre-processed. Additionally, the set of hardwired digital logic circuits may be further configured to access one or more genetic reference sequences via one of the plurality of physical electrical interconnects, receive position information designating one or more segments of the one or more reference sequences, such as from a mapper and / or aligner, and sort the plurality of reads into one or more segments of the one or more genetic reference sequences.

[0051] Thus, in one aspect, a method for sequencing genetic material may be provided, for example, to produce electronic genetic data. In certain cases, the method involves the use of a next-generation sequencer, as generally described herein and known in the art, for sequencing genomic DNA and / or RNA derived from genomic DNA. In other cases, the method involves the use of a next-generation sequencer, modified as described herein, for sequencing genomic DNA and / or RNA derived from genomic DNA. In further cases, the method involves the use of a field-effect transistor and / or CMOS sequencer, e.g., an on-chip sequencer, as described in detail below, for sequencing genomic DNA and / or RNA derived from genomic DNA. In various cases, the genetic material, once produced, can be converted into an electronic format, e.g., a digital format, which can be streamed or otherwise transferred to one or more of the pipeline modules described herein.

[0052] Additionally, once electronic, e.g., analog or digital, genetic data, such as sequencing data, is received, another aspect of the present disclosure is directed to running a sequence analysis pipeline on such genetic sequence data. The genetic data may include one or more genetic reference sequences, one or more indices of the one or more genetic reference sequences, and / or a list of one or more annotated splice sites associated therewith (e.g., in the case of RNA sequencing), and / or multiple reads of genomic data (e.g., DNA and / or RNA). The method may include one or more of receiving, accessing, mapping, aligning, and / or sorting various repeats of the genetic sequence data. For example, in some embodiments, the method may include receiving one or more of multiple reads of genomic data as input from an electronic data source to an integrated circuit, where each read of the genomic data may include a sequence of nucleotides. In such cases, the integrated circuit may be formed of a set of hardwired digital logic circuits, such as those interconnected by a plurality of physical electrical interconnects, which may include one or more of the plurality of physical electrical interconnects with inputs.

[0053] The method may further include accessing, by one or more integrated circuits over multiple physical electrical interconnects from memory, indices of one or more gene reference sequences and / or, in the case of RNA sequencing, annotated splice sites. Specifically, if annotated splice sites are provided to a mapper engine, the annotated splice sites can be utilized to improve mapping sensitivity. In such cases, a list of annotated sites can be loaded into memory so that it is accessible by the mapper engine to assist in mapping RNA genetic material. Advantageously, the annotated sites can be formatted into a table, e.g., a hash table, or an index that can be associated with a table, for easy access by the mapper engine. Thus, the method may include mapping, by a first subset of hardwired digital logic circuits of the integrated circuit, a plurality of gene reads, e.g., DNA or RNA reads, to one or more segments of one or more gene reference sequences. Additionally, the method may include accessing, by one or more integrated circuits over a plurality of physical electrical interconnects from the memory, one or more mapped reads and / or genetic reference sequences, and aligning, by a second subset of hardwired digital logic circuits of the integrated circuits, the plurality of reads, e.g., the mapped reads, to one or more segments of the one or more genetic reference sequences.

[0054] In various embodiments, the method may additionally include accessing the aligned reads by the integrated circuit at one or more of the plurality of physical electrical interconnects from memory. In such cases, the method may include sorting the aligned reads according to their location among the one or more genetic reference sequences by a third subset of hardwired digital logic circuits of the integrated circuit. In some cases, the method may further include outputting result data from the mapping and / or alignment and / or sorting, such as at one or more of the plurality of physical electrical interconnects of the integrated circuit, where the result data includes the locations of the mapped and / or aligned and / or sorted reads.

[0055] Additionally, once the genetic data has been generated and / or processed, e.g., in one or more secondary processing protocols, such as by being mapped, aligned, and / or sorted, e.g., to determine how genetic sequence data from a subject differs from one or more reference sequences, such as to produce one or more variant call files, further aspects of the present disclosure may be directed to performing one or more other analytical functions on the generated and / or processed genetic data for further processing, e.g., tertiary processing, etc. For example, a system may be configured for further processing of the generated and / or secondarily processed data, such as by running the system through one or more tertiary processing pipelines, such as one or more of a genomic pipeline, an epigenomic pipeline, a metagenomic pipeline, a joint genotyping, a MuTect2 pipeline, or other tertiary processing pipelines, such as by the devices and methods disclosed herein. Specifically, in various instances, additional layers of processing may be provided for disease diagnosis, therapeutic treatment, and / or prevention, etc., including NIPT, NICU, cancer, LDT, AgBio, and other such disease diagnosis, prevention, and / or treatment, etc., that utilize data generated by one or more of the primary and / or secondary and / or tertiary pipelines of the present disclosure. Thus, the devices and methods disclosed herein may be used to generate genetic sequence data, which may then be used to generate one or more variant call files and / or other associated data, which may be further subject to execution of other tertiary processing pipelines in accordance with the devices and methods disclosed herein, for specific and / or general disease diagnosis, and preventative and / or therapeutic treatment and / or developmental therapy, etc.

[0056] Accordingly, in various instances, implementations of various aspects of the present disclosure may include, but are not limited to, apparatuses, systems, and methods including one or more features as described in detail herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform the operations described herein. Similarly, computer systems and / or networks are described, which may include one or more processors and / or one or more memories directly or remotely coupled to one or more processors. Accordingly, computer-implemented methods consistent with one or more implementations of the present subject matter may be performed by one or more data processors present in a single computing system or in multiple computing systems, such as one or more computer clusters. Such multiple computing systems may be connected, such as via one or more connections, including but not limited to connections over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), a direct connection between one or more of the multiple computing systems, etc., and may exchange data and / or commands or other instructions, etc. The memory, which may include a computer-readable storage medium, may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein.

[0057] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While some features of the presently disclosed subject matter are described for illustrative purposes in the context of an enterprise resource software system or other business software solution or architecture, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter.

[0058] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several aspects of the subject matter disclosed herein and together with the description, help to explain some of the principles associated with the disclosed implementations. [Brief explanation of the drawings]

[0059] [Figure 1] FIG. 1 shows an RNA read showing one or more splice sites and overlapping splice sites, and a seed that intersects the splice site of the read. [Figure 2] FIG. 11 is a diagram of another exemplary RNA read, showing that short (L-base) seeds can be configured to fit more easily into short exons and accommodate short exon overhangs or exon segments truncated by edits such as SNPs. [Figure 3] FIG. 10 illustrates exemplary reference bins within the search range of a successfully mapped K-base seed that can be queried in a fixed seed hash table using, for example, an L-base seed. [Figure 4] FIG. 10 shows a comparison of left and right lead locations for suture locations. [Figure 5] FIG. 1 shows an abstract alignment rectangle with concatenated query sequences on the vertical axis and concatenated reference sequences on the horizontal axis. [Figure 6] FIG. 1 illustrates an apparatus according to an implementation of the present disclosure. [Figure 7] FIG. 1 illustrates another apparatus according to an alternative implementation of the present disclosure. [Figure 8] FIG. 1 is a block diagram of a genomics infrastructure for on-site and / or cloud-based genomics processing and analysis. [Figure 9] FIG. 9 is a block diagram of the local and / or cloud-based computing capabilities of FIG. 8 for a genomics infrastructure for on-site and / or cloud-based genomics processing and analysis. [Figure 10] FIG. 10 is a block diagram of FIG. 9 showing in more detail the computing capabilities for a genomics infrastructure for on-site and / or cloud-based genomics processing and analysis. [Figure 11] FIG. 9 is a block diagram of FIG. 8 showing in more detail third-party analysis capabilities for a genomics infrastructure for on-site and / or cloud-based genomics processing and analysis. [Figure 12] FIG. 1 is a block diagram illustrating a hybrid cloud configuration. [Figure 13] FIG. 13 is a block diagram illustrating the block diagram of FIG. 12 in more detail, showing a hybrid cloud configuration. [Figure 14] FIG. 14 is a block diagram illustrating the block diagram of FIG. 13 in more detail, showing a hybrid cloud configuration. [Figure 15] FIG. 1 is a block diagram illustrating a primary analysis pipeline, a secondary analysis pipeline, and / or a tertiary analysis pipeline as presented herein. [Figure 16] 1 is a flow diagram of the analysis pipeline of the present disclosure. [Figure 17] FIG. 1 illustrates an exemplary design and fabrication of an integrated circuit. [Figure 18] FIG. 1 is a block diagram of a hardware processor architecture according to an implementation of the present disclosure. [Figure 19] FIG. 2 is a block diagram of a hardware processor architecture according to another implementation of the present disclosure. [Figure 20] FIG. 1 shows a gene sequence analysis pipeline. [Figure 21] FIG. 1 illustrates the process steps using a gene sequence analysis hardware platform. DETAILED DESCRIPTION OF THE INVENTION

[0060] Wherever practical, like reference numerals refer to like structures, features, or elements.

[0061] To address these and possibly other problems with currently available solutions, methods, systems, articles of manufacture, etc. consistent with one or more implementations of the current subject matter can provide, among other potential advantages, a sequence analysis apparatus for running a sequence analysis pipeline on genetic sequence data.

[0062] The following provides details of various implementations of a sequencing platform, a sequence analysis pipeline, as well as a system for executing one or more tertiary processing protocols.

[0063] In its most basic form, the body consists of cells; cells form tissues; tissues form organs; organs form systems; and these systems function together to ensure that the body operates to sustain the life of an individual. Thus, the cells of the body are the building blocks of life. More specifically, each cell has a nucleus, and within each cell's nucleus are chromosomes. Chromosomes are formed from deoxyribonucleic acid, which has an organized but twisted double helix structure. DNA itself consists of two opposing but complementary strands of nucleotides; these nucleotides comprise genes that code for proteins that give cells their structure and mediate the function and control of the body's tissues and organs. Essentially, proteins perform much of the work of cells in maintaining the body's normal processes and functions.

[0064] Given the multitude of components of the body and the complexity of how they interact with each other to maintain the body's various processes and functions, there are numerous ways in which the body can become abnormal at any one of these various levels. For example, in one such case, there may be an abnormality in the way a particular gene encodes a given protein, which, depending on the protein and the nature of the abnormality, may result in the development of a disease state.

[0065] Thus, determining a subject's genetic makeup can be extremely useful in diagnosing, preventing, and / or treating such disease conditions. For example, once a person's genetic makeup, e.g., genomic makeup, is known, it can be used for diagnostic purposes and / or to determine whether the person has or is likely to develop a disease condition, and thus can be used for prevention. Similarly, knowledge of a person's genome can be useful in determining various potential treatments, such as drugs, that can or cannot be used in preventative or treatment regimens without causing harm to the user. In various instances, knowledge of a person's genome can also be utilized to determine the efficacy of drugs and / or predict and / or identify problematic side effects of such drug use. In some cases, knowledge of a person's genome can be used to produce designer drugs, such as drugs tailored and optimized according to a person's particular genetic makeup. Specifically, in one instance, designed proteins or nucleotide sequences can be tailored to an individual's unique genetic characteristics to turn gene transcription off or on to over- or under-produce proteins, thereby ameliorating a disease condition.

[0066] Thus, in some cases, the goal of bioinformatics processing is to determine people's individual genomes, which can be used in gene discovery protocols and for prevention and / or treatment plans to improve the lives of each particular individual and humanity as a whole. Furthermore, knowledge of an individual's genome can be used in drug discovery and / or FDA testing, etc., to better and more precisely predict what drugs, if any, are likely to be effective for an individual and / or have adverse side effects, such as by analyzing the individual's genome and / or protein profiles derived from the genome and comparing them with the predicted biological responses from administration of such drugs.

[0067] Such genomics and bioinformatics processes typically involve three clearly defined, but typically separate, stages of information processing. The first stage involves DNA / RNA sequencing, in which a subject's DNA / RNA is obtained and subjected to various processes to convert the subject's genetic code into machine-readable digital code, e.g., a FASTQ file. The second stage involves using the subject's generated digital genetic code to determine the individual's genetic makeup, e.g., to determine the individual's genome nucleotide sequence and / or variant call file, e.g., to determine how the individual's genome differs from the sequences of one or more reference genomes. And the third stage involves performing one or more analyses on the subject's genetic makeup to determine therapeutically useful information therefrom. These may be referred to, in order, as primary, secondary, and tertiary processes, respectively.

[0068] Preliminarily, e.g., in Phase I, or primary processing, genetic material must be preprocessed, e.g., via nucleotide sequencing, to derive usable gene sequence data. Sequencing of nucleic acids, such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), is an important part of biological discovery. Such detection is useful for a variety of purposes and is often used in scientific research and medical advances. For example, the fields of genomics and bioinformatics involve the application of information technology and computer science to the fields of genetics and / or molecular biology. Specifically, bioinformatics techniques, such as those described herein, can be applied to generate, process, and analyze various genomic data from individuals and the like to determine qualitative and quantitative information about the data, which can then be used by various practitioners in developing individual and / or holistic diagnostic, preventative, and / or therapeutic methods for detecting, preventing, and / or at least alleviating disease states, thereby improving the safety, quality, and effectiveness of medical care for individuals and / or society.

[0069] Generally, DNA / RNA analysis techniques, such as for genetic diagnosis, involve nucleic acid hybridization and detection. For example, various typical hybridization and detection techniques include the following steps: For genetic analysis, a subject's RNA or DNA sample to be analyzed can be isolated and immobilized on a substrate, and a probe for a known gene sequence, e.g., a disease marker, can be labeled and washed across the substrate. If the disease marker is present, a binding event, e.g., hybridization, occurs; because the probe is labeled, the hybridization event may or may not be detected, thereby indicating the presence or absence of the disease marker in the subject's sample. Alternatively, as noted above, if the hybridization reaction occurs next to a reaction layer configured to detect reactants and / or by-products of the reaction, such as in an appropriately configured FET device, it is not necessary to utilize a labeled probe.

[0070] Typically, for nucleotide sequencing, the unknown nucleic acid sequence to be identified, such as a single strand of DNA and / or RNA from a subject, is first isolated, amplified, and immobilized on a substrate. Next, a known nucleic acid labeled with an identifiable tag is contacted with the unknown nucleic acid sequence in the presence of a polymerase. When hybridization occurs, the labeled nucleic acid binds to complementary bases in the unknown sequence immobilized on the surface of the substrate. The binding event can then be detected, for example, optically or electrically. These steps are then repeated until the entire DNA sample has been completely sequenced.

[0071] Generally, these steps are performed manually or through an automatic sequencer such as a next-generation sequencer (NGS), and thousands to millions of sequences can be simultaneously produced in next-generation sequencing processing.However, as presented herein, a direct label-free system for DNA and / or RNA sequencing, such as on a computer chip, such as a complementary metal oxide semiconductor (CMOS) chip, is presented, and for example, various components of the sequencer or the entire detection device can be embodied in or otherwise associated with a semiconductor chip.As presented herein, such a system allows for the seamless integration of primary processing, secondary processing, and / or tertiary processing, such as within the same semiconductor chip set.

[0072] More specifically, regardless of the type of sequencing device utilized, a typical sequencing procedure involves obtaining a biological sample from a subject via venipuncture, hair, or the like, and processing the sample to isolate genetic material therefrom. Once isolated, if the genetic sample is DNA, the DNA may be denatured to separate the strands. This step may not be necessary when processing RNA, as RNA is already single-stranded. The isolated DNA and / or RNA, or portions thereof, can then be amplified, for example, via polymerase chain reaction (PCR), to construct a library of replicated strands ready to be sequenced and read, such as by an automated sequencer, which is configured to read the replicated strands, for example, by synthesis, thereby determining the nucleotide sequence comprising the DNA and / or RNA. Furthermore, in various instances, such as when constructing a library of replicated and amplified strands, it may be useful to provide overcoverage when preprocessing a given portion of DNA and / or RNA. Performing this overcoverage, for example, using PCR, may be expensive, requiring more sample preparation resources and time, but often increases the likelihood of a more accurate final result.

[0073] Once a library of replicated DNA / RNA strands is generated, they can be fed into an automated sequencer, such as an NGS, which can then read the strands, such as by synthesis, and determine the nucleotide sequence from the strands. For example, replicated single-stranded DNA or RNA can be attached to glass beads and inserted into a test vessel, such as an array. All components needed to replicate the complementary strand, including labeled nucleotides, are also added to the vessel, but in order. For example, all "A," "C," "G," and "T" that can be labeled are added either one at a time or all together to see which of the nucleotides, if labeled, will bind at location 1 of the single-stranded DNA or RNA.

[0074] After each addition, in a labeled model, light, e.g., a laser, is shone onto the array. If the mixture fluoresces, an image is produced indicating which nucleotide has bound to the target location. In a non-labeled model, a binding event can be detected by a change in resistance, such as at a gate near the reaction layer where the replicated single-stranded DNA or RNA containing glass beads is placed, e.g., a solution gate. More specifically, if nucleotides are added one at a time, a change in fluorescence or resistance is observed, indicating a binding event. If no binding event occurs, the test vessel can be washed, and the procedure is repeated until the appropriate one of the four nucleotides binds to its complement at the target location, indicating a change in conditions. If all four nucleotides are added simultaneously, each can be labeled with a different fluorescent indicator, and the nucleotide that binds to its complement at the target location can be determined by the color of the fluorescence, etc. This greatly accelerates the synthesis process.

[0075] If a binding event occurs, the mixture is then washed, and the synthesis steps are repeated for location 2. For example, a labeled or otherwise marked nucleotide "A" can be added to the reaction mixture to determine whether the complement at location 1 in the bound template molecule being sequenced is "A." If so, the labeled "A" reactant binds to the template sequence with its complement and thus fluoresces, after which the entire sample is washed to remove any excess nucleotide reactant. If a binding event occurs, the bound nucleotide is not washed. This process is repeated for all locations and all nucleotides until all oversampled nucleic acid segments, e.g., reads, have been sequenced and data collected. Alternatively, if all four nucleotides, each labeled with a different fluorescent indicator, are added simultaneously, only one nucleotide will bind to its complement at the location of interest while the others are washed; after the vessel is washed, a laser can be shone on the vessel, and which nucleotide bound to its complement can be determined, such as by the color of the fluorescence. However, when a CMOS FET sensor is utilized as described below, a binding event can be detected by a change in conductivity that occurs near an appropriately configured gate or other reaction region.

[0076] Specifically, due in part to the need to use optically detectable, for example fluorescent, labels in the sequencing reaction being carried out, the required equipment for carrying out such high-throughput sequencing may tend to be large, expensive, time-consuming and non-portable.For this reason, a new approach for direct label-free detection of DNA and / or RNA sequencing is proposed herein.For example, in various embodiments, an improved method for carrying out NGS processing is provided, while in other embodiments, an improved method and device for nucleic acid sequencing and / or processing are provided, which do not necessarily involve NGS.For example, in certain cases, a detection method based on the use of various electronic analysis devices is proposed herein.Such direct electronic detection method has several advantages over typical NGS platforms.

[0077] More specifically, sensors and / or detection devices as disclosed herein can be integrated into the substrate itself, such as by utilizing biosystem-on-a-chip devices, such as complementary metal-oxide semiconductor devices (CMOS). Specifically, when using CMOS devices in gene detection, output signals representing hybridization events, such as output signals for either hybridization and / or nucleic acid sequencing, can be acquired and processed directly on the microchip itself. In such cases, automated recognition can be achieved in real time and at lower costs than those currently achievable using typical NGS processes. Moreover, standard CMOS substrate devices can be utilized for such electronic detection, making the process simple, inexpensive, fast, and portable.

[0078] For example, for next-generation sequencing to become widely used as a diagnostic in the medical industry, sequencing equipment needs to be mass-produced with a high degree of reliability, portability, and economy. One way to achieve this is to adapt DNA / RNA sequencing in a way that fully utilizes the manufacturing infrastructure created for computer chips, such as complementary metal-oxide semiconductor (CMOS) chip fabrication, which represents the current pinnacle of large-scale, high-quality, low-cost manufacturing of advanced technologies. To achieve this, ideally, the entire sensing device of the sequencer could be embodied in a standard semiconductor chip, such as one manufactured in the same fabrication facility used for logic and memory chips.

[0079] Therefore, in another aspect of the present disclosure, a field effect transistor (FET) is provided herein, which can be fabricated on or otherwise associated with a CMOS chip configured for use in performing one or more of DNA / RNA sequencing and / or hybridization reactions. Such a FET can include a gate, a channel region connecting a source terminal and a drain terminal, and an insulating barrier that can be configured to separate the gate from the channel. The optimal operation of such a FET depends on controlling the channel conductivity, and therefore on controlling the drain current, such as by applying a voltage between the gate terminal and the source terminal.

[0080] For high speed applications and to enhance sensor sensitivity, the FETs presented herein are designed to operate at a gate voltage (V GS) can operate in a manner that rapidly responds to fluctuations in the gate and the channel. However, this requires a short gate and high-speed carriers in the channel. In light of this, the FET sensors of the present disclosure for use in nucleic acid hybridization and / or sequencing reactions, etc., are configured to have channels that can be very thin vertically and / or horizontally to enable high-speed carrier conduction and higher sensor sensitivity and accuracy, thereby providing the sensors of the present disclosure with specific advantages for nucleic acid sequencing reactions. Thus, the devices, systems, and methods utilizing them provided herein are ideal for performing genomics analyses and applications, such as for nucleic acid sequencing and / or genetic diagnostics.

[0081] Thus, one aspect of the present disclosure is a chemically sensitive transistor, such as a field effect transistor (FET), designed for analyzing biological or chemical substances, which solves many of the current problems associated with nucleic acid sequencing and genetic diagnosis. Such a FET can be fabricated on a primary structure, such as a wafer, e.g., a silicon wafer. In various cases, the primary structure can include one or more additional structures, such as an insulator material layer, e.g., in a stacked configuration. For example, the insulator material can be included on top of the primary structure and can be an inorganic material, such as silicon oxide, e.g., silicon dioxide, or silicon nitride, or an organic material, such as polyimide, BCB, or other similar materials.

[0082] For example, the primary structure and secondary structure including the insulator layer may include additional structures including one or more conductive sources and / or conductive drains, separated from one another by a space and embedded in the primary structure and / or the insulator material layer, and / or may be flush with the top surface of the insulator. In various cases, the structure may further include or be otherwise associated with a processor, such as for processing generated data, such as data derived by a sensor. Thus, the structure may be configured as or otherwise include an integrated circuit, such as those described herein, and / or may be an ASIC, structured ASIC, or FPGA.

[0083] In certain cases, the structure may be configured as a complementary metal-oxide semiconductor (CMOS), which may then be configured as a chemically sensitive FET including one or more of a conductive source, a conductive drain, a channel or well, and / or a processor. For example, the FET may include a CMOS structure having an integrated circuit fabricated on a silicon wafer, which further includes an insulator layer including a conductive source and a conductive drain embedded therein, the source terminal and drain terminal being made of a metal such as a damascene copper source and a damascene copper drain. In various cases, the structure may include a surface, e.g., a top surface, which may include a channel; for example, the surface and / or channel may be configured to extend from the conductive source to the conductive drain, thereby forming a reaction zone.

[0084] In some cases, the surface and / or channel may comprise one-dimensional transistor material, two-dimensional transistor material, three-dimensional transistor material, etc. In various cases, one-dimensional (1D) transistor material may be included, and the 1D material may comprise carbon nanotubes or semiconductor nanowires. In other cases, the chamber and / or channel is comprised of one-dimensional transistor material, such as one that comprises one or more carbon nanotubes and / or semiconductor nanowires, such as a sheet of semiconductor nanowires.

[0085] In certain cases, two-dimensional (2D) transistor materials may be included; for example, the 2D materials may be one or two atoms thick and may extend in a plane. In such cases, the 2D materials may be graphene, graphene (an allotrope of carbon consisting of a lattice of benzene rings connected by acetylene bonds), borophene (an allotrope of boron), germanene (an allotrope of germanium), germanane (another allotrope of germanium), silicene (an allotrope of silicon), stanene (an allotrope of tin), phosphorene (an allotrope of phosphorus sometimes called black phosphorus), or monoatomic layers of metals such as palladium or rhodium; molybdenum disulfide (MoS2 sometimes called molybdenite), tungsten diselenide (WS2), or other metals such as silicon dioxide (SiC). e2), tungsten disulfide (WS2), or others; MXenes (transition metal carbides and / or nitrides, typically of the formula Mn+1Xn, where M is a transition metal and X is carbon and / or nitrogen), such as Ti2C, V2C, Nb2C, Ti3C2, Ti3CN, Nb4C3, or Ta4C3 (additional MXenes may terminate with O, OH, or F to produce small bandgap semiconductors); or organometallic compounds, such as NiHITP (Ni3(2,3,6,7,10,11-hexaaminotriphenylene)2); or 2D supracrystals (a supracrystal is defined as the periodic atomic structure above, where the atoms typically found in the structure section are replaced by their symmetrical complexes). It should be noted that transition metal dichalcogenides can comprise, in ratio, one atom of any transition metal (Sc, Ti, V, Cr, Mn, Fe, Co, Ni, Cu, Zn, Y, Zr, Nb, Mo, Tc, Ru, Rh, Pd, Ag, Cd, Hf, Ta, W, Re, Os, Ir, Pt, Au, Hg, Rt, Db, Sg, Bh, Mt, Ds, or Rg) bonded to any two atoms of chalcogen (S, Se, or Te).In particular cases, the 2D material may include one or more of a graphene layer, silicene, molybdenum disulfide, black phosphorus, and / or a metal dichalcogenide. In various cases, three-dimensional (3D) material may be included at the surface and / or the channel may include a dielectric layer.

[0086] Additionally, in various cases, a reactive layer, e.g., an oxide layer, may be provided on the surface and / or channel, such as overlaid or otherwise deposited on the 1D, 2D, e.g., graphene, or 3D layer. Such an oxide layer may be a silicon oxide, such as aluminum oxide or silicon dioxide. In various cases, a passivation layer, such as overlaid or otherwise deposited on the 1D, 2D, e.g., graphene, or 3D layer, may be provided on the surface and / or channel and / or on an associated reactive layer on the surface and / or channel.

[0087] In certain cases, the primary and / or secondary structures may be fabricated or otherwise configured to include chambers or wells within and / or adjacent to the surfaces. For example, a well structure may be disposed on a surface, e.g., a portion of an exterior surface, of the primary and / or secondary structures. In some cases, the well structure may be formed on top of or otherwise include at least a portion of 1D, 2D, e.g., graphene, and / or 3D material, and / or may additionally include a reactive layer, e.g., an oxide layer, and / or a passivation layer. In various cases, the chamber and / or well structure may define an opening, such as an opening that allows access to the interior of the chamber, such as allowing direct contact of 1D, e.g., carbon nanotubes or nanowires, or 2D, e.g., graphene, with the surface and / or channel.

[0088] Thus, in various embodiments, the present disclosure is directed to a biosensor. The biosensor includes a CMOS structure that may include a metal-containing source, e.g., a damascene copper source, and a metal-containing drain, e.g., a damascene copper drain; a 1D or 2D stacked surface or channel, e.g., graphene stacked, extending from the source terminal to the drain terminal; and a well or chamber structure that may be disposed on a portion of the outer surface of the 1D, 2D, or 3D stacked well structure. In such cases, the well structure may be configured to define an opening that allows direct contact of nanotubes, nanowires, and / or graphene with the well or chamber surface. In various cases, an oxide layer and / or passivation layer may be provided within or in contact with the chamber surface. Thus, in some cases, a chemically sensitive transistor, such as a field-effect transistor (FET), may be provided that includes one or more nanowells or microwells.

[0089] In some embodiments, the chemically sensitive field-effect transistor may include multiple wells and may be configured as an array, e.g., a sensor array. Thus, a system may include an array of wells containing one or more, e.g., multiple, sensors, each of which includes a chemically sensitive field-effect transistor having a conductive source, a conductive drain, and a reactive surface or channel extending from the conductive source to the conductive drain. Such an array or arrays may be utilized for detecting the presence and / or concentration changes of various analyte types in a wide variety of chemical and / or biological processes, including DNA / RNA hybridization and / or sequencing reactions. For example, the devices described herein and / or systems including same may be utilized in methods for disease diagnosis and / or analysis of biological materials or chemicals, such as for genome-wide analysis, genome-type analysis, microarray analysis, panel analysis, exome analysis, microbiome analysis, and / or clinical analyses, such as cancer analysis, NIPT analysis, and / or UCS analysis.

[0090] In certain embodiments, the FET may be a graphene FET (gFET) array as described herein and may be utilized to facilitate DNA / RNA sequencing and / or hybridization techniques, such as based on monitoring changes in hydrogen ion concentration (pH), changes in the concentration of other analytes, and / or binding events associated with chemical processes related to DNA / RNA synthesis, such as within gated reaction chambers or wells of a gFET-based sensor. For example, chemically sensitive field-effect transistors may be configured as CMOS biosensors and / or adapted to improve the measurement sensitivity and / or accuracy of the sensor and / or associated array, such as by including one or more surfaces or wells having 1D and / or 2D and / or 3D material-layered surfaces, dielectric or reactive layers, passivation layers, etc. For example, in certain embodiments, chemically sensitive graphene field effect transistors (gFETs), such as gFETs, having a CMOS structure are provided, and the gFET sensors, e.g., biosensors, may include oxide and / or passivation layers, such as layers provided on the surfaces of wells or chambers, to improve the sensitivity and / or accuracy of measurements of the sensors and / or associated arrays. The oxide layer, when present, may be comprised of aluminum oxide, silicon oxide, silicon dioxide, etc.

[0091] The system may further include one or more of a fluidic component, such as for performing a reaction, a circuit component, such as for running a reaction process, and / or a computing component, such as for controlling and / or processing a reaction process. For example, a fluidic component may be included, where the fluidic component is configured to control one or more flows of samples across the array and / or one or more chambers of the array. Specifically, in various embodiments, the system includes multiple reaction locations, such as surfaces or wells, and also includes multiple sensors and / or multiple channels, and further includes one or more fluid sources containing fluids having multiple samples and / or analytes delivered to the one or more surfaces and / or wells for performing one or more reactions therein. In some cases, a mechanism for generating one or more electric and / or magnetic fields may also be included.

[0092] The system may additionally include circuit components, e.g., a sampling and hold circuit, an address decoder, a bias circuit, and / or at least one analog-to-digital converter. For example, the sampling and hold circuit may be configured to hold an analog value of a voltage to be applied to a selected column and / or row line of an array of a device of the present disclosure, such as during a read period. Additionally, the address decoder may be configured to generate column and / or row select signals for a column and / or row of the array to access a sensor having a given address within the array. The bias circuit may include bias components that may be coupled to one or more surfaces and / or chambers of the array and may be configured to apply lead and / or bias voltages to selected chemically sensitive field-effect transistors of the array, e.g., to the gate terminals of the transistors. The analog-to-digital converter may be configured to convert the analog value to a digital value.

[0093] Computing components may also be included, for example, the computing components may include one or more processors, such as a signal processor; a base calling module configured to determine one or more bases of one or more reads of the sequenced nucleic acid; a mapping module configured to generate one or more seeds from one or more reads of the sequenced data and perform a mapping function on the one or more seeds and / or reads; an alignment module configured to perform an alignment function on one or more mapped reads; a sorting module configured to perform a sorting function on one or more mapped and / or aligned reads; and / or a variant calling module configured to perform a variant calling function on one or more mapped, aligned, and / or sorted reads. In specific cases, the base calling module of the base calling module may be configured to correct multiple signals for phase and signal loss, normalize to a key, and / or generate multiple corrected base calls for each flow at each sensor to produce multiple sequencing reads. In various embodiments, the device and / or system may include at least one reference electrode.

[0094] Specifically, the system can be configured to perform a sequencing reaction. In such cases, the FET sequencing device can include an array of sensors having one or more associated chemically sensitive field-effect transistors. Such transistors can include cascode transistors having one or more of a source terminal, a drain terminal, or a gate terminal. In such cases, the source terminal of the transistor can be directly or indirectly connected to the drain terminal of the chemically sensitive field-effect transistor. In some cases, a one-dimensional or two-dimensional channel can be included and can extend from the source terminal to the drain terminal; for example, the 1D channel material can be a carbon nanotube or nanowire, and the 2D channel material can be made of graphene, silicene, phosphorene, molybdenum disulfide, and metal dichalcogenide. The device can further be configured to include a plurality of row or column lines coupled to the sensors in the array of sensors. In such a case, each column line among the plurality of column lines may be directly or indirectly connected to or otherwise coupled to drain terminals of transistors, e.g., cascode transistors, of corresponding plurality of pixels in the array, and similarly, each row line among the plurality of row lines may be directly or indirectly connected to or otherwise coupled to source terminals of transistors, e.g., cascode transistors, of corresponding plurality of sensors in the array.

[0095] In some cases, multiple source and drain terminals may be included with multiple reactive surfaces, e.g., channel regions, extending therebetween, e.g., each channel region comprising one, two, or three dimensions of material. In such cases, multiple first and / or second conductive layers may be coupled to first and second source / drain terminals of the chemically-sensitive field-effect transistors in respective columns and rows in the array. In addition, control circuitry may be provided and coupled to the multiple column and row lines, e.g., for reading selected sensors connected to the selected column lines and / or selected row lines. The circuitry may also include bias components, e.g., configured to apply a read voltage to the selected row lines and / or to apply a bias voltage to, e.g., gate terminals of transistors, e.g., FETs and / or cascode transistors, of the selected sensors. In certain embodiments, bias circuitry may be coupled to one or more chambers of the array and configured to apply a read bias to selected chemically-sensitive field-effect transistors via the conductive column and / or row lines. In particular, the bias circuit may be configured to apply a read voltage to the lines of the selected row, such as during a read period, and / or to apply a bias voltage to the gate terminals of transistors, e.g., cascode transistors.

[0096] A sensing circuit may be included and coupled to the array for sensing charges coupled to one or more of the gate configurations of selected chemically-sensitive field-effect transistors. The sensing circuit may also be configured to read selected sensors based on sampled voltage levels of selected row and / or column lines. In such cases, the sensing circuit may include one or more of a precharge circuit, such as for precharging selected column lines to a precharge voltage level before a read period, and a sampling circuit, such as for sampling a voltage level at a drain terminal of a selected transistor, e.g., a cascode transistor, during a read period. The sampling circuit may include a sample and hold circuit configured to hold an analog value of the voltage of the selected column line during the read period, and may further include an analog-to-digital converter for converting the analog value to a digital value.

[0097] In another aspect, the 1D, 2D, or 3D FET integrated circuits of the present invention, e.g., the gFETs, sensors, and / or arrays of the present disclosure, can be fabricated, for example, using any suitable complementary metal-oxide semiconductor (CMOS) processing techniques known in the art. In some cases, such CMOS processing techniques can be configured to improve the measurement sensitivity and / or accuracy of the sensors and / or arrays while facilitating very small sensor sizes and densely packed gFET chamber sensor areas. Specifically, the improved fabrication techniques described herein utilizing 1D, 2D, 3D, and / or oxides as reaction layers enable high-speed data acquisition from small sensors to large, densely packed arrays of sensors. In certain embodiments, an ion-selective permeable membrane is included, and the membrane layer may comprise a polymer such as a perfluorosulfone material, a perfluorocarboxylic material, PEEK, PBI, Nafion, and / or PTFE. In some embodiments, the ion-selective permeable membrane may comprise an inorganic material such as an oxide or glass. One or more of the various layers, for example the reactive layer, the passivation layer, and / or the permeable membrane layer, may be fabricated or otherwise applied by spin coating, anodization, PVD, and / or sol-gel processes.

[0098] Thus, the CMOS FET devices described herein may be utilized to sequence nucleic acid samples. In such cases, the nucleic acid sample may be bound to or adjacent to the surface of the reaction zone, e.g., a graphene-coated surface, and serve as a template for DNA / RNA synthesis and sequencing. Once immobilized, the template sequence can then be sequenced and / or analyzed by performing one or more of the following steps: For example, primers, and / or polymerases, e.g., DNA and / or RNA polymerases, and / or one or more substrates, e.g., deoxynucleotide triphosphates dATP, dGTP, dCTP, and dTTP, may be added to the reaction chamber, e.g., sequentially, after the hybridization reaction has begun, to initiate a chain extension reaction. Upon hybridization of an appropriate, e.g., corresponding, substrate to its complement in the template sequence, there is a concomitant change in a distinct, electrically characteristic voltage, e.g., source-drain voltage (Vsd), measured as a result of a new local gating effect. The sensitivity with which a binding event occurs can be amplified when a reactive layer, such as an oxide layer, is included on the 1D, 2D, or 3D surface; for example, the reactive layer can be configured to produce and / or monitor changes in hydrogen ion concentration (pH) or other analyte concentration.

[0099] Thus, there is a characteristic voltage and / or pH concentration change for each strand extension reaction involving an appropriate, e.g., complementary, substrate. For example, as described herein, a field-effect device for nucleic acid sequencing and / or gene detection can be provided in a sample chamber or well of a flow cell, and a sample solution containing, e.g., a polymerase and one or more substrates, e.g., nucleic acids, can be introduced into the sample solution chamber, e.g., via one or more of the system's fluidic components. In various embodiments, a reference electrode can be provided upstream, downstream, or within a fluid in contact with the field-effect device, and / or the source and / or drain terminals themselves can function as electrodes for hybridization detection, etc., and a gate voltage can be applied whenever needed.

[0100] Specifically, in exemplary chain extension reactions such as those described above, if the added substrate is complementary to the base sequence of the target DNA / RNA primer and / or template, a polynucleotide is synthesized. If the added substrate is not complementary to the next available base sequence in the template, hybridization does not occur and there is no chain extension. Because nucleic acids such as DNA and RNA are negatively charged in aqueous solution, hybridization leading to chain extension can additionally be determined by a change in charge density on the reaction surface and / or reaction chamber. Such detection can be improved by being able to detect an increase in ion concentration, such as by detecting a change in pH. Because the substrates are added sequentially, it is easy to determine which nucleotides have bound to the template, thereby facilitating the chain extension reaction. Therefore, as a result of chain extension, the negative charge on the graphene stacked gate surface, the insulating film surface, and / or the sidewall surface of the reaction chamber increases. This increase can then be detected, such as by a change in gate-source voltage and / or ion concentration, as described in detail herein. By determining which substrate addition resulted in a signal in the gate source voltage or a change in pH, the identity of the base sequence of the target nucleic acid can be determined and / or analyzed.

[0101] Specifically, regardless of the sequencing device used, such as NGS and / or FET-based sequencing devices as described herein, this iterative synthesis process continues until the entire DNA / RNA template strand is replicated in the container. Typically, the typical length of the sequence replicated in this manner is about 100 base pairs to about 500 base pairs, such as about 150 base pairs to about 400 base pairs, including about 200 base pairs to about 350 base pairs, such as about 250 base pairs to about 300 base pairs, depending on the sequencing protocol used. Furthermore, the nucleotide length of these template segments can be predetermined, for example, designed, for example, to comply with any specific sequencing machine and / or the protocol that the machine is operated.

[0102] The end result is a readout or read, which consists of replicated DNA / RNA segments, for example, about 100 to about 1000 or more nucleotides in length, labeled in such a way that every single nucleotide in the sequence, e.g., read, is either known by its label or determined and known by a change in a gating characteristic, such as a change in voltage and / or pH. Thus, since the human genome consists of about 3.2 billion base pairs, and various known sequencing protocols typically result in labeled replicated sequences, e.g., reads, of about 100 or 101 bases to about 250 or about 350 or about 400 bases, the total number of segments that need to be sequenced, and consequently the total number of reads generated, can be anywhere from about 10,000,000 to about 40,000,000, such as about 15,000,000 to about 30,000,000, depending on how long the labeled replicated sequences are. Thus, a sequencer may typically generate approximately 30,000,000 reads to cover a genome in one go, such as when the read length is 100 nucleotides. However, as shown herein, due to the condensed nature of the sequencing of the present invention on the chip formats presented herein, much larger read lengths may be achievable, such as 800 bases, 1000 bases, 2500 bases, 5000 bases, and up to 10,000 bases.

[0103] Furthermore, as noted above, in such procedures, it may be useful to oversample DNA / RNA by, for example, about 5x, or about 10x, or about 20x, or about 25x, or about 30x, or about 40x, or about 50x, or about 100x, or about 200x, or about 250x, or about 500x, or about 1000x, or even more than about 5000x, or about 10000x, so the amount of primary processing that needs to be performed and the time it takes to do so can be quite large. For example, with 40x oversampling, various synthesized reads are designed to overlap to some extent, and up to about 1.2 billion reads may need to be synthesized. Typically, most, if not all, of these labeled sequences can be generated in parallel. The end result is that the original biological genetic material is processed by a sequencing protocol, such as those summarized herein, to generate a digital representation of the data, which can then be subjected to a primary processing protocol.

[0104] Specifically, the genetic material of a subject may be replicated and sequenced in such a way that measurable electrical, chemical, radioactive, and / or optical signals are generated, which are then converted into a digital representation of the subject's genetic code, for example, by a sequencer and / or a processing device associated with the sequencer. More specifically, primary processing may include converting images, such as recorded luminescence or other electrical or chemical signal data, into FASTQ file data. This information is then stored in a FASTQ file, which can then be sent for further processing, for example, secondary processing. A typical FASTQ file contains a large collection of reads representing a digitally encoded nucleotide sequence, where each predicted base in the sequence has been called and assigned a probability score that the called base at the indicated location is incorrect.

[0105] In many cases, it may be useful to further process the digitally encoded sequence data obtained from the sequencer and / or sequencing protocol, such as by subjecting the digitally represented data to secondary processing. This secondary processing can be used, for example, to assemble an individual's entire genomic profile, for example, to determine the individual's overall genetic makeup, for example, by sequentially determining each nucleotide of each chromosome to identify the individual's overall genome composition. In such processing, the individual's genome can be assembled, for example, by comparison with a reference genome, such as one or more genomes obtained from the Human Genome Project, to determine how the individual's genetic makeup differs from that of the reference. This process is generally known as variant calling. Because the difference between any one person's DNA / RNA and another person's DNA / RNA is only 1 base pair in 1,000 base pairs, such variant calling can be very labor-intensive and time-consuming.

[0106] Therefore, in a typical secondary processing protocol, a subject's genetic makeup is assembled by comparison with a reference genome. This comparison involves reconstructing an individual genome from millions of short read sequences and / or comparing the entire DNA and / or RNA of an individual to an exemplary DNA and / or RNA sequence model. In a typical secondary processing protocol, a FASTQ file containing raw sequenced read data is received from a sequencer. For example, in some cases, assuming no oversampling, such as when each read is approximately 100 nucleotides in length, there may be up to 30,000,000 or more reads covering the subject's genome. Therefore, in such cases, in order to compare the subject's DNA / RNA genome with the DNA / RNA of a standard reference genome, it is necessary to determine where each of these reads maps to the reference genome, for example, how they are aligned with each other, and / or how each read can be sorted by chromosome order, in order to determine where each read belongs and which chromosome it belongs to. One or more of these functions can be performed before performing the variant calling function on the entire full-length sequence. Once it is determined where each read belongs in the genome, the full-length genetic sequence can be determined, and then the differences between the genetic code of the subject and the genetic code of the reference can be evaluated.

[0107] Since the human genome is over 3 billion base pairs in length, efficient automated sequencing protocols and machines have been developed to achieve the sequencing of such DNA / RNA genomes within a timeframe that can be clinically useful. Such innovations in automated sequencing have led to the ability to sequence entire genomes within a few hours to a few days, depending on the number of genomes being sequenced, the amount of oversampling involved, and the number of processing resources allocated to the job. Therefore, given these advances in sequencing, a large amount of sequencing data can be generated in a relatively short period of time. However, the result of these advances is an increasing bottleneck in the secondary processing stage. To help overcome this bottleneck, various software-based algorithms, such as those described herein, have been developed to help speed up the process of assembling a subject's DNA and / or RNA to be sequenced, such as through a standard-based assembly process.

[0108] For example, standard-based assembly is a typical secondary processing assembly protocol that involves comparing the sequenced genomic DNA and / or RNA of a subject with one or more standards, such as the genomic DNA and / or RNA of a known standard sequence.Various algorithms have been developed to help speed up this process.These algorithms typically involve some variation of mapping, aligning, and / or sorting millions of reads received from a digital file, such as a FASTQ file, transmitted by a sequencer to determine where each specific read corresponds or is otherwise located on each chromosome.Often, the common feature behind the function of these various algorithms is the use of indexes and / or arrays to speed up the processing function of the algorithm.

[0109] For example, in terms of mapping, a large amount of, for example, all sequenced reads can be processed to determine the possible positions that these reads can potentially align to in a reference genome.One method that can be used for this purpose is to directly compare the reads with the reference genome to find all matching locations.Another method is to use prefix or suffix arrays, or to build prefix or suffix trees, to map reads to various locations in the reference DNA / RNA genome.A typical algorithm that is useful in performing such functions is the Burrows-Wheeler transformation, which is used to map selected reads to the reference using a compression formula that compresses repetitive sequences in data.

[0110] Another method is to utilize a hash table, for example, where a selected subset of reads, a k-mer of length "k", e.g., a seed, is placed as a key in the hash table, the reference sequence is decomposed into equivalent k-mer parts, and the parts and their positions are inserted into the hash table by an algorithm at their positions in the hash table to which they map according to a hash function. A typical algorithm for performing this function is "BLAST", the Basic Local Alignment Search Tool. Such hash table-based programs compare a query nucleotide or protein sequence with one or more standard reference sequence databases and calculate the statistical significance of the match. In such a scheme, it can be determined where any given read is likely to be located relative to the reference genome. These algorithms are useful because they require less memory, fewer lookups, and therefore less processing resources and time to perform the function than in other cases, such as when a subject's genome is assembled by direct comparison without using these algorithms.

[0111] In addition, in cases where a read can be mapped to multiple locations in a genome, and this location is actually the location where the read was actually derived, such as by sequencing by the original sequencing protocol, an alignment function can be performed to determine from among all possible locations on the genome that a given read can be mapped. This function can be performed on a large number of genomic reads, and an ordered sequence of nucleotide bases can be obtained, representing a portion or the entire genetic sequence of the subject's DNA and / or RNA. Along with the ordered genetic sequence, for any given nucleotide location, a score can be assigned to each nucleotide location, representing the probability that the nucleotide predicted to be at that location, such as "A", "C", "G", "T" (or "U"), is actually the nucleotide belonging to that assigned location. Typical algorithms for performing the alignment function are Needleman-Wunsch and Smith-Waterman. In either case, these algorithms perform sequence alignment between the sequence of the subject's query genome DNA and / or RNA sequence and the sequence of the reference genome sequence, thereby comparing segments of selected portions of possible length instead of comparing the entire genome sequences with each other.

[0112] Once reads have been assigned a location relative to, for example, a reference genome, which may include identifying which chromosome the read belongs to and / or the offset of the read from the start of the chromosome, the reads can be sorted by location. This may allow downstream analysis to take advantage of the oversampling described above. All reads that overlap with a given location in the genome will be adjacent to each other after sorting, and they can be organized into pileups that can be easily examined to determine whether the majority of them match the reference value. If not, a variant can be flagged.

[0113] While these algorithms and others like them go some way to solving the bottleneck inherent in secondary processing, faster execution times and greater accuracy are still desirable. More specifically, while there have been advances in the generation of raw data, such as generated DNA / RNA sequence data, advances in information technology have not kept pace, creating a bottleneck in data analysis. While this bottleneck has been somewhat alleviated by the development of various algorithms that help accelerate these analyses, such as those described above, new technologies remain needed to handle data generation and acquisition, computation, storage, and / or analysis of such data, particularly as they relate to genome sequence analysis, such as in secondary processing stages.

[0114] For example, using standard NGS technology, sequencing a human genome can take several hours, or up to about a day, and using standard protocols for performing secondary processing on such acquired genome sequencing data, it can take three days, or even up to a week or more, to process the sequenced data to generate clinically relevant genome sequence information for an individual. Using a variety of different optimized devices, algorithms, methods, and / or systems, the time spent from primary to secondary processing can be reduced to as little as 27 to 48 hours. However, achieving such rapid results typically requires that virtually all generated reads, e.g., 30 million reads of 100 nucleotides each, be processed in parallel and simultaneously. Such parallel processing requires massive processing power, involving significant CPU resources, and still takes a relatively long time.

[0115] Furthermore, in various cases, it is desirable to improve the accuracy of the results. Such improvement in accuracy can be achieved by providing some amount of oversampling of the genome to be sequenced. For example, as explained above, it may be desirable to process the DNA of a subject in such a way that at any given position in the sequence of nucleotides, there is oversampling of that region. As indicated above, it may be desirable to oversample any given region of the genome by up to 10 times, or 15 times, or 20 times, or 25 times, or 30 times, or 40 times, or 50 times, or 100 times, or 250 times, or even 500 times, or 1000 times, or more. However, if the genome is oversampled by, for example, 40 times, the amount of reads to be processed is approximately 30 million x 40 (depending on the length of the reads), and when the entire genome is oversampled by 40 times, approximately 1.2 billion reads must be processed. Thus, such oversampling typically results in higher accuracy, but at the cost of taking longer and requiring more processing resources, as each section of the genome is covered anywhere from 1 to 40 times. Moreover, some oncology applications, such as those in which clinicians seek to distinguish the mutated genomes of cancer cells in the bloodstream from those of healthy cells, may utilize 500-fold, 1000-fold, 5000-fold, or even 10000-fold oversampling.

[0116] Accordingly, the present disclosure is directed to such new technologies that may be implemented in one or a series of genomics and / or bioinformatics protocols, e.g., pipelines, for performing genetic acquisition and / or analysis, such as primary and / or secondary processing, on acquired genome sequencing data or portions thereof. Sequencing data can be obtained directly from an automated high-throughput sequencer system, such as ROCHE's "Sequencing by Synthesis" 454 automated sequencer, ILLUMINA's HiSeq x Ten or Solexia automated sequencer, LIFE TECHNOLOGIES' "Sequencing by Oligonucleotide Ligation and Detection" (SOLiD) or Ion Torrent sequencer, and / or HELICOS GENETIC ANALYSIS SYSTEMS' "Single Molecule Fluorescent Sequencing" sequencer, by a direct link to a sequencing processing unit, or the sequencing data can be obtained directly, such as in sequencing on a chip configuration, such as a graphene-stacked FET sensor, including a CMOS sequencing chip, as described herein. Such sequencing data can also be obtained remotely, such as from a database accessible via the Internet or from other remote locations accessible via wireless communication protocols such as Wi-Fi, Bluetooth, etc.

[0117] In some aspects, these gene acquisition and / or analysis techniques may utilize improved algorithms that can be implemented by software that are less processing-intensive and / or less time-consuming and / or that perform with higher percentage accuracy. For example, in some embodiments, improved devices and methods for producing genetic sequence information, such as in primary processing protocols, and / or improved algorithms for performing secondary processing on genetic sequence information, as disclosed herein, are provided. In various specific embodiments, the improved devices, systems, methods of use thereof, and utilized algorithms are directed to more efficiently and / or accurately performing one or more of the following functions: sequencing, mapping, alignment, and / or sorting, such as for generating and / or analyzing digital representations of DNA / RNA sequence data obtained from a sequencing platform, such as an automated sequencer and / or a sequencer on a chip, such as one of those described above.

[0118] Additionally, in some embodiments, improved algorithms are provided that are directed to more efficiently and / or more accurately performing one or more of the following functions: local realignment, duplicate marking, base quality score recalibration, variant calling, compression, and / or decompression. Further, as described in more detail herein below, in some aspects, these gene production and / or analysis techniques can utilize one or more algorithms, such as improved algorithms that can be implemented by hardware, that perform in a less processing-intensive and / or less time-consuming manner and / or with a higher percentage accuracy than various software implementations for doing the same.

[0119] In certain embodiments, a technology platform for sequencing DNA / RNA to produce genetic sequence data and / or for performing genetic analysis is provided, which may include performing one or more of the following functions: sequencing, mapping, alignment, sorting, local realignment, duplicate marking, base quality score recalibration, variant calling, compression, and / or decompression, and / or may further include tertiary processing protocols as described herein. In some cases, implementations of one or more of these platform functions are for generating a consensus genome sequence for a subject and / or for performing one or more of determining and / or reconstructing; comparing the subject's genome sequence to a reference sequence, e.g., a reference or model gene sequence; determining the manner in which the subject's genomic DNA and / or RNA differs from the reference, e.g., variant calling; and / or for performing tertiary analyses on the subject's genome sequence, such as whole genome analyses, such as genome-wide variation analysis and / or genotyping analysis; gene function analyses; protein function analyses, e.g., protein binding analyses; quantitative genome and / or transcriptome analysis and / or assembly analyses; microarray analyses; panel analyses; exome analyses; microbiome analyses; and / or clinical analyses, such as cancer analyses, NIPT analyses, and / or UCS analyses, as well as for evaluating various diagnostic and / or preventative and / or therapeutic methods.

[0120] In particular, once genetic data has been generated and / or processed, e.g., in one or more primary and / or secondary processing protocols, such as by being mapped, aligned, and / or sorted, e.g., to determine how genetic sequence data from a subject differs from one or more reference sequences, such as to produce one or more variant call files, further aspects of the present disclosure may be directed to performing one or more other analytical functions on the generated and / or processed genetic data for further processing, e.g., tertiary processing, etc. For example, a system may be configured for further processing of the generated and / or secondarily processed data, such as by running the system through one or more tertiary processing pipelines, such as one or more of a genomic pipeline, an epigenomic pipeline, a metagenomic pipeline, a joint genotyping, a MuTect2 pipeline, or other tertiary processing pipelines, such as by the devices and methods disclosed herein. For example, in various instances, additional layers of processing may be provided for disease diagnosis, therapeutic treatment, and / or prevention, etc., including NIPT, NICU, cancer, LDT, AgBio, and other such disease diagnoses, preventions, and / or treatments, etc., that utilize data generated by one or more of the primary and / or secondary and / or tertiary pipelines of the present disclosure. Thus, the devices and methods disclosed herein may be used to generate genetic sequence data, which may then be used to generate one or more variant call files and / or other associated data, which may be further subject to execution of other tertiary processing pipelines in accordance with the devices and methods disclosed herein, for specific and / or general disease diagnoses, and preventative and / or therapeutic treatments and / or developmental therapies, etc.

[0121] Furthermore, in various embodiments, bioinformatics processing regimes as disclosed herein may be utilized to create one or more masks, such as a genomic reference mask, a default mask, a disease mask, and / or an iterative feedback mask, which may be added to a mapper and / or aligner, e.g., along with a reference, where the set of masks is configured to identify particular regions or objects of interest. For example, in one embodiment, the methods and apparatus disclosed herein may be utilized to create genomic reference masks, such as by creating a mask set that may be loaded into a mapper and / or aligner along with a reference, where the mask set is configured to identify regions of high importance and / or relevance to a practitioner or subject and / or to identify regions that are more error-prone. In various embodiments, the mask set can provide intelligent guidance to the mapper and / or aligner, such as which regions of the genome to focus on for quality improvement. Thus, masks may be created in a tiered manner to provide varying levels or iterative guidance based on various specific applications. Each mask may accordingly identify a region of interest and provide a minimum quality target for that region. Additionally, default masks can be used to provide guidance for identified regions of the genome, such as typical "high value" regions. Such regions can include known coding regions, regulatory regions, etc., as well as regions well known to produce errors. Furthermore, disease masks or application-specific masks can be used for mask sets that identify highly important regions, such as regions requiring a very high level of accuracy based on known markers, e.g., Cancer. Still further, iterative feedback masking can be used, such as by adding new ad hoc masks that can be specifically designed by using feedback from a tertiary analysis system (such as Cypher Genomics) that has identified problematic regions based on observed errors or discrepancies.

[0122] As indicated above, in one aspect, one or more of these platform functions, e.g., mapping, alignment, sorting, realignment, duplicate marking, base quality score recalibration, variant calling, one or more tertiary processing modules, compression, and / or decompression functions, are configured for implementation in software. In another embodiment, one or more of these platform functions, e.g., mapping, alignment, sorting, local realignment, duplicate marking, base quality score recalibration, variant calling, tertiary processing, compression, and / or decompression functions, are configured for implementation in hardware.

[0123] Thus, in some cases, methods are presented herein, and the methods involve the execution of an algorithm, such as an algorithm for performing one or more genetic analysis functions, such as mapping, alignment, sorting, realignment, duplicate marking, base quality score recalibration, variant calling, compression, and / or decompression, and the algorithm is optimized according to the manner in which it is to be implemented.Specifically, when the algorithm is to be implemented in a software solution, the algorithm and / or its associated processes are optimized to be executed faster and / or with higher accuracy by the medium.Similarly, when the functions of the algorithm are to be implemented in a hardware solution, the hardware is designed to perform these functions and / or associated processes in a manner that is optimized to be executed faster and / or with higher accuracy by the medium.For example, these methods can be used in iterative variant calling procedures, etc.

[0124] Thus, in one aspect, presented herein are systems, devices, and methods for implementing bioinformatics protocols, such as for performing one or more functions for analyzing genetic data, such as genomic data, via one or more optimized algorithms and / or one or more optimized integrated circuits, e.g., on one or more hardware processing platforms, etc. Thus, in one instance, systems and methods are provided for implementing one or more algorithms for performing one or more steps for analyzing genomic data in a bioinformatics protocol, where, for example, these steps may include performing one or more of mapping, alignment, sorting, local realignment, duplicate marking, base quality score recalibration, variant calling, compression, and / or decompression. In another instance, systems and methods are provided for implementing the functionality of one or more algorithms for performing one or more steps for analyzing genomic data in a bioinformatics protocol, as described herein, where the functionality is implemented on one or more general-purpose processors and / or hardware accelerators, which may or may not be coupled to a supercomputer.

[0125] More specifically, in some cases, a method is provided for performing secondary analysis on data regarding a subject's genetic makeup. In one case, the analysis to be performed can involve a reference-based reconstruction of the subject's genome. For example, reference-based mapping involves the use of a reference genome, which can be generated from sequencing the genomes of a single or multiple individuals, or the reference genome can be an amalgamation of DNA from various people, synthesized in such a way as to produce a prototypical standard reference genome to which any individual's DNA can be compared, for example, to determine and reconstruct the individual's genetic sequence and / or to determine differences between the individual's genetic makeup and that of a standard reference, e.g., for variant calling.

[0126] More specifically, the reason for performing secondary analysis on a subject's sequenced DNA is to determine how the subject's DNA varies from that of a reference. More specifically, to determine one, many, or all of the differences in the subject's nucleotide sequence from that of the reference. For example, the genetic sequence of any two random individuals may differ by one base in 1,000 base pairs, which, considering the more than 3 billion base pairs in the entire genome, amounts to up to 3,000,000 different base pair variations per person. Determining these differences may be useful in tertiary analysis protocols, for example, to predict the likelihood of a disease state occurring due to a genetic abnormality or the like and / or the likelihood of success of a prophylactic or therapeutic approach, based on how the prophylactic or therapeutic approach is expected to interact with the subject's DNA or proteins produced therefrom. In various cases, it may be useful to perform both de novo and reference-based reconstructions of a subject's genome to confirm one result against another and, if desired, to improve the accuracy of variant calling protocols.

[0127] In various cases, as described above, when performing a primary sequencing protocol, it may be useful to generate oversampling for one or more regions of the subject's genome.These regions can be selected based on the subject's condition and / or based on areas known to be highly variant-occurring, areas suspected of variant occurrence, etc., generally based on the entire genome.As shown above, in the basic form of sequencing based on the type of sequencing protocol being performed, sequencing produces readouts (e.g., reads) that are digital representations of the subject's genetic sequence code.The length of these reads is typically designed based on the type of sequencing machine being used. For example, ROCHE's 454 automated sequencer typically produces read lengths from 100 or 150 base pairs to about 1000 base pairs in length; at ILLUMINA, read lengths are typically designed to be from about 100 or 101 base pairs to about 150 base pairs in length for some technologies and 250 base pairs in length for other technologies; at LIFE TECHNOLOGIES, read lengths are typically designed to be about 50 to about 60 base pairs in length for SOLiD technology and 35 to 450 base pairs in length for Ion Torrent technology; and at HELICOS GENETIC ANALYSIS SYSTEMS, read lengths can vary but are typically less than 1000 nucleotides in length.

[0128] However, since the processing of DNA samples required to produce a designed read length of a particular size is labor-intensive and chemical processing-intensive, and since sequencing itself often depends on the function of sequencing machines, there is some possibility that errors may occur throughout the sequencing process, thereby causing abnormalities in the sequenced genome where the errors occur. Such errors may be particularly problematic when the purpose of reconstructing a subject's genome is to determine how the genome or at least a portion thereof varies from a standard or model reference. For example, a mechanical or chemical error that causes a single nucleotide change in, for example, one read relative to another will give a false indication of a variation that does not actually exist. This may result in an incorrect variant call, and may even result in a false indication of a disease state, etc. Therefore, due to the possibility of mechanical, chemical, and / or human errors in the execution of sequencing protocols, in many cases, it is desirable to build redundancy into the analysis system, for example, by oversampling part or the entire genome. More specifically, automated sequencers produce FASTQ files that call the sequence of a read, along with the nucleotide at a given location, along with the probability that the call, e.g., base call, for a given nucleotide at the called location is actually incorrect, and it is often desirable to utilize methods such as oversampling to ensure that base calls made by the sequencing process can be detected and corrected.

[0129] Thus, in practicing the methods described herein, in some cases, a primary sequencing protocol is performed in a manner to produce a sequenced genome, where a portion or the entire genome is oversampled by about 10x, about 15x, about 20x, about 25x, about 30x, about 40x, about 50x or more, etc. Thus, if the read length is designed to be about 50-60 base pairs in length, such as when the oversampling is about 40x, this oversampling may result in about 2 billion to about 2.5 billion reads, or if the read length is about 100 or 101 base pairs in length, the oversampling may result in about 1 billion to about 1.2 billion reads, and if the read length is about 1000 base pairs in length, about 50 million to about 100 million reads may be generated by the sequencer. More specifically, in such a case, with 40-fold oversampling, at any given point in the genome, one would expect there to be 40 reads covering any one location, but a given location could be at the beginning of one read, the middle of another, and the end of another, but would be expected to be covered approximately 40 times.

[0130] Therefore, such oversampling produces a region of the genome that is covered by a large number of reads, for example, overlaps, such as up to 40 reads when oversampling is about 40 times.These at least partial overlaps are useful for determining whether any given variation in any particular read is actually a true genome variation or a mechanical or chemical artifact.Therefore, oversampling can be used to improve the accuracy of reconstructing a subject's genome, especially in cases where the subject's genome is compared with a reference genome to determine the cases where the subject's gene sequence differs from the reference gene sequence.In this manner, as will be described in more detail below, it can be confirmed that any given variation between the reconstructed sequence and the model is actually due to the presence of a true variant, and not due to errors in the initial processing of sample DNA or read alignment software, etc.

[0131] For example, when constructing a genetic sequence of an individual's sequenced DNA, it is necessary to determine which nucleotide goes where in a growing string of nucleotides. To determine which nucleotide goes where, various reads can be organized, and pileups of reads covering overlapping positions can be built. This allows a comparison to be made between all reads covering the same position to more accurately determine whether there is a true variation at any given position, or whether any one read in the pileup may have an error at that position. For example, if only one or two out of 40 reads have a particular nucleotide at position X, all 38 or 39 other reads will agree that a different nucleotide is at that position, and the two aberrant reads can be excluded as being erroneous, at least at this particular position.

[0132] More specifically, when there are a large number of reads generated for any one position in the genome of a subject, there is a high probability that there will be multiple overlaps or pileups for any given nucleotide location.These pileups represent the coverage for any specific position and can be useful for determining the correct sequence of the genome of a subject with higher accuracy.For example, as shown, sequencing produces reads, and in various cases, the reads produced are oversampled, so that various specific reads overlap at various locations.This overlap is useful for determining the true genome of a sample with a high probability of accuracy, etc.

[0133] Thus, the goal may be to additionally scan the entire reference genome multiple times, as described in more detail herein below, to more accurately reconstruct the subject's genome; if it is desirable to determine how the subject's genome differs from a different genome, e.g., a model genome, the use of pileups can more accurately identify errors, such as chemical, mechanical, or read errors, and distinguish them from true variants. More specifically, if the subject has a true variation at location X, the majority of the reads in the pileup should corroborate, e.g., contain, that variation. Statistical analysis procedures, such as those described herein, can then be performed to determine the subject's true genetic sequence along with all variants from the reference genome.

[0134] For example, if a subject's genetic sequence is to be reconstructed with respect to the use of a reference genome, once reads, e.g., a pileup of reads, have been generated, the next step may be to map and / or align and / or sort the reads to one or more reference genomes (e.g., the more exemplary reference genomes are available as models, the better the analysis is likely to be), thereby reconstructing the subject's genome, which, if there is a match, results in a set of reads mapped and / or aligned to the reference genome at all possible locations along the strand, and at each such location, the read is given a probability score about the probability that the read actually belongs to that location.

[0135] Thus, in various instances, once reads are generated, the locations of the reads to be mapped, e.g., the possible positions in the reference genome to which the reads may map, are determined, their order is aligned, and the true gene sequence of the subject's genome can be determined, such as by performing a sort function on the aligned data. Additionally, once the true sample genome is known and compared to the reference genome, variations between the two can be determined, and a list of all variations / deviations between the reference genome and the sample genome is determined and called up. Such variations between the two gene sequences can be due to a number of reasons.

[0136] For example, there may be a single nucleotide polymorphism (SNP) in which one base in the subject's genetic sequence is replaced with another; there may be a more extensive substitution of multiple nucleotides; there may be an insertion or deletion in which one or many bases are added to or removed from the subject's genetic sequence; and / or there may be a structural variant caused, for example, by the crossing over of two chromosomal arms; and / or there may simply be an offset that causes a misalignment in the sequence. In various cases, a variant call file is generated that contains all variations in the subject's genetic sequence relative to the reference sequence. More specifically, in various embodiments, the methods of the present disclosure include generating a variant call file (VCF) that identifies one or more, e.g., all, genetic variants relative to one or more reference genomes of the individual whose DNA was sequenced. A VCF in its basic form is a list of variant locations and variant types, e.g., chromosome 3, location X, where "A" is replaced with "T," etc.

[0137] However, as noted above, to generate such a file, a subject's genome must be sequenced and reassembled before its variants can be determined. However, several problems can arise when attempting to generate such an assembly. As noted above, there can be issues with chemical errors, sequencing machine errors, and / or human errors that arise in the sequencing process. In addition, there can be genetic artifacts that make such reconstruction problematic. For example, a problem with performing such an assembly is that there can be large portions of the genome that are repeated, such as long sections of the genome that contain the same string of nucleotides. Therefore, because no genetic sequence is unique everywhere, it can be difficult to determine where identified reads actually map and align in the genome.

[0138] For example, depending on the sequencing protocol used, shorter or longer reads can be produced. Longer reads are useful in that longer reads are less likely to appear at multiple locations in the genome. Fewer potential locations to evaluate can make the system faster. However, longer reads can be more problematic because they contain true or false variations, such as SNPs, InDels (insertion or deletion), or machine errors, which increase the likelihood that the read will not match the reference genome. On the other hand, shorter reads are useful because they are less likely to cover locations that encode variants. However, the problem with shorter reads is that shorter reads are more likely to appear at multiple locations in the genome, requiring additional processing time and resources to determine which of all possible locations is most likely to be the true location to which the read aligns. Ideally, what can be achieved, such as by practicing the methods disclosed herein, is to be able to produce a variant call file, where a list of the genome to be sequenced (query sequence) is generated that shows where all the variant base pairs are located, ensuring that each variant called is a true variant and not simply a chemical or mechanical read error or other human error.

[0139] Therefore, there are two main possibilities for variation. One possibility is that there is a true variation at the specific position in question, for example, a person's genome actually differs from a reference genome at a specific position, for example, there is a natural variation due to SNPs (single base substitutions), insertions or deletions (of one or more nucleotides in length), and / or there is a structural variant, such as when DNA material from one chromosome crosses over another chromosome or arm, or when a region is copied twice in DNA. Alternatively, variation can be caused by problems in the read data, either through chemical or mechanical errors, sequencer or aligner errors, or other human errors. Therefore, the methods disclosed herein can be used in a way that compensates for these types of errors, and more specifically, to distinguish between errors in chemical, mechanical, or human variations and errors in genuine variations of the sequenced genome. More specifically, the methods, apparatus, and systems utilizing them as described herein have been developed to clearly distinguish between these two types of variation and, therefore, to more reliably ensure the accuracy of any call files generated to correctly identify true variants.

[0140] Furthermore, in various embodiments, once a subject's genome has been reconstructed and / or a VCF has been generated, such data may then be subjected to tertiary processing to interpret the data, such as to determine what the data means with respect to identifying what diseases the person may or may not be susceptible to and / or to determine what treatments or lifestyle changes the subject may wish to use to alleviate and / or prevent a disease condition. For example, the subject's genetic sequences and / or their variant call files may be analyzed to determine clinically relevant genetic markers that indicate the presence or likelihood of a disease condition and / or the potential efficacy of a proposed treatment or prevention for the subject. This data may then be used to provide the subject with one or more treatments or preventions to improve the subject's quality of life, such as treating and / or preventing a disease condition.

[0141] More specifically, medical science and technology have advanced alongside advances in information technology, which have enhanced the ability to accumulate and analyze medical data. Thus, once one or more of an individual's genetic variations have been determined, such variant call file information can be used to reveal medically useful information, which can then be used, for example, using various known statistical analysis models, health-related data, and / or medically useful information, for diagnostic purposes, such as diagnosing a disease or its likelihood and therefore its clinical interpretation (e.g., searching for markers representing disease variants), determining whether a subject should be included or excluded in various clinical trials, and other such purposes. Because the number of disease states caused by genetic anomalies is finite, certain types of tertiary processing variants, such as those known to be symptomatic of a disease state, can be queried, such as by determining whether one or more gene-based disease markers are included in the subject's variant call file.

[0142] As a result, in various instances, the methods disclosed herein may involve analyzing, e.g., scanning, the VCF and / or generated sequence against known disease sequence variants, such as in a database of genetic markers, to identify the presence of genetic markers in the VCF and / or generated sequence and, if present, make a call on the presence or likelihood of a genetically induced disease state. Because there are many known genetic variations and many individuals affected by diseases caused by such variations, in some embodiments, the methods disclosed herein may involve generating one or more databases linking genome-wide sequencing data and / or associated variant call files, e.g., from an individual or multiple individuals, with disease states, and / or searching the generated databases to determine whether a particular subject has a genetic makeup that predisposes them to such a disease state. Such searching may involve comparing an entire genome to one or more other genomes, or comparing a fragment of a genome, such as a fragment containing only a variation, to one or more fragments of one or more other genomes, such as in a reference genome or a database of fragments thereof.

[0143] It is further understood that the genetic sequences to be utilized in these methods may be DNA, ssDNA, RNA, mRNA, rRNA, tRNA, etc. Accordingly, although various references are made throughout this disclosure to various methods and devices for analyzing genomic DNA, in various instances, the systems, devices, and methods disclosed herein are equally suitable for performing their respective functions, e.g., analysis, on all types of genetic material, including, e.g., DNA, ssDNA, RNA, mRNA, rRNA, tRNA, etc. Additionally, in various instances, the methods of the present disclosure may include analyzing generated genetic sequences, e.g., DNA, ssDNA, RNA, mRNA, rRNA, tRNA, etc., from a subject and determining therefrom protein variations likely caused by the genetic sequences and / or determining and / or predicting therefrom the likelihood of a disease state due to errors in protein expression, etc. It should be noted that the gene sequence obtained may represent introns or exons, e.g., the gene sequence may be for only the coding portion of the DNA, e.g., an exome may be obtained and using known processing techniques, only the coding or non-coding regions may be sequenced, which involves more difficult sample preparation procedures but may lead to faster sequencing and / or faster processing times.

[0144] Currently, such steps and analyses described herein are typically performed in various separate and unrelated steps, often utilizing different analytical machines at different locations. Thus, in various embodiments, the methods and systems of the present disclosure are performed by a single device and / or at one location, such as in conjunction with an automated sequencer or other device configured to generate genetic sequence data. In various cases, multiple devices may be utilized at the same location or multiple remote locations, and in some cases, the methods may involve two or more processing units deployed at two or more locations.

[0145] For example, in various embodiments, a pipeline may be provided that includes performing one or more analytical functions as described herein on genomic genetic sequences of one or more individuals, such as data obtained from an automated sequencer in a digital file format, e.g., a FASTQ file format. A typical pipeline to be performed may include sequencing the genetic material, e.g., a portion or the entire genome, of one or more subjects, which may include DNA, ssDNA, RNA, rRNA, tRNA, etc., and / or in some cases, the genetic material may represent coding or non-coding regions of DNA, such as exomes, episomes, etc. The pipeline may include performing one or more base calling and / or error correction operations on the digitized genetic data, etc., and / or may include performing one or more mapping, alignment, and / or sorting functions on the genetic data. In some cases, the pipeline may include performing one or more of realignment, deduplication, base quality score recalibration, reduction and / or compression, and / or decompression on the digitized genetic data. In some cases, the pipeline may include performing a variant calling operation on the genetic data.

[0146] Thus, in various instances, a pipeline of the present disclosure may include one or more modules configured to perform one or more functions on genetic data, e.g., sequenced genetic data, such as base calling and / or error correction operations and / or mapping and / or alignment and / or sorting functions. In various instances, a pipeline may include one or more modules configured to perform one or more of local realignment, redundancy elimination, base quality score recalibration, variant calling, reduction and / or compression, and / or decompression on the genetic data. Many of these modules may be executed either by software or on hardware, or executed remotely, e.g., via software or hardware, such as in the cloud or on a remote server and / or server bank.

[0147] Additionally, many of these steps and / or modules of the pipeline are optional and / or can be arranged in any logical order and / or omitted entirely. For example, the software and / or hardware disclosed herein may or may not include base calling or sequence correction algorithms; for example, there may be concerns that such functions may introduce statistical bias. As a result, systems may or may not include base calling and / or sequence correction functions, respectively, depending on the level of accuracy and / or efficiency desired. As noted above, one or more of the pipeline functions may be utilized in generating a subject's genome sequence, such as through criteria-based genome reconstruction. Also, as noted above, in some cases, the output from the pipeline is a variant call file that indicates some or all of the variants in a genome or portion thereof.

[0148] Thus, as noted above, the output of a sequencing protocol, such as one or more of those described above, is typically a digital representation of the subject's genetic material, such as in a FASTQ file format. However, digitally transcribed autorads may also be used. More specifically, the output from a sequencing protocol may include multiple reads, each read including a sequence, e.g., a string, of nucleotides, where the location of each nucleotide is called, and a quality score represents the probability that the called nucleotide is incorrect. However, the quality of these outputs may be improved by various preprocessing protocols to achieve higher quality scores, and one or more of such protocols may be used in the methods disclosed herein.

[0149] For example, in some cases, raw FASTQ file data may be processed to clean up initial base calls obtained from a sequencer / reader, such as in a primary processing stage, prior to secondary processing, such as that described herein above. Specifically, a sequencer / reader typically analyzes sequencing data, such as fluorescence data, indicating which nucleotides are at which locations and converts the image data into base calls with quality scores, where the quality scores are based on the relative brightness of the fluorescence at each location, for example. To more accurately make appropriate base calls, special algorithms may be utilized, such as in a primary processing stage, to accurately analyze these distinctions in fluorescence. As noted above, this step may be included in a pipeline of steps and may be implemented via software, hardware, or both, but in this case is part of the primary processing platform.

[0150] Additional preprocessing steps can include error correction functions, which can involve using millions to billions of reads from a FASTQ file to attempt to correct a percentage of any machine-generated sequencing errors using available base call and quality score information before any further processing, such as mapping, alignment, and / or sorting functions. For example, the reads in a FASTQ file can be analyzed to determine whether any of the reads contain subsequences that appear in other reads, which can increase confidence that the subsequence in the read is correct due to overlapping coverage. This can be done by building a hash table containing all possible k-mers of a selected length k from each read, and storing each k-mer along with its frequency and the probability of which base immediately follows it. Each read can then be rescanned using the hash table. As each k-mer in a particular read is looked up in the hash table, an evaluation can be made as to whether the base immediately following that k-mer is likely to be correct. If not, the immediately following base can be replaced with the most likely one from the table. The subsequent k-mer of that read will then contain the corrected base as the value at that location, and the process is repeated. This is extremely effective at correcting errors, as oversampling allows for the collection of accurate statistics for predicting what comes next for each k-mer. However, as shown above, such corrections can introduce statistical bias into the system, such as by incorrectly correcting the data, so these steps can be skipped if desired.

[0151] Thus, according to aspects of the present disclosure, in various cases, the method, apparatus, and / or system of the present disclosure may include acquiring read data, either preprocessed or unprocessed, such as by directly acquiring it from an automated sequencer's FASTQ file, and subjecting the acquired data to one or more of mapping, alignment, and / or sorting functions. Performing such functions may be useful, for example, as described above, in various cases, because the sequencing data, e.g., reads, typically generated by various automated sequencers have a length significantly shorter than the entire genome sequence being analyzed. Since the human genome typically has numerous repeat sections with various repeat patterns, there may be numerous positions where any given read sequence may correspond to a segment of the human genome. As a result, due to various repeat sequences in the genome, etc., when considering all the possibilities where a given read may match a genome sequence, the raw read data may not clearly indicate which of these possibilities is actually the correct position from which the data was derived. Therefore, for each read, it is necessary to determine where the read actually maps in the genome. Additionally, it may be useful to determine the sequence alignment of the reads to determine the actual sequence identity of the subject, and / or to determine the chromosomal location for each portion of the sequence.

[0152] In various cases, the methods of the present disclosure may involve mapping, aligning, and / or sorting raw read data from a FASTQ file to find all possible locations to which a given read can be aligned, and / or to determine the actual sequence identity of a subject, and / or to determine the chromosomal location for each portion of the sequence. For example, mapping can be used to map generated reads to a reference genome, thereby finding locations where each read appears to match the genome well, for example, to find all locations where there may be a good score for aligning any given read to the reference genome. Thus, mapping can involve using one or more, for example, all, of the raw or preprocessed reads received from a FASTQ file, comparing those reads to one or more reference genomes, and determining where the reads can match the reference genome. In its basic form, mapping involves finding locations in the reference genome where one or more FASTQ reads obtained from a sequencer appear to match.

[0153] Similarly, alignment can be used to evaluate all the positional candidates of each read against a certain section of reference genome to determine where and how the read sequence best aligns with genome.However, performing alignment can be difficult due to substitution, insertion, deletion, structural variation, etc., which can prevent read from aligning exactly.Therefore, there are several different ways to obtain alignment, but it may be necessary to make changes to the read to do so, and each change that must be made to obtain proper alignment results will result in a lower reliability score.For example, any given read may have substitution, insertion, and / or deletion compared to reference genome, and these variations need to be taken into account when performing alignment.

[0154] Therefore, along with the predicted alignment, a probability score that the predicted alignment is correct can also be given. This score indicates the best alignment among multiple positions that the read may align to for any given read. For example, the alignment score is based on how well a given read matches with a potential mapping position, and can include extending, shortening, and modifying bits and fragments of the read to achieve the best alignment.

[0155] This score reflects all the ways in which the read has been modified to correspond to the reference. For example, to generate an alignment between the read and the reference, it may be necessary to insert one or more gaps into the read, with each gap insertion representing a deletion in the read relative to the reference. Similarly, it may be necessary to make deletions in the read, with each deletion representing an insertion in the read relative to the reference. In addition, various bases may need to be changed, such as by one or more substitutions. Each of these changes is made to more closely align the read to the reference, but each change is accompanied by a decrease in quality score, which is a measure of how well the entire read matches a certain region of the reference. Then, the reliability of the quality score is determined by finding all positions that the read can be mapped to the genome, comparing such quality scores at each position, and selecting the position with the highest score. More specifically, if there are multiple locations with high quality scores, the reliability is low, but if the difference between the best score and the second best score is large, the reliability is high. Finally, all proposed reads and reliability scores are evaluated, and the best match is selected.

[0156] Once the genome has been assigned to a location relative to a reference genome, consisting of identifying which chromosome the read belongs to and the offset of the read from the beginning of that chromosome, the reads can be sorted by location. This allows downstream analysis to take advantage of the various oversampling protocols described herein. All reads that overlap with a given location in the genome may be adjacent to each other after sorting, and they can be pooled and easily examined to determine whether the majority of them match the reference value.

[0157] As shown above, a FASTQ file obtained from a sequencer consists of multiple reads, for example, millions to billions or more, each consisting of a short string of nucleotide sequence data representing a portion or the entire genome of an individual. Generally, mapping involves plotting the reads to all positions in the reference genome where there is a match. For example, depending on the size of the read, there may be one or more positions where the read substantially matches the corresponding sequence on the reference genome. Therefore, mapping and / or other functions disclosed herein can be configured to determine which of all potential positions where one or more reads may match in the reference genome is actually the true position to which the read maps.

[0158] It is possible to compare each read with each location in 3.2 billion reference genomes to determine where, if any, the read matches the reference genome. This can be done, for example, when the read length reaches about 100,000 nucleotides, about 200,000 nucleotides, about 400,000 nucleotides, about 500,000 nucleotides, or even about 1,000,000 nucleotides or more. However, when the read length is significantly shorter, such as when there are 50 million or more reads, for example, 1 billion reads, this process can take a very long time and require a large amount of computing resources. Therefore, there are several methods developed to align FASTQ reads to reference genomes much faster, such as those described herein. For example, as disclosed above, one or more algorithms can be used to map one or more of the reads generated by a sequencer, for example, in a FASTQ file, and match them with a reference genome to determine where the subject's reads may map in the reference genome.

[0159] For example, in various methods, an index of the reference is generated so that reads or portions of reads can be looked up in the index, and an indication of their position in the reference is retrieved to map the read to the reference. Such an index of the reference can be constructed in various formats and queried in various manners. In some methods, the index can include a prefix and / or suffix tree. In various other methods, the index can include a Burrows / Wheeler transform of the reference. In further methods, the index can include one or more hash tables, and a hash function can be performed on one or more portions of the read in an attempt to map the read to the reference. In various cases, one or more of these algorithms can be performed sequentially or simultaneously to precisely determine where one or more reads, for example, a substantial portion or every single read, exactly match the reference genome.

[0160] Each of these algorithms may have advantages and / or disadvantages. For example, prefix and / or suffix trees and / or Burrows / Wheeler transforms may be performed on sequence data in such a way that an index of a reference genome is constructed and / or queried as a tree-like data structure, where starting with a single-base or short subsequence of a read, the subsequence is additionally extended in the read, with each additional extension triggering an access to the index, tracing a path through the tree-like data structure until the subsequence is sufficiently unique, e.g., until an optimal length is achieved, and / or until a leaf node is reached in the tree-like data structure, where the leaf tree node or the last accessed tree node indicates one or more locations in the reference genome from which the read may have originated. Thus, these algorithms typically do not have a fixed length for read subsequences that can be mapped by querying the index. However, hash functions often utilize a comparison unit of a fixed length, which may be the entire length of the read, but is often some subportion of the read in length, this subportion being called a seed. Such seeds may be longer or shorter, but unlike prefix and / or suffix trees and / or Burrows / Wheeler transforms, the seeds of reads utilized in hash functions are typically of a pre-selected fixed length.

[0161] A prefix and / or suffix tree is a data structure constructed from a reference genome, such that each link from a parent node to a child node is labeled with or associated with a nucleotide or sequence of nucleotides, and each path from the root node through various links and nodes traces a path whose associated collective nucleotide sequence matches some continuous subsequence of the reference genome. The node reached by the path is implicitly associated with the reference subsequence traced by the path from the root. Subsequences in the prefix tree proceed from the root node and extend forward in the reference genome, while subsequences in the suffix tree extend backward in the reference genome. Both prefix and suffix trees can be used in a hybrid prefix / suffix algorithm, so subsequences can extend in both directions. Prefix and suffix trees can also include additional links, such as jumping from a node associated with one reference subsequence to another node associated with a shorter reference subsequence.

[0162] For example, a tree-like data structure serving as an index to the reference genome can be queried by tracing a path through the tree corresponding to the subsequence of the mapped read, which is built by adding nucleotides to the subsequence, using the added nucleotide to select the next link for traversing the tree, and going as deep as necessary until a unique sequence is generated. This unique sequence may also be called a seed and may represent a branch and / or root of the sequence tree data structure. Alternatively, the seed may map to multiple locations in the reference genome, as descending the tree may terminate before the cumulative subsequence is completely unique. Specifically, a tree may be built for each starting location relative to the reference genome, and generated reads may be compared to the branches and / or root of the tree, and these sequences may be traversed throughout the tree to find where the read matches in the reference genome. More specifically, reads in a FASTQ file may be compared to the branches and root of the reference tree, and if there is a match, the position of the read in the reference genome may be determined. For example, the sample reads can be walked along the tree until a location is reached whereby it is determined that the accumulated subsequences are sufficiently unique to identify that the reads truly align with a particular location in the standard, e.g., walking the tree until a leaf node is reached.

[0163] However, a disadvantage of such prefix and / or suffix trees is that they are large data structures that must be accessed multiple times as the trees are traversed to map reads to the reference genome. On the other hand, an advantage of a hash table function, as described in more detail herein below, is that once constructed, it typically requires only a single lookup to determine where, if any, a match between the seed and the reference is possible. Prefix and / or suffix trees typically require multiple lookups, e.g., 5, 10, 15, 20, 25, 50, 100, 1000, or more, to determine whether and where a match exists. Furthermore, due to the double helix structure of DNA, a reverse complementary tree may also need to be constructed and searched, since the reverse complementary strand to the reference genome may also need to be found. While the data tree is described above as being constructed from a reference genome that is then compared to reads from the subject's sequenced DNA, it should be understood that the data tree may be constructed first from either or both the reference sequence and the sample reads, and then compared to each other as described above.

[0164] Instead of, or in addition to, utilizing a prefix tree or suffix tree, a Burrows / Wheeler transformation may be performed on the data. For example, the Burrows / Wheeler transformation may be used to store a tree-like data structure abstractly equivalent to a prefix tree and / or a suffix tree in a compact format, such as within the space allocated for storing the reference genome. In various cases, the stored data is not a tree-like structure; rather, the reference sequence data is in a one-dimensional list that may be scrambled into different orders to transform the reference sequence data in a very specific way to enable an accompanying algorithm to search for the reference relative to the sample reads to effectively traverse the "tree." The advantage of the Burrows / Wheeler transformation over, for example, prefix and / or suffix trees is that it typically requires less memory for storage, and the advantage over hash functions is that it supports variable-length seeds, allowing a unique sequence to be determined and searched until a match is found. For example, for prefix / suffix trees, the length of the seed is determined by how many nucleotides a given sequence needs to be unique or to map to a sufficiently small number of reference locations, whereas in hash tables, the seeds are all the same predetermined length. However, a drawback of the Burrows / Wheeler transformation is that it typically requires many lookups, such as two or more lookups for each step down the tree.

[0165] Instead of, or in addition to, utilizing one or both of a prefix / suffix tree and / or a Burrows / Wheeler transformation on the reference genome and the subject's sequence data to find where one maps to the other, another such method involves producing a hash table index and / or performing a hash function. The hash table index may be a large reference structure constructed from the sequence of the reference genome, which may then be compared to one or more portions of the read to determine where one may match to the other. Similarly, a hash table index may be constructed from portions of the read, which may then be compared to one or more sequences of the reference genome, thereby being used to determine where one may match to the other.

[0166] More specifically, in any of the mapping algorithms described herein, such as for implementation in any of the method steps disclosed herein, one or all three mapping algorithms, or others known in the art, can be used in software or hardware to map one or more sequences of a sequenced DNA sample to one or more sequences of one or more reference genomes. As described in more detail below, all of these operations can be performed via software or by being hardwired into an integrated circuit, such as on a chip as part of a circuit board. For example, one or more functions of these algorithms can be embedded on a chip, such as an FPGA (field programmable gate array), an ASIC (application-specific integrated circuit) chip, or a structured ASIC (application-specific integrated circuit) chip, and can be optimized to perform more efficiently due to such hardware implementation.

[0167] In addition, one or more, for example, two or all three, of these mapping functions can form a module, such as a mapping module, which can form part of a system, for example, a pipeline, used in a process to determine the entire or a portion of an individual's actual genome sequence. The output returned from the execution of the mapping function can be a list of possibilities for where one or more, for example, each read, maps to one or more reference genomes. For example, the output of each mapped read can be a list of possible locations, and the read can be mapped to a matching sequence in the reference genome. In various embodiments, an exact match can be found for at least a portion, if not all, of the read, for example, the seed of the read, to the reference. Thus, in various cases, it is not necessary for every portion of every read to exactly match every portion of the reference genome.

[0168] Furthermore, one or all of these functions may be programmed in a manner that allows for exact or rough matching and / or editing, such as editing the results. Accordingly, all of these processes may be configured to perform less exact matching, if desired, according to a preselected degree of dissimilarity, such as 80% match, 85% match, 90% match, 95% match, 99% match, or more. However, as described in more detail herein below, less exact matching may be much more expensive in terms of time and processing power requirements, for example, because it may require that any number of edits, e.g., one or two or three or five or more edits, which may be, for example, SNPs or insertions or deletions of one or more bases, be performed to achieve an acceptable match. Such edits are likely to be more widely used when implementing hashing protocols, or when implementing prefix and / or suffix trees, and / or when performing Burrows / Wheeler transformations.

[0169] Regarding hash tables, hash tables can be generated in many different ways. In one example, a hash table can be constructed by breaking down a reference genome into segments of standard length, e.g., seeds of about 16 to about 30 nucleotides or more, such as about 18 to about 28 nucleotides, formatting them into a searchable table, and creating an index of all reference segments to which sequenced DNA, e.g., one or more reads or portions thereof, can be compared to determine a match. More specifically, a hash table index can be generated by breaking down a reference genome into segments of known uniform length, e.g., seed nucleotide sequences, and storing them in random order in individual cubicles in a reference table. This can be done for a portion or the entire reference genome to build an actual reference index table that can be used to compare portions of one or more reads, such as from a FASTQ file, with portions of the reference genome to determine a match.

[0170] This method is then repeated in a generally similar manner for a portion, e.g., most or all, of the reads in the FASTQ file to generate seeds of an appropriate, e.g., selected, length. For example, the reads in the FASTQ file may be used to generate seeds of a predetermined length, which may be converted to binary form, passed through a hash function, and placed into a hash table index, where the binary form of the seeds may be matched with binary segments of the reference genome to provide a location in the genome where the sample seeds match.

[0171] For example, if a read is approximately 100 bases long, a typical seed may be approximately half or a third of that length, e.g., approximately 27 to 30 bases long. Thus, in such cases, for each read, multiple seeds, e.g., approximately three or four seeds depending on the length of the read and / or the length of the seed, may be generated to cover the read. Each seed may then be converted to binary format and / or fed to a hash table to obtain possible results for the seed's location relative to the reference genome. In such cases, it is not necessary to compare the entire read with every possible location in the entire reference genome; rather, only a portion of the read, e.g., one or more of the generated sample seeds for each read, may be compared with an index containing an equivalent seed portion of the reference genome. Thus, in various cases, the hash table may be configured so that a single memory lookup can typically determine where the sample seed, and therefore the read, is located relative to the reference genome. However, in some cases, it may be desirable to perform a hash function to look up one or more overlapping sections of the seed from a given read. In such cases, the seeds to be generated may be formed in such a way that at least a portion of the sequences of the seeds overlap with one another, which may be useful, for example, in addressing machine and / or human error or differences between the subject's genome and the reference genome, and may facilitate exact matching.

[0172] In some cases, the construction of the hash table and the execution of one or more of the various comparisons are performed by a hash function. A hash function is, in part, a scrambler. A hash function receives an input and provides what appears to be a random order to the input. In this case, the hash function scrambler breaks the reference genome into segments of preselected length and randomly places them in the hash table. The data can then be stored evenly throughout the storage space. Alternatively, the storage space can be segmented and / or the storage in the storage space can be weighted differently. More specifically, a hash function is a function that receives an input and provides a number, such as a binary pattern output, which can usually be random except that the same output is always returned for any one given input. Thus, even if two inputs fed into a hash table are nearly identical, they will not exactly match, and two completely random, different outputs will be returned.

[0173] Furthermore, because genetic material can consist of four basic nucleotides, e.g., "A," "C," "G," and "T" (or "U" in the case of RNA), the individual nucleotides of a sequence, e.g., a reference segment and / or read, or a portion thereof, to be fed into the hash table can be digitized and represented in binary format, e.g., each of the four bases represents a two-bit digital code, e.g., "A" = 00, "C" = 01, "G" = 11, and "T" / "U" = 10. In some cases, it is this binary "seed" value that is randomly placed in the hash table at a known location with a value equal to its binary representation. Thus, the hash function acts to resolve the reference genome into binary representations of the reference seeds and inserts each binary seed data into a random space, e.g., a cubicle, in the hash table based on its numerical value. Along with this digital binary code, e.g., an access key, each cubicle may also contain an actual entry that indicates where in the actual reference genome the segment originated, e.g., the reference location. Thus, the reference location may be a number indicating the location of the original reference seed in the genome. This may also be done for duplicate locations, which are placed in a table in random order but in known positions, such as by a hash function. In this manner, a hash table index may be generated that contains digital binary codes for some or all of multiple segments of one or more reference genomes, which may then be referenced by one or more sequences of genetic material, e.g., one or more reads, or portions thereof, from one or more individuals.

[0174] When implementing the hash table and / or function in software (e.g., where the bit width is twice the number of bases in the seed described above) and / or hardware as a module, such as a module in a pipeline of modules, as mentioned above, the hash table can be constructed such that the binary representation of the reference seed can be of any desired bit width. Since the seed can be long or short, the binary representation can be large or small, but typically the length of the seed can be chosen to be long enough to be unique, but not so long that it is too difficult to find a match between the seed of the genome reference and the seed of the sample read due to errors, variants, etc. For example, as shown above, the human genome consists of approximately 3.1 billion base pairs, and a typical read can be about 100 nucleotides long. Thus, useful seed lengths can be between lengths of less than about 16 or about 18 nucleotides to lengths of about 28 or about 30 nucleotides or more. For example, in some cases, the seed length can be the length of a 20-nucleotide segment. In other cases, the seed length can be the length of a 28-nucleotide segment.

[0175] As a result, if the seed length is a 20-nucleotide segment, each segment can be digitally represented by a 40-bit output, e.g., a 40-bit binary representation of the seed. For example, if two bits are selected to represent each nucleotide, e.g., A=00, C=01, G=10, and T=11, then a 20-nucleotide seed x 2 bits per nucleotide = 40-bit (5-byte) vector, e.g., a number. If the seed length can be 28 nucleotides long, the digital representation, e.g., binary representation, of the seed can be a 56-bit vector. Thus, if the seed length is approximately 28 nucleotides long, 56 bits can be utilized to handle the 28-nucleotide seed length. More specifically, if 56 bits represent a binary form seed of a reference genome randomly arranged in a hash table, an additional 56 bits can be used to digitally represent the seed of a read to be matched against the reference seed. These 56 bits can be passed through a polynomial that converts the 56-bit input to a 56-bit output in a 1:1 correspondence. Performing this operation without increasing or decreasing the number of bits in the output randomizes the storage locations of adjacent input values, so that different seed values ​​are evenly distributed across all potential storage locations. This also serves to minimize collisions between values ​​that hash to the same location. Specifically, in the exemplary hash table implementation described herein, only a portion of the 56 bits is used as a lookup address to select a storage location, and the remaining bits are stored in that location to check for a match. If a hash function were not used, a large number of patterns with the same address bits but different stored bits would have to share the same hash location.

[0176] More specifically, there is a similarity between how a hash table is constructed, e.g., by software and / or hardware randomly placing the seeds of a reference genome into a hash table, and how the hash table is accessed by the seeds of the reads being hashed, so that they both access the table in the same way. Thus, a reference seed and a sample read seed that are the same, e.g., have the same binary code, will end up in the same position, e.g., address, in the table because they access the hash table in the same way, e.g., for the same input pattern. This is the fastest known method for performing pattern matching. Each lookup takes an approximately constant amount of time to perform. This can be contrasted with the Burrows-Wheeler method, which may require many probes per query to find a match (this number can vary depending on how many bits are needed to find a unique pattern), or the binary search method, which requires log2(N) probes (N is the number of seed patterns in the table).

[0177] Furthermore, even if a hash function can decompose a reference genome into segments of seeds of any given length, e.g., 28 base pairs, and then convert the seeds into a digital, e.g., binary, 56-bit representation, it is not necessary that all 56 bits be accessed at exactly the same time or in the same way. For example, the hash function may be implemented in a manner such that the address of each seed specified by a number of less than 56 bits, such as about 20 bits to about 45 bits, about 25 bits to about 40 bits, about 28 bits to about 35 bits, including about 28 bits to about 30 bits, can be used as an initial key or address to access the hash table.

[0178] For example, in some instances, about 26 to about 29 bits may be used as a primary access key for the hash table, leaving about 27 to about 30 bits that may be utilized as a means to double-check the first key; for example, if both the first key and the second key arrive at the same cell in the hash table, it is relatively clear that this location is where the keys belong. Specifically, to save space and reduce the memory requirements and / or processing time of the hash module, such as when the hash table and / or hash function is implemented in hardware, about 26 to about 29 bits representing the primary access key derived from the original 56 bits representing the digitized seed of a particular sequenced read may be utilized by the hash function to construct a primary address, leaving about 27 to about 30 bits that may be used in the double-check method.

[0179] More specifically, in various instances, approximately 26 to 29 bits from the 56 bits representing the binary reference seed may be used to construct a primary address; these designated 26 to 29 bits may then be assigned a randomized location in a hash table, where a location indicating where the reference seed originated may be stored along with the remaining 27 to 30 bits of the seed so that an exact match can be confirmed. A query seed representing a subject's genomic read converted to binary format may also be hashed by the same function in a manner that is also represented by the 29 bits that constitute the primary access key. If the 29 bits representing the reference seed exactly match the 29 bits representing the query seed, they both target the same location in the hash table. If there is an exact match to the reference seed, one would expect to find an entry at that location that includes the same remaining 27 to 30 bits. In such an instance, the 29 designated address bits of the reference sequence may then be looked up to identify the location in the reference to which the query read from which the query seed was derived aligns.

[0180] However, with regard to the remaining 27 to 30 bits, these bits may also represent a secondary access key that can be brought into a hash table, such as to confirm the results of the first 26 to 29 bits of the primary access key. Because the hash table represents a perfect 1:1 scrambling of the 28 nucleotide / 56-bit sequence, and only approximately 26 to 29 of the bits are used to determine the address, these 26 to 29 bits of the primary access key are essentially confirmed, thereby determining the correct address on the first try. Therefore, this data does not need to be confirmed. However, the remaining approximately 27 to 30 bits of the secondary access key must be confirmed. Therefore, the remaining approximately 27 to 30 bits of the query seed are inserted into the hash table as a means to complete the match. Such an implementation may be shorter than storing the entire 56-bit key, thereby saving space and reducing the overall memory requirements and processing time of the module.

[0181] Thus, a hash table can be configured as an index, where known sequences of one or more reference genomes broken down into seeds of a predetermined length, such as 28 nucleotides in length, are randomly organized into a table, and one or more sequenced reads derived from sequencing a subject's genomic DNA or RNA, or "seed portions thereof," can be passed through the hash table index according to a hash function or the like to look up the seeds in the hash table index. One or more locations, such as positions in the reference genome, where the sample seeds match can be obtained from the table. Using a brute-force one-dimensional search to scan the reference genome for seed-matching locations would require identifying 3 billion locations. However, using a hashing approach, each seed lookup can be performed in approximately a fixed amount of time. Often, this location can be identified in a single access. If multiple seeds map to the same location in the table, a small number of additional accesses can be made to find the currently looked-up seed. Thus, even though there may be over 30 million potential positions to match for a given 100 nucleotide long read with respect to the reference genome, the hash table and hash function can quickly determine where the read will appear in the reference genome. Thus, by using a hash table index, it is not necessary to search the entire reference genome to determine where the read aligns.

[0182] As noted above, chromosomes have a double helix structure, consisting of two opposing, complementary strands of nucleic acid sequence that are bound together to form a double helix. For example, when the double helix structure is formed, these complementary base pairs bind to each other according to the rules that "A" binds to "T" and "G" binds to "C." This therefore results in two equally opposing strands of nucleic acid sequence that are complementary to each other. More specifically, the bases of the nucleotide sequence of one strand are mirrored by the complementary bases of the opposing strand, resulting in two complementary strands. However, DNA transcription is unidirectional, starting from one end of the DNA and proceeding toward the other end. Therefore, the conclusion is that transcription occurs in one direction for one strand of DNA, and in the opposite direction for its complementary strand. As a result, the two strands of a DNA sequence are found to be reverse complementary, i.e., when the sequence order of one strand of DNA is compared to the other, what is found are two strands in which the nucleotide letters of one strand have been switched with their complementary letters in the other strand, e.g., "A" becomes "T", "G" becomes "C", and vice versa, and the order has been reversed.

[0183] Due to the double helix structure of DNA, during sample preparation steps prior to sequencing, the chromosomes are separated, e.g., denatured, into separate strands, which are then dissolved into smaller segments of a predetermined length, e.g., 100-300 bases in length, which are then sequenced. Although it is possible to separate the strands before sequencing so that only one strand is sequenced, typically the DNA strands are not separated, so that both strands of DNA are sequenced. Therefore, in such cases, approximately half of the reads in a FASTQ file may be reverse-complemented.

[0184] Of course, both strands of the reference genome, e.g., the complementary and reverse complementary strands, can be processed and hashed as described above, but this would make the hash table twice as large and the hash function twice as timely to perform, and may require approximately twice the processing time, e.g., to compare both the complementary and reverse complementary strands of two genome sequences. Thus, to save memory space, reduce processing power, and / or speed up processing, in various instances, only one strand of the model genome DNA need be stored in the hash table as a reference.

[0185] However, according to a typical sequencing protocol, such as one in which the two strands of a subject's DNA are not separated from each other, any read generated from the sequenced DNA may be from either strand, i.e., the complementary strand or its reverse complementary strand, and it may be difficult to determine which strand, the complementary strand or the reverse complementary strand, is being processed. More specifically, in various cases, since only one strand of the reference genome needs to be used to generate a hash table, half of the reads generated by the sequencing protocol may not match a specific strand of the model genome reference, such as the complementary strand or its reverse complementary strand, because, for example, the read being processed is often the reverse complementary strand of the hashed segment of the reference genome. Thus, only the reads generated from one strand of DNA match the indexed sequence of the reference genome, while the reads generated from the other strand are theoretically the reverse complementary strand of that sequence and do not match anywhere in the reference genome. Furthermore, for any given read that is the reverse complement of a strand of a stored reference genome, matters may be further complicated by the fact that the read may nevertheless erroneously match a portion of the reference genome, such as by mere chance. In light of the above, in order for mapping to proceed efficiently, not only must it be determined in various cases where the read matches in the reference genome, but it must also be determined whether the read has been reverse complement converted. Thus, the hash table and / or function module should be constructed in a way that minimizes these complexities and / or the types of errors that may result.

[0186] For example, as shown above, in one case, a hash table can store both the complementary and reverse complementary strands of a reference genome, so that each read of a subject's sequenced DNA or its reverse complementary strand can be matched with the respective strand in the genomic reference DNA. In such a case, for any given seed in the read, assuming no errors or variations, the seed should theoretically match one strand or the other, i.e., the complementary or reverse complementary strand of the reference. However, storing both strands of the reference genome in a hash index may require approximately twice the storage space (e.g., 64 gigabytes instead of 32 gigabytes), twice the amount of processing resources, and / or twice the length of processing time. Furthermore, such a solution does not solve the problem of palindromes, which can match in both directions, e.g., the complementary and reverse complementary strands.

[0187] Thus, the hash table index can be constructed to include both strands of the genome reference sequence. In various cases, the hash table can be constructed to include only one strand of the model genome as a reference. This can be useful because storing the hash table in memory requires half the storage and / or processing resources than if both strands were stored and processed, and therefore should require less time for lookup. However, storing only one strand of the genome as a reference can be confusing because, as noted above, if the DNA of the sequenced subject is double-stranded, it is usually not known which strand any given read originated from. Therefore, in such cases, the hash table should be constructed to take into account the fact that the mapped read can be from either strand and therefore may be the complement or reverse complement of the stored segment of the reference genome.

[0188] Thus, in various cases, such as when only one orientation of the seed from the reference is stored in the hash table, when a hash function is performed on a seed generated from a read of a FASTQ file, the seed can first be looked up in its current orientation and / or then reverse complement converted and the reverse complement looked up. This may require two lookups in the hash index, e.g., twice as many lookups, assuming no errors or variations, but one of the seeds or its reverse complements should match a complementary segment in the reference genome, which should reduce overall processing resources, e.g., less memory is used, and should also reduce time, e.g., not have to compare as many sequences.

[0189] More specifically, as described above, where the seed consists of 28 nucleotides in one particular orientation, for example, digitally represented in a 56-bit binary format, the seed can be reverse-complemented, and the reverse complement can be digitally represented in a 56-bit binary format. The binary format for each representation of the seed sequence and its complement results in a number having a value, for example, an integer, and the value is represented by the number. These two values, for example, two integers, can be compared, and the number with the higher or lower value, for example, the higher or lower absolute value, can be chosen as the standard directional choice, which can be stored in a hash table and / or passed through a hash function. For example, in some cases, the number with the higher value can be selected to be processed by the hash function.

[0190] Another method that can be utilized is to construct the seeds so that each consists of an odd number of bases. The standard orientation to be selected could then be a strand in which the center base is either "A" or "G" but not "T" or "C," or vice versa. A hash function is then performed on the seeds that meet the standard orientation requirements. In such a scheme, only two bits representing the center base need to be compared to see which has the higher value, and only two bits of the sequence need be looked up. Therefore, it is only necessary to look at the bits representing the two center bases. Typically, this can work because the seeds are odd in length, so the center base is always reverse-complemented. However, while this can work for odd seed lengths, hashing seeds with higher or lower values, as explained above, should work for all seed lengths, although such a method may require processing, e.g., looking up, more bits of data.

[0191] These methods can be performed on any number of seeds of a reference, e.g., all seeds, and / or any number of seeds derived from all or a portion of the reads in a FASTQ file, e.g., all seeds. Approximately half of the time, the binary representation of a seed in a given direction, e.g., the complementary strand, will have the higher value, and approximately half of the time, the binary representation of a seed in the opposite direction, e.g., the reverse complementary strand, will have the higher value. However, looking at the binary numbers, the one with the higher value is the one that is fed into the hash table. For example, the binary integers for each read and its complementary strand can be compared, and the sequence with the first 1 is the one of the two strands selected to be stored as a strand in the hash table and / or run through the hash function. If both strands have their first 1 in the same place, the strand with the second 1 that comes first is selected, and so on. Of course, the read with the lower value can also be selected, in which case the strand with the first 0 and / or with a larger number of 0s at the beginning is selected. An indication, e.g., a flag, may also be inserted into the hash table to indicate which orientation the strand being stored and / or hashed represents, i.e., complementary or reverse complementary, e.g., a 1RC flag for reverse complementary.

[0192] More specifically, when performing a hash function to access a hash table, the seed from the genomic reference DNA and the seed derived from the sequence data read are subject to these same operations, e.g., converted to binary form and compared with their reverse complements, where the integer with the higher or lower value is selected as the standard direction, passed through the hash function, and fed into the hash table to be looked up and matched with each other. However, since this is the same operation being performed in substantially the same way on the reference sequence and the read sequence, if two sequences, i.e., the reference seed and the subject seed, initially have the same sequence, the same record will be derived and they will all be directed to the same cell in the hash table, even if one has undergone reverse complement transformation.

[0193] As a result, if a seed in a reference having a given sequence in a particular orientation is converted to binary form and hashed, a seed derived from a sample read has the same sequence but in the opposite orientation, e.g., reverse complementary, and that seed is subjected to the above protocol, and when the binary value is determined and the hash function is performed in the manner disclosed above, the lookup will be directed to the exact same address in the hash table as if the hash function had been performed on a complementary seed from the start. Thus, in this scheme, it does not matter which orientation the seed being processed is in, as it will always be directed to the same address.

[0194] Thus, in such a scheme, the method disclosed herein can hash the position of the seed in the table regardless of its orientation, thereby determining whether any given seed is reverse-complemented or not, based on a flag in the record. For example, it is known whether the seed was inverted from the reference, and whether the seed derived from the subject's read also had to be inverted. As a result, if the determination is the same for both, the orientation is the same between the read and the reference. However, if one is inverted and the other is not, it can be concluded that the read maps to a reverse-complemented version of the reference. Thus, by using a hash table, it can be determined where in the genome a given read, or a portion thereof, such as a seed, matches and / or whether it is reverse-complemented or not. Furthermore, while the above is described with respect to generating a hash table from a reference genome and performing various accompanying hash function operations on seeds generated from reads, e.g., from a FASTQ file, it should be understood that a system can be constructed such that the hash table indexes are generated from seeds derived from reads of a subject's sequenced DNA, and the various accompanying hash function operations as described herein are performed on seeds generated from the reference genome.

[0195] As described above, the advantage of using a hash table and / or hash function is that by using a seed, most of the reads of sequenced DNA can be matched with a reference genome by using a single hash lookup, and in various cases, it is not necessary to hash and / or look up all the seeds derived from the reads.Seeds can be any suitable length, such as a relatively short length, for example, less than 16 nucleotides, about 20 nucleotides, about 24 nucleotides, about 28 nucleotides, about 30 or about 40 or about 50, or 75, or about 100 nucleotides, or even 250, 500, 750, or even 999, or even about 1000 nucleotides, or a relatively long length, for example, a length of about 1000 nucleotides, or even about 10000 or more, or even about 100,000 or more, or even 1,000,000 or more nucleotides.However, as described above, there are some drawbacks to using seeds, such as in a hash table, particularly regarding selecting a seed of suitable length.

[0196] For example, any suitable seed length can be used in the mapping function, but there are advantages and disadvantages to using a relatively short or relatively long seed length.For example, the shorter the seed length, the less likely it is to contain errors or variations, which may prevent finding a match in the hash table.However, the shorter and less unique the seed length, the more matches are expected between the seed of the reference genome and the seed derived from the read of the sequenced DNA of the subject.In addition, the shorter the seed length, the more lookups must be performed by the hash function, which requires more time and processing power.

[0197] On the other hand, the longer the seed, the more unique the seed, reducing the likelihood of multiple matches between the reference seed and the query. Longer seeds also require fewer seeds in the read, resulting in fewer lookups, which takes less time and requires less processing power. However, the longer the seed, the greater the likelihood that the seed derived from the sequenced DNA may contain errors, such as sequencing errors, and / or variations compared to the reference, preventing a match from being found. Furthermore, longer seeds have the disadvantage of being more likely to hit the end of the read and / or the end of the chromosome. Thus, while a seed only 20–100 nucleotides long may have several matches in the hash table, a seed 1000 or more nucleotides long may result in far fewer matches, or even no matches at all.

[0198] There are several ways to help minimize these problems. One way is to ensure that adequate oversampling is generated in DNA processing steps prior to sequencing. For example, since it is known that there is typically at least one variation within every 1000 base pairs, the seed length can be chosen to maximize matches while simultaneously minimizing discrepancies due to the inclusion of errors and / or variants. Additionally, the use of oversampling, such as in presequencing and / or sequencing steps, can be used as an additional way to minimize various problems inherent in using seeds in hash functions, etc.

[0199] As shown above, oversampling produces pileups. A pileup is a collection of reads that generally overlap and map to the same location in the genome. For the majority of sample reads, such pileups may not be necessary; for example, the reads and / or seeds generated from the reads do not contain variants and / or do not map to multiple locations in the hash table (e.g., they are not actually overlapping in the genome). However, for reads and / or seeds that may contain variants and / or errors and / or other discrepancies between the seed and / or read and the reference genome, producing pileups for any given region of the genome may be useful. For example, even if only one exact hit is required between the seeds generated from the reads of the sample genome to enable mapping of the sample read to the reference genome, the fact that there may be machine errors or true variants in the sample DNA sequence that may prevent such an exact match from occurring between the read and the reference often makes it useful to produce overlapping pileups in the pre-sequencing and sequencing steps.

[0200] For example, in cases where a sample seed truly contains a variant or error, generating a read pileup can be useful for distinguishing between true variance and mechanical and / or chemical errors. In such cases, pileups can be used to determine whether an apparent variation is actually a true variation. For example, if 95% of the reads in a pileup indicate a "C" at a certain location, that is likely to be a correct call, even if the reference genome has a "T" at that location. In such cases, the discrepancy may be due to a SNP, such as a "T" to "C" substitution at that location in the genome, resulting in a true change in the individual's genetic code from that of the reference. In such cases, pileup depth can be used to compare the overlapping portions of pileup reads at the location where the variance exists, and based on the percentage of reads in the pileup that have the variance, it can be determined whether the variance is truly due to actual variation in the sample sequence. Thus, the actual sequence of the read that best matches the genome sequence can be determined based, in part, on what is reflected in the pileup depth. However, a drawback of using pileup is that it requires longer processing time to process all the excess reads and / or seeds generated thereby.

[0201] Another way to minimize the problems inherent in short or long reads is to utilize a secondary hash table in conjunction with or in conjunction with a first, e.g., primary, hash table. For example, a second hash table and / or hash function can be utilized for seeds that have no hits in the primary hash table or that have multiple hits in the primary hash table. For example, when comparing one seed to another, several outcomes can occur. In one case, there may be no hits, e.g., no matches, anywhere between the two sequences, suggesting a potential error or variation in the seed of the subject's read as compared to the seed derived from the reference genome. Alternatively, one or more matches may be found. However, if multiple matches are found, this can be problematic.

[0202] For example, with respect to the primary hash table, if each seed in the reference being hashed appears only a few times, e.g., once, twice, or three times, there may be no need for a secondary hash table and / or hash function. However, if one or more of the seeds occur more times, e.g., 5, 10, 15, 20, 25, 50, 100, 1000, or more times, this can be problematic. For example, there are known regions in the sequence of the human genome that have been determined to be mathematically significant in that they are repeated many times. As a result, mapping any seed to one of these locations may actually unintentionally map to many of these locations, such as when the seed constitutes a nucleotide of an overlapping sequence. In such cases, it may be difficult to determine to which of all the possibilities the seed actually aligns. However, because these repeat regions are known and / or become known, in order to avoid wasting time and processing power trying to use the primary hash function to determine something that is likely to be undecidable, any seeds that would normally map to one or more of these regions may be distinguished so that they are allocated to a secondary hash table for processing by the primary or secondary hash function.

[0203] More specifically, when comparing a seed of a genomic reference with a seed generated from a subject's genomic read, one to hundreds, or even thousands of matching locations may occur. However, the system of the present invention may be configured to handle a number of overlapping matches without requiring further processing steps, such as when the number of matches is less than about 50, or less than about 40, or less than about 30, or less than about 25 or less than about 20, or less than about 16 matches, or less than about 10 or 5 matches. However, if there are more valid hit matches returned than this, the system may be configured to implement a secondary hash function, for example, using a secondary hash table.

[0204] Thus, rather than placing such seeds known to have a high probability of redundancy in the primary hash table, such seeds may be placed in a secondary hash table or in a secondary region within the first hash table. Additionally, in some cases, a record, e.g., an extension record, may be placed in the primary hash table that conveys nothing about the multiple potential mapping locations of the seed but conveys a command for accessing the secondary hash table. For example, the extension record may be an instruction, such as an instruction to extend the length of the primary seed to the length of a longer, unique seed, by, e.g., adding one or more additional bases to the end of the seed, to make the primary seed, e.g., a non-unique or non-duplicate seed, into a longer seed sequence that can be hashed and looked up, such as in a secondary table.

[0205] The record may be configured to inform or otherwise indicate how a known redundant seed should be extended by a given amount, and may also indicate where and / or how the seed should be extended. For example, because hash tables are typically precomputed, e.g., originally constructed from seeds generated from a reference genome, which seeds generated from the reference genome will occur frequently, if any, may be known before the table is constructed. Thus, in various cases, which seeds will need to be shifted to a secondary hash table may be determined in advance. For example, when constructing a hash table index, the characteristics of the reference seed sequence being indexed into the hash table are known, so it may be determined whether each potential seed will result in a large number of hits, e.g., 10 to 10,000 hits.

[0206] More specifically, in various cases, an algorithm may be executed to determine all expected matches that a given seed and / or subject's read derived from a reference may have. If it is determined that any particular seed is likely to return a large number of matches, a flag, e.g., a record, may be generated, such as in a cell of a hash table, indicating that this particular seed is a high frequency hit. In such cases, the record may further indicate that the primary hash of this seed, and seeds like this seed, should be skipped because it is not practical to perform the large number of evaluations of such seeds, e.g., 20-10,000 or more, required to accurately determine where the seed actually maps. In such cases, the primary hash function may not be able to accurately determine which of all potential locations the seed could match is where the read truly aligns, and therefore, because the seed cannot realistically be correctly mapped at this stage, the primary hash function may not be likely to return a usable result, such as a result that accurately indicates where the seed actually matches in the genome.

[0207] In such cases, the hash function algorithm may be configured to calculate what needs to be done to make the redundant seed more unique. For example, a secondary hash function may determine by how many bases, in what order, and at what positions the seed needs to be extended to ensure that the seed is no longer redundant, but rather is suitably unique to be hashed. Thus, the record may also include instructions for extending the redundant seed, e.g., by 2, 4, 6, etc., at one or both ends of the seed, to achieve a predetermined level of uniqueness. In such a scheme, seeds that initially appear identical may be determined to be non-identical.

[0208] For example, in some cases, a typical record may indicate that overlapping seeds are to be extended by up to X odd or even bases, but in some cases, by an even number of bases, e.g., equally on each side, from about 2-4 or about 8-16, to about 32, or about 64 or more bases, etc. For example, if the extension is 64 bases, the record may indicate that 32 bases are added to each side of the seed. The number of bases by which the seed is to be extended is configurable and can be any suitable number depending on how the system is constructed. In some cases, a quadratic hash function may be utilized to determine how many bases the seed should be extended by to return a more likely number of matches. Thus, the extension may be to a point of relative uniqueness, such as to the point where there are only 1, 2, or 3 matching locations where the pattern appears, or to the point where there are as many as 16, 25, or 50 matching locations. In various cases, extending the seed equally from both sides can be useful, such as to avoid problems with reads in the opposite direction, but in various cases the seed can be extended by adding one or more bases unevenly on both sides.

[0209] More specifically, in one example, if a seed includes 28 bases and an extension record, such as an extension record located in a cell in the primary hash table, instructs the hash function to extend the seed by, for example, 64 bases, the record can further instruct the hash function on how to extend the seed, for example, by adding 32 bases on both sides of the seed. However, extension can occur at any suitable location on the read and can be performed symmetrically or asymmetrically. In some cases, the record may instruct the hash function to extend the seed symmetrically, as symmetric extension may work better in some cases, such as for reverse complements discussed herein. In such cases, when extending, the same number of bases are added to both sides of the seed, etc. However, in other cases, extension can be performed by adding an even or odd number of bases in an asymmetric format, so it is not necessary to extend the seed by the same number of bases on both sides. Typically, the primary hash table is configured so that it is not completely filled. For example, it is desirable to configure the table so that it does not exceed 80% or 90% of its capacity. This is to maintain high performance lookup rates: if there are many collisions when hashing seeds to the same location when building the table, the storage mechanism creates a chain of references to other locations, allowing the lookup mechanism to find the one assigned to the overflowed seed. The denser the table, the more collisions there are and the longer the chain that must be followed to find a real match.

[0210] In various cases, such as when an initial redundant seed is 28 bases long and the record instructs extending the seed on each opposite side of the seed from 18 to 32 to 64 bases, etc., the digital representation of the seed may be approximately 64 bases × 2 bits per base = 128 bits. Therefore, depending on how the mapping module is set up, this may be too large for the primary hash table to handle. Therefore, in some cases, to address such extensive processing needs, in some embodiments, a secondary hash module may be configured to store information associated with larger seeds. Because the number of seeds requiring extension is a fraction of the total number of seeds, the secondary hash table may be smaller than the primary hash table. However, in other cases, known redundant portions of a sequence, e.g., a primary sequence, may be replaced by a preselected variable, such as a predetermined sequence length, to reduce the processing requirements of the module, e.g., to save bits. In such cases, the redundant sequence does not need to be fully represented digitally because it is already known and identified. Rather, in various cases, all that really needs to be done is to replace the known redundant sequence with the known variable sequence; what really needs to be looked up are the extensions added on either side of the variable sequence, e.g., wings, because they are the only non-redundant parts of the new initial sequence. Thus, in some cases, the initial sequence can be replaced with a shorter unique identifier code (e.g., a 24-bit proxy instead of a 56-bit representation), an extension base such as a 36-bit extension can be added to the proxy (e.g., making it 60 bits total), and this extension can then be placed in the extension record in the primary table. In such a scheme, the drawbacks of too short and / or too long reads can be minimized while maintaining the benefit of only one or a few lookups in the hash table.

[0211] As indicated above, implementations of the hash functions described above can be performed in software and / or hardware. An advantage of implementing a hash module in hardware is that it can accelerate processing, allowing processing to occur much faster. For example, where software may include various instructions for performing one or more of these various functions, implementing such instructions often requires that data and instructions be stored, and / or fetched, and / or read, and / or interpreted, such as before execution. However, as indicated above and described in more detail herein below, a chip can be hardwired to perform these functions without the need to fetch, interpret, and / or execute one or more of a series of instructions. Rather, a chip can be hardwired to perform such functions directly. Thus, in various aspects, the present disclosure is directed to custom hardwired machines that can be configured such that part or all of the hash modules described above can be implemented by one or more network circuits, such as integrated circuits hardwired on the chip, such as an FPGA, an ASIC, or a structured ASIC.

[0212] For example, in various cases, a hash table index may be constructed and the hash function may be performed on the chip; in other cases, the hash table index may be generated off-chip, such as via software executed by a host CPU; once generated, the hash table index is loaded onto the chip and utilized by the chip, such as when executing the hash module. In some cases, the chip may include any suitable number of gigabytes, such as 8 gigabytes, 16 gigabytes, 32 gigabytes, 64 gigabytes, or approximately 128 gigabytes. In various cases, the chip may be configurable so that various processes of the hash module execute using only a portion or all of the memory resources. For example, if a custom reference genome can be constructed, a majority of the memory may be dedicated to storing the hash reference index and / or storing reads and / or reserving space for use by other functional modules; for example, 16 gigabytes may be dedicated to storing reads, 8 gigabytes may be dedicated to storing the hash index, and another 8 gigabytes may be dedicated to other processing functions. In another example, if 32 gigabytes are dedicated to storing reads, 26 gigabytes may be dedicated to storing the primary hash table, 2.5 gigabytes may be dedicated to storing the secondary table, and 1.5 gigabytes may be dedicated to the reference genome.

[0213] In some embodiments, the secondary hash table may be constructed to have a greater digital impact than the primary hash table. For example, in various instances, the primary hash table may be configured to store 8-byte hash records with 8 records each per hash bucket, for a total of 64 bytes per bucket, and the secondary hash table may be configured to store 16 hash records per bucket, for a total of 128 bytes. For each hash record that contains overflow hash bits that match the same bits of the hash key, potential matching locations in the reference genome are reported. Thus, a maximum of 8 locations may be reported for the primary hash table. A maximum of 16 locations may be reported for the secondary hash table.

[0214] Whether implemented in hardware or software, in many cases it can be useful to construct a hash table to avoid collisions. For example, due to various system artifacts, there may be multiple seeds that want to be inserted into the hash table at the same location regardless of whether there is a match there. Such cases are called collisions. Often, collisions can be avoided in part by the way the hash table is constructed. Thus, in various cases, the hash table may be constructed to avoid collisions and may therefore be configured to include one or more virtual hash buckets.

[0215] In various cases, the hash table may be constructed so that it is represented in an 8-byte, 16-byte, 32-byte, 64-byte, 128-byte format, etc. However, in various exemplary embodiments, it may be useful to represent the hash table in a 64-byte format. This may be useful, for example, when the hash function is to use accesses to memory, such as DRAM in standard DIMM or SODIMM form factors, where the minimum burst size is typically 64 bytes. In such cases, the processor design for accessing a given memory is such that the number of bytes required to form buckets in the hash table is also 64, thus achieving maximum efficiency. However, if the table were constructed in a 32-byte format, this would be inefficient because approximately half the bytes delivered in a burst would contain information not needed by the processor. This would halve the effective byte delivery rate. Conversely, if the number of bytes used to form buckets in the hash table were a multiple of the minimum burst size, e.g., 128, there would be no performance degradation as long as the processor actually needed all of the information returned in a single access. Thus, in cases where the optimal burst size of memory access is a given size, e.g., 64 bytes, the hash table can be constructed to optimally utilize the memory burst size, e.g., the bytes allocated to represent bins in the hash table and processed by the mapping function, e.g., 64 bytes, match the memory burst size. As a result, when memory bandwidth is a constraint, the hash table can be constructed to optimally utilize such constraint.

[0216] Furthermore, note that while records can be packed into 8 bytes, the hash function can be constructed to prevent 8 bytes from the table from being read to process one record, as this may be inefficient. Rather, all 8 records in a bucket, or some subportion thereof, may be read at once. This may be useful for optimizing the processing speed of the system, given the architecture described above, since processing all 8 records at the same speed takes the same amount of time as simply processing one record. Thus, in some cases, the mapping module may include a hash table, which may itself include one or more subsections, e.g., virtual sections or buckets, and each bucket may have one or more slots, such as 8 slots, so that one or more different records can be inserted into the bucket, such as to manage collisions. However, in some situations, one or more such buckets may be filled with records, and therefore means may be provided for storing additional records in other buckets and recording information in the original bucket indicating that the hash table lookup mechanism needs to look further to find a match.

[0217] Thus, in some cases, it may be useful to utilize one or more additional methods for managing collisions, etc.; one such method may include one or more of one-dimensional probing and / or hash chaining. For example, if it is not known what is actually being searched for in a hash table or a portion thereof, such as in one bucket of the hash table, and the particular bucket is full, the hash lookup function may be configured such that if a bucket is full and the desired record is not found when searched, the function can proceed to the next bucket, e.g., the +1 bucket, and then check that bucket. In such a scheme, all buckets can be searched when looking for a particular record. Thus, such a search may be performed by looking through bucket after bucket until what is being searched for is found or until it becomes clear that it is unlikely to be found, e.g., an empty slot is found in at least one of the buckets. Specifically, if each bucket is filled in order, and each bucket is searched in the order of filling, and an empty slot is found, for example, when searching the buckets in order for a particular record, the empty slot may indicate that the record is not present, because if the record were present, it would at least have been in the empty slot, even if not in the preceding bucket.

[0218] More specifically, if 64 bytes are designated for storing information in a hash bucket and it contains eight records, upon receiving the fetched bucket, the mapping processor can operate on all eight records simultaneously to determine which are matches and which are not. For example, when performing a lookup, such as a seed from a read obtained from sequenced sample DNA against a seed generated from a reference genome, the digital representation of the sample seed can be compared against the reference seed in all, say, eight, records to find a match. In such cases, several outcomes can occur. A direct match may be found. The sample seed may be entered into the hash table, and in some cases, no match is found because the seed is simply not exactly the same as any corresponding seed in the reference, for example, due to a machine or sequencing error with respect to the seed or the read from which the seed was generated, or because the person has a different genetic sequence than the reference genome. Or, the seed may be entered into the hash table, and multiple matches may be returned, for example, the sample seed may match two, three, five, ten, fifteen, twenty, or more places in the table. In such cases, multiple records may be returned, all pointing to various different locations in the reference genome that that particular seed matches, and these matching records may be in the same bucket, or multiple buckets may have to be probed to return all of the significant results, e.g., matching results.

[0219] In some cases, such as when space can be a limiting factor in a hash table, e.g., in a hash table bucket, additional mechanisms for resolving collisions and / or conserving space may be implemented. For example, a hash chaining function may be performed when space becomes limited, such as when more than eight records need to be stored in a bucket, or when this is desirable in other cases. Hash chaining may involve, for example, replacing a record containing the location of a particular location in a genome sequence with a record containing a chain pointer that, instead of pointing to a location in the genome, points to some other address, e.g., a second bucket in the current hash table, e.g., the primary hash table or the secondary hash table. This has the advantage over one-dimensional probing methods of allowing the hash lookup mechanism to directly access the bucket containing the desired record rather than sequentially checking the buckets.

[0220] Such a process can be useful considering system architecture. For example, a primary seed being hashed, such as in a primary lookup, is located in a given position in a table, e.g., its original location, while a chained seed is placed in a location that may differ from the original bucket. Thus, as shown above, a first portion of the digitally represented seed, e.g., about 26 to about 29 bits, can be hashed and looked up in a first step. Then, in a second step, the remaining about 27 to about 30 bits can be inserted into a hash table, such as in a hash chain, as a means to verify the first pass. Thus, for any seed, its original address bits can be hashed in a first step, and the secondary address bits can be used in a second verification step. Thus, a first portion of the seed can be inserted into a primary record location, and a second portion can be placed into a table at a secondary record chain location. And, as shown above, in various instances, these two different record locations can be positionally separated by a chained format record or the like. Thus, at any destination bucket in a chain, the chain format records can positionally separate entries / records for local primary first bucket access and probing, and records for the chain.

[0221] Such hash chains can continue for many lengths. An advantage of such chains is that if one or more of the buckets contains one or more, e.g., two, three, four, five, six, or more, empty record slots, these empty slots can be used to store hash chain data. Thus, in some cases, a hash chain may involve starting with an empty slot in one bucket and chaining that slot to another slot in another bucket, where the two buckets may be at distant locations in the hash table. Additional care can be taken to avoid confusion between records that are located in distant buckets as part of a hash chain and "native" records that hash directly to the same bucket. As always, the remaining approximately 27 to 30 bits of the secondary access key are verified against the corresponding approximately 27 to 30 bits stored in a record located farther away in the chained bucket, but because the chained buckets are located farther away from the original hash bucket, verifying these approximately 27 to 30 bits is not sufficient to guarantee that the matching hash record corresponds to the original seed that reached this bucket by chaining, as opposed to some other seed reaching the same bucket by direct access (e.g., verifying approximately 27 to 30 bits can be complete verification when the approximately 26 to 29 bits used to address the hash table are implicitly verified by their proximity to the initial hash bucket being accessed).

[0222] A position system may be used in chained buckets to prevent retrieving erroneous hash records without having to store the entire hash key in the record. Thus, a chained bucket must contain a chained sequence format record, which, if necessary, contains further chain pointers to continue the bucket chain. This chained sequence record must appear in a bucket slot after all "native" records corresponding to the direct hash access and before all remote records belonging to the chain. During a query, before following any chain pointer, any records that appear after the chained sequence record should be ignored, and after following any chain pointer, any records that appear before the chained sequence record should be ignored.

[0223] For example, if the buckets are about 75%-85% full, scanning eight buckets can find only 15-25 usable slots, whereas using hash chaining, these slots can be found in two, three, or four buckets. In such cases, the number of probe or chaining steps required to store hash records is important because it affects the speed of the system. At runtime, if probing is required to find records, numerous hash lookup accesses, e.g., 64-byte bucket reads, must be performed, slowing the system. Hash chaining helps minimize the average number of accesses that must be performed because more excess hash records can generally be stored per chained bucket, which can be selected from a wide range, rather than per probing bucket, which must be next in order. Thus, a given number of excess hash records can typically be stored in an array of chained buckets that is shorter than the required array of probing buckets, which in turn limits the number of accesses required to locate such excess records in a query. Nevertheless, probing remains valuable for smaller amounts of excess hash records because it does not require bucket slots to be sacrificed for chain pointers.

[0224] For example, after all possible matches have been determined for a seed relative to a reference genome, it must be determined which of all potential positions to which a given read may match is actually the correct location to which the read aligns. Thus, after mapping, there may be multiple locations in the reference genome to which one or more reads appear to match. As a result, there may be multiple seeds that appear to represent exactly the same thing, e.g., they may match the exact same location in the reference, given the location of the seed in the read.

[0225] Therefore, the actual alignment must be determined for each given read. This determination can be made in several different ways. In one example, all reads can be evaluated to determine the correct alignment with respect to the reference genome based on the location indicated by each and every seed from the read that returned location information during the hash lookup process. However, in various cases, a seed chain filtering function can be performed on one or more of the seeds before performing the alignment.

[0226] For example, in some cases, seeds associated with a given read that appear to map to the same general location relative to the reference genome can be aggregated into a single chain that references the same region. All seeds associated with a read can be grouped into one or more seed chains, with each seed being a member of only one chain. It is such chains that the read is then aligned to each indicated location in the reference genome. Specifically, in various cases, all seeds with the same supporting evidence indicating that they all belong to the same general location in the reference can be collected to form one or more chains. Thus, seeds that group together, or at least appear to be close to each other in the reference genome, for example, within a certain range, are grouped into a seed chain, and those outside this range are placed into a different seed chain.

[0227] Once these various seeds are aggregated into one or more various seed chains, it can be determined which of the chains actually represent the correct chain to be aligned. This can be done, at least in part, by using a filtering algorithm, which is a heuristic designed to remove weak seed chains that are highly unlikely to be correct. Generally, the longer the seed chain, in terms of the length spanned in the read, the more likely it is to be correct, and furthermore, seed chains that contribute more seeds are more likely to be correct. In one example, a heuristic can be applied, where relatively strong "good" seed chains, for example, those with long or many seeds, remove relatively weak "poor" seed chains, for example, those with short or few seeds.

[0228] In one variation, the length of the poor linkage determines a threshold length, e.g., twice its length, whereby a good linkage of at least the threshold length can remove the poor linkage. In another variation, the seed count of the poor linkage determines a threshold seed count, e.g., five times its length, whereby a good linkage of at least the threshold seed count can remove the poor linkage. In another variation, the length of the poor linkage determines a threshold seed count, e.g., twice its seed count minus its seed length, whereby a good linkage of at least the threshold seed count can remove the poor linkage. In some variations, such as when chimeric alignment of reads is desired, only good seed links that substantially overlap with poor seed links in the reads can remove the poor seed linkage.

[0229] This process removes seeds that are unlikely to have identified regions of the reference genome where high-quality read alignments can be found. This can be useful because it reduces the number of alignments that need to be performed for each read, thereby increasing processing speed and saving time. Thus, this process can be used, in part, as a tuning function, whereby when higher speed is desired, e.g., in a high-speed mode, more detailed seed linkage filtering is performed, and when higher overall accuracy is desired, e.g., in a high-accuracy mode, less seed linkage filtering is performed, e.g., all seed linkages are evaluated.

[0230] In various embodiments, seed editing can be performed, such as before the seed chain filtering step. For example, if all of the seeds for each read are subjected to a mapping function and none of them return a hit, it may be highly likely that there were one or more errors in the read, for example, errors created by the sequencer. In such cases, if no match results are returned, an editing function such as a single-change editing process, e.g., an SNP editing process, can be performed on each seed. For example, at location X, the single-change editing function may indicate that a specified nucleotide be replaced with one of three other nucleotides, and making that change, e.g., an SNP substitution, determines whether a hit, e.g., a match, is obtained. This single-change editing can be performed on every single location in the seed and / or every single seed of a read in the same manner, e.g., replacing each location in the seed with a respective alternative base. In addition, if a single change is made in a seed, the effect that the change has on every other overlapping seed can be determined in light of the single change.

[0231] Such editing may also be performed for insertions, such as when one of four nucleotides is added at a given insertion location X, and making a substitution determines whether a hit is obtained. This may be performed for all four nucleotides, and / or for all locations in the seed (X, X+1, X+2, X+3, etc.), and / or for all seeds in the read. Such editing may also be performed for deletions, such as when one of four nucleotides is deleted at a given location X in the seed, and making the deletion determines whether a hit is obtained. This may then be repeated for all locations X+1, X+2, X+3, etc. However, such editing may result in a lot of extra processing work and time, such as by requiring a large number of additional lookups, such as two, three, four, five, ten, fifty, one hundred, or two hundred, etc. Nevertheless, if such editing allows an actual hit to be determined when there was no match before, for example, if a match is made, such extra processing and time may be useful. In such cases, it is usually possible to determine that an error has occurred and then correct the error, thereby reusing the lead.

[0232] Additionally, further heuristics can be utilized to determine whether an edit function should be performed, whereby the algorithm performs calculations to determine the probability of obtaining a hit if such an edit is performed. If a certain threshold probability is met, for example, with an 85% chance, such seed chain edits can be performed. For example, the system can generate various statistics about the seed chains, such as calculating how frequently hits exist and / or how many seed chains contain frequent hits, to determine whether editing the seed chains is likely to make a difference in determining a match. For example, if it is determined that there is a large percentage of frequent hits, then seed chain edits can be skipped in such cases because it is unlikely that various sequences will be unique enough to provide a hit within a reasonable number of hash table lookups, such as 100 or fewer, 50 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, or 10 or fewer. Such statistics can be reviewed, and then a decision can be made as to whether to perform seed edits. For example, if statistics show that for any one read, half of the locations show no matches and the other half show frequent matches, then it is probably worth doing a seed edit since there is a high probability of an error if no matches are returned, but if there are a large number of frequent matches then it may simply not be worth doing a seed edit.

[0233] The result of performing one or more of these mapping, filtering, and / or editing functions is a list of reads, each of which contains a list of all the potential positions that the read can match with respect to the reference genome.Therefore, mapping function can be performed to quickly determine where the reads of the FASTQ file obtained from sequencer map to the reference genome, for example, where various reads map in the whole genome.However, if there is an error or genetic variation in any of the reads, it may not get an exact match with the reference, and / or there may be some places where one or more reads appear to match.Therefore, it is necessary to determine where various reads actually align with respect to the genome as a whole.

[0234] Thus, after mapping and / or filtering and / or editing, the locations of the positions of a large number of reads have been determined, where for some of an individual's reads, the locations of multiple positions have been determined, and it is now necessary to determine which of all the potential positions is in fact the true or most likely position to which the various reads align. Such alignment may be performed by one or more algorithms, such as dynamic programming algorithms, that match the mapped reads to a reference genome and perform an alignment function against the reference genome.

[0235] An exemplary alignment function compares reads to a reference, such as by arranging one or more, e.g., all, of the reads in a graphical relationship to one another, such as in a table, e.g., a virtual array or matrix, where the sequence of one of the reference genome or the mapped reads is arranged in one dimension or axis, e.g., the horizontal axis, and the other is arranged in an opposing dimension or axis, such as the vertical axis. A conceptual scoring wavefront is then passed over the array to determine the alignment of the reads to the reference genome, such as by calculating an alignment score for each cell in the matrix.

[0236] A scoring wavefront represents one or more, e.g., all cells or a portion of such cells, of a matrix, which may be scored independently and / or simultaneously according to dynamic programming rules applicable in alignment algorithms, such as the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, and / or related algorithms. For example, taking the origin of the matrix (corresponding to the start of the lead and / or the start of the reference interval of the conceptual scoring wavefront) as the upper left corner, first, only the upper left cell at coordinates (0,0) of the matrix can be scored, e.g., a one-cell wavefront. Next, two cells to the right and bottom at coordinates (0,1) and (1,0) can be scored, e.g., a two-cell wavefront. Next, three cells at (0,2), (1,1), and (2,0) can be scored, e.g., a three-cell wavefront. These exemplary wavefronts can then extend diagonally along a straight line from lower left to upper right, with the movement of the wavefront from step to step being diagonally from upper left to lower right through the matrix. The alignment scores may be calculated sequentially, by calculating all scores from left to right in the top row, followed by all scores from left to right in the next row, etc., or in other orders. In this way, the diagonal sweep of the diagonal wavefront represents an optimal arrangement of a bunch of scores, which are calculated simultaneously or in parallel in a series of wavefront steps.

[0237] For example, in one embodiment, the interval of the reference genome containing the segment to which the read is mapped is arranged on the horizontal axis, and the read is arranged on the vertical axis. In this manner, an array or matrix, for example, a virtual matrix, is generated, which allows the nucleotide at each position in the read to be compared with the nucleotide at each position in the reference interval. As the wavefront passes over the array, all possible ways of aligning the read to the reference interval are considered, including when a sequence needs to be modified to match the read with the reference sequence, such as by changing one or more nucleotides of the read to other nucleotides, or by inserting one or more new nucleotides into the sequence, or by deleting one or more nucleotides from the sequence.

[0238] An alignment score is generated that represents the degree of modification that needs to be made to achieve strict alignment, and this score and / or other related data can be stored in a given cell of the array.Each cell of the array corresponds to the probability that the nucleotide at the cell's location on the read axis aligns with the nucleotide at the cell's location on the reference axis, and the score generated for each cell represents the partial alignment that ends at the cell's location in the read and reference interval.The highest score generated in any cell represents the overall best alignment of the read to the reference interval.In various cases, alignment can be global, where the entire read must be aligned to some part of the reference interval, such as using the Needleman-Wunsch algorithm or a similar algorithm, and in other cases, alignment can be local, where only a portion of the read can be aligned to a part of the reference interval, such as using the Smith-Waterman algorithm or a similar algorithm.

[0239] The size of the reference interval can be any suitable size. For example, a typical read can be about 100 nucleotides to about 1000 nucleotides in length, so in some cases, the length of the reference interval can be about 100 nucleotides to 1000 nucleotides in length or longer. However, in some cases, the length of the read can be longer, and / or the length of the reference interval can be longer, such as about 10,000 nucleotides, 25,000 nucleotides, 50,000 nucleotides, 75,000 nucleotides, 100,000 nucleotides, 200,000 nucleotides in length or longer. It can be advantageous to pad the reference interval to be somewhat longer than the read, such as by including 32, 64, 128, 200, or even 500 extra nucleotides in the reference interval beyond the end of the segment of the reference genome to which the read is mapped, such as to allow insertions and / or deletions near the end of the read to be fully evaluated. For example, if only a portion of the read is mapped to a segment of the reference, extra padding can be applied to the reference interval corresponding to the unmapped portion of the read, or by some factor, such as 10%, 15%, 20%, 25%, or even 50% or more, to allow the unmapped portion of the read space to be fully aligned to the reference interval. However, in some cases, the length of the reference interval may be selected to be shorter than the length of the read, such that a long portion of the read, such as 1000 or so nucleotides at one end of the read, is not mapped to the reference, e.g., to focus on alignment at the mapped portion.

[0240] The alignment wavefront may be unlimited in length, or may be limited to any suitable fixed length, or may be variable in length. For example, all cells along a generally diagonal line of each wavefront step extending completely from one axis to the other may be scored. Alternatively, a limited length, such as 64 cells wide, may be scored for each wavefront step, such as by tracking a 64-cell wide range diagonally across the matrix of cells to be scored, and leaving cells outside this range unscored. In some cases, calculating scores far from the range around the true alignment path may be unnecessary, and considerable work can be saved by using a fixed-length scoring wavefront and calculating scores only within a limited range of widths, as described herein.

[0241] Thus, in various instances, an alignment function may be performed on data obtained from, e.g., a mapping module. Thus, in various instances, the alignment function may form a module, e.g., an alignment module, that may form part of a system, e.g., a pipeline, used in addition to, e.g., a mapping module, in a process to determine an individual's actual genome sequence in its entirety, or a portion thereof. For example, output returned from execution of a mapping function, e.g., from, e.g., a list of possibilities for where one or more or all of the reads map to one or more locations in one or more reference genomes, may be utilized by the alignment function to determine the true sequence alignment of the subject's sequenced DNA.

[0242] Such alignment function can sometimes be useful, because, as explained above, sequenced reads are often not always exactly the same as the reference genome for various reasons.For example, one or more of the reads may have SNPs (single nucleotide polymorphisms), such as a substitution of one nucleotide for another at a single location; there may be "indels", i.e., an insertion or deletion of one or more bases along one or more of the read sequences, which are not present in the reference genome; and / or there may be sequencing errors (e.g., errors in sample preparation and / or sequencer reads and / or sequencer output, etc.) that cause one or more of these apparent variations.Therefore, when a read varies from the reference due to SNPs, indels, etc., this may be because the reference is different from the true DNA sequence sampled, or the read may be different from the true DNA sequence sampled.The problem is to find out how to correctly align the read to the reference genome, given the fact that the two sequences will likely differ from each other in many different ways.

[0243] Thus, in various cases, the input to the alignment function, such as from a mapping function, such as a prefix / suffix tree, or a Burrows / Wheeler transformation, or a hash table and / or hash function, can be a list of possibilities for where one or more reads may match one or more locations in one or more reference sequences. For example, for any given read, the read may match any number of locations in the reference genome, such as 1 or 16 or 32 or 64 or 100 or 500 or 1000 or more positions to which the given read maps in the genome. However, every individual read is derived, e.g., sequenced, from only one particular portion of the genome. Thus, to find the true location from which a given particular read was derived, an alignment function, such as a Smith-Waterman gapped alignment, a Needleman-Wunsch alignment, or the like, can be performed to determine where in the genome one or more of the reads actually derived from, such as by comparing all of the potential locations where a match occurs and determining which of all the possibilities is the most likely location in the genome from which the read was sequenced based on which location has the highest alignment score.

[0244] As shown, algorithms are typically used to perform such alignment functions. For example, Smith-Waterman and / or Needleman-Wunsch alignment algorithms can be used to align two or more sequences to each other. In this case, they can be used in such a way that for any given location where a read maps to a reference genome, the probability that this mapping is actually the location from which the read originates is determined. Typically, these algorithms are configured to be executed by software, but in various cases, such as those presented herein, one or more of these algorithms can be configured to be executed by hardware, as will be described in more detail below.

[0245] Specifically, the alignment function operates to at least partially align one or more, e.g., all, of the reads to a reference genome, despite the presence of one or more mismatches, e.g., SNPs, insertions, deletions, structural artifacts, etc., to determine where the reads are likely to match correctly within the genome. For example, one or more reads are compared to the reference genome, and the best possible match of the read to the genome is determined, taking into account substitutions and / or indels and / or structural variants. However, to better determine which of the modified versions of the reads best matches the reference genome, the proposed modifications must be taken into account, and thus a scoring function may also be performed.

[0246] For example, the scoring function can be performed, for example, as part of the overall alignment function, whereby the alignment module performs its function and introduces one or more changes to the sequences to be compared with each other, for example, to achieve a better or best match between the change and the reference, and for each change made to achieve a better alignment, a number is subtracted from the starting score, for example, a perfect score or a starting score of 0, and this is done in such a way that as the alignment is performed, the score of the alignment is also determined, for example, the score increases if a match is found and is penalized for each change introduced, so that the best match for the potential alignment can be determined, for example, by finding which of all potential modified reads matches the genome with the highest score. Thus, in various cases, the alignment function can be configured to determine the best combination of changes that need to be made to the reads to achieve the alignment with the highest score, and this alignment can then be determined to be the correct or most likely alignment.

[0247] Considering the above, there are therefore at least two goals that can be achieved by performing an alignment function. One is a report of the best alignment, including a location in the reference genome and a description of what changes are required to align the read with the reference segment at that location; the other is an alignment quality score. For example, in various cases, the output from the alignment module can be a Compact Idiosyncratic Gapped Alignment Report, e.g., a CIGAR string, where the CIGAR string output is a report detailing all the changes made to the read to achieve the best alignment, e.g., detailed alignment instructions showing how the query actually aligns with the reference. Such a CIGAR string readout can be useful in further stages of processing to better determine whether the predicted variations for a given subject's genomic nucleotide sequence as compared to the reference genome are actually true variations or merely due to machine, software, or human error.

[0248] As described above, in various embodiments, alignment is typically performed sequentially, where an algorithm receives read sequence data for a read and one or more potential locations to which the read may map in one or more reference genomes, such as from a mapping module, and also receives genome sequence data for one or more locations in one or more reference genomes to which the read may map, such as from one or more memories. Specifically, in various embodiments, the mapping module processes reads, such as from a FASTQ file, and maps each of them to one or more locations in the reference genome to which the read may align. An aligner then receives these predicted locations and uses the predicted locations to align the read to the reference genome, such as by building a virtual array with which the read can be compared to the reference genome.

[0249] In performing this function, the aligner evaluates each mapped location for each individual read, specifically, evaluates reads that map to multiple possible locations in the reference genome, and scores the probability that each location is the correct location. The aligner then compares the best scores, e.g., the two best scores, to make a decision about where a particular read actually aligns. For example, when comparing the first and second best alignment scores, the aligner notes the difference between these scores; if the difference is large, the confidence score that the one with the larger score is correct is high. However, if the difference is small, e.g., 0, the confidence score that can distinguish which of the two locations the read actually derives from is low, and further processing may be useful to unambiguously determine the true location in the reference genome from which the read is derived. Thus, in making a call, the aligner looks for the difference between the highest and second-highest confidence scores that a given read maps to a given location in the reference genome. Ideally, the score of the best possible choice of alignment is much greater than the score of the second best alignment for that sequence.

[0250] There are many different ways in which alignment scoring methods can be implemented, for example, according to the methods disclosed herein, each cell of the array can be scored, or a sub-portion of cells can be scored. Typically, each alignment match, corresponding to a diagonal step in the alignment matrix, contributes a positive score, such as +1, if the corresponding read and reference nucleotides match, and a negative score, such as -4, if the two nucleotides do not match. Furthermore, each deletion from the reference, corresponding to a horizontal step in the alignment matrix, contributes a negative score, such as -7, and each insertion into the reference, corresponding to a vertical step in the alignment matrix, contributes a negative score, such as -7.

[0251] In various cases, the scoring parameters for nucleotide matches, nucleotide mismatches, insertions, and deletions can have any of a variety of positive, negative, or zero values. In various cases, these scoring parameters can be modified based on available information. For example, in some cases, alignment gaps (insertions or deletions) are penalized by an affine function of the gap length, e.g., −7 for the first deleted (and correspondingly inserted) nucleotide, but only −1 for each additional deleted (and correspondingly inserted) nucleotide in the contiguous sequence. In various implementations, the affine gap penalty can be achieved by dividing the gap (insertion or deletion) penalty into two components, such as a gap opening penalty, e.g., −6, applied to the first step in the gap, and a gap extension penalty, e.g., −1, applied to each or further steps in the gap. The affine gap penalty can produce more accurate alignments, such as by allowing alignments with long insertions or deletions to achieve a moderately high score. Furthermore, each lateral movement may have the same or a different cost, such as the same cost per step, and / or if a gap occurs, the cost of a lateral movement of the aligner may be lower than the cost of the gap, as such gaps may be associated with higher or lower costs. Thus, in various embodiments, affine gap scoring may be performed, but this may be expensive in software and / or hardware, as it typically requires multiple scores, e.g., three, to be scored for each cell, and therefore, in various embodiments, affine gap scoring is not performed.

[0252] In various cases, scoring parameters may also be affected by "base quality scores" corresponding to nucleotides in a read. Some sequenced DNA read data, such as in a format like FASTQ, may include a base quality score associated with each nucleotide, which indicates the estimated probability that the nucleotide is incorrect, for example, due to a sequencing error. In some read data, the base quality score may indicate the probability that an insertion and / or deletion sequencing error is present in or adjacent to each location, or additional quality scores may provide this information separately. Therefore, more accurate alignment may be achieved by varying scoring parameters, including any or all of the nucleotide match score, nucleotide mismatch score, gap (insertion and / or deletion) penalty, gap opening penalty, and / or gap extension penalty, according to the base quality score associated with the nucleotide or location of the current read. For example, score bonuses and / or penalties may be smaller when the base quality score indicates a high probability that a sequencing or other error exists. Base quality influenced scoring may be performed, for example, using a fixed or configurable lookup table that is accessed using the base quality score, which returns the corresponding scoring parameters.

[0253] In a hardware implementation in an integrated circuit such as an FPGA, ASIC, or structured ASIC, a scoring wavefront may be implemented as a one-dimensional array of scoring cells, such as 16 cells, 32 cells, 64 cells, or 128 cells. Each of the scoring cells may be constructed with digital logic elements in a wired configuration for calculating an alignment score. Thus, for each step of the wavefront, e.g., a clock period or some other fixed or variable unit of time, each of the scoring cells, or a portion of the cells, calculates one or more scores required for a new cell in the virtual alignment matrix. Conceptually, various scoring cells may be thought of as being at various locations in the alignment matrix corresponding to the scoring wavefront as discussed herein, e.g., along a line extending from the lower left to the upper right in the matrix. As is well understood in the field of digital logic circuit design, the physical scoring cells and their constituent digital logic circuits need not be physically arranged in a similar manner on the integrated circuit.

[0254] Thus, as the wave front takes multiple steps to sweep through the virtual alignment matrix, the conceptual locations of the scoring cells update each cell accordingly, conceptually "moving," e.g., a step to the right, or, e.g., a step down in the alignment matrix. All scoring cells make the same relative conceptual movement, leaving the diagonal wave front placement intact. Each time the wave front moves to a new location, e.g., with a vertical step down in the matrix or a horizontal step to the right, the scoring cell arrives at the new conceptual location and calculates the alignment score for the virtual alignment matrix cell that the scoring cell entered.

[0255] In such an implementation, adjacent scoring cells in the one-dimensional array are coupled to convey the query (read) nucleotide, the reference nucleotide, and the previously calculated alignment score. The nucleotides of the reference interval may be sequentially fed to one end of the wavefront, for example, the top right scoring cell in the one-dimensional array, and sequentially shifted downward from there by the length of the wavefront, so that at any given time, there are segments of reference nucleotides in cells of length equal to the number of scoring cells, with one consecutive nucleotide in each consecutive scoring cell.

[0256] Thus, with each horizontal advance of the wavefront, another reference nucleotide is fed into the top right cell, and other reference nucleotides are shifted down and left across the wavefront. This shift of reference nucleotides may be the reality behind the conceptual movement of the wavefront of scoring cells to the right across the alignment matrix. Thus, as the nucleotides of the read are sequentially fed to the opposite end of the wavefront, for example, to the bottom left scoring cell in a one-dimensional array, and from there may be sequentially shifted up the length of the wavefront, so that at any given time, there are segments of query nucleotides in cells of length equal to the number of scoring cells, with one consecutive nucleotide in each consecutive scoring cell.

[0257] Similarly, each time the wavefront advances vertically, another query nucleotide is fed into the bottom-left cell, and other query nucleotides are shifted up and to the right across the wavefront. This shift of query nucleotides is the reality behind the conceptual movement of the wavefront of scoring cells downward across the alignment matrix. Thus, commanding a shift of the reference nucleotide can move the wavefront one horizontal step, and commanding a shift of the query nucleotide can move the wavefront one vertical step. Thus, to produce a generally diagonal wavefront movement, e.g., to follow a typical alignment of query and reference sequences without insertions or deletions, the wavefront can be commanded to step alternately vertically and horizontally.

[0258] Thus, neighboring scoring cells in a one-dimensional array can be combined to communicate previously calculated alignment scores. In various alignment scoring algorithms, such as Smith-Waterman or Needleman-Wunsch, or variations thereof, the alignment score in each cell of a hypothetical alignment matrix can be calculated using previously calculated scores in other cells of the matrix, such as the three cells located immediately to the left of the current cell, immediately above the current cell, and immediately diagonally above and to the left of the current cell. When a scoring cell calculates a new score for another matrix location in which it resides, it must retrieve such previously calculated scores corresponding to such other matrix locations. These previously calculated scores can be obtained from a storage of previously calculated scores in the same cell and / or from a storage of previously calculated scores in one or two adjacent scoring cells in the one-dimensional array. This is because the three contributing score locations in the hypothetical alignment matrix (immediately to the left, above, and diagonally above and to the left) have been scored by the current scoring cell or by one of the adjacent scoring cells in the one-dimensional array.

[0259] For example, the cell immediately to the left in the matrix is ​​scored by the current scoring cell if the most recent wave front step was horizontal (to the right), or by its neighboring cell to the lower left in the 1-D array if the most recent wave front step was vertical (to the bottom). Similarly, the cell immediately above in the matrix is ​​scored by the current scoring cell if the most recent wave front step was vertical (to the bottom), or by its neighboring cell to the upper right in the 1-D array if the most recent wave front step was horizontal (to the right). Similarly, the cell diagonally to the top left in the matrix is ​​scored by the current scoring cell if the two most recent wavefront steps are in different directions, e.g., down then right or right then down, or by its neighboring cell to the top right in the one-dimensional array if the two most recent wavefront steps are both horizontal (to the right), or by its neighboring cell to the bottom left in the one-dimensional array if the two most recent wavefront steps are both vertical (to the bottom).

[0260] Thus, by considering information about the direction of the last one or two wave front steps, a scoring cell can select appropriate previously calculated scores, accessing them within itself and / or in adjacent scoring cells using adjacent cell-to-cell coupling. In one variation, scoring cells at the two ends of a wave front may have outgoing score inputs hardwired to invalid, or zero, or minimum-valued scores, so that those inputs do not affect the calculation of new scores at these end cells.

[0261] With a wavefront thus implemented in a one-dimensional array of scoring cells, with couplings for shifting the reference and query nucleotides in opposite directions across that array to conceptually move the wavefront in vertical and horizontal steps, and couplings for accessing scores previously calculated by neighboring cells to calculate an alignment score at a new virtual matrix cell location where the wavefront enters, it becomes possible to score a range of cells in the virtual matrix, i.e., the width of the wavefront, such as by directing successive steps of the wavefront to sweep the wavefront across the matrix. Thus, for new read and reference intervals to be aligned, the wavefront may begin to be positioned inside the scoring matrix, or may advantageously gradually enter the scoring matrix from the outside, for example, starting to the left of, or above, or diagonally above and to the left of the upper left corner of the matrix.

[0262] For example, a wave front may begin at an upper-left scoring cell located immediately to the left of the upper-left cell of a hypothetical matrix; the wave front may then sweep rightward into the matrix in a series of horizontal steps to score a horizontal range of cells in the upper-left region of the matrix. Once the wave front reaches an expected alignment relationship between the reference and the query, or a match is detected from increasing alignment scores, the wave front may begin to sweep diagonally down and to the right by alternating vertical and horizontal steps to score a diagonal swath of cells through the center of the matrix. Once the scoring cell of the lower-left wave front reaches the bottom of the alignment matrix, the wave front may again begin to sweep rightward in successive horizontal steps to score a horizontal swath of cells in the lower-right region of the matrix until some or all of the wave front cells have swept outside the boundary of the alignment matrix.

[0263] In one variation, improved efficiency can be achieved from the alignment wavefront by sharing scoring cells between two consecutive alignment operations. Because the next alignment matrix is ​​pre-established, as the top-right portion of the wavefront exits the bottom-right region of the current alignment matrix, it can enter the top-right region of the next alignment matrix immediately, or after passing a minimal gap, such as one or three cells. In this way, the horizontal sweep of the wavefront out of one alignment matrix can be the same movement as the horizontal sweep of the wavefront into the next alignment matrix. Doing this can include entering the next alignment matrix with the reference and query bases of the next alignment to be fed into the scoring cells, reducing the average time per alignment consumed by the time required to perform a number of wavefront steps approximately equal to the number of alignment cells in the wavefront, e.g., 64, 63, or 61 steps, which can require, for example, 64, 63, or 61 clock cycles.

[0264] The number of scoring cells in an implementation of an alignment wavefront can be selected to balance various factors, including alignment accuracy, maximum insertion and deletion length, area, cost, and power consumption of the digital logic, clock frequency of the aligner logic, and overall integrated circuit performance. Longer wavefronts are particularly desirable for good alignment accuracy, since a wavefront of N cells can align over indels approximately N nucleotides long or slightly shorter. However, longer wavefronts require more logic, which consumes more power. Furthermore, longer wavefronts can increase the complexity and delay of wiring routing on the integrated circuit, leading to lower maximum clock frequencies and lower net aligner performance. Furthermore, when integrated circuit size or power consumption is limited, using a longer wavefront may require less logic to be implemented elsewhere on the IC, such as replicating the entire wavefront or other aligner or mapper logic components less, which reduces the net performance of the IC. In one particular embodiment, 64 scoring cells in the wavefront may provide an acceptable balance of these factors.

[0265] Thus, if the wavefront is X, say 64, scoring cells wide, then the scored swath in the alignment matrix is ​​likewise 64 cells wide (measured diagonally). As long as the optimal (best-scoring) alignment path through the matrix stays within the scored swath, matrix cells outside this swath do not necessarily need to be processed or scored. Thus, for relatively small matrices used to align relatively short reads, e.g., 100- or 250-nucleotide reads, this can be a safe assumption, for example, if the wavefront sweeps completely diagonally along the predicted aligned locations of the reads.

[0266] However, in some cases, such as in large alignment matrices used to align long reads (e.g., 1000 nucleotides, 10,000 nucleotides, or 100,000 nucleotides), there may be a significant risk of accumulating indels that cause the true alignment to deviate from the scored swath, resulting in a significant risk of deviation from a perfect diagonal orientation. In such cases, it may be useful to manipulate the wavefront so that the highest set of scores is near the center of the wavefront. As a result, as the wavefront sweeps, if the highest scores begin to move in one direction or another, for example, from left to right, the wavefront shifts to track this movement. For example, if the highest score is observed in a scoring cell that is substantially to the right and above the center of the wavefront, the wavefront can be manipulated straight to the right for some distance by successive horizontal steps until the highest score returns near the center of the wavefront.

[0267] Therefore, an automatic steering mechanism can be implemented in the wavefront control logic circuit to determine a steering target location within the length of the wavefront based on the current and past scores observed in the wavefront scoring cells, and to steer the wavefront toward this target if the wavefront is off-center. More specifically, the highest-scoring location among the most recently scored wavefront locations can be used as the steering target. This is an effective method in some cases. However, in some cases, this highest-scoring location can be a poor steering target. For example, with some combinations of alignment scoring parameters, as a long indel begins and the scores begin to drop accordingly, a pattern of two higher-scoring peaks separated by a low-scoring valley can form along the wavefront, with these two peaks separating as the indel continues.

[0268] Because it is not easy to determine whether an ongoing event is an insertion or a deletion, it is important for the wavefront to track diagonally—some distance to the right for deletions, or some distance down for insertions—until successful matching begins again. However, when two divergent peaks of scores form, it is likely that one of them is slightly higher than the other, pulling the automated operation in that direction and potentially causing the wavefront to lose alignment if the actual indel were in the other direction. Therefore, a more robust approach might be to determine a threshold score by subtracting a delta value from the highest observed wavefront score, identify two end scoring cells at least equal to this threshold score, and use the midpoint between these end cells as the operational target. This tends to drive the diagonal direction between the two-peaked scoring patterns. However, other operational criteria can easily be applied that work to keep higher scores closer to the center of the wavefront. If there is a delayed response between obtaining a score from the wavefront scoring cell and making the corresponding operational decision, hysteresis can advantageously be applied to compensate for the operational decision made in the intervening time to avoid oscillating patterns of automatic wavefront operation.

[0269] One or more of these alignment procedures can be performed by any suitable alignment algorithm, such as the Needleman-Wunsch and / or Smith-Waterman alignment algorithms, which may be modified to accommodate the functions described herein. Generally, these algorithms and similar algorithms are generally performed in a similar manner in some cases. For example, as described above, these alignment algorithms typically construct virtual arrays in a similar manner, so that in various cases, the top horizontal boundary can be configured to represent a genomic reference sequence, which can be laid out across the top row of the array according to its base pair composition. Similarly, the vertical boundary can be configured to represent sequenced and mapped query sequences arranged in order down the first column, so that the order of the nucleotide sequences of these query sequences is generally matched with the nucleotide sequences of the reference sequences to which they map. A score for the probability that the relevant base of the query at a given location is located at that position relative to the reference can then be stored in the intervening cell. In performing this function, one can move a diagonal band across the matrix storing the scores in the intervening cells, and determine the probability that each base in the query is at the indicated location.

[0270] For the Needleman-Wunsch alignment function, which generates an optimal global (or semi-global) alignment to align the entire read sequence to some segment of the reference genome, the operation of the wavefront can typically be configured to sweep from the top to the bottom of the alignment matrix. Once the wavefront sweep is complete, the best score at the bottom of the alignment matrix (corresponding to the end of the read) is selected, and the alignment is back-traced to the cell at the top of the matrix (corresponding to the beginning of the read). In various cases disclosed herein, the reads can be of any length or size, and there is no need for extensive read parameters for how the alignment is performed; for example, in various cases, the reads can be the same length as a chromosome. However, in such cases, memory size and chromosome length can be limiting factors.

[0271] Regarding the Smith-Waterman algorithm, which generates an optimal local alignment to align the entire read sequence or a portion of the read sequence to some segment of the reference genome, the algorithm can be configured to find the best possible score based on the complete or partial alignment of the read. Thus, in various cases, the wavefront-scoring swath may not extend to the top and / or bottom of the alignment matrix, such as when a very long read has only a seed in its center mapped to the reference genome; however, the wavefront can generally still score from top to bottom of the matrix. Local alignment is typically achieved through two adjustments: first, alignment scores are never allowed to fall below 0 (or some other minimum value), and if an otherwise calculated cell score is negative, a score of 0 is substituted to represent the start of a new alignment. Second, the highest alignment score produced in any cell in the matrix, not necessarily along the bottom, is used as the end of the alignment. The alignment is traced back up and left through the matrix from this highest score to a score of 0, which is used as the starting point for the local alignment even if it is not in the top row of the matrix.

[0272] In view of the above, there are several different potential paths through the virtual array. In various embodiments, the wavefront starts at the upper left corner of the virtual array and moves downward toward the identifier with the highest score. For example, the results of all potential alignments can be collected, processed, correlated, and scored to determine the highest score. Once the edge of the boundary or end of the array is reached and / or the computational result that results in the highest score for all of the processed cells is determined (e.g., the highest overall score is identified), a backtrace can be performed to find the path taken to achieve that highest score.

[0273] For example, a path that results in a predicted highest score can be identified, and once identified, an audit can be performed to determine how that highest score was derived by moving backwards according to a best score alignment arrow that traces the path that resulted in the achievement of the identified highest score, e.g., as calculated by a wavefront scoring cell, etc. This backward reconstruction or backtracing involves starting from the determined highest score and working all the way backward through the table through previous cells along the path of the cell whose score resulted in the achievement of the highest score, to the first boundary, such as the start of the array or a score of 0 in the case of local alignment.

[0274] During backtracing, when a particular cell in the alignment matrix is ​​reached, the next backtracing step is to the adjacent cell immediately to the left, or immediately above, or immediately diagonally above and to the left, that contributed to the best score selected to construct the score for the current cell. In this way, the evolution of the best score can be determined, thereby discovering how the best score was achieved. A backtrace may end at a corner, edge, or boundary, or may end at a score of 0, such as in the upper left corner of the array. It is such backtracing that thus identifies the appropriate alignment, thereby producing a CIGAR string readout, e.g., 3M, 2D, 8M, 4I, 16M, etc., that represents how the sample genomic sequence derived from the individual, or a portion thereof, matches or otherwise aligns with the genomic sequence of the reference DNA.

[0275] Thus, once it has been determined where each read maps, and further once it has been determined where each read aligns, e.g., each associated read has been given a location and a quality score reflecting the probability that the location is the correct alignment, the nucleotide sequence of the subject's DNA is known, and the order of the subject's various reads and / or genomic nucleic acid sequence can then be verified, such as by performing a backtrace function that works backwards through the array to determine the identity of each and every nucleic acid that is in the proper order in the sample genomic sequence. Consequently, in some embodiments, the present disclosure is directed to a backtrace function, such as part of an alignment module that performs both alignment and backtrace functions, such as a module that may be part of a pipeline of modules, such as a pipeline that aims to take raw sequence read data, such as from a genomic sample from an individual, and map and / or align that data, which can then be sorted.

[0276] To aid in backtracing, it is useful to store a scoring vector for each scored cell in the alignment matrix, encoding the score selection decision. In classical Smith-Waterman and / or Needleman-Wunsch scoring with one-dimensional gap penalties, the scoring vector can encode four possibilities, optionally stored as 2-bit integers between 0 and 3, e.g., 0 = new alignment (null score selected), 1 = vertical alignment (score from selected cell above modified by gap penalty), 2 = horizontal alignment (score from selected cell to the left modified by gap penalty), and 3 = diagonal alignment (score from selected cell above and to the left modified by nucleotide match or mismatch score). Optionally, the calculated score for each scored matrix cell can also be stored (in addition to the best achieved alignment score, which is typically stored), but this is generally not necessary for backtracing and can consume large amounts of memory. Performing a backtrace then becomes a matter of following the scoring vector: when the backtrace reaches a given cell in the matrix, the next backtrace step is determined by that cell's stored scoring vector, e.g., 0 = end backtrace, 1 = backtrace up, 2 = backtrace left, 3 = backtrace diagonally up and to the left.

[0277] Such scoring vectors may be stored in a two-dimensional table ordered according to the dimensions of the alignment matrix, with only entries corresponding to cells scored by the wavefront being stored. Alternatively, to conserve memory, more easily record the scoring vectors as they are generated, and more easily accommodate alignment matrices of various sizes, the scoring vectors may be stored in a table sized to store scoring vectors from a single wavefront of scoring cells, e.g., with each row having 128 bits to store 64 2-bit scoring vectors from a 64-cell wavefront, and the number of rows equal to the highest number of wavefront steps in the alignment operation.

[0278] Additionally, with this option, the various wave front step directions can be recorded, e.g., by storing an extra bit, e.g., 129, in each row of the table that encodes a 0 for vertical wave front steps preceding this wave front location and a 1 for horizontal wave front steps preceding this wave front location. This extra bit can be used during backtracing to track which virtual scoring matrix location the scoring vector in each table row corresponds to, so that the appropriate scoring vector can be retrieved after each successive backtrace step. When the backtrace step is vertical or horizontal, the next scoring vector should be retrieved from the previous table row, but when the backtrace step is diagonal, the next scoring vector should be retrieved from two rows back, since the wave front had to take two steps to move from scoring any one cell to scoring the cell diagonally below and to the right of that cell.

[0279] For affine gap scoring, the scoring vector information can be expanded, for example, to four bits per scored cell. For example, in addition to the two-bit score selection direction indicator, two one-bit flags can be added: a vertical expansion flag and a horizontal expansion flag. According to the method of extending affine gap scoring to Smith-Waterman or Needleman-Wunsch or similar alignment algorithms, for each cell, in addition to a primary alignment score representing the best-scoring alignment terminating at that cell, a "vertical score" corresponding to the highest alignment score reached in the vertical step should be generated, and a "horizontal score" corresponding to the highest alignment score reached in the horizontal step should be generated. When calculating any of these three scores, the vertical step to the cell can be calculated using the larger of the primary score from the cell above minus the gap opening penalty and the vertical score from the cell above minus the gap extension penalty, and the horizontal step to the cell can be calculated using the larger of the primary score from the cell to the left minus the gap opening penalty and the horizontal score from the cell to the left minus the gap extension penalty. If the vertical score minus the gap extension penalty is selected, the vertical extension flag in the scoring vector should be set, e.g., "1"; otherwise, it should not be set, e.g., "0". If the horizontal score minus the gap extension penalty is selected, then the horizontal extension flag in the scoring vector should be set, e.g., "1", otherwise it should not be set, e.g., "0". During backtracing for affine gap scoring, whenever the backtrace takes a vertical step up from a given cell, if the vertical extension flag in that cell's scoring vector is set, then the subsequent backtrace step must also be vertical, regardless of the scoring vector of the cell above.Similarly, whenever a backtrace takes a horizontal step to the left from a given cell, if the horizontal extension flag in that cell's scoring vector is set, then the subsequent backtrace step must also be horizontal, regardless of the scoring vector of the cell to the left.

[0280] Thus, such a table of scoring vectors with some number of rows, NR, e.g., 129 bits per row for 64 cells using one-dimensional gap scoring, or 257 bits per row for 64 cells using affine gap scoring, is suitable to support backtracing after alignment scoring, where the scoring wavefront takes NR or fewer steps. For example, when aligning 300-nucleotide reads, the number of wavefront steps required may always be less than 1024, so the table may be 257 × 1024 bits, or roughly 32 kilobytes, which may often be reasonable local memory within the IC. However, when very long reads, e.g., 100,000 nucleotides, are to be aligned, the memory requirements for the scoring vectors may be quite large, e.g., 8 megabytes, which may be prohibitively expensive to include as local memory within the IC. For such support, the scoring vector information can be stored in bulk memory external to the IC, e.g., DRAM, but this can result in excessive bandwidth requirements, e.g., 257 bits per clock cycle per aligner module, which can create a bottleneck and significantly degrade aligner performance.

[0281] Therefore, it would be desirable to have a way to dispose of the scoring vector before completing the alignment, e.g., to perform additional backtraces, so that the storage requirements of the scoring vector can be kept limited, e.g., by generating additional partial CIGAR strings from earlier parts of the scoring vector's history for the alignment, so that such earlier parts of the scoring vector can then be discarded. The problem is that because the backtrace is expected to start at the highest scoring cell at the end of the alignment, which is not known until alignment scoring is complete, any backtrace started before the alignment is complete may start at an incorrect cell that is not on the final, optimal alignment path.

[0282] Thus, a method is provided for performing additional backtraces from partial alignment information, e.g., with partial scoring vector information for previously scored alignment matrix cells. Starting from a currently completed alignment boundary, e.g., the location of a particular scored wavefront, a backtrace is initiated from all cell locations on that boundary. Such backtraces from all boundary cells may be performed sequentially, or advantageously, particularly in hardware implementations, all backtraces may be performed together. It is not necessary to extract alignment notations, e.g., CIGAR strings, from these multiple backtraces; merely determine which alignment matrix locations are passed through during the backtraces. In implementing simultaneous backtraces from scoring boundaries, a number of 1-bit registers corresponding to the number of alignment cells may be utilized, e.g., all initialized to "1," indicating whether any of the backtraces pass through the corresponding locations. For each step of the simultaneous backtrace, the scoring vectors corresponding to all current "1"s in these registers, for example from a row of the scoring vector table, can be examined to determine the next backtrace step corresponding to each "1" in the register, which yields the subsequent location for each "1" in the register for the next simultaneous backtrace step.

[0283] Importantly, it is readily possible for multiple "1"s in a register to merge at a common location, which corresponds to multiple concurrent backtraces merging together onto a common backtrace path. Once two or more concurrent backtraces merge together, they remain merged forever because they utilize scoring vector information from the same cell from that point forward. Empirically and for theoretical reasons, it has been observed that with such high probability, all of the concurrent backtraces merge onto a single backtrace path in a relatively small number of backtrace steps, which may be a small multiple of the number of scoring cells in the wavefront, e.g., 8 times. For example, in a 64-cell wavefront, there is a high probability that all backtraces from the boundary of a given wavefront will merge onto a single backtrace path within 512 backtrace steps. Alternatively, it is possible and not uncommon for all backtraces to terminate within that number, e.g., 512 backtrace steps.

[0284] Thus, multiple simultaneous backtraces can be performed from a scoring boundary, e.g., the location of the scored wavefront, far enough back that they terminate or merge into a single backtrace path, e.g., in 512 or fewer backtrace steps. When they all merge into a single backtrace path, additional backtraces are possible from the location in the scoring matrix where they merge, or from any distance further back on the single backtrace path, using partial alignment information. Further backtraces from the merge point, or from any distance further back, are initiated by the usual single backtrace method, including recording the corresponding alignment notation, e.g., a partial CIGAR string. This additional backtrace, and e.g., the partial CIGAR string, must be part of any possible final backtrace, e.g., the complete CIGAR string, that results after the alignment is complete; this is true as long as such final backtrace does not end before reaching the scoring environment where the concurrent backtrace began, since when it reaches a scoring boundary, it must follow one of the concurrent backtrace paths and merge into a single backtrace path that is now extracted additively.

[0285] Thus, for example, all scoring vectors for matrix regions corresponding to additionally extracted backtraces in all table rows for wavefront locations preceding the start of the extracted single backtrace can be safely discarded. When the final backtrace is performed from the highest scoring cell, if the final backtrace terminates before reaching a scoring boundary (or alternatively, terminates before reaching the start of the extracted single backtrace), the additional alignment notation, e.g., a partial CIGAR string, can be discarded. If the final backtrace continues to the start of the extracted single backtrace, the alignment notation, e.g., a CIGAR string, of the final backtrace can then be spliced ​​into the additional alignment notation, e.g., a partial CIGAR string.

[0286] Furthermore, in very long alignments, the process of running simultaneous backtraces from a scoring boundary, e.g., the location of a scored wavefront, until all backtraces terminate or merge, and then running a single backtrace with extracted alignment representations, can be repeated multiple times from various successive scoring boundaries. The additional alignment representations from each successive additional backtrace, e.g., the partial CIGAR string, can then be spliced ​​into the accumulated previous alignment representation unless the new simultaneous or single backtrace terminates early, in which case the accumulated previous alignment representation can be discarded. The final, final backtrace similarly splices alignment representations into the latest accumulated alignment representation for the complete backtrace description, e.g., the CIGAR string.

[0287] Thus, in this scheme, memory for storing scoring vectors can be kept limited, assuming that concurrent backtraces always merge together in a limited number of steps, e.g., 512 steps. In the rare cases where concurrent backtraces do not merge or terminate in a limited number of steps, various exceptional actions can be taken, including failing the current alignment, or repeating the alignment with higher or no restrictions, possibly by different or conventional methods, such as storing all scoring vectors for a complete alignment, such as all for a complete alignment, in external DRAM, etc. In some variations, failing such an alignment may be appropriate because it is extremely rare, and even more rare that such a failed alignment is the best-scoring alignment that will be used in reporting the alignment.

[0288] In an optional variation, the storage of the scoring vectors can be physically or logically divided into a number of separate blocks, e.g., 512 lines each, with the last line in each block being used as the scoring boundary for starting the simultaneous backtraces. Optionally, simultaneous backtraces may be required to terminate or merge within a single block, e.g., 512 steps. Optionally, if simultaneous backtraces merge in fewer steps, the merged backtraces may still continue throughout the block until starting to extract a single backtrace in the previous block. Thus, after the scoring vectors are completely written to block N, writing to block N+1 may begin, and simultaneous backtraces may begin in block N, followed by extraction of a single backtrace and alignment notation in block N-1. If the speeds of simultaneous backtracing, single backtracing, and alignment scoring are all similar or identical and can be executed simultaneously, for example in parallel hardware in an IC, then a single backtrace in block N-1 can be simultaneous with the scoring vector filling block N+2, and block N-1 can be freed and reused when block N+3 is filled.

[0289] Thus, in such an implementation, a minimum of four scoring vector blocks may be utilized, and may be utilized cyclically. Thus, the overall scoring vector storage for the aligner module may be, for example, four blocks of 257×512 bits each, or approximately 64 kilobytes. In one variation, if the current best alignment score corresponds to a block earlier than the current wavefront location, this block and the previous block may be saved rather than reused, so that the final backtrace can start from this location if it remains the best score. Having two extra blocks left saved in this manner results in a minimum of, for example, six blocks. In another variation, to support overlapping alignments, additional blocks, for example, one or two additional blocks, for example, eight blocks total, for example, approximately 128 kilobytes, may be utilized as the scoring wavefront gradually passes from one alignment matrix to the next as described above. Thus, if such a limited number of blocks, e.g., four or eight blocks, are used periodically, alignment and backtracing of reads of any length, e.g., 100,000 nucleotides or entire chromosomes, is possible without using external memory for scoring vectors.

[0290] As explained above, certain regions of DNA are genes, which code for proteins or functional RNA. Each gene exists on one strand of a double-stranded DNA duplex as a series of exons (coding segments), often separated by introns (non-coding segments). Some genes have only a single exon, but most have several exons (separated by introns), and some have hundreds or even thousands of exons. Exons are generally several hundred nucleotides long but can be as short as a single nucleotide or as long as tens or hundreds of thousands of nucleotides. Introns are generally several thousand nucleotides long, with some exceeding one million nucleotides.

[0291] Genes can be transcribed by the enzyme RNA polymerase into messenger RNA (mRNA) or other types of RNA. The immediate RNA transcript is a single-stranded copy of the gene, except that thymine (T) bases in DNA are transcribed into uracil (U) bases in RNA. However, shortly after this copy is produced, copies of introns are usually excised by the spliceosome, leaving copies of exons joined together at "splice sites" (which are not immediately apparent after this point). RNA splicing does not always occur in the same way; one or more exons may be excised, and splice sites may not fall on the most common intron / exon boundaries. Thus, a single gene can produce multiple different transcribed RNA segments, a process sometimes known as alternative splicing.

[0292] The spliced ​​mRNA leaves the cell nucleus (in eukaryotes) and is transported to the ribosome, which decodes the mRNA into a protein, with each group of three RNA nucleotides (a codon) coding for one amino acid. In this way, genes in DNA serve as the original instructions for making proteins.

[0293] RNA splicing tends to occur at consistent exon / intron boundaries, which are characterized by typical sequence content, especially near the end of the intron. Specifically, the first two or last two bases of an intron, called the intron motif, follow one of only three "canonical" intron motifs in the vast majority of cases (approximately 99.9%). The most common canonical intron motif is "GT / AG," meaning that the first two bases of the intron are "G" and "T" and the last two bases are "A" and "G." The GT / AG motif occurs in approximately 98.8% of cases. Other canonical intron motifs are GC / AG, which occurs in approximately 1.0% of cases, and AT / AC, which occurs in approximately 0.1% of cases. These canonical motifs and their occurrence are reasonably consistent across multiple species but may not be universal.

[0294] Not all genes are transcribed, and those that are transcribed may be transcribed at different frequencies. Many factors can influence whether and how frequently a given gene is transcribed into RNA. Some of these factors are hereditary, some vary from tissue to tissue depending on cellular specialization, and some change over time due to environmental conditions or disease. Thus, two cells with exactly the same DNA may produce significantly different types and amounts of proteins and functional RNA. This means that sequencing (reading) the RNA present in one or more cells provides different information than sequencing the DNA. A more complete picture of cellular conditions and activity is provided by combining DNA and RNA sequencing.

[0295] Transcriptome-wide RNA sequencing is generally carried out by first selecting target RNA, such as protein-coding RNA, and then using reverse transcriptase to convert the RNA segment back into a complementary DNA (cDNA) strand. This DNA can be amplified using polymerase chain reaction (PCR) and / or fragmented into the length of a desired sequence distribution. The DNA fragments are then sequenced using a DNA sequencer, such as a "shotgun" next-generation sequencer.

[0296] The resulting DNA reads are either reverse-complemented copies of the original RNA strand, except that "U" is again replaced with "T," or forward copies. While some library preparation and sequencing protocols can maintain or flag the relative orientation of the sequenced DNA strand to the original RNA, in typical protocols, approximately 50% of the sequenced DNA is reverse-complemented to the original RNA, with no direct (but indirect) indication of orientation.

[0297] DNA reads from RNA-seq protocols differ from other methods of genome-wide or exon-wide DNA sequencing. First, aside from contamination, only transcribed RNA is sequenced, so non-coding DNA and inactive genes are generally not represented. Second, the amount of sequenced reads corresponding to various genes is related to the biological transcription rates of those genes. Third, due to intron splicing, RNA-seq reads tend to skip intronic (non-coding) segments within genes.

[0298] RNA-seq reads are typically processed quite differently from DNA reads. Both types of reads are typically mapped and aligned to a reference genome, but the techniques for mapping and aligning DNA and RNA are different (see next section). After mapping and alignment, reads are typically sorted by their mapped reference location for both DNA and RNA. Duplicate marking, which is optional for DNA processing, is not typically used for RNA-seq data.

[0299] After this, the DNA reads are typically processed by a variant caller to identify differences between the sampled DNA and the reference genome. RNA-seq reads are not typically used for variant calling, although this is sometimes done. More commonly, aligned and sorted RNA reads are analyzed to determine which genes are expressed and in what relative abundance, or which of various alternatively spliced ​​transcripts are produced and in what relative abundance. This analysis typically involves counting how many reads align to various genes, exons, etc., and may also involve transcript assembly (reference-based or de novo) to infer from relatively short RNA-seq reads how longer RNA transcripts are likely spliced ​​from the DNA.

[0300] Gene, exon, or transcript expression analysis is often extended to differential expression analysis, in which RNA-seq data from multiple samples, often from two or more different classes (subpopulations or phenotypes), are compared to quantify how differentially expressed a gene, exon, or transcript is in the different classes. This can involve calculating the likelihood of the "null hypothesis" that corresponding expression levels in the different classes were the same, as well as estimating "fold changes" in expression between samples, e.g., differences of 8 or more or 10 or more folds.

[0301] In many applications of DNA or RNA sequencing, an early processing step is to map and align reads to a reference genome. Typically, a DNA-oriented reference genome with "T" but no "U" is used for both DNA and RNA sequencing, especially considering that RNA-seq usually involves reve...

Claims

1. 1. A genomics infrastructure for on-site or cloud-based DNA or RNA processing and analysis, comprising: a platform application programming interface (API) defining inputs for receiving result data from secondary processing of multiple reads of DNA or RNA sequence data; a plurality of user-selectable DNA or RNA processing pipelines, each having an input defined in accordance with the platform API to receive the result data from the secondary processing, wherein the plurality of DNA or RNA processing pipelines have a common pipeline API that defines tertiary processing operations on the result data from the secondary processing received in accordance with the platform API, each of the plurality of DNA or RNA processing pipelines configured to perform a subset of the tertiary processing operations, and a user-selected set of the DNA or RNA processing pipelines configured to output the result data from the tertiary processing in accordance with the pipeline API.

2. 2. The genomic infrastructure of claim 1, further comprising a plurality of user-selectable DNA or RNA analysis applications stored in one or more application repositories, wherein each of a selected set of the plurality of DNA or RNA analysis applications is accessible by a computer via electronic media from an on-site or cloud-based application repository for execution by a computer processor to perform targeted analysis of DNA or RNA data from the result data of the tertiary processing, and wherein each of the plurality of genomic analysis applications is defined by the application API for receiving the result data of the tertiary processing, performing the targeted analysis of the DNA or RNA data from the result data of the tertiary processing, and outputting the result data from the targeted analysis to one of one or more genomic databases in accordance with the application API.

3. 2. The genome infrastructure of claim 1, wherein the plurality of user-selectable genome processing pipelines are selected from a set of DNA or RNA pipelines consisting of a genome processing pipeline, an epigenome processing pipeline, a metagenomics processing pipeline, a joint genotyping processing pipeline, and a Genome Analysis Toolkit (GATK) processing pipeline.

4. 3. The genomic infrastructure of claim 2, wherein the plurality of user-selectable genomic analysis applications are selected from a set of genomic analysis applications consisting of a non-invasive prenatal testing application, a neonatal intensive care unit application, a cancer analysis application, a home-test (LDT) application, and an agricultural and biological analysis application.

5. a pipeline application programming interface (API) defining inputs for receiving result data from secondary and / or tertiary processing of multiple reads of genomic sequence data; a plurality of user-selectable genomic analysis applications stored in one or more application repositories, each of a selected set of the plurality of genomic analysis applications being accessible by a computer from the application repository via electronic media for execution by a computer processor to perform targeted analysis of genomic data from result data of the secondary and / or tertiary processing, and each of the plurality of genomic analysis applications defined by the application API for receiving the result data of the secondary and / or tertiary processing, performing the targeted analysis of the genomic data from the result data of the tertiary processing, and outputting the result data from the targeted analysis to one of one or more genomic databases in accordance with the application API.

6. 6. The genome analysis platform of claim 5, further comprising a plurality of user-selectable genome processing pipelines, each having an input defined according to a platform API for receiving the result data from the secondary processing, the plurality of genome processing pipelines having a common pipeline API that defines tertiary processing operations on the result data from the secondary processing received according to the platform API, each of the plurality of genome processing pipelines configured to perform a subset of the tertiary processing operations, and a user-selected set of the genome processing pipelines configured to output result data from the tertiary processing according to the pipeline API.

7. 6. The genomic analysis platform of claim 5, wherein the one or more genomic databases include an electronic medical record database.

8. 6. The genome analysis platform of claim 5, wherein the one or more genome databases include a government database.

9. 6. The genome analysis platform of claim 5, wherein at least one of the one or more genome databases is a cloud-based data repository.

10. 1. A genomics infrastructure for on-site or cloud-based genome processing and analysis, comprising: a platform application programming interface (API) defining an input for receiving result data from secondary processing of multiple reads of genomic sequence data; a plurality of user-selectable genome processing pipelines, each having an input defined in accordance with the platform API for receiving the result data from the secondary processing, the plurality of genome processing pipelines having a common pipeline API that defines tertiary processing operations on the result data from the secondary processing received in accordance with the platform API, each of the plurality of genome processing pipelines configured to perform a subset of the tertiary processing operations, and a user-selected set of the genome processing pipelines configured to output the result data of the tertiary processing in accordance with the pipeline API; 1. A genomic infrastructure comprising: a plurality of user-selectable genomic analysis applications stored in one or more application repositories, each of a selected set of the plurality of genomic analysis applications accessible from the application repository by a computer via electronic media for execution by a computer processor to perform targeted analysis of genomic data from the results data of the tertiary processing, wherein each of the plurality of genomic analysis applications is defined by the application API for receiving the results data of the tertiary processing, performing the targeted analysis of the genomic data from the results data of the tertiary processing, and outputting the results data from the targeted analysis to one of one or more genomic databases in accordance with the application API.

11. The secondary treatment is receiving the plurality of reads of genomic sequence data and one or more gene reference sequences; and processing the plurality of reads of genome sequence data to map and align at least some of the plurality of reads of genome sequence data according to the one or more genetic reference sequences.

12. The genomic infrastructure of claim 10, wherein the resulting data from the secondary processing comprises genomic data reads.

13. 13. The genomic infrastructure of claim 12, wherein the resulting data from the secondary processing comprises mapped and aligned reads from the plurality of reads of genomic data.

14. 14. The genomic infrastructure of claim 13, wherein the resulting data from the secondary processing comprises one or more variant call files generated from the mapped and aligned reads.

15. 1. A genomics infrastructure for on-site or cloud-based genome processing and analysis, comprising: a bioinformatics processing platform having a memory storing one or more DNA or RNA reference sequences and one or more indexes for said one or more DNA or RNA reference sequences; and an integrated circuit formed of a set of pre-configured hardwired digital logic circuits interconnected by a plurality of physical electrical interconnects, said integrated circuit having an input for receiving a plurality of reads of DNA or RNA data and a memory interface for accessing said one or more DNA or RNA reference sequences and said indexes, said hardwired digital logic circuits arranged as a set of processing engines each formed of a subset of said hardwired digital logic circuits for performing one pre-configured step of secondary processing on said plurality of reads of DNA or RNA data according to said DNA or RNA reference sequences and said indexes, said integrated circuit further having an output for outputting result data from said secondary processing according to a platform application programming interface (API); a plurality of user-selectable genome processing pipelines, each having an input defined in accordance with the platform API for receiving the result data from the secondary processing by the bioinformatics processing platform, wherein the plurality of DNA or RNA processing pipelines have a common pipeline API that defines tertiary processing operations on the result data from the secondary processing, each of the plurality of genome processing pipelines configured to perform a subset of the tertiary processing operations, and a user-selected set of the genome processing pipelines configured to output the result data of the tertiary processing in accordance with the pipeline API; 1. A genomic infrastructure comprising: a plurality of user-selectable DNA or RNA analysis applications stored in one or more application repositories, each of a selected set of the plurality of genomic analysis applications accessible from the application repository by a computer via electronic media for execution by a computer processor to perform targeted analysis of genomic data from the result data of the tertiary processing, wherein each of the plurality of genomic analysis applications is defined by the application API for receiving the result data of the tertiary processing, performing the targeted analysis of the DNA or RNA data from the result data of the tertiary processing, and outputting the result data from the targeted analysis to one of one or more genomic databases in accordance with the application API.

16. 16. The genome infrastructure of claim 15, wherein the resulting data from the bioinformatics processing platform comprises mapped and aligned reads from the plurality of reads of DNA or RNA data.

17. 17. The genomic infrastructure of claim 16, wherein the resultant data from the bioinformatics processing platform comprises one or more variant call files generated from the mapped and aligned reads.

18. 1. A system for performing a portion of a sequence analysis pipeline on a plurality of reads of DNA or RNA data using DNA or RNA reference sequence data, wherein each read of the plurality of reads of DNA or RNA data and the DNA or RNA reference sequence data represents a sequence of nucleotides, and wherein an integrated circuit comprises: a memory for storing the plurality of reads of DNA or RNA data and the DNA or RNA reference sequence data; and a field programmable gate array (FPGA), wherein the FPGA comprises a set of pre-configured hardwired digital logic circuits, the hardwired digital logic circuits interconnected by a plurality of physical electrical interconnects, one or more of the plurality of physical electrical interconnects comprising a memory interface for accessing the memory; the hardwired digital logic circuits arranged as a set of processing engines, each processing engine formed of a subset of the hardwired digital logic circuits for performing one or more steps in the sequence analysis pipeline on the plurality of reads of DNA or RNA data; and the set of processing engines comprises a first hardwired configuration mapping module for accessing one or more of the plurality of reads of DNA or RNA data and the DNA or RNA reference sequence data, comparing the sequence of nucleotides in at least one of the plurality of reads of DNA or RNA data to the sequence of nucleotides in the DNA or RNA reference sequence data, and mapping the one or more of the plurality of reads of DNA or RNA data to the DNA or RNA reference sequence data to produce one or more mapped DNA or RNA reads.

19. 19. The system of claim 18, wherein the FPGA further comprises a second hardwired configuration for accessing at least one of the mapped reads of DNA or RNA data and the DNA or RNA reference sequence data, comparing the sequence of nucleotides in the at least one mapped read of DNA or RNA data to the sequence of nucleotides in the DNA or RNA reference sequence data, and aligning the one or more mapped reads of DNA or RNA data to the DNA or RNA reference sequence data.

20. 20. The system of claim 19, wherein the FPGA further comprises a third hardwired sorting module for sorting the mapped and aligned DNA or RNA reads.

Citation Information

Patent Citations

  • Apparatus for searching primer set, method and program for searching primer set

    JP2011062085A

  • System and method of genome data processing using in-memory database system and real-time analysis

    JP2014146318A

  • Genome read alignment of high efficiency in in-memory database

    JP2014146319A

  • Lead molecule cross-reaction prediction and optimization system

    JP2015072691A

  • Integrated consumer genomic services

    US20150227697A1