Flexible seed extension for hash table genomic mapping

The hash table index with interval and extension records in genome mapping systems optimizes seed extension, reducing resource consumption and improving accuracy by dynamically determining accurate matching positions, addressing inefficiencies and inaccuracies in conventional methods.

JP2025097984AActive Publication Date: 2025-07-01ILLUMINA INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025026056
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-24
Filing Date
2025-02-20
Publication Date
2025-07-01
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

Conventional seed extension methods in genome mapping and alignment systems consume excessive computing resources and power, and are prone to unmapped reads, high-accuracy mis-mapping, and fixed maximum match problems, leading to inefficiencies and inaccuracies.

Method used

A hash table index is constructed with interval and extension records to facilitate flexible seed extension, reducing resource consumption and improving accuracy by dynamically determining accurate matching positions and optimizing seed extension processes.

Benefits of technology

The solution reduces computing resource usage and power consumption while enhancing the accuracy of genome mapping and alignment, addressing the inefficiencies and inaccuracies of conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097984000001_ABST
    Figure 2025097984000001_ABST
Patent Text Reader

Abstract

To improve the performance of genomic mapping and aligning systems.SOLUTION: Methods, systems, and apparatuses, including computer programs for generating and using a hash table configured to improve mapping of reads are disclosed that include: acquiring a first seed of K nucleotides from a reference sequence; generating a seed extension tree having nodes, each node of the nodes corresponds to (i) an extended seed that is an extension of the first seed and has a nucleotide length of K* and (ii) one or more locations, in a seed extension table, that include data describing reference sequence locations that match the extended seed; and for each node, storing interval information at a location of the hash table that corresponds to the index key for the extended seed, the interval information references one or more locations in the seed extension table that include reference sequence locations that match the extended seed associated with the node.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 852,965, filed May 24, 2019, which is hereby incorporated by reference in its entirety.

Background Art

[0002] A nucleic acid sequencer is an instrument configured to automate the process of nucleic acid sequencing. Nucleic acid sequencing is the process of determining the order of nucleotides in a nucleic acid sequence. Nucleic acids can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

[0003] A nucleic acid sequencer is configured to receive a nucleic acid sample and generate output data called one or more "reads" that represent the order of nucleotides in the nucleic acid sample. Nucleotides in a DNA sample can include one or more nucleotide bases in any combination of guanine (G), cytosine (C), adenine (A), and thymine (T). Nucleotides in an RNA sample can include one or more bases in any combination of G, C, A, and uracil (U).

[0004] Reads generated by a DNA sequencer can be mapped to the known sequence of nucleotides of a reference genome using a mapping and alignment engine. The mapping of a read to the sequence of nucleotides of a reference genome can be achieved by a mapping and alignment engine using a hash table index.

Summary of the Invention

[0005] This disclosure describes the construction and use of a hash table index to facilitate flexible seed extension in order to improve the performance of genome mapping and alignment systems. In particular, the disclosure is used to perform flexible seed extension in a manner that (i) reduces the consumption of computing resources and power and (ii) solves problems associated with conventional seed extension methods described herein. To achieve these advantages, the disclosure provides, among other things, "interval records" that can be stored at hash table positions.

[0006] Aspects of the disclosure enable a mapping and alignment unit to use interval records alone or in combination with one or more extension records to reduce the number of matching positions processed by the mapping and alignment unit through seed extension, while at the same time using dynamic seed extension to determine whether the matching reference positions identified are accurate or, in some cases, whether seed extension using one or more extension records should even be performed. This provides the mapping and alignment unit with the flexibility to do so. This results in a mapping and alignment unit that uses less power and fewer computing resources while also being more accurate than other mapping and alignment units that utilize the conventional seed extension technique itself.

[0007] In one aspect, the disclosure provides a method for generating a hash table for mapping sample reads to a reference. In one aspect, the method comprises obtaining, by a computer system, a first seed of nucleotides from a reference sequence, wherein the first seed has a length of K nucleotides, determining, by the computer system, that the first seed matches more than a predetermined number of reference sequence positions, and generating, by the computer system, a seed extension tree having a plurality of nodes based on determining that the first seed matches more than a predetermined number of reference sequence positions, wherein each node of the plurality of nodes is (i) an extension of the first seed and K* An extended seed having a nucleotide length of K * wherein K is one or more nucleotides greater than K, an extended seed, and (ii) one or more positions corresponding to one or more positions in the seed extension table that include data describing a reference array position that matches the extended seed, generating, for each node of a plurality of nodes, storing interval information at a position in a hash table corresponding to an index key of the extended seed by a computer system, the interval information including data describing a reference array position that matches the extended seed associated with the node, referring to one or more positions in the seed extension table, and storing.

[0008] Other aspects include corresponding systems, devices, and computer programs for performing the actions of the methods as disclosed herein, as defined by instructions encoded on a computer-readable storage device.

[0009] These and other aspects may optionally include one or more of the following features. For example, in some implementations, each of the matching reference array positions includes K nucleotides of the first seed.

[0010] In some implementations, the method further includes obtaining, by a computer system, a second seed of nucleotides from a reference array different from the first seed, determining, by the computer system, that the second seed does not match more than a predetermined number of reference array positions, and based on determining, by the computer system, that the second seed does not match more than a predetermined number of reference array positions, obtaining, by the computer system, data describing each of the reference array positions that match the second seed, and storing, by the computer system, the data describing the reference array positions that match the second seed at a second position in a hash table corresponding to an index key of the second seed.

[0011] In some implementations, one or more positions within a seed extension table that contain data describing a reference array position matching an extended seed can include a contiguous interval of positions within the seed extension table that contain data describing a reference array position matching the extended seed.

[0012] In some implementations, one or more positions within a seed extension table that contain data describing a reference array position matching an extended seed associated with a node can include a contiguous interval within the extension table of positions that match the reference array position of the extended seed associated with the node.

[0013] In some implementations, obtaining, by a computer system, a first seed of nucleotides from a reference array, where the first seed represents an array of nucleotides having a nucleotide length of K nucleotides, can include the computer system determining a position of a seed access window within the reference array and the computer system obtaining a subset of the reference array identified by the seed access window.

[0014] In some implementations, the method includes, by a computer system, adjusting a seed extension window forward along the reference array by only K nucleotides to identify a second seed of nucleotides from the reference array having a nucleotide length of K nucleotides, obtaining, by the computer system, the second seed from the reference array, determining that the second seed matches more than a predetermined number of reference array positions, and based on determining that the second seed matches more than a predetermined number of reference array positions, generating, by the computer system, a second seed extension tree having a plurality of second nodes, where each second node of the plurality of second nodes is (i) an extension of the second seed and has a second extended seed having a nucleotide length of K * and having a nucleotide length of K *a second extended seed that is one or more nucleotides larger than K, and (ii) one or more second positions including data in the second seed extension table that describes a reference sequence position that matches the second extended seed, and generating, corresponding thereto, for each second node of the plurality of second nodes, storing, by a computer system, second interval information at a position in a hash table corresponding to an index key of the second extended seed, the second interval information including data that describes a reference sequence position that matches the second extended seed associated with the second node, referring to one or more positions in the second seed extension table, and storing, can further be included.

[0015] In some implementations, the method further includes, for each node of the plurality of nodes, determining, by a computer system, whether the node of the seed extension tree is a leaf node, and, based on determining, by the computer system, that the node of the extension tree is not a leaf node, storing, by the computer system, an extension record at a position in a hash table corresponding to an index key of the extended seed.

[0016] In some implementations, the extension record includes one or more instructions that, when executed by a computer system, cause the computer system to add one or more additional nucleotides to the seed associated with the extension record.

[0017] In some implementations, the method can further include not storing, by the computer system, an extension record at a position in a hash table corresponding to an index key of the extended seed, based on determining, by the computer system, that the node extension tree node is a leaf node.

[0018] In some implementations, the method can further include generating, by a computer system, a seed extension table. In such an implementation, generating the seed extension table can include, by the computer system, identifying each seed of a reference array that matches a first seed, and storing, in the seed extension table, data that identifies the identified seeds by the computer system.

[0019] In some implementations, the method can further include sorting, by a computer system, the identified seeds within the seed extension table.

[0020] In some implementations, the method can further include generating, by a computer system, a hash table installation package, the hash table installation package including instructions that cause one or more computers that receive the hash table installation package to install a hash table in a memory accessible by a programmable logic circuit.

[0021] In some implementations, the hash table installation package can include the seed extension table, and the hash table installation package can include instructions that cause (i) a programmable logic circuit or (ii) another computer to store the seed extension table in a memory device accessible by the programmable logic circuit.

[0022] In some implementations, a computer system provides the hash table installation package to another computer.

[0023] In some implementations, the other computer can include (i) a computer configured to communicate with a programmable logic circuit, or (ii) the programmable logic circuit itself.

[0024] In some implementations, the computer system can include multiple computers.

[0025] In another aspect, the present disclosure provides a method for improving the mapping of sample reads to a reference array using a hash table. In one aspect, the method is to execute a query of the hash table by a mapping and alignment unit, the query including a first seed, the first seed including a subset of nucleotides obtained from a particular read of the sample read, and to obtain, by the mapping and alignment unit, a response to the executed query including information stored at a position in the hash table determined to be responsive to the query, and to determine, by the mapping and alignment unit, whether the response to the executed query includes (i) an extension record, (ii) a gap record, or (iii) one or more matching reference array positions, and to determine, by the mapping and alignment unit, based on determining that the response to the query executed by the mapping and alignment unit includes (i) an extension record and (ii) a gap record, whether the extension table is accessed to obtain one or more matching reference array positions in the extension table referenced by the gap record, and to determine, by the mapping and alignment unit, based on determining that the extension table is not accessed, whether to store, in the memory device, as information describing the best gap candidate, first information describing the gap record, and to generate, by the mapping and alignment unit, a first extended seed that is an extension of the first seed using the extension record, and to generate, by the mapping and alignment unit, a subsequent hash query including the first extended seed, and to execute, by the mapping and alignment unit, the subsequent hash query of the hash table.

[0026] Other versions include corresponding systems, devices, and computer programs for performing actions of a method defined by instructions encoded on a computer-readable storage device.

[0027] These and other aspects of the present disclosure can optionally include one or more of the following features. For example, in some implementations, the method includes accessing, by a mapping and alignment unit, an extension table based on determining that the extension table is accessed, and obtaining, by the mapping and alignment unit, one or more matching reference array positions in the extension table referenced by an interval record; and adding, by the mapping and alignment unit, the one or more matching reference array positions to a seed match set.

[0028] In some implementations, the method includes determining, by a mapping and alignment unit, that a response to an executed query includes one or more matching reference array positions; and adding, by the mapping and alignment unit, the one or more matching reference array positions to a seed match set based on determining that the response to the executed query includes one or more matching reference array positions.

[0029] In some implementations, the mapping and alignment unit can include determining, by the mapping and alignment unit, whether to store, in a memory device, first information describing an interval record as information describing a best interval candidate; determining, by the mapping and alignment unit, that there is no previous information describing an interval record as a best interval candidate for a particular read; and storing, by the mapping and alignment unit, the first information describing the interval record as information describing a best interval candidate in the memory device.

[0030] In some implementations, the method includes obtaining, by a mapping and alignment unit, a response to a subsequent executed query that includes information stored at a position in a hash table determined to respond to a query; determining, by the mapping and alignment unit, that the response to the subsequent executed query includes (i) a second extended record, (ii) a second interval record, or (iii) one or more matching reference array positions; determining, by the mapping and alignment unit, based on determining that the response to the subsequent executed query includes (i) a second extended record or (ii) a second interval record, whether the extension table is accessed by the mapping and alignment unit to obtain one or more matching reference array positions within the extension table referenced by the second interval record; determining, by the mapping and alignment unit and using one or more heuristic rules, whether second information describing the second interval record or first information describing a best interval candidate is used as the best interval candidate based on determining that the extension table is not accessed; generating, by the mapping and alignment unit, a second extended seed that is an extension of a first extended seed using the second extended record; generating, by the mapping and alignment unit, a third hash query that includes the second extended seed; and executing, by the mapping and alignment unit, a third query of a hash table that includes the second extended seed.

[0031] In some implementations, the mapping and alignment unit, using one or more heuristic rules, determines whether the second information describing the second interval record or the first information describing the best interval candidate is used as the best interval, and selects either the second information describing the second interval record or the first information describing the best interval candidate record based on a plurality of factors including: (i) the number of matching reference array positions returned by each of the interval record and the second interval record; (ii) a predetermined threshold level of the reference array positions; or (iii) the length of each seed that reaches the hash position storing the interval record and the second interval record. This can include selection based on these factors.

[0032] In some implementations, the interval record refers to a plurality of positions including data describing reference array positions in the seed extension table that match the first seed of the query.

[0033] In some implementations, the plurality of positions including data describing reference array positions in the seed extension table that match the first seed of the query can include consecutive intervals of reference array positions in the extension table that match the first seed of the query.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0035] These and other aspects of the present disclosure are discussed in more detail in the following detailed description with reference to the accompanying drawings and the claims.

Brief Description of the Drawings

[0036]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0037] This disclosure describes the construction and use of a hash table index to facilitate flexible seed extension in order to improve the performance of a genome mapping and alignment system. As used herein, the term "seed" refers to a subset of contiguous nucleotides present in a nucleic acid read ("read") or nucleic acid reference sequence ("reference sequence"). By way of example, a short seed for a read can have, for example, 21 bases or nucleotides extracted from a 150-base or nucleotide read generated by, for example, a nucleic acid sequencer ("sequencer") based on a biological sample input to the sequencer. Such a short seed can match positions in hundreds, thousands, hundreds of thousands, or even more reference sequences. A seed of a reference sequence can include a subset of contiguous nucleotides from the reference sequence that represents a reference sequence position. Identification of such a large number of reference sequence positions that match a particular short seed of a read can occur for multiple reasons, including the occurrence of repetitive sequences such as "...ATATAT..." that can occur at many positions within the reference sequence. Alternatively, or in addition, such a large number of matching reference sequence positions can occur because many nearby copies of the genomic sequence can occur within the reference sequence.

[0038] These large numbers of reference sequence positions that match a particular short seed can cause strain on conventional mapping and alignment units using a conventional hash table index because the mapping and alignment engine can be forced to process a large number of matches, resulting in unnecessary consumption of computing resources, including overuse of processing resources and memory resources, and waste of power used to power the processing resources, memory devices, and cooling units used to cool the processing resources and memory resources, or any combination thereof.

[0039] Conventional methods have been used to address potential problems arising from the identification and processing of a large number of reference array positions that match a short seed. For example, conventional methods that iteratively extend a short seed by utilizing extended records stored at hash table positions have been used. As such a method, for example, there is the one described in U.S. Patent No. 10,083,276, which is incorporated herein by reference, and which can return an "extended record" stored at the position of the hash table corresponding to the seed of the hash query. By using the extended record to symmetrically increase the length of the seed in the received hash query by adding one or more bases or nucleotides to each end of the seed, an extended seed can be created. The conventional system can then re-query the hash table using another hash query that includes the extended seed. This other hash query with the extended seed is likely to correspond to a hash position that identifies fewer reference positions that match the extended seed because the extended seed is longer. This iterative process can continue until (i) the resulting set of matches shrinks sufficiently such that it contains fewer than a threshold number of reference array positions that match the extended seed, (ii) the set of matches becomes empty, (iii) the maximum seed extension is reached, or (iv) the extension moves beyond the edge of the read based on the short seed such that the next extension is not possible. Strictly speaking, in the conventional system, the mapping and alignment unit can obtain a non-empty set of matching reference positions only when the iterative process ends in the above-described manner (i).

[0040] These conventional methods can be useful for reducing the amount of reference array positions that match a short seed. However, these conventional methods are plagued by three distinct problems.

[0041] First, the conventional method is likely to fall into the "unmapped read problem". The unmapped read problem occurs when the conventional seed extension method returns zero matches for an extended seed. Such a set of zero match results can occur when the extended seed incorporates a variant such as an SNP, or when the extended seed overruns the edge of the corresponding read. If such a scenario occurs for each seed location of a read using the conventional method, the read may not be mapped.

[0042] Second, the conventional method is likely to fall into the "high-accuracy mis-mapping problem". Such a high-accuracy mis-mapping problem occurs when the extended seed contains a variant such as an SNP but still matches one or more reference positions. Such a mapping can be characterized by a high-accuracy score such as a high MAPQ score even if the extended seed is inappropriately mapped. If this occurs for each seed location of a read using the conventional method, the read can potentially be mis-mapped with high accuracy. Due to such a mapping, conversely, evidence may be lost. High-accuracy mis-mapping may be more harmful to the overall mapper accuracy than low-accuracy mis-mapping. The MAPQ score can include a quality score that quantifies the probability that a mapped read is mis-placed.

[0043] Third, the conventional method is likely to fall into the "fixed maximum match problem". Generally, the hash table constructed for seed extension uses a maximum match parameter M such as M = 16. This parameter ensures that the leaf nodes of the seed extension tree do not exceed the parameter of M. However, some applications require a different maximum match parameter M such as M = 64 *Benefits can be obtained from using it. In the case of the conventional seed extension method, the seed is continuously extended iteratively until the leaf node is reached. Therefore, in an application using the conventional method, when the hash table is not reconstructed such that the maximum match parameter M is set to 64, the extension of the seed cannot be stopped when a matching set with M = 64 is achieved.

[0044] Using the innovative aspects of the present disclosure, a flexible seed extension can be performed in a manner that (i) reduces the consumption of computing resources and power as described herein and (ii) solves problems associated with conventional seed extension methods such as those problems described above. To achieve these advantages, the present disclosure provides, among other things, "interval records" that can be stored at hash table positions. An interval record identifies, for a particular seed, a contiguous set of reference array positions that match the particular seed and are stored in the seed extension table. When executing a hash query that identifies a particular seed, the mapping and alignment unit can, based on the content of the hash position in response to the query, (i) extend the seed based on the seed extension record stored at the hash position, (ii) store an interval record that identifies the reference position that matches the particular seed within the seed extension table, or (iii) determine whether to access the reference array position identified by the interval record stored at the hash position within the seed extension table. In some implementations, combinations of these operations, such as extending the seed and storing the interval, can be performed.

[0045] Thus, by using one or more extension records in conjunction with the interval records, the mapper and aligner are able to reduce the number of matching positions processed by the mapping and alignment unit through seed extension, while at the same time providing the mapping and alignment unit with the flexibility to determine whether the matching reference positions identified using dynamic seed extension are accurate, or in some cases, whether seed extension using one or more extension records should even be performed. This results in a mapping and alignment unit that uses less power and fewer computational resources while also being more accurate than other mapping and alignment units that utilize the conventional seed extension technique itself. Generation of Hash Table Index for Flexible Encapsulated Extension

[0046] FIG. 1 is a context diagram of a system 100 for generating a hash table index that facilitates flexible seed extension for hash table genomic mapping. System 100 includes a computer 110, a memory 112, and a memory 130. Although memories 112 and 130 are depicted in FIG. 1 as separate memory devices, the present disclosure need not be so limited. Instead, in some implementations, memories 112 and 130 can be the same memory device. For example, memories 112 and 130 can simply refer to two separate storage locations on a single memory device. Alternatively, memories 112 and 130 can each be stored in a separate memory device, such as a separate hard disk accessible by computer 110. As another example, memory 112 can be a memory device of a cloud-based server that stores a library of reference genomes, and memory 130 can be the local memory of computer 110. Thus, the depiction of memories 112 and 130 as separate memories in FIG. 1 is not limited to memories 112, 130 themselves or their contents, and these memories need not be organized or stored in any particular implementation of the present disclosure.

[0047] The computer 110 can include a computer or multiple computers each including one or more processing units configured to execute operations by executing one or more software instructions. The one or more processing units can include one or more central processing units (CPUs), one or more graphical processing units (GPUs), or any combination thereof. The computer 110 can be configured to directly interact with the memory 112, the memory 130, or the programmable circuit 162 via one or more buses, one or more USB cables, one or more USB-C cables, etc., or any combination thereof. Alternatively, or in addition, the computer 110 can be configured to interact with the memory 112, the memory 130, or the programmable circuit 162 via one or more networks. The one or more networks can include a wired Ethernet network, a wireless network, an optical network, a LAN, a WAN, a cellular network, the Internet, or any combination thereof.

[0048] As an example, one implementation form can include a computer 110 configured to (i) interact with a memory 112 and a memory 130 stored in one or more local memory devices accessible to the computer 110 to generate a seed extension table 132 and a hash table 140, and (ii) use one or more networks to transmit the generated seed extension table 132 and hash table 140 to a programmable circuit 162 integrated with another device 160. The other device 160 can include a nucleic acid sequencer, a cloud-based server(s), or any other computer. In some implementations, the programmable circuit 162 can be integrated with other devices using an expansion card such as a PCI card. In such an implementation, the programmable circuit 162 can be housed on the logic board of a PCI card inserted into the motherboard of a sequencer, a cloud-based server, or another computer using a PCI port on the motherboard.

[0049] The programmable circuit 162 can include one or more programmable integrated circuits such as one or more Field Programmable Gate Arrays (FPGAs). A Field Programmable Gate Array is an integrated circuit including a plurality of hardware digital logic gates, hardware digital logic circuits, etc., which can be dynamically configured to implement one or more processing modules, such as a genome analysis module of a genome analysis pipeline like a mapping and alignment unit 170, or a part of a processing module like a hash table 140. The FPGA can be programmed using a hardware description language (HDL) such as Very High Speed Integrated Circuit Hardware Description Language (VHDL), Verilog. The FPGA is flexible in that a pre-programmed FPGA including one or more genome analysis modules of a genome analysis pipeline or a part of this genome analysis module can be dynamically reconfigured to include updates for one or more genome analysis modules, other different genome analysis modules, etc.

[0050] To implement the functionality of the programmable circuit 162 described herein, other types of integrated circuits can be used instead of, or in addition to, the programmable circuit 162. For example, one or more application specific integrated circuits (ASICs) can be used to implement the functionality of the programmable circuit 162, or a portion of the functionality. An ASIC is a custom integrated circuit that includes a plurality of hardware digital logic gates, a plurality of digital logic circuits, etc., configured at the time of manufacture. An ASIC is similar to the FPGAs described herein in that the hardware digital logic gates, or the plurality of digital logic circuits of the ASIC, can be described or designed using a hardware description language such as VHDL, Verilog, etc. The ASIC can then be manufactured or printed to include the digital logic circuits or digital circuits described by the HDL. However, once manufactured or printed, the ASIC cannot be reconfigured dynamically like an FPGA. The embodiments described herein describe programmable circuits or custom circuits, but the disclosure need not be so limited. In some implementations, for example, other types of integrated circuits can be used to implement the functionality described as being performed by the programmable circuit 162.

[0051] The memory 112 can store one or more reference sequences 114. The reference sequences can include (i) a full reference genome representative of a species, (ii) a portion of the reference genome representative of a species, or (iii) complete and / or partial reference genomes representative of multiple species. The reference sequences include a continuous list of bases or nucleotides. The continuous list of bases or nucleotides that make up the reference sequences can be organized in the memory 112 as a digital nucleic acid sequence database. As representative of a species, a particular reference sequence can be assembled from a plurality of different donors of a particular species by a human, a computer, or both.

[0052] In some embodiments, a particular reference array can be assembled as representative of a particular population, where the particular population is a subset of a species having a particular nucleic acid sequence that can uniquely distinguish the particular population from other populations within the species. The species can include any species including, for example, humans, non-human mammals, worms, fish, insects, plants, bacteria, viruses, etc. The reference array can be generated from samples of non-extinct species such as humans, or from currently extinct species such as populations of dinosaurs or mammoths. The reference array of an extinct species such as a dinosaur can be assembled using samples obtained from biological materials contained within the remains of the extinct species preserved by fossilization, freezing, or other methods. The reference array of an extinct species can be assembled from a combination of (i) sequencing of biological remnants obtained from fossilized remains of the extinct species and (ii) sequencing of biological samples from non-extinct species. The entire reference genome can include many contiguous bases or nucleotides. For example, the human reference genome can include up to three hundred million contiguous bases or nucleotides.

[0053] Computer 110 is configured to generate a hash table 140 that facilitates flexible seed elongation. Computer 110 begins generating hash table 140 by accessing reference array 114 stored in memory 112 and obtaining seeds 114-1, 114-2, 114-3~114-n of the reference array, where n is any non-zero integer greater than 0. In some implementations, computer 110 can use a seed access window to identify and obtain seeds 114-1, 114-2, 114-3~114-n of reference array 114. Computer 110 can initialize the seed access window to a seed length K, where K is the number of bases or nucleotides included in each seed, and K is any non-zero integer greater than 0. At the start of the reference array, computer 110 can position a seed access window of length K and begin accessing the seeds of reference array 114 to include a first set of K nucleotides in a reference array seed such as seed "GTTTA" 114-1. In this example, K is equal to 5, but K is not limited to such nucleotide lengths. Instead, K can be made equal to any non-zero integer greater than 0, and in some implementations, for example, it can be equal to 7, 10, 12, 15, 18, 20, 21, 25, or more bases or nucleotides. Seeds 114-1, 114-2, and 114-3~114-n are merely examples of seeds of reference array 114, and in this example, it is not necessary to correspond to a set of four array seeds of reference array 114.

[0054] To generate the hash table 140, the computer 110 is configured to access each of the seeds 114-1, 114-2, and 114-3~114-n of the reference array 114 and execute a set of operations on each of the seeds 114-1, 114-2, and 114-3~114-n. The set of operations is designed to generate information for storing at the hash position 144 of the hash table 140 corresponding to the index key 142 of the hash table 140. Each index key 142 can correspond to each of the plurality of seeds 114-1, 114-2, 114-3~114-n of the reference array 114, the reverse complement of each of the seeds 114-1, 114-2, 114-3~114-n, one or more extended seeds of the plurality of seeds 114-1, 114-2, 114-3~114-n, or the reverse complement of each of the extended seeds. Each of the index seeds 142 can be mapped to the hash position 144 using the hash function 143.

[0055] After each respective seed has been accessed and used to execute the set of operations, the computer 110 can identify each of the plurality of seeds 114-1, 114-2, and 114-3~114-n and access each such seed by advancing the K positions of the seed access window within the reference array 114. The set of operations executed for each of the seeds 114-1, 114-2, and 114-3~114-n is described in more detail below. The set of operations can include the generated information to form the population of the hash table 140. Alternatively, the population of the hash table 140 can occur after the set of operations has been completed for each seed.

[0056] The set of operations of the computer 110 begins by the computer 110 obtaining the seed identified by the seed access window on each of the seeds 114-1, 114-2, and 114-3~114-n of the reference array 114. In the example of FIG. 1, assume that the seed of the reference array 114 identified by the seed access window is "GTTTA" 114-1.

[0057] The computer 110 can determine whether the obtained seed "GTTTA" 114-1 matches more than a predetermined number of positions of the reference sequences 114. The matching reference sequence positions can include a subset of the reference sequences 114 that includes the seed 114-1. The subset of the reference sequences 114 can include a set of consecutively ordered nucleotides that are K or more nucleotides in the obtained seed. In some implementations, the predetermined number of matching reference sequence positions can include one matching reference sequence position. However, in other implementations, the predetermined number can be set to two or more matching reference sequence positions.

[0058] If the seed 114-1 is less than or equal to the predetermined number of reference sequence positions, or if the computer 110 determines a match, the computer can fill the hash position 144 reached by the seed "GTTTA" 114-1 with the reference position(s) that match the seed 114-1. If the seed 114-1 matches the hash key 142 that the hash function 143 maps to the hash position, a seed such as the seed 114-1 can "reach" the hash position 144. Alternatively, if the computer 110 determines that the predetermined number of matching reference sequence positions is more than the predetermined number of reference sequence positions, the computer 110 can generate a seed extension tree for the seed 114-1. In the example of FIG. 1, the computer 110 determines that the seed "GTTTA" 114-1 matches more than the predetermined number of reference sequence positions. Accordingly, the computer 110 generates a seed extension tree 120 for the seed 114-1.

[0059] Computer 110 can generate a seed extension tree 120 for seed 114-1 for each node starting from root node 120. The seed extension tree 120 can be generated such that the set of matching reference positions identified by the leaf nodes does not exceed a predetermined match limit when further seed extension is not possible. Each node of the seed extension tree 120 can include a seed and an interval of consecutive addresses within the seed extension table 132. In some implementations, the seed extension table 132 includes a centrally lexicographically sorted list of reference array positions 131-1 to 131-6 that match a seed such as seed 114-1 obtained by computer 110 using a seed access window. The central lexicographical sort can include, for example, establishing a priority order of symbol locations and then alternately swapping to the left and right outside from the central symbol. Alternatively, the central lexicographical sort can include, for example, establishing a priority order of symbol locations and then alternately swapping to the right and left outside from the central symbol. Additionally, other variations can also be used.

[0060] In the embodiment of FIG. 1, the seed extension table 132 is centrally lexicographically sorted at 133 based on the seed 114-1 of "GTTTA". This embodiment assumes the normal alphabetical nucleotide order (i.e., A, C, G, T) with the left being the first to achieve the central lexicographical sort order shown in FIG. 1. Computer 110 can generate a seed extension table such as seed extension table 132 for each seed 114-1, 114-2, 114-3, 114-n that is determined to have more matching reference array positions than a predetermined threshold number. In some implementations, the seed extension table 132 for each eligible seed can be generated for a particular seed after computer 110 accesses the particular seed using a seed access window and before the seed extension tree 120 for the seed is generated.

[0061] The description of the nodes of the above seed expansion tree 120 indicates that the intervals between the addresses of each node are continuous. However, the present disclosure need not be so limited. Instead, the intervals between the addresses of the nodes may not be continuous. For example, a particular implementation may use intervals for describing multiple different sets of one or more consecutive positions of the seed expansion table, or multiple different sets of other data structures stored in one or more memory devices, and each consecutive set of one or more consecutive positions may be non - consecutive with respect to each other. That is, it is possible that there is a break in continuity between each set.

[0062] The seed expansion table for each eligible seed can be stored in the memory 130. Thereby, n seed expansion tables, that is, one for each of the n seeds of the reference array 114 can be provided. Alternatively, the number of seed expansion tables may be less than n, such as when the seed expansion table is generated and stored only for seeds having a matching reference array position greater than a predetermined threshold number. After the generation of each of the seed expansion tables, each set 132A of the seed expansion tables can be provided to a device 160 housing the programmable circuit 162 and stored in a memory 180 accessible by the programmable circuit 162. The memory 180 can include a DRAM memory, an SRAM memory, a NAND memory, etc. In some implementations, the set 132A of the seed expansion tables can be provided to the device 160 housing the programmable circuit 162 as individual seed expansion tables. In other implementations, the set 132A of the seed expansion tables can be provided as a single master seed expansion table composed of the concatenation of each of the respective seed expansion tables of each seed. The set 132A of the seed expansion tables can be provided in any number of formats. In some implementations, the set 132A of the seed expansion tables can be compressed by the computer 110 to reduce the size of the seed expansion table file provided to the device 160 and then decompressed by the device 160, the programmable circuit 162, etc. for storage in the memory 180.

[0063] The computer 110 can generate a root node 121 of a seed extension tree 120 for including the seed "GTTTA" 121a and the interval A 121b. The interval A 121b identifies a consecutive interval of positions in a seed extension table 132 that stores reference array positions matching the seed "GTTTA" 121a represented by the root node 121. In this example, the interval A identifies positions in the seed extension table 132 that span from 131-1 to 132-6 and include "TAGTTTACT", "TAGTTTATC", "GAGTTTATG", "ACGTTTAGT", "TCGTTTAGT", and "ACGTTTAGC". The computer 110 can determine an appropriate interval or intervals for a particular seed of a node, such as node 121, by accessing the seed extension table 132 and determining the positions in the seed extension table 132 that have reference array positions matching the seed of node 121.

[0064] In some implementations, the interval 121b for a particular seed of a node, such as node 121, can be described using the start position address of the interval within the seed extension table 132 and the end position address of the interval within the seed extension table 132. In other implementations, the interval 121b for a particular seed of a node, such as node 121, can be described using the start position address of the interval within the seed extension table 132 and an offset from the start position address. In such implementations, the interval can be calculated later using the start and end addresses of the interval, or the start address and offset of the interval. However, the present disclosure need not be so limited. Instead, it is understood that the interval record can be represented within the hash table position 144 using any form of information, structured or unstructured, in any suitable manner. For example, in some implementations, the interval record can be implemented using one record of a fixed size and format. In other implementations, the interval record can be implemented by selecting from multiple formats of different sizes, including a record count, etc., to optimize the memory space consumed by the hash table 140, enable compression of the hash table 140, improve the efficiency of hash queries for other interval record formats, and the like.

[0065] Computer 110 can continue to generate the seed extension tree 120 by extending the number of bases or nucleotides for the seed 121a identified by the root node. For example, computer 110 can extend the seed length of the root node from 5 bases or nucleotides to 7 bases or nucleotides and identify the largest subset of reference sequence positions within the seed extension table having 7 matching bases or nucleotides. In the embodiment of FIG. 1, computer 110 can determine that the largest subset of reference sequence positions having 7 matching nucleotides is "CGTTTAG". Interval B identifies a consecutive interval of positions within the seed extension table 132 that stores the reference sequence positions matching the seed "CGTTTAG". In this embodiment, interval B spans positions 132-4 to 132-6 of the seed extension table 132 and identifies positions of the seed extension table 132 that include "ACGTTTAGT", "TCGTTTAGT", and "ACGTTTAGC". Computer 110 can generate node 122 using the information determined using the seed extension table 132. For example, computer 110 can generate node 122 that includes the seed "CGTTTAG" 122a and interval B 122b.

[0066] Referring to the embodiment of FIG. 1, the computer 110 can continue to generate the seed extension tree 120 by determining whether there are other reference sequence positions in the seed extension table that have seven matching bases or nucleotides. If there are other reference sequence positions in the seed extension table 132 that have seven matching bases or nucleotides, the computer 110 uses the next largest set of reference sequence positions that have seven matching bases or nucleotides to generate the next node of the seed extension tree. In the embodiment of FIG. 1, the computer 110 can determine that the next largest subset of reference sequence positions has seven matching nucleotides. The interval E identifies the consecutive intervals of positions within the seed extension table 132 that store the reference sequence positions that match the seed "AGTTTAT". In this embodiment, the interval E spans from 132-2 to 132-3 and identifies the positions in the seed extension table 132 that include "TAGTTTATC" and "GAGTTTATG". The computer 110 can generate the node 123 using the information determined using the seed extension table 132. For example, the computer 110 can generate a node 123 that includes the seed "AGTTTAT" 123a and the interval E 123b.

[0067] Referring to the embodiment of FIG. 1, the computer 110 can continue to generate the seed extension tree 120 by determining whether there are other reference sequence positions in the seed extension table that have seven matching bases or nucleotides. If other reference sequence positions in the seed extension table are identified as having seven matching bases or nucleotides, the computer 110 can generate a new node of the seed extension table 120 using the next largest set of reference sequence positions having seven matching bases or nucleotides, as described above. However, in the embodiment of FIG. 1, there are no other reference sequence positions in the seed extension table 132 that have seven matching bases or nucleotides. Therefore, the computer 110 determines to extend the number of bases of the nucleotide from seven to nine and can continue to analyze the reference sequence positions in the seed extension table 132.

[0068] Referring to the embodiment of FIG. 1, the computer 110 can identify the largest subset of reference sequence positions having nine matching nucleotides. In this embodiment, there are multiple subsets of reference sequence positions having nine matching nucleotides. In such an example, the computer 110 can determine to create a node of the seed extension tree for each set of reference sequence positions having nine matching reference sequence nucleotides. In some implementations, the computer 110 can randomly determine the order of creating the seed extension tree nodes. In other implementations, the computer 110 can begin to generate subsequent extension tree nodes based on their lexicographical order.

[0069] Regardless of their order of creation, computer 110 can continue by generating nodes of the seed extension table for each subset of reference sequence positions having nine matching nucleotides. By way of example, computer 110 can generate a node 124 of the seed extension tree 120 such that it includes an extended short seed "TCGTTTAGT" 124a and a spacer C 124b. The spacer C 124b identifies consecutive intervals of positions within the seed extension table 132 that store reference sequence positions matching the short seed "TCGTTTAGT" 124a. In this example, the spacer C identifies positions in the seed extension table 132 spanning 132 - 5 and including "TCGTTTAGT". Computer 110 can access the seed extension table 132 and determine the appropriate spacer for a particular short seed of a node such as node 124 by determining the position in the seed extension table 132 having a reference sequence position matching the short seed of node 124.

[0070] Referring to the embodiment of FIG. 1, computer 110 can continue by generating nodes of the seed extension table for each subset of reference sequence positions having nine matching nucleotides. By way of example, computer 110 can generate a node 125 of the seed extension tree 120 such that it includes an extended short seed "ACGTTTAGC" 125a and a spacer D 125b. The spacer D 125b identifies consecutive intervals of positions within the seed extension table 132 that store reference sequence positions matching the short seed "ACGTTTAGC" 125a. In this example, the spacer D identifies positions in the seed extension table 132 spanning 132 - 6 and including "ACGTTTAGC". Computer 110 can access the seed extension table 132 and determine the appropriate spacer for a particular short seed of a node such as node 125 by determining the position in the seed extension table 132 having a reference sequence position matching the short seed of node 125.

[0071] Referring to the embodiment of FIG. 1, computer 110 can continue by generating nodes of the seed extension table for each subset of reference array positions having nine matching nucleotides. By way of example, computer 110 can generate a node 126 of the seed extension tree 120 such that it includes the extended short seed "TAGTTTATC" 126a and the interval F 126b. The interval F 126b identifies consecutive intervals of positions within the seed extension table 132 that store reference array positions matching the short seed "TAGTTTATC" 126a. In this embodiment, the interval F identifies positions in the seed extension table 132 that span 132-2 and include "TAGTTTATC". Computer 110 can access the seed extension table 132 and determine the appropriate interval for a particular short seed of a node, such as node 126, by determining the position in the seed extension table 132 that has a reference array position matching the short seed of node 126.

[0072] The present disclosure describes embodiments that construct a seed extension table in a particular regular order that proceeds from the largest set of matching bases to the smallest set of matching bases. However, the present disclosure need not be limited to the use of a seed extension tree constructed in this manner. Instead, any process for constructing the seed extension table can be used as long as the result of the seed extension table structure process creates the seed extension table. For example, the seed extension tree can be generated from the smallest set of matching bases to the largest set of matching bases, or without any particular order at all. In some implementations, a previously generated seed extension table can be generated and used by system 100 without the seed extension table having to be constructed by system 100.

[0073] Computer 110 can use the generated seed extension tree 120 to fill the hash position 144 of the hash table 140 where a specific seed input corresponding to a specific hash index key 142 reaches. As an example, computer 110 can determine whether node 121 is a leaf node. Based on the determination that node 121 is not a leaf node, computer 110 can use root node 121 to fill hash position 144-y, where y is any non-zero integer. Filling hash position 144-y using root node 121 can include storing interval record 153b at hash table position 144-y where seed 121a reaches. Interval record 153b identifies interval 121b for node 121. Hash table 140 can include hash table index keys 142 for each seed 114-1, 114-2, 114-3~114-n, the reverse complement of each seed 114-1, 114-2, 114-3~114-n, or combinations thereof. Each hash table index key 142 can be mapped to one or more hash positions 144 using hash function 143. Each hash position 144 can be implemented using one or more storage buckets, where the storage buckets correspond to a set of one or more storage positions of a memory device. Each of the one or more storage positions of the memory device can be memory positions that are consecutive or non-consecutive.

[0074] The embodiment of FIG. 1 shows only a portion of the hash table 140 having keys 142 corresponding to the forward seeds of seeds 121, 122, 123, and 125. However, the present disclosure need not be so limited. For example, in some implementations, the seeds can be hashed using the hash function 143 in such a way that the reverse complement nucleotide sequence of any seed yields the same hash as the original forward seed. The reverse complement of a nucleotide sequence can be determined by reversing the order of the original nucleotide sequence and exchanging As for Ts, Ts for As, Cs for Gs, and Gs for Cs. As an example, the hash key 142 for the original forward seed GTTTA 121a can have the same hash as the hash key 142 for the reverse complement of the seed GTTTA, which is TAAAC. In such implementations, when the matching reference sequence positions are stored at the hash position 144 or as entries in the seed extension table 132, their sequence orientations can be annotated, for example, using a reverse-complement (RC) flag. However, in other implementations, the reverse complement of a seed may yield a different hash, and it may not be necessary to annotate the orientation at the matching reference sequence positions stored at the hash position 144 of the hash table 140 or the seed extension table 132.

[0075] Filling the position 144 of the hash table 140 can also include determining whether the extension record is to be filled at the hash position 144. Determining whether the extension record should be filled at the hash position 144 can include determining whether the node of the seed extension tree 120 used to fill the hash position is a leaf node. If the node is determined to be a leaf node, the computer 110 can determine not to store the extension record at the hash position reached by the seed associated with the node. Instead, if the node is determined not to be a leaf node, the computer 110 can generate an extension record and store the generated extension record at the hash table position 144. Referring to the embodiment of FIG. 1, the computer 110 can determine or has previously determined that the node 121 is not a leaf node. In such an example, the computer 110 can generate and store the extension record 153a at the hash table position 144-y reached by the seed 121a. Thus, the hash position 144-y can include the extension record 153a and the interval record 153b.

[0076] When an extension record is executed by a computer such as a central processing unit (CPU), a graphics processing unit (GPU), or a programmable circuit 162 that executes software instructions, the CPU, GPU, or programmable circuit 162 can reach a hash position that stores the extension record by one or more nucleotides and extend the seed used in the hash query. In some implementations, the extension record can be generated such that the computer is instructed to extend the seed symmetrically on each end of the seed. Thus, by way of example, the extension record can be generated such that the computer, such as a CPU, GPU, or programmable circuit 162, is instructed to extend the seed by two nucleotides, four nucleotides, six nucleotides, etc. In such implementations, the symmetric extension of the seed can be achieved by extending the seed by one nucleotide on each end of the seed, two nucleotides on each end of the seed, three nucleotides on each end of the seed, etc. In the embodiment of FIG. 1, the extension record 153a is configured to symmetrically extend the initial seed 121a by two bases. The computer 110 can determine the extension length to include in the extension record based on various factors, including (i) the number of reference array positions that match the seed, (ii) the number of desired runtime seed extension iterations, (iii) the number of matching reference array positions required for each iteration, etc. The runtime flexible seed extension using the hash table 140 is described in more detail below in connection with FIG. 3.

[0077] Nucleotide seeds have generally been described as consisting of a contiguous set of contiguous nucleotides. Similarly, extension records have been described as sequentially extending one or more additional nucleotides of a contiguous set of contiguous nucleotides in a manner that can be contiguous, whether symmetric or asymmetric. However, the present disclosure is not limited to the use of contiguous sets of contiguous nucleotides. Instead, the seed of a read or reference sequence can be a non-contiguous seed pattern from the read or reference sequence. Similarly, an extension record can include instructions to extend an initial seed to incorporate non-contiguous neighboring bases or nucleotides into a CPU, GPU, or programmable circuit 162 when processed by the CPU, GPU, or programmable circuit 162. In such an implementation, the matching reference sequence positions for each root node seed can be sorted lexicographically within the seed extension table 132 in a manner commensurate with the use of non-contiguous seeds.

[0078] Computer 110 can continue to fill the information at hash position 144 for each of the remaining nodes 122, 123, 124, 125, 126 of the seed extension tree 120. As an example, computer 110 can determine whether node 122 is a leaf node. Based on the determination that node 122 is not a leaf node, computer 110 can use node 122 to fill hash position 144-3. Filling hash position 144-3 using node 122 can include storing interval record 152b at the hash table position 144-3 reached by seed 122a. Interval record 152b identifies interval 122b for node 122. Computer 110 determines or has previously determined that node 122 is not a leaf node and generates an extension record 152a for storage at hash position 144-3. In the embodiment of FIG. 1, extension record 152a includes instructions to symmetrically extend seed 122a by two bases or nucleotides. These instructions of extension record 152a can be executed at runtime, for example, if interval B is not accessed in response to a query for seed 122a.

[0079] However, the present disclosure is not so limited, and other extended record scans can be generated that instruct the CPU, GPU, or programmable circuit 162 to extend the seed by different additional nucleotide lengths (e.g., 2, 4, 6, 8, etc.) or in different ways (e.g., asymmetrically using additional nucleotide lengths of 1, 3, 5, etc.). The example of FIG. 1 shows a single extended record within hash position 144-3, but the present disclosure is not so limited. Instead, in some implementations, multiple extended records can be stored at a single hash position 144-3. For example, computer 110 can also store one or more additional extended records at hash position 144-3 configured to extend the initial seed 122a by 4 bases. In such an implementation, the CPU, GPU, or programmable circuit 162 can attempt to first extend the initial seed 122a by 4 bases at runtime. If such a seed extension fails, subsequent queries of the hash table 140 will not create a matching reference array position at runtime, so the CPU, GPU, or programmable circuit can obtain another extended record 152a that includes an instruction to extend the initial base by only 2 bases. This can increase the likelihood that a matching reference array position will be returned.

[0080] Computer 110 can continue to fill the hash positions 144 for each node 123, 124, 125, 126 of the seed extension tree 120. As an example, the computer can determine whether node 123 is a leaf node. Based on the determination that node 123 is not a leaf node, computer 110 can use node 123 to fill hash position 144-1. Filling hash position 144-1 can include storing the interval record 150b at the hash table position 144-1 reached by seed 123a. The interval record 150b identifies the interval 123b for node 123. Computer 110 determines, or has previously determined, that node 123 is not a leaf node and generates an extension record 150a for storage at hash position 144-1. In this embodiment, the extension record 150a includes instructions to symmetrically extend the seed 123a by two bases or nucleotides. These instructions of the extension record 150a can be executed at runtime, for example, if interval E is not accessed in response to a query for seed 123a.

[0081] Computer 110 can continue to fill the hash positions 144 for each node 124, 125, 126 of the seed extension tree 120 with information. As an example, computer 110 can determine whether node 125 is a leaf node. Based on the determination that node 125 is a leaf node, computer 110 can determine to fill the hash table 140 by storing the matching reference array position 155 identified by the interval D 125b that matches the seed "ACGTTTAGC" at the hash position 144-2. Alternatively, in other implementations, computer 110 can determine to store the interval record at the hash position 144-2 that identifies the interval D 125b. Such a determination can be made by computer 110 in some implementations based on whether the storage of each matching reference array position at the hash table position 144 for the leaf node is an optimal use of memory resources. Thus, if it is determined that the storage of the matching reference array position at the hash table position 144 for the leaf node does not meet a predetermined threshold usage of memory resources, computer 110 can store the matching reference array position at the hash position that the seed of the leaf node of the seed extension tree reaches. Otherwise, if this memory resource usage threshold is exceeded, computer 110 can store the interval record at the hash position 144 that the seed of the leaf node of the seed extension tree reaches. Computer 110 has determined or determines that node 125 is a leaf node and does not generate an extension record for storage at hash position 144-2. Thus, in this example, there is no further extension of the seed "ACGTTTAGC" as would occur at runtime.

[0082] As described above, the hash position 144-2 can match the seed 125a and store only the matching reference array positions corresponding to the hash key 142-1. This is because in this embodiment, the seed 125a is the seed of the leaf node 125 that cannot be extended. However, the population of reference array positions without one or both of the extension record or the interval record is not limited to the hash position 144 reached by the seed of the leaf node. Instead, the computer 110 can determine to fill the hash position 144 having the matching reference array positions without one or both of the extension record or the interval record in other instances. For example, in some implementations, when the computer 110 determines that the seed extension table 132 for a particular seed identifies only intervals of matching reference array positions that are less than the threshold number of matching reference array positions, the computer 110 can fill the hash position 144 reached by the particular seed having the matching reference array positions without one or both of the extension record or the interval record.

[0083] Other types of information can be stored at the hash position 144 of the hash table 140. For example, the computer 110 can receive an instruction to insert one or more "stop" records at the hash position 144 of the hash table 140. Such a "stop" hash record can result in a particular hash position 140 that stores either (i) a gap record or (ii) a set of one or more matching reference positions that have no further extension of the seed used to reach the hash position, in order to return either the gap record or the set of one or more matching reference positions. In other implementations, the computer 110 can receive an instruction to insert a "stop" record at a hash position that already contains an extended record. In such an implementation, when the CPU, GPU, or programmable circuit 162 encounters a "stop" record, the CPU, GPU, or programmable circuit 162 can conditionally determine whether to (i) discard the extended record and return either (i) a gap record or (ii) a set of one or more matching reference positions that match the seed used to reach the hash position having the "stop" record, or (ii) execute the seed extension described by the extended record. In some implementations, the conditional determination can be made based on one or more factors such as (i) a gap record or (ii) the number of matching reference arrays identified by a set of one or more matching reference array positions. Thus, the insertion of one or more "stop" records at specific hash positions in response to each input seed can be used as a design tool to avoid a fixed maximum false match problem without reconstructing a hash table such as the hash table 140.

[0084] The computer 110 can continue to iteratively fill the hash positions 144 for each of the remaining nodes of the seed extension tree 120, such as nodes 124, 126. Since nodes 124, 126 are leaf nodes like node 125, the entries for each of these leaf nodes can be filled using the method described above for node 125.

[0085] In addition, computer 110 can continue to apply the process described above iteratively in the embodiment of FIG. 1 with reference to seed "GTTTA" 114-1 to each seed of reference array 114. For example, when seed "GTTTA" 114-1 is processed as described above, computer 110 can advance the seed access window to the next subsequent seed within the reference array, access the seed, and then iteratively execute the process described above with reference to seed "GTTTA" 114-1 for each of the n seeds of the reference array. These processes can include obtaining the seed identified by the seed access window, determining whether the seed has more than a predetermined number of matching reference array positions, generating a seed extension tree if there are more than a predetermined number of matching reference array positions, and then filling hash table 144 using the seeds and intervals identified by the nodes of the seed extension tree. In some implementations, computer 110 can also iteratively execute the same process described above with reference to seed "GTTTA" 114-1 for the reverse complement of each of the n seeds of reference array 114. Culturing these iterative processes for each reference seed and each reverse complement can result in a hash table 140 having x index entries and y hash positions, where x and y are each on the order of 100 million or even 1 billion for a particular reference array such as the human genome.

[0086] In some implementations, one or more CPUs, GPUs, or combinations thereof are used to execute software instructions that, when executed, cause one or more CPUs, GPUs, or combinations thereof to execute the processes described with respect to FIGS. 3 and 4. The hash table 140 can be used by a computer such as computer 110 in software to execute a hash query against the hash table 140, thereby performing run-time flexible seed extension. In other implementations, computer 110 can generate a hash table installation package that includes software instructions for installing the hash table 140 and a set 132A of seed extension tables on another computer. For example, the hash table installation package can include software instructions that, when executed, perform the operations described by process 200 of FIG. 2. Computer 110 can provide the hash table installation package including the software instructions to another computer. The other computer can receive the hash table installation package and install the hash table 140 and the set 132A of seed extension tables. The other computer can then perform run-time flexible seed extension by executing, in software, a hash query against the hash table 140 using one or more CPUs, GPUs, or combinations thereof to execute software instructions that, when executed, cause one or more CPUs, GPUs, or combinations thereof to execute the processes described with respect to FIGS. 3 and 4.

[0087] However, in some implementations, computer 110 can generate a hash table install package 146 that includes hardware programming language instructions that can configure programmable circuit 162 to implement mapping and alignment unit 170 to a hardware digital logic circuit. The hardware programming language instructions can be in the form of a file, such as a binary bitstream file. The binary bitstream file can be generated prior to being included in hash table install package 146 by compiling hardware programming language code, such as VHDL, Verilog, that describes the circuit mechanism implemented by programmable circuit 162. The hardware programming language instructions of the hash table install package, when processed by programmable circuit 162, can program the dynamically configurable logic circuits of the programmable logic circuit to implement flexible seed stretching by performing a hash query on hash table 140 in hardware using the processes described with respect to FIGS. 3 and 4. Hash table install package 146 can also include a set of seed extension tables 132A and instructions for installing the set of seed extension tables 132A in memory 180 accessible by programmable circuit 162. Hash table install package 146 can also include hash table 140 and instructions for installing hash table 140 in memory 180 accessible by programmable circuit 162. Programmable circuit 160 can be programmed to use hash table 140 as part of mapping / alignment unit 170 to perform a mapping to a reference array of short seeds, as discussed in more detail herein with respect to FIG. 3.Computer 110 can provide a hash table installation package to device 160, which can be a desktop computer, a laptop computer, a tablet computer, a smartphone, a cloud-based server, a sequencer, or another device that houses programmable circuit 160 using one or more networks, one or more buses, direct connections such as USB cables, USB-C cables, or any combination thereof. Device 160 can receive the hash table installation package and program programmable circuit 162 to implement mapping and alignment unit 170 within the hardware logic gates of programmable circuit 162 using the hardware programming language instructions of the hash table installation package.

[0088] Thus, each hash table installation package 146 can be used to manage the installation, use, and even removal of the hash table 140 and the seed extension table in a variety of different ways. For example, in some implementations, the set 132A of the hash table 140 and the seed extension table can each be stored as a file on a hard disk or other storage medium, and then each can be loaded into common memory such as DRAM that includes one or more components or modules for implementing runtime-flexible seed extension as described herein with respect to the processes described in FIGS. 3 and 4 prior to runtime access. However, in other implementations, the hash table 140 or the set 132A of the seed extension table can be stored together or separately as one or more distinct contiguous portions, or non-contiguous portions, within the memory device. Similarly, the hash table 140 or the set 132A of the seed extension table can be compressed or uncompressed, stored on a common or separate storage medium and / or memory, or cached or uncached, as long as there is some means and method for the runtime mapping or otherwise the runtime mapping and alignment unit 170 to access selected portions of both the hash table 140 and the set 132A of the seed extension table. In yet other implementations, the hash table 140 can be fully implemented within the hardware logic circuit of the programmable circuit 162, and the set 132A of the seed extension table can be stored in the memory 180 accessible by the programmable logic circuit 162 such as a DRAM memory unit. In yet other implementations, the hash table 140 can be stored in the memory 180 accessible by the programmable logic circuit 162 such as a DRAM memory unit, and the set 132A of the seed extension table can be fully implemented within the hardware logic circuit of the programmable circuit 162.

[0089] In some implementations, computer 110 can also generate an installation package that includes a hash table and a seed extension builder as described herein. Computer 110 can provide the installation package to another computer via a network. Using the installation package, a party that receives and installs the hash table and the seed extension builder can construct its own hash table and seed extension table from its own selected reference array and with its own selected settings, so that the hash table and the seed extension builder can be installed on another computer or a different computer. Thus, the recipient of the hash table and seed extension builder installation package can construct its own hash table at any time from its own selected reference, store the hash table on disk, load the hash table into memory 180 accessible by programmable circuit 162, and use programmable circuit 162 to perform mapping and alignment.

[0090] FIG. 2 is a flowchart of a process 200 for generating a hash table index that facilitates flexible seed extension for hash table genome mapping. Generally, process 200 involves obtaining, by a computer system, a particular seed of nucleotides from a reference array, where the particular seed represents an array of nucleotides having a nucleotide length of K nucleotides (210), determining, by the computer system, that the particular seed matches more than a predetermined number of reference array positions (220), and generating, by the computer system, a seed extension tree having a plurality of nodes based on the determination that the particular seed matches more than a predetermined number of reference array positions, where each node of the plurality of nodes is (i) an extension of the particular seed and has an extended seed having a nucleotide length of K * of nucleotides, where K *(i) an extended seed that is one or more nucleotides larger than K, and (ii) a plurality of positions including data describing a reference sequence position in the seed extension table that matches the extended seed, corresponding to generating (230); for each node of the plurality of nodes, storing, by a computer system, interval information at a position in a hash table corresponding to an index key of the extended seed, the interval information including data describing a reference sequence position that matches the extended seed associated with the node, referring to a plurality of positions in the seed extension table, storing (240); By including, a hash table can be generated. The process 200 is described in more detail below as being executed by a computer system such as the computer 110.

[0091] More specifically, the computer system can start the execution of the process 200 by obtaining a specific seed of nucleotides from the reference sequence by the computer system, the specific seed representing an array of nucleotides having a nucleotide length of K nucleotides (210). In some implementations, obtaining a specific seed can include determining, by the computer system, the position of a seed access window within the reference sequence. The computer system can then obtain a subset of the reference sequence identified by the seed access window. The computer system can include one or more computers.

[0092] The computer system can continue the execution of process 200 (220) by determining, using the computer system, whether a particular seed matches more than a predetermined number of reference array positions. If the computer system determines that a particular seed does not match more than a predetermined number of reference array positions, the computer system can determine not to generate a seed extension tree for the particular seed. Instead, the computer system can obtain data describing each of the reference array positions that match the second seed. Then, the computer system can store the data describing the reference array positions that match the particular seed at a second position in the hash table corresponding to the index key of the particular seed.

[0093] Alternatively, if the computer system determines that a particular seed matches more than a predetermined number of reference array positions, the computer system can generate a seed extension tree having a plurality of nodes (230). Each node of the plurality of nodes can include data representing (i) an extension of the particular seed and an extended seed having a nucleotide length of K, where K is one or more nucleotides greater than K, and (ii) a plurality of positions in a seed extension table that include data describing the reference array positions that match the extended seed. In some implementations, the plurality of positions can include consecutive intervals in the extension table of the reference array positions that match the extended seed associated with the node. * of the extended seed, where K * is one or more nucleotides greater than K, and data describing the reference array positions that match the extended seed.

[0094] The computer system can continue the execution of process 200 by storing interval information at the hash positions of the hash table for each node of the seed extension tree. In some implementations, the computer system can generate a hash table (240) by storing interval information at the hash positions of the hash table corresponding to the index keys of the extended seeds for each node of the seed extension tree. The interval information can include references to a plurality of seed extension positions including data that describes the reference array positions that match the extended seeds associated with the nodes. In some implementations, the plurality of seed extension table positions described by the interval information can include consecutive intervals of positions within the seed extension table that include data that describes the reference array positions that match the extended seeds. Run-Time Flexible Seed Extension Using Hash-Table Genome Mapping

[0095] Figure 3 is a context diagram of a runtime system 300 for performing run-time flexible seed extension for hash-table genome mapping. The runtime system 300 includes a programmable logic circuit 162, a mapping and alignment unit 170, a hash table 140, a memory 18, and a plurality of seed extension tables such as a seed extension table 132 stored in the memory 180. The example of FIG. 3 describes a mapping and alignment unit 170 and a hash table 140 implemented in hardware using the hardware logic circuit of the programmable logic unit 162, but the present disclosure is not so limited. Instead, the mapping and alignment unit 170 may be a software application implemented using software instructions executed by one or more CPUs, GPUs, or combinations thereof that access the hash table 140 stored in the memory unit.

[0096] The execution of runtime-flexible seed extension for hash table genome mapping by system 300 can be initiated by the mapping and alignment unit 170 accessing the current read 305. The current read 305 can be generated by a nucleic acid sequencer that has performed a primary analysis of a biological sample. The primary analysis can include the nucleic acid sequencer receiving a biological sample, such as a blood sample, a tissue sample, or sputum, and generating output data, such as one or more reads 305 that represent the order of nucleotides in the nucleic acid sequence of the received biological sample. In some implementations, the biological sample can include a DNA sample, and the nucleic acid sequencer can include a DNA sequencer. In such implementations, the order of the sequenced nucleotides in the read 305 generated by the nucleic acid sequencer can include one or more of guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. In other implementations, the nucleic acid sequencer can include an RNA sequencer, and the biological sample can include an RNA sample. In such implementations, the order of the sequenced nucleotides in the read generated by the nucleic acid sequencer can include one or more of G, C, A, and uracil (U) in any combination. Thus, while the example of FIG. 3 describes the processing of reads consisting of G, C, A, and T generated by a DNA sequencer based on a DNA sample, the present disclosure is not so limited. Instead, other implementations can process reads consisting of C, G, A, and U generated by an RNA sequencer based on an RNA sample.

[0097] Generally, the mapping and alignment unit 170 can be configured to be agnostic to the type of reads that the mapping and alignment unit 170 receives, maps, and aligns. For example, in some implementations, the same binary code can be used to represent both “T” and “U”. Reads received by the mapping and alignment unit 170 can include DNA, cDNA, and / or RNA, and the reference can be DNA, cDNA, and / or RNA. In such implementations, read bases T and / or U can share a single binary code such that read T and / or U match reference T and / or U.

[0098] In some implementations, the nucleic acid sequencer can include a next generation sequencer (NGS) configured to generate sequence reads, such as read 305 for a given sample, in a manner that achieves high throughput, scalability, and speed through the use of massively parallel sequencing technology. NGS enables rapid whole-genome sequencing, zooming in on deeply sequenced target regions, discovering novel RNA variants and splice sites using RNA sequencing (RNA-Seq), or analyzing epigenetic factors such as gene expression analysis, genome-wide DNA methylation, and DNA-protein interactions, sequencing cancer samples to study rare somatic mutations and tumor subclones, and quantifying mRNA for the study of microbial diversity in humans or the environment.

[0099] Sequence reads, such as read 305 generated by a nucleic acid sequencer, can be accessed and processed by a secondary analysis unit, such as mapping and alignment unit 170. In some implementations, the secondary analysis unit, such as mapping and alignment unit 170, can be implemented in hardware, such as a digital logic circuit, using a programmable circuit 162, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC). In other implementations, the secondary analysis unit, such as mapping and alignment unit 170, can be implemented using one or more CPUs, GPUs, or a combination of both to implement the functionality of the mapping and alignment unit 170. The hash table 140 can be implemented in the hardware logic circuit of the programmable circuit 162 in some implementations, such as when the mapping and alignment unit 170 is implemented using the programmable circuit 162, but the present disclosure is not so limited. Instead, the hash table 140 can be stored in a memory device and accessed as needed by (i) a CPU, GPU, or a combination of both that executes software instructions that implement the functionality of the mapping and alignment unit 170, or (ii) the mapping and alignment unit 170 implemented in a hardware digital logic circuit.

[0100] In some implementations, the programmable circuit 162 can be integrated with a nucleic acid sequencer that generated the read 305. In such an implementation, for example, the programmable circuit 162 can be housed on an expansion card, such as a Peripheral Component Interconnect (PCI) expansion card, and installed in the nucleic acid sequencer. In other implementations, for example, each of the programmable circuits 162 can be part of a separate computer that is directly connected to the nucleic acid sequencer using, for example, an Ethernet cable, a USB cable, a USB-C cable, etc., different from the nucleic acid sequencer. In yet other implementations, for example, the programmable circuit 162 can be integrated with a cloud-based server that is remotely accessible by a nucleic acid sequencer that generated the read 305 using one or more wired or wireless networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, the Internet, or combinations thereof.

[0101] The mapping and alignment unit 170 can receive a first hash query 310 that includes an initial seed "GTTTA" 310a. In some implementations, the hash query can simply consist of the seed of a sample read, such as the current read 305, which is used as an input to the hash table 140. In other implementations, additional data, metadata, etc. can be added to the seed of the sample read to convert the sample into a format that can be used to search the hash table 140.

[0102] In the example of FIG. 3, the initial seed "GTTTA" 310a included in the hash query 310 is the current read identified using the seed access window 305a [Table 1] It is obtained from the first part of 305. The mapping and alignment unit 170 can execute the hash query 310 using the hash table 140 to map the short initial seed 310a to the hash position 144 using the hash function 143. In the embodiment of FIG. 3, the execution of the hash query 310 can determine that the seed "GTTTA" 310a matches the hash index key "GTTTA" 142-2 that is mapped to the hash position 144-y by the hash function 143.

[0103] The mapping and alignment unit 170 can generate a response 310b to the hash query 310 using the hash table 140. The response 310b can include the content of the hash position 144-y reached by the seed 310a of the hash query 310. The mapping and alignment unit 170 evaluates the response 310b to determine whether the content includes a set of matching reference array positions, an extension record, an interval record, or a combination thereof. If the response 310b includes only a set of matching reference array positions without an extension record or an interval record, the mapping and alignment unit 170 can store the set of matching reference array positions in the seed match set 352 together with the metadata associating the matching reference positions with the seed of the received query. Alternatively, if the mapping and alignment unit 170 determines that the response includes an interval record, an extension record, or both, the mapping and alignment unit 170 must determine at 320 whether to use the matching reference seed identified by the interval record or to proceed with the extension of the query seed.

[0104] In the embodiment of FIG. 3, the evaluation of response 310b indicates that (i) the response does not include a set of matching reference array positions, and (ii) response 310b includes extension record 153a and interval record 153b. Based on response 310b, mapping and alignment unit 170 can determine at 320 whether the matching reference position identified by interval record 153b is accessed. In some implementations, mapping and alignment unit 170 is configured such that a response such as 310b to hash query 310 does not access the matching reference array position identified by an interval record such as interval record 153b when the response includes extension record 153a.

[0105] However, in other implementations, the mapping and alignment unit 170 can be configured to evaluate the number of matching reference array positions identified by the spacing record 153b before extending the seed 310a using the extension record 153b. In such an implementation, if the number of matching reference array positions is below a predetermined threshold, the mapping and alignment unit 170 can output the matching reference array positions of interval A identified by the spacing record at 310d. Outputting the matching reference array positions can include the mapping and alignment unit 170 accessing the matching reference array positions stored in interval A of the seed extension table 132 in the memory 180 and storing the accessed matching reference array positions in the seed match set storage 352. When the accessed matching reference array positions are stored in the seed match set 352, the process described by FIG. 3 can end without further extension of the seed 310a. Then, the seed access window 305a can be adjusted forward by one or more nucleotides along the current read 305. When the seed access window 305a is adjusted, the process described in connection with FIG. 3 starts again and can continue iteratively until the entire current read 305 has been queried. On the other hand, in this alternative implementation, if it is determined that the number of matching reference array positions is not below a predetermined threshold, the seed 310a can be extended using the extension record 152a.

[0106] Returning to the example of FIG. 3, the mapping and alignment unit 170 does not apply the aforementioned threshold to the match identified by the spacing record 153b. Instead, the mapping and alignment unit 170 determines at 320 not to use the matching reference array positions identified by the spacing record 153b because the output 310b includes the extension record 153a. Thus, in this scenario, the mapping and alignment unit 170 determines to extend the seed 310a.

[0107] Before proceeding to execute subsequent queries based on the extended seed, the mapping and alignment unit can store information describing interval A 310c in the "best interval" storage 350. Since no other intervals are identified and evaluated at this point in the process, interval A can be considered the "best interval" for the matching reference array position within the seed extension table 132 for seed 310a. However, in subsequent iterations of the process described by FIG. 3, each subsequent interval identified can be heuristically evaluated to determine whether the interval is better than an existing interval stored in the best interval storage for the initial seed 310a or the extended seed of the initial seed 310a. By storing the information describing interval A 310c in the best interval storage 340, the matching reference array position of interval A regressed in the event extension of the initial seed 310a can cause mapping failures such as an unmapped mapping problem or a high-precision mapping problem. The information describing interval A 310c can include data describing the start and end positions of a contiguous list of reference array positions matching the initial seed. In some implementations, the information describing interval A 310c can also include data identifying the seed whose reference array positions identified by interval A match.

[0108] The execution of flexible seed extension by the mapping and alignment unit 170 can continue with the mapping and alignment unit 170 generating a first extended seed 312a that is an extension of the initial seed 310a using the extension record 153a. In the embodiment of FIG. 3, the extension record 153a can include one or more instructions that instruct the mapping and alignment unit 170 to symmetrically extend the initial seed 310a by two bases or nucleotides. In the embodiment of FIG. 3, symmetrically extending the initial seed "GTTTA" 310a by two bases or nucleotides results in the extended seed "CGTTTAG" 312a of the read 305. In some implementations, the additional nucleotides "C" and "G" used to extend the initial seed 310a can be obtained from the next nucleotides of the read 305 that are on both sides of the initial seed 310a identified by the seed access window 305a.

[0109] In other implementation forms, such as when the seed access window is at the beginning of the lead 305, additional seeds for facilitating this seed extension exist on each side of the seed access window, but the extension can cause an extension of the initial seed beyond the boundary 305 of the lead. In such an implementation form, the seed extension may fail, and the process of mapping the initial seed to the corresponding reference array position using the hash table 140 may end without adding any corresponding reference array positions to the seed match set 352 for the query cycle starting from the initial seed. However, in such an implementation form, the seed access window 305a can be adjusted by one or more nucleotides in the forward direction along the lead 305, and the next seed of the lead 305 identified by the adjusted seed access window can be obtained for use as the initial seed of the hash query for a new query cycle using the hash table 140. Execution of a new query cycle for the next seed, and then using each of the seeds until each of the seeds is processed, can update the best interval storage 350, store one or more sets of corresponding reference array positions in the seed match set storage 352, or both, and evaluate this to identify the optimal set of corresponding reference array positions of the lead 305 as described with reference to FIG. 5, thus solving the problem of unmapped leads that may exist in the conventional method even in the case of a failed seed extension.

[0110] When the seed access window 305a moves towards both ends of the lead 305, for the same reason, a similar seed extension failure may occur. The present disclosure similarly solves these seed extension failures by evaluating the best interval storage 250, the seed match set 353, or both, from the previous iteration before the hash query for the lead, as described with reference to FIG. 5.

[0111] Returning to the embodiment of FIG. 3, the mapping and alignment unit 170 can generate a subsequent hash query 312 that includes a first extended seed 312a. The mapping and alignment unit 170 can obtain the first extended seed 312a from the hash query 312 and use the hash table to map the first extended short seed 312a to the hash position 144 using the hash function 143. In some implementations, the generation of the hash query 312 using the first extended seed 312a can include providing the first extended seed 312a to the mapping and alignment unit 170 as an input for seed mapping using the hash table 140 without generating a query. In the embodiment of FIG. 3, the execution of the hash query 312 determines that the seed "CGTTTAG" 312a matches the hash index key "CGTTTAG" 142-x that is mapped to the hash position 144-3 by the hash function 143.

[0112] The mapping and alignment unit 170 can generate a response 312b to the hash query 312 using the hash table 140. The response 312b can include the content of the hash position 144-3 that the seed 312a of the hash query 312 reaches. The mapping and alignment unit 170 can evaluate the response 312b to the hash query 312 and determine that the response 312b (i) does not include a set of matching reference array positions and (ii) includes an extended record 152a and a gap record 152b. Based on the response 312b, the mapping and alignment unit 170 can determine at 330 whether the matching reference position identified by the gap record 152b is accessed. In some implementations, the mapping and alignment unit 170 is configured not to access the matching reference array position identified by a gap record such as the gap record 152b when a response such as 312b to the hash query 312 includes an extended record 152a.

[0113] However, in other implementations, the mapping and alignment unit 170 can be configured to evaluate the number of matching reference array positions identified by the spacing record 152b before extending the seed 312a using the extension record 152b. In such an implementation, if the number of matching reference array positions identified by the spacing record 152b is below a predetermined threshold, the mapping and alignment unit 170 can output the matching reference array positions of interval B identified by the spacing record 152b at 312d. Outputting the matching reference array positions can include the mapping and alignment unit 170 accessing the matching reference array positions stored in interval B of the seed extension table 132 in the memory 180 and storing the accessed matching reference array positions in the seed match set storage 352. When the accessed matching reference array positions are stored in the seed match set storage 352, the process described by FIG. 3 can end without further extension of the seed 312a. Subsequently, the seed access window 305a can be adjusted by one or more nucleotides along the current read 305. When the seed access window 305a is adjusted, the process described in connection with FIG. 3 can start again and continue iteratively until the entire current read 305 has been queried. On the other hand, in this alternative implementation, if it is determined that the number of matching reference array positions is not below the predetermined threshold, the seed 312a can be extended using the extension record 152a.

[0114] Returning to the example of FIG. 3, the mapping and alignment unit 170 does not apply the aforementioned threshold to the match identified by the spacing record 152b. Instead, the mapping and alignment unit 170 determines at 330 not to use the matching reference array positions identified by the spacing record 152b because the output 312b includes the extension record 152a. Accordingly, the mapping and alignment unit 170 determines to extend the seed 312a.

[0115] Before proceeding to execute subsequent queries based on the extended seed, the mapping and alignment unit can determine whether to store the information describing interval B 312c as the "best interval" in the best interval storage 350. Determining whether to store the information describing interval B 312c as the "best interval" heuristically determines whether interval B is a better interval than the interval currently stored in the best interval storage 352 for the previous iteration of the first extended seed, which is interval A in this example. In one implementation, the best interval among a plurality of intervals can be determined by evaluating the number of target hits returned for each interval. In such an implementation, the "best" interval can be selected according to the multi-part rule. As an example, the mapping and alignment unit 170 can assign a first priority to an interval that encloses at least a predetermined number of matching reference array positions, which can be referred to as a threshold such as intvl-target-hits(32) matches. However, if each interval has fewer matches than intvl-target-hits(32), the interval with the most matches is stored as the best interval. Further, the mapping and alignment unit 170 can assign a second priority to an interval associated with a longer extended seed because such an interval is preferred. Also, if the mapping and alignment unit 170 determines that at least one interval has at least intvl-target-hits(32) matches, the best interval is selected based on the interval associated with the longest extended seed among all intervals that satisfy at least intvl-target-hits(32) matches. The examples in this specification refer to the threshold intvl-target-hits(32) having 32 matches, but the present disclosure need not be so limited. Instead, the threshold intvl-target-hits() can be sent to any number of matching reference array positions to implement this multi-part heuristic rule.

[0116] In the embodiment of FIG. 3, the previously stored interval A as the best interval in the best interval storage 350 identifies six matching reference array positions 132-1 to 132-6, and the interval B identifies three matching reference array positions 132-4 to 132-6. Applying an exemplary intvl-target-hit(10) threshold of ten matches, the mapping and alignment unit 170 can apply a multi-part heuristic rule and determine that the interval meets the intvl-target-hit(10) threshold. Thus, in accordance with the multi-part heuristic rule, the mapping and alignment unit 170 can select interval A as the best interval because interval A has the most matches, i.e., six matches, between interval A and interval B. Based on the application of this exemplary multi-part heuristic rule, the information describing interval B 321c can be discarded and interval A remains stored as the best interval. However, under other embodiments that apply different heuristic rules that do not need to be multi-part heuristic rules, it is possible for interval B to be selected as the best interval and stored in the best interval storage 350 to replace interval A. Such results can ultimately be left to specific design configurations such as the setting of the intvl-target-hits() threshold, the design of one or more heuristic rules, etc.

[0117] In the embodiment of FIG. 3, the previously stored interval A in the best interval storage 350 is compared with the interval B included in the response 312b to the query 312 using the aforementioned heuristic rules. However, the present disclosure need not be so limited. For example, in some implementations, the response to a hash query may include a plurality of interval records stored at the hash position 144 where a particular seed of the hash query reached. In such an implementation, the mapping and alignment unit 170 can apply the aforementioned heuristic rules to determine which of the plurality of interval records should be accessed. Similarly, the mapping and alignment unit 170 can also use such heuristic rules to determine the best interval for storage in the best interval storage 450 from each of the interval records returned in the query response. As another example, the mapping and alignment unit 170 can also use such heuristic rules to determine the best interval for storage in the best interval storage 450 from among each interval record returned in the query response and another interval previously stored in the best interval storage 350 for a previous iteration of the seed used in a query that returns a plurality of intervals.

[0118] In some implementations, the system 300 can facilitate the storage of two or more best intervals in the best interval storage 350. For example, in some implementations, up to two best intervals may be tracked. In some implementations, up to N best intervals may be tracked. In such an implementation, when N>1 best intervals are stored, the criteria for determining which intervals are being held can involve an evaluation of the relationship between interval candidates, the associated extended seeds of the interval candidates, or both, such that the N best intervals are associated with extended seeds that do not overlap each other within a lead.

[0119] The execution of flexible seed extension by the mapping and alignment unit 170 can continue with the mapping and alignment unit 170 generating a second extended seed 314a that is an extension of the first extended seed 312a using the extension record 152a. In the embodiment of FIG. 3, the extension record 152a can include one or more instructions that instruct the mapping and alignment unit 170 to symmetrically extend the first extended seed 312a by two bases or nucleotides. In the embodiment of FIG. 3, symmetrically extending the first extended seed "CGTTTAG" 312a by two bases or nucleotides results in the second extended seed "ACGTTTAGC" 314a of the read 305. In some implementations, the additional nucleotides "A" and "C" used to extend the first extended seed 312a can be obtained from the next nucleotides of the read 305 that are on both sides of the first extended seed "CGTTTAG" 312a.

[0120] Returning to the embodiment of FIG. 3, the mapping and alignment unit 170 can generate a subsequent hash query 314 that includes the second extended seed 314a. The mapping and alignment unit 170 can obtain the second extended seed 314a from the hash query 314 and use the hash table to map the second extended short seed 314a to the hash position 144 using the hash function 143. In some implementations, the generation of the hash query 314 using the second extended seed 314a can include providing the second extended seed 314a to the mapping and alignment unit 170 as an input for seed mapping using the hash table 140 without generating a query. In the embodiment of FIG. 3, the execution of the hash query 314 determines that the seed "ACGTTTAGC" 314a matches the hash index key "ACGTTTAGC" 142-1 that is mapped to the hash position 144-2 by the hash function 143.

[0121] The mapping and alignment unit 170 can generate a response 314b to the hash query 314 using the hash table 140. The response 314b can include the content of the hash position 144-2 that the second extended seed 314a of the hash query 314 reaches. The mapping and alignment unit 170 evaluates the response 314b to the hash query 314 and determines that the response 314b (i) includes a set of matching reference array positions 155, (ii) does not include an extended record, and (iii) does not include an interval record. Based on the response 314b, the mapping and alignment unit 170 can determine that the matching reference array positions 155 should be stored in the seed match set storage 352.

[0122] Since the response 314b does not include an extended record, the runtime flexible seed extension process for the seed "GTTTA" 310a of the read 305 ends. The seed access window 305a can be continuously advanced along the read 305 by one or more nucleotides until each process described with respect to FIG. 3 is executed for each respective seed of the read 305. This process is also described with respect to the flowchart of FIG. 4. As described above, when the seed access window 305a extends towards the end of the read 305, an attempt to extend the seed input to the mapping and alignment unit 170 may fail, potentially causing an un-mapped read problem. However, the present disclosure can use one or more intervals stored in the best interval storage, one or more reads stored in the seed match set 352, or a combination of both to identify a set of matching reference array positions of the read 305, at least as described with respect to FIG. 5.

[0123] Figure 4 is a flowchart of a process 400 for performing runtime flexible seed extension for hash table genomic mapping. Process 400 is described below as being performed by a computer system of one or more computers. The one or more computers can include, for example, a mapping and alignment unit 170. For the purposes of the present disclosure, the one or more computers can include a CPU or GPU configured to obtain and execute software instructions to implement the particular programmed functionality described by the software instructions. Alternatively, or in addition to this, the one or more computers can include a programmable circuit configured such that a hardware digital logic circuit of the programmable circuit is configured to implement the particular programmed functionality in hardware.

[0124] The computer system can start the execution of process 400 by executing a query of hash table 405. The query can include a nucleotide seed. The nucleotide seed can include a subset of nucleotides obtained from a read. The read can include a set of nucleotides generated by a nucleic acid sequencer based on a biological sample input to the nucleic acid sequencer. The biological sample can include, for example, a blood sample, a tissue sample, sputum, and the like.

[0125] As an example, reads generated by a nucleic acid sequencer based on a biological sample can include a series of nucleotides such as "ACGTTTAGC". This example includes a 9-nucleotide read. However, the use of 9-nucleotide reads is used only as an example. Instead of being limited to 9 nucleotides, reads as described by the present disclosure can be of any nucleotide length including 5 bases or nucleotides, 10 bases or nucleotides, 12 bases or nucleotides, 15 bases or nucleotides, 18 bases or nucleotides, 21 bases or nucleotides, 25 bases or nucleotides, 35 bases or nucleotides, 50 bases or nucleotides, 100 bases of nucleotides, 150 bases or nucleotides, 1,000 bases or nucleotides, millions of bases or nucleotides, or even more bases or nucleotides, but are not limited thereto. The seed of the query can include a portion of the read such as "GTTTA". The seed obtained from the read for use in the first hash query during the first iteration of process 400 can be of any length K, where K is less than the number of bases or nucleotides in the read. In some implementations, K can be made substantially smaller than the read nucleotide length, such as 1 / 100 of the read length, 1 / 10 of the read length, 1 / 5 of the read length, etc.

[0126] A computer system can execute a query containing a seed by obtaining the seed and comparing the seed to a hash key of a hash table. The hash key can correspond to each reference array seed, the reverse complement of each reference array seed, each extended seed of the reference array, and the reverse complement of each extended seed of the reference array. The reference array can include, for example, a reference genome or a portion of a reference genome for a species such as a human or other animal. When a hash key matching the seed of the query is identified by the computer system, the computer system can use a hash function to map the hash key to one or more hash positions. In some aspects of the present disclosure, the one or more hash positions can store (i) an extended record, (ii) a gap record, or (iii) one or more reference array positions. The computer system can generate a response to the query that includes the content of the one or more hash positions reached by the seed of the query.

[0127] The computer system can continue execution of process 400 by obtaining (410) a response to the executed query that includes information stored by one or more positions of the hash table determined to be reached by the query. The query is determined to reach one or more positions of the hash table when the seed of the query is determined to match a hash key mapped to the one or more positions using a hash function.

[0128] The computer system can continue the execution of process 400 by determining (415) whether the response to the executed query includes (i) an extended record (ii) a gap record, or (iii) one or more matching reference array positions. Determining (415) by the computer system whether the response to the executed query includes (i) an extended record (ii) a gap record, or (iii) one or more matching reference array positions can include parsing the received response and analyzing the parsed response data. The computer system can determine, based on the parsed data, whether the parsed data represents (i) an extended record, (ii) a gap record, or (iii) one or more matching reference array positions. In other implementations, the response to the executed query can include one or more data flags indicating whether the response includes (i) an extended record, (ii) a gap record, or (iii) one or more matching reference array positions.

[0129] In some examples, the computer system can continue the execution of process 400 by determining at step 415 that the response does not include an extended record, a gap record, or one or more matching reference array positions. If the computer system determines that the response does not include (i) an extended record, (ii) a gap record, or (iii) one or more matching reference array positions, the process ends at step 420 without adding any matching reference array positions to the seed match set for the seed of the query. As an example, if the seed is an extended seed and there is a seed extension error, the received response to the query containing the seed may not include (i) an extended record, (ii) a gap record, or (iii) one or more matching reference array positions. Such a seed extension error can occur, for example, when the computer system attempts to extend the seed beyond the end of the read from which the seed was obtained.

[0130] Alternatively, in other examples, the computer system can continue the execution of process 400 by determining at stage 415 that the response to the executed query includes (i) an extended record, (ii) an interval record, or (iii) both. In such examples, the computer system can continue the execution of process 400 by determining (430) whether the extended table is accessed to obtain one or more matching reference array positions in the extended table referred to by the interval record.

[0131] In some examples, the computer system can continue the execution of process 400 by determining that the seed extended table is accessed to obtain one or more matching reference array positions in the extended table re-referred to by the interval record. For example, in some implementations, the computer system can be configured to access the seed extended table to obtain one or more matching reference array positions identified by the interval record if the number of matching reference array positions is below a predetermined threshold. Alternatively, or in addition, the computer system can be configured to access the seed extended table to obtain one or more matching reference array positions identified by the interval record if the response to the executed query also includes a "stop" record stored at the hash position where the hash query seed reached. The "stop" record can preferentially instruct the computer system not to perform further seed extension of the seed in the query and not to access one or more matching reference array positions identified by the interval record, such as when the number of matching reference array positions is below a predetermined threshold.

[0132] In such an example where it is determined that the seed extension table is accessed at stage 430, the computer system can continue the execution of process 400 by accessing the seed extension table and obtaining one or more reference array positions within the seed extension table (450). The computer system can identify a particular set of one or more matching reference array positions for obtaining from the seed extension table by using interval records. The interval records can include information referring to multiple positions within the seed extension table, including data describing the reference array positions that match the query seed. In some implementations, the information referring to multiple positions can include consecutive intervals of reference array positions within the extension table that match the extended seed of the query. Alternatively, in other implementations, the information referring to multiple positions can include one or more discontinuous intervals of reference array positions within the extension table that match the query seed.

[0133] In such an example, the computer system can obtain one or more matching reference array positions from a seed extension table identified using interval records. The one or more obtained reference array positions can be added to a seed match set (455). In some implementations, adding one or more matching reference array positions to a seed match set can include obtaining and storing data representing the one or more matching reference array positions at a location in a memory device allocated for seed match set storage. In other implementations, adding one or more matching reference array positions to a seed match set can include storing data, such as a pointer, that references an interval (s) of a seed extension table that stores the one or more reference array positions. Thus, a seed match set can be a storage location that stores a set of identified and obtained matching reference array positions. Alternatively, a seed match set can include one or more storage locations that store references to one or more matching reference array positions. When the computer system adds one or more matching reference array positions identified by interval records to a seed match, this example of process 400 can end at 460.

[0134] In another example, after the computer system determines (415) that the response includes at least (i) an extended record, (ii) a spacing record, or (iii) both, the computer system can determine (430) that the seed extension table is not accessed to obtain one or more matching reference array positions. The determination by the computer system that the seed extension table is not accessed to obtain one or more matching reference array positions can be based on various factors. As an example, in some implementations, the computer system can determine not to access the seed extension table to obtain the matching reference array positions identified by the spacing record when the response returns an extended record. Such a determination may be preferred because the extended seed is likely to create a set of matching reference array positions that is smaller than the set of matching reference array positions identified by the spacing record.

[0135] As another example, in other implementations, the computer system can determine not to access the seed extension table to obtain the matching reference array positions identified by the spacing record when the computer system determines that the number of matching reference array positions exceeds a predetermined threshold number of matching reference array positions. Similarly, in such implementations, the computer system can determine not to access the seed extension table when the matching reference array positions identified by the spacing exceed the matching threshold.

[0136] If the computer system determines at 430 that it cannot access the seed extension table, the computer system can continue the execution of process 400 by determining at 465 whether the acquired response includes an interval record and an extension record. If the computer system determines at 465 that the acquired response includes an interval record and an extension record, the computer system can determine at 435 whether to store, as the best interval candidate, information describing the interval record or the interval record included in the response to the executed query. During the first iteration of process 400 for a query having an initial seed that has not yet been extended, the computer system can determine to store, in the best interval storage of the memory device, the interval record or the information describing the interval record as the best interval candidate. Such an interval record is encountered during the initial iteration of process 400 for a query having an unextended initial seed, and there are no other interval records encountered in response to other queries for one or more subsequent extended seeds. Therefore, the first interval returned in response to a query having an unextended initial seed must be the "best interval" because there are no other intervals identified for comparison.

[0137] However, for subsequent interactions by process 400 after a response is received for a query having an extended seed, the computer system can obtain a second interval record from the response to the query having the extended seed. In such an example, the computer system can heuristically determine whether the second interval record should be used to replace a previously stored best interval candidate in the best interval storage. The determination of whether to maintain the previously stored best interval candidate or replace the best interval candidate with the second interval or the information describing the interval can be made by applying one or more heuristic rules as described with reference to the embodiment of FIG. 3. In some implementations, the heuristic rules can include one or more multi-part heuristic rules.

[0138] Some implementations of the present disclosure can be aimed at repeatedly evaluating the comparison between each interval record returned later and the previously stored best interval candidate in order to determine the single best interval to be stored for the current lead based on the query seed, but the present disclosure need not be so limited. Instead, in some implementations, all intervals can be stored in interval storage and evaluated later for use in complementing the seed match set.

[0139] The computer system can continue the execution of process 400 by generating (440) an extended seed. The extended seed can be generated based on instructions included in the extended record returned in response to the query. As an example, when the extended record is executed by a computer such as a central processing unit (CPU), a graphics processing unit (GPU), or a programmable circuit 162 that executes software instructions, the CPU, GPU, or programmable circuit can reach the hash position storing the extended record for one or more nucleotides and extend the seed used in the hash query. In some implementations, an extended record can be generated such that the computer is instructed to extend the seed symmetrically on each end of the seed. Thus, as an example, the extended record can be generated such that the computer, such as a CPU, GPU, or programmable circuit 162, is instructed to extend the seed by two nucleotides, four nucleotides, six nucleotides, etc. In such implementations, the symmetric extension of the seed can be achieved by extending the seed by one nucleotide on each end of the seed, two nucleotides on each end of the seed, three nucleotides on each end of the seed, etc. However, the present disclosure should not be limited to the symmetric extension of the seed. Instead, the asymmetric extension of the seed is also contemplated by the present disclosure.

[0140] The computer system can continue the execution of process 400 by generating a hash query containing an extended seed at 445. Next, at stage 405, the computer system executes another iteration of process 400 by executing a query with the extended query, and then (a) continues the execution of process 400 until the process ends at 427 or 460 by adding one or more matching reference array positions to the seed match set, and the process ends at 475 after determining whether to store the interval record as the best interval candidate, or (c) the process ends at stage 420 as a result of one or more errors, such as a seed extension error that results in a query that does not receive a response to the executed query, including (i) an extended record, (ii) an interval record, or (iii) one or more matching reference array positions.

[0141] Alternatively, at stage 465, if the computer system determines that the acquired response does not include both the interval record and the extended record, the computer system can continue the execution of process 400 by determining whether the acquired response includes an extended record.

[0142] If the computer system determines that the acquired response includes an extended record, the computer system continues the execution of process 400 at stage 440 by generating an extended seed, generates a hash query 445 containing the extended seed, and can execute another iteration of process 400 by executing a query with the extended query at stage 405. Next, the computer system can (a) continue the execution of process 400 until the process ends at 427, 420, 460, 475.

[0143] On the one hand, if the computer system determines that the acquired response does not include an extended record, the computer system can continue the execution of process 400 at step 470 by determining whether to store the interval record or the information describing the interval record as the best interval candidate. The computer system can use the same process described for determining whether to store the interval record as the best interval candidate at step 435 to determine at step 470 whether to store the interval record as the best interval candidate. Regardless of whether the computer system determines at step 470 to store the interval record as the best interval candidate, process 400 ends at step 475.

[0144] At least one variation of process 400 can be implemented, and instead, the computer system determines at step 470 whether the acquired response includes an interval record. In such an example, logically, if the computer system determines that the acquired response includes an interval record, the computer system can continue the execution of the process at step 470. Instead, if the computer system determines that the acquired response does not include an interval record, the process continues at step 440 by generating an extended seed. Other variations of the process flow of process 400 can be implemented similarly and can fall within the spirit and scope of the present disclosure.

[0145] FIG. 5 is a flowchart of a process 500 for performing iterative runtime flexible seed extension for hash table genomic mapping for each seed of a read. Generally, process 500 includes obtaining (505) a read generated by a nucleic acid sequencer and determining a position of a seed access window, the seed access window identifying, determining (510) a seed of the read, generating (515) a hash query including the seed identified by the seed access window, starting (520) execution of process 400 as depicted in FIG. 4 at stage 410 by executing hash queries generated until process 400 ends and continuing iterative execution of process 400, and determining (525) whether the read includes another seed. If it is determined (525) that the read includes another seed, the seed access window is adjusted to identify another seed (530) and stage 515 is executed to generate a hash query using the other seed (515).

[0146] Process 500 can continue to execute the processing loops of steps 515, 520, 525, and 530 in step 525 until it is determined that the read obtained in step 505 does not include another seed to be mapped and aligned using process 400. In such an example, it can be determined (535) whether to complement the current set of seed matches for the read using the best interval. If it is determined in step 535 to complement the current seed match using the best interval, process 500 can continue by processing the best interval (540), complement the current set of seed matches (545) using one or more matching reference array positions obtained from a portion of the seed extension table identified using the best interval, and determine (550) whether there is another read that is ready for mapping and alignment using process 500. If there is no other read that is ready for mapping and alignment, process 500 ends at step 555. Instead, if there is another read that is ready for mapping and alignment using process 500, process 500 continues by obtaining in step 505 another read that is ready for mapping and alignment. Then, process 500 can continue to execute process 500 iteratively until it is determined in step 550 that there is no other read that is ready for mapping and alignment using process 500.

[0147] The process 500 is described in more detail below as being executed by a computer system of one or more computers. The one or more computers can include, for example, a mapping and alignment unit 170. For the purposes of the present disclosure, the one or more computers can include a CPU or GPU configured to obtain and execute software instructions to implement certain programmed functionality described by the software instructions. Alternatively or in addition, the one or more computers can include a programmable circuit configured such that hardware digital logic circuitry of the programmable circuit is configured to implement certain programmed functionality in hardware.

[0148] A computer system can initiate execution of process 500 by obtaining (505) data representing nucleic acid reads (also referred to herein as "reads") generated by a nucleic acid sequencer. Reads can be received by the computer system as input from the nucleic acid sequencer after the reads are generated by the nucleic acid sequencer. Alternatively, or in addition, reads generated by the nucleic acid sequencer may be stored in a memory device accessible by the computer system. The computer system 500 can then obtain the stored read(s) by accessing the memory to retrieve one or more reads from the memory device. By way of example, a read can include a set of nucleotides such as "ACGTTTAGC". This example includes a 9-nucleotide read. However, the use of a 9-nucleotide read is used only as an example. Instead of being limited to 9 nucleotides, reads as described by this disclosure can be of any nucleotide length including 5 bases or nucleotides, 10 bases or nucleotides, 12 bases or nucleotides, 15 bases or nucleotides, 18 bases or nucleotides, 21 bases or nucleotides, 25 bases or nucleotides, 35 bases or nucleotides, 50 bases or nucleotides, 100 bases of nucleotides, 150 bases or nucleotides, 1,000 bases or nucleotides, millions of bases or nucleotides, or even more bases or nucleotides, but is not limited thereto.

[0149] The computer system can continue the execution of process 500 by determining (510) the position of the seed access window. Using the seed access window, a nucleotide seed composed of a subset of the nucleotides of the read can be identified. An example of a seed is the set of consecutive nucleotides "GTTTA", which is the seed of the read "ACGTTTAGC". The set of consecutive nucleotides "GTTTA" represents an example of a consecutive seed of the read "ACGTTTAGC", but the present disclosure need not be so limited. Instead, in some implementations, non-consecutive seeds can be obtained and analyzed using the systems and processes described by the present disclosure. For example, a non-consecutive seed such as "G_T_A" can also be obtained from the read "ACGTTTAGC" and analyzed using the systems and methods described herein. In such implementations, the systems and methods of the present disclosure can handle skipped positions represented by an underscore "_" as wildcards that can match any base or nucleotide.

[0150] The seed access window can be configured to be any base or nucleotide length shorter than the read length. The seed access window can be configured to move in the forward or reverse direction along the consecutive read to identify the seed of the read for processing. If non-consecutive seeds are utilized, the seed access window can be configured accordingly. As an example, the seed access window can be configured to identify a non-consecutive seed of nine nucleotides with wildcards inserted at nucleotide position 6 and nucleotide position 8.

[0151] The computer system can continue the execution of process 500 by generating (515) a hash query that includes a seed identified by a seed access window. In some implementations, the hash query can simply consist of a seed of a read, such as "GTTTA" from a read of "ACGTTTAGC". In other implementations, additional data, metadata, etc. can be added to the sample seed to convert the seed into a format that can be used to search a hash table.

[0152] The computer system can continue the execution of process 500 by executing (520) the process 400 described by FIG. 4 and mapping and aligning the generated query seed to one or more reference array positions. The computer system starts the execution of process 400 by executing the hash query generated in step 515 at step 410. Then, the computer system can continue the iterative execution of process 400 until process 400 ends at steps 420, 427, 460, or 475, optionally adding reference array positions that match the seed match set at steps 425 or 455.

[0153] After process 400 ends, the computer system can determine (525) whether the read obtained at step 505 contains another seed. In some implementations, determining whether a read contains another seed includes considering all possible seed access window positions in the read. Alternatively, determining whether a read contains another seed can include considering only a predetermined subset of all possible seed access window positions, such as only even seed access window positions or only odd seed access window positions. Thus, the present disclosure does not require that each seed of a read be evaluated using process 500. Instead, in some implementations, those computer systems can determine at step 505 whether there is another seed in a predetermined subset of the seeds of the reads that are evaluated using process 500.

[0154] If the computer system determines at step 525 that the read contains another seed, the computer system can adjust the seed access window to identify the other seed (530), and the computer system can perform step 515 to generate a hash query using the other seed identified by the adjusted seed access window (515). Adjusting the seed access window can include, for example, moving the seed access window in the forward direction along the read obtained at step 505 by the position of one or more bases or nucleotides. The computer system can continue to execute the processing loop of steps 515, 520, 525, and step 530 at step 525 until it determines that the read obtained at step 505 does not contain another seed that is mapped and aligned using process 400.

[0155] When the computer system determines that the read obtained at step 505 does not contain another seed that is mapped and aligned, the computer system can determine (535) whether to complement the current set of seed matches for the read using the best interval. In some examples, if the computer system determines that the set of seed matches should not be complemented, the computer system can determine (550) whether there is another read that is ready for mapping and alignment using process 500. In such an example, if the computing system determines at step 550 that there is another read that is ready for mapping and alignment, the computer system can continue to execute process 500 by obtaining at step 505 another read that is ready for mapping and alignment. The computer system can then repeatedly execute process 500 until it is determined at step 550 that there is no other read that is ready for mapping and alignment.

[0156] Alternatively, in other examples, the computer system can determine that the current set of seed matches for the read should be complemented using one or more matching reference array positions identified by the best interval. The computer system can determine that the current set of seed matches should be complemented using one or more matching reference array positions identified by the best interval by applying one or more heuristic rules to (i) the seed length of the extended read that created the best interval, (ii) the seed length of one or more matching reference array positions, (iii) the number of generated seed chains, or a combination thereof. In some implementations, the heuristic rules can specify one or more independent trigger conditions that cause the computer system to process the best interval when triggered.

[0157] As an example, a first independent trigger condition that can trigger the processing of the best interval by the computer system is to determine whether the seed length of the extended read that created the best interval by the query is at least intvl-seed-length(60) bases or nucleotides. In this example, the threshold intvl-seed-length(60) is a predetermined threshold that can be used by the computer system to evaluate the length of the extended read that creates the best interval. In this example, the computer system checks the seed length of the extended read that created the best interval it examines for 60 nucleotides. However, the present disclosure need not be so limited. Instead, the threshold intvl-seed-length() can be set to any nucleotide length. If the computer system determines that the intvl-seed-length() threshold is not met, the computer system can evaluate other trigger conditions to determine whether the best interval is processed.

[0158] As another example, a second independent trigger condition that can trigger processing by a best interval computer system is to determine whether the seed length of the extended seed from which the query created the best interval is greater than the longest matching reference array position processed by at least intvl-seed-longer(8) bases or nucleotides. In this example, the threshold intvl-seed-longer(8) is a predetermined threshold that can be used by a computer system to evaluate the comparison between (i) the seed length of the extended seed from which the query created the best interval and (ii) the longest matching reference array position. In this example, if the computer system determines that the seed length of the extended seed from which the query created the best interval is 8 or more bases or nucleotides larger than any matching seed, the processing of the best interval is triggered.

[0159] As another example, a third independent trigger condition that can trigger processing by a best interval computer system is to determine whether the number of seed chains is less than intvl-min-chains(8). A seed chain can include a group of matches of similarly arranged reference array positions. In this example, the threshold intvl-min-chains(8) is a predetermined threshold that can be used to evaluate the number of generated seed chains. In this example, if fewer than 8 seed chains are generated, the processing of the best interval is triggered.

[0160] Examples of three independent trigger conditions for triggering processing of the best interval to complement the seed match set are described, but the present disclosure need not be so limited. Instead, other trigger conditions can be constructed to trigger processing of the best interval as may be required by a particular computer system. For example, if the computer system determines to complete the seed match at stage 535 because one or more thresholds of the trigger condition for processing the best interval are met, the computer system can determine to complete the current set of seed matches at stage 535 using the best interval. Completing the current set of seed matches using the best interval can include the computer system processing the best interval (540). Processing the best interval can include applying one or more heuristic rules to the best interval to identify one or more matching reference array positions identified by the best interval and stored in the seed extension table.

[0161] As an example, the computer system can determine to process all of the one or more reference array positions identified by the best interval if the number of reference array positions identified by the best interval is less than or equal to intvl - max - hits (64). In this example, if the computer system determines that the best interval identifies less than or equal to 64 matching reference array positions, the computer system can obtain all of the matching reference array positions identified by the best interval from the seed extension table using the best interval. Alternatively, if the computer system determines that the best interval identifies more than 64 matching reference array positions, the computer system can randomly obtain intvl - sample - hits (32) matching reference arrays from the set of matching reference array positions identified by the best interval.

[0162] Randomly obtaining a threshold amount of 32 matching reference array positions can include randomly obtaining, or obtaining by deterministic pseudo-random selection, a threshold amount of 32 matching reference array positions from a seed extension table using a best interval. The best interval can include data identifying (i) one or more stop and start positions of the seed extension table, (ii) one or more start positions and one or more offsets, or combinations thereof. Examples of thresholds such as 64 matching reference positions and 32 randomly sampled hits are described, but the present disclosure need not be so limited. Instead, other thresholds with other numerical values can be used to achieve the advantages of the present disclosure.

[0163] The matching reference array positions obtained using the best interval can be used to complement the current seed match set 545. Such a complement of the seed match set using the best interval can address problems such as unmapped read problems or high-confidence mis-mapping problems that may not result in, or may result in only a very small number of, the matching reference array positions stored in the seed match set. The matching reference array positions may be obtained from, or be obtainable from, a portion of the seed extension table identified by the best interval (540).

[0164] Once the seed match set is replenished, the computer system can determine whether there is another read for which mapping and alignment using process 500 are ready. If there is another read for which mapping and alignment are ready, the computer system continues execution of process 500 by obtaining the other read. Alternatively, if there is no other read for which mapping and alignment are ready, process 500 can end at 555.

[0165] In the example described with reference to process 500, the best interval is evaluated to determine whether the best interval or a portion of the best interval can be used to complement the set of seed matches. However, it is not necessary to store only a single best interval in the best interval storage. In some implementations, the computer system can facilitate the storage of two or more best intervals in the best interval storage. For example, in some implementations, up to two best intervals may be tracked. In some implementations, up to N best intervals may be tracked. In such implementations, when N > 1 best intervals are stored, the criteria for determining which interval is being held may involve an evaluation of the relationship between interval candidates, the associated extended seeds of the interval candidates, or both, such that the N best intervals are associated with extended seeds that do not overlap each other within the read. In some implementations, the computer system may even be capable of selecting a matching reference array position from among a plurality of different best intervals. Such selection of a matching reference array position from among a plurality of different best intervals can be performed randomly, pseudo-randomly, or by applying one or more heuristics. System components

[0166] Figure 6 is a diagram of system components that can be used to implement the system described herein related to flexible seed extension for hash table genome mapping.

[0167] Computing device 600 is intended to represent various forms of digital computers such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. In addition to this, computing device 600 or 650 can include a Universal Serial Bus (USB) flash drive. The USB flash drive can store an operating system and other applications. The USB flash drive can include input / output components such as a wireless transmitter or a USB connector that can be inserted into a USB port of another computing device. The components shown in this specification, the connections and relationships of this component, and the functions of this component are meant to be merely examples and are not meant to limit the implementation forms of the invention described and / or claimed in this document.

[0168] The computing device 600 includes a processor 602, a memory 604, a storage device 608, a high-speed interface 608 that connects to the memory 604 and a high-speed expansion port 610, and a low-speed interface 612 that connects to a low-speed bus 614 and the storage device 608. Each of the components 602, 604, 608, 608, 610, and 612 is interconnected using various buses and can be implemented on a common motherboard or in other suitable ways. The processor 602 processes instructions for execution within the computing device 600, including instructions stored in the memory 604 or the storage device 608, and can display graphical information regarding a GUI on an external input / output device such as a display 616 coupled to the high-speed interface 608. In other implementations, multiple processors and / or multiple buses can be used, along with multiple memories and multiple types of memories, as appropriate. Also, multiple computing devices 600 can be connected so that each device provides a portion of the required operations, for example, as a server bank, a blade server group, or a multiprocessor system.

[0169] The memory 604 stores information within the computing device 600. In one implementation, the memory 604 is a volatile memory unit or multiple volatile memory units. In another implementation, the memory 604 is a non-volatile memory unit or multiple non-volatile memory units. The memory 604 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0170] The memory device 608 can provide large-scale storage for the computing device 600. In one implementation, the memory device 608 can be, or can contain, a computer-readable medium such as an array of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other device within another configuration. The computer program product can be tangibly embodied in an information carrier. The computer program product can also contain instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable or machine-readable medium such as the memory 604, the memory device 608, or the memory on the processor 602.

[0171] The high-speed controller 608 manages the bandwidth-intensive operations of the computing device 600, while the low-speed controller 612 manages the low-bandwidth-intensive operations. Such a functional assignment is merely an example. In one implementation, the high-speed controller 608 is coupled, for example, to the memory 604, the display 616, and a high-speed expansion port 610 that can receive various expansion cards (not shown) via a graphics processor or an accelerator. In this implementation, the low-speed controller 612 is coupled to the storage device 608 and a low-speed expansion port 614. The low-speed expansion port can include various communication ports, such as USB, Bluetooth, Ethernet, and wireless Ethernet, and can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a microphone / speaker pair, a scanner, or a networking device such as a switch or a router, via, for example, a network adapter. The computing device 600 can be implemented in several different forms, as shown in the figure. For example, the computing device 600 can be implemented as a standard server 620 or multiple times within a group of such servers. The computing device 600 can also be implemented as part of a rack server system 624. In addition, the computing device 600 can be implemented in a personal computer such as a laptop computer 622. Alternatively, the components from the computing device 600 can be combined with other components within a mobile device (not shown) such as the device 650. Each of such devices can include one or more of the computing devices 600, 650, and the entire system can be composed of multiple computing devices 600, 650 that communicate with each other.

[0172] As shown in the figure, computing device 600 can be implemented in several different forms. For example, computing device 600 can be implemented as a standard server 620 or multiple times within a group of such servers. Computing device 600 can also be implemented as part of a rack server system 624. In addition, computing device 600 can be implemented in a personal computer such as a laptop computer 622. Alternatively, components from computing device 600 can be combined with other components within a mobile device (not shown) such as device 650. Each of such devices can include one or more of computing devices 600, 650, and the entire system can be composed of multiple computing devices 600, 650 that communicate with each other.

[0173] Computing device 650 includes, among other components, a processor 652, a memory 664, and input / output devices such as a display 654, a communication interface 666, and a transceiver 668. Device 650 can also include a storage device such as a microdrive or other device to provide additional storage. Each of components 650, 652, 664, 654, 666, and 668 are interconnected using various buses, and some of the components can be implemented on a common motherboard or in other suitable ways as appropriate.

[0174] Processor 652 can execute instructions within computing device 650, including instructions stored in memory 664. The processor can be implemented as a chipset of chips including separate and multiple analog and digital processors. In addition to this, the processor can be implemented using any of several architectures. For example, processor 610 can be a CISC (Complex Instruction Set Computers) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimal Instruction Set Computer) processor. The processor can provide coordination of other components of device 650, such as, for example, control of the user interface, applications executed by device 650, and wireless communication by device 650.

[0175] Processor 652 can communicate with a user via a control interface 658 and a display interface 656 coupled to a display 654. The display 654 can be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display, an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 656 can include appropriate circuitry for driving the display 654 to present graphical information and other information to the user. The control interface 658 can receive commands from the user and convert them for submission to the processor 652. Additionally, an external interface 662 can be provided to communicate with the processor 652 to enable proximity area communication between the device 650 and other devices. The external interface 662 can provide, for example, wired communication in some implementations or wireless communication in other implementations, and multiple interfaces can also be used.

[0176] Memory 664 stores information within computing device 650. Memory 664 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. Also, for example, an expansion memory 674 can be provided and connected to device 650 via an expansion interface 672 that can include, for example, a SIMM (Single In Line Memory Module) card interface. Such an expansion memory 674 can provide additional storage space for device 650 or can also store applications or other information for device 650. Specifically, expansion memory 674 can include instructions to execute or complement the processes described above and can also include secure information. Thus, for example, expansion memory 674 can be provided as a security module for device 650 and can be programmed with instructions to enable secure use of device 650. Additionally, secure applications can be provided with additional information, such as by placing identification information on the SIMM card in a hack-proof manner via the SIMM card.

[0177] The memory can include, for example, flash memory and / or non-volatile random-access memory (NVRAM) memory, as described later. In one implementation, a computer program product is tangibly implemented in an information carrier. The computer program product, when executed, includes instructions to execute one or more methods such as those described above. The information carrier is a computer-readable medium or machine-readable medium such as, for example, memory 664, expansion memory 674, or memory on processor 652 that can be received via transceiver 668 or external interface 662.

[0178] Device 650 can wirelessly communicate via a communication interface 666 that can include a digital signal processing circuit as needed. The communication interface 666 can provide communication under various modes or protocols such as, among others, GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA (registered trademark), CDMA2000, or GPRS. Such communication can be performed, for example, via a radio transceiver 668. In addition, short-range communication can be performed, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 670 can provide additional navigation-related and location-related wireless data to device 650, which can be used as appropriate by applications operating on device 650.

[0179] Device 650 can also communicate audibly using an audio codec 660, which can receive speech information from a user and convert this speech information into usable digital information. The audio codec 660 can also generate audible sounds for the user, such as via a speaker, for example, within the handset of device 650. Such sounds can include sounds from a voice telephone call, can include recorded sounds, such as voice messages, music files, etc., and can also include sounds generated by applications operating on device 650.

[0180] Computing device 650 can be implemented in several different forms, as shown in the figure. For example, computing device 650 can be implemented as a mobile phone 680. Also, computing device 650 can be implemented as part of a smartphone 682, a personal digital assistant, or other similar mobile device.

[0181] Various implementations of the systems and methods described herein can be realized in digital electronic circuitry, integrated circuitry, ASICs (application specific integrated circuits) specifically designed therefor, computer hardware, firmware, software, and / or combinations of such implementations. These various implementations can be either special-purpose or general-purpose, and can be executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device. They can include implementations in one or more computer programs that are executable and / or interpretable on such programmable system.

[0182] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” “computer-readable medium” refer to any computer program product, apparatus, and / or device, for example, magnetic disks, optical disks, memory, Programmable Logic Devices (PLDs) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0183] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a pointing device by which the user can provide input to the computer, such as a mouse or trackball. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, including acoustic input, speech input, or tactile input.

[0184] The systems and techniques described herein can be implemented in a computing system that includes backend components, such as a data server, or in a computing system that includes middleware components, such as an application server, or in a computing system that includes frontend components, such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein, or in any combination of such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.

[0185] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is created by computer programs that operate on respective computers and have a client-server relationship with each other. Example

[0186] The present disclosure is further described in the following examples, which do not limit the scope of the claims. Example 1: Comparison of the Percentage of Unmapped Reads between a System Using Flexible Seed Extension and a System Not Using Flexible Seed Extension

[0187] In this example, different nucleic acid sequencers including the HiSeq® 2500 sequencer, HiSeq® X sequencer, and NovaSeq® sequencer were used to array a specific sample. Next, using the DRAGEN™ platform, reads generated by each sequencer were mapped with and without using flexible seed extension as described herein. Once mapped, the computer system determined the percentage of unmapped reads resulting from each mapping operation for each sequencer.

[0188] The DRAGEN™ platform is a mapping and alignment unit implemented in the hardware circuitry of a field programmable gate array (FPGA). The DRAGEN™ v7 platform does not currently utilize flexible seed extension as described herein, whereas the DRAGEN™ v8 platform utilizes flexible seed extension. Although the DRAGEN™ platform used herein was implemented in an FPGA, generally, the DRAGEN™ platform can also be implemented in other integrated circuits such as an application specific integrated circuit (ASIC).

[0189] Specifically, the "DNA_Nexus_hiseq2500" sample was sequenced using the HiSeq (registered trademark) 2500 sequencer, the "DNA_Nexus_hiseqX" sample was sequenced using the HiSeq (registered trademark) X sequencer, and the "DNA_Nexus_NovaSeq" sample, "NovaSeq_NA12878_rep1 sample", "NovaSeq_TruSeq-nano-550 sample", and "AWS_HG005_40x" sample were sequenced using the NovaSeq (registered trademark) sequencer. "AWS_HG005_40X" was derived from subject HG005. All of the other samples were derived from subject HG001.

[0190] FIG. 7 is an explanatory diagram of a bar graph 700 that displays data representing test results in the form of the percentage of unmapped reads in a system that uses a flexible seed extension method as described herein, compared to a system that does not use a flexible seed extension method. Bar graph 700 is a graphical display of test results 710, 720, 730, 740, 750, and 760 that compare the results of mapping operations performed on genomic reads generated by different sequencing devices of Illumina, Inc.

[0191] In the first example, test result 710 shows that the percentage of unmapped reads that occur at 710b when the HiSeq (registered trademark) 2500 sequencer sequences the "DNA_Nexus_hiseq2500" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the percentage of unmapped reads that occur at 710a when the HiSeq (registered trademark) 2500 sequencer sequences the "DNA_Nexus_hiseq2500" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0192] In the second embodiment, test result 720 indicates that the proportion of unmapped reads that occur at 720b when the NovaSeq® sequencer sequences the "DNA_Nexus_NovaSeq" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the proportion of unmapped reads that occur at 720a when the NovaSeq sequencer sequences the "DNA_Nexus_NovaSeq" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0193] In the third embodiment, test result 730 indicates that the proportion of unmapped reads that occur at 730b when the HiSeq® X sequencer sequences the "DNA_Nexus_hiseqX" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the proportion of unmapped reads that occur at 730a when the HiSeq X sequencer sequences the "DNA_Nexus_hiseqX" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0194] In the fourth embodiment, test result 740 indicates that the proportion of unmapped reads that occur at 740b when the NovaSeq® sequencer sequences the "NovaSeq_NA12878_rep1" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the proportion of unmapped reads that occur at 740a when the NovaSeq sequencer sequences the "NovaSeq_NA12878_rep1" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0195] In the fifth embodiment, test result 750 indicates that the percentage of unmapped reads that occur at 750b when the NovaSeq™ sequencer sequences the "NovaSeq_TruSeq-nano-550" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the percentage of unmapped reads that occur at 750a when the NovaSeq™ sequencer sequences the "NovaSeq_TruSeq-nano-550" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0196] In the sixth embodiment, test result 760 indicates that the percentage of unmapped reads that occur at 760b when the NovaSeq™ sequencer sequences the "AWS_HG005_40X" sample and utilizes flexible seed extension as described in one or more implementations herein during mapping is significantly smaller than the percentage of unmapped reads that occur at 760a when the NovaSeq™ sequencer sequences the "AWS_HG005_40X" sample without utilizing flexible seed extension as described in one or more implementations herein during mapping.

[0197] Accordingly, the implementation of flexible seed extension using the hash table described herein achieves a significant performance improvement in reducing unmapped reads when compared to conventional methods that do not generate or use the hash table described herein. Example 2: Comparison of read mapping accuracy between a system using flexible seed extension and a system not using flexible seed extension

[0198] In this Example 2, the DRAGEN™ platform was used to map the reads generated by the nucleic acid sequencer to the reference sequence. Each DRAGEN™ platform mapped the same read set to the same reference sequencer. Once the mapping was complete, the computer system determined the read mapping accuracy of each mapping operation as a function of the mapping error rate.

[0199] The DRAGEN™ platform is a mapping and alignment unit implemented in the hardware circuitry of a field programmable gate array (FPGA). The DRAGEN™ v7 platform does not currently utilize flexible seed extension as described herein, whereas the DRAGEN™ v8 platform and the DRAGEN™ v8 hi-effort platform utilize flexible seed extension.

[0200] The DRAGEN™ platform used herein was implemented in an FPGA, but in general, the DRAGEN™ platform can be implemented in other integrated circuits such as application specific integrated circuits (ASICs).

[0201] The differences between the DRAGEN™ v8 platform and the DRAGEN™ v8 hi-effort platform are the heuristic and other parameter settings. The DRAGEN™ v8 platform uses the following heuristics: intvl-target-hits = 32, intvl-max-hits = 16, and intvl-sample-hits = 16. Each of these heuristics is described herein. In addition, the DRAGEN™ v8 platform uses other parameters: max-hifreq-hits = 16, rescue-hifreq = 0, and sw-extra-intvl = 1. The max-hifreq-hits parameter indicates the maximum number of random sample matches taken from the match intervals reached before a failed seed extension (e.g., one sample per failed extension until a limit is reached). The rescue-hifreq parameter determines whether an expensive rescue scan operation is utilized for matches found only by random samples from the match intervals. The rescue scan is a method for searching for possible fitting read alignments near read alignment candidates. The sw-extra-intvl parameter determines the policy for utilizing the expensive Smith-Waterman alignment for matches by accessing the best ("extra") intervals or by randomly sampling the match intervals. Smith-Waterman is generally not used if the gapless alignment is not clipped, but may be utilized depending on the heuristic that includes this setting if the gapless alignment is clipped. A setting of "1" means that Smith-Waterman can be used for candidates from the extra / best match intervals that are accessed in their entirety but not by random sampling. A setting of "2" means that Smith-Waterman can also be used for candidates from random sampling of the match intervals.For a setting of 「0」, it means that Smith-Waterman is not applied for candidates from extra / best interval processing or from random sampling of match intervals.

[0202] On the other hand, the DRAGEN™ v8 hi-effort platform uses the following heuristics, namely, intvl-target-hits = 32, intvl-max-hits = 64, and intvl-sample-hits = 48. In addition, the DRAGEN™ v8 hi-effort platform uses other parameters of max-hifreq-hits = 32, rescue-hifreq = 0, and sw-extra-intvl = 2. Therefore, the DRAGEN™ v8 hi-effort platform has a more comprehensive set of heuristics than the DRAGEN™ v8 platform.

[0203] FIG. 8 is an explanatory diagram of a line graph 800 that displays data representing test results in the form of read mapping accuracy in a system using a flexible seed extension method as disclosed herein, compared to a system that does not use a flexible seed extension method. In particular, graph 800 shows the trade-off between false detection and missed detection when data is stratified using an accuracy metric, using an accuracy curve in the form of a receiver operating characteristic (「ROC」) curve (or line). In the explanatory diagram of FIG. 8, a curve (or line) close to the upper wall and the left wall of graph 800 means better read mapping accuracy.

[0204] Curve 810 representing the read mapping accuracy of the DRAGEN™ v7 platform without using flexible seed extension as described in one or more implementations herein during mapping is depicted. Curve 820 representing the read mapping accuracy of the DRAGEN™ v8 platform using flexible seed extension as described in one or more implementations herein during mapping is depicted. By comparing curve 810 and curve 820, it is clear that curve 820 is closer to the upper wall and the left wall than curve 810. Therefore, the improvement in read mapping accuracy is achieved by simply implementing flexible seed extension as described in one or more implementations herein with some capacity.

[0205] FIG. 8 further depicts curve 830 representing the “hi-effort” DRAGEN™ implementation form v8. Similar to the DRAGEN™ v8 implementation form, the DRAGEN™ v8 hi-effort implementation form also utilizes a flexible seed extension method as described herein during mapping. However, as described above, the heuristics utilized by the DRAGEN™ v8 “hi-effort” implementation form are more substantial than the heuristics used to utilize the DRAGEN™ v8 implementation form whose performance is represented by curve 820. The “hi-effort” v8 version of DRAGEN™ is assigned parameters (e.g., sw-extra-intvl = 2) that increase the willingness to perform more Smith-Waterman alignment work downstream for the DRAGEN™ v8 implementation form (e.g., sw-extra-intvl = 1). As shown in FIG. 8, curve 830 exhibits a significant performance gain in read mapping accuracy by the DRAGEN™ v8 h-effort implementation form by being closer to the upper wall and the left wall than both curves 810 and 820.

[0206] FIG. 8 also depicts a curve 840 representing the read mapping accuracy achieved by the BWA-MEM software mapping tool. The BWA-MEM software mapping tool uses the Burrows-Wheeler Transform (BWT) of the reference genome as an index of the BWA-MEM software mapping tool. This way of representing the reference genome can inherently provide similar benefits as those provided by flexible seed extension, such as the ability to extract a complete set of matches corresponding to matches of any length. As depicted by curves 830 and 840, the DRAGEN™ v8 “hi-effort” implementation can achieve the same read mapping accuracy as the software-based BWA software mapping tool. Thus, it is important that DRAGEN™ v8 “hi-effort” can achieve a read mapping accuracy level equivalent to that of the software-based BWA mapping tool because DRAGEN™ v8 “hi-effort” also utilizes other benefits of the DRAGEN™ platform, including, for example, fewer memory accesses for mapping seeds. However, prior to the hardware-based flexible seed extension implementation described herein, the DRAGEN™ platform was able to achieve the same level of read mapping accuracy as that achieved by the BWA software mapping tool.

[0207] Thus, the flexible seed extension implementation using the hash table described herein achieves a significant performance improvement in terms of read mapping accuracy when compared to conventional methods that do not generate or use the hash table described herein. Other embodiments

[0208] Some embodiments have been described. Nevertheless, it will be understood that various changes can be made without departing from the spirit and scope of the present invention. Additionally, the logical flow depicted in the figures does not require the specific order or sequential order shown to achieve the desired result. Additionally, other steps can be provided from the described flow, or steps can be eliminated, and other components can be added to or removed from the described system. Accordingly, other embodiments are within the scope of the following claims.

Explanation of Signs

[0209] 100 System 110 Computer 112 Memory 114 Reference Array 120 Seed Extension Tree 121 Node 122 Node 123 Node 124 Node 125 Node 126 Node 130 Memory 132 Seed Extension Table 140 Hash Table 142 Index Key 143 Hash Function 144 Hash Location 146 Hash Table Installation Package 150 Record 152 Record 153 Record 155 Reference Array Location 160 Device 162 Circuit 170 Mapping and Alignment Unit 180 Memory 300 Runtime System 305 Read 310 Hash Query 340 Best Interval Storage 350 Optimal Interval Storage 600 Computing Device 602 Processor 604 Memory 610 High-Speed Expansion Port 612 Low-Speed Controller 614 Low-Speed Expansion Port 616 Display 620 Standard Server 622 Laptop Computer 624 Rack Server System 650 Device 652 Processor 654 Display 656 Display Interface 658 Control Interface 660 Audio Codec 662 External Interface 664 Memory 666 Communication Interface 668 Transceiver 670 Receiver Module 672 Expansion Interface 674 Expansion Memory 680 Mobile Phone 682 Smartphone

Claims

1. 1. A method for generating a hash table for mapping sample reads to a reference sequence, comprising: Obtaining, by a computer system, a first seed of nucleotides from a reference sequence, the first seed having a length of K nucleotides; determining, by the computer system, that the first seed matches more than a predetermined number of reference sequence positions; generating, by the computer system, a seed extension tree having a plurality of nodes based on determining that the first seed matches more than a predetermined number of reference sequence positions, wherein each node of the plurality of nodes comprises: an extension of the first seed, and * an extended seed having a nucleotide length of K * is one or more nucleotides greater than K; and generating, corresponding to one or more positions in a seed extension table, the positions including data describing reference sequence positions that match the extended seeds; For each node of the plurality of nodes, storing, by the computer system, interval information at a location in the hash table corresponding to an index key of the decompressed seed, the interval information referencing one or more locations in the seed extension table that contain data describing a reference array location that matches the decompressed seed associated with the node.

2. The method of claim 1 , wherein each of the matching reference sequence positions comprises the K nucleotides of the first seed.

3. obtaining, by the computer system, a second seed from the reference sequence that is a different nucleotide from the first seed; determining, by the computer system, that the second seed does not match more than the predetermined number of reference sequence positions; based on the computer system determining that the second seed does not match more than the predetermined number of reference sequence positions, obtaining, by the computer system, data describing each of the reference sequence positions that match the second seed; 2. The method of claim 1, further comprising: storing, by the computer system, the data describing the reference array location that matches the second seed in a second location of the hash table that corresponds to an index key of the second seed.

4. 4. A method according to any one of claims 1 to 3, wherein the one or more locations in the seed extension table containing data describing reference sequence positions that match the extended seed comprises a contiguous interval of locations in the seed extension table containing data describing reference sequence positions that match the extended seed.

5. 4. A method according to any one of claims 1 to 3, wherein the one or more locations in the seed extension table containing data describing reference sequence positions that match the extended seed comprise intervals in the extension table of reference sequence positions that match the extended seed that are non-contiguous.

6. obtaining, by a computer system, a first seed of nucleotides from a reference sequence, the first seed representing a sequence of nucleotides having a nucleotide length of K nucleotides; determining, by the computer system, a location of a seed access window within a reference sequence; and obtaining, by the computer system, a subset of the reference sequences identified by the seed access window.

7. adjusting, by the computer system, the seed extension window forward along the reference sequence by K nucleotides to identify a second seed of nucleotides from the reference sequence having a nucleotide length of K nucleotides; obtaining, by the computer system, the second seed from the reference sequence; determining, by the computer system, that the second seed matches more than a predetermined number of reference sequence positions; generating, by the computer system, a second seed extension tree having a plurality of second nodes based on determining that the second seed matches more than a predetermined number of reference sequence positions, wherein each second node of the plurality of second nodes (i) is an extension of the second seed and * a second extended seed having a nucleotide length of K * is one or more nucleotides greater than K, and (ii) one or more second positions in a second seed extension table that includes data describing reference sequence positions that match said second extended seed; For each second node of the plurality of second nodes, 7. The method of claim 6, further comprising: storing, by the computer system, second interval information at a location in the hash table corresponding to an index key of the second decompressed seed, the second interval information referencing one or more locations in the second seed decompression table that contain data describing a reference array location that matches the second decompressed seed associated with the second node.

8. For each node of the plurality of nodes, determining, by the computer system, whether the node of the seed extension tree is a leaf node; 8. The method of claim 1, further comprising: storing, by the computer system, a decompression record at the location of the hash table that corresponds to the index key of the decompressed seed based on the computer system determining that the node of the decompressed tree is not a leaf node.

9. 9. The method of claim 8, wherein the extension record comprises one or more instructions that, when executed by the computer system, cause the computer system to add one or more additional nucleotides to a seed associated with the extension record.

10. 8. The method of claim 1, further comprising: determining, by the computer system, based on the computer system determining that the node decompression tree is a leaf node, not to store a decompression record at the location of the hash table corresponding to the index key of the decompressed seed.

11. generating, by the computer system, the seed extension table, said generating comprising: identifying, by the computer system, each seed of the reference sequence that matches the first seed; The method of any one of claims 1 to 7, further comprising storing, by said computer system, data identifying said identified seed in said seed extension table.

12. The method of any one of claims 1 to 11, further comprising sorting, by the computer system, the identified seeds in the seed extension table.

13. 13. The method of any one of claims 1 to 12, further comprising generating, by the computer system, a hash table install package that includes instructions that, when processed by one or more computers that receive the hash table install package, cause the one or more computers to install the hash table in memory accessible by a programmable logic circuit.

14. the hash table installation package includes the seed extension table; 14. The method of claim 13, wherein the hash table installation package includes instructions to instruct (i) the programmable logic circuit or (ii) another computer to store the seed extension table in the memory device accessible to the programmable logic circuit.

15. 15. The method of claim 13, further comprising providing, by the computer system, the hash table installation package to another computer.

16. The method of claim 15 , wherein the other computer is (i) a computer configured to communicate with the programmable logic circuit, or (ii) includes the programmable logic circuit.

17. The method of any one of claims 1 to 16, wherein the computer system comprises a plurality of computers.

18. 1. A system for generating a hash table mapping sample reads to a reference sequence, comprising: One or more computers and one or more storage devices storing instructions operable, when executed by the one or more computers, to cause the one or more computers to: obtaining, by one or more computers, a first seed of nucleotides from a reference sequence, the first seed having a length of K nucleotides; determining, by the one or more computers, that the first seed matches more than a predetermined number of reference sequence positions; generating, by the one or more computers, a seed extension tree having a plurality of nodes based on determining that the first seed matches more than a predetermined number of reference sequence positions, wherein each node of the plurality of nodes (i) is an extension of the first seed and * an extended seed having a nucleotide length of K * is one or more nucleotides greater than K, and (ii) one or more positions in a seed extension table that includes data describing reference sequence positions that match said extended seed; For each node of the plurality of nodes, and storing, by the one or more computers, interval information at a location in the hash table corresponding to an index key of the decompressed seed, the interval information referencing one or more locations in the seed extension table that contain data describing a reference array location that matches the decompressed seed associated with the node.

19. 20. The system of claim 18, wherein each of the matching reference sequence positions comprises the K nucleotides of the first seed.

20. The operation, obtaining, by the one or more computers, a second seed from the reference sequence that is a different nucleotide from the first seed; determining, by the one or more computers, that the second seed does not match more than the predetermined number of reference sequence positions; based on determining, by the one or more computers, that the second seed does not match more than the predetermined number of reference sequence positions; obtaining, by the one or more computers, data describing each of the reference sequence positions that match the second seed; 20. The system of claim 18, further comprising: storing, by the one or more computers, the data describing the reference array location that matches the second seed in a second location of the hash table that corresponds to an index key of the second seed.

21. 21. A system according to any one of claims 18 to 20, wherein the one or more locations in the seed extension table containing data describing reference sequence positions that match the extended seed comprises a contiguous interval of locations in the seed extension table containing data describing reference sequence positions that match the extended seed.

22. 21. A system according to any one of claims 18 to 20, wherein the one or more locations in the seed extension table containing data describing reference sequence locations that match the extended seed comprise a non-contiguous interval in the extension table of reference sequence locations that match the extended seed.

23. obtaining, by one or more computers, a first seed of nucleotides from a reference sequence, the first seed representing a sequence of nucleotides having a nucleotide length of K nucleotides; determining, by the one or more computers, a location of a seed access window within a reference sequence; and obtaining, by the one or more computers, a subset of the reference sequences identified by the seed access window.

24. The operation, adjusting, by the one or more computers, the seed extension window forward along the reference sequence by K nucleotides to identify a second seed of nucleotides from the reference sequence having a nucleotide length of K nucleotides; obtaining, by the one or more computers, the second seed from the reference sequence; and determining, by the one or more computers, that the second seed matches more than a predetermined number of reference sequence positions; generating, by the one or more computers, a second seed extension tree having a plurality of second nodes based on determining that the second seed matches more than a predetermined number of reference sequence positions, wherein each second node of the plurality of second nodes (i) is an extension of the second seed and * a second extended seed having a nucleotide length of K * is one or more nucleotides greater than K, and (ii) one or more second positions in a second seed extension table that includes data describing reference sequence positions that match said second extended seed; For each second node of the plurality of second nodes, 24. The system of claim 23, further comprising: storing, by the one or more computers, second interval information at a location in the hash table corresponding to an index key of the second extended seed, the second interval information referencing one or more locations in the second seed extension table that contain data describing a reference array location that matches the second extended seed associated with the second node.

25. The operation, For each node of the plurality of nodes, determining, by the one or more computers, whether the node of the seed extension tree is a leaf node; 25. The system of claim 18, further comprising: storing, by the one or more computers, a decompression record at the location in the hash table that corresponds to the index key of the decompressed seed based on the one or more computers determining that the node of the decompressed tree is not a leaf node.

26. 26. The system of claim 25, wherein the extension record comprises one or more instructions that, when executed by the one or more computers, cause the one or more computers to add one or more additional nucleotides to a seed associated with the extension record.

27. The operation, 26. The system of claim 18, further comprising: determining, by the one or more computers, based on determining, by the one or more computers, that the node extension tree is a leaf node, not to store an extension record in the location of the hash table that corresponds to the index key of the extended seed.

28. The operation, generating, by the one or more computers, the seed extension table, identifying, by the one or more computers, each seed of the reference sequence that matches the first seed; and storing, by said one or more computers, data identifying said identified seeds in said seed extension table.

29. The operation, The system of any one of claims 18 to 28, further comprising sorting, by the one or more computers, the identified seeds in the seed extension table.

30. The operation, 29. The system of any one of claims 18 to 28, further comprising generating, by the one or more computers, a hash table installation package that includes instructions that, when processed by one or more computers that receive the hash table installation package, cause the one or more computers to install the hash table in memory accessible by a programmable logic circuit.

31. the hash table installation package includes the seed extension table; 31. The system of claim 30, wherein the hash table installation package includes instructions to instruct (i) the programmable logic circuit or (ii) another computer to store the seed extension table in a memory device accessible to the programmable logic circuit.

32. The operation, 32. The system of claim 30 or 31, further comprising providing, by the one or more computers, the hash table installation package to another computer.

33. 33. The system of claim 32, wherein the other computer is (i) a computer configured to communicate with the programmable logic circuitry; or (ii) the programmable logic circuitry.

34. The system of any one of claims 18 to 32, wherein the one or more computers comprises a plurality of computers.

35. A non-transitory computer-readable medium storing software including instructions executable by one or more computers, the instructions, when executed, causing the one or more computers to: obtaining, by one or more computers, a first seed of nucleotides from a reference sequence, the first seed having a length of K nucleotides; determining, by the one or more computers, that the first seed matches more than a predetermined number of reference sequence positions; generating, by the one or more computers, a seed extension tree having a plurality of nodes based on determining that the first seed matches more than a predetermined number of reference sequence positions, wherein each node of the plurality of nodes (i) is an extension of the first seed and * an extended seed having a nucleotide length of K * is one or more nucleotides greater than K, and (ii) one or more positions in a seed extension table that includes data describing reference sequence positions that match said extended seed; For each node of the plurality of nodes, and storing, by the one or more computers, interval information at a location in the hash table that corresponds to an index key of the decompressed seed, the interval information referencing one or more locations in the seed extension table that contain data describing a reference sequence location that matches the decompressed seed associated with the node.

36. 36. The computer-readable medium of claim 35, wherein each of the matching reference sequence positions comprises the K nucleotides of the first seed.

37. The operation, obtaining, by the one or more computers, a second seed from the reference sequence that is a different nucleotide from the first seed; determining, by the one or more computers, that the second seed does not match more than the predetermined number of reference sequence positions; based on determining, by the one or more computers, that the second seed does not match more than the predetermined number of reference sequence positions; obtaining, by the one or more computers, data describing each of the reference sequence positions that match the second seed; 36. The computer-readable medium of claim 35, further comprising: storing, by the one or more computers, the data describing the reference array location that matches the second seed in a second location of the hash table that corresponds to an index key of the second seed.

38. 38. The computer readable medium of claim 35, wherein the one or more locations in the seed extension table that contain data describing reference sequence positions that match the extended seed comprise a contiguous interval of locations in the seed extension table that contain data describing reference sequence positions that match the extended seed.

39. 38. The computer readable medium of claim 35, wherein the one or more locations in the seed extension table that contain data describing reference sequence locations that match the extended seed comprise a non-contiguous interval in the extension table of reference sequence locations that match the extended seed.

40. obtaining, by one or more computers, a first seed of nucleotides from a reference sequence, the first seed representing a sequence of nucleotides having a nucleotide length of K nucleotides; determining, by the one or more computers, a location of a seed access window within a reference sequence; and obtaining, by the one or more computers, a subset of the reference sequences identified by the seed access window.

41. The operation, adjusting, by the one or more computers, the seed extension window forward along the reference sequence by K nucleotides to identify a second seed of nucleotides from the reference sequence having a nucleotide length of K nucleotides; obtaining, by the one or more computers, the second seed from the reference sequence; and determining, by the one or more computers, that the second seed matches more than a predetermined number of reference sequence positions; generating, by the one or more computers, a second seed extension tree having a plurality of second nodes based on determining that the second seed matches more than a predetermined number of reference sequence positions, wherein each second node of the plurality of second nodes (i) is an extension of the second seed and * a second extended seed having a nucleotide length of K * is one or more nucleotides greater than K; and (ii) one or more second positions in a second seed extension table that includes data describing reference sequence positions that match said second extended seed. For each second node of the plurality of second nodes, 41. The computer-readable medium of claim 40, further comprising: storing, by the one or more computers, second interval information at a location in the hash table corresponding to an index key of the second decompressed seed, the second interval information referencing one or more locations in the second seed decompression table that contain data describing a reference array location that matches the second decompressed seed associated with the second node.

42. The operation, For each node of the plurality of nodes, determining, by the one or more computers, whether the node of the seed extension tree is a leaf node; 42. The computer-readable medium of claim 35, further comprising: storing, by the one or more computers, a decompression record at the location of the hash table that corresponds to the index key of the decompressed seed based on the one or more computers determining that the node of the decompressed tree is not a leaf node.

43. 43. The computer readable medium of claim 42, wherein the extension record comprises one or more instructions that, when executed by the one or more computers, cause the one or more computers to add one or more additional nucleotides to a seed associated with the extension record.

44. The operation, 43. The computer-readable medium of claim 35, further comprising: determining, by the one or more computers, based on determining, by the one or more computers, that the node extension tree is a leaf node, not to store an extension record in the location of the hash table that corresponds to the index key of the extended seed.

45. The operation, generating, by the one or more computers, the seed extension table, identifying, by the one or more computers, each seed of the reference sequence that matches the first seed; and storing, by the one or more computers, data identifying the identified seed in the seed extension table.

46. The operation, The computer readable medium of any one of claims 35 to 45, further comprising sorting, by the one or more computers, the identified seeds in the seed extension table.

47. The operation, 46. ​​The computer readable medium of any one of claims 35-45, further comprising generating, by the one or more computers, a hash table install package that includes instructions that, when processed by one or more computers that receive the hash table install package, cause the one or more computers to install the hash table in memory accessible by a programmable logic circuit.

48. the hash table installation package includes the seed extension table; 48. The computer readable medium of claim 47, wherein the hash table installation package includes instructions to instruct (i) the programmable logic circuit or (ii) another computer to store the seed extension table in a memory device accessible to the programmable logic circuit.

49. The operation, 49. The computer readable medium of claim 47 or 48, further comprising providing, by the one or more computers, the hash table installation package to another computer.

50. 50. The computer readable medium of claim 49, wherein the other computer is (i) a computer configured to communicate with the programmable logic circuitry; or (ii) includes the programmable logic circuitry.

51. The computer readable medium of any one of claims 35 to 50, wherein the one or more computers comprises a plurality of computers.

52. 1. A method for using a hash table for mapping sample reads to a reference sequence, comprising: performing, by a mapping and aligning unit, a query of a hash table, the query including a first seed, the first seed including a subset of nucleotides obtained from a particular one of the sample reads; obtaining a response to the performed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the executed query comprises (i) an extension record, (ii) an interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query includes (i) an extension record and (ii) an interval record; determining, by said mapping and aligning unit, whether an extension table is accessed to obtain one or more matching reference sequence positions in said extension table referenced by said interval record; based on determining that the decompression table is not accessed, determining, by said mapping and aligning unit, whether to store said first information describing said interval record in a memory device as information describing a best interval candidate; generating, by the mapping and aligning unit, a first extended seed using the extended record, the first extended seed being an extension of the first seed; generating, by the mapping and aligning unit, a subsequent hash query that includes the first extended seed; performing, by the mapping and aligning unit, the subsequent query of the hash table.

53. The method further comprising: based on determining that the decompression table is accessed, accessing, by said mapping and aligning unit, said extension table to obtain said one or more matching reference sequence positions in said extension table referenced by said interval record; 53. The method of claim 52, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

54. The method further comprising: determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; 54. The method of claim 52 or 53, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

55. determining, by the mapping and aligning unit, whether to store the first information describing the interval record in a memory device as information describing a best interval candidate; determining, by the mapping and aligning unit, that there is no prior information describing the interval record as the best interval candidate for the particular read; 55. A method according to any one of claims 52 to 54, comprising storing, by said mapping and aligning unit, said first information describing said interval record in said memory device as information describing a best interval candidate.

56. The method further comprising: obtaining a response to the subsequent executed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the subsequent executed query includes (i) a second extension record, (ii) a second interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the subsequent executed query includes (i) the second extension record and (ii) the second interval record; determining whether an extension table is accessed by the mapping and aligning unit to obtain one or more matching reference sequence positions in the extension table referenced by the second interval record; based on determining that the decompression table is not accessed, determining, by the mapping and aligning unit and using one or more heuristic rules, whether second information describing the second interval record or the first information describing the best interval candidate is used as the best interval candidate; generating, by the mapping and aligning unit, a second extended seed using the second extended record, the second extended seed being an extension of the first extended seed; generating, by the mapping and aligning unit, a third hash query that includes the second extended seed; 56. The method of claim 52, further comprising: performing, by the mapping and aligning unit, the third query of the hash table with the second extended seed.

57. determining, by the mapping and aligning unit and using one or more heuristic rules, whether the second information describing the second interval record or the first information describing the best interval candidate is to be used as the best interval; 57. The method of claim 56, comprising selecting either the second information describing the second interval record or the first information describing the best interval candidate record based on a number of factors including: (i) a number of matching reference sequence positions returned by each of the interval record and the second interval record, (ii) a predetermined threshold level of reference sequence positions, or (iii) a respective seed length of the respective seeds that arrived at the hash position storing the interval record and the second interval record.

58. 58. A method according to any one of claims 52 to 57, wherein the interval record references one or more positions in the seed extension table that contain data describing reference sequence positions that match the first seed of the query.

59. the one or more locations in the seed extension table that contain data describing reference sequence locations that match the first seed of the query; 60. The method of claim 58, further comprising: a contiguous interval of reference sequence positions in an extension table that match the first seed of the query.

60. 1. A system for improving the mapping of sample reads to reference sequences using a hash table, comprising: One or more computers and one or more storage devices storing instructions operable, when executed by the one or more computers, to cause the one or more computers to: performing, by a mapping and aligning unit, a query of a hash table, the query including a first seed, the first seed including a subset of nucleotides obtained from a particular one of the sample reads; obtaining a response to the performed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the executed query comprises (i) an extension record, (ii) an interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query includes (i) an extension record and (ii) an interval record; determining, by said mapping and aligning unit, whether an extension table is accessed to obtain one or more matching reference sequence positions in said extension table referenced by said interval record; based on determining that the decompression table is not accessed, determining, by said mapping and aligning unit, whether to store said first information describing said interval record in a memory device as information describing a best interval candidate; generating, by the mapping and aligning unit, a first extended seed using the extended record, the first extended seed being an extension of the first seed; generating, by the mapping and aligning unit, a subsequent hash query that includes the first extended seed; performing, by the mapping and aligning unit, the subsequent hash query of the hash table.

61. The operation, based on determining that the decompression table is accessed, accessing, by said mapping and aligning unit, said extension table to obtain said one or more matching reference sequence positions in said extension table referenced by said interval record; 61. The system of claim 60, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

62. The operation, determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; 62. The system of claim 60 or 61, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

63. determining, by the mapping and aligning unit, whether to store the first information describing the interval record in a memory device as information describing a best interval candidate; determining, by the mapping and aligning unit, that there is no prior information describing the interval record as the best interval candidate for the particular read; Storing, by the mapping and aligning unit, the first information describing the interval record in the memory device as information describing a best interval candidate.

64. The operation, obtaining a response to the subsequent executed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the subsequent executed query includes (i) a second extension record, (ii) a second interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the subsequent executed query includes (i) the second extension record and (ii) the second interval record; determining whether an extension table is accessed by the mapping and aligning unit to obtain one or more matching reference sequence positions in the extension table referenced by the second interval record; based on determining that the decompression table is not accessed, determining, by the mapping and aligning unit and using one or more heuristic rules, whether second information describing the second interval record or the first information describing the best interval candidate is used as the best interval candidate; generating, by the mapping and aligning unit, a second extended seed using the second extended record, the second extended seed being an extension of the first extended seed; generating, by the mapping and aligning unit, a third hash query that includes the second extended seed; 64. The system of claim 60, further comprising: performing, by the mapping and aligning unit, the third query of the hash table with the second extended seed.

65. determining, by the mapping and aligning unit and using one or more heuristic rules, whether the second information describing the second interval record or the first information describing the best interval candidate is to be used as a best interval; 65. The system of claim 64, further comprising selecting either the second information describing the second interval record or the first information describing the best interval candidate record based on a number of factors including: (i) a number of matching reference sequence positions returned by each of the interval record and the second interval record, (ii) a predetermined threshold level of reference sequence positions, or (iii) a respective seed length of the respective seeds that reached the hash position storing the interval record and the second interval record.

66. 66. The system of claim 60, wherein the interval record references one or more positions in the seed extension table that contain data describing reference sequence positions that match the first seed of the query.

67. the one or more locations in the seed extension table that contain data describing reference sequence locations that match the first seed of the query; 67. The system of claim 66, further comprising a contiguous interval of reference sequence positions in an extension table that match the first seed of the query.

68. A non-transitory computer-readable medium storing software including instructions executable by one or more computers, the instructions, when executed, causing the one or more computers to: performing, by a mapping and aligning unit, a query of a hash table, the query including a first seed, the first seed including a subset of nucleotides obtained from a particular one of the sample reads; obtaining a response to the performed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the executed query comprises (i) an extension record, (ii) an interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query includes (i) an extension record and (ii) an interval record; determining, by said mapping and aligning unit, whether an extension table is accessed to obtain one or more matching reference sequence positions in said extension table referenced by said interval record; based on determining that the decompression table is not accessed, determining, by said mapping and aligning unit, whether to store said first information describing said interval record in a memory device as information describing a best interval candidate; generating, by the mapping and aligning unit, a first extended seed using the extended record, the first extended seed being an extension of the first seed; generating, by the mapping and aligning unit, a subsequent hash query that includes the first extended seed; performing, by the mapping and aligning unit, the subsequent query of the hash table.

69. The operation, based on determining that the decompression table is accessed, accessing, by said mapping and aligning unit, said extension table to obtain said one or more matching reference sequence positions in said extension table referenced by said interval record; 69. The computer readable medium of claim 68, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

70. The operation, determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the executed query contains one or more matching reference sequence positions; 70. The computer readable medium of claim 68 or 69, further comprising adding, by the mapping and aligning unit, the one or more matching reference sequence positions to a seed match set.

71. determining, by the mapping and aligning unit, whether to store the first information describing the interval record in a memory device as information describing a best interval candidate; determining, by the mapping and aligning unit, that there is no prior information describing the interval record as the best interval candidate for the particular read; Storing, by the mapping and aligning unit, the first information describing the interval record in the memory device as information describing a best interval candidate.

72. The operation, obtaining a response to the subsequent executed query including information stored by the location of the hash table determined by the mapping and aligning unit to be responsive to the query; determining, by the mapping and aligning unit, whether the response to the subsequent executed query includes (i) a second extension record, (ii) a second interval record, or (iii) one or more matching reference sequence positions; based on determining, by the mapping and aligning unit, that the response to the subsequent executed query includes (i) the second extension record and (ii) the second interval record; determining whether an extension table is accessed by the mapping and aligning unit to obtain one or more matching reference sequence positions in the extension table referenced by the second interval record; based on determining that the decompression table is not accessed, determining, by the mapping and aligning unit and using one or more heuristic rules, whether second information describing the second interval record or the first information describing the best interval candidate is used as the best interval candidate; generating, by the mapping and aligning unit, a second extended seed using the second extended record, the second extended seed being an extension of the first extended seed; generating, by the mapping and aligning unit, a third hash query that includes the second extended seed; 72. The computer-readable medium of claim 68, further comprising: performing, by the mapping and aligning unit, the third query of the hash table including the second extended seed.

73. determining, by the mapping and aligning unit and using one or more heuristic rules, whether the second information describing the second interval record or the first information describing the best interval candidate is to be used as the best interval; 73. The computer-readable medium of claim 72, further comprising selecting either the second information describing the second interval record or the first information describing the best interval candidate record based on a number of factors including: (i) a number of matching reference sequence positions returned by each of the interval record and the second interval record, (ii) a predetermined threshold level of reference sequence positions, or (iii) a respective seed length of the respective seeds that reached the hash position storing the interval record and the second interval record.

74. 74. The computer readable medium of any one of claims 68 to 73, wherein the interval record references one or more locations in the seed extension table that contain data describing reference sequence positions that match the first seed of the query.

75. the one or more locations in the seed extension table that contain data describing reference sequence locations that match the first seed of the query; 75. The computer readable medium of claim 74, comprising a contiguous interval of reference sequence positions in an extension table that match the first seed of the query.

Citation Information

Patent Citations

  • Quick comparing and positioning method for gene sequence segments on reference genome

    CN105243297A

  • Method and device for quick contrast and analysis of short sequence for second-generation sequencing

    CN106295250A

  • Optimal alignment computation device and program

    JP2012032975A

  • Assembling pre-treatment apparatus, program for assembling pre-treatment apparatus, and method for controlling computer cluster system

    JP2013106567A

  • Genomics infrastructure for on-site or cloud-based DNA and RNA processing and analysis

    JP2019510323A