HARDWARE-ACCELERATED K-MERO GRAPH GENERATION
Patent Information
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2021-12-15
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for generating K-mer graphs are computationally intensive and time-consuming, occupying valuable software resources and limiting the efficiency of genomic data processing.
Utilizing a programmable logic device with hardware-accelerated K-mer graph generation through a control machine managing non-segmented hardware logical units, which generates and updates graph description data to facilitate parallel processing and accelerate the K-mer graph creation.
Significantly reduces the time required for K-mer graph generation, freeing up software resources for other genomic tasks while enhancing productivity by processing multiple segments simultaneously.
Smart Images

Figure MX433829B0
Abstract
Description
HARDWARE-ACCELERATED K-MERO GRAPH GENERATION Cross-reference to related applications This application claims the benefit of U.S. provisional patent application no. 63 / 006,668, filed on April 7, 2020, the contents of which are incorporated herein by reference in their entirety. Background K-mer grates can be used to represent a plurality of readings of the sequence. Summary According to an innovative aspect of the present description, a method for the hardware-accelerated generation of a K-mer graph using a programmable logic device is described. In one aspect, the method may include actions for obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence; generating, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer graph using the first set of nucleic acid sequences obtained, wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations; wherein each node of the K-mer graph represents a K-mer, and each edge of the K-mer graph represents a link between a pair of K-mers.and each weight of each edge of the K-number graph represents a number of occurrences of a sequence of K-numbers represented by a pair of K-numbers; and, during the generation of the K-number graph: periodically updating, with a control machine, the graph description data of the K-number graph after the performance of one or more operations by each logical hardware unit used to generate at least a portion of the K-number graph, wherein the graph description data represents (i) an identifier of the K-number graph and (ii) state information of the K-number graph, wherein the control machine creates a workflow of operations using the logical hardware units, Lzoczn / zznz / q / YiAi not segmented by activating the realization of one or more operations of each respective hardware logic unit during the generation of the K-numbers. Other versions include corresponding systems and devices configured to perform the actions of the methods mentioned above defined by hardware logic circuits of a hardware-accelerated gratos generation unit. These and other versions may optionally include one or more of the following features. For example, in some implementations, the output of each hardware logical unit in the plurality of hardware logical units is stored through a hash table cache. In some implementations, the control machine is implemented using a hardware logic unit of the programmable logic device. In some implementations, the control machine is implemented using one or more central processing units (CPUs) or graphics processing units (GPUs) to execute software instructions to perform the control machine functionality. In some implementations, the operations may further include providing the generated K-number graph to a variant call unit, wherein the variant call unit processes the K-number graph to determine candidate variants from one or more of the plurality of reads and the reference sequence. In some implementations, software instructions may be executed by one or more CPUs or G PUs to perform one or more variant call unit functions. In some implementations, the programmable logic device is used to accelerate one or more functions of the variant call unit. In some implementations, graph description data further include (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed hardware logic on the K-mer graph or the nucleic acid sequences of the stack associated with the identifier of the K-mer graph. According to another innovative aspect of the present description, a system for the hardware-accelerated generation of a K-mer gram using a programmable logic device is described. In one aspect, the system may include a hardware-accelerated gram generation unit comprising digital hardware logic circuits arranged to perform operations. In some implementations, the operation may comprise: obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence; generating, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer gram using the first set of nucleic acid sequences obtained;wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-numbers array represents a K-number, each edge of the K-numbers array represents a link between a pair of K-numbers, and each weight of each edge of the K-numbers array represents a number of occurrences of a sequence of K-numbers represented by a pair of K-numbers, and, during the generation of the K-numbers array: periodically updating, with a control machine, the description of the data of the K-numbers array after the performance of one or more operations by each hardware logic unit used to generate at least a portion of the K-numbers array, wherein the description data of the array represents (i) an identifier of K-numbers arrays and (ii) state information of K-numbers arrays,where the control machine creates a workflow of operations using the non-segmented hardware logic units by activating the execution of one or more operations of each respective hardware logic unit during the generation of the K-number graph. Other versions include corresponding methods and apparatus for performing the operations mentioned above. These and other versions may optionally include one or more of the following features. For example, in some implementations, the output of each hardware logical unit in the plurality of hardware logical units is stored through a hash table cache. I 7QC7n / 77n7 / 3 / YILI In some implementations, the operations may further include providing the generated K-numbers array to a variant call unit, wherein the variant call unit is configured to process the K-numbers array to determine candidate variants among one or more of the plurality of reads and the reference sequence. In some implementations, the system may further include one or more computers and one or more memory devices that store instructions which, when executed by the one or more computers, cause the one or more computers to perform secondary operations of a variant call unit. In some implementations, the secondary operations of the variant call unit may include obtaining, by means of the variant call unit, the generated K-mer sequence and identifying, based on the processing of the generated K-mer sequence by the variant call unit, one or more candidate variants, where a candidate variant is a difference between a base call from one or more reads in the read stack and a nucleotide of a reference genome at a particular location in the reference genome. In some implementations, the operations may also include obtaining, through a variant call unit, the generated K-mer grade and identifying, based on the processing of the generated K-mer grade by the variant call unit, one or more candidate variants, where a candidate variant is a difference between a base call of one or more reads in the read stack and a nucleotide of a reference genome at a particular location in the reference genome. In some implementations, the grate description data may further include (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer grate or nucleic acid sequences of the stack associated with the K-mer grate identifier. According to another innovative aspect of the present description, a hardware-accelerated gene generation unit is described. In one aspect, the hardware-accelerated gene generation unit may include digital logic hardware circuits arranged to perform operations. In some implementations, the operations may include obtaining a first set of nucleic acid sequences, where the first set of sequences The nucleic acid LZQCZn / ZZnZ / q / YIAI includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence, generating, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer graph using the first set of nucleic acid sequences obtained, wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-mer graph represents a K-mer, each edge of the K-mer graph represents a link between a pair of K-mers, and each edge weight of the K-mer graph represents a number of occurrences of a K-mer sequence represented by a pair of K-mers, and, during the generation of the K-mer graph: periodically updating, with a control machine,The description of the K-number graph data after the performance of one or more operations by each hardware logic unit used to generate at least a portion of the K-number graph, wherein the graph description data represents (i) a K-number graph identifier and (ii) K-number graph state information, wherein the control machine creates a workflow of operations using the non-segmented hardware logic units by triggering the performance of one or more operations by each respective hardware logic unit during the generation of the K-number graph. Other implementations may include methods and systems that are configured to perform hardware circuit operations of the hardware-accelerated gratos generation unit. These and other versions may optionally include one or more of the following features. For example, in some implementations, the output of each hardware logical unit in the plurality of hardware logical units is stored through a hash table cache. In some implementations, the operations may further include providing the generated K-number graph to a variant call unit, wherein the variant call unit is configured to process the K-number graph to determine candidate variants from one or more of the plurality of reads and the reference sequence. In some implementations, the operations may also include obtaining, through a variant call unit, the generated K-number graph and identifying, based on the I 7QC7n / 77n7 / 3 / YILI processing of the K-mer grad generated by the variant call unit, one or more candidate variants, wherein a candidate variant is a difference between a base call of one or more reads in the read array and a nucleotide of a reference genome at a particular location in the reference genome. In some implementations, the grate description data may further include (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer grate or nucleic acid sequences of the stack associated with the K-mer grate identifier. According to another innovative aspect of the present description, a method for hardware-accelerated generation of a K-mer array in a programmable logic device is described. In one aspect, the method may include actions for obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence, for each particular nucleic acid sequence in the first set of nucleic acid sequences; generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a array node for each K-mer of the particular nucleic acid sequence; and detecting, by means of a control machine,that the first hardware logic unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence, configure, by means of the control machine, a second hardware logic unit to perform the generation of graph edges for the generated graph nodes, and for one or more pairs of the generated graph nodes: generate, by means of the second hardware logic unit and for storage in the graph hash table, data representing graph edges between one or more pairs of the generated graph nodes generated by the first hardware logic unit, wherein the data representing the graph node for each K-mer stored in the hash table cache and the data representing graph edges stored in the hash table cache represent a K-mer graph of the first set of nucleic acid sequences. Other versions include corresponding systems and devices configured to perform the actions of the methods mentioned above defined by hardware logic circuits of a hardware-accelerated gratos generation unit. I 7QC 70 / 7707 / 3 / YILI These and other versions may optionally include one or more of the following features. For example, in some implementations, the method may also include periodically storing, by the control machine and in a memory unit accessible by the control machine, grate description data for an instance of the Kmeros grate, where the grate description data represents (i) an identifier of the Kmeros grate and (ii) state information of the Kmeros grate. In some implementations, the first hardware logic unit can be further configured to: determine whether one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence and, based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence, store data that marks the one or more particular K-mers as non-unique K-mers. In some implementations, the second hardware logic is further configured to: assign an edge weight to each edge of the K-numbers array. In some implementations, the method may also include instructing a third logical hardware unit of the programmable logic device to execute hardware logic configured to: obtain data representing the K-number graph from the hash table cache, and provide the obtained data representing the K-number graph to a variant call unit. In some implementations, the method may also include instructing a third logical hardware unit of the programmable logic device to execute hardware logic configured to: selectively remove data representing graph nodes and data representing graph edges from the K-number graph of the hash table cache. In some implementations, the control machine is implemented using a third hardware logic unit of the programmable logic device. I 7QC 70 / 7707 / 3 / YILI In some implementations, the hash table cache is implemented using a third hardware logic unit of the programmable logic device. In some implementations, the control machine is implemented using one or more CPUs or GPUs that execute software instructions to perform the control machine functionality. In some implementations, the graph description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic in the K-mer graph or the nucleic acid sequences of the stack associated with the K-mer graph identifier. In some implementations, the method may also include evaluating the K-number graph to check for cycles. In such implementations, if a cycle is detected during the evaluation, the process may include terminating the generation of the K-number graph. Alternatively, if no cycle is detected during the evaluation, the method may include retrieving data from the hash table cache that describes the structure of the K-number graph and providing this data to a variant designation module. According to another innovative aspect of the present description, a system for the hardware-accelerated generation of a K-mer graph using a programmable logic device is described. The system may include a hardware-accelerated graph generation unit comprising digital logic circuits arranged to perform operations. In one aspect, the operations may include obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence, for each particular nucleic acid sequence in the first set of nucleic acid sequences; generating, for storage in a hash table cache and by means of a first hardware logic unit,data representing a graph node for each K-mer of the particular nucleic acid sequence, detecting, by means of a control machine, that the first hardware logic unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence, configuring, by means of the control machine, a second hardware logic unit to perform the generation of graph edges for the generated graph nodes, and for one or more pairs of the generated graph nodes: generating, by means of the second hardware logic unit and for storage in the graph hash table, data that, I 7QC7n / 77n7 / 3 / YILI represent grato edges between one or more pairs of grato nodes generated by the first logical hardware unit, wherein the data representing the grato node for each K-mer stored in the hash table cache and the data representing grato edges stored in the hash table cache represent a K-mer graph of the first set of nucleic acid sequences. Other versions include corresponding methods and apparatus for performing the operations mentioned above. These and other versions may optionally include one or more of the following features. For example, in some implementations, the operations may also include periodically storing, by the control machine and in a memory unit accessible by the control machine, graph description data for an instance of the Kmeros graph, where the graph description data represents (i) an identifier of the Kmeros graph or (ii) state information of the Kmeros graph. In some implementations, the first hardware logic unit is further configured to: determine whether one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence and, based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence, store data that marks the one or more particular K-mers as non-unique K-mers. In some implementations, the second hardware logic is further configured to: assign an edge weight to each edge of the K-numbers array. In some implementations, the operations may also include instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: obtain data representing the K-numbers range from the hash table cache, and provide the obtained data representing the K-numbers range to a variant call unit. In some implementations, the operations may also include instructing a third hardware logic unit of the programmable logic device to execute the logic of I 7QC 70 / 7707 / 3 / YILI hardware configured to: selectively remove data representing graph nodes and data representing graph edges from the K-number graph of the hash table cache. In some implementations, the hash table cache is implemented using a third hardware logic unit of the programmable logic device. In some implementations, the graph description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic in the K-mer graph or the nucleic acid sequences of the stack associated with the K-mer graph identifier. In some implementations, the method may also include evaluating the K-number graph to check for cycles. In such implementations, if a cycle is detected during the evaluation, the process may include terminating the generation of the K-number graph. Alternatively, if no cycle is detected during the evaluation, the method may include retrieving data from the hash table cache that describes the structure of the K-number graph and providing this data to a variant designation module. According to another innovative aspect of the present description, a hardware-accelerated graph generation unit is described. The hardware-accelerated graph generation unit may include hardware digital logic circuits arranged to perform operations. In one aspect, the operations may include obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence, for each particular nucleic acid sequence in the first set of nucleic acid sequences; generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a graph node for each K-mer of the particular nucleic acid sequence; and detecting, by means of a control machine,that the first logical hardware unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence, configure, by means of the control machine, a second logical hardware unit to perform the generation of graph edges for the generated graph nodes, and for one or more pairs of the generated graph nodes:, LZQCZn / ZZnZ / q / YIAI generate, by means of the second hardware logic unit and for storage in the scatter table of the grate, data representing grate edges between one or more pairs of the grate nodes generated by the first hardware logic unit, wherein the data representing the grate node for each K-mer stored in the scatter table cache and the data representing grate edges stored in the scatter table cache represent a grate of K-mers from the first set of nucleic acid sequences. Other implementations may include methods and systems that are configured to perform hardware circuit operations of the hardware-accelerated gratos generation unit. These and other versions may optionally include one or more of the following features. For example, in some implementations, the operations may also include periodically storing, by the control machine and in a memory unit accessible by the control machine, grate description data for an instance of the Kmeros grate, where the grate description data represents (i) an identifier of the Kmeros grate and (ii) state information of the Kmeros grate. In some implementations, the first hardware logic unit can be further configured to determine whether one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence and, based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence matches another K-mer of the particular nucleic acid sequence, store data that marks the one or more particular K-mers as non-unique K-mers. In some implementations, the second hardware logic is further configured to: assign an edge weight to each edge of the K-number graph. In some implementations, the operations may also include instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: obtain data representing the K-number graph from the hash table cache, and provide the obtained data representing the K-number graph to a variant call unit. I 70070 / 7707 / 3 / YILI In some implementations, the operations may also include instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: selectively remove data representing the grate nodes and data representing the grate edges of the K-number grate from the hash table cache. In some implementations, the hash table cache is implemented using a third hardware logic unit of the programmable logic device. In some implementations, the grate description data may also include (i¡¡) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer grate or nucleic acid sequences of the stack associated with the K-mer grate identifier. In some implementations, the operations may also include evaluating the K-meros grate to check for the existence of grate cycles and, if a grate cycle is detected during the evaluation: terminate the generation of the K-meros grate or, if no grate cycle is detected during the evaluation: obtain data from the hash table cache that describes the structure of the k-meros grate, and provide, to a variant call module, the obtained data that describes the structure of the k-meros grate. These and other aspects of the present description are described in greater detail in the following detailed description with reference to the attached figures. Brief description of the figures Figure 1 is an example of a system for hardware-accelerated generation of a K-number. Figure 2 is a flowchart of an example of a process for hardware-accelerated generation of a K-number. Figure 3 is a flowchart of another example of a process for hardware-accelerated generation of a K-number. LZQCZn / ZZnZ / q / YIAI Figure 4 is an example of a K-number graph. Figure 5 is a block diagram of an example of system components that can be used for the hardware-accelerated K-number graph. Detailed description This description refers to the hardware-accelerated generation of a K-meros graph. Generating the K-meros graph using hardware circuitry significantly reduces the time required and transfers the computationally intensive K-meros graph generation process from a software processor to the hardware logic of an integrated circuit, such as a field-programmable gate array or an application-specific integrated circuit (ASIC). This frees up the software processor's resources, which can then be used for other genomic data processing tasks. Hardware-accelerated generation of K-number graphs can be achieved using a control machine configured to manage a workflow of operations performed by a plurality of non-segmented hardware logic units. Specifically, the control machine can abstractly achieve high-level segmented functionality using these non-segmented hardware logic units. This functionality is achieved by storing and updating graph description data, which includes (i) a K-number graph identifier that identifies an instance of a K-number graph and (ii) K-number graph state information.The state information of a K-number graph can include, for example, data indicating the last hardware logic unit that operated on raw graph data for a particular instance of a K-number graph, data indicating whether the last hardware logic unit aborted the operation, data indicating a length in K-numbers, data indicating a list of K-number nodes, data indicating a list of pointers that can be used to identify K-number nodes in a cache, the length of a list of K-number nodes, data indicating a list of non-unique K-numbers, or any subset or combination thereof. The control machine manages the generation of K-number graphs by calling a particular hardware logic unit to perform an operation on raw graph data and providing an updated set of graph description data to the called hardware logic unit. This storage and updating of grade description data allows the control machine to manage the parallel processing of non-segmented hardware logic units in a way that enables each non-segmented hardware logic unit to operate on data corresponding to a different K-number grade. Consequently, in addition to the increased speed benefits achieved by using hardware logic instead of executing software instructions to generate K-number grades, this description achieves further accelerated operation by increasing productivity by generating segments of different K-number grades simultaneously using different hardware logic units managed by the control machine. Figure 1 is an example of a system 100 for the hardware-accelerated generation of a K-mer gram. In some implementations, the system 100 may include a nucleic acid sequencer 110, a reference sequence database 120, a hardware-accelerated gram generation unit 130, a plurality of hardware logic units 131, 132, 133, 134, 135, 136, 137, 138, a control machine 140, a gram hash table cache 150, dynamic random-access memory (DRAM) 160, and a variant call unit 180. In some implementations, the hardware-accelerated gram generation unit 130 may be implemented using a programmable logic circuit, such as a field-programmable gate array (FPGA).In other implementations, the hardware-accelerated grating unit 130 can be implemented using an application-specific integrated circuit (ASIC). In either of these situations, the functionality described with respect to the hardware-accelerated grating unit 130, and each of the components implemented therein, is implemented using hardware logic circuits arranged to perform the functionality described herein without executing software instructions to perform the functionality. The term “unit” is used in this specification to describe a software module, a hardware module, or a combination of both, used to perform a specified function. A determination of whether a particular “unit” described herein is hardware, software, or a combination of both can be made based on the context of its use. For example, an “input unit” 131, a “graph node unit” 132, a “graph edge unit” 133, or the like, residing in a hardware-accelerated graph generation unit 130 implemented using an FPGA or ASIO, is a hardware unit whose functionality is realized by means of hard-wired digital logic gates or hard-wired digital logic blocks arranged to perform the functionality described herein with respect to the particular “unit.” By way of another example, a “variant call unit” 180 not implemented using a hardware-accelerated graph generation unit 130 in Figure 1 is a software module whose functionality is realized by means of one or more computers executing software instructions that define the functionality of the “variant call unit” 180.As another example, a computer or processing unit can be a hardware device that performs its functionality by processing software instructions, and therefore the functionality of the computer or processing unit is a combination of hardware and software. Although examples of one or more components in Figure 1 are provided in this description as hardware implementations, such as the “control machine” 140, because the “control machine” is depicted in Figure 1 as implemented in the hardware-accelerated grade generation unit 130, this description is not limited to such examples. Instead, other implementations may be used where the “control machine” 140 is implemented in software as a software module or a combination of hardware and software with a computer or processing unit executing software instructions to perform the functionality of the “control machine” 140 described herein. Similarly, there may be implementations of this description where certain components described as software with respect to Figure 1, such as the “variant call unit” 180, are implemented as a hardware implementation. The nucleic acid sequencer 110 is a device configured to perform primary analysis. Primary analyses may include receiving a biological sample 105, such as a blood sample, tissue sample, sputum, or nucleic acid sample, using the nucleic acid sequencer 110, and generating output data such as one or more reads 112, each representing a nucleotide sequence order from the nucleic acid sequence of the received biological sample. In some implementations, sequencing using the nucleic acid sequencer 110 may be performed in multiple read cycles, with the first read cycle generating one or more The first reads include a string of base calls representing a nucleotide sequence from one end of a nucleic acid sequence fragment, and a second read cycle generates one or more respective second reads, each including a string of base calls representing a nucleotide sequence from the other ends of one of the nucleic acid sequence fragments. In some implementations, the reads can be generated using clonal amplification. In the example in Figure 1, the one or more reads can include a pilot stack of reads for a particular location in the reference genome, with the reference genome location composed of multiple sequential locations in the reference genome. Therefore, each read corresponds to data representing a portion of a nucleic acid genome for an organism, such as an animal, insect, plant, or the like. Assuming short fragments of the nucleic acid sequence of approximately 600 base calls, a first read might represent 150 nucleotides ordered for the first end of the nucleic acid sequence fragment, and a second read might represent 150 nucleotides ordered for the other end of the nucleic acid sequence fragment. However, these numbers are merely examples, and any nucleic acid sequencer can be configured to generate reads that can be processed by a hardware-accelerated read generation unit as described herein using any sequencing method. Such reads may be of different lengths than those mentioned herein.For example, in some implementations, the present description can be used to generate hardware-accelerated K-mer graphs for reads generated from fragments of nucleic acid sequences that are up to 1000 nucleotides or more in length, where each read has, for example, 50 base calls, 75 base calls, 150 base calls, 200 base calls, 300 base calls, 500 base calls, or more from the end of each fragment. Each base call can correspond to one nucleotide. The present description can also be used to generate hardware-accelerated K-mer graphs for long reads. Accordingly, the hardware-accelerated graph generation unit 130 can be used to generate K-mer graphs for any read generated in any way by any type of nucleic acid sequencer. In some implementations, the biological sample 105 may include a DNA sample, and the nucleic acid sequencer 110 may include a DNA sequencer. In such implementations, the order of nucleotides sequenced in a read generated by the The LZQCZn / ZZnZ / q / YIAI nucleic acid sequencer may include one or more guanine (G), cytosine (C), adenine (A), and thymine (T) sequences in any combination. In some implementations, the 110 nucleic acid sequencer may be used to sequence RNA samples. In some implementations, this may occur using RNA sequencing protocols. For example, an RNA sample may be preprocessed using reverse transcription to form complementary DNA (cDNA) using a reverse transcriptase enzyme. In other implementations, the 110 nucleic acid sequencer may include an RNA sequencer, and the biological sample may include an RNA sample. Consequently, although the example in Figure 1 describes a nucleic acid sequencer that produces reads comprising G, C, A, and T generated by a DNA sequencer based on a DNA sample, this description is not limited to it.In contrast, other implementations can process reads comprising C, GA, and U that are generated by an RNA sequencer based on an RNA sample. In some implementations, the DNA or RNA reads generated by the nucleic acid sequencer may include an N-base call, where N is indicative of an unknown base call generated by the nucleic acid sequencer. In some implementations, the nucleic acid sequencer 110 may include a next-generation sequencer (NGS) configured to generate sequence reads such as reads 112 from a given sample in a manner that achieves ultra-high productivity, scalability, and speed through the use of massively parallel sequencing technology. NGSs enable rapid sequencing of entire genomes, the ability to zoom in on deeply sequenced target regions, the use of RNA sequencing (Sec-RNA) to discover novel RNA variants and splice sites, or the quantification of mRNAs for gene expression analysis, the analysis of epigenetic factors such as genome-wide DNA methylation and DNA-protein interactions, the sequencing of cancer samples to study rare somatic and tumor subclones, and the study of microbial diversity in humans or the environment. The nucleic acid sequencer 110 can obtain a reference genome 122 from the reference genome database 122. In some implementations, only a portion of the reference genome 122 is obtained. The portion of the reference genome 122 that is obtained may correspond to the reference locations of the reference genome 122 with respect to which the read stack 112 is mapped and aligned. The reference genome database 122 may include a data store that stores a plurality of I 7QC 70 / 7707 / 3 / YILI different reference genomes. In some implementations, the particular reference genome 122 selected from the reference genome database may be based on the DNA sample type 105. In some implementations, the reference genome type 122 selected from the reference genome database 120 may be selected based on a user's input to the nucleic acid sequencer 110. In such implementations, the user may, for example, select a reference genome identifier 120 that can be used, by the nucleic acid sequencer 110, to select a particular reference genome 122 from the reference genome database 120. The reference genome 122 may include, for example, a nucleic acid sequence assembled as a representative example of a gene set for a particular species. The combination of the read stack 112 generated by the nucleic acid sequencer 110 and the obtained reference genome 122 can be provided as inputs to the hardware-accelerated graph generation unit 130. These inputs can be processed by one or more of the hardware logic units 131 to 138 of the hardware-accelerated graph generation unit 130 to generate an instance of a K-mer graph. For example, each hardware logic unit of the hardware logic units 131 to 138 can be configured to perform its respective operations for each read from the read stack 112 included as an input for the hardware-accelerated graph generation unit 130. The system 100 in Figure 1 is described herein as including a nucleic acid sequencer. In some implementations, such as the one described with reference to Figure 1, the system may include a sequencer 110, and the hardware-accelerated sequence generation unit 130 and other components of system 100 may be integrated within the nucleic acid sequencer 110. However, this description is not limited to integration within the nucleic acid sequencer 110. Instead, in some implementations, the hardware-accelerated sequence generation unit 130 may be implemented in a programmable logic device or ASIC embedded within or housed within a computer that is remote from the nucleic acid sequencer 110 and communicatively coupled to the nucleic acid sequencer 110, such as by using one or more wired or wireless networks.Similarly, the database 120, the variant call unit 180, or both, can be implemented outside of the nucleic acid sequencer. 110. Likewise, the system 100 does not necessarily have to include a nucleic acid sequencer at all. 110 Instead, in some. In implementations of 7QC7n / 77n7 / Σ1 / YILI, the hardware-accelerated read generation unit 130 and other components of Figure 1 can be implemented in a computer system that does not include a nucleic acid sequencer 110. In such implementations, the hardware-accelerated read generation unit can obtain the read stack 112, the reference sequence 122, or both, over a network, from one or more storage locations of one or more memory devices, or the like. Accordingly, system 100 represents an example of the present description but does not limit the present description to any particular configuration of the system components. The one or more hardware logic units of the hardware-accelerated grate generation unit 130 may include an input unit 131, a grate node unit 132, a grate edge unit 133, a backpropagation unit 134, a cycle unit 135, a trimming unit 136, the grate output unit 137, and an erase unit 138. In some implementations, the generation of a K-number grate by the hardware-accelerated grate generation unit 130 may include the control machine 140, which activates and configures each of the hardware logic units 131 to 138 to execute their respective hardware logic operations on a set of data stored in the cache 150 or DRAM 160.In other implementations, the control machine 140 can only activate and configure a subset of the hardware logic units 131 to 138 to execute their respective hardware logic operation on a set of raw data stored in the cache 150 or DRAM 160. As an example, in some implementations, the hardware-accelerated graph generation unit 130 can be used to generate a specialized form of a De Bruijn graph. This specialized form of a De Bruijn graph can be optimized so that non-unique K-numbers are represented using multiple respective nodes in a graph, each with a single edge, rather than being represented by a single node with multiple edges. This can be achieved, in part, by using a graph unit 132 to identify non-unique K-numbers and mark them for further processing. A non-unique K-number can be defined as a sequence of K-numbers that appears at least twice in any single read or at least twice in the reference sequence. A unique K-number does not appear more than once in the same read, but may still appear in multiple reads. I 7QC7n / 77n7 / 3 / YILI However, in other implementations, De Bruijn graphs can be generated without distinguishing between unique and non-unique K-numbers. Therefore, in some implementations, it is not necessary to implement a graph node unit 132 that can identify non-unique K-numbers. In yet another example, backpropagation unit 134 does not need to be used to generate all K-number graphs. Instead, backpropagation unit 134 can be limited to implementations where it can improve performance. For example, backpropagation unit 134 can be used to improve the quality of edge weights when a generated K-number graph is to be transformed into a sequence graph. The K-number graph generation example described with respect to the example in Figure 1 shows each hardware logic unit 131 to 138 that can be activated and configured by the control machine 140. In this description, although each hardware logic unit 131 to 138 is generally described as being activated by the control machine 140, configured by the control machine 140 using, for example, graph description data stored by the control machine 140, obtaining raw graph data, performing one or more specific processing operations on the obtained raw graph data or other data, and then updating the raw graph data, the graph description data, or both, this description is not limited to such implementations.In contrast, in some implementations, each hardware logic unit 131 to 138 can be configured to run multiple instances of its respective functionality. For example, the graph node unit 132 can be configured to accept up to 3 distinct sets of raw graph data and simultaneously perform operations on them, the cycle hardware logic unit 135 can be configured to accept up to 3 distinct sets of raw graph data and simultaneously perform operations on them, and the PRU can be configured to accept up to 2 distinct sets of raw graph data and simultaneously perform operations on them.The number of distinct sets or sets of raw graph data, where each corresponds to different K-number graphs, that can be received and processed by a particular hardware logic unit 131 to 138 is limited only by the hardware resources available to a system 100. For example, assuming sufficient levels of DRAM and FPGA logic units are available for use, the number of distinct sets of raw graph data that can be received and processed simultaneously by a hardware logic unit 131 to 138 may be greater than 3. Similarly, one or more hardware logic units 131 to 138 may be configured to receive and process fewer sets simultaneously. I 7QC7n / 77n7 / 3 / YILI quantities of raw graph datasets if such resources are not readily available or if a particular hardware logic unit is not expected to be used intensively. In still other implementations, there is no requirement that there be only one instance of each hardware logical unit 131 to 138, where each is capable of simultaneously processing distinct sets of raw graph data, each corresponding to different K-number graphs. Instead, in some implementations, multiple instances of each hardware-accelerated graph generation unit 130 can be configured to include multiple instances of one or more of the hardware logical units 131 to 138. In such cases, the control machine 140 can be configured to monitor the status and availability of each hardware logical unit 131 to 138, and then activate and configure each hardware logical unit in such a way that the load balances the processing operations across each respective hardware logical unit.For example, in some implementations, a hardware-accelerated graph generation unit 130 can be configured to have 3 instances of a graph node unit 132, where each is capable of receiving up to 3 distinct sets of raw graph data and simultaneously performing its operations on them, 3 instances of a graph edge unit 133, where each is capable of receiving up to 3 distinct sets of raw graph data and simultaneously performing its operations on them, 2 backpropagation units 134, where each is capable of receiving up to 2 distinct sets of raw graph data and simultaneously performing its operations on them, and 3 cycle units 135, where each is capable of receiving up to 2 distinct sets of raw graph data and simultaneously performing its operations on them.The activation / deactivation of each hardware logic unit, configuration of each hardware logic unit, inputs for each hardware logic unit, outputs for each hardware logic unit, and the updating of graph description data by each hardware logic unit are managed and directed by the control machine 140. Input unit The input unit 131 can receive input data that includes the slow stack generated from reads 112 and the selected reference genome 122, which may be referred to herein as raw graph data. The selected reference genome 122 may include a portion of a reference genome. The raw graph data may include, for example, Lzoczn / zznz / q / YiAi example, data that is processed by one or more hardware logical units 131 to 138 during the generation of an instance of a K-number graph. While raw graph data includes, for example, generated reads 112 and the selected reference genome 122, raw graph data may also include, for example, K-number nodes generated by the graph node unit 132, edges generated by the graph edge unit 133, and the like. The input unit 131 can format the generated reads 112 and the obtained reference genome 122 for storage in DRAM 160. Formatting the generated reads can include, for example, encoding the reads for storage in DRAM 160. In some implementations, encoding the read can include encoding each base call corresponding to a nucleotide of the read into 4-bit values.For example, an A might be encoded as 0000, a C as 0001, a G as 0010, a T as 0011, and an N as 0100, where N is an unknown base call. In some implementations, the encoded data might also include data representing a MAPQ score, a read number, a sequence length, SAM flags, the base or nucleotide calls of the read, one or more read quality indicators other than the MAPQ score, or any combination thereof. The encoded read data can range from a 16-bit value to a 64-bit value, or more, describing the read. Input unit 131 can write the generated reads to DRAM 160. The control machine 140 can detect the receipt of raw input data, activate input unit 131, and initiate the graph description data corresponding to an instance of a K-number graph to be generated based on the raw input data. Activating input unit 131 may involve the control machine 140 sending one or more control messages to input unit 131 that instruct input unit 131 to perform system-defined operations on the raw data provided as input to input unit 131. In some implementations, activating a hardware logic unit such as input unit 131 may also involve the control machine providing the hardware logic unit with graph description data that can be used to configure the hardware logic unit to perform its operations.Configuring the hardware logic unit might include, for example, providing pointers to cache storage locations that store K-numbers, K-number nodes, providing information describing the length of K-numbers, or the like, on which the hardware logic unit needs to operate. I 70070 / 7707 / 3 / YILI Initializing graph description data may include, for example, the control machine 140 generating a K-number graph identifier for the raw input data, generating a graph state information data structure, or a combination thereof. The K-number graph identifier comprises a data string of one or more characters, one or more numbers, or a combination thereof, which can be used to identify an instance of a K-number graph throughout the K-number graph generation process, from the time the raw graph data is received by input unit 131 until at least the time the erase unit 138 is used to remove data related to the K-number graph identifier from cache 150, DRAM 160, or both, after the complete K-number graph generation for a particular set of raw graph data.In some implementations, the K-number graph identifier may include a number, such as a 6-bit number with a value between 0 and 63. In some implementations, the K-number graph identifier may even be used to refer to the K-number graph after the erase unit is used to remove the aforementioned data from cache 150, DRAM 160, or both. The graph state data structure is a data structure that has one or more fields that store data describing the current state of an instance of a K-number graph that will be generated for a particular set of raw input data.State information may include, for example, data indicating a last logical hardware unit that operated on raw graph data for a particular instance of a K-number graph, data indicating whether the last logical hardware unit aborted the operation, data indicating a length of K-numbers, data indicating a list of K-number nodes, data indicating a list of pointers that can be used to identify K-number nodes in a cache, a length of a list of K-number nodes, data indicating a list of non-unique K-numbers, data indicating the locations of raw input data in the cache or DRAM, data indicating a base address in DRAM for nodes of the instance of a K-number graph, or any subset or combination of these. Input unit 131 can format input reads 112 and reference genome 122 and write input reads 112 and reference genome to DRAM 160. Control machine 140 can detect when input unit 131 has completed formatting and writing input reads 112 and reference genome 122 to DRAM 160. After control machine 140 detects the completion of formatting and writing input reads 112 and reference genome 122 to DRAM, control machine 140 can update the graph status information to indicate that input unit 131 has completed I 7QC7n / 77n7 / =l / YILI its operations on the first raw graph data for a first instance of a K-number graph. Once the initial raw graph data has been entered, formatted, and stored in DRAM 160, the control machine 140 can determine the next logical hardware unit to be activated and configured. For example, the control machine 140 can activate and configure a graph node unit 132 to generate K-mer nodes based on the reference genome portion 122 and the stack of formatted reads stored in DRAM 160. Graph node unit The control machine 140 can activate and configure the graph node unit 132 to continue generating the first instance of a K-number graph by processing the formatted reads 112 and the reference genome 122 stored in DRAM. This can include, for example, sending a control signal to the graph node unit 132, providing graph description data to the graph node unit 132, or a combination of these. The graph description data can be used to configure the graph node unit 132 for operation. For example, by providing the graph description data to the graph node unit 132, the control machine 140 can configure the graph node unit 132 to identify the particular size K-numbers defined by the graph description data.Other fields in the graph description data described herein can be used to configure a logical hardware unit, such as graph node unit 132, in a similar manner. Furthermore, in a substantially parallel manner, control machine 140 can detect that input unit 131 has received the second raw graph data as input. Control machine 140 can then activate input unit 131, instruct it to format the reads and reference genome of the second raw graph data, and generate the second graph description data for a second instance of a K-number graph to be generated based on the second raw graph data. Thus, control machine 140 can achieve high levels of productivity by simultaneously managing the throughput of different hardware logic units 131 and 132 performing K-number graph generation processes on different sets of raw graph data at different processing stages. Control machine 140 is I 7QC 70 / 7707 / 3 / YILI is configured to manage this parallel functionality through each of the hardware logic units 131, 132, 133, 134, 135, 136, 137, 138 in such a way that, at any particular point in time, there can be up to eight hardware logic units operating on eight different sets of raw graph data, where the hardware-accelerated graph generation unit 130 works to generate eight different K-number graphs simultaneously. The control machine 140, using the graph description data, manages this entire process by activating and configuring each respective hardware logic unit to achieve high-level segmented functionality of the non-segmented hardware logic units 131, 132, 133, 134, 135, 136, 137, 138, which do not have a direct, physical input / output connection between each respective hardware logic unit.Although an example of eight simultaneous K-number graph generations being performed at the same time is illustrated, the present description can be configured to achieve many more simultaneous K-number graph generations such as by implementing multiple hardware-accelerated graph generation units 130 at once, multiple instances of multiple hardware logic units in one or more hardware-accelerated graph generation units 130, or a combination of these. The graph node unit 132 can analyze each read from a read algorithm 112 to identify each of the K-numbers in the read. This might include, for example, sliding a K-number access window along each position of each read to identify each particular K-number of the respective read. The graph node unit 132 can store data representing a node in a K-number graph, for each identified K-number from each read, in cache 150. Similarly, the graph node unit 132 can also generate and store, in DRAM, a list of node pointers in a node pointer data structure, where each node pointer points to a location in the cache of a K-number node. The graph node unit 132 can also generate and store information indicating the location and length of the node pointer list for each K-number graph in the graph description data maintained by the control machine.These pointers can be used as graph state information by the control machine 140 to configure another logical hardware unit during a later portion of the K-number graph generation process. Cache 150 can use one or more cache consistency policies, such as a Least Recently Used (LRU) cache consistency policy, which is configured to discard the oldest objects from cache 150, where the oldest objects are determined based on when the object was written to cache 150. LZQCZn / ZZnZ / q / YIAI An example of data generated by the grato node unit 132 and stored in cache 150, DRAM 160, or both, is shown with reference to Figure 4. Figure 4 shows a portion of a reference genome 410, a read 420, and a De Bruijn grato 400 is provided. The De Bruijn grato 400, which is described in more detail below, includes a node for each K-mer in the genome portion 410 and read 420 and an edge between each pair of K-mer nodes that links a pair of nodes having overlapping k-1 nucleotides. With reference to the example in Figure 4, the 132 grad node unit can generate, based on the reception of a portion of the reference genome 410 and read 420, data representing nodes 431, 432, 433, 434, 435, 436, 437, 438 of a first path 430 of the De Bruijn grad 400, and nodes 441, 442, 443, 444 can be generated. First, the 132 grad node unit can align the overlapping portions of the reference genome 410a, 410b and the overlapping portions of read 420a, 420b to identify the overlapping regions as shown in Figure 4. The 132 grad node unit can identify each of the K-mers of the portion of the genome 410 and the reading 420.This can be achieved by using an access window of length k, which is equal to 4 in this example, at a first position of the reference genome portion 410 that captures the K-mer identified by the access window, generating data representing a grate node that includes the captured K-mer, storing the data representing the node in cache 150, advancing the access window by one nucleotide, and then iteratively repeating this process. In this example, the grate node unit 132 can identify the K-mers ATCG, TCGC, CGCC, GCCT, CCTA, CTAG, TAGA, and AGAA for the reference genome portion 410 and generate a respective node 431, 432, 433, 434, 435, 436, 437, 438, where one of these nodes corresponds to a respective K-mer. Each node is created to have a length k, which is 4 in this example, and has an overlapping number of k-1 k-numbers with the next adjacent node.The grato node unit 132 can store the generated nodes in cache 150, DRAM 160, or both. In some implementations, the cache may include a hash table cache. In such implementations, the nodes may be stored as keys in a hash table. The grato node unit 132 can store data describing the pointers to the K-number node locations in the grato description data maintained by the control machine 140. Graph node unit 132 can perform the same operations for read 420. With respect to read 420, graph node unit 132 can identify the K-mers ATCG, TCGC, CGCG, GCGT, CGTA, GTAG, TAGA, and AGAA. This can be similarly achieved by using an access window of length k, which is equal to 4 in this example, at a first position in the portion of read 420 that captures the K-mer identified by the access window, and then advancing the access window and repeating the process. Graph node unit 132 can begin by identifying each K-mer for genome portion 410. In some implementations, graph node unit 132 can generate a corresponding node for each of the K-mers. In other implementations, the graph node unit 132 can only generate 431, 432, 433, 434, 435, 436, 437, 438 corresponding to the identified K-meres that differ from the K-mer nodes of the reference genome portion410.In each scenario, each node is created to have a length k, which is 4 in this example, and has an overlap of k-1 K-mers with the next adjacent node. This can continue until a node is created for each K-mer of the reference genome portion 410 and stored in cache 150 or DRAM 160. In some implementations, graph node unit 132 can also be configured to identify non-unique K-mers. In such implementations, graph node unit 132 can, for each particular read from a first read stack, determine whether an identified K-mer is a unique K-mer or a non-unique K-mer. If graph node unit 132 determines that a particular K-mer is a unique K-mer, then graph node unit 132 can advance the access window by a single nucleotide to evaluate the next K-mer. Alternatively, if graph node unit 132 determines that a particular K-mer is not a unique K-mer, then graph node unit 132 can store data indicating that the particular K-mer is not a unique K-mer.For example, graph node unit 132 can store a data indicator in the graph data description maintained by the control machine for a particular instance of a K-number graph, indicating that the K-number is a non-unique K-number. However, such data can be stored by any other component of the hardware-accelerated graph generation unit 130, stored in any other memory unit of the hardware-accelerated graph generation unit 130, or a combination of these. Subsequent hardware logic units can then perform operations that address the non-unique K-number to reduce or eliminate cycles in the K-number graph instance. At this point in the process, the graph node unit 132 stores the data representing a node in a K-number graph for each K-number in cache 150, DRAM 160, or both. That is, I 70070 / 7707 / 3 / YILI The hardware-accelerated grate generation unit 130 has not yet generated the grate edges 431a, 432a, 433a, 434a, 435a, 436a, 437a, 432b, 441a, 442a, 443a, 444a, the grate edge weights, or the like. These features of this instance of a K-number grate can be generated by one or more hardware logic units of the hardware-accelerated grate generation unit 130. The control machine 140 can monitor the operation of the grato node unit 132. Once the grato node unit 132 generates data representing a K-number node for each K-number of each read of the first raw grato data for this first instance of a K-number grato, the control machine can update the grato description data to indicate that the hardware-accelerated grato generation unit 130 has completed the operations of the grato node unit 132 on the first raw grato data. Furthermore, the control machine 140 can also store grato description data that includes, for example, a flag identifying each of the non-unique K-numbers, storage locations for the K-number grato nodes, and data indicating that the grato node unit 132 has completed its operations. Once the K-number grato nodes have been generated and stored in cache 150, the control machine 140 can determine the next logical hardware unit to be activated and configured. For example, the control machine 140 can activate and configure a grato edge unit 133 to generate, weight, or both, the grato edges between pairs of nodes. Pleasant edge unit The grato edge unit 133 can generate grato edges between pairs of K-number nodes generated by the grato node unit 132. The control machine 140 can activate the grato edge unit 133 once it is determined that the grato node unit 132 for the first instance of a K-number grato is complete and that the grato edge unit 133 is available. In some implementations, the control machine 140 can provide, or otherwise make accessible, the locations of the grato edge unit 133 that store the K-number grato nodes generated by the grato node unit 132 for a particular instance of a K-number grato. For example, control machine 140 can access the K-number graph description data generated and stored by graph node unit 132 during K-number node generation for an instance of a K-number graph.The accessed K-number graph description data can indicate, or otherwise describe, a list of K-numbers. Once the graph edge unit 133 has obtained the location of the K-numbers graph nodes for the first instance of the K-numbers graph, the graph edge unit can begin generating one or more graph edges between the data representing the K-numbers nodes. In some implementations, the graph edge unit 133 can access the data representing the graph node for each of the K-numbers from the particular hash table cache read. The graph edge unit 133 can generate, for storage in the hash table cache, data representing a graph edge between the graph nodes for the K-numbers. For example, the data representing the graph edge can be stored in the hash table cache as part of a graph node record for the edge's source node.In some implementations, the graph edge unit 133 can assign an edge weight to each edge in the K-number graph. For example, the graph edge unit 133 can add a weight of +1, or some other weight, for each occurrence of the graph edge that links the respective pair of K-numbers. As an example, graph edge unit 133 can identify adjacent nodes using data descriptions obtained from control machine 140. Nodes can be determined to be adjacent based on a variety of factors, including whether the nodes share k-1 overlapping nucleotides and whether the nodes are observed at two consecutive positions in the K-mer sliding access window. For instance, graph edge unit 133 can slide a K-mer access window along each position of every read of the raw graph data. In some implementations, graph edge unit 133 can create an edge or increment an edge weight after determining that two successive K-mers, overlapping by all but one base, are observed in a read.Graph node 133 creates an edge or increases an edge weight in such a situation because this situation involves an edge between the graph nodes that correspond to those two successive K-numbers. In some implementations, the data representing graph nodes can be stored as hash keys in a hash table. In such implementations, the graph edge unit 133 can generate an edge from a first node (or hash key) to I 7QC7n / 77n7 / 3 / YILI a second node (or hashing key) by accessing a hashing location to which the first node (or hashing key) is mapped and generating a pointer to be stored at the hashing location that points to the second node (or hashing key). Subsequent edges can be generated in the same way, which creates a path 430 or 440 through a graph that can be traversed using one or more graph traversal algorithms. With reference to the example in Figure 4, the graph edge unit 133 can generate data representing one or more edges between the pairs of nodes 431, 432, 433, 434, 435, 436, 437, 438 of a first path 430, the pairs of nodes 441, 442, 443, 444 of a second path 440, or one or more pairs of nodes in the first path 430 and the second path 440. An example of such edges is shown in Figure 4 as the edges 431a, 432a, 433a, 434a, 435a, 436a, 437a, 432b, 441a, 442a, 443a, 444a. In this example, the nucleotide sequence 420 is referred to as a read. However, in some implementations, the removal of the low-quality base can occur in such a way that the nucleotide sequence 420 is a contigo, or a portion of a read. In such implementations, graph edges linking a pair of nodes will create only a path of one or more links within a particular contigo, and not from the K-mer node of a first contigo to a K-mer node of a second configuration. A contigo can include a nucleotide sequence that results from the removal of a low-quality base. Backpropagation unit The hardware-accelerated graph generation unit 130 may include a backpropagation unit 134. However, the control machine 140 can only activate and configure the backpropagation unit 134 in certain implementations. For example, the control machine 140 can activate the backpropagation unit 134 when it determines that the K-number graph generated from a current set of raw graph data will later be transformed into a sequence graph. When activated, the backpropagation module 134 can receive graph description data from the control machine 140. In such an implementation, adjustments to the graph edge weights can make the weights more reliable when inherited by the sequence graph. LZQCZn / ZZnZ / q / YIAI The edge unit 133 can create edges between K-number contiguous nodes by identifying the corresponding K-number nodes in cache 150, generating an edge that links the K-number nodes, and then incrementing the edge weight by +1 for each occurrence. In some implementations, after the K-number nodes of a contiguous node have been added to a K-number contiguous node and weighted using edge unit 134, backpropagation unit 134 can be used to backpropagate weight increments of +1 at k-1 steps through a linear chain of edges from the contiguous node to the "left" of the initial contiguous node in the K-number contiguous node. In this context, the “left” of the initial node of the graph in the K-numbers graph is a direction in a K-numbers graph that is opposite the directed edge of the K-numbers graph. To illustrate this concept, to the “left” of a node 434 in the De Bruijn graph 400 would be nodes 433, 432, and 431. For example, in some implementations, the backpropagation module 134 can access, in the hash table cache, data representing the K-numbers of a gram, including its K-numbers of nodes, corresponding gram edges, and so on, and then adjust the edge weights for K-1 nodes that occur before a new configuration begins. The backpropagation unit 135 can locate the appropriate K-numbers of edges and nodes for adjustment based on gram description data received from the control machine 160, which includes pointers to locations in the cache that store this information. The gram description data can be updated with any changes that occur during backpropagation. In some implementations, it can be advantageous to perform the backpropagation described above because a base N contiguity can only increment a certain number of edge weights (NK). However, if this K-number graph is subsequently transformed into a sequence graph with internal (N-1) edges corresponding to this contiguity, then the first (K-1) edge weights will not inherit the appropriately incremented edge weights. The backpropagation described above addresses most instances of this problem. Therefore, backpropagation can be used to address this issue in order to increase the reliability of inherited edge weights when the K-number graph is transformed into a sequence graph. Cycle unit The hardware-accelerated gratos generation unit 130 may include a cycle unit 135. In some implementations, the cycle unit 135 can be enabled and configured. I 7QC7n / 77n7 / 3 / YILI using the control machine 140 to detect cycles in an instance of a K-numbered grate. For example, the cycle unit 135 can evaluate the K-numbered nodes and K-numbered edges of the raw grate data of an instance of the K-numbered grate that has been generated by one or more of the input unit 131, the grate node unit 132, the grate edge unit 133, and the backpropagation unit 134. The cycle unit 135 is configured to receive grate description data from the control machine 140. The cycle unit 135 can iteratively mark principal nodes for deletion. A principal node can include a node that does not include any incoming edges. Where an incoming edge is an edge that points from a first node to the first node itself.After each principal node is marked for deletion, cycle unit 135 can determine if any nodes pointed to by outgoing edges have become principal nodes as a result of a node's deletion. If such nodes are found, they are marked for deletion. An outgoing edge is a grid edge that points from a first node to another node. Cycle unit 135 can continue performing this process until no principal nodes remain. After determining that no principal nodes remain, cycle unit 135 can determine whether the grate is empty, where an empty grate means that all nodes in the grate are marked for deletion. If the grate with no principal nodes is empty, then there was no cycle. Alternatively, if the grate with no principal nodes is not empty, then the grate must contain a cycle. Cycles are resistant to this type of deletion because no node in a cycle becomes a principal node through the deletion of nodes outside the cycle. After making any of these determinations, cycle unit 135 can provide an indication to control machine 140 as to whether or not a cycle was detected. Then, based on the indication provided by cycle unit 135, control machine 140 can determine which hardware logic unit should be activated and configured next. For example, if a cycle was detected and the generation of the K-numbers gram instance must be aborted, then control machine 140 can activate and configure erase unit 138. In such cases, the erase unit can remove raw gram data from cache 150 and the corresponding DRAM 160 for the aborted K-numbers gram instance.If, alternatively, the generation of the K-numbers grate instance is to continue, then the control machine 140 can activate and configure another hardware logic unit to perform further operations on the raw grate data to generate the K-numbers grate instance. For example, if the generation of the K-numbers grate were to continue, the machine... LZQCZn / ZZnZ / q / YIAI control 140 can activate and configure the clipping unit 136 or the graph output unit 137. Although cycle unit 135 can be used by the hardware-accelerated graph generation unit 130 to detect cycles in a generated K-number graph instance, cycle unit 135 can be selectively activated and configured as backpropagation unit 134. This is because some types of K-number graphs are expected to contain cycles. However, for certain types of K-number graphs, it may be beneficial to avoid graphs with cycles. Consequently, the hardware-accelerated graph generation unit 130 can be configured, for example, to provide the control machine 140 with input indicating whether cycle unit 135 should be performed for a particular graph instance. Cutting unit The hardware-accelerated graph generation unit 130 may include a trimming unit 135. Like the backpropagation unit 134 and the cycling unit 135, the trimming unit 136 may be selectively activated and configured by the control machine 160. If activated and configured, the trimming unit 135 may evaluate a weight for each graph edge on raw graph data for an instance of a K-number graph that has been generated up to this point by the hardware-accelerated graph generation unit 130. In some implementations, if the trimming unit 136 determines that the weight value of a graph edge does not satisfy a predetermined threshold, then the trimming unit 136 may remove the edge from the graph and any K-number nodes that occur after the identified graph edges.Alternatively, if clipping unit 136 determines that the weight value of a graph edge satisfies the predetermined threshold, then clipping unit 136 will leave the graph edge intact. In other implementations, the clipping unit 136 can identify linear chains, where linear chains are maximum paths through the graph, in which every internal node between the start node and the end node has exactly one incoming edge and one outgoing edge. In such implementations, the clipping unit 136 can determine whether each internal edge of the linear chain fails to satisfy the clipping threshold. If such a situation occurs, the clipping unit 136 can remove the entire linear chain, including all internal edges and all internal nodes. I 7QC7n / 77n7 / 3 / YILI except the initial node and / or final node of the chain, which the trimming unit 136 can retain if they have any non-internal edges. Graph output unit The hardware-accelerated graph generation unit 130 may include a graph output unit 137. The graph output unit 137 can be used, via the hardware-accelerated graph generation unit 130, to generate a final version 170 of the K-meros graph instance, since the K-meros graph instance is described by graph description data and cached data. For example, the graph output unit 137 can obtain data representing the K-meros graph from the hash table cache 150 using graph description data that includes, for example, pointers to locations in the hash table cache that stores the K-meros graph data. Next, the graph output unit 137 can provide the data obtained from the hash table cache 150 describing the final version of the K-number graph 170 to a variant call unit 180.The variant call unit 180 can perform a variant call analysis on the final version of the K-mer graph 170 to produce a variant set 190. A variant set 190 can include one or more candidate variants. A variant is an alteration in an organism's genomic data. A candidate variant is a determination made by a variant call unit that is inferred by the variant call unit based on the processing of a K-mer graph 170. In some implementations, the candidate variant may have a threshold level of error in the variant determination. The variant call unit 180 can identify candidate variants by processing the K-mer graph 180. In some implementations, for example, the variant call unit 180 can identify a candidate variant when a base or nucleotide designation of one or more reads in the read configuration and a nucleotide in a reference genome at a particular location in the reference genome are different. Data describing the variant set 190 can be generated or determined in various ways. For example, in some implementations, variant call operations can be performed as described in more detail in, for example, U.S. Publication No. 2016 / 0180019; U.S. Publication No. 2016 / 0306922; and U.S. Publication No. 2019-0259468, the content of which is incorporated into this description as a reference in its entirety.Data describing the set of variants 190 can be provided for the. LZQCZn / ZZnZ / q / YIAI output in several different ways. For example, data describing the variant set 190 can be displayed on a nucleic acid sequencer screen 110, displayed on a different computer screen, audibly output through one or more speakers of a computer device, output through a printer, or any combination of these. Erasing unit The erase unit 138 can be used to perform memory reclamation tasks after a K-series grato instance has been completed and issued by the hardware-accelerated grato generation unit 130. For example, the erase unit can delete all raw grato data related to the particular K-series grato instance that the hardware-accelerated grato generation unit 130 has completed and produced. Alternatively, or in addition, the erase unit 138 can delete all data related to the particular instance of the graphics unit 130 that is stored by the control machine, delete all data related to the particular instance of the graphics unit 130 that is stored in DRAM 160, or similar.Therefore, erase unit 138 can selectively delete the data representing the grato nodes and the data representing the grato edges of the K-number grato from the hash table cache. Such deletion is selective because only a portion of the cache contents, control machine, or DRAM needs to be removed. Furthermore, when the grato is stored as a hash table, the hash table may often be sparsely populated, and it is faster for erase unit 138 to selectively erase only the occupied hash table entries than to erase the entire hash table, resulting in performance improvements. However, in some implementations, non-hashtable data associated with a gram does not need to be deleted item by item. In such implementations, the deletion unit 138 can either set the list lengths to zero in the gram description data or the allocated memory space can simply be freed for reuse without deleting the contents. Figure 2 is a flowchart of an example of a 200 process for the hardware-accelerated generation of a K-mer array. Generally, the 200 process may include obtaining a first set of nucleic acid sequences, where the first set of I 7QC 70 / 7707 / 3 / YILI nucleic acid sequences include (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence, (210), generating a K-mer graph using the first set of nucleic acid sequences obtained and using a plurality of non-segmented hardware logic units of a programmable logic device, wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-mer graph represents a K-mer, each graph edge represents a link between a pair of K-mers, and each weight of each edge of the K-mer graph represents a number of occurrences of a K-mer sequence represented by a pair of K-mers (220), during the generation of the K-mer graph: periodically updating, with a control machine,the description of the K-number graph data after the performance of one or more operations by each hardware logic unit used to generate at least a portion of the K-number graph, wherein the graph description data represents (i) a K-number graph identifier and (ii) K-number graph state information, wherein the control machine creates a workflow of operations using the non-segmented hardware logic units by triggering the performance of one or more operations by each respective hardware logic unit during the generation of the K-number graph (230), and provides the K-number graph to a variant call module, wherein the variant call unit processes the K-number graph to determine one or more candidate variants from one or more of the plurality of reads and the reference sequence. Figure 3 is a flowchart of another example of a 300 process for the hardware-accelerated generation of a K-mer graph. Generally, the 300 process may include obtaining a first set of nucleic acid sequences, wherein the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence (310), for each particular nucleic acid sequence in the first set of nucleic acid sequences; generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a graph node for each K-mer of the particular nucleic acid sequence (320); detecting, by means of a control machine, that the first hardware logic unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence (330);configure, using the control machine, a second hardware logic unit to perform graph edge generation for the generated graph nodes (340) and, for one, LZQCZn / ZZnZ / q / YIAI or more pairs of the generated graph nodes: generate, by means of the second logical hardware unit and for storage in the graph hash table, data representing graph edges between one or more pairs of the generated graph nodes generated by the first logical hardware unit, wherein the data representing the graph node for each K-mer stored in the hash table cache and the data representing graph edges stored in the hash table cache represent a K-mer graph of the first set of nucleic acid sequences (350). Figure 4 is an example of a K-mer 400 graph. In this example, the K-mer 400 graph is generated based on at least a portion of a reference genome 410 and a read 420. In this example, the K-mer 400 graph is a De Bruijn graph. The K-mer graph 400 is generated using a plurality of nodes and one or more edges between pairs of nodes. Each node represents a K-mer of length k, where in this example k=4. Each edge provides an indication that there is an overlap of k-1 nucleotides of the K-mers joined by the edge. In the K-mer graph 400, path 430 includes a plurality of nodes and edges representing each K-mer of the reference sequence portion 410. Path 440 then includes a plurality of nodes and edges representing portions of a read 420 that differ from the reference genome portion 410. Figure 5 is a block diagram of an example of system 500 components that can be used for the hardware-accelerated K-number graph. The Computing Device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The Computing Device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, cell phones, smartphones, and other similar computing devices. Additionally, the Computing Device 500 or 550 may include Universal Serial Bus (USB) flash drives. USB flash drives can store operating systems and other applications. USB flash drives may include input / output components, such as a wireless transmitter or USB connector that can be inserted into a USB port on another computing device. The components shown here, their The connections and relationships, and their functions, are intended to be examples only, and are not intended to limit the implementations of the inventions described and / or claimed in this document. The computing device 500 includes a processor 502, memory 504, a storage device 506, a high-speed interface 508 connected to the memory 504 and high-speed expansion ports 510, and a low-speed interface 512 connected to the low-speed bus 514 and storage device 506. Each of the components 502, 504, 506, 508, 510, and 512 is interconnected using various buses and can be mounted on a common motherboard or in other ways as appropriate. The processor 502 can process instructions for execution on the computing device 500, including instructions stored in memory 504 or storage device 506 to display graphical information for a GUI on an external input / output device, such as the display 516 connected to the high-speed interface 508.In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and memory types. Furthermore, multiple 500 computing devices can be connected, with each device providing portions of the necessary operations, for example, as a server bank, a group of blade servers, or a multi-processor system. Memory 504 stores information on the computer device 500. In one implementation, memory 504 is a volatile memory unit or units. In another implementation, memory 504 is a non-volatile memory unit or units. Memory 504 can also be another form of computer-readable media, such as a magnetic or optical disk. The 506 storage device is capable of providing mass storage for the 500 computing device. In one implementation, the 506 storage device may be or contain a computer-readable medium, such as a floppy disk drive, hard disk drive, optical disk drive, or tape drive, flash memory or other similar solid-state memory device, or a collection of devices, including devices in a storage area network or other configurations. A computer program product may be tangibly embodied in a data carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The data carrier is a computer-readable or machine-readable medium, such as I 7QC7n / 77n7 / 3 / YILI as memory 504, storage device 506, or memory in processor 502. The high-speed controller 508 handles bandwidth-intensive operations for the computing device 500, while the low-speed controller 512 handles lower-bandwidth-intensive operations. This role assignment is just one example. In one implementation, the high-speed controller 508 is coupled to memory 504, display 516, for example, via a graph processor or accelerator, and to high-speed expansion ports 510, which can accommodate various expansion cards (not shown). In the implementation, the low-speed controller 512 is coupled to storage device 506 and low-speed expansion port 514.The low-speed expansion port, which can include various communication ports, such as USB, Bluetooth, Ethernet, or wireless Ethernet, can be connected to one or more input / output devices, such as a keyboard, a pointing device, a microphone / speaker pair, a scanner, or a network device such as a switch or router, for example, via a network adapter. The Computing Device 500 can be implemented in a number of different ways, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. In addition, it can be implemented in a personal computer, such as a laptop computer 522. Alternatively, the components of the Computing Device 500 can be combined with other components in a mobile device (not shown), such as the Device 550.Each of such devices may contain one or more of a 500, 550 computing device, and an entire system may be composed of multiple 500, 550 computing devices that communicate with each other. The Computing Device 500 can be implemented in a number of different ways, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. In addition, it can be implemented in a personal computer, such as a laptop computer 522. Alternatively, the components of the Computing Device 500 can be combined with other components in a mobile device (not shown), such as the Device 550. Each of such devices can contain one or more Computing Devices 500, 550, and an entire system can be composed of multiple Computing Devices 500, 550 communicating with each other. LZQCZn / ZZnZ / q / YIAI The 550 computer device includes a 552 processor, 564 memory, and an input / output device, such as a 554 display, a 566 communication interface, and a 568 transceiver, among other components. The 550 device can also be provided with a storage device, such as a micro disk drive or other device, to provide additional storage. Each of the 550, 552, 564, 554, 566, and 568 components is interconnected using various buses, and several of the components can be mounted on a common motherboard or in other ways, as appropriate. The 552 processor can execute instructions on the 550 computing device, including instructions stored in memory 564. The processor can be implemented as a chipset comprising separate and multiple analog and digital processors. Additionally, the processor can be implemented using any of several architectures. For example, the 510 processor can be a CISC (Complex Instruction Set Computer), a RISC (Reduced Instruction Set Computer), or a MISC (Minimal Instruction Set Computer) processor. The processor can provide, for example, coordination for other components of the 550 device, such as the centroid user interfaces, applications run by the 550 device, and wireless communication via the 550 device. The processor 552 can communicate with a user through the control interface 558 and the display interface 556 coupled to a display 554. The display 554 can be, for example, a thin-film transistor (TFT) liquid crystal display, an organic light-emitting diode (OLED) display, or other suitable display technology. The display interface 556 can comprise circuitry suitable for driving the display 554 to present graphical and other information to a user. The control interface 558 can receive commands from a user and convert them for transmission to the processor 552. In addition, an external interface 562 can be provided in communication with a processor 552 to enable near-area communication of the device 550 with other devices.The external interface 562 can provide, for example, wired communication in some implementations, or wireless communication in other implementations, and multiple interfaces can also be used. Memory 564 stores information on the 550 computing device. Memory 564 can be implemented as one or more of a computer-readable medium or drive, or I 7QC 70 / 7707 / 3 / YILI volatile memory units, or one or more non-volatile memory units. Expansion memory 574 can also be provided and connected to the device 550 via expansion interface 572, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such expansion memory 574 can provide additional storage space for the device 550, or it can store applications or other information for the device 550. Specifically, expansion memory 574 can include instructions for performing or completing the processes described above, and it can also include security information. Therefore, for example, expansion memory 574 can be provided as a security module for the device 550 and can be programmed with instructions to ensure the secure use of the device 550.In addition, secure applications can be provided through SIMM cards, along with additional information, such as storing identification information on the SIMM card in a way that cannot be legally extracted. Memory may include, for example, flash memory and / or non-volatile random-access memory (NVRAM), as described later. In an implementation, a computer program product is tangibly embodied in a data carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The data carrier is a computer-readable medium, such as memory 564, expansion memory 574, or memory in the processor 552, which may be received, for example, through the transceiver 568 or external interface 562. The 550 device can communicate wirelessly via the 566 communication interface, which may include a digital signal processing circuit when required. The 566 communication interface can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication can occur, for example, via the 568 radio frequency transceiver. Short-range communication, such as Bluetooth, Wi-Fi, or another transceiver of this type (not shown), can also occur. Furthermore, the 570 GPS (Global Positioning System) receiver module can provide additional wireless navigation and location data to the 550 device, which can be used as appropriate by applications running on the 550 device. I 70070 / 7707 / 3 / YILI The 550 device can also communicate audibly using a 560 audio codec, which can receive spoken information from a user and convert it into usable digital information. The 560 audio codec can likewise generate audible sound for a user, such as through a speaker, for example, on a mobile phone connected to the 550 device. This sound can include audio from voice calls, recorded audio (such as voicemails, music files, etc.), and audio generated by applications running on the 550 device. The 550 computing device can be implemented in a number of different ways, as shown in the figure. For example, it can be implemented as a 580 cell phone. It can also be implemented as part of a 582 smartphone, personal digital assistant, or other similar mobile device. Various implementations of the systems and methods described herein may be carried out in a digital electronic circuit, integrated circuit, specially designed ASIC (application-specific integrated circuit), computer hardware, firmware, software, and / or combinations of such implementations. These various implementations may include implementation in one or more computer programs that are executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer product, apparatus, and / or device, such as magnetic disks, optical disks, memory, and programmable logic devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to I 7QC7n / 77n7 / 3 / YILI any signal used to provide machine instructions and / or data to a programmable processor. To facilitate interaction with a user, the systems and techniques described here can be implemented on a computer that has a display device, for example, a CRT (cathode ray tube) or LCD (liquid crystal display) monitor to show information to the user, and a keyboard and pointing device, for example, a mouse or scroll ball, by which the user can provide input data to the computer. Other types of devices can be used to enable interaction with a user; for example, the feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and user input can be received in any form, including acoustic, voice, or tactile input. The systems and techniques described here can be implemented in a computer system that includes a management component, such as a data server, or a middleware component, such as an application server, or a user interface component, such as a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such management, middleware, or user interface components. The system components can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet. A computer system can include clients and servers. A client and a server are generally located remotely and interact, typically, through a communication network. The relationship between the client and the server arises from computer programs running on their respective computers, which have a client-server relationship with each other. Lzoczn / zznz / q / Yi Other modalities Several embodiments have been described. However, it is understood that various modifications may be made without departing from the spirit and scope of the invention. Furthermore, the logical flows depicted in the figures do not require the particular order shown, or sequential order, to achieve the desired results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A method for hardware-accelerated generation of a K-mer graph using a programmable logic device, the method comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence;generate, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer graph using the first set of nucleic acid sequences obtained, wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-mer graph represents a K-mer, each edge of the K-mer graph represents a link between a pair of K-mers, and each weight of each edge of the K-mer graph represents a number of occurrences of a K-mer sequence represented by a pair of K-mers;and during the generation of the K-number graph: periodically updating, with a control machine, the graph description data for the K-number graph after the performance of one or more operations by each hardware logic unit used to generate at least a portion of the K-number graph, wherein the graph description data represents (i) an identifier of the K-number graph and (ii) state information of the K-number graph, wherein the control machine creates a workflow of operations using the non-segmented hardware logic units by activating the performance of one or more operations of each respective hardware logic unit during the generation of the K-number graph.
2. The method of claim 1, characterized in that the output of each hardware logic unit of the plurality of hardware logic units is stored through a hash table cache.
3. The method of claim 1, characterized in that the control machine is implemented using a hardware logic unit of the programmable logic device. I 7QC7n / 77n7 / 3 / YILI 4. The method of claim 1, characterized in that the control machine is implemented using one or more CPUs or GPUs to execute software instructions to perform the functionality of the control machine.
5. The method of claim 1, the operations further comprise: providing the generated K-number graph to a variant call unit, characterized in that the variant call unit processes the K-number graph to determine candidate variants from one or more of the plurality of reads and the reference sequence.
6. The method of claim 5, characterized in that the software instructions are executed by one or more CPUs or GPUs to perform one or more variant call unit functions.
7. The method of claim 5, characterized in that the programmable logic device is used to accelerate one or more functions of the variant call unit.
8. The method of claim 1, characterized in that the graph description data further includes (ii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer graph or nucleic acid sequences of the stack associated with the K-mer graph identifier.
9. A system for the hardware-accelerated generation of a K-mer graph using a programmable logic device, the system comprising: a hardware-accelerated graph generation unit including hardware digital logic circuits arranged to perform operations; the operations comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence;generate, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer graph using the first set of nucleic acid sequences obtained, wherein each hardware logic unit LZQCZn / ZZnZ / q / YIAI comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-mer graph represents a K-mer, each edge of the K-mer graph represents a link between a pair of K-mers, and each weight of each edge of the K-mer graph represents a number of occurrences of a K-mer sequence represented by a pair of K-mers;and during the generation of the K-numbers grade: periodically updating, with a control machine, the grade description data for the K-numbers grade after the performance of one or more operations by each hardware logic unit used to generate at least a portion of the K-numbers grade, wherein the grade description data represents (i) an identifier of the K-numbers grade and (ii) status information of the K-numbers grade, wherein the control machine creates a workflow of operations using the non-segmented hardware logic units by activating the performance of one or more operations of each respective hardware logic unit during the generation of the K-numbers grade.
10. The system of claim 9, characterized in that the output of each hardware logic unit of the plurality of hardware logic units is stored through a hash table cache.
11. The system of claim 9, characterized in that the operations further include: providing the generated K-numbers to a variant call unit, wherein the variant call unit is configured to process the K-numbers to determine candidate variants from one or more of the plurality of reads and the reference sequence.
12. The system of claim 9, the system further comprising: one or more computers; and one or more memory devices storing instructions which, when executed by the one or more computers, cause the one or more computers to perform second operations of a variant call unit, the second operations comprising obtaining, by means of the variant call unit, the generated K-mer graph; and identifying, based on the processing of the generated K-mer graph by the variant call unit, one or more candidate variants, characterized in that a candidate variant is a difference between a base call of one or more reads in the read stack and a nucleotide of a reference genome at a particular location of the reference genome.
13. The system of claim 9, characterized in that the operations further comprise: obtaining, by means of a variant call unit, the generated K-mer graph; and identifying, based on the processing of the generated K-mer graph by the variant call unit, one or more candidate variants, characterized in that a candidate variant is a difference between a base call of one or more reads in the read stack and a nucleotide of a reference genome at a particular location of the reference genome.
14. The system of claim 9, characterized in that the graph description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer graph or nucleic acid sequences of the stack associated with the K-mer graph identifier.
15. A method for hardware-accelerated generation of a K-mer graph on a programmable logic device, the method comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence; for each particular nucleic acid sequence in the first set of nucleic acid sequences: generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a graph node for each K-mer of the particular nucleic acid sequence; detecting, by means of a control machine, that the first hardware logic unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence;I 7QC7n / 77n7 / 3 / YILI configure, by means of the control machine, a second hardware logic unit to perform graph edge generation for the generated graph nodes; and for one or more pairs of the generated graph nodes: generate, by means of the second hardware logic unit and for storage in the graph hash table, data representing the graph edges between one or more pairs of the generated graph nodes generated by the first hardware logic unit, wherein the data representing the graph node for each K-mer stored in the hash table cache and the data representing the graph edges stored in the hash table cache represent a K-mer graph of the first set of nucleic acid sequences.; 16. The method of claim 15, the method further comprises: periodically storing, by means of the control machine and in a memory unit accessible by the control machine, graph description data for an instance of the K-number graph, characterized in that the graph description data represents (i) an identifier of the K-number graph and (ii) state information of the K-number graph.
17. The method of claim 15, characterized in that the first logical hardware unit is further configured to: determine whether one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence; and based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence, store data that marks the one or more particular K-mers as non-unique K-mers.
18. The method of claim 15, characterized in that the second hardware logic is further configured to: assign an edge weight to each edge of the K-number graph.
19. The method of claim 15, the method further comprises: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: I 70070 / 7707 / 3 / YILI obtain data representing the K-number range of the hash table cache; and provide the obtained data representing the K-number range to a variant call unit.
20. The method of claim 15, the method further comprises: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: selectively delete the data representing the grato nodes and the data representing the grato edges of the K-number graph of the hash table cache.
21. The method of claim 15, characterized in that the control machine is implemented using a third logic unit of the programmable logic device hardware.
22. The method of claim 15, characterized in that the hash table cache is implemented using a third hardware logic unit of the programmable logic device.
23. The method of claim 15, characterized in that the control machine is implemented using one or more CPUs or GPUs that execute software instructions to perform the functionality of the control machine.
24. The method of claim 15, characterized in that the grate description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer grate or nucleic acid sequences of the stack associated with the K-mer grate identifier.
25. The method of claim 15, the method further comprises: evaluating the k-number graph to check for the existence of cycles in the graph; if a cycle is detected during the evaluation: terminating the generation of the k-number graph; or if no cycle is detected during the evaluation: obtaining data from the hash table cache describing the structure of the k-number graph; and providing, to a variant call module, the obtained data describing the structure of the k-number graph.
26. A system for hardware-accelerated generation of a K-mer graph using a programmable logic device, the system comprising: a hardware-accelerated graph generation unit including hardware digital logic circuits arranged to perform operations; the operations comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence; for each particular nucleic acid sequence of the first set of nucleic acid sequences: generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a graph node for each K-mer of the particular nucleic acid sequence;detect, by means of a control machine, that the first hardware logic unit has completed the generation of a graph node for each K-mer of the particular nucleic acid sequence; configure, by means of the control machine, a second hardware logic unit to perform the generation of graph edges for the generated graph nodes; and for one or more pairs of the generated graph nodes: generate, by means of the second hardware logic unit and for storage in the graph hash table, data representing the graph edges between one or more pairs of the generated graph nodes generated by the first hardware logic unit, wherein the data representing the graph node for each K-mer stored in the hash table cache and the data representing the graph edges stored in the hash table cache represent a K-mer graph of the first set of nucleic acid sequences.
27. The system of claim 26, the operations further comprise: periodically storing, by means of the control machine and in a memory unit accessible by the control machine, graph description data for an instance of the K-number graph, characterized in that the graph description data represents (i) an identifier of the K-number graph and (ii) state information of the K-number graph.
28. The system of claim 26, characterized in that the first logical hardware unit is further configured to: determine whether one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence; and based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence, store data that marks the one or more particular K-mers as non-unique K-mers.
29. The system of claim 26, characterized in that the second hardware logic is further configured to: assign an edge weight to each edge of the K-number graph.
30. The system of claim 26, the operations further comprise: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: obtain data representing the K-number graph from the hash table cache; and provide the obtained data representing the K-number graph to a variant call unit.
31. The system of claim 26, the operations further comprise: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: selectively delete data representing graph nodes and data representing graph edges from the K-number graph of the hash table cache.
32. The system of claim 26, characterized in that the hash table cache is implemented using a third hardware logic unit of the programmable logic device. LZQCZn / ZZnZ / q / YIAI 33. The system of claim 26, characterized in that the grate description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer grate or nucleic acid sequences of the stack associated with the K-mer grate identifier.
34. The system of claim 26, the operations further comprise: evaluating the k-numbers array to check for the existence of array cycles; if an array cycle is detected during the evaluation: terminating the generation of the k-numbers array; or if no array cycle is detected during the evaluation: obtaining data from the hash table cache describing the structure of the k-numbers array; and providing, to a variant call module, the obtained data describing the structure of the k-numbers array.
35. A hardware-accelerated gene generation unit comprising hardware digital logic circuits arranged to perform operations; the operations comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence;generate, using a plurality of non-segmented hardware logic units of a programmable logic device, a K-mer array using the first set of nucleic acid sequences obtained, wherein each hardware logic unit comprises a different hardware logic circuit configured to perform one or more operations, wherein each node of the K-mer array represents a K-mer, each edge of the K-mer array represents a link between a pair of K-mers, and each weight of each edge of the K-mer array represents a number of occurrences of a K-mer sequence represented by a pair of K-mers;and during the generation of the K-number graph: periodically updating, with a control machine, the graph description data for the K-number graph after the performance of one or more operations by each logical hardware unit used to generate at least one LZQCZn / ZZnZ / q / YIAI portion of the K-number graph, wherein the graph description data represents (i) an identifier of the K-number graph and (ii) state information of the K-number graph, wherein the control machine creates a workflow of operations using the non-segmented logical hardware units by activating the performance of one or more operations of each respective logical hardware unit during the generation of the K-number graph.
36. The hardware-accelerated gratos generation unit of claim 35, characterized in that the output of each hardware logic unit of the plurality of hardware logic units is stored through a hash table cache.
37. The hardware-accelerated graph generation unit of claim 35, characterized in that the operations further include: providing the generated K-number graph to a variant call unit, wherein the variant call unit is configured to process the K-number graph to determine candidate variants from one or more of the plurality of reads and the reference sequence.
38. The hardware-accelerated graph generation unit of claim 35, characterized in that the operations further include: obtaining, by means of a variant call unit, the generated K-mer graph; and identifying, based on the processing of the generated K-mer graph by the variant call unit, one or more candidate variants, characterized in that a candidate variant is a difference between a base call of one or more reads in the read stack and a nucleotide of a reference genome at a particular location of the reference genome.
39. The hardware-accelerated graph generation unit of claim 35, characterized in that the graph description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer graph or nucleic acid sequences of the stack associated with the K-mer graph identifier. LZQCZn / ZZnZ / q / YIAI 40. A hardware-accelerated grate generation unit comprising hardware digital logic circuits arranged to perform operations; the operations comprising: obtaining a first set of nucleic acid sequences, characterized in that the first set of nucleic acid sequences includes (i) a plurality of reads corresponding to an active region of a reference sequence and (ii) a portion of the reference sequence; for each particular nucleic acid sequence in the first set of nucleic acid sequences: generating, for storage in a hash table cache and by means of a first hardware logic unit, data representing a grate node for each K-mer of the particular nucleic acid sequence; detecting, by means of a control machine, that the first hardware logic unit has completed the generation of a grate node for each K-mer of the particular nucleic acid sequence;Configure, using the control machine, a second hardware logic unit to perform the generation of grato edges for the generated grato nodes; and for one or more pairs of the generated grato nodes: generate, using the second hardware logic unit and for storage in the grato hash table, data representing the grato edges between one or more pairs of the grato nodes generated by the first hardware logic unit, wherein the data representing the grato node for each K-mer stored in the hash table cache and the data representing the grato edges stored in the hash table cache represent a grato of Kmers from the first set of nucleic acid sequences.
41. The hardware-accelerated grate generation unit of claim 40, the operations further comprising: periodically storing, by means of the control machine and in a memory unit accessible by the control machine, grate description data for an instance of the K-number grate, characterized in that the grate description data represents (i) an identifier of the K-number grate and (ii) state information of the K-number grate. I 7QC 70 / 7707 / 3 / YILI 42. The hardware-accelerated graph generation unit of claim 40, characterized in that the first hardware logic unit is further configured to: determine whether one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence; and based on a determination that one or more of the particular K-mers of the particular nucleic acid sequence match another K-mer of the particular nucleic acid sequence, store data marking the one or more particular K-mers as non-unique K-mers.
43. The hardware-accelerated graph generation unit of claim 40, characterized in that the second hardware logic is further configured to: assign an edge weight to each edge of the K-number graph.
44. The hardware-accelerated graph generation unit of claim 40, the operations further comprising: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: obtain data representing the K-number graph from the hash table cache; and provide the obtained data representing the K-number graph to a variant call unit.
45. The hardware-accelerated graph generation unit of claim 40, the operations further comprising: instructing a third hardware logic unit of the programmable logic device to execute hardware logic configured to: selectively remove data representing graph nodes and data representing graph edges from the K-number graph of the hash table cache.
46. The hardware-accelerated grating unit of claim 40, characterized in that the hash table cache is implemented using a third hardware logic unit of the programmable logic device. I 7QC 70 / 7707 / 3 / YILI 47. The hardware-accelerated graph generation unit of claim 40, characterized in that the graph description data further includes (iii) data representing a last hardware logic unit of the plurality of hardware logic units that executed the hardware logic on the K-mer graph or nucleic acid sequences of the stack associated with the K-mer graph identifier.
48. The hardware-accelerated graph generation unit of claim 40, the operations further comprising: evaluating the K-number graph to check for the existence of graph cycles; 10 if a graph cycle is detected during the evaluation: terminating the generation of the k-number graph; or if no graph cycle is detected during the evaluation: obtaining data from the hash table cache describing the structure of the k-number graph; and 15 providing, to a variant call module, the obtained data describing the structure of the k-number graph.