RAPID DETECTION OF GENE FUSIONS

MX431843BActive Publication Date: 2026-02-25ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2021012019
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-05
Filing Date
2021-09-30
Publication Date
2026-02-25
Estimated Expiration
2040-12-04

AI Technical Summary

Technical Problem

Current methods for detecting gene fusions are inefficient and resource-intensive, requiring the processing and scoring of all candidates, which leads to high computing resource consumption and prolonged lead times.

Method used

A method utilizing a filtering engine to narrow down gene fusion candidates, combined with a machine learning model trained to identify valid gene fusions based on feature data from read alignment units, reducing the number of candidates processed and optimizing computing resources.

Benefits of technology

This approach achieves high-precision detection of gene fusions with reduced execution time, computing resources, and power consumption, while significantly decreasing processing overhead.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Methods, systems, and devices are described, including computer programs for identifying a gene fusion in a biological sample.The method may include actions to obtain first data representing a plurality of aligned reads, identify a plurality of fusion candidates included within the first data obtained, filter the plurality of fusion candidates to determine a filtered set of fusion candidates, for each particular fusion candidate in the filtered set of fusion candidates: generate, by one or more computers, input data for input to a machine learning model that includes extracted feature data to represent the particular fusion candidate, provide the generated input data as input to the machine learning model that has been trained to generate output data representing a probability that a fusion candidate is a valid gene fusion, and determine whether the particular fusion candidate corresponds to a valid gene fusion based on the output data.
Need to check novelty before this filing date? Find Prior Art

Description

RAPID DETECTION OF GENE FUSIONS Cross-reference to related request This application claims the benefit of U.S. provisional patent application no. 62 / 944,304 filed on December 5, 2019, which is incorporated herein by reference in its entirety. Background Gene fusions can be used as oncogenic drivers that are important diagnostic and therapeutic targets in the treatment of diseases such as cancer. Summary According to an innovative aspect of the present description, a computer-implemented method for identifying one or more gene fusions in a biological sample is described. In one aspect, the method may include actions to obtain, using one or more computers, the first data representing a plurality of aligned reads from a read alignment unit; to identify, using one or more computers, a plurality of gene fusion candidates included within the first data obtained; and to filter, using one or more computers, the plurality of gene fusion candidates to determine a filtered set of gene fusion candidates.for each particular gene fusion candidate in the filtered set of gene fusion candidates: generate, by means of one or more computers, input data for input to a machine learning model, wherein generating the input data comprises extracting feature data to represent the particular gene fusion candidate from data that include: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit;to provide, by means of one or more computers, the generated input data as an input to the machine learning model, wherein the machine learning model has been trained to generate output data that represent a probability that a gene fusion candidate is a valid gene fusion based on the machine learning model processing input data that represent (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results; RLHZ Ln / Lznz / E / Yli of the read alignment unit; obtain, using one or more computers, output data generated by the machine learning model based on the machine learning model that processes the generated input data; and determine, using one or more computers, whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data. Other versions include the corresponding systems, devices, and computer programs to perform the actions of methods defined by instructions encoded on computer-readable storage devices. These and other versions may optionally include one or more of the following features. For example, in some implementations, generating the input data also involves extracting feature data, including annotation data that describes annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit.In such implementations, the machine learning model has been trained to generate output data that represent a probability that a gene fusion candidate is a valid gene fusion candidate based on input data from the machine learning model's processing that represent: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, (ii) annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit. In some implementations, identifying, using one or more computers, a plurality of gene fusion candidates included within the first data obtained may include identifying, using one or more computers, a plurality of split-read alignments. In some implementations, identifying, using one or more computers, a plurality of gene fusion candidates included within the first data obtained comprises identifying, using one or more computers, a plurality of non-matching read pair alignments. In some implementations, the read alignment unit is implemented using a set of one or more processing engines configured with hardware logic circuits physically arranged to perform operations, using the hardware logic circuits, to: (i) RLHZ Ln / Lznz / E / Yli receive data representing a first read, (i) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence locations, (iii) generate one or more alignment scores corresponding to each of the matching reference sequence locations for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) produce data representing a candidate alignment for the first read. In some implementations, the read alignment unit is implemented using a set of one or more processing engines using one or more central processing units (CPUs) or one or more graphics processing units (GPUs) to execute software instructions that cause the one or more CPUs or one or more GPUs to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more reference sequence locations that match the first read, (iii) generate one or more alignment scores corresponding to each of the reference sequence locations that match the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores,and (v) produce data that represent a candidate alignment for the first read. In some implementations, the method may further include receiving, via the read alignment unit, a plurality of reads that are not yet aligned, aligning, via the read alignment unit, a first subset of the plurality of reads, and storing, via the read alignment unit, the first subset of aligned reads in a memory device. In such implementations, obtaining, via one or more computers, the first data representing a plurality of aligned reads from a read alignment unit may include obtaining, via one or more computers, the first subset of aligned reads from the memory device and performing one or more of the operations of claim 1, while the read alignment unit aligns a second subset of the plurality of reads that are not yet aligned. In some implementations, the data generated based on the read alignment unit result may include one or more of a variant allele frequency count, a count of unique read alignments, a read coverage across the transcript, a MAPQ score, or data indicating homology between precursor genes. RLHZ Ln / Lznz / E / Yli In some implementations, determining whether a particular fusion candidate corresponds to a valid gene fusion candidate based on output data may include determining, using one or more computers, whether the output data satisfies a predetermined threshold, and based on the determination that the output data satisfies the predetermined thresholds, determining that the particular fusion candidate corresponds to a valid gene fusion candidate. In some implementations, determining whether a particular fusion candidate corresponds to a valid gene fusion candidate based on output data may include: determining, by one or more computers, whether the output data satisfies a predetermined threshold, and based on the determination that the output data does not satisfy the predetermined thresholds, determining that the particular fusion candidate does not correspond to a valid gene fusion candidate. These and other innovative aspects of the present description are readily apparent from the detailed description, the accompanying figures, and the claims. Brief description of the figures Figure 1 is a block diagram of an example of a system for the rapid detection of valid gene fusions. Figure 2 is a flowchart of an example of a process for performing rapid detection of valid gene fusions. Figure 3 is a block diagram of another example of a system for the rapid detection of valid gene fusions. Figure 4 is a block diagram of system components that can be used to implement a system for the rapid detection of valid gene fusions. Detailed description This description pertains to systems, methods, devices, software programs, or any combination thereof, for the rapid detection of gene fusions. The presence of certain gene fusions can constitute important indicators of a particular disease, an indicator suggesting the use of a particular therapeutic agent for a particular disease, or RLnz Ln / lzoz / b / yl or a combination of these. For example, certain gene fusions may be indicators of a particular type of cancer, such as acute and chronic myeloid leukemias, myelodysplastic syndromes (MDS), soft tissue sarcomas, or treatments for these. The present description can rapidly detect accurate gene fusions by using a filtering engine to narrow down a number of gene fusion candidates (also referred to as “fusion candidates” in this description) and process them to determine whether each fusion candidate is a valid gene fusion.This filtering engine allows for high-precision selection of fusion candidates for further analysis while also achieving a reduction in the computing resources that need to be consumed to identify valid gene fusions, since only the filtered subset of candidate gene fusions can be promoted for further processing as described herein. The reduced set of candidate gene fusions also provides other technological advantages. For example, the methods and systems described herein offer reduced execution time compared to conventional methods that process and qualify all gene fusion candidates. The reduced execution time for performing these operations also directly results in reduced processing resource expenditure (e.g., CPU or GPU resources), memory usage, and power consumption. While a filtering engine provides reduced execution time compared to conventional methods, the methods and systems described herein can also provide other ways to reduce execution time.For example, in some implementations, even further reductions in runtime can be achieved by using a hardware-accelerated read alignment unit to perform the mapping, alignment, and metadata generation used to process candidate gene fusions. Figure 1 is a block diagram of an example of a system 100 for the rapid detection of valid gene fusions. The system 100 may include a nucleic acid sequencing device 110, a memory 120, a secondary analysis unit 130, a fusion candidate identification module 140, a fusion candidate filtering module 150, a feature set generation module 160, a machine learning model 170, a gene fusion determination module 180, an output application programming interface (API) module 190, and an output display 195. In the example in Figure 1, each of these components is described as being implemented within the nucleic acid sequencing device 110. However, this description is not limited to such modalities. RLHZ Ln / Lznz / E / Yll In contrast, in some implementations, one or more of the components described in Figure 1 may run on a computer outside the nucleic acid sequencing device 110. For example, in some implementations, the secondary analysis modules may be implemented within the nucleic acid sequencing device 110, and the fusion candidate identification module 140, a fusion candidate filtering module 150, a feature set generation module 160, a machine learning model 170, a gene fusion determination module 180, and an output application programming interface (API) module 190 may be implemented on one or more different computers. In these implementations, the one or more different computers and the nucleic acid sequencing device may be coupled communicatively using one or more wired networks, one or more wireless networks, or a combination of these. For the purposes of this specification, the term “module” includes one or more software components, one or more hardware components, or any combination thereof, that can be used to perform the functionality attributed to a respective module by this specification. Generally, a “module,” as described herein, uses one or more processors to execute software instructions to perform the module functionality described herein. A processor may include a central processing unit (CPU), a graphics processing unit (GPU), or the like. Similarly, the term “unit,” as used in this specification, includes one or more software components, one or more hardware components, or any combination thereof, that can be used to perform the functionality attributed to a respective unit by this specification. Generally, a “unit,” as described herein, uses one or more hardware components, such as hardwired digital logic gates or hardwired digital logic blocks arranged as processing engines to perform operations that accomplish the functionality of the unit described herein. Such hardwired digital logic gates or hardwired digital logic circuits may include a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIO), or the like. The nucleic acid sequencing device 110 (also referred to herein as the sequencing device 110) is configured to perform primary nucleic acid sequencing analysis. Performing primary analysis may include receiving a biological sample 105, such as a blood sample, tissue sample, sputum sample, or nucleic acid sample, via the sequencing device 110, and generating output data such as one or more reads 112, each representing RLHZ Ln / Lznz / E / Yli, a nucleotide sequence order from a nucleic acid sequence of the received biological sample. In some implementations, sequencing using the 110 nucleic acid sequencer can be performed in multiple read cycles, with a first read cycle, “Read 1,” generating one or more first reads representing a nucleotide sequence order from one end of a nucleic acid sequence fragment, and a second read cycle, “Read 2,” generating one or more second reads, respectively, representing a nucleotide sequence order from the other ends of one of the nucleic acid sequence fragments. In some implementations, the reads may be short reads of approximately 80 to 120 nucleotides in length. However, this description is not limited to reads of any particular nucleotide length. Instead, this description can be used to read any nucleotide length.In some implementations, the biological sample 105 may include a DNA sample, and the nucleic acid sequencer 110 may include a DNA sequencer. In such implementations, the sequenced nucleotide order in a read generated by the nucleic acid sequencer may include one or more guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. In some implementations, the nucleic acid sequencer 110 may be used to produce RNA reads from a biological sample 105. In such implementations, this may occur using RNA sequencing protocols. For example, a biological sample 105 may be preprocessed using reverse transcription to form complementary DNA (cDNA) using a reverse transcriptase enzyme. In other implementations, the nucleic acid sequencer 110 may include an RNA sequencer, and the biological sample may include an RNA sample.RNA reads produced using cDNA or through an RNA sequencer may comprise C, G, A, and uracil (II). The example in Figure 1 described herein refers to the generation and analysis of RNA reads. However, the present invention can be used to produce and analyze any type of nucleic acid sequence reads, including DNA or RNA reads. The 110 sequencing device can include a next-generation sequencer (NGS) configured to generate sequence reads such as 1121, 112-2, 112-n, where “n” is any positive integer greater than 0, for a given sample in a manner that achieves exceptionally high throughput, scalability, and speed using massively parallel sequencing technology. NGSs enable rapid sequencing of entire genomes, the ability to zoom in on deeply sequenced target regions, and the use of RNA sequencing (RNA sec.) to discover novel RNA variants and splice sites. RLHZ Ln / Lznz / E / Yli or quantify mRNAs for gene expression analysis, analysis of epigenetic factors such as genome-wide DNA methylation and DNA-protein interactions, sequencing cancer samples to study rare somatic and tumor subclones, and studying microbial diversity, e.g., in humans or the environment. The sequencing device 110 can sequence the biological sample 105 and generate a corresponding set of reads represented using A, C, T, and G. The sequencing device can then perform reverse transcription to generate a cDNA sequence representing the corresponding RNA sequence. These RNA sequence reads 112-1, 112-2, 112-n are produced by the sequencing device 110 and stored in the memory device 120. In some implementations, the RNA sequence reads 112-1, 112-2, 112-n can be compressed into smaller data records before being stored in the memory device 120.The memory device 120 can be accessed by each of the components in Figure 1, which include the secondary analysis unit 130, the fusion candidate identification module 140, the fusion candidate filtering module 150, the feature set generation module 160, the machine learning model 170, the gene fusion determination module 180, and the output API module 190. Although the respective modules can be represented as providing a result from a first module to a second module, the practical implementation of such a feature may include the first module storing the results in a memory device such as memory 120, and the second module accessing the stored results from the memory device and processing the accessed results as input to the second module. The secondary analysis unit 130 can access reads 112-1, 112-2, and 112-n stored in memory device 120 and perform one or more secondary analysis operations on reads 112-1, 112-2, and 112-n. In some implementations, reads 112-1, 112-2, and 112-n may be stored in memory device 120 in compressed data registers. In such implementations, the secondary analysis unit may perform decompression operations on the compressed read registers before performing secondary analysis operations on the read registers. Secondary analysis operations may include mapping one or more reads to a reference genome, aligning one or more reads to the reference genome, or both. In some implementations, secondary analysis operations may also include variant designation operations.In addition to performing secondary analysis operations, the 130 secondary analysis unit can also be configured to perform classification operations. Classification operations can include, for example, sorting readings that have been aligned by the analysis unit. RLHZ Ln / Lznz / E / Yli secondary based on the position in the reference genome to which the aligned reads were mapped. In some implementations, such as the example in Figure 1, the secondary analysis unit 130 may include a memory 132 and a programmable logic device 134. The programmable logic device 134 may have hardware logic circuits that can be dynamically configured to include one or more secondary analysis operational units, such as a read alignment unit 136, and may be used to perform one or more secondary analysis operations using hardware logic circuits.Dynamically configuring the programmable logic device 134 to include a secondary analysis operational unit, such as a read alignment unit 136, may include, for example, providing one or more instructions to the programmable logic device 134 that cause the programmable logic device 134 to arrange the hardware logic gates of the programmable logic device 134 into a coded digital logic configuration that is configured to perform the functionality, in hardware logic, of the read alignment unit 136. The one or more operations that activate the dynamic configuration of the programmable logic device 134 may include compiled hardware description language code, one or more instructions for the programmable logic device 134 to configure itself given in the compiled hardware description language code, or the like. Such operations that activate the dynamic configuration of the programmable logic device 134 may be generated and implemented on the programmable logic device 134 by a control program executed by the sequencing device 110, or another computer hosting the control program. In some embodiments, the control program may be a software module whose instructions reside in a memory device, such as memory 120.The functionality of the control program to generate and implement hardware description language code instructions or other instructions to configure the programmable logic device 134 can be accomplished by running the control program software module using one or more processors such as one or more CPUs or one or more GPUs. The functionality of the read alignment unit 136 may include obtaining one or more first reads, such as RNA reads 112-1, 112-2, 112-n, stored in memory 120 by the sequencing device 110, mapping the obtained first reads 112-1, 112-2, 112-n to one or more reference sequence locations within a reference sequence, and then aligning the mapped first reads 112-1, 112-2, 112-n to the reference sequence. That is, the mapping step may identify a set of candidate reference sequence locations for each particular first read obtained that match the RLnZLn / LZnZ / E / Yli particular read. Then, the alignment stage can mark each location of the candidate reference sequence and select the location of the particular reference sequence that has the highest alignment score as the correct alignment for the particular read. A reference sequence can include an organized series of nucleotides that correspond to a known genome. Arranging the hardware logic gates of the programmable logic device 134, which responds to one or more instructions from the control program, may include configuring logic gates such as AND gates, OR gates, ÑOR gates, XOR gates, or any combination thereof, to execute digital logic functions of a read alignment unit 136. Alternatively, or in addition, arranging the hardware logic gates may include dynamically configured logic blocks comprising customizable hardware logic units to perform complex computational operations including addition, multiplication, comparisons, or the like. The precise arrangement of the hardware logic gates, logic blocks, or a combination thereof, is defined by the instructions received from the control program.The instructions received may include, or be derived from, a compiled hardware description language (HDL) program code that was written by an entity and defines the schematic design of the secondary analysis operational unit to be programmed into the programmable logic device 134. The HDL program code may include program code written in a language such as a very high-speed integrated circuit hardware description language (VHDL), Verilog, or a similar language. The entity may include one or more human users who wrote the HDL program code, one or more artificially intelligent agents that generated the HDL program code, or a combination of these. The programmable logic device 134 can include any type of programmable logic device. For example, the programmable logic device 134 can include one or more field-programmable gate arrays (FPGAs), one or more complex programmable logic devices (CPLDs), or one or more programmable logic arrays (PLAs), or a combination thereof, which are dynamically configurable and reconfigurable, as required, by the control program to execute a particular workflow. For example, in some implementations, it may be desirable to use the programmable logic device 134 as a read alignment unit 136, as described above.However, in other implementations, it may be desirable to use the programmable logic device 134 to perform variant designation functions or functions to support variant designation, such as a Hidden Markov Model (HMM) unit. Even in other implementations... The programmable logic device 134 (RLHZ Ln / Lznz / E / Yli) can also be dynamically configured to support general computing tasks, such as compression and decompression, because the hardware logic of the programmable logic device 134 is capable of performing these tasks, as well as the other tasks identified above, much faster than the performance of the same tasks using software instructions executed by one or more processing units 150. In some implementations, the programmable logic device 134 can be dynamically reconfigured during runtime to perform different operations. As an example, in some implementations, the programmable logic device 134 can be implemented using an FPGA that is dynamically configured as a decompression unit to access data representing a compressed version of the first reads 112-1, 1122, 112-n stored in memory device 120 or 132. The secondary analysis unit 130 can use the decompression unit to decompress the compressed data representing the first reads 112-1, 112-2, 122-n (e.g., if the reads received from the nucleic acid sequencer are compressed). The decompression unit can store decompressed reads in memory 120 or 132. In such implementations, the FPGA can then be dynamically reconfigured as a read alignment unit 136 and used to perform mapping and alignment of the first decompressed reads 112-1, 112-2, 112-n now stored in memory 132 or 120.The read alignment unit 136 can then store data representing the mapped and aligned reads in memory 132 or 120. Although a series of operations, including decompression, mapping, and alignment, are described, this description is not limited to performing those operations or only those operations. Instead, the programmable logic device 134 can be dynamically configured to perform the functionality of any operational unit in any order, as required, to accomplish the functionality described herein. The example in Figure 1 describes a secondary analysis unit 130 that uses a hardware logic device in the form of a programmable logic device 134 to implement a read alignment unit 136. However, the present description is not limited to the use of programmable logic devices to implement the read alignment unit 136. Instead, other types of integrated circuits can be used to implement a read alignment unit 136 in a digital logic circuit of the secondary analysis unit 130. For example, in some implementations, a secondary analysis unit 143 can be configured to use one or more application-specific integrated circuits (ASICs) to implement the functionality of one or more secondary analysis operational units.Although they are not reprogrammable, one or more ASICs can be designed with custom hardware logic from one or more secondary analysis operational units. RLHZ Ln / Lznz / E / Yli, such as a read alignment unit 136, a variant designation unit, a variant designation computational support unit, or the like, to accelerate and match the performance of secondary analysis operations. In some implementations, using one or more ASICs as hardwired logic circuits of the secondary analysis unit 130, which performs the functionality of one or more secondary analysis operation units, can be even faster than using a programmable logic device such as an FPGA. Accordingly, an experienced technician would understand that an ASIC could be used in place of a programmable logic device, such as an FPGA, in any of the modes described herein. For implementations where ASICs are to be used, it would be necessary to use a dedicated ASIC or dedicated logic groups of a single ASIC for each secondary analysis operation unit that an ASIC is to perform.For example, one or more ASICs for read alignment, one or more ASICs for decompression, one or more ASICs for compression, or a combination of these. Alternatively, the same functionality could also be achieved with dedicated logic groups within the same ASIC. Furthermore, the examples in this description that refer to systems 100 and 300 in Figures 1 and 3, respectively, are described using a hardware implementation of a read alignment unit 136 on a programmable logic device. It was also noted earlier that one or more ASICs can be used to implement the read alignment engine or other secondary analysis operations. However, this description is not limited to the use of hardware units to implement such secondary analysis operations. Instead, in some implementations, any of the operations described herein as being performed by the programmable logic device, such as read alignment, compression, or decompression, can also be implemented using one or more software modules. With reference to the example in Figure 1, the execution of system 100 can begin with the sequencing device 110 sequencing the biological sample 105. The sequencing of the biological sample can include the generation, by the sequencing device 110, of read sequences that are a data representation of the ordered nucleotide sequences present in the biological sample 105. If system 100 is configured to process DNA reads, then the reads generated by the sequencing device 110 can be stored in memory 120. Alternatively, in some implementations, if system 100 is configured to process RNA reads, sequencing device 110 can be configured to perform preprocessing of the RLHZ Ln / Lznz / E / Yli biological sample 110 uses reverse transcription to form complementary DNA (cDNA) using a reverse transcriptase enzyme. In implementations such as the one in the example in Figure 1, the reads generated by the sequencing device 110 include RNA reads 112-1, 112-2, and 112-n. In other implementations, the nucleic acid sequencer 110 may include an RNA sequencer, and the biological sample may include an RNA sample. Regardless of whether the RNA reads are produced by a DNA sequencing device using cDNA or via an RNA sequencer, each RNA read includes a nucleotide sequence composed of C, G, A, and U. The 112-1, 112-2, and 112-n reads can be stored in memory 120 in either a compressed or uncompressed format. The execution of system 100 can continue with the secondary analysis unit 130 retrieving reads 112-1, 112-2, and 112-n stored in memory 120. In some implementations, the secondary analysis unit 130 can access reads 112-1, 112-2, and 112-n from memory device 120 and store the accessed reads 112-1, 112-2, and 112-n in memory 132 of the secondary analysis unit 130. In other implementations, after a control program determines that the sequencing of reads 112-1, 112-2, and 112-n is complete and that the secondary analysis unit 130 is available to perform secondary analysis operations, the control program can load reads 112-1, 112-2, and 112-n. in memory 132 of secondary analysis unit 130. If reads 112-1, 112-2, 112-n are compressed, the secondary analysis unit 130 can dynamically configure the programmable logic device 134 as a decompression unit to access reads 112-1, 112-2, 112-n in memory 132 or 120, decompress reads 112-1, 112-2, 112-n, and then store the decompressed reads 112-1, 112-2, 112-n in memory 1320 or 120. In some implementations, the secondary analysis unit can dynamically reconfigure the programmable logic device and perform the decompression in response to instructions from a control program. If reads 112-1, 112-2, and 112-n are uncompressed, the secondary parser unit 130 can access memory reads 132 or 120 and perform read alignment operations. In some implementations, the secondary parser unit 130 may receive an instruction from a control program that instructs it to configure or reconfigure the programmable logic device 134 to include a read alignment unit 136 and then use the read alignment unit 136 to perform the alignment of reads 112-1, 112-2, and 112-n. Alternatively, in other implementations, the programmable logic device may have been configured to include a read alignment unit. RLHZ Ln / Lznz / E / Yli 136 and use the read alignment unit 136 to perform the alignment of reads 112-1, 112-2, 112-n. In still other implementations, the secondary analysis unit 130 may include an ASIC that is configured to perform read alignment and then use the ASIC to perform the alignment of reads 112-1, 112-2, 112-n. The secondary analysis unit 130 can be configured to perform read alignment operations in parallel with gene fusion analysis. For example, the secondary analysis unit 140 can obtain an initial batch of reads generated by the sequencing device 110 that are not aligned, use the read alignment unit 136 to align the initial batch of reads, use a classification engine that can be implemented in a hardware configuration of the programmed logic device 136 or implemented in software by executing program instructions to classify the aligned reads, and then produce the initial batch of aligned and classified reads for storage in a memory device 132,130.In some implementations, memory 132 can function as a local cache for the secondary analysis unit 132, which loads the data to be processed by the read alignment unit and then downloads the data produced by the read alignment unit 136. Therefore, once the read alignment unit 136 produces the first batch of aligned reads to memory 132, this first batch of aligned reads can be sorted and then produced to memory 120. The merge candidate identification module 140 can then access the first batch of aligned and sorted reads from memory 120 and begin processing them, while the secondary analysis unit 130 performs alignment operations on a second batch of reads generated by the sequencing device 110 that have not been previously aligned.This process can be repeated until each batch of reads is processed through System 100. Although this example is described as having batches that are aligned and sorted, there is no requirement in this description that the aligned read batches also be sorted. Instead, aligned and sorted reads can be used in either System 100 or System 300 to achieve performance improvements such as reduced runtime, as described later. The fusion candidate identification module 140 can obtain a batch of aligned and classified reads that were aligned using the read alignment unit 136 and determine whether the batch of aligned and classified reads includes one or more gene fusion candidates. In some implementations, if the received batch includes aligned and classified reads, then the fusion candidate identification module 140 can evaluate the classified reads in a batch where the genomic range corresponding to the batch overlaps with a breakpoint of at least one fusion candidate. This can reduce the number of fusion candidates requiring further analysis. In other implementations, if the received batch includes aligned reads that were not classified, then the fusion candidate identification module 140 can evaluate each of the aligned reads in the batch to determine whether the aligned read is a fusion candidate.In some implementations, the operation to determine, using the merge candidate identification module 140, whether the batch of reads includes one or more merge candidates includes determining, using the merge candidate identification module 140, where the batch of reads includes one or more split read alignments, one or more mismatched read pairs, one or more soft trim alignments, or a combination of these. In some implementations, the Fusion Candidate Identification Module 140 can be configured to identify split-read alignments as fusion candidates. This module identifies split-read alignments by analyzing the genes of a reference sequence to which each read in a batch of aligned reads was aligned. If the Fusion Candidate Identification Module 140 determines that a read maps to only one gene, then the read is not a split read. Alternatively, if the Fusion Candidate Identification Module 140 determines that a read aligns to two different genes, then the read is a split read. In such implementations, the split read is identified as a fusion candidate.A read can be determined to align with two different reads if, for example, a first subset of nucleotides in the read aligns with a first precursor gene in the reference genome, and a second subset of nucleotides in the read aligns with a second precursor gene in the reference genome. In some implementations, the first subset of nucleotides may be a prefix to the read, and the second subset of nucleotides may be a suffix to the read. If the merge candidate identification module 140 is configured to identify split reads, the data identifying the split reads, if any, can be stored in memory device 120. In some implementations, the Fusion Candidate Identification Module 140 can be configured to identify mismatched read pairs as fusion candidates. The Fusion Candidate Identification Module 140 can identify mismatched read pairs by analyzing genes from a reference sequence to which each particular read pair was aligned in a batch of aligned reads. If the read pair aligns to a reference sequence, and the orientation and range of the alignment are expected, then the read pair is determined not to be a mismatched read. Alternatively, if the read pair aligns to a reference sequence, and the orientation or range of the alignment is unexpected, then it is determined to be a mismatched read. RLHZ Ln / Lznz / E / Yll determines that the read pair is a mismatched read pair. In such implementations, if one read in the pair maps to one precursor gene and the other maps to a different precursor gene, the mismatched read can be identified as a fusion candidate. If the fusion candidate identification module 140 is configured to identify mismatched reads, the data identifying the mismatched reads, if any, can be stored in memory device 120. In some implementations, the Fusion Candidate Identification Module 140 can be configured to identify soft-clipping alignments. This module can identify soft-clipping alignments by analyzing the genes of a reference sequence to which each particular aligned read in a batch of aligned reads was aligned. In some implementations, the Fusion Candidate Identification Module 140 can determine whether the read is aligned to only one location in the entire reference genome. If the Fusion Candidate Identification Module 140 determines that the read is aligned to only one location in the entire reference genome, then the Fusion Candidate Identification Module 140 can determine that the read is not a soft-clipping read.Alternatively, if fusion candidate identification module 140 determines that only a portion of the read aligns with the reference genome, then the fusion candidate identification module 140 can determine that the read is a soft-cut read. If the aligned portion of the read maps to a precursor gene and the unaligned portion is determined to have a sequence similar to another precursor gene, then the soft-cut read is determined to be a fusion candidate. If fusion candidate identification module 140 is configured to identify soft-cut reads, the data identifying the soft-cut reads, if any, can be stored on memory device 120 as a gene fusion candidate. The fusion candidate filtering module 150 can obtain data describing a set of fusion candidates identified by the fusion candidate identification module 140. In some implementations, the fusion candidate filtering module can access memory device 120 and obtain data describing the fusion candidates on memory device 120. In other implementations, the fusion candidate filtering module can receive data describing fusion candidates from the output of a previous module, such as the fusion candidate identification module 140. The fusion candidate filtering module 150 can use one or more filters to filter the data describing the set of fusion candidates to identify a filtered set of gene fusion candidates that is smaller than the full set of gene fusion candidates. In some implementations, these filters are applied in a single step.For example, each of one or more filters can be applied, and each merge candidate in the set of merge candidates can be evaluated against each of the one or more filters. However, in other implementations, they may not. ML / a / ZUZl / U1 ZU I and multi-stage filtering approaches can be used. In such implementations, a first set of one or more filters is applied to the initial set of merge candidates identified by the merge candidate identification module 140. Then, a second set of one or more filters is applied to the first set of filtered merge candidates remaining after the application of the first filtering stage. Additional filtering stages can also be applied as needed to achieve an optimal filtered set of merge candidates. In some implementations, the fusion candidate filter module 150 can filter the fusion candidate pool to account for duplicate fusion candidates originating from the high depths of coverage used during short-read sequencing. For example, an accumulation that occurs from 30x sequencing might cause the fusion candidate identification module 140 to identify up to 30 duplicate fusion candidates. The fusion candidate filter module 150 can eliminate such duplicate fusion candidates by applying a filter to the fusion candidate features to check for duplicates. For example, the fusion candidate filter module 150 can determine whether multiple fusion candidates align to the same precursor gene, align to a portion of the reference genome spanning the same or a similar breakpoint, or a combination of these.If the Fusion Candidate Filter Module 150 identifies multiple fusion candidates that align to the same precursor gene, align to a portion of the reference genome spanning the same or a similar breakpoint, or a combination thereof, the Fusion Candidate Filter Module 150 can determine that the fusion candidates are duplicates and select only one of the fusion candidates as a representative fusion candidate. In such cases, the remaining fusion candidates that align to the same precursor gene, align to a portion of the reference genome spanning the same or a similar breakpoint, or a combination thereof, can be discarded without further analysis. The representative fusion candidate can then be added to a pool of filtered fusion candidates on a memory device, such as the Memory Device 120. Alternatively, or in addition, the merge candidate filter module 150 can filter the set of merge candidates based on one or more rule conditions. For example, the merge candidate filter module 150 can analyze each merge candidate and determine whether the merge candidate has one or more attributes that satisfy the one or more rule conditions used by the filter modules 150. In some implementations, the one or more rule conditions might include the alignment position of each portion of a merge candidate, the alignment overlap distance with respect to a breakpoint spanned by the merge candidate, the alignment orientation of the merge candidate, and the alignment quality. RLHZ Ln / Lznz / E / Yli merge candidate reading, an additional merge candidate mapping location, or any combination of these. As an example, one or more rule conditions can be used by the Merge Candidate Filter 150 module to filter merge candidates based on alignment position. In some implementations, for example, the Merge Candidate Filter 150 module can be configured to use a rule condition that filters merge candidates whose read is aligned to a reference sequence such that the extent of the alignment crosses a merge breakpoint by more than a predetermined number of nucleotides. In some implementations, the default number of nucleotides for this rule condition can be 8 nucleotides.Alternatively, or in addition, the 150 fusion candidate filter module can be configured to filter fusion candidates that have a read aligned to a reference sequence such that the portion of the alignment at the reference sequence does not reach a predetermined threshold number of nucleotides from the fusion breakpoint. In some implementations, the predetermined threshold number of nucleotides for this rule condition may be 50 nucleotides. Alternatively, or in addition, the 150 fusion candidate filter module can be configured to use a rule condition that filters fusion candidates that have a read aligned to a reference sequence such that the aligned portions of the read at the two fusion breakpoints share at least a predetermined number of nucleotides. In some implementations, the predetermined number of shared nucleotides may include at least 8 nucleotides. As another example, one or more rule conditions can be used by the fusion candidate filter module 150 to filter fusion candidates based on orientation. In some implementations, for example, the fusion candidate filter module 150 can be configured to use a rule condition that filters fusion candidates that have an alignment orientation indicating that a nucleotide sequence from at least one of the precursor genes is inverted in the fusion transcript. As another example, the merge candidate filter module 150 can use one or more rule conditions to filter merge candidates based on mapping quality. In some implementations, for example, the merge candidate filter module 150 can be configured to use a rule condition that filters merge candidates whose read alignment has a mapping quality score that does not meet a predetermined threshold. RLHZ Ln / Lznz / E / Yli As another example, the Fusion Candidate Filter 150 module can use one or more rule conditions to filter fusion candidates based on additional mapping locations. In some implementations, for example, the Fusion Candidate Filter 150 module can be configured to use a rule condition that filters fusion candidates based on a determination that a portion of the fusion candidate read maps to multiple locations in the reference sequence. In some implementations, the Fusion Candidate Filter 150 module can be configured to exclude locations annotated to be homologous genes. Merge candidates that satisfy each of one or more rule conditions can be added to a filtered pool of merge candidates in a memory device such as memory device 120. Merge candidates that do not satisfy each of one or more rule conditions can be discarded without further analysis. In some implementations, filtering of merge candidates based on rule conditions can be applied as a second-stage filter after the application of a first-stage deduplication filter. In other implementations, filtering of merge candidates based on rule conditions can be applied as the first stage of filtering, and then the deduplication filter can be applied as a second-stage filter. In still other implementations, filtering based on rule conditions can be applied as a single-stage filter without prior deduplication filtering.Filtering merger candidates based on one or more of these rule conditions can significantly reduce the number of merger candidates that need to be processed subsequently and additionally. Subsequent processing can be performed on each fusion candidate in the filtered set of fusion candidate output by the fusion candidate filtering module 150. This subsequent processing includes running the feature set generation module 160, the machine learning model 170, the gene fusion determination module 180, and the output API module 190. Such subsequent processing can be used to determine whether a fusion candidate corresponds to a valid gene fusion. The feature set generation module 160 can leverage data from multiple data sources to identify the set of data attributes on which to perform feature extraction. These data sources include in-memory attribute data 120 on the fusion candidate, which includes (i) the fusion candidate read(s), (ii) portion(s) of the reference sequence locations to which the fusion candidate reads were aligned, and (iii) annotations of the reference genome segments to which the particular gene fusion candidate was aligned. In some implementations, the RLHZ LO / L7R7 / E / YI annotations may include gene exon annotations, annotations indicating the presence of homologous genes, annotations indicating a list of enriched genes, or a combination of these. The data sources that feature set generation module 160 can use may also include data generated by read alignment unit 136 during the alignment process. In some implementations, feature set generation module 160 can derive feature data from the data generated by read alignment unit 136 during fusion candidate alignment. For example, feature set generation module 160 can derive information from data generated by read alignment unit 136, such as a variant allele frequency count, a count of unique read alignments, read coverage across the transcript, a MAPQ score, data indicating homology between precursor genes, or a combination thereof. The feature set generation module 160 can be used to generate feature data representing one or more of the attributes mentioned above from a merge candidate leverage from multiple data sources, and to encode the feature data into one or more data structures 162 for input to the machine learning model 170. For example, in some implementations, the entire feature set extracted from merge candidate attributes can be encoded into a single vector 162 that is incorporated into the machine learning module 170. For example, in the split-read or soft-clip alignment scenario, each of the features extracted from the attributes of these types of merge candidates can be encoded into a single vector 162. In other implementations, feature data may be extracted from the attributes of merge candidates, which can be multiple encoded input vectors. In such a scenario, input vector 162 may comprise a pair of input vectors 162a and 162b. For example, in the scenario of a split-read merge candidate, each of the features extracted from the split-read prefix-related attributes—including features representing the nucleotides of the split-read prefix, features representing the segment of the reference sequence to which the prefix aligns, and any other features extracted from the aforementioned prefix-related attributes, or any combination thereof—may be encoded in input vector 162a.Similarly, in such an implementation, each of the features extracted from attributes related to the split-read suffix includes the features that represent nucleotides of the split-read suffix, the features that. RLHZ Ln / Lznz / E / Yli represent the segment of the reference sequence to which the suffix aligns, and any other extracted features from the aforementioned attributes related to the suffix, or any combination thereof, can be encoded in input vector 162b. As another example, when a mismatched read pair is identified as a merge candidate, then the extracted features representing the first read of the mismatched read pair, the extracted features representing the portion of the reference sequence to which it aligned, the extracted features from the attributes related to the first read of the mismatched read pair, or any combination thereof, can be encoded in input vector 162a.Similarly, in such an example, the extracted features representing the second read of the mismatched read pair, the extracted features representing the portion of the reference sequence to which it aligned, the extracted features of attributes related to the second read of the mismatched read pair, or any combination of these, can be encoded in the input vector 162b. Each of the one or more vectors 162 can numerically represent the generated feature data, with the feature data including any of the features extracted from the merge candidate or any of the features extracted from data received from the read alignment unit 136 related to the merge candidate and stored in memory 120. For example, each vector 162 or 162a, 162b can include a plurality of fields, each of which corresponds to a particular feature from a particular read of a particular merge candidate. Depending on the particular merge candidate, this can result in one or more input vectors, as described above.The feature set generation module 160 can determine a numerical value for each field that describes the extent to which the particular feature was expressed in the attributes of the particular merge candidate reading. The numerical values ​​determined for each field can be used to encode the generated feature data representing attributes of the merge candidate readings into one or more respective vectors 162. The one or more generated vectors 162a, 162b, which numerically represent the corresponding merge candidate readings, are provided as inputs to the machine learning model 170. In some implementations, even if multiple conceptual vectors are generated for a merge candidate, the multiple conceptual vectors can be contacted with a single vector 162 that can be fed into the machine learning model 170.In such implementations, if multiple vectors were justified in (i) certain split-read implementations where the prefix features are assigned to a first vector and the suffix features are assigned to a second vector or (ii) in mismatched pair implementations, a first portion of the unique. RLHZ Ln / Lznz / E / Yli vector may correspond to the first conceptual vector and a second portion of the single vector could correspond to the second conceptual vector. The machine learning model 170 may include a deep neural network that has been trained to generate a probability that a fusion candidate corresponds to a valid gene fusion based on the processing of one or more input vectors 162 that represent features of a fusion candidate. A valid gene fusion is a chimeric transcript containing a sequence of multiple genes due to a rearrangement in the genome that connects a prefix of one precursor gene to the suffix of another precursor gene. In the context of this description, a valid gene fusion will be determined to have been predicted by the model 170 if, for example, the output data 178 generated by the machine learning model satisfy a predetermined threshold.The machine learning model 170 may include an input layer 172 for receiving input data, one or more hidden layers 174a, 174b, 174c for processing the input data received through the input layer 172, and an output layer 176 for providing output data 178. Each hidden layer 174a, 174b, 174c includes one or more weights or other parameters. The weights or other parameters of each respective hidden layer 174a, 174b, 174c may be adjusted, during training, so that the trained deep neural network produces the desired target output 178, which indicates a probability that the one or more input vectors 162 represent a valid gene fusion based on the machine learning model 170 processing the one or more input vectors 162. The machine learning model 170 can be trained in several different ways. In one implementation, the machine learning model 170 can be trained to distinguish between (i) one or more input vectors representing features extracted from attributes of valid fusion candidates and (ii) one or more input vectors representing features extracted from attributes of invalid fusion candidates. In some implementations, such training can be achieved using labeled pairs of training vectors. Each training vector can represent a training fusion candidate and comprise the same types of feature data as the one or more input vectors 162 above. In such implementations, one or more input vectors 162 representing features extracted from attributes of fusion candidates can be labeled as either a valid gene fusion or an invalid gene fusion.In some implementations, the valid gene fusion label or the invalid gene fusion label may be represented as a numeric value. For example, in some implementations, a valid gene fusion label may be a Ί” and an invalid gene fusion label may be “0”. In other implementations, for example, the valid gene fusion label may be a number between “0” and “1” that satisfies a predetermined threshold and a The invalid gene fusion label RLHZ Ln / Lznz / E / Yli can be a number between 0 and Ί that does not satisfy a predetermined threshold. In such implementations, the degree to which the number satisfies or fails to satisfy the predetermined threshold indicates a level of confidence that the training pair of input vectors represents a valid or invalid gene fusion. In some implementations, satisfying a predetermined threshold may include exceeding the predetermined threshold. However, implementations can also be configured to satisfy a threshold halfway between the predetermined threshold and the predetermined threshold. Such implementations might include, for example, implementations where both the comparator and the parameters are negated. During training, each labeled set of one or more training vectors is provided as input to the machine learning model 170, processed by the machine learning model 170, and then the training output generated by the machine learning model 170 is used to determine a predicted label for each labeled set of one or more training vectors. The predicted label generated by the machine learning model 170, based on the machine learning model's processing of the one or more training vectors that correspond to a pair of reads for a training merge candidate, can be compared to a training label for the one or more training vectors that correspond to the one or more reads (or portions of reads) for the training merge candidate.Next, the parameters of the machine learning model 170 can be adjusted based on the differences between the predicted labels and the training labels. This process can continue iteratively for each of a plurality of labeled training vector(s) corresponding to a respective training merge candidate until the predicted merge candidate labels produced by the machine learning model 170, based on processing a set of one or more training vectors corresponding to a training merge candidate, match, within a predetermined error level, the training labels of the set of one or more training vectors corresponding to the respective training merge candidate. In some implementations, the flagged training merge candidates can be obtained from a library of training merge candidates that have been reviewed and flagged by one or more human users. However, in other implementations, the flagged training merge candidates may include training merge candidates that have been generated and flagged by a simulator. In such implementations, the simulator can be used to create distributions of different categories of training merge candidates that can be used to train the machine learning model. Generally, if the model If the runtime machine learning module RLHZ Ln / Lznz / E / Ylj 170 accepts a single input vector 162, with each of the extracted features for a merge candidate being encoded by the single input vector 162, then the machine learning model 170 must be trained using a single input vector with the same features as the input vector 162 used in the previous training process. Similarly, if the runtime machine learning module 170 accepts two training vectors 162a and 162b, as described above, then the machine learning model 170 must be trained using two input vectors that each have the same corresponding features as the input vectors 162a and 162b above.In other words, the type of input vectors that will be processed at runtime is the same type of vectors that will be used to train model 170, using the training process described above. During the processing of input data 162, which correspond to features extracted from attributes of a fusion candidate, the output of each hidden layer 174a, 174b, 174c may include an activation vector. The resulting activation vector for each respective hidden layer can be propagated through subsequent layers of the deep neural network and used by the output layer to produce output data 178. In the example in Figure 1, the machine learning model 170 is trained to produce output data 178, which represents a combined score generated by the machine learning model 170 based on the machine learning model's processing of the separate input vectors 162a, 162b, each of which corresponds to one of the fusion candidate reads.This combined score 178 is finally produced by the output layer 176 of the machine learning model trained based on calculations performed by the output layer 176 of the machine learning model trained 170 on an activation vector received from the final hidden layer 174c. The output data 178 generated by the trained machine learning model 170 can be evaluated by a gene fusion determination module 180 to determine whether it indicates that the fusion candidate corresponding to one or more input vectors 162 is a valid fusion candidate. In some implementations, the output data 178 can be provided to the gene fusion determination module 180 by the trained machine learning model 170. In other implementations, the system 100 can store the output 178 of the trained machine learning model 170 in a memory device such as memory device 120 for later access by the gene fusion determination module 180. The gene fusion determination module 180 can obtain the output data 178 generated by the machine learning model 170 and evaluate the output data 178 to determine, based RLHZ Ln / Lznz / E / Yli in the output data 178, if the fusion candidate corresponding to the pair 162 of input vectors 162a, 162b is a valid gene fusion. In some implementations, the gene fusion determination module 180 can determine whether the fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion by comparing the output data 178 generated by the machine learning model with a predetermined threshold. If the gene fusion determination module 180 determines that the output data 178 satisfies the predetermined threshold, then the gene fusion determination module 180 can determine that the fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion.Alternatively, if the gene fusion determination module 180 determines that the output data 178 do not satisfy the predetermined threshold, then the gene fusion determination module 180 may determine that the fusion candidate corresponding to the one or more input vectors 162 is not a valid gene fusion. In some implementations, the gene fusion determination module 180 can generate output data 182 that indicate the results of the determination performed by the gene fusion determination module 180 based on the evaluation by the gene fusion determination module 180 of the output data 178 produced by the machine learning model 170. This output data 182 can include data that identifies the gene fusion candidate corresponding to one or more input vectors 162 and data that identifies the determination by the gene fusion determination module 180. The data that identifies the determination by the gene fusion determination module 180 can include data that indicates whether the gene fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion or an invalid gene fusion.In some implementations, output data 182 may only indicate the list of valid gene fusions identified based on output data 178, a list of invalid gene fusions identified based on output data 178, data indicating that no valid gene fusions were identified, or any combination thereof. In some implementations, this output data 182 may be stored in memory 182 for later use by another computer module, for subsequent output to a user device, or similar purposes. Alternatively, or in addition, the gene fusion determination module 180 can generate output data 184 that can be provided as input to the output application programming interface (API) module 190. The output data 184 can instruct the output API to display an output screen indicating whether the gene fusion candidate corresponding to the one or more input vectors 162 is a valid or invalid gene fusion. In some implementations, the instructions can cause the output API module 190 to access the output data 182 stored in memory device 120 and generate the RLHZ Ln / Lznz / E / Yli data representation that, when rendered by a computing device coupled to output display 195, causes output display 195 to display (i) data identifying the fusion candidate corresponding to one or more input vectors 162 and (ii) data indicating whether the identified fusion candidate is a valid gene fusion or an invalid gene fusion. This may include causing output display 195 to display any of the output data 182 stored in memory 184. In some implementations, this output may be displayed in the form of a report. In some implementations, the gene fusion determination module 180 stores the output data 182 for each gene fusion candidate in memory device 120 based on the performance of the subsequent processing performed on each fusion candidate in the filtered set of gene fusion candidates. In such implementations, the gene fusion determination module 180 can only instruct the output API module 190 to produce the gene fusion analysis results stored in memory 120 for each fusion candidate in the filtered set of gene fusion candidates once the subsequent processing of each fusion candidate is complete. In such a scenario, the results 192 provided for display on the output screen 195 would include a list of valid gene fusions, a list of invalid gene fusions, or both.In other implementations, the gene fusion determination module 180 can cause the output API module 190 to produce results data indicating a list of identified gene fusions, if any, upon completion of subsequent processing for that particular fusion candidate. The output API module 190 can provide other types of outputs 192. For example, in some implementations, outputs 192 might be data that causes another device, such as a printer, to output a report that includes (i) data identifying the fusion candidate corresponding to one or more vectors 162 and (ii) data indicating whether the identified fusion candidate is a valid gene. In other implementations, this output data 192 might cause a speaker to output audio data that includes (i) data identifying the fusion candidate corresponding to one or more vectors 162 and (ii) data indicating whether the identified fusion candidate is a valid gene. Other types of output data can also be triggered by output API modules 190. In some implementations, output display 195 may be a display panel of the sequencing device 110. In other implementations, output display 195 may be a display panel of a user device that connects to the sequencing device 110 using one or more networks. In fact, the sequencing device 110 can be used to communicate output data 192 to any device that has any display. RLHZ Ln / Lznz / E / Yll Figure 2 is a flowchart of an example of a process 200 for performing rapid detection of valid gene fusions. A system, such as system 100, can begin the execution of process 200 by using one or more computers to obtain initial data representing a plurality of aligned reads from a read alignment unit (210). The system can identify a plurality of gene fusion candidates included within the initial data obtained (220). The system can filter the plurality of gene fusion candidates to determine a filtered set of gene fusion candidates (230). The system can obtain a particular gene fusion candidate from the filtered set of gene fusion candidates (240). The system can generate input data for input to a machine learning model, wherein generating the input data includes extracting feature data to represent the particular gene fusion candidate from data that include (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit (250). The system can provide the generated input data as input to the machine learning model, wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion based on the machine learning model processing input data representing (i) segments of a reference genome to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit (260). The system can obtain output data generated by the machine learning model based on the machine learning model processing the input data (270). The system can determine whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data (280). Once step 280 is completed, the system can determine whether another merger candidate from the filtered set of merger candidates will be evaluated (290). If the system determines that another merger candidate exists from the filtered set of merger candidates to be evaluated, then the system can continue process execution 200 at step 240. Alternatively, if the system determines that there is no other merger candidate from the filtered set of merger candidates to be evaluated, then the system can terminate process execution at step 295. Another merger candidate may exist in the filtered set of merger candidates if the set of merger candidates has not been exhausted. RLHZ Ln / Lznz / E / Yli Figure 3 is a block diagram of another example of a System 300 for the rapid detection of valid gene fusions. System 300 performs the same functions as System 100, except that System 300 uses a sequencing device 110 to generate RNA (or DNA) sequence reads 112, uses a secondary analysis unit 130 to align the RNA sequence reads 112 with a reference sequence, uses a fusion candidate identification module 140 to identify fusion candidates, uses a fusion candidate filtering module 150 to determine a filtered set of fusion candidates for subsequent analysis, and then performs subsequent analysis of the filtered set of fusion candidates to identify valid gene fusions using a feature set generation module 160, a machine learning model 170, a gene fusion determination module 190, and an output API module 190.Each of these functional units, modules, or models performs the same functions as those attributed in the system description 100 of Figure 1. The difference between System 300 and System 100 is that fusion candidate identification, fusion candidate filtering, and subsequent analysis of the filtered set of fusion candidates are performed on a different computer, 320, and not within the sequencing device, 110. Consequently, the differences between System 300 and System 100 lie in how aligned reads are packaged and communicated to the computer, 320, for gene fusion analysis using the network, 310, and unpacked by the computer, and how the gene fusion results are packaged and transmitted to another device with a corresponding display for output. In more detail, the sequencing device 110 can sequence the biological sample 105 and generate RNA reads 112-1, 112-2, 112-n, where “n” is any positive integer greater than 0 as described with reference to system 100. Although RNA reads are used as an example, the system can also perform the same processes on DNA reads. The sequencing device 110 can store the reads 112-1, 112-2, 112-n in memory 120. In some implementations, the reads 112-1, 112-2, 112-n may be in a compressed format. The secondary analysis unit 130 can obtain reads 112-1, 112-2, 112-n and store reads 112-1, 112-2, 112-n in memory 132 of the secondary analysis unit 130. In some implementations, this may include a sequencing device control program 110 that transmits reads 112-1, 112-2, 112-n to memory 132 of the secondary analysis unit 130. In other implementations, the secondary analysis unit 130 may request reads 112-1, 112-2, 112-n. If reads 112-1, 112-2, 112-n are compressed, the programmable logic device 134 of the secondary analysis unit 130 can be configured in state B as a unit of RLHZ Ln / Lznz / E / Yli decompression 138 and be used to decompress readings 112-1, 112-2, 112-n. Afterward, programmable logic device 134 can be reconfigured in state A as a reading alignment unit and used to align readings 112-1, 112-2, 112-n with a reference sequence. The secondary analysis unit 130 can be reconfigured to state B as a compression unit and used to compress the aligned reads in preparation for transmission to the computer 320. In this example, compressing the first batch of aligned reads includes compressing not only the aligned reads but also the data generated by the read alignment unit 136 related to the aligned reads that will be used for gene fusion analysis. This data is described with reference to system 100 in Figure 1 and may include, for example, a variant allele frequency count, a count of unique read alignments, read coverage across the transcript, a MAPQ score, data indicating homology between precursor genes, or a combination thereof.In addition, other data that can be compressed into the first batch of aligned reads may include (i) the fusion candidate reads, (ii) the portion of the reference sequence locations to which the fusion candidate reads were aligned, and (iii) annotations of the reference genome segments to which the particular gene fusion candidate was aligned. In some implementations, the annotations may include gene exon annotations, annotations indicating the presence of homologous genes, annotations indicating a list of enriched genes, or a combination thereof. After compressing the aligned reads, the secondary analysis unit 130 can store the first batch of compressed reads in memory 120. The sequencing device 110 can then transmit the first batch 125 of aligned reads to the computer 320 over the network 310 for gene fusion analysis. The network 310 can include one or more wired networks, one or more wireless networks, or a combination thereof. In different implementations, the network 310 can be one or more of a wired Ethernet network, a wired optical network, a LAN, a WAN, a cellular network, the Internet, or a combination thereof. In some implementations, the computer 320 can be a remote cloud server. However, in other implementations, the computer 320 can connect to the sequencing device 110 via a direct connection, such as a direct Ethernet connection, a USB-C connection, or similar.Although the first batch of reads is compressed before communication in this example in Figure 300, there is no need to use compression. Instead, compression is provided as a method to reduce network bandwidth consumption and minimize storage costs, which can provide significant technological benefits and reduced costs when dealing with large genome datasets. RLHZ Ln / Lznz / E / Yli In some implementations, the first batch of aligned reads includes the entire set of reads generated for sample 105. In other implementations, the first batch of aligned reads is only a portion of the complete set of reads generated for sample 105, and a batch processing system can be used to facilitate parallel processing. For example, in some implementations, after the secondary analysis unit stores the first batch of aligned reads in memory 120, the secondary analysis unit 130 fetches a second batch of reads that are not yet aligned for storage in memory 132. The secondary analysis unit 130 can then perform decompression (if the second batch of reads was compressed) and alignment of the second batch of reads while computer 320 performs gene fusion analysis of the first batch of reads.Such parallel processing facilitated through batch processing of the readings can significantly reduce the system run time required to determine valid gene fusions for readings from a sample. Computer 320 can receive the first batch of reads 125 via network 310 and store the first batch of reads in memory 320. If the first batch of reads 125 is compressed, computer 320 can use compression / decompression module 325 to decompress the first batch of reads and store the first batch of reads in memory 320. Computer 320 can then execute the gene fusion analysis pipeline of fusion candidate identification module 140, fusion candidate filtering module 150, feature set generation module 160, machine learning model 170, gene fusion determination module 180, and output API module 190 in the same manner as described with reference to system 100 in Figure 1. The results 192 can be provided to several different devices via the 310 network. For example, the output data can be transmitted to the sequencing device for display on a sequencer screen 195. Alternatively, or in addition, the results 192 can be provided for display on a user device screen 330 via the 310 network. The user device 330 can include a smartphone, tablet, laptop, desktop computer, or any other computer with a screen. Alternatively, or in addition, the results 192 can also be provided for output via a printer 340 via the 310 network. In such implementations, the output can be a hard copy report of the determined valid gene fusions. Figure 4 is a block diagram of system components that can be used to implement a system for rapid detection of gene fusions. RLHZ Ln / Lznz / E / Yli The Computing Device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The Computing Device 450 is intended to represent various forms of mobile devices, such as personal digital assistants, cell phones, smartphones, and other similar computing devices. In addition, the Computing Device 400 or 450 may include Universal Serial Bus (USB) flash drives. USB flash drives can store operating systems and other applications. USB flash drives may include input / output components, such as a wireless transmitter or USB connector that can be inserted into a USB port on another computing device.The components shown here, their connections and relationships, and their functions, are intended only as examples, and are not intended to limit the implementations of the inventions described and / or claimed in this document. The Computing Device 400 includes a processor 402, memory 404, a storage device 406, a high-speed interface 408 connected to memory 404 and high-speed expansion ports 410, and a low-speed interface 412 connected to the low-speed bus 414 and storage device 408. Each of the components 402, 404, 406, 408, 410, and 412 is interconnected using various buses and can be mounted on a common motherboard or in other ways as appropriate. The processor 402 can process instructions for execution on the Computing Device 400, including instructions stored in memory 404 or storage device 408 to display graphical information for a GUI on an external input / output device, such as the display 416 connected to the high-speed interface 408.In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and memory types. Furthermore, multiple 400 computing devices can be connected, with each device providing portions of the necessary operations, for example, as a server bank, a group of blade servers, or a multi-processor system. Memory 404 stores information on the computer device 400. In one implementation, memory 404 is a volatile memory unit or units. In another implementation, memory 404 is a non-volatile memory unit or units. Memory 404 can also be another form of computer-readable media, such as a magnetic or optical disk. RLHZ Ln / Lznz / E / Yli The 408 storage device is capable of providing mass storage for the 400 computing device. In one implementation, the 408 storage device may be or contain a computer-readable medium, such as a floppy disk drive, hard disk drive, optical disk drive, or tape drive, flash memory or other similar solid-state memory device, or a set of devices, including devices in a storage area network or other configurations. A computer program product may be tangibly embodied in a data carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The data carrier is a computer-readable or machine-readable medium, such as the 404 memory, the 408 storage device, or the memory in the 402 processor.The high-speed controller 408 handles bandwidth-intensive operations for the computing device 400, while the low-speed controller 412 handles lower-bandwidth-intensive operations. This role assignment is just one example. In one implementation, the high-speed controller 408 is coupled to memory 404, display 416, for example, via a graphics processor or accelerator, and to high-speed expansion ports 410, which can accommodate various expansion cards (not shown). In another implementation, the low-speed controller 412 is coupled to storage device 408 and low-speed expansion port 414.The low-speed expansion port, which can include various communication ports, such as USB, Bluetooth, Ethernet, or wireless Ethernet, can be connected to one or more input / output devices, such as a keyboard, pointing device, microphone / speaker pair, scanner, or a network device such as a switch or router, for example, via a network adapter. The Computing Device 400 can be deployed in a number of different ways, as shown in the figure. For example, it can be deployed as a standard server (420), or multiple times in a group of such servers. It can also be deployed as part of a rack server system (424). Furthermore, it can be deployed in a personal computer, such as a laptop (422). Alternatively, the components of the Computing Device 400 can be combined with other components in a mobile device (not shown), such as the Device 450.Each of such devices may contain one or more of a 400, 450 computing device, and an entire system may be composed of multiple 400, 450 computing devices that communicate with each other. The 400 computing device can be deployed in a number of different ways, as shown in the figure. For example, it can be deployed as a standard 420 server, or multiple times in a group of such servers. Additionally, it can be deployed as part of a system of RLHZ LO / L7R7 / E / YI rack server 424. In addition, it can be deployed in a personal computer, such as a laptop 422. Alternatively, the components of the computing device 400 can be combined with other components in a mobile device (not shown), such as the device 450. Each of such devices can contain one or more of a computing device 400, 450, and an entire system can be composed of multiple computing devices 400, 450 communicating with each other. The 450 computer device includes a 452 processor, 464 memory, and an input / output device, such as a 454 display, a 466 communication interface, and a 468 transceiver, among other components. The 450 device can also be provided with a storage device, such as a micro disk drive or other device, to provide additional storage. Each of the 450, 452, 464, 454, 466, and 468 components is interconnected using various buses, and several of the components can be mounted on a common motherboard or in other ways, as appropriate. The 452 processor can execute instructions on the 450 computing device, including instructions stored in memory 464. The processor can be implemented as a chipset comprising separate and multiple analog and digital processors. Additionally, the processor can be implemented using any of several architectures. For example, the 452 processor can be a CISC (Complex Instruction Set Computer), a RISC (Reduced Instruction Set Computer), or a MISC (Minimal Instruction Set Computer) processor. The processor can provide, for example, coordination for other components of the 450 device, such as controlling user interfaces, running applications on the 450 device, and enabling wireless communication with the 450 device. The processor 452 can communicate with a user through the control interface 458 and the display interface 456 coupled to a display 454. The display 454 can be, for example, a TFT (thin-film transistor liquid crystal display) or an OLED (organic light-emitting diode) display, or other suitable display technology. The display interface 456 can comprise appropriate circuitry for driving the display 454 to present graphical and other information to a user. The control interface 458 can receive commands from a user and convert them for transmission to the processor 452. In addition, an external interface 462 can be provided in communication with a processor 452 to enable near-area communication of the device 450 with other devices. The external interface 462 can provide, for example, wired communication in some RLHZ Ln / Lznz / E / Yli implementations, or for wireless communication in other implementations, and multiple interfaces can also be used. Memory 464 stores information on the computing device 450. Memory 464 can be implemented as one or more computer-readable media, a volatile memory unit or units, or a non-volatile memory unit or units. Expansion memory 474 can also be provided and connected to the device 650 via expansion interface 472, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such expansion memory 474 can provide additional storage space for the device 450, or it can store applications or other information for the device 450. Specifically, expansion memory 474 can include instructions for performing or completing the processes described above, and it can also include secure information.Therefore, for example, the 474 expansion memory can be provided as a security module for the 450 device, and can be programmed with instructions that enable the secure use of the 450 device. In addition, secure applications can be provided through SIMM cards, along with additional information, such as storing identification information on the SIMM card in a way that prevents its illegal extraction. Memory may include, for example, flash memory and / or non-volatile random-access memory (NVRAM), as described later. In an implementation, a computer program product is tangibly embodied in a data carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The data carrier is a computer-readable medium, such as memory 464, expansion memory 474, or memory in the processor 452, which may be received, for example, through the transceiver 468 or external interface 462. The 450 device can communicate wirelessly via the 466 communication interface, which may include a digital signal processing circuit when required. The 466 communication interface can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication can occur, for example, via the 468 radio frequency transceiver. In addition, short-range communication, such as Bluetooth, Wi-Fi, or another transceiver of this type (not shown), can occur. Furthermore, the 470 GPS (Global Positioning System) receiver module can RLHZ Ln / Lznz / E / Ylj provide additional wireless navigation and location-related data to the 450 device, which can be used as appropriate by applications running on the 450 device. The 450 device can also communicate audibly using an audio codec 460, which can receive spoken information from a user and convert it into usable digital information. The audio codec 460 can likewise generate audible sound for a user, such as through a speaker, for example, on a mobile phone connected to the 450 device. This sound can include audio from voice calls, recorded audio (such as voicemails, music files, etc.), and audio generated by applications running on the 450 device. The computing device 450 can be implemented in a number of different ways, as shown in the figure. For example, it can be implemented as a cell phone 480. It can also be implemented as part of a smartphone 482, personal digital assistant, or other similar mobile device. Various implementations of the systems and methods described herein can be carried out in a digital electronic circuit, integrated circuit, specially designed ASIC (application-specific integrated circuit), computer hardware, firmware, software, and / or combinations of such implementations. These various implementations may include implementation in one or more computer programs that are executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer product, apparatus, and / or device, such as magnetic disks, optical disks, memory, and programmable logic devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. RLHZ Ln / Lznz / E / Yli “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. To facilitate user interaction, the systems and techniques described here can be implemented on a computer that has a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, to show information to the user, and a keyboard and pointing device, such as a mouse or scroll ball, through which the user can provide input data to the computer. Other types of devices can be used to enable user interaction; for example, the feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and user input can be received in any form, including acoustic, voice, or tactile input. The systems and techniques described here can be implemented in a computer system that includes a management component, such as a data server, or a middleware component, such as an application server, or a user interface component, such as a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such management, middleware, or user interface components. The system components can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet. A computer system can include clients and servers. A client and a server are generally located remotely and interact, typically, through a communication network. The relationship between the client and the server arises from computer programs running on their respective computers, which have a client-server relationship with each other. Other modalities Several embodiments have been described. However, it is understood that various modifications may be made without departing from the spirit and scope of the invention. Furthermore, the logical flows represented in the figures do not require the particular order shown, or sequential order, to achieve the desired results. In addition, other steps may be provided, or steps may be eliminated. ML / a / ZUZl / U1 ZU I and the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method for identifying one or more gene fusions in a biological sample, the method comprising: obtaining, by means of one or more computers, first data representing a plurality of reads aligned from a read alignment unit; identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained; filtering, by means of one or more computers, the plurality of gene fusion candidates to determine a filtered set of gene fusion candidates;for each particular gene fusion candidate from the filtered set of gene fusion candidates: generate, by means of one or more computers, input data for input to a machine learning model, characterized in that generating the input data comprises extracting feature data to represent the particular gene fusion candidate from data that include: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit;to provide, by means of one or more computers, the generated input data as an input to the machine learning model, wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion based on the machine learning model processing input data representing (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit; to obtain, by means of one or more computers, output data generated by the machine learning model based on the machine learning model processing the generated input data;and determine, using one or more computers, whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data. RLHZ Ln / Lznz / E / Yli; 2. The method of claim 1, characterized in that generating the input data further comprises extracting feature data that includes annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit;and wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion candidate based on input data from the machine learning model's processing representing: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, (ii) annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit, and (iii) generated data based on the results of the read alignment unit.

3. The method of any one of the preceding claims, characterized in that identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of split-read alignments.

4. The method of any one of the preceding claims, characterized in that identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of non-coincident read pair alignments.

5. The method of any one of the preceding claims, characterized in that the read alignment unit is implemented using a set of one or more processing engines configured using hardware logic circuits physically arranged to perform operations, using hardware logic circuits, to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence locations, RLHZ Ln / Lznz / E / Yll, (iii) generate one or more alignment scores corresponding to each matching reference sequence location for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) produce data representing a candidate alignment for the first read.

6. The method of any one of claims 1-4, characterized in that the read alignment unit is implemented using a set of one or more processing engines using one or more central processing units (CPUs) or one or more graphics processing units (GPUs) to execute software instructions causing the one or more CPUs or one or more GPUs to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more locations in the reference sequence that are matched for the first read, (iii) generate one or more alignment scores corresponding to each location in the reference sequences that are matched for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores,and (v) produce data that represent a candidate alignment for the first read.

7. The method of any one of the preceding claims, the method further comprising: receiving, by means of the read alignment unit, a plurality of reads that are not yet aligned; aligning, by means of the read alignment unit, a first subset of the plurality of reads; and storing, by means of the read alignment unit, the first subset of aligned reads in a memory device; characterized in that obtaining, by means of one or more computers, first data representing a plurality of aligned reads from a read alignment unit comprises obtaining, by means of one or more computers, the first subset of aligned reads from the memory device and performing one or more of the operations of claim 1 while the read alignment unit aligns a second subset of the plurality of reads that are not yet aligned. RLHZ Ln / Lznz / E / Yli 8. The method of any one of the preceding claims, characterized in that the data generated based on the results of the read alignment unit include one or more of a variant allele frequency count, a count of unique read alignments, a read coverage across the transcript, a MAPQ score, or data indicating homology between precursor genes.

9. The method of any one of the preceding claims, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining, by means of one or more computers, whether the output data satisfy a predetermined threshold; and based on the determination that the output data satisfy the predetermined thresholds, determining that the particular fusion candidate corresponds to a valid gene fusion candidate.

10. The method of any one of the preceding claims, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining, by means of one or more computers, whether the output data satisfy a predetermined threshold; and based on the determination that the output data do not satisfy the predetermined thresholds, determining that the particular fusion candidate does not correspond to a valid gene fusion candidate.

11. A system for identifying one or more gene fusions in a biological sample comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: obtaining, by means of one or more computers, first data representing a plurality of reads aligned from a read alignment unit; identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained; filtering, by means of one or more computers, the plurality of gene fusion candidates to determine a filtered set of gene fusion candidates;for each particular gene fusion candidate from the filtered set of gene fusion candidates: RLHZ Ln / Lznz / E / Yli generate, by means of one or more computers, input data for input to a machine learning model, characterized in that generating the input data comprises extracting feature data to represent the particular gene fusion candidate from data that include: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit;to provide, by means of one or more computers, the generated input data as an input to the machine learning model, wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion based on the machine learning model processing input data representing (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit; to obtain, by means of one or more computers, output data generated by the machine learning model based on the machine learning model processing the generated input data;and determine, using one or more computers, whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data.

12. The system of claim 11, characterized in that generating the input data further comprises extracting feature data that includes annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit;and wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion candidate based on input data from the machine learning model's processing representing: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, RLHZ Ln / Lznz / E / Yli; (ii) annotation data describing annotations of the segments of the reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit; and (iii) generated data based on the results of the read alignment unit.

13. The system of any one of claims 11-12, characterized in that identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of split-read alignments.

14. The system of any one of claims 11-13, characterized in that identifying, by means of one or more computers, a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of non-coincident read pair alignments.

15. The system of any one of claims 11-14, characterized in that the read alignment unit is implemented using a set of one or more processing engines configured using hardware logic circuits physically arranged to perform operations, using hardware logic circuits, to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence locations, (iii) generate one or more alignment scores corresponding to each matching reference sequence location for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) produce data representing a candidate alignment for the first read.

16. The system of any one of claims 11-14, characterized in that the read alignment unit is implemented using a set of one or more processing engines using one or more central processing units (CPUs) or one or more graphics processing units (GPUs) to execute software instructions causing the one or more CPUs or one or more GPUs to: (i) receive data representing a first read, RLHZ Ln / Lznz / E / Yli; (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more locations in the reference sequence that are matched for the first read; (iii) generate one or more alignment scores corresponding to each location in the reference sequences that are matched for the first read; (iv) select one or more candidate alignments for the first read based on the one or more alignment scores;and (v) produce data that represent a candidate alignment for the first read.

17. The system of any one of claims 11-16, the operations further comprising: receiving, by means of the read alignment unit, a plurality of reads that are not yet aligned; aligning, by means of the read alignment unit, a first subset of the plurality of reads; and storing, by means of the read alignment unit, the first subset of aligned reads in a memory device; characterized in that obtaining, by means of one or more computers, the first data representing a plurality of aligned reads from a read alignment unit comprises obtaining, by means of one or more computers, the first subset of aligned reads from the memory device and performing one or more of the operations of claim 11 while the read alignment unit aligns a second subset of the plurality of reads that are not yet aligned.

18. The system of any one of claims 11-17, characterized in that the generated data based on the results of the read alignment unit include one or more of a variant allele frequency count, a count of unique read alignments, a read coverage across the transcript, a MAPQ score, or data indicating homology between precursor genes.

19. The system of any one of claims 11-18, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining, by means of one or more computers, whether the output data satisfy a predetermined threshold; and RLHZ Ln / Lznz / E / Yli based on the determination that the output data satisfy the predetermined thresholds, determining that the particular fusion candidate corresponds to a valid gene fusion candidate.

20. The system of any one of claims 11-19, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining, by means of one or more computers, whether the output data satisfy a predetermined threshold; and based on the determination that the output data do not satisfy the predetermined thresholds, determining that the particular fusion candidate does not correspond to a valid gene fusion candidate.

21. A non-transient, computer-readable medium that stores software comprising instructions executable by one or more computers, which, after such execution, cause the one or more computers to perform operations comprising: obtaining first data representing a plurality of reads aligned from a read alignment unit; identifying a plurality of gene fusion candidates included within the first data obtained; filtering the plurality of gene fusion candidates to determine a filtered set of gene fusion candidates;for each particular gene fusion candidate from the filtered set of gene fusion candidates: generate input data for input to a machine learning model, characterized in that generating the input data comprises extracting feature data to represent the particular gene fusion candidate from data that include: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit;to provide the generated input data as an input to the machine learning model, wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion based on the machine learning model processing input data representing (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, and (ii) generated data based on the results of the read alignment unit; to obtain output data generated by the machine learning model based on the machine learning model processing the generated input data; and to determine whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data.

22. The computer-readable means of claim 21, characterized in that generating the input data further comprises extracting feature data including annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit;and wherein the machine learning model has been trained to generate output data representing a probability that a gene fusion candidate is a valid gene fusion candidate based on input data from the machine learning model's processing representing: (i) one or more segments of a reference sequence to which the particular gene fusion candidate was aligned by the read alignment unit, (ii) annotation data describing annotations of the reference sequence segments to which the particular gene fusion candidate was aligned by the read alignment unit, and (iii) generated data based on the results of the read alignment unit.

23. The computer-readable means of any one of claims 21-22, characterized in that identifying a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of split-read alignments.

24. The computer-readable means of any one of claims 21-23, characterized in that identifying a plurality of gene fusion candidates included within the first data obtained comprises identifying, by means of one or more computers, a plurality of non-coincident read-pair alignments. RLHZ Ln / Lznz / E / Yli 25. The computer-readable medium of any one of claims 21-24, characterized in that the read alignment unit is implemented using a set of one or more processing engines configured using hardware logic circuits physically arranged to perform operations, using the hardware logic circuits, to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence locations, (iii) generate one or more alignment scores corresponding to each matching reference sequence location for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) produce data representing a candidate alignment for the first read.

26. The computer-readable medium of any one of claims 21-24, characterized in that the read alignment unit is implemented using a set of one or more processing engines using one or more central processing units (CPUs) or one or more graphics processing units (GPUs) to execute software instructions causing the one or more CPUs or one or more GPUs to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more locations in the reference sequence that are matched for the first read, (iii) generate one or more alignment scores corresponding to each location in the reference sequences that are matched for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores,and (v) produce data that represent a candidate alignment for the first read.

27. The computer-readable medium of any one of claims 21-26, the operations further comprising: receiving, by means of the read alignment unit, a plurality of reads that are not yet aligned; aligning, by means of the read alignment unit, a first subset of the plurality of reads; and storing, by means of the read alignment unit, the first subset of aligned reads in a memory device; characterized in that obtaining first data representing a plurality of aligned reads from a read alignment unit comprises obtaining the first subset of aligned reads from the memory device and performing one or more of the operations of claim 21 while the read alignment unit aligns a second subset of the plurality of reads that are not yet aligned.

28. The computer-readable medium of any one of claims 21-27, characterized in that the generated data based on the results of the read alignment unit include one or more of a variant allele frequency count, a count of unique read alignments, a read coverage across the transcript, a MAPQ score, or data indicating homology between the precursor genes.

29. The computer-readable means of any one of claims 21-28, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining whether the output data satisfies a predetermined threshold; and based on the determination that the output data satisfies the predetermined thresholds, determining that the particular fusion candidate corresponds to a valid gene fusion candidate.

30. The computer-readable means of any one of claims 21-29, characterized in that determining whether the particular fusion candidate corresponds to a valid gene fusion candidate based on the output data comprises: determining whether the output data satisfies a predetermined threshold; and based on the determination that the output data does not satisfy the predetermined thresholds, determining that the particular fusion candidate does not correspond to a valid gene fusion candidate.