Rapid detection of gene fusions
The gene fusion detection method implemented by computers uses read segment comparison units and machine learning models to quickly identify and screen gene fusion candidates, solving the problems of low detection efficiency and high resource consumption in the existing technology, and achieving efficient and accurate gene fusion detection.
Patent Information
- Application Number
- CN202080021779.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-05
- Filing Date
- 2020-12-04
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-12-04
AI Technical Summary
The prior art is difficult to detect gene fusion quickly and accurately, especially in the diagnosis and treatment of diseases such as cancer, which have problems such as low detection efficiency and high resource consumption.
Through a computer-implemented method, the data of multiple aligned reads is obtained using the read segment alignment unit, the gene fusion candidates are identified, the candidates are screened, the input data is generated, and the gene fusion candidates are inputted into the machine learning model to determine the effectiveness of the gene fusion candidates.
Rapid detection gene fusion is achieved, reducing the consumption of processing resources, shortening the running time, and improving the accuracy of detection.
Smart Images

Figure CN113574603B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 944,304, filed December 5, 2019, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to genetic testing, and more particularly to systems and methods for rapid detection of gene fusions. Background Art
[0004] Gene fusions can be driving factors for cancer and are therefore important diagnostic and therapeutic targets in the treatment of diseases such as cancer. Summary of the invention
[0005] According to an innovative aspect of the present disclosure, a computer-implemented method for identifying one or more gene fusions in a biological sample is disclosed. In one aspect, the method may include the following actions: obtaining, by one or more computers, first data representing a plurality of aligned reads from a read alignment unit, identifying, by one or more computers, a plurality of gene fusion candidates included in the obtained first data, screening, by one or more computers, the plurality of gene fusion candidates to determine a screened set of gene fusion candidates, and for each specific gene fusion candidate in the screened set of gene fusion candidates: generating, by one or more computers, input data for input into a machine learning model, wherein generating the input data includes extracting feature data from the data to represent the specific gene fusion candidate, the data including: (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) based on the read alignment unit, The data generated by the output of the segment alignment unit is provided by one or more computers as input to the machine learning model, wherein the machine learning model has been trained to generate output data representing the possibility that the gene fusion candidate is a valid gene fusion based on the input data processed by the machine learning model, and the input data represents (i) one or more segments of the reference sequence to which the read segment alignment unit aligns the specific gene fusion candidate, and (ii) based on the data generated by the output of the read segment alignment unit, one or more computers obtain the output data generated by the machine learning model based on the input data processed by the machine learning model, and the one or more computers determine whether the specific fusion candidate corresponds to a valid gene fusion candidate based on the output data.
[0006] Other versions include corresponding systems, apparatus, and computer programs that perform the actions of the method defined by instructions encoded on a computer-readable storage device.
[0007] These and other versions may optionally include one or more of the following features. For example, in some embodiments, generating input data also includes extracting feature data, the feature data including annotation data, the annotation data describing annotations of segments of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate. In such embodiments, the machine learning model has been trained to generate output data representing the likelihood that the gene fusion candidate is a valid gene fusion candidate based on processing the input data by the machine learning model, the input data representing: (i) one or more segments of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, (ii) annotation data describing annotations of segments of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (iii) data generated based on the output of the read alignment unit.
[0008] In some implementations, identifying, by one or more computers, a plurality of gene fusion candidates included in the obtained first data may include identifying, by one or more computers, a plurality of segmented read alignments.
[0009] In some implementations, identifying, by one or more computers, a plurality of gene fusion candidates included in the obtained first data includes identifying, by one or more computers, a plurality of discordant read alignment pairs.
[0010] In some embodiments, a read alignment unit is implemented using a set of one or more processing engines configured using hardware logic circuitry that is physically arranged to perform operations using the hardware logic circuitry to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions, (iii) generate one or more alignment scores corresponding to each matching reference sequence position for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) output data representing candidate alignments for the first read.
[0011] In some embodiments, a read alignment unit is implemented using a set of one or more processing engines by using one or more central processing units (CPUs) or one or more graphics processing units (GPUs) to execute software instructions, which software instructions cause the one or more CPUs or the one or more GPUs to: (i) receive data representing a first read, (ii) map the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions, (iii) generate one or more alignment scores corresponding to each matching reference sequence position for the first read, (iv) select one or more candidate alignments for the first read based on the one or more alignment scores, and (v) output data representing the candidate alignments for the first read.
[0012] In some implementations, the method may further include receiving, by the read alignment unit, a plurality of reads that have not yet been aligned, aligning, by the read alignment unit, a first subset of the plurality of reads, and storing, by the read alignment unit, the aligned reads of the first subset in a memory device. In such implementations, obtaining, by the one or more computers, first data representing the plurality of aligned reads from the read alignment unit may include obtaining, by the one or more computers, the aligned reads of the first subset from the memory device, and performing, when the read alignment unit aligns a second subset of the plurality of reads that have not yet been aligned, one or more operations of the computer-implemented method for identifying one or more gene fusions in a biological sample.
[0013] In some implementations, the data generated based on the output of the read alignment unit may include any one or more of variant allele frequency counts, unique read alignment counts, read coverage across transcripts, MAPQ scores, or data indicating homology between parental genes.
[0014] In some embodiments, determining whether a particular fusion candidate corresponds to a valid gene fusion candidate based on the output data may include determining, by one or more computers, whether the output data satisfies a predetermined threshold, and determining that the particular fusion candidate corresponds to a valid gene fusion candidate based on determining that the output data satisfies the predetermined threshold.
[0015] In some embodiments, determining whether a particular fusion candidate corresponds to a valid gene fusion candidate based on the output data may include determining, by one or more computers, whether the output data satisfies a predetermined threshold, and determining that the particular fusion candidate does not correspond to a valid gene fusion candidate based on determining that the output data does not satisfy the predetermined threshold.
[0016] These and other innovative aspects of the disclosure will be apparent from the detailed description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram of an example of a system for rapidly detecting valid gene fusions.
[0018] Figure 2 is a flow chart of an example of a process for performing rapid detection of valid gene fusions.
[0019] Figure 3 is a block diagram of another example of a system for rapid detection of valid gene fusions.
[0020] Figure 4 is a block diagram of system components that may be used to implement a system for rapid detection of valid gene fusions. DETAILED DESCRIPTION
[0021] The present disclosure relates to systems, methods, devices, computer programs, or any combination thereof for rapid detection of gene fusions. The presence of certain gene fusions may be an important indicator of a particular disease, an indicator that a particular treatment is recommended for a particular disease, or a combination thereof. For example, certain gene fusions may be indicators of specific types of cancer, such as acute and chronic myeloid leukemia, myelodysplastic syndrome (MDS), soft tissue sarcoma, or their treatment. The present disclosure uses a screening engine to reduce the number of gene fusion candidates (also referred to herein as "fusion candidates") that are processed to determine whether each fusion candidate is a valid gene fusion, thereby enabling rapid detection of accurate gene fusions. The screening engine enables high-accuracy selection of candidate fusions for subsequent analysis, while also achieving a reduction in the computational resources that need to be consumed in order to identify valid gene fusions, because only a subset of the screened candidate gene fusions can be used for further downstream processing as described herein.
[0022] The reduced set of candidate gene fusions also provides other technical advantages. For example, compared to conventional methods that process all gene fusion candidates and perform scoring, the methods and systems disclosed in the present invention provide shortened runtimes. Shortening the runtime of executing operations also directly leads to reduced consumption of processing resources (e.g., CPU or GPU resources), memory usage, and power consumption. Although the screening engine provides a shortened runtime compared to conventional methods, the methods and systems disclosed in the present invention can also provide other ways to shorten runtimes. For example, in some specific implementations, further shortening of runtimes can be achieved by using a hardware-accelerated read alignment unit to perform mapping, alignment, and generation of metadata for processing candidate gene fusions.
[0023] Figure 11 is a block diagram of an example of a system 100 for rapidly detecting effective gene fusions. The system 100 may include a nucleic acid sequencing device 110, a memory 120, a secondary analysis unit 130, a fusion candidate identification module 140, a fusion candidate screening module 150, a feature set generation module 160, a machine learning model 170, a gene fusion determination module 180, an output application program interface (API) module 190, and an output display 195. Figure 1 In the example of , each of these components is described as being implemented within the nucleic acid sequencing device 110. However, the present disclosure is not limited to such embodiments.
[0024] On the contrary, in some specific implementations, Figure 1 One or more of the components described in the above may be executed on a computer external to the nucleic acid sequencing device 110. For example, in some specific implementations, the secondary analysis module may be implemented within the nucleic acid sequencing device 110, and the fusion candidate identification module 140, the fusion candidate screening module 150, the feature set generation module 160, the machine learning model 170, the gene fusion determination module 180, and the output application program interface (API) module 190 may be implemented in one or more different computers. In such specific implementations, the one or more different computers and the nucleic acid sequencing device may be communicatively coupled using one or more wired networks, one or more wireless networks, or a combination thereof.
[0025] For the purposes of this specification, the term "module" includes one or more software components, one or more hardware components, or any combination thereof that can be used to implement the functions attributed to the corresponding module by this specification. Generally speaking, a "module" as described herein uses one or more processors to execute software instructions to implement the functions of the module described herein. The processor may include a central processing unit (CPU), a graphics processing unit (GPU), etc.
[0026] Likewise, the term "unit" as used in this specification includes one or more software components, one or more hardware components, or any combination thereof that can be used to implement the functions attributed to the corresponding unit by this specification. Generally speaking, as described herein, a "unit" uses one or more hardware components such as hardwired digital logic gates or hardwired digital logic blocks arranged as a processing engine to perform operations that implement the functions of the unit described herein. Such hardwired digital logic gates or hardwired digital logic circuits may include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.
[0027] Nucleic acid sequencing device 110 (also referred to herein as sequencing device 110) is configured to perform primary nucleic acid sequence analysis. Performing preliminary analysis may include receiving biological samples 105 such as blood samples, tissue samples, sputum or nucleic acid samples by sequencing device 110, and generating output data such as one or more reads 112 by sequencing device 110, each read representing the nucleotide sequence of the nucleic acid sequence of the received biological sample. In some specific implementations, sequencing performed by nucleic acid sequencer 110 may be performed in multiple read segments cycles, wherein the first read segment cycle "read segment 1" generates one or more first read segments representing the nucleotide sequence from the first end of the nucleic acid sequence fragment, and the second read segment cycle "read segment 2" generates one or more second read segments representing the nucleotide sequence from the other end of one of the nucleic acid sequence fragments. In some specific implementations, the read segment may be a short read segment of about 80 to 120 nucleotides in length. However, the present disclosure is not limited to read segments of any specific nucleotide length. On the contrary, the present disclosure can be used for read segments of any nucleotide length.
[0028] In some embodiments, the biological sample 105 may include a DNA sample, and the nucleic acid sequencer 110 may include a DNA sequencer. In such embodiments, the order of ordered nucleotides in the reads generated by the nucleic acid sequencer may include one or more of guanine (G), cytosine (C), adenine (A), and thymine (T) in any combination. In some embodiments, the nucleic acid sequencer 110 may be used to generate RNA reads of the biological sample 105. In such embodiments, this may occur using an RNA-seq protocol. By way of example, the biological sample 105 may be pre-treated using reverse transcription to form complementary DNA (cDNA) using a reverse transcriptase. In other embodiments, the nucleic acid sequencer 110 may include an RNA sequencer, and the biological sample may include an RNA sample. RNA reads generated using cDNA or via an RNA sequencer may be composed of C, G, A, and uracil (U). The generation and analysis of reference RNA reads described herein are described. Figure 1 However, the present invention can be used to generate and analyze any type of nucleic acid sequence reads, including DNA or RNA reads.
[0029] The sequencing device 110 may include a next generation sequencer (NGS) configured to generate sequence reads of a given sample such as reads 112-1, 112-2, 112-n, where "n" is any positive integer greater than 0, by using massively parallel sequencing technology in a manner that achieves ultra-high throughput, scalability, and speed. NGS enables rapid sequencing of entire genomes, the ability to scale up to target regions for deep sequencing, the discovery of novel RNA variants and splice sites using RNA sequencing (RNA-Seq), or quantification of mRNA for gene expression analysis, analysis of epigenetic factors such as genome-wide DNA methylation and DNA-protein interactions, sequencing of cancer samples to study rare somatic variants and tumor subclones, and the study of, for example, microbial diversity in humans or the environment.
[0030] The sequencing device 110 can sequence the biological sample 105 and generate a corresponding set of reads represented by A, C, T, and G. The sequencing device can then perform reverse transcription to generate a cDNA sequence representing the corresponding RNA sequence. These RNA sequence reads 112-1, 112-2, 112-n are output by the sequencing device 110 and stored in the memory device 120. In some specific implementations, before the RNA sequence reads 112-1, 112-2, 112-n are stored in the memory device 120, the reads 112-1, 112-2, 112-n can be compressed into data records of smaller size. The memory device 120 can be composed of Figure 1 Each of the components of the system 100 includes a secondary analysis unit 130, a fusion candidate identification module 140, a fusion candidate screening module 150, a feature set generation module 160, a machine learning model 170, a gene fusion determination module 180, and an output API module 190. Although the corresponding modules may be depicted as providing the output of the first module to the second module, an actual specific implementation of such features may include the first module storing the output in a memory device such as the memory 120, and the second module accessing the stored output from the memory device and processing the accessed output as the input of the second module.
[0031] The secondary analysis unit 130 may access the reads 112-1, 112-2, 112-n stored in the memory device 120 and perform one or more secondary analysis operations on the reads 112-1, 112-2, 112-n. In some specific implementations, the reads 112-1, 112-2, 112-n may be stored in the memory device 120 in the form of compressed data records. In such specific implementations, the secondary analysis unit may perform a decompression operation on the compressed read records before performing the secondary analysis operation on the read records. The secondary analysis operation may include mapping one or more reads to a reference genome, aligning one or more reads with the reference genome, or both. In some specific implementations, the secondary analysis operation may also include a variant calling operation. In addition to performing the secondary analysis operation, the secondary analysis unit 130 may also be configured to perform a sorting operation. The sorting operation may include, for example, sorting the aligned reads based on the positions in the reference genome to which the reads aligned by the secondary analysis unit are mapped.
[0032] In some specific implementations, such as Figure 1 In an example of, the secondary analysis unit 130 may include a memory 132 and a programmable logic device 134. The programmable logic device 134 may have a hardware logic circuit that may be dynamically configured to include one or more secondary analysis operation units (such as a read segment alignment unit 136), and may be used to perform one or more secondary analysis operations using the hardware logic circuit. Dynamically configuring the programmable logic device 134 to include a secondary analysis operation unit (such as a read segment alignment unit 136) may include, for example, providing one or more instructions to the programmable logic device 134, the one or more instructions causing the programmable logic device 134 to arrange the hardware logic gates of the programmable logic device 134 into a hardwired digital logic configuration that is configured to implement the functionality of the read segment alignment unit 136 in the hardware logic.
[0033] The one or more operations that trigger the dynamic configuration of the programmable logic device 134 may include compiled hardware description language code, one or more instructions for causing the programmable logic device 134 to configure itself based on the compiled hardware description language code, etc. Such operations that trigger the dynamic configuration of the programmable logic device 134 may be generated and deployed to the programmable logic device 134 by a control program executed by the sequencing device 110 or other computer hosting the control program. In some specific implementations, the control program may be a software module whose instructions reside in a memory device such as the memory 120. The function of the control program to generate and deploy instruction hardware description language code or other instructions to configure the programmable logic device 134 may be implemented by executing the control program software module using one or more processors such as one or more CPUs or one or more GPUs.
[0034] The functions of the read segment alignment unit 136 may include: obtaining one or more first read segments such as RNA read segments 112-1, 112-2, 112-n stored in the memory 120 by the sequencing device 110, mapping the obtained first read segments 112-1, 112-2, 112-n to one or more reference sequence positions of a reference sequence, and then aligning the mapped first read segments 112-1, 112-2, 112-n with the reference sequence. That is, the mapping stage can identify a set of candidate reference sequence positions that match the specific read segment for each specific read segment of the obtained first read segment. The alignment stage can then score each of the candidate reference sequence positions and select the specific reference sequence position with the highest alignment score as the correct alignment for the specific read segment. The reference sequence may include an organized series of nucleotides corresponding to a known genome.
[0035] In response to one or more instructions from the control program, the programmable Logical Devices The hardware logic gates of 134 may include configured logic gates (such as AND gates, OR gates, NOR gates, XOR gates, or any combination thereof) to perform digital logic functions of the read segment comparison unit 136. Alternatively or in addition, the arrangement of the hardware logic gates may include dynamically configured logic blocks, which include customizable hardware logic units to perform complex calculation operations including addition, multiplication, comparison, etc. The precise arrangement of the hardware logic gates, logic blocks, or combinations thereof is defined by instructions received from the control program. The received instructions may include compiled hardware description language (HDL) program code or be derived from compiled HDL program code, which is written by the entity and defines the schematic layout of the secondary analysis operation unit to be programmed into the programmable logic device 134. The HDL program code may include program code written in a language such as Very High Speed Integrated Circuit Hardware Description Language (VHDL), Verilog, etc. The entity may include one or more human users who draft the HDL program code, one or more artificial intelligence agents that generate the HDL program code, or a combination thereof.
[0036] The programmable logic device 134 may include any type of programmable logic device. For example, the programmable logic device 134 may include one or more field programmable gate arrays (FPGAs), one or more complex programmable logic devices (CPLDs), or one or more programmable logic arrays (PLAs), or a combination thereof, which may be dynamically configured and reconfigured by a control program as needed to perform a specific workflow. For example, in some implementations, it may be advantageous to use the programmable logic device 134 as a read segment alignment unit 136, as described above. However, in other implementations, it may be advantageous to use the programmable logic device 134 to perform variant calling functions or functions that support variant calling, such as a hidden Markov model (HMM) unit. In other implementations, the programmable logic device 134 may also be dynamically configured to support general computing tasks, such as compression and decompression, because the hardware logic of the programmable logic device 134 is able to perform these tasks and other tasks mentioned above much faster than using software instructions executed by one or more processing units 150 to perform the same tasks. In some implementations, the programmable logic device 134 may be dynamically reconfigured during runtime to perform different operations.
[0037] By way of example, in some implementations, the programmable logic device 134 may be implemented using an FPGA that is dynamically configured as a decompression unit to access data representing a compressed version of the first reads 112-1, 112-2, 112-n stored in the memory device 120 or 132. The secondary analysis unit 130 may use the decompression unit to decompress the compressed data representing the first reads 112-1, 112-2, 112-n (e.g., if the reads received from the nucleic acid sequencer are compressed). The decompression unit may store the decompressed reads in the memory 120 or 132. In such implementations, the FPGA may then be dynamically reconfigured as a read alignment unit 136 and used to perform mapping and alignment of the decompressed first reads 112-1, 112-2, 112-n now stored in the memory 132 or 120. The read alignment unit 136 may then store data representing the mapped and aligned reads in the memory 132 or 120. Although a series of operations are described including decompression and mapping and comparison operations, the present disclosure is not limited to performing these operations or only performing these operations. Instead, the programmable logic device 134 can be dynamically configured to perform the functions of any operating unit in any order as needed to achieve the functions described herein.
[0038] Figure 1The example of describes a secondary analysis unit 130 that implements a read alignment unit 136 using a hardware logic device in the form of a programmable logic device 134. However, the present invention is not limited to using a programmable logic device to implement the read alignment unit 136. Instead, other types of integrated circuits may be used to implement the read alignment unit 136 in the hardwired digital logic of the secondary analysis unit 130. For example, in some implementations, the secondary analysis unit 143 may be configured to use one or more application specific integrated circuits (ASICs) to implement the functionality of one or more secondary analysis operation units. Although not reprogrammable, one or more ASICs may be designed with customized hardware logic of one or more secondary analysis operation units (such as the read alignment unit 136, variant calling unit, variant calling computational support unit, etc.) to accelerate and parallelize the execution of secondary analysis operations. In some implementations, using one or more ASICs as a hardwired logic circuit of the secondary analysis unit 130 that implements the functionality of one or more secondary analysis operation units may be even faster than using a programmable logic device such as an FPGA. Therefore, a skilled person will understand that an ASIC may be used in place of a programmable logic device, such as an FPGA in any of the embodiments described herein. For specific implementations that are to employ ASICs, dedicated ASICs or dedicated logic groups within a single ASIC will be required for each secondary analysis operation unit to be performed by the ASIC. By way of example, one or more ASICs for read alignment, one or more ASICs for decompression, one or more ASICs for compression, or a combination thereof. Alternatively, dedicated logic groups within the same ASIC may be used to implement the same functionality.
[0039] In addition, reference Figure 1 and Figure 3 The examples of the present disclosure discussed in connection with systems 100 and 300 are described in conjunction with hardware implementations of a read alignment unit 136 in a programmable logic device, respectively. In addition, it is noted above that one or more ASICs may be used to implement a read alignment engine or other secondary analysis operation units. However, the present disclosure is not limited to the use of hardware units to implement such secondary analysis operations. Rather, in some implementations, any operations described herein as being performed by a programmable logic device, such as read alignment, compression, or decompression, may also be implemented using one or more software modules.
[0040] refer to Figure 1 , execution of the system 100 may begin with sequencing the biological sample 105 by the sequencing device 110. Sequencing the biological sample may include generating, by the sequencing device 110, a sequence of reads that are data representations of an ordered sequence of nucleotides present in the biological sample 105. If the system 100 is configured to process DNA reads, the reads generated by the sequencing device 110 may be stored in the memory 120.
[0041] Alternatively, in some implementations, if the system 100 is configured to process RNA reads, the sequencing device 110 can be configured to perform pre-processing of the biological sample 110 using reverse transcription to form complementary DNA (cDNA) using a reverse transcriptase. Figure 1 In the specific implementation of the example, the read segments generated by the sequencing device 110 include RNA read segments 112-1, 112-2, and 112-n. In other specific implementations, the nucleic acid sequencer 110 may include an RNA sequencer, and the biological sample may include an RNA sample. Regardless of whether the RNA read segments are generated by a DNA sequencing device using cDNA or via an RNA sequencer, the RNA read segments each include a nucleotide sequence consisting of C, G, A, and U. The read segments 112-1, 112-2, and 112-n can be stored in the memory 120 in a compressed or uncompressed format.
[0042] Execution of the system 100 may continue as the secondary analysis unit 130 obtains the reads 112-1, 112-2, 112-n stored in the memory 120. In some implementations, the secondary analysis unit 130 may access the reads 112-1, 112-2, 112-n in the memory device 120 and store the accessed reads 112-1, 112-2, 112-n in the memory 132 of the secondary analysis unit 130. In other implementations, the control program may load the reads 112-1, 112-2, 112-n into the memory 132 of the secondary analysis unit 130 when the control program determines that sequencing of the reads 112-1, 112-2, 112-n is complete and the secondary analysis unit 130 is available to perform secondary analysis operations.
[0043] If the read segments 112-1, 112-2, 112-n are compressed, the secondary analysis unit 130 may dynamically configure the programmable logic device 134 as a decompression unit to access the read segments 112-1, 112-2, 112-n in the memory 132 or 120, decompress the read segments 112-1, 112-2, 112-n, and then store the decompressed read segments 112-1, 112-2, 112-n in the memory 132 or 120. In some implementations, the secondary analysis unit may dynamically reconfigure the programmable logic device and perform decompression in response to instructions from the control program.
[0044] If the reads 112-1, 112-2, 122-n are not compressed, the secondary analysis unit 130 may access the reads from the memory 132 or 120 and perform a read alignment operation. In some implementations, the secondary analysis unit 130 may receive instructions from the control program that instruct the secondary analysis unit 130 to configure or reconfigure the programmable logic device 134 to include a read alignment unit 136 and then perform an alignment of the reads 112-1, 112-2, 112-n using the read alignment unit 136. Alternatively, in other implementations, the programmable logic device may have been configured to include the read alignment unit 136 and perform an alignment of the reads 112-1, 112-2, 112-n using the read alignment unit 136. In other implementations, the secondary analysis unit 130 may include an ASIC configured to perform read alignment, and then use the ASIC to perform alignment of the reads 112 - 1 , 112 - 2 , 112 - n .
[0045] The secondary analysis unit 130 may be configured to perform read alignment operations in parallel with the gene fusion analysis. For example, the secondary analysis unit 140 may obtain the first batch of unaligned reads generated by the sequencing device 110, use the read alignment unit 136 to align the first batch of reads, use a sorting engine, which may be implemented in the hardware configuration of the programming logic device 136 or implemented in software by executing program instructions to sort the aligned reads, and then output the first batch of aligned and sorted reads for storage in the memory device 132, 130. In some specific implementations, the memory 132 may be used as a local cache of the secondary analysis unit 132, which loads data to be processed by the read alignment unit and then unloads data that has been output by the read alignment unit 136. Therefore, once the read alignment unit 136 outputs the first batch of aligned reads to the memory 132, the first batch of aligned reads can be sorted and then output to the memory 120. Then, the fusion candidate identification module 140 can access the first batch of aligned and sorted reads from the memory 120 and begin processing the first batch of aligned and sorted reads while the secondary analysis unit 130 performs an alignment operation on a second batch of reads generated by the sequencing device 110 and not previously aligned. This process can be performed iteratively until the system 100 has processed each batch of reads. Although this example is described as having aligned and sorted batches, the present disclosure does not require that the batches of aligned reads are also sorted. Instead, the aligned and sorted reads can be used in the system 100 or system 300 in an effort to obtain performance enhancements, such as reduced run time, as described below.
[0046] The fusion candidate identification module 140 may obtain a batch of aligned and sorted reads aligned by the read alignment unit 136, and determine whether the batch of aligned and sorted reads includes one or more gene fusion candidates. In some specific implementations, if the received batch includes aligned and sorted reads, the fusion candidate identification module 140 may evaluate the sorted reads of the batch, wherein the genomic interval corresponding to the batch overlaps with the breakpoint of at least one fusion candidate. This can reduce the number of fusion candidates that require downstream analysis. In other specific implementations, if the received batch includes unsorted aligned reads, the fusion candidate identification module 140 may evaluate each aligned read in the batch to determine whether the aligned read is a fusion candidate. In some specific implementations, the operation of determining by the fusion candidate identification module 140 whether a batch of reads includes one or more fusion candidates includes determining by the fusion candidate identification module 140 whether the batch of reads includes one or more split read alignments, one or more inconsistent read pairs, one or more soft cut alignments, or a combination thereof.
[0047] In some specific implementations, the fusion candidate identification module 140 may be configured to identify the split read alignment as a fusion candidate. The fusion candidate identification module 140 may identify the split read alignment by analyzing the gene of the reference sequence aligned to each specific read in a batch of aligned reads. If the fusion candidate identification module 140 determines that the read is mapped to a single gene, the fusion candidate identification module 140 may determine that the read is not a split read. Alternatively, if the fusion candidate identification module 140 determines that the read is aligned to two different genes, the read may be determined to be a split read. In such specific implementations, the split read may be determined as a fusion candidate. If, for example, a first subset of nucleotides of the read is aligned to a first parent gene of a reference genome, and a second subset of nucleotides of the read is aligned to a second parent gene of the reference genome, it may be determined that the read is aligned to two different reads. In some specific implementations, the first subset of nucleotides may be a prefix of the read, and the second subset of nucleotides may be a suffix of the read. If the fusion candidate identification module 140 is configured to identify split reads, data identifying the split reads, if any, may be stored in the memory device 120 .
[0048] In some specific implementations, the fusion candidate identification module 140 may be configured to identify inconsistent read pairs as fusion candidates. The fusion candidate identification module 140 can identify inconsistent read pairs by analyzing the genes of the reference sequence to which each specific read pair in a batch of aligned reads is aligned. If the read pair is aligned to the reference sequence, and the orientation and range of the alignment are the expected orientation and range, it is determined that the read pair is not an inconsistent read. Alternatively, if the read pair is aligned to the reference sequence, and the orientation or range of the alignment is beyond expectations, it is determined that the read pair is an inconsistent read pair. In such specific implementations, if one read in the pair is mapped to one parental gene and the other read is mapped to another parental gene, the inconsistent read can be determined to be a fusion candidate. If the fusion candidate identification module 140 is configured to identify inconsistent reads, the data identifying the inconsistent reads (if any) can be stored in the memory device 120.
[0049] In some specific implementations, the fusion candidate identification module 140 may be configured to identify soft-cut alignments. The fusion candidate identification module 140 may identify soft-cut alignments by analyzing the genes of the reference sequence to which each specific aligned read in a batch of aligned reads is aligned. In some specific implementations, the fusion candidate identification module 140 may determine whether the read is aligned as a whole to a single position in the reference genome. If the fusion candidate identification module 140 determines that the read is aligned as a whole to a single position in the reference genome, the fusion candidate identification module 140 may determine that the read is not a soft-cut read. Alternatively, if the fusion candidate identification module 140 determines that only a portion of the read is aligned to the reference genome, the fusion candidate identification module 140 may determine that the read is a soft-cut read. If the aligned portion of the read is mapped to one parental gene and the unaligned portion is determined to have a sequence similar to another parental gene, the soft-cut read is determined to be a fusion candidate. If the fusion candidate identification module 140 is configured to identify soft-clipped reads, data identifying the soft-clipped reads, if any, may be stored in the memory device 120 as gene fusion candidates.
[0050] The fusion candidate screening module 150 may obtain data describing a set of fusion candidates identified by the fusion candidate identification module 140. In some implementations, the fusion candidate screening module may access the memory device 120 and obtain data describing the fusion candidates from the memory device 120. In other implementations, the fusion candidate screening module may receive data describing the fusion candidates from the output of a previous module (such as the fusion candidate identification module 140). The fusion candidate screening module 150 may use one or more filters to filter the data describing the set of fusion candidates so as to identify a filtered set of gene fusion candidates that is smaller than the entire set of gene fusion candidates. In some implementations, these filters are applied in a single stage. For example, each of one or more filters may be applied, and each fusion candidate in the set of fusion candidates may be evaluated according to each of the one or more filters. However, in other implementations, a multi-stage screening method may be adopted. In such implementations, a first set of one or more filters is applied to the initial set of fusion candidates identified by the fusion candidate identification module 140. Then, a second set of one or more filters is applied to the first set of filtered fusion candidates that remain after applying the first screening stage. Additional screening stages can also be applied as needed to achieve an optimally screened set of fusion candidates.
[0051] In some specific implementations, the fusion candidate screening module 150 may screen the group of fusion candidates to consider repeated fusion candidates caused by the high coverage depth used during short read sequencing. For example, the accumulation from 30x sequencing may cause the fusion candidate identification module 140 to identify up to 30 repeated fusion candidates. The fusion candidate screening module 150 may remove such repeated fusion candidates by applying a filter to the characteristics of the fusion candidate to check for duplication. For example, the fusion candidate screening module 150 may determine whether multiple fusion candidates are aligned to the same parental gene, aligned to a portion of the reference genome spanning the same or similar breakpoints, or a combination thereof. If the fusion candidate screening module 150 identifies multiple fusion candidates aligned to the same parental gene, aligned to a portion of the reference genome spanning the same or similar breakpoints, or a combination thereof, the fusion candidate screening module 150 may determine that the fusion candidate is repeated and only select one fusion candidate as a representative fusion candidate. In such cases, the remaining fusion candidates aligned to the same parental gene, aligned to a portion of the reference genome spanning the same or similar breakpoints, or a combination thereof, may be discarded without further downstream analysis. The representative fused candidate may then be added to a set of filtered fused candidates in a memory device, such as memory device 120 .
[0052] Alternatively or in addition, the fusion candidate screening module 150 may screen the set of fusion candidates based on one or more rule conditions. For example, the fusion candidate screening module 150 may analyze each fusion candidate and determine whether the fusion candidate has one or more attributes that satisfy the one or more rule conditions employed by the screening module 150. In some implementations, the one or more rule conditions may include the alignment position of each portion of the fusion candidate, the distance of the overlap of the alignment relative to the breakpoint spanned by the fusion candidate, the orientation of the alignment of the fusion candidate, the read alignment quality of the fusion candidate, the additional mapping position of the fusion candidate, or any combination thereof.
[0053] By way of example, the fusion candidate screening module 150 may use one or more rule conditions to screen fusion candidates based on the alignment position. In some specific implementations, for example, the fusion candidate screening module 150 may be configured to use a rule condition that screens out fusion candidates with reads that are aligned to the reference sequence so that the span of the alignment spans across the fusion breakpoint by more than a predetermined number of nucleotides. In some specific implementations, the predetermined number of nucleotides of the rule condition may be 8 nucleotides. Alternatively or in addition, the fusion candidate screening module 150 may be configured to screen out fusion candidates with reads that are aligned to the reference sequence so that the span of the alignment on the reference sequence does not reach the fusion breakpoint within a predetermined threshold number of nucleotides. In some specific implementations, the predetermined threshold number of nucleotides for the rule condition may be 50 nucleotides. Alternatively or in addition, the fusion candidate screening module 150 may be configured to use a rule condition that screens out fusion candidates with reads that are aligned to the reference sequence so that the aligned portions of the reads at the two fusion breakpoints share at least a predetermined number of nucleotides. In some implementations, the predetermined number of shared nucleotides can include at least 8 nucleotides.
[0054] For another example, the fusion candidate screening module 150 may use one or more rule conditions to screen fusion candidates based on orientation. In some specific implementations, for example, the fusion candidate screening module 150 may be configured to use a rule condition that screens out fusion candidates having an alignment orientation indicating that the nucleotide sequence of at least one parental gene is reversed in the fusion transcript.
[0055] As another example, the fusion candidate screening module 150 may use one or more rule conditions to screen fusion candidates based on mapping quality. In some specific implementations, for example, the fusion candidate screening module 150 may be configured to use a rule condition that screens out fusion candidates with read alignments whose mapping quality scores do not meet a predetermined threshold.
[0056] As another example, the fusion candidate screening module 150 may use one or more rule conditions to screen fusion candidates based on additional mapping positions. In some specific implementations, for example, the fusion candidate screening module 150 may be configured to use rule conditions that screen out fusion candidates based on determining that a portion of the read segment of the fusion candidate is mapped to multiple positions of the reference sequence. In some specific implementations, the fusion candidate screening module 150 may be configured to exclude positions annotated as homologous genes.
[0057] Fusion candidates that satisfy each of the one or more rule conditions may be added to a set of filtered fusion candidates in a memory device, such as memory device 120. Fusion candidates that do not satisfy each of the one or more rule conditions may be discarded without further downstream analysis. In some implementations, rule condition-based filtering of fusion candidates may be applied as a second-stage filter after applying a first-stage deduplication filter. In other implementations, rule condition-based filtering of fusion candidates may be applied as a first-stage filtering, and then a deduplication filter may be applied as a second-stage filter. In other implementations, rule condition-based filtering may be applied as a single-stage filter without prior deduplication filtering. Filtering fusion candidates based on one or more of these rule conditions may significantly reduce the number of fusion candidates that need to be further processed downstream.
[0058] Downstream processing may be performed on each fusion candidate in the screened set of fusion candidates output by the fusion candidate screening module 150. The downstream processing includes executing the feature set generation module 160, the machine learning model 170, the gene fusion determination module 180, and the output API module 190. Such downstream processing may be used to determine whether the candidate fusion candidate corresponds to a valid gene fusion.
[0059] The feature set generation module 160 can extract data from multiple data sources to identify data attribute groups for which feature extraction is to be performed. These data sources include attribute data about fusion candidates stored in the memory 120, which attribute data includes (i) reads of fusion candidates, (ii) portions of reference sequence positions to which the reads of fusion candidates are aligned, and (iii) annotations of fragments of the reference genome to which specific gene fusion candidates are aligned. In some specific implementations, the annotations may include gene exon annotations, annotations indicating the presence of homologous genes, annotations indicating enriched gene lists, or combinations thereof.
[0060] The data source used by the feature set generation module 160 may also include data generated by the read alignment unit 136 during the alignment process. In some implementations, the feature set generation module 160 may derive feature data from data generated by the read alignment unit 136 during the fusion candidate alignment. For example, the feature set generation module 160 may derive information such as variant allele frequency counts, counts of unique read alignments, read coverage across transcripts, MAPQ scores, data indicating homology between parental genes, or a combination thereof from the data generated by the read alignment unit 136.
[0061] The feature set generation module 160 may be used to generate feature data representing one or more of the above-mentioned attributes of the fusion candidate extracted from multiple data sources, and encode the feature data into one or more data structures 162 for input to the machine learning model 170. For example, in some implementations, the entire set of features extracted from the attributes of the fusion candidate may be encoded into a single vector 162 that is combined into the machine learning module 170. For example, in the case of segmented reads or soft-clipped alignments, each of the features extracted from the attributes of these types of fusion candidates may be encoded into a single vector 162.
[0062] In other specific implementations, the feature data extracted from the attributes of the fusion candidate can be multiple encoded input vectors. In this case, the input vector 162 can be composed of a pair of input vectors 162a, 162b. For example, in the scenario of the split read fusion candidate, each feature extracted from the attributes related to the split read prefix, including the features of the nucleotides representing the prefix of the split read, the features representing the fragment of the reference sequence to which the prefix is aligned, and any other features extracted from the above-mentioned prefix-related attributes or any combination thereof can be encoded into the input vector 162a. Similarly, in such a specific implementation, each feature extracted from the attributes related to the split read suffix, including the features of the nucleotides representing the suffix of the split read, the features representing the fragment of the reference sequence to which the suffix is aligned, and any other features extracted from the above-mentioned suffix-related attributes or any combination thereof can be encoded into the input vector 162b. For another example, when an inconsistent read pair is identified as a fusion candidate, the extracted features representing the first read of the inconsistent read pair, the extracted features representing the portion of the reference sequence aligned therewith, the features extracted from the attributes associated with the first read of the inconsistent read pair, or any combination thereof may be encoded into the input vector 162a. Similarly, in such an example, the extracted features representing the second read of the inconsistent read pair, the extracted features representing the portion of the reference sequence aligned therewith, the features extracted from the attributes associated with the second read of the inconsistent read pair, or any combination thereof may be encoded into the input vector 162b.
[0063] Each of the one or more vectors 162 may represent the generated feature data with a number, wherein the feature data includes any feature extracted from the fusion candidate or any feature extracted from the data related to the fusion candidate received from the read segment alignment unit 136 and stored in the memory 120. For example, each vector 162 or 162a, 162b may include multiple fields, each field corresponding to a specific feature of a specific read segment of a specific fusion candidate. Depending on the specific fusion candidate, this may produce one or more input vectors, as described above. The feature set generation module 160 may determine a numerical value for each field that describes the extent to which a specific feature is expressed in the attributes of the specific read segment of the fusion candidate. The determined numerical value of each field may be used to encode the generated feature data representing the read segment attributes of the fusion candidate into one or more corresponding vectors 162. The generated one or more vectors 162a, 162b (which represent the corresponding read segment of the fusion candidate with a number) are provided as input to the machine learning model 170. In some specific implementations, even if multiple concept vectors are generated for the fusion candidate, the multiple concept vectors may be shrunk into a single vector 162 that can be input into the machine learning model 170. In such embodiments, if multiple vectors are required in (i) certain segmented read embodiments in which features of the prefix are assigned to a first vector and features of the suffix are assigned to a second vector or (ii) in discordant pair embodiments, then the first portion of the single vector may correspond to the first concept vector and the second portion of the single vector may correspond to the second concept vector.
[0064] The machine learning model 170 may include a deep neural network that has been trained to generate the likelihood that a fusion candidate corresponds to an effective gene fusion based on processing of one or more input vectors 162 representing features of the fusion candidate. An effective gene fusion is a chimeric transcript that contains sequences from multiple genes due to a rearrangement in the genome that connects a prefix of one parental gene to a suffix of another parental gene. In the context of the present disclosure, if, for example, the output data 178 generated by the machine learning model meets a predetermined threshold, it will be determined that the model 170 has predicted an effective gene fusion. The machine learning model 170 may include an input layer 172 for receiving input data, one or more hidden layers 174a, 174b, 174c for processing input data received via the input layer 172, and an output layer 176 for providing output data 178. Each hidden layer 174a, 174b, 174c includes one or more weights or other parameters. During training, the weights or other parameters of each respective hidden layer 174a, 174b, 174c may be adjusted so that the trained deep neural network produces a desired target output 178 indicating the likelihood that one or more input vectors 162 represent a valid gene fusion based on processing of the one or more input vectors 162 by the machine learning model 170.
[0065] The machine learning model 170 can be trained in a variety of different ways. In one implementation, the machine learning model 170 can be trained to distinguish between (i) one or more input vectors representing features extracted from the attributes of a valid fusion candidate and (ii) one or more input vectors representing features extracted from the attributes of an invalid fusion candidate. In some implementations, such training can be achieved using labeled training vector pairs. Each training vector can represent a training fusion candidate and can be composed of feature data of the same type as the one or more input vectors 162 described above. In such implementations, one or more input vectors 162 representing features extracted from the attributes of a fusion candidate can be labeled as a valid gene fusion or an invalid gene fusion. In some implementations, a valid gene fusion label or an invalid gene fusion label can be represented as a numerical value. For example, in some implementations, a valid gene fusion label can be "1" and an invalid gene fusion label can be "0". In other implementations, for example, a valid gene fusion label can be a number between "0" and "1" that meets a predetermined threshold, and an invalid gene fusion label can be a number between "0" and "1" that does not meet a predetermined threshold. In this type of implementation, the training of the numerical value indication input vector that meets or does not meet a predetermined threshold value is to representing the confidence level of effective gene fusion or invalid gene fusion. In some implementations, meeting a predetermined threshold value may include exceeding this predetermined threshold value. However, the specific implementation may also be configured to make meeting the threshold value mean not exceeding this predetermined threshold value. This type of implementation may include, for example, the specific implementation of both comparator and parameter being invalid.
[0066] During training, each labeled set of one or more training vectors is provided as an input to the machine learning model 170, processed by the machine learning model 170, and then the training output generated by the machine learning model 170 is used to determine a predicted label for each labeled set of one or more training vectors. The predicted label generated by the machine learning model 170 based on the machine learning model's processing of the labeled one or more training vectors for a pair of reads corresponding to the training fusion candidate can be compared with the training labels of the one or more training vectors corresponding to the one or more reads (or read portions) of the training fusion candidate. The parameters of the machine learning model 170 can then be adjusted based on the difference between the predicted label and the training label. This process can continue iteratively for each of the multiple labeled training vectors corresponding to the corresponding training fusion candidate until the predicted fusion candidate label generated by the machine learning model 170 based on the processing of the set of one or more training vectors corresponding to the training fusion candidate matches the training label of the set of one or more training vectors corresponding to the corresponding training fusion candidate within a predetermined error level.
[0067] In some implementations, the labeled training fusion candidates may be obtained from a library of training fusion candidates that have been reviewed and labeled by one or more human users. However, in other implementations, the labeled training fusion candidates may include training fusion candidates that have been generated and labeled by a simulator. In such implementations, the simulator may be used to create a distribution of different categories of training fusion candidates that can be used to train the machine learning model 170. Generally speaking, if the runtime machine learning model 170 will accept a single input vector 162, where each of the extracted features for the fusion candidate is encoded as a single input vector 162, the machine learning model 170 will be trained using the above-described training process, using a single input vector having the same features as the input vector 162. Similarly, if the runtime machine learning module 170 accepts two training vectors 162a, 162b as described above, the machine learning model 170 will be trained using two input vectors, each of which has the same corresponding features as the above-described input vectors 162a, 162b. That is, the type of input vector that will be processed at runtime is the same as the time when the vector will be used to train the model 170 using the above-described training process.
[0068] During processing of input data 162 corresponding to features extracted from attributes of fusion candidates, the output of each hidden layer 174a, 174b, 174c may include an activation vector. The activation vector output by each corresponding hidden layer may be propagated through subsequent layers of the deep neural network and used by the output layer to generate output data 178. Figure 1 In the example of , the machine learning model 170 is trained to produce output data 178, which represents a combined score generated by the machine learning model 170 based on the processing of the machine learning model on separate input vectors 162a, 162b, each input vector corresponding to a read segment of the fusion candidate. The combined score 178 is ultimately generated by the output layer 176 of the trained machine learning model based on the calculations performed by the output layer 176 of the trained machine learning model 170 on the received activation vectors from the final hidden layer 174c.
[0069] The output data 178 generated by the trained machine learning model 170 can be evaluated by the gene fusion determination module 180 to determine whether the output data indicates that the fusion candidate corresponding to the one or more input vectors 162 is a valid fusion candidate. In some implementations, the output data 178 can be provided to the gene fusion determination module 180 by the trained machine learning model 170. In other implementations, the system 100 can store the output 178 of the trained machine learning model 170 to a memory device such as the memory device 120 for subsequent access by the gene fusion determination module 180.
[0070] The gene fusion determination module 180 may obtain the output data 178 generated by the machine learning model 170, and evaluate the output data 178 to determine whether the fusion candidate corresponding to the pair 162 of the input vectors 162a, 162b is a valid gene fusion based on the output data 178. In some specific implementations, the gene fusion determination module 180 may determine whether the fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion by comparing the output data 178 generated by the machine learning model with a predetermined threshold. If the gene fusion determination module 180 determines that the output data 178 meets the predetermined threshold, the gene fusion determination module 180 may determine that the fusion candidate corresponding to the one or more input vectors 162 is a valid gene fusion. Alternatively, if the gene fusion determination module 180 determines that the output data 178 does not meet the predetermined threshold, the gene fusion determination module 180 may determine that the fusion candidate corresponding to the one or more input vectors 162 is not a valid gene fusion.
[0071] In some implementations, the gene fusion determination module 180 may generate output data 182 indicating the result of the determination made by the gene fusion determination module 180, which is based on the evaluation of the output data 178 generated by the machine learning model 170 by the gene fusion determination module 180. The output data 182 may include data identifying the gene fusion candidate corresponding to one or more input vectors 162 and data identifying the determination made by the gene fusion determination module 180. The data identifying the determination made by the gene fusion determination module 180 may include data indicating whether the gene fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion or an invalid gene fusion. In some implementations, the output data 182 may only indicate a list of valid gene fusions identified based on the output data 178, a list of invalid gene fusions identified based on the output data 178, data indicating that no valid gene fusions were identified, or any combination thereof. In some implementations, the output data 182 may be stored in the memory 182 for subsequent use by another computing module, for subsequent output to a user device, etc.
[0072] Alternatively or in addition, the gene fusion determination module 180 may generate output data 184, which may be provided as input to an output application programming interface (API) module 190. The output data 184 may indicate that the output API causes the output display to generate an output indicating whether the gene fusion candidate corresponding to one or more input vectors 162 is a valid gene fusion or an invalid gene fusion. In some specific implementations, the instruction may cause the output API module 190 to access the output data 182 stored in the memory device 120 and generate rendering data, which, when rendered by a computing device coupled to the output display 195, causes the output display 195 to display (i) data identifying the fusion candidate corresponding to the one or more input vectors 162, and (ii) data indicating whether the identified fusion candidate is a valid gene fusion or an invalid gene fusion. This may include causing the output display 195 to display any output data 182 stored in the memory 184. In some specific implementations, the output may be displayed in the form of a report.
[0073] In some implementations, the gene fusion determination module 180 stores output data 182 for each gene fusion candidate in the memory device 120 based on the performance of downstream processing performed on each fusion candidate in the screened set of gene fusion candidates. In such implementations, once the downstream processing of each fusion candidate is completed, the gene fusion determination module 180 may simply instruct the output API module 190 to output the results of the gene fusion analysis stored in the memory 120 for each fusion candidate in the screened set of gene fusion candidates. In this case, the output 192 provided for display on the output display 195 will include a list of valid gene fusions, a list of invalid gene fusions, or both. In other implementations, the gene fusion determination module 180 may cause the output API module 190 to output result data indicating a list of identified gene fusions (if any) when downstream processing of a particular fusion candidate is completed.
[0074] The output API module 190 may provide other types of outputs 192. For example, in some implementations, the output 192 may be data that causes another device, such as a printer, to output a report that includes (i) data identifying fusion candidates corresponding to one or more vectors 162, and (ii) data indicating whether the identified fusion candidates are valid genes. In other implementations, the output data 192 may cause a speaker to output audio data that includes (i) data identifying fusion candidates corresponding to one or more vectors 162, and (ii) data indicating whether the identified fusion candidates are valid genes. Other types of output data may also be triggered by the output APIR module 190.
[0075] In some implementations, the output display 195 may be a display panel of the sequencing device 110. In other implementations, the output display 195 may be a display panel of a user device connected to the sequencing device 110 using one or more networks. In fact, the sequencing device 110 may be configured to transmit the output data 192 to any device having any display.
[0076] Figure 2 is a flow chart of an example of a process 200 for performing rapid detection of valid gene fusions. A system (such as system 100) can begin performing process 200 by obtaining first data representing a plurality of aligned reads from a read alignment unit using one or more computers (210). The system can identify a plurality of gene fusion candidates included in the obtained first data (220). The system can screen the plurality of gene fusion candidates to determine a screened set of gene fusion candidates (230).
[0077] The system can obtain a specific gene fusion candidate from the group of gene fusion candidates that have been screened (240). The system can generate input data for input to the machine learning model, wherein generating the input data includes extracting feature data from the data to represent the specific gene fusion candidate, the data including (i) one or more fragments of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on the output of the read alignment unit (250).
[0078] The system may provide the generated input data as input to a machine learning model, wherein the machine learning model has been trained to generate output data representing the likelihood that a gene fusion candidate is a valid gene fusion based on processing the input data by the machine learning model, the input data representing (i) a fragment of a reference genome to which a read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on the output of the read alignment unit (260). The system may obtain output data (270) generated by the machine learning model based on processing the input data by the machine learning model. The system may determine whether the specific fusion candidate corresponds to a valid gene fusion candidate based on the output data (280).
[0079] Upon completion of stage 280, the system may determine whether to evaluate another fusion candidate in the filtered set of fusion candidates (290). If the system determines that there is another fusion candidate in the filtered set of fusion candidates to be evaluated, the system may continue to perform process 200 at stage 240. Alternatively, if the system determines that there is not another fusion candidate in the filtered set of fusion candidates to be evaluated, the system may terminate performance of the process at stage 295. If the filtered set of fusion candidates has not been exhausted, there may be another fusion candidate in the set of fusion candidates.
[0080] Figure 3 is a block diagram of another example of a system 300 for rapidly detecting effective gene fusions. System 300 performs the same functions as system 100, in that system 300 generates RNA (or DNA) sequence reads 112 using a sequencing device 110, aligns RNA sequence reads 112 with a reference sequence using a secondary analysis unit 130, identifies fusion candidates using a fusion candidate identification module 140, determines a screened set of fusion candidates for downstream analysis using a fusion candidate screening module 150, and then performs downstream analysis on the screened set of fusion candidates to identify effective gene fusions using a feature set generation module 160, a machine learning model 170, a gene fusion determination module 190, and an output API module 190. Each of these functional units, modules, or models performs the same functions as described above. Figure 1 The same functions are attributed to them in the description of system 100.
[0081] The difference between system 300 and system 100 is that fusion candidate identification, fusion candidate screening, and downstream analysis of the screened set of fusion candidates are performed on another computer 320 rather than within sequencing device 110. Therefore, the difference between system 300 and system 100 is how the aligned reads are packaged and transmitted to computer 320 using network 310 for gene fusion analysis, unpacked by computer 320, and how the gene fusion results are packaged and transmitted to another device with a corresponding display for output.
[0082] In more detail, the sequencing device 110 can sequence the biological sample 105 and generate RNA reads 112-1, 112-2, 112-n, where "n" is any positive integer greater than 0, as described with reference to the system 100. Although RNA reads are used as an example, the system can also perform the same process on DNA reads. The sequencing device 110 can store the reads 112-1, 112-2, 112-n in the memory 120. In some specific implementations, the reads 112-1, 112-2, 112-n can be in a compressed format.
[0083] The secondary analysis unit 130 may obtain the reads 112-1, 112-2, 112-n and store the reads 112-1, 112-2, 122-n in a memory 132 of the secondary analysis unit 130. In some implementations, this may include a control program of the sequencing device 110 that streams the reads 112-1, 112-2, 112-n to the memory 132 of the secondary analysis unit 130. In other implementations, the secondary analysis unit 130 may request the reads 112-1, 112-2, 122-n. If the reads 112-1, 112-2, 112-n are compressed, the programmable logic device 134 of the secondary analysis unit 130 may be configured in state B as a decompression unit 138 and may be used to decompress the reads 112-1, 112-2, 112-n. The programmable logic device 134 may then be reconfigured to state A as a read alignment unit and used to align the reads 112 - 1 , 112 - 2 , 112 - n to a reference sequence.
[0084] The secondary analysis unit 130 can be reconfigured back to state B as a compression unit and use the compression unit to compress the aligned reads to prepare the aligned reads for transmission to the computer 320. In this example, the compression of the first batch of aligned reads includes not only compressing the aligned reads, but also compressing the data generated by the read alignment unit 136 and related to the aligned reads to be used for gene fusion analysis. Figure 1 The data can be described using the system 100, and can include, for example, variant allele frequency counts, counts of unique read alignments, read coverage across transcripts, MAPQ scores, data indicating homology between parental genes, or a combination thereof. In addition, other data that can be compressed into the first batch of aligned reads can include (i) reads of fusion candidates, (ii) portions of reference sequence positions to which the reads of fusion candidates are aligned, and (iii) annotations of fragments of the reference genome to which specific gene fusion candidates are aligned. In some specific implementations, the annotations can include gene exon annotations, annotations indicating the presence of homologous genes, annotations indicating enriched gene lists, or a combination thereof.
[0085] After compressing the aligned reads, the secondary analysis unit 130 may store the first batch of compressed reads in the memory 120. Then, the sequencing device 110 may transmit the first batch 125 of aligned reads to the computer 320 via the network 310 for gene fusion analysis. The network 310 may include one or more wired networks, one or more wireless networks, or a combination thereof. In different implementations, the network 310 may be one or more of a wired Ethernet, a wired optical network, a LAN, a WAN, a cellular network, the Internet, or a combination thereof. In some implementations, the computer 320 may be a remote cloud server. However, in other implementations, the computer 320 may be connected to the sequencing device 110 via a direct connection (such as a direct Ethernet connection, a USB-C connection, etc.). Although in this example of the system 300, the first batch of reads is compressed before transmission, compression does not necessarily need to be used. Instead, compression is provided as a method to reduce network bandwidth consumption and minimize storage costs, which can provide significant technical benefits and reduce costs when processing large amounts of genomes.
[0086] In some implementations, the first batch of aligned reads includes the entire set of reads generated for the sample 105. In other implementations, the first batch of aligned reads is only a portion of the entire set of reads generated for the sample 105, and the batch processing system can be used to facilitate parallel processing. For example, in some implementations, after the secondary analysis unit stores the first batch of aligned reads in the memory 120, the secondary analysis unit 130 obtains a second batch of reads that have not yet been aligned to store in the memory 132. Then, if the second batch of reads is compressed, the secondary analysis unit 130 can perform decompression and perform alignment of the second batch of reads while the computer 320 performs gene fusion analysis of the first batch of reads. This parallel processing facilitated by batching reads can significantly reduce the operating time of the system 300 required to determine the valid gene fusions of the reads of the sample 105.
[0087] The computer 320 may receive the first batch of reads 125 via the network 310 and store the first batch of reads in the memory 320. If the first batch of reads 125 are compressed, the computer 320 may use the compression / decompression module 325 to decompress the first batch of reads and store the first batch of reads in the memory 320. The computer 320 may then use the same compression / decompression module as the reference numeral 325 to store the first batch of reads in the memory 320. Figure 1 The gene fusion analysis pipeline of the fusion candidate identification module 140, the fusion candidate screening module 150, the feature set generation module 160, the machine learning model 170, the gene fusion determination module 180 and the output API module 190 is executed in the same manner as described in the system 100.
[0088] Output 192 may be provided to a plurality of different devices via network 310. By way of example, the output data may be transmitted to a sequencing device for output on a display 195 of a sequencer. Alternatively or in addition thereto, output 192 may be provided for display on a display of a user device 330 via network 310. User device 330 may include a smart phone, tablet computer, laptop computer, desktop computer, or any other computer with a display. Alternatively or in addition thereto, output 192 may also be provided for output from a printer 340 via network 310. In such specific implementations, the output may be a hard copy report of the determined valid gene fusions.
[0089] Figure 4 is a block diagram of system components that may be used to implement a system for rapid detection of gene fusions.
[0090] Computing device 400 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. Computing device 450 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, and other similar computing devices. In addition, computing device 400 or 450 may include a universal serial bus (USB) flash drive. The USB flash drive can store operating systems and other applications. The USB flash drive may include input / output components, such as a wireless transmitter or a USB connector that can be inserted into a USB port of another computing device. The components shown here, their connections and relationships, and their functions are intended only as examples and are not intended to limit the specific implementation of the present invention described and / or claimed in this document.
[0091] The computing device 400 includes a processor 402, a memory 404, a storage device 406, a high-speed interface 408 connected to the memory 404 and a high-speed expansion port 410, and a low-speed interface 412 connected to a low-speed bus 414 and the storage device 408. Each of the components 402, 404, 406, 408, 410, and 412 is interconnected using various buses and can be installed on a common motherboard or installed in other ways as appropriate. The processor 402 can process instructions for execution within the computing device 400, including instructions stored in the memory 404 or on the storage device 408, to display graphical information of a GUI on an external input / output device (such as a display 416 coupled to the high-speed interface 408). In other specific implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple types of memories as appropriate. In addition, multiple computing devices 400 can be connected, each device providing some parts of the necessary operations, for example, as a server library, a group of blade servers, or a multi-processor system.
[0092] The memory 404 stores information within the computing device 400. In one implementation, the memory 404 is one or more volatile memory units. In another implementation, the memory 404 is one or more non-volatile memory units. The memory 404 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0093] Storage device 408 can provide mass storage for computing device 400. In one specific implementation, storage device 408 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. A computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 404, storage device 408, or a memory on processor 402.
[0094] The high-speed controller 408 manages bandwidth-intensive operations of the computing device 400, while the low-speed controller 412 manages bandwidth-less intensive operations. This functional allocation is only an example. In a specific implementation, the high-speed controller 408 is coupled to the memory 404, the display 416, and is coupled to the high-speed expansion port 410, which can accept various expansion cards (not shown) through a graphics processor or accelerator. In this specific implementation, the low-speed controller 412 is coupled to the storage device 408 and the low-speed expansion port 414. The low-speed expansion port (which may include various communication ports, such as USB, Bluetooth, Ethernet, wireless Ethernet) can be coupled to one or more input / output devices, such as keyboards, pointing devices, microphone / speaker pairs, scanners or networking devices such as switches or routers, for example, through a network adapter. The computing device 400 can be implemented in a variety of different forms, as shown in the figure. For example, the computing device can be implemented as a standard server 420, or it can be implemented multiple times in a group of such servers. It can also be implemented as a part of a rack server system 424. In addition, the computing device can be implemented in a personal computer such as a laptop computer 422. Alternatively, components from computing device 400 may be combined with other components in a mobile device (not shown), such as device 450. Each of such devices may contain one or more of computing devices 400, 450, and the entire system may be composed of multiple computing devices 400, 450 in communication with each other.
[0095] The computing device 400 can be implemented in a variety of different forms, as shown. For example, the computing device can be implemented as a standard server 420, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 424. In addition, the computing device can be implemented in a personal computer such as a laptop computer 422. Alternatively, components from the computing device 400 can be combined with other components in a mobile device (not shown) such as device 450. Each of such devices can contain one or more of the computing devices 400, 450, and the entire system can be composed of multiple computing devices 400, 450 that communicate with each other.
[0096] Computing device 450 includes processor 452, memory 464, and input / output devices such as display 454, communication interface 466, and transceiver 468, among other components. Device 450 may also be provided with a storage device, such as a microdrive or other device, to provide additional storage. Each of components 450, 452, 464, 454, 466, and 468 are interconnected using various buses, and several of these components may be mounted on a common motherboard or otherwise as appropriate.
[0097] Processor 452 can execute instructions within computing device 450, including instructions stored in memory 464. The processor can be implemented as a chipset including a plurality of independent analog processors and digital processor chips. In addition, the processor can be implemented using any of a variety of architectures. For example, processor 410 can be a CISC (complex instruction set computer) processor, a RISC (reduced instruction set computer) processor, or a MISC (minimum instruction set computer) processor. The processor can provide, for example, coordination of other components of device 450, such as control of a user interface, applications run by device 450, and wireless communications performed by device 450.
[0098] The processor 452 can communicate with the user through a control interface 458 and a display interface 456 coupled to the display 454. The display 454 can be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display or other appropriate display technology. The display interface 456 may include appropriate circuits for driving the display 454 to present graphics and other information to the user. The control interface 458 can receive commands from the user and convert these commands to submit to the processor 452. In addition, an external interface 462 in communication with the processor 452 can be provided to enable close range area communication of the device 450 with other devices. The external interface 462 can, for example, provide wired communication in some specific implementations, or provide wireless communication in other specific implementations, and multiple interfaces can also be used.
[0099] Memory 464 stores information within computing device 450. Memory 464 may be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. An expansion memory 474 may also be provided and connected to device 450 via expansion interface 472, which may include, for example, a SIMM (single in-line memory module) card interface. Such expansion memory 474 may provide additional storage space for device 450, or may also store applications or other information for device 450. Specifically, expansion memory 474 may include instructions for executing or supplementing the above-mentioned processes, and may also include security information. Thus, for example, expansion memory 474 may be provided as a security module for device 450, and may be programmed with instructions for implementing secure use of device 450. In addition, secure applications may be provided via a SIMM card together with additional information, such as placing identification information on a SIMM card in an unbreakable manner.
[0100] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer readable medium or machine readable medium, such as memory 464, expansion memory 474, or memory on processor 452 that can be received by, for example, transceiver 468 or external interface 462.
[0101] Device 450 can communicate wirelessly via communication interface 466, which may include digital signal processing circuitry as needed. Communication interface 466 can provide communication in various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000 or GPRS, etc. Such communication can occur, for example, via radio frequency transceiver 468. In addition, short-range communication can occur, such as using Bluetooth, Wi-Fi or other such transceivers (not shown). In addition, GPS (global positioning system) receiver module 470 can provide additional navigation and location-related wireless data to device 450, which can be used as appropriate by applications running on device 450.
[0102] Device 450 may also communicate audibly using audio codec 460, which may receive verbal information from a user and convert it into usable digital information. Audio codec 460 may also generate audible sounds for the user, such as through a speaker (e.g., in a handheld terminal of device 450). Such sounds may include sounds from voice phone calls, may include recorded sounds, such as voice messages, music files, etc., and may also include sounds generated by applications operating on device 450.
[0103] The computing device 450 can be implemented in many different forms, as shown. For example, the computing device can be implemented as a cellular phone 480. The computing device can also be implemented as part of a smart phone 482, a personal digital assistant, or other similar mobile device.
[0104] Various implementations of the systems and methods described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations of such implementations. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose processor, coupled to receive data and instructions from and send data and instructions to a storage system, at least one input device, and at least one output device.
[0105] These computer programs (also referred to as programs, software, software applications or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly language / machine language. As used herein, the terms "machine-readable medium", "computer-readable medium" refer to any computer program product, apparatus and / or device, such as a disk, optical disk, memory, programmable logic device (PLD), for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0106] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input.
[0107] The systems and techniques described herein may be implemented in a computing system that includes a back-end component (e.g., as a data server) or includes a middleware component (e.g., an application server) or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser) through which a user may interact with a specific implementation of the systems and techniques described herein, or with any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0108] The computing system may include clients and servers. Clients and servers are generally remote from each other and generally interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0109] Other Implementations
[0110] A number of embodiments have been described. However, it should be appreciated that various modifications may be made without departing from the spirit and scope of the invention. Furthermore, the logic flows shown in the accompanying drawings do not require the particular order or ordered order shown to achieve the desired results. Additionally, other steps may be provided in the flow, or steps may be eliminated, and other components may be added to or removed from the system. Therefore, other embodiments are also within the scope of the following claims.
Claims
1. A computer-implemented method for identifying one or more gene fusions in a biological sample, the method comprising: include: For each of the millions of reference sequence positions: Obtaining, by one or more computers, first data corresponding to accumulated aligned reads from a read alignment unit, the aligned reads being generated using short read sequencing and having a high depth of coverage at positions of the reference sequence; Determining, by one or more computers, whether one or more gene fusion candidates are included in the obtained first data; Based on determining that one or more gene fusion candidates are included in the obtained first data, (i) removing duplicate fusion candidates caused by high coverage depth at the reference sequence position by one or more computers, and (ii) adding the remaining gene fusion candidates to the screened set of gene fusion candidates; For each specific gene fusion candidate in the screened set of gene fusion candidates: Generating input data for input to a machine learning model by one or more computers, wherein generating the input data comprises extracting feature data to represent the specific gene fusion candidate from data comprising: (i) one or more segments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit comprises one or more of variant allele frequency counts, counts of unique read alignments, or a MAPQ score; Providing the generated input data as input to the machine learning model by one or more computers, wherein the machine learning model has been trained using labeled training vectors to generate output data representing the likelihood that a gene fusion candidate is a valid gene fusion based on processing the input data by the machine learning model, each labeled training vector representing a training fusion candidate and including (i) one or more segments of a reference sequence to which the training fusion candidate is aligned, and (ii) one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores, the input data representing (i) one or more segments of a reference sequence to which the read alignment unit aligns the particular gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit includes one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores; obtaining, by one or more computers, output data generated by the machine learning model based on input data generated by processing the machine learning model; as well as A determination is made by one or more computers based on the output data as to whether the particular gene fusion candidate corresponds to a valid gene fusion candidate.
2. The method according to claim 1, wherein generating the input data further comprises extracting feature data, the feature data comprising annotation data, the annotation data describing annotations of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate; and wherein the machine learning model has been trained to generate output data representing a likelihood that a gene fusion candidate is a valid gene fusion candidate based on processing input data by the machine learning model, the input data representing: (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, (ii) annotation data describing an annotation of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (iii) data generated based on the output of the read alignment unit.
3. The method of claim 1, wherein determining by one or more computers whether one or more gene fusion candidates are included in the obtained first data comprises identifying by one or more computers a plurality of segmented read alignments.
4. The method of claim 1, wherein determining, by one or more computers, whether one or more gene fusion candidates are included in the obtained first data comprises identifying, by one or more computers, a plurality of inconsistent read alignment pairs.
5. The method of claim 1 , wherein the read alignment unit is implemented using a set of one or more processing engines configured to use hardware logic circuits that have been physically arranged to perform operations using the hardware logic circuits to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
6. The method of claim 1, wherein the read alignment unit is implemented using a set of one or more processing engines by executing software instructions using one or more central processing units (CPUs) or one or more graphics processing units (GPUs), the software instructions causing the one or more CPUs or one or more GPUs to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions for the first read, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
7. The method according to claim 1, further comprising: include: Wherein obtaining, by one or more computers, first data representing the accumulated aligned reads from a read alignment unit comprises obtaining, by one or more computers, the accumulated aligned reads from a memory device, and performing one or more operations in the method according to claim 1 when the read alignment unit aligns a second accumulated read that has not yet been aligned.
8. The method of claim 1, wherein determining whether the specific gene fusion candidate corresponds to a valid gene fusion candidate based on the output data include: determining, by one or more computers, whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data satisfies the predetermined threshold, it is determined that the particular gene fusion candidate corresponds to a valid gene fusion candidate.
9. The method of claim 1, wherein determining whether the specific gene fusion candidate corresponds to a valid gene fusion candidate based on the output data include: determining, by one or more computers, whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data does not satisfy the predetermined threshold, it is determined that the particular gene fusion candidate does not correspond to a valid gene fusion candidate.
10. The method of claim 1, wherein the high depth of coverage at a reference sequence position is 30x coverage.
11. The method of claim 1, wherein the data generated based on the output of the read alignment unit comprises data indicating homology between parental genes.
12. The method according to claim 11, wherein the generated input data is provided as input to the machine learning model by one or more computers. include: The generated input data is provided as input to the machine learning model by one or more computers, wherein the machine learning model has been trained using labeled training vectors to generate output data representing the likelihood that a gene fusion candidate is a valid gene fusion based on processing the input data by the machine learning model, each labeled training vector representing a training fusion candidate and including (i) one or more fragments of a reference sequence to which the training fusion candidate is aligned, and (ii) one or more of variant allele frequency counts, counts of unique read alignments, MAPQ scores, or data indicating homology between parental genes, the input data representing (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on the output of the read alignment unit, wherein the data generated based on the output of the read alignment unit includes one or more of variant allele frequency counts, counts of unique read alignments, MAPQ scores, or data indicating homology between parental genes.
13. The method according to claim 1, further comprising: include: Based on determining that one or more gene fusion candidates are included in the obtained first data, determining not to add any gene fusion candidates to the screened set of gene fusion candidates.
14. A system for identifying one or more gene fusions in a biological sample, include: One or more computers, and one or more storage devices storing instructions, the instructions, when executed by the one or more computers, being operable to cause the one or more computers to perform operations comprising: For each of the millions of reference sequence positions: obtaining, by one or more computers, first data representing accumulated aligned reads from a read alignment unit, the aligned reads being generated using short read sequencing and having a high depth of coverage at positions of the reference sequence; Determining, by one or more computers, whether one or more gene fusion candidates are included in the obtained first data; Based on determining that one or more gene fusion candidates are included in the obtained first data, (i) removing duplicate fusion candidates caused by high coverage depth at the reference sequence position by one or more computers, and (ii) adding the remaining gene fusion candidates to the screened set of gene fusion candidates; For each specific gene fusion candidate in the screened set of gene fusion candidates: Generating input data for input into a machine learning model by one or more computers, wherein generating the input data comprises extracting feature data from the data to represent the specific gene fusion candidate, the data comprising: (i) one or more segments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit comprises one or more of variant allele frequency counts, counts of unique read alignments, and MAPQ scores; The generated input data is provided as input to the machine learning model by one or more computers, wherein the machine learning model has been trained using labeled training vectors to generate output data representing the likelihood that a gene fusion candidate is a valid gene fusion based on processing the input data by the machine learning model, each labeled training vector representing a training fusion candidate and comprising (i) one or more reference sequences to which the training fusion candidate is aligned; fragments, and (ii) one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores, the input data representing (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit comprises one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores; obtaining, by one or more computers, output data generated by the machine learning model based on input data generated by processing the machine learning model; and A determination is made by one or more computers based on the output data as to whether the particular gene fusion candidate corresponds to a valid gene fusion candidate.
15. The system according to claim 14, wherein generating the input data further comprises extracting feature data, the feature data comprising annotation data, the annotation data describing annotations of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate; and wherein the machine learning model has been trained to generate output data representing a likelihood that a gene fusion candidate is a valid gene fusion candidate based on processing input data by the machine learning model, the input data representing: (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, (ii) annotation data describing an annotation of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (iii) data generated based on the output of the read alignment unit.
16. The system of claim 14, wherein determining, by one or more computers, whether one or more gene fusion candidates are included within the obtained first data comprises identifying, by one or more computers, a plurality of segmented read alignments.
17. The system of claim 14, wherein determining, by one or more computers, whether one or more gene fusion candidates are included within the obtained first data comprises identifying, by one or more computers, a plurality of discordant read alignment pairs.
18. The system of claim 14, wherein the read alignment unit is implemented using a set of one or more processing engines configured to use hardware logic circuits that have been physically arranged to perform operations using the hardware logic circuits to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
19. The system of claim 14, wherein the read alignment unit is implemented using a set of one or more processing engines by executing software instructions using one or more central processing units (CPUs) or one or more graphics processing units (GPUs), the software instructions causing the one or more CPUs or one or more GPUs to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions for the first read, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
20. The system of claim 14, wherein the operation further comprises: include: Wherein obtaining, by one or more computers, first data representing the accumulated aligned reads from a read alignment unit comprises obtaining, by one or more computers, the accumulated aligned reads from a memory device, and performing one or more of the operations according to claim 14 when the read alignment unit aligns a second accumulated read that has not yet been aligned.
21. The system of claim 14, wherein determining whether the particular gene fusion candidate corresponds to a valid gene fusion candidate is based on the output data. include: determining, by one or more computers, whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data satisfies the predetermined threshold, it is determined that the particular gene fusion candidate corresponds to a valid gene fusion candidate.
22. The system of claim 14, wherein determining whether the particular gene fusion candidate corresponds to a valid gene fusion candidate is based on the output data. include: determining, by one or more computers, whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data does not satisfy the predetermined threshold, it is determined that the particular gene fusion candidate does not correspond to a valid gene fusion candidate.
23. A non-transitory computer-readable medium storing software, the software comprising instructions executable by one or more computers, the instructions, when subjected to such execution, causing the one or more computers to perform operations, the operations include: For each of the millions of reference sequence positions: obtaining first data representing accumulated aligned reads from a read alignment unit, the aligned reads generated using short read sequencing and having a high coverage depth at the reference sequence positions; determining whether one or more gene fusion candidates are included in the obtained first data; Based on determining that one or more gene fusion candidates are included in the obtained first data, (i) removing duplicate fusion candidates arising from a high coverage depth at a reference sequence position, and (ii) adding the remaining gene fusion candidates to a screened set of gene fusion candidates; For each specific gene fusion candidate in the screened set of gene fusion candidates: Generating input data for input into a machine learning model, wherein generating the input data comprises extracting feature data from the data to represent the specific gene fusion candidate, the data comprising: (i) one or more segments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit comprises one or more of variant allele frequency counts, counts of unique read alignments, and MAPQ scores; Providing the generated input data as input to the machine learning model, wherein the machine learning model has been trained using labeled training vectors to generate output data representing the likelihood that a gene fusion candidate is a valid gene fusion based on processing the input data by the machine learning model, each labeled training vector representing a training fusion candidate and including (i) one or more segments of a reference sequence to which the training fusion candidate is aligned, and (ii) one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores, the input data representing (i) one or more segments of a reference sequence to which the read alignment unit aligns the particular gene fusion candidate, and (ii) data generated based on an output of the read alignment unit, wherein the data generated based on the output of the read alignment unit includes one or more of variant allele frequency counts, counts of unique read alignments, or MAPQ scores; obtaining output data generated by the machine learning model based on input data generated by processing the machine learning model; as well as A determination is made based on the output data as to whether the particular gene fusion candidate corresponds to a valid gene fusion candidate.
24. The computer readable medium of claim 23, wherein generating the input data further comprises extracting feature data, the feature data comprising annotation data, the annotation data describing annotations of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate; and wherein the machine learning model has been trained to generate output data representing a likelihood that a gene fusion candidate is a valid gene fusion candidate based on processing input data by the machine learning model, the input data representing: (i) one or more fragments of a reference sequence to which the read alignment unit aligns the specific gene fusion candidate, (ii) annotation data describing an annotation of the segment of the reference sequence to which the read alignment unit aligns the specific gene fusion candidate, and (iii) data generated based on the output of the read alignment unit.
25. The computer-readable medium of claim 23, wherein determining whether one or more gene fusion candidates are included within the obtained first data comprises identifying, by one or more computers, a plurality of segmented read alignments.
26. The computer-readable medium of claim 23, wherein determining whether one or more gene fusion candidates are included within the obtained first data comprises identifying, by one or more computers, a plurality of discordant read alignment pairs.
27. The computer-readable medium of claim 23, wherein the read alignment unit is implemented using a set of one or more processing engines configured to use hardware logic circuits that have been physically arranged to perform operations using the hardware logic circuits to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
28. The computer-readable medium of claim 23, wherein the read alignment unit is implemented using a set of one or more processing engines by executing software instructions using one or more central processing units (CPUs) or one or more graphics processing units (GPUs), the software instructions causing the one or more CPUs or one or more GPUs to: (i) receiving data representing a first read segment, (ii) mapping the data representing the first read to one or more portions of a reference sequence to identify one or more matching reference sequence positions for the first read, (iii) generating one or more alignment scores corresponding to each of said matched reference sequence positions for said first read, (iv) selecting one or more candidate alignments for the first read based on the one or more alignment scores, and (v) outputting data representing candidate alignments of the first read.
29. The computer readable medium of claim 23, wherein the operation further comprises: include: Wherein obtaining first data representing the accumulated aligned reads from the read segment alignment unit comprises obtaining the accumulated aligned reads from a memory device, and when the read segment alignment unit aligns the second accumulated reads that have not been aligned, one or more operations of the operations according to claim 23 are performed.
30. The computer-readable medium of claim 23, wherein determining whether the particular gene fusion candidate corresponds to a valid gene fusion candidate is based on the output data. include: determining whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data satisfies the predetermined threshold, it is determined that the particular gene fusion candidate corresponds to a valid gene fusion candidate.
31. The computer-readable medium of claim 23, wherein determining whether the particular gene fusion candidate corresponds to a valid gene fusion candidate is based on the output data. include: determining whether the output data satisfies a predetermined threshold; as well as Based on determining that the output data does not satisfy the predetermined threshold, it is determined that the particular gene fusion candidate does not correspond to a valid gene fusion candidate.
Citation Information
Patent Citations
System and method for detecting gene fusion
US20180341746A1