Method and system for determining fusion events

A computer-implemented system for de novo assembly of sequence reads enhances the specificity of fusion event detection in cancer genetics by using alignment, breakpoint identification, and criteria application, addressing the challenge of false positives in existing methods.

JP2025134743APending Publication Date: 2025-09-17GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025094835
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-14
Filing Date
2025-06-06
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Existing methods for detecting fusion events in cancer genetics face challenges in achieving high specificity without compromising sensitivity, especially with ultra-deep coverage assays, leading to false positives due to technical artifacts.

Method used

A computer-implemented system and method for de novo assembly of sequence reads, involving alignment, breakpoint identification, grouping, contig assembly, and applying criteria to detect fusion events, including error correction and footprint tests to enhance specificity.

Benefits of technology

The method significantly improves the specificity of fusion event detection while maintaining sensitivity, enabling accurate identification of candidate fusion events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134743000001_ABST
    Figure 2025134743000001_ABST
Patent Text Reader

Abstract

To provide a method and system for determining fusion events of gene sequences, the method and system having an improved capability to detect fusion events with high sensitivity and specificity by using de novo assembly of input sequence reads before calling the fusion events.SOLUTION: A method 100 for determining a fusion event includes: the step 110 of extracting nucleic acids from a test sample; the step 120 of preparing a sequencing library; the step 130 of sequencing the nucleic acids in the sequencing library to generate sequence reads; and the step 140 of processing the sequence reads by computer analysis to call a fusion event.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of the priority date of U.S. Provisional Patent Application No. 62 / 976,884, filed February 14, 2020, which is incorporated by reference in its entirety for all purposes. [Background technology]

[0002] background Cancer is one of the leading causes of death worldwide and is a heterogeneous and complex disease with multiple genes in diverse pathways involved in its development, uncontrolled growth, invasion, and metastasis. One hallmark of cancer is genetic instability, which can lead to chromosomal translocations, insertions, duplications, deletions, and inversions. These genetic mutations often cause gene fusions, which are subsequently transcribed into fusion mRNAs or fusion transcripts. However, de novo detection of such fusion events can be challenging, especially when high specificity is required. This is because technical artifacts introduced at both the assay and analytical levels can result in false positives. This is exacerbated when the input data contains sequences generated by ultra-deep coverage assays.

[0003] Thus, there is a need for improved systems and methods for detecting fusion events that significantly increase specificity without adversely affecting overall sensitivity.It is therefore an object of the present invention to provide computer-implemented systems and methods that have an improved ability to detect fusion events by de novo assembly of input sequence reads prior to calling the fusion events. Summary of the Invention [Means for solving the problem]

[0004] Abstract It is to be understood that both the following general description and the following detailed description are exemplary and explanatory only and are not restrictive. Methods, systems, and apparatus for determining fusion events are described herein.

[0005] In an embodiment, a method is described that includes the steps of aligning a plurality of sequence reads to a reference sequence, determining one or more breakpoints in the alignment of at least one sequence read of the plurality of sequence reads to the reference sequence, identifying any sequence read associated with the one or more breakpoints in the alignment as a candidate fusion sequence read, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints, grouping the candidate fusion sequence reads based on the one or more common breakpoints, assembling the candidate fusion sequence reads within the group into one or more contigs, aligning the contigs from the group to the reference sequence, determining one or more candidate fusion events based on the alignment of the contigs from the group, applying one or more criteria to the one or more candidate fusion events, and determining one or more fusion events based on the step of applying the one or more criteria to the one or more candidate fusion events.

[0006] In another embodiment, a method is described that includes the steps of aligning a plurality of sequence reads to a reference sequence; determining one or more candidate fusion sequence reads from the plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence; grouping the one or more candidate fusion sequence reads into one or more container data structures based on one or more common breakpoints; for each container data structure, assembling the one or more candidate fusion sequence reads into one or more contigs; for each container data structure, aligning the one or more contigs to the reference sequence; and determining one or more aligned contigs that indicate a fusion event based on one or more criteria.

[0007] In certain embodiments, identifying any sequence reads associated with one or more breakpoints in an alignment as candidate fusion sequence reads includes discarding logical alignments. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints includes determining that at least two candidate fusion sequence reads include breakpoints on the same chromosome and in the same orientation. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints includes determining that at least two candidate fusion sequence reads include breakpoints at the same position. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints includes determining that at least two candidate fusion sequence reads include breakpoints within a threshold number of bases from a certain position. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints includes determining that at least two candidate fusion sequence reads include breakpoints on the same chromosome and in the same orientation. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads include multiple breakpoints at the same position. In certain embodiments, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads each include multiple breakpoints within a threshold number of bases from multiple positions.

[0008] In certain embodiments, grouping candidate fusion sequence reads based on one or more common breakpoints includes generating a de Bruijn graph for the group. In certain embodiments, assembling candidate fusion sequence reads within the group into one or more contigs includes linearizing the de Bruijn graph to generate contigs for the group. In certain embodiments, assembling candidate fusion sequence reads within the group into one or more contigs includes performing one or more error correction procedures. In certain embodiments, the one or more error correction procedures include eliminating mismatches between the candidate fusion sequence reads and a reference sequence. In certain embodiments, the one or more error correction procedures include inserting padding between at least two candidate fusion sequence reads. In certain embodiments, the one or more error correction procedures include discarding one or more candidate fusion sequence reads having unaligned portions exceeding a threshold.

[0009] In certain embodiments, determining one or more candidate fusion events based on an alignment of contigs from the group comprises applying one or more of a footprint test or a variability test. In certain embodiments, applying the footprint test comprises determining that a threshold number of families of candidate fusion sequence reads that support the contig span the breakpoint. In certain embodiments, applying the variability test comprises determining that a threshold amount of variability exists between at least two families of candidate fusion sequence reads that support the contig and span the breakpoint.

[0010] In certain embodiments, applying one or more criteria to the one or more candidate fusion events comprises determining, for the candidate fusion event, the distance between the breakpoint of one or more aligned contigs and the location of at least one probe of the panel; and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that do not contain a breakpoint that is less than a threshold distance from the location of at least one probe of the panel. In certain embodiments, applying one or more criteria to the one or more candidate fusion events comprises determining one or more genes of interest; and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that do not contain a breakpoint associated with the one or more genes of interest. In certain embodiments, applying one or more criteria to the one or more candidate fusion events comprises determining, for the candidate fusion event, that the breakpoint of one or more aligned contigs is a deletion; and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that contain a deletion located within several bases of another deletion. In certain embodiments, applying one or more criteria to the one or more candidate fusion events includes determining, for the candidate fusion event, that the breakpoint of one or more aligned contigs is a deletion; and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that contain a deletion that includes fewer than a threshold number of bases. In certain embodiments, applying one or more criteria to the one or more candidate fusion events includes discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that contain an insertion or deletion that is completely buried in an intron region.In certain embodiments, applying one or more criteria to one or more candidate fusion events includes determining a molecular to read ratio for one or more aligned contigs for the candidate fusion event; and discarding any candidate fusion events associated with the aligned contigs of one or more contigs that are associated with a molecular to read ratio above a threshold but are not associated with double-stranded support molecules. In certain embodiments, applying one or more criteria to one or more candidate fusion events includes determining, for a breakpoint pair of one or more aligned contigs for the candidate fusion event, sequences adjacent to the breakpoint of the breakpoint pair; aligning the sequences adjacent to the breakpoint of the breakpoint pair; determining an alignment score for the alignment of the sequences adjacent to the breakpoint of the breakpoint pair; and discarding any candidate fusion events associated with the aligned contigs of one or more contigs based on an alignment score above a threshold. In certain embodiments, applying one or more criteria to one or more candidate fusion events includes, for the candidate fusion event, determining, for one or more aligned contig breakpoint pairs, sequences centered at the breakpoints of the breakpoint pairs; aligning the sequences centered at the breakpoints with each other; determining an alignment score for the alignment of the sequences centered at the breakpoints; and discarding any candidate fusion event associated with an aligned contig of one or more contigs based on an alignment score exceeding a threshold.

[0011] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report may be in paper or electronic format. For example, fusion events determined by the methods and systems disclosed herein may be displayed directly in such a report. Alternatively or additionally, diagnostic information or therapeutic recommendations based on the determination of fusion events may be included in the report.

[0012] The various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographic locations, e.g., countries, and / or by the same or different persons.

[0013] In some embodiments, methods of treating a subject are described, comprising administering one or more therapeutic agents to the subject, where the subject has been determined to have a fusion event using the disclosed methods of determining a fusion event. In some embodiments, methods of treating a subject are described, comprising administering a different therapeutic agent to the subject than was previously administered, where the subject has been determined to have a fusion event using the disclosed methods of determining a fusion event. In some embodiments, methods of treating a subject are described, comprising discontinuing administration of a therapeutic agent to the subject, where the subject has been determined to have a fusion event using the disclosed methods of determining a fusion event.

[0014] Additional advantages will be set forth in part in the description which follows, or may be learned by practice. The advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, serve to explain the principles of the methods and systems described herein. [Brief explanation of the drawings]

[0015] [Figure 1] Figure 1 shows an example of the method. [Figures 2A-2C] Figures 2A-2C show an example of the stitching and trimming process to generate fragments. [Figure 3] Figure 3 shows an example of an artifact from the stitching process. [Figure 4] FIG. 4 shows an example of the method. [Figure 5] FIG. 5 shows examples of cut points. [Figure 6]FIG. 6 shows the selection of candidate fusion sequence reads. [Figure 7] FIG. 7 shows the identification of a common breakpoint between two candidate fusion sequence reads. [Figure 8] FIG. 8 shows the identification of a common breakpoint between two candidate fusion sequence reads. [Figure 9A-9B] 9A-B show minimal examples of de Bruijn graphs and compact de Bruijn graphs. [Figure 10] FIG. 10 shows an example of the use of an adjacency list for each vertex in a graph data structure. [Figure 11] FIG. 11 shows an example of the use of adjacency lists for each vertex and edge of a graph data structure. [Figure 12] FIG. 12 shows the error correction procedure. [Figure 13] FIG. 13 shows the error correction procedure. [Figure 14] FIG. 14 shows the error correction procedure. [Figure 15] FIG. 15 shows the error correction procedure. [Figure 16] FIG. 16 shows the determination of candidate fusion events. [Figure 17] FIG. 17 shows the determination of candidate fusion events. [Figure 18] Figure 18 shows the prevalence of FGFR2 / 3 fusion partners in a broad cancer cohort. Frequency of FGFR2 and FGFR3 fusion partners detected in a broad cancer cohort. IGR: intergenic region. FGFR2 as a partner gene to itself represents a long deletion or insertion. [Figure 19] Figure 19 shows the prevalence of FGFR3 fusion partners in advanced urothelial carcinoma (aUC). Some aUC patients with FGFR3 fusions were detected by partner genes. IGR: intergenic region. FGFR3 as a partner gene to itself represents a long deletion or insertion. [Figure 20]Figure 20 shows mutations that occur with FGFR2 / 3 fusions in a broad cancer cohort. Mutations occurring in at least three FGFR2 or FGFR3 fusion-positive patients in a broad cancer cohort are shown. Variants marked with a triangle show significant enrichment in the fusion-positive population (▼ p<1×10-4, ▼▼ p<1×10-10, chi-squared test, Bonferroni correction). [Figure 21] FIG. 21 shows an example of a computing device. [Figure 22] FIG. 22 shows an example of the method. [Figure 23] FIG. 23 shows an example of the method. DETAILED DESCRIPTION OF THE INVENTION

[0016] Detailed Description As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, the other configuration includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms the other configuration. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint.

[0017] "Optionally" and "as needed" mean that the subsequently described event or circumstance may or may not occur, and that the description includes cases where said event or circumstance occurs and cases where it does not occur.

[0018] Throughout the description and claims of this specification, the word "comprise" and variations of the word, such as "comprising" and "comprises," mean "including but not limited to." "Exemplary" means "an example of" and is not intended to convey an indication of a preferred or ideal configuration. "E.g., such as" is used for illustrative purposes and not limiting.

[0019] The term "subject" may refer to an animal, such as a mammalian species (preferably a human) or an avian (e.g., a bird) species. More specifically, the subject may be a vertebrate, e.g., a mammal, such as a mouse, a primate, a monkey, or a human. Animals include livestock, sport animals, and pets. The subject may be a healthy individual, an individual with symptoms or signs, or suspected of having a disease or predisposed to a disease, or an individual in need of treatment or suspected of needing treatment. In some embodiments, the subject is a human, e.g., a human with or suspected of having cancer.

[0020] The phrase "cell-free nucleic acids" can refer to unencapsulated nucleic acids obtained from a subject's body fluids (e.g., blood, urine, CSF, etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-binding RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or partially double-stranded and single-stranded. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Some are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA. Cell-free fetal DNA (cffDNA) is the fetal DNA that circulates freely in maternal bloodstream.Cell-free nucleic acid can have one or more related epigenetic modifications, for example, acetylation, 5-methylation, ubiquitination, phosphorylation, sumoylation, ribosylation and / or citrullination.In some embodiments, cell-free nucleic acid is cfDNA, which usually comprises double-stranded cfDNA.

[0021] The terms "alignment" and "aligning" can refer to arranging DNA or RNA sequences to identify regions of similarity. Similarity can be related to functional, structural, and / or evolutionary relationships between sequences. Aligning DNA sequences involves aligning the genomic DNA of one sequence with the genomic DNA of at least one other sequence. Such alignment can exclude non-genomic DNA, such as molecular barcodes and padding bases. For example, the genomic DNA of a sequence read can be aligned to the genomic DNA of a reference DNA sequence, excluding any molecular tags that may be attached to the sequence read.

[0022] As used herein, a statement that a nucleotide "corresponds to" a nucleotide in a sequence refers to the nucleotide identified upon alignment with the sequence to maximize identity using a standard alignment algorithm, such as the GAP algorithm.

[0023] As used herein, "sequence identity," "sequence homology," or "identity" refers to the number of identical or similar nucleotide bases in an alignment between two or more polynucleotide sequences. In one non-limiting example, "at least 90% identical to" refers to a percent identity of 90-100% relative to a reference polynucleotide. Identity at a level of 90% or higher refers to the fact that 10% or less (i.e., 10 out of 100) of the nucleotides in the test polynucleotide differ from those in the reference polynucleotide, assuming, for illustrative purposes, that test and reference polynucleotides of 100 nucleotide lengths are being compared. Such differences may be expressed as point mutations randomly distributed throughout the entire length of the nucleotide sequence, or they may be clustered in one or more locations of variable length up to the maximum allowable, e.g., 10 / 100 nucleotide difference (approximately 90% identity). Differences are defined as nucleic acid substitutions, insertions, or deletions.

[0024] Sequence identity can be determined by aligning nucleic acid sequences to identify regions of similarity or identity. For purposes herein, sequence identity is generally determined by alignment to identify identical bases. Alignment can be local or global. Matches, mismatches, and gaps can be identified between the compared sequences. A gap is a null nucleotide inserted between the bases of the aligned sequences, so that identical or similar bases are aligned. Generally, internal and terminal gaps can exist. Sequence identity can be determined by taking gaps into account as the number of identical bases / shortest sequence length x 100. When gap penalties are used, sequence identity can be determined without penalty for end gaps (e.g., no penalty is imposed on terminal gaps). Alternatively, sequence identity can be determined without taking gaps into account as the number of identical positions / (total length of aligned sequences) x 100.

[0025] As used herein, a "global alignment" refers to an alignment of two sequences from beginning to end, with each base in each sequence aligned only once. The alignment is generated regardless of whether there is similarity or identity between the sequences. For example, 50% sequence identity based on a "global alignment" means that 50% of the bases are the same across the entire sequence of two compared sequences, each 100 nucleotides in length. It will be understood that global alignment can also be used to determine sequence identity even if the lengths of the aligned sequences are not the same. Differences at the ends of the sequences are taken into account when determining sequence identity unless "no end gap penalty" is selected. Generally, global alignment is used for sequences that share significant similarity over most of their length. An exemplary algorithm for performing global alignment is the Needleman-Wunsch algorithm (Needleman et al. J. Mol. Biol. 48: 443 (1970)). An exemplary program for performing the alignment is Glo.Alignment, which is publicly available and available at the National Center for Biotechnology Information (NCBI) website (ncbi.nlm.nih.gov / ). bal Sequence Alignment Tool, and programs available at deepc2.psi.iastate.edu / aat / align / align.html.

[0026] As used herein, a "local alignment" is an alignment of two sequences, but only aligns portions of the sequences that share similarity or identity. Thus, a local alignment determines whether a subsegment of one sequence exists in another sequence. If there is no similarity, no alignment will be returned. Local alignment algorithms include BLAST or the Smith-Waterman algorithm (Adv. Appl. Math. 2: 482 (1981)). For example, 50% sequence identity based on a "local alignment" means that, in a full sequence alignment of two compared sequences of any length, a region of similarity or identity of 100 nucleotides in length has 50% of the bases that are the same within that region of similarity or identity.

[0027] The phrase "nucleic acid tag" refers to a short nucleic acid (e.g., less than 500, 100, 50, or 10 nucleotides in length) used to label a nucleic acid molecule to distinguish it from other nucleic acid molecules in different samples (e.g., representing a sample index) or the same sample that has undergone different types or different treatments (e.g., representing a molecular barcode). Tags can be single-stranded, double-stranded, or at least partially double-stranded. Tags can be the same length or have various lengths. Tags can be blunt-ended or have overhangs. Tags can be attached to one or both ends of a nucleic acid. Nucleic acid tags can be decoded to reveal information such as the sample of origin, type, or treatment of the nucleic acid. Tags can be used to enable the pooling and parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indexes, and the nucleic acids are then deconvoluted by reading the molecular barcodes. Additionally or alternatively, nucleic acid tags can be used to distinguish different molecules in the same sample (i.e., molecular barcodes). This includes both uniquely tagging different molecules in a sample or non-uniquely tagging molecules in a sample. In the case of non-unique tagging, molecules can be tagged using a limited number of different tags, and thus, in combination with at least one tag, different molecules can be distinguished based on their start and / or stop positions (i.e., genomic coordinates) located on a reference genome. Typically, a sufficient number of different tags are then used so that the probability that any two molecules with the same start / stop also have the same tag is low (e.g., <10%, <5%, <1%, or <0.1%). Some tags contain multiple identifiers to label samples, forms of molecules within the sample, and molecules within forms with the same start and stop points. Such tags can exist in type A1i (where the letter indicates the same sample type, the Arabic numerals indicate the forms of molecules within the sample, and the Roman numerals indicate the molecules within the forms).

[0028] The term "adapter" refers to a short nucleic acid (e.g., less than 500, 100, or 50 nucleotides in length), typically at least partially double-stranded, for ligation to either or both ends of a sample nucleic acid molecule. The adapter may include primer binding sites to enable amplification of the nucleic acid molecule flanked by the adapter at both ends, and / or a sequencing primer binding site, including a primer binding site for next-generation sequencing (NGS). The adapter may also include a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support. The adapter may also include a tag, as described above. The tag is preferably positioned relative to the primer and sequencing primer binding sites so that the tag is included in the amplicon and sequencing reads of the nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. Adapters of the same sequence, but with different barcodes, may also be ligated to each end. A preferred adaptor is a Y-shaped adaptor, one end of which is blunt-ended or tailed for joining to a nucleic acid molecule, which is also blunt-ended or tailed with one or more complementary nucleotides. Another preferred adaptor is a bell-shaped adaptor, which also has a blunt or tailed end for joining to a nucleic acid to be analyzed.

[0029] As used herein, the term "sequencing" or "sequencer" refers to any of several techniques used to determine the sequence of a biomolecule, e.g., a nucleic acid, e.g., DNA or RNA. Exemplary sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-speed sequencing, and the like. The sequencing method includes but is not limited to high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR (COLD-PCR) at lower denaturation temperature, multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, single-nucleotide synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD sequencing, MS-PET sequencing, and combinations thereof.In some embodiments, sequencing can be carried out by genetic analyzer, for example, the genetic analyzer commercially available from Illumina or Applied Biosystems.

[0030] The phrase " next-generation sequencing " or NGS refers to sequencing technology that has increased throughput compared with the traditional Sanger and capillary electrophoresis-based approach, for example, has the ability to simultaneously generate hundreds of thousands of relatively short sequence reads.Some examples of next-generation sequencing techniques include but are not limited to single-nucleotide synthesis, sequencing by ligation, and sequencing by hybridization.

[0031] The term "DNA (deoxyribonucleic acid)" refers to a chain of nucleotides containing deoxyribonucleosides, each containing one of four nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). The term "RNA (ribonucleic acid)" refers to a chain of nucleotides containing four types of ribonucleosides, each containing one of four nucleobases: A, uracil (U), G, and C. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand that is composed of nucleotides complementary to those in the first strand, the two strands bind to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence," or "nucleic acid sequencing read" refers to any information or data that indicates the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid such as DNA or RNA (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment). It should be understood that the present teachings contemplate sequence information obtained using any available type of technique, platform, or technology, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0032] A "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to several hundred monomeric units. Whenever a polynucleotide is represented by a sequence of letters, such as "ATGCCTG," it will be understood that the nucleotides are in 5'→3' order from left to right, and that "A" represents adenosine, "C" represents cytosine, "G" represents guanosine, and "T" represents thymidine, unless otherwise noted. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide that includes the base, as is common in the art.

[0033] The phrase "reference sequence" refers to a known sequence that is used for the purpose of comparison with experimentally determined sequences.For example, the known sequence can be a whole genome, a chromosome, or any segment thereof.The reference typically comprises at least 20, 50, 100, 200, 250, 300, 350, 400, 450, 500, 1000 or more nucleotides.The reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise discontinuous segments that are aligned with different regions of a genome or chromosome.In some embodiments, the reference sequence is a human genome.Reference human genomes include, for example, hG19 and hG38.

[0034] The phrase "biological sample," as used herein, generally refers to a tissue or fluid sample derived from a subject. A biological sample can be obtained directly from a subject. A biological sample can be or contain one or more nucleic acid molecules, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) molecules. A biological sample can be derived from any organ, tissue, or biological fluid. A biological sample can include, for example, a body fluid or a solid tissue sample. An example of a solid tissue sample is a tumor sample, for example, from a solid tumor biopsy. Body fluids include, for example, blood, serum, plasma, tumor cells, saliva, urine, lymph, prostatic fluid, semen, breast milk, sputum, feces, tears, and derivatives thereof. In some embodiments, the biological sample is blood or derived from blood.

[0035] The phrase "fusion sequence read" in the context of nucleic acid sequence information refers to a sequencing read that includes subsequences located in different, non-contiguous regions or loci of a given reference sequence. A "candidate fusion sequence read" is a sequence read that may be a fusion sequence read. In certain embodiments, for example, a first subsequence of a given fusion sequence read is located in a first exon of a given gene of the reference sequence, while a second subsequence of the given fusion sequence read is located in a second exon of the same gene of the reference sequence, and these first and second exons are separated by an intervening intron of the same gene of the reference sequence. In some of these embodiments, such a fusion sequence read indicates the presence of an intragenic fusion in the genome of the subject from which the given fusion sequence read was obtained. In other exemplary embodiments, a first subsequence of a given fusion sequence read is located in an exon of a first gene of the reference sequence, while a second subsequence of the given fusion sequence read is located in an exon of a different second gene of the reference sequence, and these exons are non-contiguous with each other in the reference sequence. In some of these embodiments, such fusion sequence reads indicate the presence of intragenic fusions within the genome of the subject from which the given fusion sequence read was obtained.

[0036] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. Sequence reads can be obtained by various methods known in the art.

[0037] The term "breakpoint" in the context of a nucleic acid fusion molecule or corresponding sequencing read refers to the terminal nucleotide position at the junction between the fused subsequences of a nucleic acid fusion or represented in the corresponding sequencing read. For example, a given split sequence read may include a first subsequence that is contiguous with and 5'-side of a second subsequence in the split sequence read, and the first subsequence is located at a first locus in the reference sequence that is discontinuous with a second locus in the reference sequence where the second subsequence is located. In this example, the first subsequence of the split sequence read includes a breakpoint at its 3'-terminal nucleotide, while the second subsequence of the split sequence read includes a breakpoint at its 5'-terminal nucleotide. In certain applications, breakpoints, for example, these breakpoints, are referred to as "breakpoint pairs."

[0038] The term "fusion event" refers to the fusion between two separate genes at a specific location. Examples of causes of fusion events include translocations, interstitial deletions, or chromosomal inversion events.

[0039] The terms "abfusion," "de novo fusion caller," "fusion caller," or "de novo method" refer to a fusion caller, either a DNA fusion caller or an RNA fusion caller, that identifies fusion events de novo, i.e., without prior knowledge such as that which can be obtained from a database of previously known gene fusion events.

[0040] The phrase "about" or "approximately," when applied to one or more values ​​or elements of interest, refers to a value or element that is similar to the stated reference value or element. In certain embodiments, unless otherwise stated or apparent from the context, the term "about" or "approximately" refers to a range of values ​​or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction of (greater than or less than) the stated reference value or element (except where such number exceeds 100% of possible values ​​or elements).

[0041] Where combinations, subsets, interactions, groups, etc. of components are described, it will be understood that each is specifically contemplated and described herein, even though specific reference to each of the various individual and collective combinations and permutations may not be expressly described. This applies to all parts of this application, including, but not limited to, steps in described methods. Thus, where there are various additional steps that may be performed, it will be understood that each of these additional steps may be performed in any particular configuration or combination of configurations of the described methods.

[0042] As will be appreciated by those skilled in the art, the implementation may be hardware, software, or a combination of software and hardware. Additionally, a computer program product on a computer-readable storage medium (e.g., non-transitory) having processor-executable instructions (e.g., computer software) embodied in the storage medium. Any suitable computer-readable storage medium may be utilized, including a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, a resistive memory, a non-volatile random access memory (NVRAM), a flash memory, or a combination thereof.

[0043] Throughout this application, reference will be made to block diagrams and flowcharts. It will be understood that each block of the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, respectively, can be implemented by processor-executable instructions. These processor-executable instructions can be loaded into a general-purpose computer, special-purpose computer, or other programmable data processing device to produce a machine such that a device implements the function(s) specified in the flowchart block(s) by the processor-executable instructions executing on the computer or other programmable data processing device is created.

[0044] These processor-executable instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner to produce an article of manufacture that includes processor-executable instructions for implementing the functions specified in the flowchart block(s) by the processor-executable instructions stored in the computer-readable memory. The processor-executable instructions may also be loaded into a computer or other programmable data processing apparatus to cause the computer or other programmable apparatus to perform a series of operational steps to create a computer-implemented process, such that the processor-executable instructions running on the computer or other programmable apparatus provide the steps for implementing the functions specified in the flowchart block(s).

[0045] The blocks in the block diagrams and flowcharts represent combinations of devices for performing the specified functions, combinations of steps for performing the specified functions, and program instruction means for performing the specified functions. It will also be understood that each block in the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, can be implemented by a computer system based on dedicated hardware that performs the specified functions or steps, or a combination of dedicated hardware and computer instructions.

[0046] FIG. 1 illustrates an example method 100 for processing a test sample obtained from an individual to call fusion events. The test sample can be obtained from a patient. In step 110, nucleic acid (DNA or RNA) can be extracted from the test sample. In certain embodiments, the nucleic acid comprises cell-free nucleic acid. In various embodiments, the test sample can be a sample selected from one or more of blood, plasma, serum, urine, feces, saliva, and / or combinations thereof. Alternatively, the biological sample can comprise a sample selected from one or more of whole blood, blood fraction, tissue biopsy, pleural fluid, pericardial fluid, cerebrospinal fluid, and ascites. In one embodiment, the test sample can comprise cell-free nucleic acid, examples of which are cell-free DNA and / or cell-free RNA. For example, the test sample can be a cell-free nucleic acid sample collected from the subject's blood. In one embodiment, the cell-free nucleic acid sample can be extracted from a test sample obtained from a subject known to have cancer (e.g., a cancer patient) or suspected of having cancer.

[0047] The following description of fusion calling can be applied to both DNA and RNA types of nucleic acid sequences.In various embodiments, nucleic acid is extracted from test sample by purification process.Generally, any known method in the art can be used to purify nucleic acid.For example, nucleic acid can be isolated by pelleting and / or precipitating nucleic acid in a tube.In some embodiments, nucleic acid can be further processed.For example, the cell-free nucleic acid extracted from test sample can be RNA, and then use reverse transcriptase to convert this RNA into DNA.

[0048] In some embodiments, method 100 includes step 110. In some embodiments, method 100 may begin at step 120 with nucleic acid obtained from a test sample.

[0049] Method 100 may include preparing a sequencing library at step 120. During library preparation, for example, adapters containing one or more sequencing oligonucleotides (e.g., known P5 and P7 sequences used in single-base synthesis (SBS) (Illumina, San Diego, Calif.)) for use in subsequent cluster generation and / or sequencing can be ligated to the ends of nucleic acid molecules by adapter ligation. In one embodiment, molecular barcodes can be added to the extracted nucleic acids during adapter ligation. In some embodiments, molecular barcodes are degenerate base pairs that serve as unique tags that can be used to identify sequence reads obtained from the nucleic acids. In other embodiments, molecular barcodes are selected from a limited set of molecular barcodes (e.g., 2 to 1,000,000; 2 to 100,000; 2 to 10,000; 2 to 1,000 different molecular barcode sequences). In some embodiments, the number of molecular barcodes in the set of molecular barcodes is less than the number of polynucleotides in the sample. In some embodiments with a limited number of molecular barcodes in a set, the molecular barcodes may contain non-degenerate base pairs that can be used to distinguish different molecules based on sequence information from the molecular barcodes and genomic coordinate information based on where the sequence reads are located relative to a reference sequence. In some embodiments, molecular barcodes are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of nucleic acids during adapter ligation. The attached molecular barcodes can be further replicated during amplification along with the nucleic acid, providing a means of identifying sequence reads arising from the same original nucleic acid segment in downstream analysis.

[0050] In some embodiments, step 120 may optionally include using hybridization probes to hybridize nucleic acids and / or enrich for nucleic acid fragments. For example, this is the case when generating sequence reads through a targeted gene panel or by whole-exome sequencing. Conversely, using hybridization probes to hybridize nucleic acids and / or enrich for nucleic acid fragments is not performed when generating sequence reads by whole-genome sequencing. Hybridizing nucleic acids using hybridization probes may include using hybridization probes to enrich a sequencing library for a selected set of nucleic acids. Hybridization probes can be designed to target and hybridize with target nucleic acid sequences to pull down and enrich target nucleic acid molecules that may provide information about the presence or absence of cancer (or disease), the state of the cancer, or the classification of the cancer (e.g., the type or tissue of origin of the cancer). Following this step, multiple hybridization pull-down probes can be used for a given target sequence or gene. The probes can range in length from about 40 to about 160 base pairs (bp), about 60 to about 120 bp, or about 70 to about 100 bp. In one embodiment, the probes cover overlapping portions of the target region or gene. For targeted gene panel sequencing, hybridization probes can be designed to target and pull down nucleic acid molecules derived from specific gene sequences included in the targeted gene panel. For whole exome sequencing, hybridization probes can be designed to target and pull down nucleic acid molecules derived from exon sequences in the reference genome. The hybridized nucleic acid molecules can then be enriched. For example, the hybridized nucleic acids can be captured and amplified using PCR. The target sequences are enriched to obtain enriched sequences, which can then be sequenced.For example, as is well known in the art, a biotin moiety can be added to the 5' end of a probe (i.e., biotinylation) to facilitate pull-down of the target probe-nucleic acid complex using a streptavidin-coated surface (e.g., streptavidin-coated beads). This can improve the sequencing depth of sequence reads. However, PCR is imperfect, and it introduces artifacts (e.g., skew and new hybrid or erroneous sequences) into the pool of amplified DNA molecules. For example, template crossover, the process by which two templates combine to form a new chimeric product during amplification, can generate artifacts. PCR template crossover generates a hybrid sequence of two sequences already present in the input. DNA polymerase can jump from one template to another within the region of complementarity during PCR without interrupting the nascent DNA strand. Thus, this nascent strand has a new hybrid sequence, one piece complementary to the old template and the other piece complementary to the new template. Similarly, nascent transcripts may be terminated before completion, but then serve as primers in subsequent cycles of PCR, again resulting in new hybrid species.

[0051] In some embodiments, method 100 includes steps 110 and 120. In some embodiments, method 100 may begin at step 120 using nucleic acids obtained from a test sample. In some embodiments, method 100 may begin at step 130 using a previously prepared sequence library. In some embodiments, a previously prepared sequence library may be purchased.

[0052] Method 100 may include, in step 130, sequencing the nucleic acids in the sequencing library to generate sequence reads. Sequence reads can be obtained by means known in the art. For example, some techniques and platforms directly obtain sequence reads from millions of individual nucleic acid (e.g., DNA, e.g., cfDNA or gDNA, or RNA, e.g., cfRNA) molecules in parallel. Such techniques may be suitable for performing any of targeted gene panel sequencing, whole exome sequencing, whole genome sequencing, targeted gene panel bisulfite sequencing, and whole genome bisulfite sequencing.

[0053] As a first example, single-base synthesis techniques rely on the detection of fluorescent nucleotides, which are incorporated into nascent strands of DNA complementary to the template to be sequenced. In one method, oligonucleotides 30–50 bases in length are covalently attached at their 5′ ends to a glass coverslip. These attached strands serve two functions. First, they serve as capture sites for the target template strands, if the template is constructed with a capture tail complementary to the surface-bound oligonucleotide. They also serve as primers for template-directed primer extension, which is the basis for sequence reading. The capture primers serve as fixed sites for sequencing using multiple cycles of synthesis, detection, and chemical cleavage of the dye-linker to remove the dye. Each cycle consists of adding a polymerase / labeled nucleotide mixture, rinsing, imaging the dye, and cleavage.

[0054] In an alternative method, the polymerase is modified with a fluorescent donor molecule and immobilized on a glass slide, while each nucleotide is color-coded with an acceptor fluorescent moiety attached to the gamma-phosphate. The system detects the interaction of the fluorescently tagged polymerase with the fluorescently modified nucleotide as the nucleotide is incorporated into the de novo strand.

[0055] Any suitable single-nucleotide synthesis platform can be used to identify mutations. Single-nucleotide synthesis platforms include the Genome Sequencers from Roche / 454 Life Sciences, the GENOME ANALYZER from Illumina / SOLEXA, the SOLID system from Applied BioSystems, and the HELISCOPE system from Helicos Biosciences. Single-nucleotide synthesis platforms are also described by VisiGen Biotechnologies. In some embodiments, multiple nucleic acid molecules to be sequenced are attached to a support (e.g., a solid support). To immobilize nucleic acids on a support, a capture sequence / universal priming site can be added to the 3' and / or 5' end of the template. Nucleic acids can be attached to a support by hybridizing the capture sequence to a complementary sequence covalently attached to the support. A capture sequence (also called a universal capture sequence) is a nucleic acid sequence complementary to the sequence attached to the support, which can double as a universal primer.

[0056] As an alternative to capture sequences, a member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair) can be linked to each molecule captured on a surface coated with the second member of the coupling pair. After capture, the sequence can be analyzed by single-molecule detection / sequencing, for example, using template-dependent single-base synthesis. In single-base synthesis, surface-bound molecules are exposed to multiple labeled nucleotide triphosphates in the presence of a polymerase. The sequence of the template is determined by the order of labeled nucleotides incorporated at the 3' end of the growing strand. This can be done in real time or in a step-and-repeat fashion. For real-time analysis, each nucleotide can be incorporated with a different optical label, and multiple lasers can be used for stimulation of the incorporated nucleotides.

[0057] Massively parallel sequencing or next-generation sequencing (NGS) techniques include synthesis technology, pyrosequencing, ion semiconductor technology, single molecule real-time sequencing, ligation sequencing or paired-end sequencing.Examples of massively parallel sequencing platforms are Illumina HISEQ or MISEQ, ION PERSONAL GENOME MACHINE, PACBIO RSII sequencer or SEQUEL System, Qiagen's GENEREADER and Oxford MINION.Other similar current massively parallel sequencing technologies and future generations of these techniques can be used.

[0058] In various embodiments, a sequence read may consist of a read pair denoted R1 and R2. For example, a first read R1 may be sequenced from a first end of a nucleic acid molecule, while a second read R2 may be sequenced from a second end of the nucleic acid molecule.

[0059] In some embodiments, in step 130, the sequence reads can be subjected to further processing. In some embodiments, rather than generating sequence reads through steps 110-130, the sequence reads can be obtained, downloaded, determined, received, etc. from any available data source. The sequence reads can be obtained, downloaded, determined, received, etc., from, for example, whole exome sequencing (WES) data (DNA-seq), whole genome sequencing (WGS) data (DNA-seq), and / or transcriptome sequencing (RNA-seq) data. The described methods and systems can obtain sequence reads in one of a variety of formats (e.g., FASTA, FASTQ, and / or other proprietary formats), depending, for example, on the sequencing platform used to generate the sequence reads. Thus, obtaining sequence reads from a sequencing platform can include standardizing the read format so that the sequence reads can be used for further processing and analysis as described herein. One non-limiting example of standardizing the sequence format is adjusting the quality score format of the sequence reads. In some embodiments, the structure of a data file containing sequence reads can be optimized to improve (e.g., speed up or make more efficient) searching of the data file.

[0060] Further processing may include, for example, a pre-filtering step to remove sequence reads, stitching of read pairs, and / or overhang trimming of read pairs. Pre-filtering may include removing sequence reads that meet one or more criteria. Examples of criteria include, but are not limited to, identifying whether a sequence read is a singleton, identifying whether a sequence read is hard-clipped, filtering based on template length (TLEN) (e.g., a threshold TLEN), filtering based on alignment scores (e.g., a threshold alignment score), or filtering based on base quality scores (e.g., a median or mean base quality score threshold). Another criterion includes determining to keep a sequence read pair and not filter it out if the reads of the read pair meet the criterion that the reads of the read pair are from different chromosomes. Further examples of criteria include filtering based on bit flags, cigars, edit distance (e.g., minimum or maximum edit distance), suboptimal alignment scores, or complementary alignment measures.

[0061] 2A, 2B, and 2C depict an example of a stitching and trimming process for generating fragment s205 from read pair r1210A and r2210B, according to an embodiment.

[0062] As shown in Figures 2A, 2B, and 2C, r1210A and r2210B are represented as arrows pointing toward each other, indicating forward and reverse complementary strands. The read pair (r1, r2) is evaluated to determine whether they need to be stitched to the same fragment s205, i.e., r1 and r2 are decomposed into kmers, and each common kmer anchors the suffix-prefix alignment of r1210A and r2210B (Figure 2A). If the alignment similarity passes a certain threshold, stitching is applied. As shown in Figure 2A, the overlap region 220 between the read pair indicates one of the shared kmers (e.g., overlap) between them, which is the anchor for the suffix-prefix alignment. Thus, the stitched fragment s205 is a concatenation of the prefix, overlap, and suffix of r1210A and r2210B. Occasionally, the stitching code fuses long molecules with perfect repeats, resulting in artifacts that resemble fusions. As shown in Figure 3, read mates are stitched de novo, but adjacent perfect repeats can cause long molecules to be stitched incorrectly.

[0063] In another scenario, if the 3' end of r1 / r2 extends beyond the 5' end of r1 / r2 (an overhang), fragment s205 becomes an overlap region. This is the scenario shown in Figure 2B, where r1210A and / or r2210B extend beyond the 5' region of the other read. The overhang is trimmed, and fragment s205 is an overlap.

[0064] In another scenario, as shown in Figure 2C, if r1210A and r2210B could not be stitched because they either did not overlap and / or there were too many sequencing errors, the paired reads would be joined to form fragment s205, where the reverse complement r2210B would convert both reads to the same strand. Non-alphabetic characters not contained in either kmer are arbitrarily selected to prevent the generation of non-existent kmers from the data.

[0065] Method 100 may include using computational analysis to process sequence reads to call fusion events in step 140. Such computational analysis is now described with respect to Figure 4, which depicts a method 400 of identifying fusion events according to an embodiment. Generally, the computational analysis is a de novo fusion caller configured to predict the presence of fusion events in an individual without prior knowledge.

[0066] Method 400 may include determining candidate fusion sequence reads in step 410, generating contigs from the candidate fusion sequence reads in step 420, determining candidate fusion events in step 430, and determining fusion events in step 440.

[0067] Determining candidate fusion sequence reads in step 410 may include aligning multiple sequence reads to a reference sequence. The reference sequence may include the DNA sequence of an entire genomic region, such as a chromosome. A reference sequence including the DNA sequence of an entire genomic region can be used to identify candidate fusion events affecting that particular genomic region. The reference sequence may include exon DNA sequences. Thus, the reference sequence can be used to identify candidate fusion events affecting exon DNA sequences. In some embodiments, the reference sequence may include intron DNA sequences in addition to exon DNA sequences. Thus, the reference sequence can be used to identify candidate fusion events affecting both exon and intron DNA sequences. In some embodiments, the reference sequence may include a combination of exon DNA sequences, intron DNA sequences, and additional nucleotide bases in a padding region. The padding region may be a nucleic acid sequence known to be unlikely to be associated with gene fusion events, such as a repetitive nucleic acid sequence or other intron region. Thus, the reference sequence can be used to identify candidate fusion events affecting exon DNA sequences, intron DNA sequences, as well as junctions between exon / intron DNA sequences.

[0068] Alignment of multiple sequence reads and a reference sequence can include any alignment technique known in the art. Examples of alignment techniques include, but are not limited to, pairwise alignment and multiple sequence alignment. Pairwise alignment can include, for example, exhaustive or heuristic (e.g., non-exhaustive) pairwise alignment. Exhaustive pairwise alignment, sometimes referred to as a "brute force" approach, calculates an alignment score for every possible alignment between every possible pair of sequences in a set. Multiple sequence alignment can include progressive alignment, such as that implemented by the program ClustalW (see, for example, Thompson, et al., Nucl. Acids. Res., 22:4673-80 (1994)). The alignment result can include one or more binary alignment map (BAM) files.

[0069] Determining candidate fusion sequence reads in step 410 may further include determining one or more breakpoints in the alignment of at least one sequence read of the plurality of sequence reads to the reference sequence. Any sequence read associated with one or more breakpoints in the alignment can be identified as a candidate fusion sequence read. A breakpoint may be a region or point where the sequence read varies from the reference sequence. The alignment of each sequence read may contribute to one or more breakpoints. A breakpoint may be an orientation position on a chromosome. The presence of a breakpoint in the alignment may indicate either an error in the sequencing process or a genuine signal for a true fusion event. Figure 5 shows an example of a sequence read 510 determined to be a candidate fusion sequence read. The sequence read 510 is aligned to a reference sequence 520. A first portion 530 of the sequence read 510 is successfully aligned to the reference sequence 520, but a second position 540, starting at a breakpoint 550, is not successfully aligned to the reference sequence 520. A sequence read 510 can be considered a candidate fusion sequence read based on the presence of breakpoint 550. Although not shown in Figure 5, additional breakpoints are generated from other alignments of the same sequence read 510.

[0070] In some embodiments, one or more BAM files can be queried to determine sequence reads that should be discarded and / or considered as candidate fusion sequence reads. BAM files can be scanned, and any logical sequence reads can be discarded. Logical sequence reads can include reads that do not appear to contain fusion events (e.g., not hard-clipped or soft-clipped). In some embodiments, a minimum alignment length and / or a maximum alignment length can be used to identify logical sequence reads. The minimum alignment length can be, for example, 1 to 100, inclusive. In some embodiments, the minimum alignment length can be 40. The maximum alignment length can be, for example, 600 to 1000, inclusive. In some embodiments, the maximum alignment length can be 800. Any sequence reads aligned to a reference sequence that contain several bases less than the minimum alignment length or more than the maximum alignment length are not considered logical sequence reads and can be retained for further analysis. In some embodiments, sequence reads associated with a low mapping quality score (MAPQ) can be discarded. A low mapping quality score can be, for example, anywhere between 0 and 60, inclusive. In one embodiment, a low mapping quality score can be 50 or less. Sequence reads containing indels longer than a threshold can be retained as candidate fusion sequence reads. The threshold can be, for example, anywhere between 15 and 30 bases, inclusive. In one embodiment, the threshold can be 24 bases. FIG. 6 shows an example of a sequence read 610 determined to be a candidate fusion sequence read. The sequence read 610 has two alignments to a reference sequence 620: a primary alignment 630 (soft-clipped bases), in which portions of the sequence read 610 do not match the reference sequence 620 well on either side of the sequence read 610; and a secondary alignment 640 (hard-clipped bases), in which the sequence read 610 may align reasonably well to more than one position in the reference sequence 620 and includes portions of the sequence read 610 that were removed prior to alignment.

[0071] Returning to FIG. 4 , generating contigs from candidate fusion sequence reads in step 420 may include grouping candidate fusion sequence reads into groups (or "containers" or "packets") based on one or more common breakpoints, and assembling candidate fusion sequence reads within each packet into one or more contigs. Candidate fusion sequence reads that share the same or adjacent breakpoints (e.g., common breakpoints) can be placed in the same packet / container. In an embodiment, a common breakpoint may be 1) a breakpoint in each of two candidate fusion sequence reads that are present on the same chromosome in the same orientation, and / or 2) a breakpoint in each of two candidate fusion sequence reads at the same position or within a threshold number of bases (e.g., within a threshold of any of 1 to 40 bases (inclusive), e.g., 12 bases) and having the same orientation. In another embodiment, a compatibility test for two vectors of breakpoints can be performed.

[0072] 7 illustrates a scenario in which one candidate fusion sequence read includes a single breakpoint and another candidate fusion sequence read includes multiple breakpoints. The first candidate fusion sequence read includes breakpoint 710, and the second candidate fusion sequence read includes breakpoint 720, breakpoint 730, and breakpoint 740. Breakpoints 720 and 740 are not located within a threshold number of bases from the position of breakpoint 710, and therefore do not contribute to grouping the first candidate fusion sequence read and the second candidate fusion sequence read. However, the positions of breakpoints 710 and 730 are within a threshold number of bases and can serve as a basis for grouping the first candidate fusion sequence read and the second candidate fusion sequence read into the same packet.

[0073] FIG. 8 illustrates a scenario in which one candidate fusion sequence read contains multiple breakpoints, and another candidate fusion sequence read also contains multiple breakpoints. The first candidate fusion sequence read contains breakpoint 810, breakpoint 820, and breakpoint 830. The second candidate fusion sequence read contains breakpoint 840, breakpoint 850, and breakpoint 860. A comparison can be made between each breakpoint of the first candidate fusion sequence read and each breakpoint of the second candidate fusion sequence read. As shown in FIG. 8, breakpoint 810 and breakpoint 840 are located within a threshold number of bases, and breakpoint 830 and breakpoint 860 are located within a threshold number of bases. These paired breakpoints can serve as the basis for grouping the first candidate fusion sequence read and the second candidate fusion sequence read into the same packet. However, breakpoint 820 and breakpoint 860 are not within the threshold number of bases of any other breakpoints, and therefore do not contribute to the grouping of the first candidate fusion sequence read and the second candidate fusion sequence read.

[0074] In some embodiments, packets of candidate fusion sequence reads can be computationally generated by constructing one or more container data structures. In some embodiments, the one or more container data structures can include one or more graph data structures. The graph data structure can include nodes representing candidate fusion sequence reads and edges connecting nodes representing matching candidate fusion sequence reads. Each connected node can be considered part of a packet. Graph data structure construction can be parallelized given the computationally intensive nature of such construction.

[0075] A graph data structure may include a type of data structure in which pairs of vertices (also called nodes) are connected by edges. In one embodiment, the graph data structure is stored in a memory subsystem ( FIG. 21 , memory 2107), which may include a pointer to identify the physical location in memory 2107 where each vertex is stored. Typically, each node in a graph data structure represents an element in a set, while an edge represents a relationship between the elements. Graph data structures may include directed graphs, trees, and / or directed acyclic graphs (DAGs), etc. A directed graph is a graph in which the edges have directions. A tree is a type of directed graph data structure that has a root node and several additional nodes, each of which is either an interior node or a leaf node. The root node and interior nodes each have one or more "child" nodes, each of which is called the "parent" of its child nodes. Leaf nodes do not have any child nodes. Edges in a tree are conventionally directed from parent to child. In a tree, a node has exactly one parent. A generalization of a tree known as a directed acyclic graph (DAG) allows a node to have multiple parents, but does not allow edges to form cycles.

[0076] In one embodiment, the graph data structure may represent a de Bruijn graph. Bruijn graphs reduce computational effort by breaking down reads into smaller DNA sequences called k-mers, where the parameter k indicates the length of these sequences in base pairs. In de Bruijn graphs, all reads are broken down into k-mers (all subsequences of length k within a read) and paths between the k-mers are calculated. Assembly using this method represents reads as paths through the k-mers. The de Bruijn graph captures the overlap of length k-1 between these k-mers, not between the actual reads. Thus, for example, the sequence CATGGA can be represented as a path by the following 2-mers: CA, AT, TG, GG, and GA. Other k-mers, such as 1-mer, 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, and 8-mer, are contemplated. The de Bruijn graph approach deals well with redundancy, making computation of complex paths tractable. By reducing the entire dataset to k-mer overlaps, de Bruijn graphs reduce the high redundancy in short-read datasets. The most efficient k-mer size for a particular assembly can be determined by the read length and error rate. The value of the parameter k has a significant impact on the quality of the assembly. A good value can be estimated before assembly, or an optimal value can be found by testing a small range of values.

[0077] In an embodiment, each of the candidate fusion sequence reads can include a string of symbols. For example, the string s can be an alphabet [ka] The length of s is denoted by |s|. A substring of s is a string that occurs in s, has a starting position i and length l, and is denoted by s(i,l). A substring of length l is also denoted as an l-mer. In the following, [ka] is the DNA alphabet [Chem.] is assumed, and these symbols have complements, and (A,T) and (C,G) are complementary pairs. The reverse complement string [Chem.] <00003�18> is the reverse order of the complementary symbols in s. The canonical string [Chem.] is the lexicographically smallest of s and its reverse complement [Chem.] . The minimum solution of l-mer x is the g-mer y that exists in x, so g < l, and y is the lexicographically smallest of all g-mers in x. The lexicographical order is cumbersome to use because g-mers of polyA occur naturally in sequencing data and are often replaced in a random order. The simplest way to obtain a random order is to compute the hash value for each g-mer in x with a computer and select the g-mer with the smallest hash value as the minimum solution. In certain embodiments, the minimum solution resulting from random ordering can be used.

[0078] The de Bruijn graph (dBG) can be a directed graph G = (V,E) where each vertex v ∈ V represents a k-mer. A directed edge e ∈ E from vertex v representing k-mer x to vertex v’ representing k-mer x’ exists if and only if x(2,k - 1) = x’(1,k - 1). Each k-mer x has [Chem.] possible next nodes [Chem.]<� , where [ka] and [ka] is the concatenation operator. In the definition of the combination of elements in dBG, the alphabet [ka] Note that while all possible k-mers for v exist in the graph, in this embodiment the definition is restricted to a subset of the de Bruijn graph that represents the k-mers in the input. A path in the graph is a set of distinct connected vertices p=(v,...,v m ) A path p has an end vertex v1 that may have more than one incoming edge and a source vertex v1 that may have more than one outgoing edge. mA graph is unbranched if all its vertices have in-degree and out-degree of 1, except for . An unbranched path is maximal if it cannot be expanded in the graph without branching. A compressed de Bruijn graph (cdBG) merges all maximal unbranched paths of η vertices from a dBG into a single vertex, called a unitig, representing word length k + η − 1. Examples of the minimum dBG and cdBG are provided in Figures 9A and 9B, respectively. Traditional techniques for generating graph data structures include Bloom filters. However, Bloom filter data structures trade poor data locality for memory usage and time complexity with a reduced false positive rate because bits corresponding to one element are scattered across a bitmap, resulting in some CPU cache misses during insertion and querying. To overcome these technical limitations, some embodiments may use a rolling hash function to select a g-mer as the smallest solution within a single k-mer. Because overlapping k-mers may share a minimal solution, a minimum-up approach can be used to compute the minimal solution at amortized O(1) cost, such that iteration of the minimal solution for adjacent k-mers in a sequence is linear in the sequence length. Another optimization that can be implemented is to restrict the computation of the minimal solution to a subset of the k-mer's g-mers, i.e., excluding the first and last g-mers from the list of candidates for the minimal solution. This ensures that for a given k-mer, all of its preceding and succeeding neighboring k-mers share the same minimal solution. In addition to the high likelihood that k-mer x and its neighbor x' will share a minimal solution, this neighborhood hashing approach ensures that when all preceding and succeeding neighbors of x are searched, they will all have the same minimal solution and will be stored in the same block, thereby minimizing cache misses.

[0079] In one embodiment, an adjacency technique is used to store a graph data structure (e.g., representing a dBG or cdBG) in a memory subsystem (e.g., FIG. 21, memory 2107), which may include pointers to identify the physical location in memory 2107 where each vertex is stored. In one embodiment, an adjacency list is used to store the graph data structure in memory 2107. In some embodiments, there is an adjacency list for each vertex.

[0080] 10 shows a graph data structure 1000, including vertex objects 1005 and edge objects 1009. Portions of a sequence (e.g., a k-mer) are identified as blocks, and the blocks are converted into objects 1005 that are stored in a tangible memory device. Note that this object may potentially be stored using a single byte of information. For example, if A=00, C=01, G=10, and T=11, then the block representing the string "AGTT" contains 00101111 (1 byte). The objects 1005 are connected to create paths such that there is a path for each of the candidate fusion sequences. The paths are directed, in the sense that the direction of each path corresponds to the 5' to 3' orientation of the nucleic acid. Note, however, that it may be convenient or desirable to represent the sequence in a 3' to 5' direction, and doing so does not depart from the scope of the present invention. The connections that make up a path can themselves be implemented as objects, such that blocks are represented by vertex objects 1005 and connections are represented by edge objects 1009. Thus, a directed graph includes vertex and edge objects stored in a tangible memory device. The graph data structure 1000 can represent multiple candidate fused sequences, since each of the original candidate fused sequences can be retrieved by reading the path in the path's direction. However, the graph data structure 1000 differs from the original candidate fused sequences in that at least the portions of the sequences that match each other when aligned have been converted into a single object. Candidate fused sequence strings can be stored in either vertex objects 1005 or edge objects 1009 (nodes and vertices are used synonymously). As used herein, node objects 1005 and edge objects 1009 refer to objects created using a computer system.

[0081] 10 further illustrates the use of an adjacency list 1001 for each vertex 1005. The disclosed method and system can use a processor to create a graph data structure 1000 that includes vertex objects 1005 and edge objects 1009 through the use of adjacencies, e.g., adjacency lists or index-free adjacency. For example, the processor can use index-free adjacency to create a graph data structure 1000 in which a vertex 1005 includes a pointer to another vertex 1005 to which it is connected, the pointer identifying a physical location on the memory device 1807 where the connected vertex is stored. The graph data structure 1000 can be implemented using adjacency lists such that each vertex or edge stores a list of such objects to which it is adjacent. Each adjacency list includes a pointer to a specific physical location in the memory device for the adjacent object.

[0082] The graph data structure 1000 is typically stored on a physical device in the memory subsystem 1807 in a manner that provides very fast traversal. In that sense, the bottom portion of Figure 10 depicts objects being stored in specific physical locations on a tangible portion of the memory subsystem 1807. Each node 1005 is stored in a physical location, and that location is referenced by a pointer in any adjacency list 1001 that references that node. Each node 1005 has an adjacency list 1001 that contains every adjacent node in the graph data structure 1000. The entries in the list 1001 are pointers to adjacent nodes.

[0083] In one particular embodiment, there is an adjacency list for each vertex and edge, and the adjacency list for a vertex or edge lists the edges or vertices to which that vertex or edge is adjacent.

[0084] Figure 11 illustrates the use of an adjacency list 1101 for each vertex 1005 and edge 1009. As shown in Figure 11, the disclosed method and system can create a graph data structure 1000 using an adjacency list 1001 for each vertex and edge, where the adjacency list 1001 for a vertex 1005 or edge 1009 lists the edges or vertices to which that vertex or edge is adjacent. Each entry in the adjacency list 1101 is a pointer to an adjacent vertex or edge.

[0085] Each pointer identifies a physical location in the memory subsystem where the adjacent object is stored. In a preferred embodiment, a pointer, or native pointer, can be manipulated as a memory address because it points to a physical location in memory and allows access to the intended data by dereferencing the pointer. That is, a pointer is a reference to data stored somewhere in memory, and obtaining that data is a matter of dereferencing the pointer. The characteristic that separates a pointer from other kinds of references is that the pointer's value is interpreted, at a low or hardware level, as a memory address. Such a graph representation provides a means for fast random access, modification, and data retrieval.

[0086] In some embodiments, graph object storage is implemented with index-free contiguity, where every element contains direct pointers to its neighboring elements, thereby eliminating the need for index lookups and making traversal very quick, supporting fast random access. Index-free contiguity is another example of low-level, or hardware-level, memory references for data retrieval. Specifically, index-free contiguity can be implemented such that the pointers contained within elements are references to physical locations in memory.

[0087] Because technical implementations that use physical memory addressing, such as native pointers, can access and use data in such a lightweight manner without requiring a separate index table or other intervening lookup step, the performance of a given computer, e.g., any modern consumer-grade desktop computer, is extended to enable full operation of genome-scale graphs (e.g., container data structures such as graph data structure 1000 representing a set of candidate fusion sequences). Thus, by storing graph elements (e.g., nodes and edges) using a library of objects with native pointers, or other implementations that provide index-free adjacency, the capabilities of technologies that provide storage, retrieval, and alignment of genomic information are actually improved, as this uses the computer's physical memory in a specific way.

[0088] In some embodiments, an error correction procedure can be performed on candidate fusion sequence reads in a given packet / container. The error correction procedure is designed to reduce the likelihood that a non-fusion event will be identified as a fusion event. In some embodiments, indels exceeding or equal to a threshold number of bases can be exempt from the error correction procedure. The threshold number of bases can be anywhere between 20 and 30 bases, inclusive. In some embodiments, the threshold number of bases can be 24 bases. Figure 12 illustrates an error correction procedure that replaces mismatches or local differences (e.g., variants) with corresponding bases from a reference sequence. Figure 13 illustrates the error correction procedure applied to two candidate fusion sequence reads that align to a reference sequence within the threshold number of bases. One candidate fusion sequence read includes several padding bases. Gaps between the two candidate fusion sequence reads can be filled using bases from the reference sequence at the same positions as the gap. In some embodiments, the padding bases can be retained or replaced with bases from the reference sequence at the same positions as the padding bases. Several padding bases can be inserted between two candidate fusion sequence reads to join the two candidate fusion sequence reads into a single read. Figure 14 illustrates an error correction procedure that discards candidate fusion sequence reads with unaligned portions exceeding a threshold. For example, any candidate fusion sequence reads with unaligned portions exceeding or equal to a threshold percentage of the candidate fusion sequence reads can be excluded. In one embodiment, the threshold percentage can be anywhere from 1% to 99%, inclusive. In one embodiment, the threshold percentage can be 10%, meaning that any candidate fusion sequence reads with 10% or more unaligned bases can be discarded. The actual result can be the exclusion of candidate fusion sequence reads containing soft-clipped bases. Figure 15 further illustrates the error correction procedure of Figure 14, in which candidate fusion sequence reads with unaligned portions exceeding a threshold are excluded.

[0089] Assembling the remaining candidate fused sequence reads in each packet / container into one or more contigs can include any known contig assembly method. For example, assembly by alignment can proceed by aligning sequence reads to each other or to a reference. For example, each read can be aligned to a reference genome in turn to place all of the reads in relation to each other to create an assembly. In one embodiment, the container data structure for each packet can include a graph data structure representing a de Bruijn graph, and assembling the candidate fused sequence reads for each packet into contigs can include linearizing the de Bruijn graph to output contigs for each packet. For example, a greedy algorithm can be used to select the edges of the de Bruijn graph that are most represented by the sequence reads.

[0090] Returning to FIG. 4 , determining candidate fusion events in step 430 may include aligning contigs from each packet to a reference sequence and determining one or more candidate fusion events based on the alignment. In an embodiment, contigs from a packet may be aligned to a reference sequence (with a decoy), and candidate fusion sequence reads for the packet may be aligned to the contigs. Candidate fusion sequence reads for a packet may be clustered into families. A family may include candidate fusion sequence reads associated with the same molecule. A family may be determined based on molecular barcoding. Candidate fusion sequence reads containing the same molecular barcode may be grouped into the same family. In an embodiment, sequence reads containing the same molecular barcode and whose alignment begins within a number of bases (e.g., 30-50 bases) of each other may be grouped into the same family. One or more tests may be applied to the resulting alignment to determine candidate fusion events. The one or more tests may include a footprinting test and / or a variability test. The footprinting test may include determining that a threshold number of a family of candidate fusion sequence reads supporting a contig spans a breakpoint. The threshold value can be, for example, anywhere from 2 to 5 families (inclusive). In some embodiments, the threshold value can be 2 families. In some embodiments, the threshold value can be 3 families. The variability test can include determining that a threshold amount of variability exists between sequence reads of at least two families of candidate fusion sequence reads that support the contig and span the breakpoint. In some embodiments, the variability test includes aligning each sequence read to the contig. Then, for each sequence read, start and stop coordinates on the contig for the first and last base are calculated by computer. The average and standard deviation for all start points for each sequence read are calculated to create an average start point and start standard deviation. The average and standard deviation for all stop points for each sequence read are calculated to create an average stop point and stop standard deviation.The variability can then be defined as the smallest or lowest standard deviation between the start and stop standard deviations. It will be understood, therefore, that in some embodiments, only the standard deviation is used to define the variability test. The threshold for the variability test can be 1 to 15 bases inclusive. In some embodiments, the threshold can be 8 bases. If the variability is less than 8, the fusion fails the variability test and is discarded. In some embodiments, the threshold can be 7 bases. In some embodiments, the threshold can be 6 bases. In some embodiments, the threshold can be 5 bases.

[0091] The footprinting test is shown in Figure 16. Figure 16 shows contig 1610 aligned to a first portion of reference sequence 1620 and a second portion of reference sequence 1630. Breakpoint 1640 exists between the aligned portions. Candidate fusion sequence reads supporting the contig are shown as candidate fusion sequence read 1650, candidate fusion sequence read 1660, candidate fusion sequence read 1670, and candidate fusion sequence read 1680. Candidate fusion sequence read 1650 belongs to a first family, candidate fusion sequence read 1660 belongs to a second family, and candidate fusion sequence read 1670 and candidate fusion sequence read 1680 belong to a third family. As shown in Figure 16, at least two families of candidate fusion sequence reads supporting the contig span breakpoint 1640, resulting in breakpoint 1640 being identified as a candidate fusion event.

[0092] The variability test is shown in Figure 17. As shown, for each sequence read 1650-1680, the start and stop coordinates on contig 1610 for the first and last bases can be determined. The mean and standard deviation for all of the start points for each sequence read 1650-1680 can be determined, resulting in the average start point and start standard deviation. Similarly, the mean and standard deviation for all of the stop points for each sequence read 1650-1680 can be determined, resulting in the average stop point and stop standard deviation. The variability (1710, 1720) can then be defined as the minimum or lowest standard deviation between the start standard deviation and the stop standard deviation. The threshold for the variability test can be 1 to 15 bases (inclusive). In one embodiment, the threshold can be 8 bases. If the variability (1710, 1720) is less than 8, the fusion fails the variability test and is discarded. In one embodiment, the threshold can be 7 bases. In one embodiment, the threshold may be 6 bases.

[0093] 4, determining fusion events in step 440 may include applying one or more criteria to one or more candidate fusion events and determining one or more fusion events based on the application of the one or more criteria. Any candidate fusion events remaining after application of the one or more criteria may be identified as fusion events.

[0094] The one or more criteria may include, for example, the proximity of the candidate fusion event to the probe. At least one candidate fusion event (e.g., breakpoint) must be within the distance of the probe used in the sample enrichment step, or the candidate fusion event is discarded. By way of example, the distance may be anywhere from 250 to 500 bases, inclusive. In some embodiments, the distance may be 300 bases. In some embodiments, the distance may be 350 bases. In some embodiments, the distance may be 400 bases. In some embodiments, the distance may be 450 bases.

[0095] The one or more criteria may include, for example, applying a whitelist. A whitelist of genes may be determined. If a candidate fusion event (e.g., a breakpoint) is not associated with one of the genes in the whitelist, the candidate fusion event is discarded.

[0096] The one or more criteria may include, for example, the application of a blacklist. A blacklist of genes may be determined. If a candidate fusion event (e.g., a breakpoint) is associated with one of the genes in the blacklist, the candidate fusion event is discarded.

[0097] The one or more criteria may include, for example, filtering out certain indels. If a candidate fusion event (e.g., a breakpoint) is an indel completely buried within an intron region, the candidate fusion event is discarded. If a candidate fusion event (e.g., a breakpoint) is a deletion and is shorter than a threshold number of bases, the candidate fusion event is discarded. The threshold number of bases may be anywhere from 10 to 100 bases, inclusive. In one embodiment, the threshold number of bases may be 50 bases. If a candidate fusion event (e.g., a breakpoint) is a deletion and is within a threshold distance of another deletion, the candidate fusion event is discarded. The threshold distance may be anywhere from 10 to 100 bases, inclusive. In one embodiment, the threshold distance may be 49 bases. In one embodiment, the threshold distance may be 48 bases. In one embodiment, the threshold distance may be 47 bases. In one embodiment, the threshold distance may be 46 bases. In one embodiment, the threshold distance may be 45 bases.

[0098] The one or more criteria may include, for example, determining whether the ratio of molecules to reads exceeds a threshold and whether there are double-stranded support molecules (double-stranded support molecules are defined as molecules with two or more reads on each strand). The threshold may be anywhere from 0.5 to 0.9, inclusive. In some embodiments, the threshold may be 0.8. In some embodiments, the threshold may be 0.7. In some embodiments, the threshold may be 0.6. In some embodiments, the threshold may be 0.5. If the ratio associated with a candidate fusion event is greater than and / or equal to the threshold, the candidate fusion event is discarded.

[0099] The one or more criteria may include, for example, determining whether a candidate fusion event is a stitching artifact. A stitching artifact may be a long molecule stitched across a short repeat (introducing an artificial deletion event). The stitching process may fuse a long molecule with a perfect repeat, resulting in a stitching artifact that may be classified as a candidate fusion event. As shown in Figure 3, adjacent perfect repeats on two sequence reads may cause a long molecule to be stitched incorrectly. To address this issue, several bases of the reference sequence flanking the breakpoint may be aligned with each other, and if the alignment score is greater than or equal to a threshold score, the candidate fusion event may be discarded. The number of bases may be anywhere between 80 and 160, inclusive. In one embodiment, the number of bases may be 120. The threshold score may be anywhere between 60 and 80, inclusive. In one embodiment, the threshold score may be 70.

[0100] One or more criteria may include, for example, determining whether a candidate fusion event is a template crossover artifact. Template crossover is an artifact that occurs during sequence library preparation due to sequence similarity. This problem is similar to stitching artifacts. To address this problem, several bases of the references centered around the two breakpoints can be aligned with each other, and if the alignment score is greater than or equal to a threshold score, the candidate fusion event can be discarded. The threshold score can be anywhere from 10 to 30 (inclusive). In one embodiment, the threshold score can be 20.

[0101] Determining alignment score is well known in the art.Sequence alignment can use algorithm to establish the similarity between two sequences.For example, a positive number can be assigned to each match of sequence, and a negative number can be assigned to each mismatch of sequence.The sum of these numbers can then be used as alignment score.Programs such as Basic Local Alignment Search Tool (BLAST), MUSCLE, Mauve, MAFFT, Clustal Omega, Jotun Hein, Wilbur-Lipman, Martinez Needleman-Wunsch, Lipman-Pearson, Kalign, MView and EMBOSS Cons can be used to determine alignment score.

[0102] The one or more criteria can include, for example, determining that the candidate fusion event contains a suitable number of non-singleton support molecules. Singleton support molecules are sequence molecules with a family size of 1, and the compatibility test can check for the presence of one or more non-singleton molecules, or for the presence of two or more non-singleton molecules, or for the presence of a predefined number or more non-singleton molecules.

[0103] The above-mentioned method and system for determining fusion events differs from the typical technique of relying only on the alignment of input reads to reference genome to identify mismatched alignments that may be the result of fusion events.If relying only on alignment, once fusion-supporting reads are misaligned, they can no longer be recovered downstream, thereby leading to false-positive fusion calls.Furthermore, the present method and system can quickly and accurately identify fusion events, and can reduce time and complexity compared to previous systems.

[0104] Fusion detection is an important aspect of the oncology pipeline. It is known that tumors rearrange portions of their genome to either enhance tumor functions they require or suppress the functionality of tumor suppressor genes. Some drugs are specifically designed to address certain tumors driven by certain fusions. The identification of these fusions has a significant impact on the identification and selection of treatments for a given patient.

[0105] The described methods and systems generate clinically meaningful gene fusion data containing low false positives for gene fusion detection based on a subject's DNA sequence information (DNA-SEQ) and / or RNA sequence information (RNA-SEQ) dataset. The resulting annotated gene fusion data contains clinically meaningful information and highly specific gene fusion identification (e.g., low false positives) that can be used in clinical and / or R&D settings.

[0106] Methods are disclosed that use information determined by the disclosed methods (e.g., identification of a fusion event). For example, a method of treating a subject is disclosed, comprising administering a cancer therapeutic to the subject, wherein the subject has been determined to have a fusion event using one or more of the disclosed methods. In some embodiments, the subject has been determined to have cancer based on the identification of a fusion event using one or more of the disclosed methods. In some embodiments, the cancer can be any cancer associated with a fusion event. The cancer associated with a fusion event can be any cancer caused by a fusion event. For example, the cancer associated with a fusion event can be, but is not limited to, advanced urothelial cancer, prostate cancer, breast cancer, lung cancer, colon cancer, glioblastoma, liver cancer, or ovarian cancer. In some embodiments, the cancer therapeutic can be a known cancer therapeutic used to treat a specific cancer. For example, if a subject is determined to have an FGFR2 / 3 fusion event, the FDA-approved drug erdafitinib can be administered to the subject. Thus, in some embodiments, the cancer therapeutic is specific for a fusion event. A fusion event-specific cancer therapeutic can be a cancer therapeutic previously determined to effectively treat a cancer associated with a particular fusion event.

[0107] In some embodiments, a subject has previously been diagnosed with cancer (before the fusion event is known), and in this case, the identification of the fusion event using the disclosed method can allow the subject to be administered a specific cancer therapeutic drug.Thus, the identification of the fusion event using the disclosed method can enable personalized medicine.

[0108] Performance evaluation of the disclosed methods and systems relied on proxies, including AV samples and samples from healthy donors. A software package in an existing production pipeline with a fusion caller function was thoroughly validated (rather than as a de novo caller) on a selected set of fusion events. The sensitivity of abfusion is comparable to that of the fusion caller function, but abfusion is only performed on a very limited set of fusion cases.

[0109] In one example, a de novo fusion caller was used to identify FGFR2 / 3 fusions from clinical cfDNA. FGFR2 / 3 rearrangements are therapeutic targets, particularly in advanced urothelial carcinoma (aUC) treated with FDA-approved erdafitinib. Liquid biopsies are an attractive noninvasive method for identifying these fusions, but detection in cfDNA is technically challenging due to low tumor shedding levels, short molecules, and wide diversity of gene partners. To address this, a de novo fusion caller was used. A cohort of 17,718 patients with mixed cancer types (including 795 aUC patients, as well as breast, cholangiocarcinoma, colorectal, and gastric cancers) plus 276 healthy control samples previously tested with a cfDNA NGS-based assay were reanalyzed using the de novo fusion caller. Median unique molecular coverage was approximately 3,000 molecules, sequenced to a read depth of 15,000x. The samples were reanalyzed in silico using a novel algorithm: Briefly, reads aligned to candidate fusion breakpoints were assembled into a de Bruijn graph. The resulting contigs were aligned to the reference, and filters were applied to remove technical artifacts. The majority of FGFR2 (85%) and FGFR3 (66%) fusion partners in the mixed cancer cohort were observed only once, consistent with previous reports (Figure 18). FGFR3-TACC3 was the most common fusion, present in 59% of FGFR3 fusion-positive patients. De novo caller-detected partners in 36% of FGFR2 fusion-positive patients had not been previously described. In the aUC cohort, FGFR3 fusions were detected in 3.1% of patients, with 8 / 10 (80%) partner genes / intergenic regions present only once, consistent with previous reports (Figure 19). Fusions were not identified in 276 healthy control samples.In the mixed cancer cohort, common mutations that co-occurred with FGFR2 fusions enriched in patients with these fusions were FGFR2 N549K (7.1%), FGFR2 N549D (3.2%), and FGFR2 V564I (2.6%). Common mutations that co-occurred with FGFR3 fusions enriched in patients with these fusions included KRAS Q61H, which was observed in 30.6% of patients with FGFR3 fusions (Figure 20). Thus, the FGFR3 fusion prevalence observed in cfDNA from aUC patients, comparable to previous reports on tissue testing, demonstrates the feasibility of capturing targetable genomic rearrangements with plasma-based NGS. The FGFR2 / 3 fusion partners detected by the highly specific assembly-based de novo fusion caller were heterogeneous and individually low in frequency, highlighting the importance of de novo approaches.

[0110] 21 is a block diagram depicting an environment 2100 including non-limiting examples of a computing device 2101 and a server 2102 connected by a network 2103. In certain embodiments, some or all steps of any described method can be performed on a computing device described herein. The computing device 2101 can include one or more computers configured to store a fusion caller module 2104 and one or more of sequence data 2105 (e.g., sequence reads, contigs, reference sequences, standards, container data structures, graph data structures, etc.). The server 2102 can include one or more computers configured to store the fusion caller module 2104 and one or more of sequence data 2105 (e.g., sequence reads, contigs, reference sequences, standards, etc.) for remote access. The multiple servers 2102 can communicate with the computing device 2101 via the network 2103.

[0111] In terms of hardware architecture, the computing device 2101 and the server 2102 may generally be digital computers including a processor 2106, a memory system 2107, an input / output (I / O) interface 2108, and a network interface 2109. These components (2106, 2107, 2108, and 2109) are communicatively coupled by a local interface 2110. The local interface 2110 may be, for example, but not limited to, one or more buses or other wired or wireless connections as known in the art. The local interface 2110 may have additional elements, such as controllers, buffers (caches), drivers, repeaters, and receivers, which are omitted for simplicity, to enable communication. Furthermore, the local interface may include address, control, and / or data connections to enable appropriate communication between the aforementioned components.

[0112] The processor 2106 may be a hardware device for executing software, particularly that stored in the memory system 2107. The processor 2106 may be any custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with the computing device 2101 and the server 2102, a semiconductor-based microprocessor (in the form of a microchip or chipset), or generally any device for executing software instructions. When the computing device 2101 and / or the server 2102 are in operation, the processor 2106 can be configured to execute software stored in the memory system 2107, to communicate data to and from the memory system 2107, and to generally control the operation of the computing device 2101 and the server 2102 in accordance with the software.

[0113] The I / O interface 2108 can be used to receive user input from one or more devices or components and / or provide system output to one or more devices or components. User input can be provided, for example, by a keyboard and / or a mouse. System output can be provided by a display device and a printer (not shown). The I / O interface 2108 can include, for example, a serial port, a parallel port, a small computer system interface (SCSI), an infrared (IR) interface, a radio frequency (RF) interface, and / or a universal serial bus (USB) interface.

[0114] The network interface 2109 can be used to transfer and receive data from the computing device 2101 and / or the server 2102 over the network 2103. The network interface 2109 can include, for example, a 10BaseT Ethernet Adaptor, a 100BaseT Ethernet Adaptor, a LAN PHY Ethernet Adaptor, a Token Ring Adaptor, a wireless network adapter (e.g., WiFi, cellular, satellite), or any other suitable network interface device. The network interface 2109 can include address, control, and / or data connections to enable appropriate communication over the network 2103.

[0115] The memory system 2107 may include any one or combination of volatile memory elements (e.g., random access memory (RAM, e.g., DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, hard drive, tape, CD-ROM, DVD-ROM, etc.). Additionally, the memory system 2107 may incorporate electronic, magnetic, optical, and / or other types of storage media. It should be noted that the memory system 2107 may have a distributed architecture where various components are remote from each other but can be accessed by the processor 2106.

[0116] The software in memory system 2107 may include one or more software programs, each containing an ordered list of executable instructions for implementing a logical function. In the example of Figure 21, the software in memory system 2107 of computing device 2101 may include a fusion caller module 2104 (or subcomponents thereof), array data 2105, and a suitable operating system (O / S) 2111. Operating system 2111 essentially controls the execution of other computer programs and provides scheduling, input-output control, file and data management, memory management, and communication management and related services.

[0117] For purposes of illustration, application programs and other executable program components, such as the operating system 2111, are illustrated herein as separate blocks, with the understanding that such programs and components may reside at various times in different storage components of the computing device 2101 and / or the server 2102. An implementation of the fusion caller module 2104 may be stored on or transmitted via some form of computer-readable media. Any of the disclosed methods may be performed by computer-readable instructions embodied on a computer-readable medium. Computer-readable media may be any available medium that can be accessed by a computer. By way of example, and not limitation, computer-readable media may include “computer storage media” and “communications media.” “Computer storage media” may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Exemplary computer storage media may include RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer.

[0118] In an embodiment, the fusion caller module 2104 can be configured to access the sequence data 2105 and perform the method 2200 shown in Figure 22. The method 2200 can be performed, in whole or in part, by a single computer device, multiple electronic devices, etc. The method 2200 can include, in step 2201, aligning multiple sequence reads to a reference sequence.

[0119] The method 2200 may include, at step 2202, determining one or more breakpoints in an alignment of at least one sequence read of the plurality of sequence reads to a reference sequence.

[0120] Method 2200 may include, in step 2203, identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read. Identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read may include discarding alignments having a mappability score below a threshold. Identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read may include discarding alignments that are logical.

[0121] Method 2200 may include, at step 2204, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include a breakpoint on the same chromosome in the same orientation. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include a breakpoint at the same position. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include a breakpoint within a threshold number of bases from a position. The threshold number of bases from the position may be, for example, 1 to 40 bases. In one embodiment, the threshold number of bases from the position may be 10 bases. In one embodiment, the threshold number of bases from the position may be 11 bases. In one embodiment, the threshold number of bases from the position may be 12 bases. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include multiple breakpoints in the same chromosome and in the same orientation. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include multiple breakpoints at the same location. Determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints may include determining that two candidate fusion sequence reads include multiple breakpoints within a threshold number of bases from a plurality of positions. The threshold number of bases from a plurality of positions may be, for example, 1 to 40 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 10 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 11 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 12 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 13 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 14 bases. In an embodiment, the threshold number of bases from a plurality of positions may be 15 bases.

[0122] Method 2200 may include grouping candidate fusion sequence reads based on one or more common breakpoints at step 2205. Grouping candidate fusion sequence reads based on one or more common breakpoints may include generating a de Bruijn graph for the groups (e.g., for each group).

[0123] Method 2200 may include, in step 2206, assembling candidate fusion sequence reads within the groups (e.g., for each group) into one or more contigs. Assembling candidate fusion sequence reads within the groups into one or more contigs may include linearizing each de Bruijn graph to generate contigs for the group. Assembling candidate fusion sequence reads within the groups into one or more contigs may include performing one or more error correction procedures. The one or more error correction procedures may include resolving mismatches between the candidate fusion sequence reads and a reference sequence. The one or more error correction procedures may include inserting padding between at least two candidate fusion sequence reads. The one or more error correction procedures may include discarding one or more candidate fusion sequence reads having unaligned portions exceeding a threshold.

[0124] The method 2200 may include, at step 2207, aligning the contigs from the groups (e.g., for each group) to a reference sequence.

[0125] Method 2200 may include, at step 2208, determining one or more candidate fusion events based on alignment of contigs from the groups (e.g., for each group). Determining one or more candidate fusion events based on alignment of contigs from the groups may include applying one or more of a footprint test or a variability test. Applying a footprint test may include determining that a threshold number of families of candidate fusion sequence reads that support the contig span the breakpoint. Applying a variability test may include determining that a threshold amount of variability exists between at least two families of candidate fusion sequence reads that support the contig and span the breakpoint.

[0126] The method 2200 may include, at step 2209, applying one or more criteria to the one or more candidate fusion events.

[0127] Applying one or more criteria to one or more candidate fusion events may include, for the candidate fusion events (e.g., for each candidate fusion event), determining the distance between the breakpoint of one or more aligned contigs and the location of at least one probe of the panel, and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs that do not contain a breakpoint that is less than a threshold distance from the location of at least one probe of the panel. By way of example, the distance may be between 1 and 1,000 bases. In one embodiment, the distance may be 350 bases. The sequence reads used to determine the candidate fusion events (step 2201) may be derived from the enriched DNA for the panel.

[0128] Applying one or more criteria to the one or more candidate fusion events may include determining one or more genes of interest and discarding any candidate fusion events associated with aligned contigs of one or more contigs that do not contain breakpoints associated with the one or more genes of interest.

[0129] Applying one or more criteria to the one or more candidate fusion events may include determining, for the candidate fusion event, that the breakpoint of one or more aligned contigs is a deletion, and discarding any candidate fusion event associated with an aligned contig of one or more contigs that contains a deletion located within a number of bases away from another deletion.

[0130] Applying one or more criteria to the one or more candidate fusion events may include determining, for the candidate fusion event, that the breakpoint of one or more aligned contigs is a deletion, and discarding any candidate fusion event associated with the aligned contig of the one or more contigs that contains a deletion that includes fewer than a threshold number of bases.

[0131] Applying the one or more criteria to the one or more candidate fusion events may include discarding any candidate fusion events associated with an aligned contig of one or more contigs that contains an insertion or deletion that is entirely buried in an intron region.

[0132] Applying one or more criteria to the one or more candidate fusion events may include determining a molecule to read ratio for one or more aligned contigs for the candidate fusion event, and discarding any candidate fusion event associated with an aligned contig of one or more contigs that is associated with a molecule to read ratio that exceeds a threshold but is not associated with a double-stranded support molecule.

[0133] Applying one or more criteria to the one or more candidate fusion events may include, for the candidate fusion event, determining, for one or more aligned contig breakpoint pairs, sequences adjacent to the breakpoints of the breakpoint pairs, aligning the sequences adjacent to the breakpoints of the breakpoint pairs, determining an alignment score for the alignment of the sequences adjacent to the breakpoints of the breakpoint pairs, and discarding any candidate fusion events associated with the aligned contigs of the one or more contigs based on an alignment score exceeding a threshold.

[0134] Applying one or more criteria to the one or more candidate fusion events may include, for the candidate fusion event, determining, for a breakpoint pair of one or more aligned contigs, sequences centered at the breakpoint of the breakpoint pair, aligning the sequences centered at the breakpoints to each other, determining an alignment score for the alignment of the sequences centered at the breakpoints, and discarding any candidate fusion event associated with the aligned contig of the one or more contigs based on an alignment score exceeding a threshold.

[0135] The method 2200 may include determining one or more fusion events based on applying one or more criteria to the one or more candidate fusion events at step 2210. Any remaining candidate fusion events may be determined as one or more fusion events.

[0136] In an embodiment, the fusion caller module 2104 can be configured to access the sequence data 2105 and perform the method 2300 shown in Figure 23. The method 2300 can be performed, in whole or in part, by a single computer device, multiple electronic devices, etc. The method 2300 can include, in step 2310, aligning the multiple sequence reads to a reference sequence.

[0137] Method 2300 may include, at step 2320, determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence. Determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads include breakpoints in the same chromosome and in the same orientation. Determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads include breakpoints at the same position. Determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads include breakpoints within a threshold number of bases from a certain position. The threshold number of bases from a position may be, for example, 1 to 40 bases. In an embodiment, the threshold number of bases from a position may be 10 bases. In an embodiment, the threshold number of bases from a position may be 11 bases. In some embodiments, the threshold number of bases from a position may be 12 bases. Determining one or more candidate fusion sequence reads from a plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads contain multiple breakpoints in the same chromosome and in the same orientation. Determining one or more candidate fusion sequence reads from a plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads contain multiple breakpoints at the same position. Determining one or more candidate fusion sequence reads from a plurality of sequence reads based on one or more breakpoints in the alignment of the sequence reads to the reference sequence may include determining that two candidate fusion sequence reads contain multiple breakpoints within a threshold number of bases from a plurality of positions. The threshold number of bases from a plurality of positions may be, for example, 1 to 40 bases. In some embodiments, the threshold number of bases from a position may be 10 bases.In some embodiments, the threshold number of bases from a position may be 11. In some embodiments, the threshold number of bases from multiple positions may be 12.

[0138] Method 2300 may include, at step 2330, grouping the one or more candidate fusion sequence reads into one or more container data structures based on one or more common breakpoints. Breakpoints from different alignments can be assigned to a common container data structure. The one or more candidate fusion sequence reads into the one or more container data structures according to de Bruijn graph techniques.

[0139] Method 2300 may include, at step 2340, assembling one or more candidate fusion sequence reads into one or more contigs for the container data structure (e.g., for each container data structure). Assembling one or more candidate fusion reads into one or more contigs may include assembling one or more candidate fusion sequence reads into a graph data structure for the container data structure (e.g., for each container data structure) and linearizing the graph data structure to generate one or more contigs. Assembling one or more candidate fusion sequence reads into one or more contigs may include performing one or more error correction procedures. The one or more error correction procedures may include resolving mismatches between the candidate fusion sequence reads and a reference sequence. The one or more error correction procedures may include inserting padding between two or more candidate fusion sequence reads. The one or more error correction procedures may include discarding one or more candidate fusion sequence reads having unaligned portions exceeding a threshold.

[0140] Method 2300 may include, at step 2350, aligning one or more contigs to a reference sequence for the container data structure (e.g., for each container data structure). Method 2300 may further include determining one or more candidate fusion events based on the alignment of the contigs from the container data structure, which may include applying one or more of a footprint test or a variation test. Applying the footprint test may include determining that a threshold number of families of candidate fusion sequence reads that support the contig span the breakpoint. Applying the variation test may include determining that a threshold amount of variation exists between at least two families of candidate fusion sequence reads that support the contig and span the breakpoint.

[0141] Method 2300 may include, in step 2360, determining one or more aligned contigs indicative of a fusion event based on one or more criteria. Any remaining candidate fusion events may be determined as one or more fusion events. Determining one or more aligned contigs indicative of one or more fusion events based on one or more criteria may include determining the distance between the breakpoint of one or more aligned contigs and the location of at least one probe of the panel, and discarding any aligned contigs of the one or more contigs that do not contain a breakpoint that is less than a threshold distance from the location of at least one probe of the panel. By way of example, the distance may be between 1 and 1,000 bases. In one embodiment, the distance may be 350 bases. The sequence reads used to determine the candidate fusion events (step 2310) may be derived from the enriched DNA for the panel. Determining one or more aligned contigs that indicate a fusion event based on one or more criteria can include determining one or more genes of interest and discarding any aligned contigs of the one or more contigs that do not contain a breakpoint associated with the one or more genes of interest. Determining one or more aligned contigs that indicate a fusion event based on one or more criteria can include determining that the breakpoint of one or more aligned contigs is a deletion and discarding any aligned contigs of the one or more contigs that contain a deletion located within a few bases away from another deletion. Determining one or more aligned contigs that indicate a fusion event based on one or more criteria can include determining that the breakpoint of one or more aligned contigs is a deletion and discarding any aligned contigs of the one or more contigs that contain a deletion that includes fewer than a threshold number of bases.Determining one or more aligned contigs indicative of a fusion event based on one or more criteria can include discarding any aligned contigs of the one or more contigs that contain insertions or deletions completely buried in intron regions. Determining one or more aligned contigs indicative of a fusion event based on one or more criteria can include determining a molecular to read ratio for one or more aligned contigs, and discarding any aligned contigs of the one or more contigs that are associated with a molecular to read ratio that exceeds a threshold but are not associated with a double-stranded support molecule. Determining one or more aligned contigs indicative of a fusion event based on one or more criteria can include, for breakpoint pairs of one or more aligned contigs, determining sequences adjacent to the breakpoints of the breakpoint pairs, aligning the sequences adjacent to the breakpoints of the breakpoint pairs, determining an alignment score for the alignment of the sequences adjacent to the breakpoints of the breakpoint pairs, and discarding any aligned contigs of the one or more contigs based on an alignment score that exceeds a threshold. Determining one or more aligned contigs that are indicative of a fusion event based on one or more criteria may include, for a breakpoint pair of one or more aligned contigs, determining sequences centered at the breakpoint of the breakpoint pair, aligning the sequences centered at the breakpoints to each other, determining an alignment score for the alignment of the sequences centered at the breakpoints, and discarding any aligned contigs of the one or more contigs based on an alignment score above a threshold.

[0142] The method 2300 may further include generating a notification indicating a problem associated with the library preparation based on discarding any aligned contigs of the one or more contigs.

[0143] Although specific configurations have been described, the configurations herein are intended in all respects to be possible configurations, not exclusive, and are not intended to limit the scope to the specific configurations shown. Unless expressly stated otherwise, in no way is it intended that any method set forth herein be construed as requiring its steps to be performed in a particular order. Accordingly, unless a method claim actually recites the order in which the steps are to be followed, or unless the claims or the specification otherwise specifically state that the steps should be limited to a particular order, no order is intended to be inferred in any respect. This applies to all possible implicit bases of interpretation, including questions of logic regarding the arrangement of steps or operational flow; apparent meaning derived from grammatical structure or punctuation; and the number or type of structure set forth in the specification.

[0144] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of the present invention. Other configurations will be apparent to those skilled in the art from consideration of the specification and practice described herein. It is intended that the specification and described configurations be considered exemplary only, with the true scope and spirit being indicated by the following claims. In certain embodiments, for example, the following are provided: (Item 1) aligning the plurality of sequence reads to a reference sequence; determining one or more breakpoints in an alignment of the plurality of sequence reads to the reference sequence; identifying any sequence reads associated with the one or more breakpoints in the alignment as candidate fusion sequence reads; determining candidate fusion sequence reads associated with a common breakpoint among one or more breakpoints; Steps to be followed: grouping the candidate fusion sequence reads based on one or more common breakpoints; assembling the candidate fusion sequence reads in the group into one or more contigs; aligning the contigs from the group of a plurality of groups to the reference sequence; determining one or more candidate fusion events based on the alignment of the contigs from the group; applying one or more criteria to the one or more candidate fusion events; and determining one or more fusion events based on applying the one or more criteria to the one or more candidate fusion events. A method comprising: (Item 2) 2. The method of claim 1, wherein the step of identifying any sequence reads associated with the one or more breakpoints in the alignment as candidate fusion sequence reads comprises discarding alignments having a mappability score below a threshold. (Item 3) 3. The method of any one of items 1 to 2, wherein the step of identifying any sequence reads associated with the one or more breakpoints in the alignment as candidate fusion sequence reads comprises discarding alignments that are logical. (Item 4) 4. The method of any one of items 1 to 3, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads comprise breakpoints on the same chromosome and in the same orientation. (Item 5) 5. The method of any one of items 1 to 4, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads include a breakpoint located at the same position. (Item 6) 6. The method of any one of items 1 to 5, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads include a breakpoint that is within a threshold number of bases from a certain position. (Item 7) 7. The method of any one of items 1 to 6, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads comprise multiple breakpoints on the same chromosome and in the same orientation. (Item 8) 8. The method of any one of items 1 to 7, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads include multiple breakpoints located at the same position. (Item 9) 9. The method of any one of items 1 to 8, wherein the step of determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints comprises determining that at least two candidate fusion sequence reads each include a plurality of breakpoints that are within a threshold number of bases from a plurality of positions. (Item 10) A sequence for grouping the candidate fusion sequence reads based on one or more common breakpoints. 10. The method of any one of items 1 to 9, wherein step 10 comprises generating a de Bruijn graph for the group. (Item 11) Item 11. The method of item 10, wherein assembling the candidate fusion sequence reads in the group into one or more contigs comprises linearizing the de Bruijn graph to generate contigs for the group. (Item 12) 12. The method of any one of items 1 to 11, wherein assembling the candidate fusion sequence reads in the group into one or more contigs comprises performing one or more error correction procedures. (Item 13) Item 13. The method of item 12, wherein the one or more error correction procedures comprise resolving mismatches between candidate fusion sequence reads and the reference sequence. (Item 14) 14. The method of any one of items 12 to 13, wherein the one or more error correction procedures comprise inserting padding between at least two candidate fusion sequence reads. (Item 15) 15. The method of any one of items 12 to 14, wherein the one or more error correction procedures comprise discarding one or more candidate fusion sequence reads having unaligned portions exceeding a threshold. (Item 16) 16. The method of any one of items 1 to 15, wherein determining one or more candidate fusion events based on the alignment of the contigs from the group comprises applying one or more of a footprint test or a variability test. (Item 17) 17. The method of claim 16, wherein applying the footprinting test comprises determining that a threshold number of a family of candidate fusion sequence reads supporting the contig spans the breakpoint. (Item 18) 18. The method of any one of items 16 to 17, wherein applying the variability test comprises determining that a threshold amount of variability exists between at least two families of candidate fusion sequence reads that support the contig and span the breakpoint. (Item 19) applying one or more criteria to the one or more candidate fusion events determining the distance between the breakpoint of the one or more aligned contigs and the position of at least one probe of the panel for the candidate fusion event; and Discarding any candidate fusion events associated with an aligned contig of said one or more contigs that does not contain a breakpoint that is less than a threshold distance from the location of at least one probe of the panel. 19. The method according to any one of items 1 to 18, comprising: (Item 20) applying one or more criteria to the one or more candidate fusion events Determining one or more genes of interest; and 20. The method of any one of items 1 to 19, comprising discarding any candidate fusion events associated with aligned contigs of said one or more contigs that do not contain breakpoints associated with said one or more genes of interest. (Item 21) applying one or more criteria to the one or more candidate fusion events determining, for said candidate fusion event, that the breakpoint of said one or more aligned contigs is a deletion; and Discarding any candidate fusion events associated with an aligned contig of said one or more contigs that contains a deletion located within a number of bases away from another deletion. 21. The method of any one of items 1 to 20, comprising: (Item 22) applying one or more criteria to the one or more candidate fusion events determining, for said candidate fusion event, that the breakpoint of said one or more aligned contigs is a deletion; and Discarding any candidate fusion events associated with aligned contigs of said one or more contigs that contain a deletion comprising fewer than a threshold number of bases. 22. The method according to any one of items 1 to 21, comprising: (Item 23) applying one or more criteria to the one or more candidate fusion events Discarding any candidate fusion events associated with aligned contigs of said one or more contigs that contain insertions or deletions that are completely buried in intron regions. 23. The method of any one of items 1 to 22, comprising: (Item 24) applying one or more criteria to the one or more candidate fusion events determining a molecular to read ratio for the one or more aligned contigs for the candidate fusion event; and Discarding any candidate fusion events associated with aligned contigs of said one or more contigs that are associated with a ratio of molecules to reads above a threshold but are not associated with double-stranded support molecules. 24. The method according to any one of items 1 to 23, comprising: (Item 25) applying one or more criteria to the one or more candidate fusion events determining, for said candidate fusion event, for a breakpoint pair of said one or more aligned contigs, the sequence flanking said breakpoint of said breakpoint pair; aligning the sequences adjacent to the breakpoints of the breakpoint pair; determining an alignment score for the alignment of the sequences flanking the breakpoints of the breakpoint pair; and discarding any candidate fusion events associated with aligned contigs of said one or more contigs based on said alignment score exceeding a threshold. 25. The method according to any one of items 1 to 24, comprising: (Item 26) applying one or more criteria to the one or more candidate fusion events determining, for said breakpoint pair of said one or more aligned contigs for said candidate fusion event, a sequence centered at said breakpoint of said breakpoint pair; aligning the sequences centered around the breakpoints to each other; determining an alignment score for the alignment of the sequences centered around the breakpoint; and discarding any candidate fusion events associated with aligned contigs of said one or more contigs based on said alignment score exceeding a threshold. 26. The method of any one of items 1 to 25, comprising: (Item 27) aligning the plurality of sequence reads to a reference sequence; determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence; grouping the one or more candidate fusion sequence reads into one or more container data structures based on one or more common breakpoints; assembling the one or more candidate fusion sequence reads into one or more contigs for the container data structure; aligning the one or more contigs to the reference sequence for the container data structure; and determining, based on one or more criteria, one or more aligned contigs that represent a fusion event; A method comprising: (Item 28) 28. The method of claim 27, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads comprise breakpoints in the same chromosome and in the same orientation. (Item 29) 29. The method of any one of items 27 to 28, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads comprise breakpoints that are located at the same position. (Item 30) 30. The method of any one of items 27 to 29, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads include a breakpoint that is within a threshold number of bases from a certain position. (Item 31) 31. The method of any one of items 27 to 30, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads comprise multiple breakpoints in the same chromosome and in the same orientation. (Item 32) 32. The method of any one of items 27 to 31, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads include multiple breakpoints located at the same position. (Item 33) 33. The method of any one of items 27 to 32, wherein determining one or more candidate fusion sequence reads of the plurality of sequence reads based on one or more breakpoints in the alignment of sequence reads to the reference sequence comprises determining that at least two candidate fusion sequence reads include a plurality of breakpoints that are within a threshold number of bases from a plurality of positions. (Item 34) 34. The method of any one of items 27 to 33, wherein cut points from different alignments are assigned to a common container data structure. (Item 35) for the group, assembling the one or more candidate fused reads into one or more contigs, for said group, assembling said one or more candidate fusion sequence reads into a graph data structure; and linearizing the graph data structure to generate one or more contigs; 35. The method according to any one of items 27 to 34, comprising: (Item 36) Assembling the one or more candidate fusion sequence reads into one or more contigs. 36. The method of any one of items 27 to 35, wherein the step of performing comprises performing one or more error correction procedures. (Item 37) 37. The method of claim 36, wherein the one or more error correction procedures comprise resolving mismatches between candidate fusion sequence reads and the reference sequence. (Item 38) 38. The method of any one of items 36 to 37, wherein the one or more error correction procedures comprise inserting padding between at least two candidate fusion sequence reads. (Item 39) 39. The method of any one of items 36 to 38, wherein the one or more error correction procedures comprise discarding one or more candidate fusion sequence reads having unaligned portions exceeding a threshold. (Item 40) 40. The method of any one of items 27 to 39, further comprising determining one or more candidate fusion events based on the alignment of the contigs from the group, the determining step comprising applying one or more of a footprint test or a variability test. (Item 41) 41. The method of claim 40, wherein applying the footprint test comprises determining that a threshold number of a family of candidate fusion sequence reads supporting the contig spans the breakpoint. (Item 42) 42. The method of any one of items 40 to 41, wherein applying the variability test comprises determining that a threshold amount of variability exists between at least two families of candidate fusion sequence reads that support the contig and span the breakpoint. (Item 43) determining, based on the one or more criteria, the one or more aligned contigs indicative of one or more fusion events, determining the distance between the breakpoint of said one or more aligned contigs and the location of at least one probe of the panel; and 43. The method of any one of items 27 to 42, comprising discarding any aligned contig of the one or more contigs that does not contain a breakpoint that is less than a threshold distance from the location of at least one probe of the panel. (Item 44) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, Determining one or more genes of interest; and Discarding any aligned contigs of said one or more contigs that do not contain breakpoints associated with said one or more genes of interest. 44. The method according to any one of items 27 to 43, comprising: (Item 45) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, determining that the breakpoints in the one or more aligned contigs are deletions; and Discarding any aligned contigs of said one or more contigs that contain a deletion located within a number of bases away from another deletion. 45. The method according to any one of items 27 to 44, comprising: (Item 46) Based on the one or more criteria, the one or more alarms indicative of the fusion event are determining the inserted contigs, determining that the breakpoints in the one or more aligned contigs are deletions; and Discarding any aligned contigs of said one or more contigs that contain a deletion containing fewer than a threshold number of bases. 46. ​​The method according to any one of items 27 to 45, comprising: (Item 47) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, Discarding any aligned contigs of said one or more contigs that contain an insertion or deletion that is completely buried in an intron region. 47. The method of any one of items 27 to 46, comprising: (Item 48) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, determining a molecular to read ratio for the one or more aligned contigs; and 48. The method of any one of items 27 to 47, comprising discarding any aligned contigs of the one or more contigs that are associated with a ratio of molecules to reads that exceeds a threshold but are not associated with double-stranded support molecules. (Item 49) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, for breakpoint pairs of said one or more aligned contigs, determining the sequences flanking said breakpoints of said breakpoint pairs; aligning the sequences adjacent to the breakpoints of the breakpoint pair; determining an alignment score for the alignment of the sequences flanking the breakpoints of the breakpoint pair; and discarding any aligned contigs of said one or more contigs based on said alignment score exceeding a threshold. 49. The method according to any one of items 27 to 48, comprising: (Item 50) determining the one or more aligned contigs indicative of the fusion event based on the one or more criteria, determining, for the breakpoint pairs of the one or more aligned contigs, a sequence centered at the breakpoint of the breakpoint pair; aligning the sequences centered around the breakpoints to each other; determining an alignment score for the alignment of the sequences centered around the breakpoint; and discarding any aligned contigs of said one or more contigs based on said alignment score exceeding a threshold. 50. The method of any one of items 27 to 49, comprising: (Item 51) generating a notification indicating a problem associated with the library preparation based on discarding any aligned contigs of said one or more contigs. 51. The method of any one of items 27 to 50, further comprising: (Item 52) one or more processors; a memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to perform the method of any of items 1 to 51; 1. An apparatus comprising: (Item 53) 52. A non-transitory computer-readable medium storing processor-executable instructions that, when executed by at least one computing device, cause the at least one computing device to perform the method of any of items 1 to 51. (Item 54) 52. A system including at least one computing device configured to perform the method according to any of items 1 to 51. (Item 55) 52. A method of treating a subject, comprising administering a therapeutic agent to the subject, wherein the subject has been determined to have a fusion event using one or more of the methods described in items 1 to 51. (Item 56) 56. The method of claim 55, wherein the subject determined to have a fusion event has been diagnosed with cancer. (Item 57) 57. The method of item 56, wherein the cancer is a cancer associated with a fusion event. (Item 58) 58. The method of claim 57, wherein the cancer associated with a fusion event is selected from the group consisting of advanced urothelial cancer, prostate cancer, breast cancer, lung cancer, colon cancer, glioblastoma, liver cancer, and ovarian cancer. (Item 59) 59. The method of any one of items 55 to 58, wherein the therapeutic agent is a cancer therapeutic agent. (Item 60) 60. The method of claim 59, wherein the cancer therapeutic agent is specific for the cancer with which the subject has been diagnosed. (Item 61) 61. The method of any one of items 59 to 60, wherein the cancer therapeutic agent is specific for the fusion event.

Claims

[Claim 1] The invention described in the specification.