Real-time genomic characterization of cancer

A rapid assay using long-read whole-genome sequencing with adaptive sampling addresses the complexity and cost of current leukemia classification methods, enabling accurate detection of clinically relevant genomic subtypes for improved cancer diagnosis and treatment.

WO2026050126A1PCT designated stage Publication Date: 2026-03-05THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/043292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-26
Filing Date
2025-08-25
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current methods for classifying pediatric acute leukemias like B-ALL and AML are complex, costly, and lack comprehensiveness, requiring multiple laborious techniques such as karyotype analysis, FISH, PCR testing, and targeted sequencing panels, which are time-consuming and do not adequately address clinically relevant genomic subtypes for prognosis and treatment.

Method used

A rapid and sensitive assay method using long-read whole-genome sequencing with adaptive sampling to detect karyotype abnormalities and translocations/gene fusions within 24 hours, involving DNA extraction, library preparation, sequencing, and computational analysis to identify clinically relevant genomic subtypes.

Benefits of technology

Enables accurate and sensitive detection of clinically relevant genomic subtypes in cancer patients, enhancing diagnosis and treatment by classifying cancers based on translocations and gene fusions in a rapid and cost-effective manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025043292_05032026_PF_FP_ABST
    Figure US2025043292_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for classifying cancers to aid in diagnosis and treatment of the cancer, in particular acute leukemia and solid tumors such as pediatric sarcomas, in a rapid manner (less than 24 hours). Methods comprising the use of long-read whole genome sequencing of libraries comprising high molecular weight DNA to identify copy number variations and / or presence or absence of a mutation such as a karyotype abnormality or translocation / gene fusion in a panel of targets and treatment of patients based upon cancer classification.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.5470.980.WO REAL-TIME GENOMIC CHARACTERIZATION OF CANCER CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit, under 35 U.S.C. § 119 (e), of U.S. Provisional Application No.63 / 686,935 filed on August 26, 2024, the entire content of which is incorporated by reference herein. STATEMENT OF GOVERNMENT SUPPORT

[0002] This invention was made with government support under Grant No. CA259926 awarded by the National Institutes of Health. The government has certain rights in the invention. STATEMENT REGARDING ELECTRONIC FILING OF A SEQUENCE LISTING

[0003] A Sequence Listing in XML format, entitled 5470-980WO_ST26.xml, 2,723 bytes in size, generated on August 20, 2025, and filed herewith, is hereby incorporated by reference in its entirety for its disclosures. FIELD OF THE INVENTION

[0004] This invention relates to methods for classifying cancers to aid in diagnosis and treatment of the cancer. The methods comprise optimized sample handling, sequencing, and computational analysis steps to determine karyotype abnormalities and translocations / gene fusions in a rapid manner (less than 24 hours). BACKGROUND OF THE INVENTION

[0005] B-cell acute lymphoblastic leukemia (B-ALL) is the most common type of pediatric cancer. B-ALL shares genomic underpinnings across the age spectrum into late adulthood, albeit with increased relative frequency of high-risk genomic subtypes. Acute Myeloid Leukemia (AML) is a lineage of acute leukemia that accounts for 15-20% of pediatric acute leukemia cases and is defined by large-scale structural variation, which, similar to B-ALL, is detectable with whole- genome sequencing (de Rooij et al. (2015) J. Clin. Med.4(1):127). Currently, B-ALL and AML cases are classified into genomic subtypes through a cascade of internal and external testing that is complex, costly, and lacks comprehensiveness. These tests include karyotype analysis,Attorney Docket No.5470.980.WO fluorescence in situ hybridization (FISH), polymerase chain reaction (PCR) testing, and targeted sequencing panels and genomic tests (Arber et al. (2016) Blood 127(20):2391-2405; Iacobucci & Mullighan (2017) J. Clin. Oncol.35(9):975-83; Narayanan & Weinberg (2020) Intl. J. Laboratory Hematol. 42(1):3-15). These multiple distinct, costly, and laborious techniques are critical to classify pediatric acute leukemia into clinically relevant genomic subtypes, with implications for prognosis, selection of treatment intensity, and precision medicine approaches.

[0006] Needed in the art is a rapid, low-cost assay to classify cancers based on clinically relevant genomic variation and tumorigenic drivers. The present invention addresses this need in the art. SUMMARY OF THE INVENTION

[0007] The present invention is based on the development of a rapid and sensitive assay method for determining karyotype abnormalities and translocations / gene fusions of cancer cells using long-read whole-genome sequencing with adaptive sampling that can be carried out in less than 24 hours. The method enables accurate and sensitive detection of clinically-relevant genomic subtypes, enhancing diagnosis and treatment of cancer patients.

[0008] Thus, this invention provides a method of classifying cancer in a subject in need thereof, comprising extracting high molecular weight DNA from a biological sample from the subject; preparing a library from the extracted high molecular weight DNA for long-read (e.g., nanopore) sequencing; performing long-read whole genome sequencing of the library to generate sequencing data; identifying the presence or absence of translocations or other mutations or structural or copy number variations in a panel of targets in the sequencing data; and classifying the cancer based on the presence or absence of the panel of targets. In some embodiments, the method of the invention further comprises treatment steps selected based on the determined classification. BRIEF DESCRIPTION OF THE FIGURES

[0009] FIGS.1A-D. Large-scale copy-number characterization (karyotyping) based on relative chromosome-wide sequencing depth (FIG. 1A). Small-scale / intragenic copy number variation detection based on sequencing depth and split-read mapping (FIG. 1B). Single nucleotide variation (SNV) calling within enriched gene targets (FIG.1C). Fusion gene identification based on split alignment of long reads (FIG.1D).Attorney Docket No.5470.980.WO

[0010] FIGS.2A-2C. Points show the number of reads in non-overlapping 1Mbp bins and lines indicate the median of the entire chromosome. FIG.2A. Patient sample 0052 represents a case of B-ALL with a high hyperdiploid genomic subtype. Our inferred digital karyotype is 55, XY, +4, +5, +6, +10, +14, +17, +18, +21, +21. FIG. 2B. Sample 0225 represents B-ALL with a near haploid genomic subtype. Our inferred digital karyotype is 27, XY, +10, +18, +21. FIG. 2C. Sample 0223 is B-ALL with low hypodiploidy, inferred to be 38, XX, -3, -4, -7, -9, +11q, -15, - 16, -17, -18. In all cases, we only annotate whole-chromosome and arm-level changes for the purposes of gross aneuploidy detection, but smaller sub-chromosomal copy-number changes are also conspicuous.

[0011] FIGS.3A-3C. Long reads aligning to ETV6 (FIG.3A) and RUNX1 (FIG.3B); Dark bars represent matched fragments overlapping the putative breakpoint; light bars are singly-mapping. Relative sequencing depth is also shown above, indicating discontinuous coverage representing small indels at the putative translocation breakpoint resulting from imperfect double-strand break repair. FIG. 3C. A sample of reads supporting the putative ETV6::RUNX1 breakpoint, showing segments mapping to ETV6 (filled bars) and RUNX1 (unfilled) indicating an inverted translocation consistent with ETV6 intron 5-6 fused to RUNX1 intron 1-2, matching their respective coding orientation.

[0012] FIGS. 4A-4C. DUX4-IGH rearrangement where dark bars indicate reads aligning to DUX4 (FIG.4A) and IGH (FIG.4B) regions with a focal deletion in ERG (FIG.4C) characterized by a drop in sequencing depth, and reads split across the deletion boundaries (dark bars).

[0013] FIG. 5. Recapitulating commonly used FLT3-ITD PCR and capillary electrophoretic analysis using in silico primer capture—sample 0141 with clinically reported AR 0.65. We identified 46 reads spanning FLT3 primers, 28 consisting of 300-309nt and 18 containing 377- 392nt, inferring an ITD:WT ratio of 0.64.

[0014] FIG. 6. DNA extraction and library preparation took 3 hours, and sample classification took between 3 hours and 6 hours. Complete sample classification (from sample receipt to classification) took between 6 hours and 9 hours.

[0015] FIGS.7A-7B. FIG.7A. When using adaptive sampling and running samples individually (hatched bars), 19 / 25 (76%) of translocations are identified in under 3 hours, and 23 / 25 (92%) are identified in under 6 hours. Samples run multiplexed (four on one chip; unfilled bars) or without adaptive sampling (dotted bars) show the expected longer sequencing times to result. FIG. 7B.Attorney Docket No.5470.980.WO Sequencing time required to classify a chromosome abnormality into a specific aneuploid category (such as high hyperdiploid, low hypodiploid, or near haploid). Hatched bars represent samples that were sequenced individually on a P2, unfilled bars indicate samples that were multiplexed, and one sample (dotted bars) was sequenced on a low-throughput Flongle flow cell. Bars with horizontal lines indicate samples sequenced individually where karyotype abnormalities represented subclones secondary to a translocation-driven primary clone (BCR::ABL1).

[0016] FIGS.8A-8D. Sequencing depth profiles illustrating karyotype variation at early and late timepoints. Sample 0227 depth profile is shown when it is called low hypodiploid after 2 minutes (FIG.8A) and at the end of sequencing (FIG.8B). Sample 0173 is shown when it is called high hyperdiploid after 5 minutes (FIG.8C) and at the end of sequencing (FIG.8D). DETAILED DESCRIPTION

[0017] The present invention will now be described in more detail with reference to the accompanying drawings, in which preferred embodiments of the invention are shown. This invention may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In addition, any references cited herein are incorporated by reference in their entireties.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of skill in the art to which this invention belongs. The terminology used in the description of the invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. All publications, patent applications, patents, patent publications and other references cited herein are incorporated by reference in their entireties for the teachings relevant to the sentence and / or paragraph in which the reference is presented.

[0019] Except as otherwise indicated, standard methods known to those skilled in the art may be used for cloning genes, amplifying and detecting nucleic acids, and the like. Such techniques are known to those skilled in the art. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual 4th Ed. (Cold Spring Harbor, NY, 2012); Ausubel et al. Current Protocols in Molecular Biology (Green Publishing Associates, Inc. and John Wiley & Sons, Inc., New York).Attorney Docket No.5470.980.WO

[0020] Unless the context indicates otherwise, it is specifically intended that the various features of the invention described herein can be used in any combination.

[0021] Moreover, the present invention also contemplates that in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted.

[0022] To illustrate, if the specification states that a complex comprises components A, B and C, it is specifically intended that any of A, B or C, or a combination thereof, can be omitted and disclaimed singularly or in any combination.

[0023] As used in the description of the invention and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells.

[0024] Also as used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0025] The term “about,” as used herein when referring to a measurable value such as an amount of polypeptide, dose, time, temperature, enzymatic activity or other biological activity and the like, is meant to encompass variations of 10%, 5%, 1%, 0.5%, 0.1%, or even 0.01% of the specified amount.

[0026] As used herein, the transitional phrase “consisting essentially of” (and grammatical variants) is to be interpreted as encompassing the recited materials or steps and those that do not materially affect the basic and novel characteristic(s) of the claimed invention. Thus, the term “consisting essentially of” as used herein should not be interpreted as equivalent to “comprising.”

[0027] The term “consists essentially of” (and grammatical variants), as applied to polynucleotide sequence of this invention, means a polynucleotide that consists of both the recited sequence (e.g., SEQ ID NO) and a total of ten or less (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) additional nucleotides on the 5’ and / or 3’ ends of the recited sequence such that the function of the polynucleotide is not materially altered. The total of ten or less additional nucleotides includes the total number of additional nucleotides on both ends added together.

[0028] A “subject” or “patient” may be any vertebrate organism in various embodiments. A subject may be an individual from whom a biological sample is obtained, or on whom a procedure is performed, or to whom an agent is administered, e.g., for experimental, diagnostic, and / or therapeutic purposes. In some embodiments a subject is a mammal, e.g., a human, non-humanAttorney Docket No.5470.980.WO primate, lagomorph (e.g., rabbit), or rodent (e.g., mouse, rat). In some embodiments a human subject is a neonate, child, adult or geriatric subject. In some embodiments, a patient or subject is a subject having (e.g., a subject diagnosed with cancer), suspected of having (e.g., a subject exhibiting signs or symptoms of cancer, but not diagnosed by a clinician), or at risk of having cancer (e.g., a subject predisposed to having cancer based upon genetic (inherited), lifestyle, and / or environmental factors). In some embodiments, a patient or subject is a subject having, suspected of having, or at risk of having an acute leukemia. In some embodiments, a patient or subject is a subject having, suspected of having, or at risk of having a pediatric acute leukemia, such as an acute lymphoblastic leukemia (ALL) such as a B-cell lymphoblastic leukemia / lymphoma (B-ALL) or T-cell lymphoblastic leukemia (T-ALL); or an acute myeloid leukemia (AML). In some embodiments, a patient or subject is a subject having, suspected of having, or at risk of having a solid tumor. In some embodiments, a patient or subject is a subject having, suspected of having, or at risk of having a solid tumor that is a pediatric sarcoma such as an osteosarcoma, Ewing sarcoma, rhabdomyosarcoma, synovial sarcoma, desmoplastic small round cell tumor, alveolar soft part sarcoma, desmoid tumor, dermatofibrosarcoma protuberan, infantile fibrosarcoma, leiomyosarcoma, or liposarcoma.

[0029] As used herein, the term “biological sample” refers to a fluid or solid containing cells and compounds of biological origin, and may include blood, bone marrow, stool or feces, lymph, urine, serum, plasma, pus, saliva, seminal fluid, tears, urine, bladder washings, colon washings, sputum or fluids from the respiratory, alimentary, circulatory, or other body systems. In certain aspects of the current method, any medical professional such as a doctor, nurse or medical technician may obtain a biological sample for testing. Alternatively, a patient or subject may obtain a biological sample for testing without the assistance of a medical professional, such as obtaining a whole blood sample, a urine sample, a fecal sample, a buccal sample, or a saliva sample. A sample may also comprise in vitro cell culture constituents (including but not limited to conditioned medium resulting from the growth of cells in cell culture medium, recombinant cells and cell components). In some embodiments, a “biological sample” contains or is presumed to contain genomic DNA. In some embodiments, a “biological sample” is a peripheral blood sample, bone marrow sample, resected tissue sample (e.g., a resected tumor tissue sample), tissue biopsy sample (e.g., a tumor tissue biopsy sample) or cell fraction thereof, e.g., mononuclear cell fraction thereof. A “bone marrow mononuclear cell fraction” or “peripheral blood mononuclear cell fraction” refers to a cellAttorney Docket No.5470.980.WO fraction that is enriched for mononuclear cells such as B cells, T cells, NK lymphocytes, early myeloid cells, and a small fraction (e.g., less than 10%) of endothelial progenitors, hematopoietic stem / progenitor and / or mesenchymal stromal cells.

[0030] As used herein, “bone marrow sample” may either be a bone marrow aspirate or from bone pieces, such as cancellous bone pieces. In one embodiment, a bone marrow sample is a bone marrow aspirate. A bone marrow sample may be collected by any conventional method and may comprise a mixture of different cell types. In some embodiments, a bone marrow sample is a bone marrow mononuclear cell sample or fraction.

[0031] In some aspects, a biological sample may be a fresh biological sample or a cryopreserved biological sample. A “fresh biological sample” refers to a biological sample that is collected from a subject and used in the method described herein without being cryopreserved (frozen). A “cryopreserved biological sample” refers to a biological sample that is collected from a subject and frozen (e.g., at -20°C, -70°C, or -80°C) prior to being used in the method described herein.

[0032] The term “genomic DNA” or “DNA” herein refers to DNA of a cellular genome. The genomic DNA can be cellular, i.e., contained within a cell, or it may be cell-free. In particular embodiments, genomic DNA used in accordance with the method therein is cell-free DNA (cfDNA).

[0033] “High molecular weight DNA,” “HMW DNA,” or “High molecular weight genomic DNA” refers to DNA molecules in excess of at least about 15 kb, e.g., DNA molecules of at least about 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 100 kb, 150 kb, 200 kb, 250 kb, 300 kb, 350 kb, 400 kb, 450 kb, 500 kb or more.

[0034] The term “library” refers to a collection or plurality of template molecules, i.e., target DNA duplexes, which share common sequences (e.g., adaptors) at their 5′ ends and / or common sequences (e.g., adaptors) at their 3′ ends. Use of the term “library” to refer to a collection or plurality of template molecules should not be taken to imply that the templates making up the library are derived from a particular source, or that the “library” has a particular composition. By way of example, use of the term “library” should not be taken to imply that the individual templates within the library must be of different nucleotide sequence or that the templates must be related in terms of sequence and / or source.

[0035] The term “read” as used herein refers to the raw or processed output of sequencing systems, such as massively parallel sequencing. In some embodiments, the output of the methodsAttorney Docket No.5470.980.WO described herein is reads. In some embodiments, these reads may need to be trimmed, filtered, and aligned, resulting in raw reads, trimmed reads, aligned reads.

[0036] The term “sequencing,” as used herein, refers to a method by which the identity of at least about 10 consecutive nucleotides (e.g., the identity of at least about 20, at least about 50, at least about 100 or at least about 200 or more consecutive nucleotides) of a polynucleotide is obtained.

[0037] The term “long-read sequence” refers to reads of at least about 0.5, 1, 2, 3, 5, 10, 15, or 20 kb in length, and in some embodiments averaging about 0.5, 1, 2, 3, 5, 10, 15, or 20 kb in length. Long-read sequencing methods may include nanopore sequencing methods such as that commercialized by Oxford Nanopore Technologies, electronic detection methods such as Ion Torrent technology commercialized by Life Technologies, and single-molecule fluorescence- based methods such as that commercialized by Pacific Biosciences.

[0038] “Whole genome sequencing,” “WGS,” “full genome sequencing,” “complete genome sequencing,” or “entire genome sequencing” is the process of determining the complete DNA sequence of an organism's genome at a single time. This entails sequencing all of an organism's chromosomal DNA as well as DNA contained in the mitochondria. “Long-read whole genome sequencing,” or grammatical variations thereof, refers to the process of determining the complete DNA sequence of an organism's genome using methods that obtain reads of at least about 0.2, 0.5, 1, 2, 3, 5, 10, 15, or 20 kb in length, e.g,. averaging about 0.5, 1, 2, 3, 5, 10, 15, or 20 kb in length.

[0039] A “mutation” herein refers to a change introduced into a reference sequence, including, but not limited to, substitutions, indels (insertions or deletions including truncations), translocations, fusions, and / or single nucleotide variations relative to the reference sequence. Mutations may involve large sections of DNA (e.g., copy number variation). Mutations may involve whole chromosomes (e.g., aneuploidy). Mutations may involve small sections of DNA. Examples of mutations involving small sections of DNA include, e.g., point mutations or single nucleotide polymorphisms (SNPs), multiple nucleotide polymorphisms, insertions (e.g., insertion of one or more nucleotides at a locus but less than the entire locus), multiple nucleotide changes, deletions (e.g., deletion of one or more nucleotides at a locus), and inversions (e.g., reversal of a sequence of one or more nucleotides). The consequences of a mutation include, but are not limited to, the creation of a new characteristic, property, function, phenotype or trait not found in the product (e.g., mRNA or protein) encoded by the reference sequence. In some embodiments, the reference sequence may be a parental sequence. In some embodiments, the reference sequenceAttorney Docket No.5470.980.WO may be a reference human genome. In some embodiments, the reference sequence may be derived from a non-cancer (or non-tumor) sequence. In some embodiments, the mutation may be inherited. In some embodiments, the mutation may be spontaneous or de nova.

[0040] A “translocation” refers to a change in position of a chromosomal segment from a first region to a second region on either the same chromosome or to a second chromosome. In some embodiments, the term “chromosomal segment” may refer to at least 2, 5, 50, 100, 250, 500, 1000, 2500, 5000, 10,000, 25,000, 50,000, 100,000, 250,000, 500,000, 1,000,000, 2,500,000, 5,000,000, 10,000,000, 25,000,000, or at least 50,000,000 contiguous nucleotides of a chromosome.

[0041] A “fusion” or “fusion gene” refers to a mutation in which two physically separated genes in an organism's genome are spliced into a new gene. For example, the BCR-ABL1 fusion gene is a well-known example of a fusion gene that has been found in some leukemia cases.

[0042] A “single nucleotide variation” or “SNV” refers to a difference in a single nucleotide (A, T, C, or G) at a specific position in the genome, compared to a reference sequence. An SNV may be either common, like single nucleotide polymorphisms (SNPs), or rare. An SNV may occur in a coding and non-coding region of the genome and may have various effects on gene expression and protein function.

[0043] As used herein, “indel” refers to an insertion or deletion of one or more nucleotide bases in a nucleic acid. A “small insertion” or “small deletion” refers to the insertion or deletion of between one and 20 extra nucleotides in a nucleic acid. These insertions or deletions may result in frameshift mutations within the coding region of the gene.

[0044] A “copy number variation” or “CNV” refers to any duplication or deletion of a genomic segment. A “copy number loss variant” or “CNLV” refers to a deletion of a genomic segment of more than about 100 base pairs.

[0045] A “target,” “target sequence,” or “genomic target sequence” refers to a selected target polynucleotide, e.g., a sequence present in a cfDNA molecule, whose presence, amount, and / or nucleotide sequence, or changes in these, are desired to be determined. Target sequences are interrogated for the presence or absence of a mutation (e.g., a somatic mutation) and or copy number variation. The target sequence may be a region of gene (e.g., target gene) associated with a disease. In some embodiments, the region is an exon. In some embodiments, the disease is a cancer such as leukemia (e.g., an acute leukemia) or a solid tumor (e.g., a pediatric sarcoma).Attorney Docket No.5470.980.WO

[0046] A “panel of targets” or “panel of biomarkers,” or grammatical variants thereof, refers to a specified set of targets (e.g., target genes) known to have mutations therein or copy number variations that are responsible for or correlate with cancer. Exemplary panels of targets are set forth in Tables 3, 4, 5, 11, and 12.

[0047] A “gene,” for the purposes of the present disclosure, includes a DNA region encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding sequences, transcribed sequences, and combinations of coding and transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites and locus control regions.

[0048] The term “modulate,” “modulates,” or “modulation” refers to enhancement (e.g., an increase) or inhibition (e.g., a decrease) in the specified level or activity.

[0049] The term “enhance” or “increase” refers to an increase in the specified parameter of at least about 1.25-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 8-fold, 10-fold, twelve-fold, or even fifteen-fold and / or can be expressed in the enhancement and / or increase of a specified level and / or activity of at least about 1%, 5%, 10%, 15%, 25%, 35%, 40%, 50%, 60%, 75%, 80%, 90%, 95% or more.

[0050] The term “inhibit” or “reduce” or grammatical variations thereof as used herein refers to a decrease or diminishment in the specified level or activity of at least about 1, 5, 10, 15%, 25%, 35%, 40%, 50%, 60%, 75%, 80%, 90%, 95% or more. In particular embodiments, the inhibition or reduction results in little or essentially no detectible activity (at most, an insignificant amount, e.g., less than about 10% or even 5%).

[0051] “Pharmaceutical composition” means a mixture of substances suitable for administering to a subject. For example, a pharmaceutical composition may comprise a cancer therapeutic of the invention in a pharmaceutically acceptable carrier.

[0052] “Pharmaceutically acceptable” or “physiologically tolerable” and grammatical variations thereof, as they refer to compositions, carriers, diluents and reagents, are used interchangeably and represent that the materials are capable of administration to or upon a subject (e.g., a mammal such as a mouse, rat, rabbit, or a primate such as a human), without the production of therapeutically prohibitive undesirable physiological effects.Attorney Docket No.5470.980.WO

[0053] Grammatical variations of “administer,” “administration,” and “administering” to a subject include any route of introducing or delivering to a subject an agent. Administration can be carried out by any suitable route, including oral, topical, intravenous, subcutaneous, transcutaneous, transdermal, intramuscular, intra-joint, parenteral, intra-arteriole, intradermal, intraventricular, intracranial, intraperitoneal, intralesional, intranasal, rectal, vaginal, by inhalation, via an implanted reservoir, parenteral (e.g., subcutaneous, intravenous, intramuscular, intra-articular, intra-synovial, intrasternal, intrathecal, intraperitoneal, intrahepatic, intralesional, and intracranial injections or infusion techniques), and the like. “Concurrent administration,” “administration in combination,” “simultaneous administration,” or “administered simultaneously” as used herein, means that the compounds are administered at the same point in time, overlapping in time, or one following the other. In the latter case, the two compounds are administered at times sufficiently close that the results observed are indistinguishable from those achieved when the compounds are administered at the same point in time. “Systemic administration” refers to the introducing or delivering to a subject an agent via a route which introduces or delivers the agent to extensive areas of the subject’s body (e.g., greater than 50% of the body), for example through entrance into the circulatory or lymph systems. By contrast, “local administration” refers to the introducing or delivery to a subject an agent via a route which introduces or delivers the agent to the area or area immediately adjacent to the point of administration and does not introduce the agent systemically in a therapeutically significant amount. For example, locally administered agents are easily detectable in the local vicinity of the point of administration but are undetectable or detectable at negligible amounts in distal parts of the subject's body. Administration includes self-administration and the administration by another.

[0054] “Treat,” “treating” and similar terms as used herein in the context of treating a subject refer to providing medical and / or surgical management of a subject. Treatment may include, but is not limited to, administering an agent or composition (e.g., a pharmaceutical composition) to a subject. Treatment is typically undertaken in an effort to alter the course of a disease (which term is used to indicate any disease, disorder, syndrome or undesirable condition warranting or potentially warranting therapy) in a manner beneficial to the subject. The effect of treatment may include reversing, alleviating, reducing severity of, delaying the onset of, curing, inhibiting the progression of, and / or reducing the likelihood of occurrence or recurrence of the disease or one or more symptoms or manifestations of the disease. A therapeutic agent may be administered to aAttorney Docket No.5470.980.WO subject who has a disease or is at increased risk of developing a disease relative to a member of the general population. In some embodiments a therapeutic agent may be administered to a subject who has had a disease but no longer shows evidence of the disease. The agent may be administered e.g., to reduce the likelihood of recurrence of evident disease. A therapeutic agent may be administered prophylactically, i.e., before development of any symptom or manifestation of a disease. “Prophylactic treatment” refers to providing medical and / or surgical management to a subject who has not developed a disease or does not show evidence of a disease in order, e.g., to reduce the likelihood that the disease will occur, delay the onset of the disease, or to reduce the severity of the disease should it occur. The subject may have been identified as being at risk of developing the disease (e.g., at increased risk relative to the general population or as having a risk factor that increases the likelihood of developing the disease.

[0055] As described herein, the present invention provides a method of classifying cancer in a subject in need thereof, comprising: extracting high molecular weight DNA from a biological sample from the subject; preparing a library from the extracted high molecular weight DNA for long-read sequencing; performing long-read whole genome sequencing of the library to generate sequencing data; identifying the presence or absence of translocations or other mutations or structural or copy number variations in a panel of targets in the sequencing data; and classifying the cancer based on the presence or absence of the panel of targets.

[0056] In some embodiments, the biological sample is peripheral blood, bone marrow, or a mononuclear cell fraction thereof. In some embodiments, the biological sample is peripheral blood, bone marrow, or a mononuclear cell fraction thereof, isolated or obtained from a subject with an acute leukemia such as a pediatric acute leukemia. In some embodiments, the biological sample is a fresh or cryopreserved peripheral blood, bone marrow, or a mononuclear cell fraction sample isolated or obtained from a subject with an acute leukemia such as a pediatric acute leukemia.

[0057] In some embodiments, the biological sample is a resected tissue (tumor tissue) or tissue biopsy sample. In some embodiments, the biological sample is a resected tissue or tissue biopsy sample from a solid tumor such as a pediatric sarcoma. In some embodiments, the biological sample is a fresh or cryopreserved resected tissue or tissue biopsy sample isolated or obtained from a subject with a solid tumor such as a pediatric sarcoma.

[0058] High molecular weight DNA may be obtained or extracted from a biological sample as described herein or using any suitable method that provides for the isolation of DNA molecules inAttorney Docket No.5470.980.WO excess of at least about 15 kb, in particular DNA molecules of at least about 20 kb, e.g., at least about 25 kb, 35 kb, 40 kb, 45 kb, 50 kb, 100 kb, 150 kb, 200 kb, 250 kb, 300 kb, 350 kb, 400 kb, 450 kb, 500 kb or more. By way of illustration, high molecule weight DNA may be extracted using commercially available kits such as, e.g., a ZymoBIOMICS™ MagBead DNA / RNA kit (Zymo Research), WIZARD® HMW DNA extraction kit (PROMEGA®), MAGATTRACT® HMW DNA Kit (QIAGEN®). The quality of the extracted DNA may or may not be assessed prior to the step of preparing a library from the DNA. In some embodiments, the quality of the extracted DNA is determined by gel electrophoresis.

[0059] In some embodiments, the high molecular weight DNA extracted from the biological sample (hereinafter referred to as “extracted DNA”) may be subjected to shearing or fragmentation prior to preparing a library from the extracted DNA. Shearing or fragmentation may be achieved by, e.g., hydrodynamic shear or other mechanical force, or fragmented by chemical or enzymatic digestion, such as restriction digestion. The sheared or fragmented DNA may further be separated by size to remove small background molecules and obtain DNA molecules of an optimal size for sequence analysis. Size selection may be accomplished, e.g., using Beckman Coulter AMPURE® beads or the method described, e.g., in WO 2019 / 006321 A1, incorporated herein by reference in its entirety.

[0060] In accordance with the method herein, a library for long-read sequencing is subsequently prepared from the extracted DNA and optionally sheared and size selected DNA. In some embodiments, the library is suitable for sequencing using a nanopore-based method. In some embodiments, a library may be produced by a combination of steps including repairing the ends of the extracted DNA, modifying the ends of the repaired DNA, and optionally amplifying the modified DNA. After fragmentation, ends of the extracted DNA may be damaged or blunt-ended and end repair techniques are used to create a more suitable substrate for adapter ligation. In one embodiment, ends of the extracted DNA are repaired by dA-tailing, which adds a single adenosine base to the 3' end of the extracted DNA fragments. Subsequently, the DNA molecules are subjected to additional modification, resulting in the attachment of oligonucleotides to the DNA molecules. The oligonucleotides may comprise an adapter sequence or a molecular barcode (or both). In some embodiments, the adapter sequence is common to all oligonucleotides in a plurality of oligonucleotides that are used to form the library. By way of example, an oligonucleotide may be attached to the DNA molecules by ligation. Sequencing adapters facilitate binding of the DNA toAttorney Docket No.5470.980.WO the sequencing device and initiate the sequencing process. In an optional embodiment, the resulting library of adapter-modified DNA may be amplified, e.g., by PCR, to increase the library concentration.

[0061] The library of modified and optionally amplified DNA molecules is subsequently subjected to long-read whole genome sequencing to generate sequencing data. Long-read whole genome sequencing methods share the characteristic of generating molecules larger than 200 nucleotides in a sequencing reaction. Typically, molecules are greater than 500 nucleotides and even several thousand nucleotides in length. Long-read sequencing approaches offer several advantages over other types of sequencing, including decreased turn-around time, decreased computational complexity, decreased cost, and improved structural variation detection (Jeck et al. (2019) J. Mol. Diagnost. 21(1):58-69; Liu et al. (2020) BMC Genomics 21(11):793; Oikonomopoulos et al. (2016) Sci. Rep.6:31602; Jain et al. (2018) Nat. Biotechnol.36(4):338-45). Existing platforms for long-read, real-time, single-molecule sequencing include, e.g., the Pacific Biosciences systems (RSII and Sequel) and the Oxford Nanopore systems (BASE™, GridION™, MinION™, and PromethION™), which are further described in U.S. Pat. Nos. 8,324,914 and 5,795,782; and U.S. Patent Publication Nos. 2010 / 0025249, 2010 / 0148126, and 2010 / 0174625, incorporated herein by reference in their entireties. These platforms allow read lengths over 55 kb and even longer.

[0062] In some embodiments, the long-read whole genome sequencing is performed using adaptive sampling. Adaptive sampling is a nanopore-specific in silico enrichment technique that optimizes sequencing by selectively enriching for fusion oncogene targets as described herein, while maintaining broad genomic coverage necessary for detecting large-scale chromosomal alterations. Adaptive sampling is performed by modulating the current of nanopores during sequencing to keep DNA fragments of interest and physically ejecting fragments not matching a predefined list of targets. The signal generated by each pore is analyzed continuously to identify the source (location) of the DNA in the genome. After one second of sequencing time, the first “chunk” of signal data (one second of 5 Khz signal data corresponding to about 400 nucleotides) is sent to the attached computer. The signal chunk undergoes base calling and alignment to the reference genome and a decision is made to either keep (if the sequence is near a gene of interest) or eject each read. If an “eject” decision is made, the sequencer is signaled and the voltage bias on the corresponding nanopore is reversed, effectively removing the DNA from the pore and allowingAttorney Docket No.5470.980.WO another to enter. In practice, analysis of a chunk and signaling ejection is accomplished in under one second following the second of data collection, resulting in an average ejected read length of about 600-700 nucleotides (1.5-2.0 seconds total). This selective enrichment, resulting in a 10-fold to 20-fold increase in sequencing depth over enrichment targets, does not rely on targeted sample preparation (e.g., biotin probes or selective PCR amplification), and also maintains the sequencing breadth necessary to determine copy-number variation status, and the sensitivity crucial to identifying gene fusions (Weilguny et al. (2023) Nat. Biotechnol. 41(7):1018-25; Martin et al. (2022) Genome Biol.23(1):11). In some embodiments, adaptive sampling in accordance with the method herein enriches at least 50 of the target genes listed in any one of Tables 3, 4, 5, or 11. In some embodiments, adaptive sampling in accordance with the method herein enriches at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 of the target genes listed in any one of Tables 4, 5, or 11. In some embodiments, adaptive sampling in accordance with the method herein enriches at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, or 220 of the target genes listed in Table 4 or Table 11. In some embodiments, adaptive sampling in accordance with the method herein enriches at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250 or 260 of the target genes listed in Table 11. In some embodiments, adaptive sampling in accordance with the method herein enriches at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 of the target genes listed in Table 12.

[0063] In some embodiments, the long-read whole genome sequencing system is a single- molecule real-time sequencing system. Single-molecule systems do not rely on clonal populations of amplified DNA fragments to generate a detectable signal. These systems fix sequence- determining proteins at specific locations and allow nucleic acid strands to advance through the protein. The existing Pacific Bioscience system uses polymerases, while the Oxford Nanopore system currently uses membrane channel proteins. In a preferred embodiment, the sequencing method is a single molecule real-time (SMRT) sequencing method, which provides sequencing data that is analyzed continuously in real-time.

[0064] The sequence data generated from the long-read whole genome sequence is subsequently analyzed to identify the presence or absence of mutations or copy number variations in a panel of targets (e.g., target genes) and based upon the identified presence or absence of mutations or copy number variations in the panel of targets, the cancer is classified. In some embodiments, the panel of targets comprises genes with a known mutation associated with cancer, optionally wherein theAttorney Docket No.5470.980.WO known mutation is a translocation, fusion, single nucleotide variation, and / or indel. In some embodiments, a target gene with a known mutation associated with cancer may be selected from the genes listed in any one of Tables 3, 4, 5, 11, or 12. In some embodiments, the panel of targets may differ for each type of cancer.

[0065] In some embodiments, a target gene is one having a known mutation associated with leukemia, in particular an acute leukemia such as a pediatric acute leukemia. In some embodiments, a target gene with a known mutation associated with leukemia, in particular an acute leukemia such as a pediatric acute leukemia, may be selected from the genes listed in any one of Tables 3, 4, 5, or 11. In some embodiments, the cancer is acute leukemia and the panel of targets includes at least 50 of the target genes listed in any one of Tables 3, 4, 5, or 11. In some embodiments, the cancer is acute leukemia and the panel of targets includes at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 of the target genes listed in any one of Tables 4, 5, or 11. In some embodiments, the cancer is acute leukemia and the panel of targets includes at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, or 220 of the target genes listed in Table 4 or Table 11. In some embodiments, the cancer is acute leukemia and the panel of targets includes at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250 or 260 of the target genes listed in Table 11.

[0066] In some embodiments, a target gene is one having a known mutation associated with a solid tumor, in particular a pediatric tumor such as a pediatric sarcoma. In some embodiments, a target gene with a known mutation associated with a solid tumor, in particular a pediatric sarcoma, may be selected from the genes listed in Table 12. In some embodiments, the cancer is a solid tumor and the panel of targets includes at least 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 of the target genes listed in Table 12.

[0067] In some embodiments, copy number variations may be determined by assessing relative copy number across the genome at the chromosome level (“digital karyotype”) and inferred based on relative sequencing depth by assessing genome-wide and chromosome-level depth of coverage as described herein. In some embodiments, translocations may be characterized by counting reads for which multiple alignments exist to two independent genes in the enrichment set. Exemplary genes known in the art to be associated with translocations in include, but are not limited to, ETV6, RUNX1, BCR, ABL1, CRLF2, P2RY8, MEF2D, CSF1R, EBF1, PDGFRB, PAX5,ZCCHC7, TCF3, PBX1, DUX4, NUP98, NSD1, KMT2A, MLLT1, EP300, ZNF384, USP2,Attorney Docket No.5470.980.WO RUNXIT1, MLLT10, HNRNRPUL1, BCL9, and CREBBP. See, also Table 3. In some aspects, smaller-scale (sub-chromosomal) copy number variation (e.g., iAMP21, CDKN2A, and ERG deletion, CRLF2-P2RY8 interstitial deletion) may be determined by a combination of sequencing depth and split read alignment evidence. Focal deletions may be characterized by a haploid (or multiple thereof) drop in read depth and reads split-mapped to either side of the deletion breakpoints.

[0068] In certain embodiments, the sequence information analyzed comprises replicate sequence information. For example, replicate sequence reads may be generated by repeatedly sequencing the same molecules, sequencing templates comprising multiple copies of a target sequence, sequencing multiple individual biomolecules all of which contain the sequence of interest or “target” sequence, or a combination of such approaches. Replicate sequence reads need not begin and end at the same position in a biomolecule sequence, as long as they contain at least a portion of the target sequence. Examples of methods of generating replicate sequence information from a single molecule are provided, e.g., in U.S. Pat. No.7,476,503; U.S. Patent Application Publication No. 2009 / 0298075; U.S. Patent Application Publication No. 2010 / 0075309; U.S. Patent Application Publication No. 2010 / 0075327; and U.S. Patent Application Publication No. 2010 / 0081143.

[0069] Once the presence or absence of a mutation(s) and / or copy number variation in a panel of targets is identified in the sequence data, the presence or absence of a mutation(s) and / or copy number variation in a panel of targets is used to classify the cancer. Such classifications include, but are not limited to, classifying an acute leukemia as an ALL or AML; classifying an ALL as a T-ALL or B-ALL; classifying a B-ALL as a high hyperdiploidy (>50 chromosomes) subtype, low hypodiploidy (31-39 chromosomes), near-haploidy (24-30 chromosomes), iAMP21 subtype, ETV6::RUNX1 subtype, TCF3::PBX1 subtype, Ph subtype, Ph-like subtype, MEF2Dr subtype, ZNF384r subtype, DUX4r subtype; classifying an AML as a RUNX1::RUNX1T1 subtype, KMT2Ar subtype, NUP98::NSD1 subtype, FLT3-ITD subtype, NPM1-mutated subtype or PML-RARA subtype, CBFB::MYH11 subtype, DEK::NUP214 subtype, RBM15::MRTFA subtype, BCR::ABL1 subtype, MECOM rearrangement subtype, NUP98 rearrangement subtype, CEBPA mutation subtype; or classifying the acute leukemia as any of the subtypes in Table 2. Likewise, a solid tumor, in particular a pediatric sarcoma may be classified as an osteosarcoma, Ewing sarcoma, rhabdomyosarcoma, synovial sarcoma, desmoplastic small round cell tumor, alveolar soft partAttorney Docket No.5470.980.WO sarcoma, desmoid tumor, dermatofibrosarcoma protuberan, infantile fibrosarcoma, leiomyosarcoma, or liposarcoma.

[0070] In some embodiments, the method as described herein is completed in less than 24 hours, e.g., less than 24, 22, 20, 18, 16, 14 or 12 hours.

[0071] Classification of a cancer as a particular type or subtype of cancer allows for improved risk-adapted therapeutic selection and precision therapy. For example, a subject classified with a known mutation associated with a relatively favorable prognosis (e.g., isolated 9p / CDKN2A- CDKN2B deletions) may receive a different treatment than those with an intermediate prognosis (e.g., 6q deletions), unfavorable prognosis (e.g., t(9;22) / BCR / ABL1, t(4;11) / MLL / AF4, and t(1;19) / TCF3 / PBX1) or high risk prognosis (e.g., t(9;22)(q34;q11)). Accordingly, the method of the invention further provides for treating a subject based on the classification. Such treatment may include, but is not limited to, chemotherapy, targeted therapy (e.g., antibody therapy), immunotherapy, hormone therapy, radiotherapy, surgery, or any combination thereof.

[0072] Examples of chemotherapies include, but are not limited to, antimetabolites (e.g., folic acid, purine, and pyrimidine derivatives) and alkylating agents (e.g., nitrogen mustards, nitrosoureas, platinum, alkyl sulfonates, hydrazines, triazenes, aziridines, spindle poison, cytotoxic agents, toposimerase inhibitors and others). Exemplary agents include Aclarubicin, Actinomycin, Alitretinon, Altretamine, Aminopterin, Aminolevulinic acid, Amrubicin, Amsacrine, Anagrelide, Arsenic trioxide, Asparaginase, Atrasentan, Belotecan, Bexarotene, endamustine, Bleomycin, Bortezomib, Busulfan, Camptothecin, Capecitabine, Carboplatin, Carboquone, Carmofur, Carmustine, Celecoxib, Chlorambucil, Chlormethine, Cisplatin, Cladribine, Clofarabine, Crisantaspase, Cyclophosphamide, Cytarabine, Dacarbazine, Dactinomycin, Daunorubicin, Decitabine, Demecolcine, Docetaxel, Doxorubicin, Efaproxiral, Elesclomol, Elsamitrucin, Enocitabine, Epirubicin, Estramustine, Etoglucid, Etoposide, Floxuridine, Fludarabine, Fluorouracil (5FU), Fotemustine, Gemcitabine, Gliadel implants, Hydroxycarbamide, Hydroxyurea, Idarubicin, Ifosfamide, Irinotecan, Irofulven, Ixabepilone, Larotaxel, Leucovorin, Liposomal doxorubicin, Liposomal daunorubicin, Lonidamine, Lomustine, Lucanthone, Mannosulfan, Masoprocol, Melphalan, Mercaptopurine, Mesna, Methotrexate, Methyl aminolevulinate, Mitobronitol, Mitoguazone, Mitotane, Mitomycin, Mitoxantrone, Nedaplatin, Nimustine, Oblimersen, Omacetaxine, Ortataxel, Oxaliplatin, Paclitaxel, Pegaspargase, Pemetrexed, Pentostatin, Pirarubicin, Pixantrone, Plicamycin, Porfimer sodium, Prednimustine,Attorney Docket No.5470.980.WO Procarbazine, Raltitrexed, Ranimustine, Rubitecan, Sapacitabine, Semustine, Sitimagene ceradenovec, Strataplatin, Streptozocin, Talaporfin, Tegafur-uracil, Temoporfin, Temozolomide, Teniposide, Tesetaxel, Testolactone, Tetranitrate, Thiotepa, Tiazofurine, Tioguanine, Tipifarnib, Topotecan, Trabectedin, Triaziquone, Triethylenemelamine, Triplatin, Tretinoin, Treosulfan, Trofosfamide, Uramustine, Valrubicin, Verteporfin, Vinblastine, Vincristine, Vindesine, Vinflunine, Vinorelbine, Vorinostat, Zorubicin, and other cytostatic or cytotoxic agents.

[0073] Targeted therapy constitutes the use of agents specific for a deregulated protein of a cancer cell. Small molecule targeted therapy drugs are generally inhibitors of enzymatic domains on mutated, overexpressed, or otherwise critical proteins within the cancer cell. Prominent examples are the tyrosine kinase inhibitors such as axitinib, bosutinib, cediranib, desatinib, erlotinib, imatinib, gefitinib, lapatinib, lestaurtinib, nilotinib, semaxanib, sorafenib, sunitinib, and vandetanib, and also cyclin-dependent kinase inhibitors such as alvocidib and seliciclib. Monoclonal antibody therapy is another strategy in which the therapeutic agent is an antibody which specifically binds to a protein on the surface of the cancer cells. Examples include the anti- CD20 antibody rituximab or tositumomab typically used in a variety of B-cell malignancies. Targeted therapy may also involve small peptides as “homing devices,” which can bind to cell surface receptors or affected extracellular matrix surrounding the tumor. Radionuclides which are attached to these peptides (e.g., RGDs) eventually kill the cancer cell if the nuclide decays in the vicinity of the cell. An example of such therapy includes BEXXAR®.

[0074] Immunotherapies may include, e.g., CAR T-cell therapies such as brexucabtagene autoleucel or obecabtagene autoleucel used in the treatment of adults with B-cell ALL, and tisagenlecleucel used in the treatment of children and young adults up to age 25 with B-cell ALL.

[0075] By way of illustration, the treatment regimen for subjects with ALL may be determined by classification of the Philadelphia chromosome status of the leukemia and the age of the patient. Patients classified as Philadelphia chromosome-positive (Ph+) ALL may be treated with a tyrosine kinase inhibitor (TKI) in combination with chemotherapy. Examples of suitable TKIs include, but are not limited to imatinib, dasatinib, nilotinib, bosutinib and ponatinib. Ph+ ALL Patients aged 15-39 years are referred to as "AYA" (adolescent and young adult) and may be treated with more intensive pediatric-style treatment regimens including induction (e.g., combinations of drugs, including vincristine, prednisone, cyclophosphamide, doxorubicin, and asparaginase), consolidation (multiagent therapy including cytarabine and methotrexate), and maintenanceAttorney Docket No.5470.980.WO therapy (including 6-mercaptopurine, methotrexate, steroids, and vincristine). Older adult Ph+ ALL patients (age ≥40 y) may be treated with chemotherapy (e.g., hyper-CVAD) plus TKI. Patient classified as Philadelphia chromosome-negative (Ph-) ALL as an older adult (age ≥40 y) may be treated with a multiagent chemotherapy regimen (e.g., CALGB 8811 [daunorubicin, vincristine, prednisone, pegaspargase, and cyclophosphamide]. Patients classified as Philadelphia chromosome-negative B-cell ALL may be treated with blinatumomab and patients classified as having relapsed or refractory CD22-positive B-cell ALL may be treated with Inotuzumab. Patients classified as relapsed or refractory acute leukemia that have a KMT2A gene translocation may be treated with revumenib. Patients classified as having AML may receive treatment with combination chemotherapies such as cytarabine and anthracyclines. In addition, targeted therapies may be used. For example, a patient classified with an FLT3 gene mutation may receive a therapy that includes an FLT3 inhibitor and patients with IDH1 or IDH2 gene mutations may receive a therapy that includes IDH inhibitors (e.g., ivosidenib, enasidenib, or vorasidenib).

[0076] The administration of a therapeutic described herein may be, for example, by injection, intravenously, intraarterially, subdermally, intraperitoneally, intramuscularly, or subcutaneously; or orally, buccally, nasally, transmucosally, topically, in an ophthalmic preparation, or by inhalation, with a dosage ranging from about 0.02 to about 100 mg / kg of body weight, alternatively dosages between 1 mg and 1000 mg / dose, every 4 to 120 hours, or according to the requirements of the particular drug. The methods herein contemplate administration of an effective amount of a therapeutic to achieve the desired or stated effect. Such administration may be used as a chronic or acute therapy. Lower or higher doses than those recited above may be required. Specific dosage and treatment regimens for any particular patient will depend upon a variety of factors, including the activity of the specific compound employed, the age, body weight, general health status, sex, diet, time of administration, rate of excretion, drug combination, the severity and course of the disease, condition or symptoms, the patient's disposition to the disease, condition or symptoms, and the judgment of the treating physician.

[0077] Upon improvement of a patient's condition, a maintenance dose of a therapeutic agent may be administered, if necessary. Subsequently, the dosage or frequency of administration, or both, may be reduced, as a function of the symptoms, to a level at which the improved condition is retained when the symptoms have been alleviated to the desired level. Patients may, however, require intermittent treatment on a long-term basis upon any recurrence of disease symptoms.Attorney Docket No.5470.980.WO

[0078] The following non-limiting examples are provided to further illustrate the present invention. EXAMPLES Example 1: Material and Methods

[0079] Samples. We performed ONT whole genome sequencing (WGS) on fifty-seven (57) acute leukemia specimens representing diverse clinically diagnosed genomic subtypes (retrospective sampling; n = 39) or new diagnoses before clinical genomic subtyping (real-time sampling; n = 18) (Table 1). Table 1. The number, lineage, and genomic subtype of specimens sequenced in this study using nanopore whole-genome sequencing with adaptive sampling. Number of Samples LineageaGenotype Subtype 9 B-ALL High hyperdiploidy (>50 chromosomes)

[0080] DNA from these specimens was extracted from peripheral blood mononuclear cells, bone marrow mononuclear cells, or whole blood (Table 2). Samples were obtained from the University of North Carolina at Chapel Hill (UNC) and St. Jude Children’s Research Hospital (SJCRH) with approval by their respective Institutional Review Boards. Illumina RNA sequencing was available for samples from SJCRH. Clinical diagnosis at UNC was determined by G-banding karyotype analysis, FISH, and sometimes microarray as part of the standard of care.Attorney Docket No.5470.980.WO Table 2. Sample metadata and associated translocations. Nanopore Determined Fusions / Small Fusions / Small Sample Gene Panel Lineage Genomic Subtype Alterations Alterations A T TAttorney Docket No.5470.980.WO 0171d152RTB-ALL ZNF384r EP300::ZNF384 EP300::ZNF384 0172e152RTAML KMT2Ar KMT2A::USP2 KMT2A::USP2 0173d152RTB-ALL High hyperdiploidy no abnormalities no abnormalities dT, real-time.

[0081] DNA extraction and shearing. DNA was extracted from cryopreserved or fresh samples using the ZymoBIOMICS™ MagBead DNA / RNA kit following manufacturer’s instructions (Zymo Research). Fragment sizes in excess of 20 Kbp (without significant degradation) were verified by gel electrophoresis. Extracted DNA was sheared using a 26G-1″ needle for a total of seven passes to obtain a fragment size distribution for optimal nanopore sequencing throughput. Size selection was performed using 0.4 volumes of AMPURE® XP Beads (Beckman Coulter). DNA was quantified using the QUBIT® fluorometer with the QUBIT® dsDNA Quantification, High Sensitivity Assay Kit (ThermoFisher Scientific).

[0082] Library preparation and sequencing. Ligation-based library preparation of native DNA was performed using the following ONT library preparation kits: SQK-LSK 109, SQK-LSK 110, SQK-LSK 112, and SQK-LSK 114, SQK-NBD 114. All library preparation was conducted following the manufacturer’s instructions, with the exception of SQK-LSK 114, which was modified in several ways, including the exclusion of 1 μL DNA CS, and a reduction in the final room temperature incubation time of 5 minutes instead of the listed 10  minutes. Samples wereAttorney Docket No.5470.980.WO sequenced either multiplex or singly on FLO-MIN106, FLO-PRO002, FLO-PRO112, or FLO- PRO114M flow cells for up to 72 hours or until the available pores were exhausted. Except four samples, samples were sequenced on a PROMETHION™ 2 Solo (P2) machine or MINION™ device using adaptive sampling. Initial sequencing using adaptive sampling enriched for 59 genes frequently involved in B-ALL and AML translocations and fusions (Table 3). Table 3. List and coordinates of genes used in 59-gene set enrichment. GRCh38 GRCh38 Chromosome Start.. End Gene Chromosome Start.. End GeneAttorney Docket No.5470.980.WO NC_000021.938330027..150063839.. 38711780 ERG NC_000005.10 150205872PDGFRB11599674128569969 1Attorney Docket No.5470.980.WO NC_000005.10150003291..34293315.. 150163372 CSF1R NC_000015.10 3440NUTM177371462581181738601

[0083] Subsequently, two larger gene panels were employed: a 152-gene panel was used for six samples, and an expanded 223-gene panel for four samples. These enhanced panels were designed to capture additional genes and genetic regions implicated in B-ALL, AML, and T-ALL (Table 4 and Table 5). To inform the scope of relevant fusions, we referenced previous work detailing the landscape of ALL and AML genomic subtypes (Brady et al. (2022) Nat. Genet. 54:1376-89; Umeda et al. (2024) Nat. Genet. 56:281-93). First, single partner genes were parsed from gene fusions detected by RNA-seq or WGS. Genes that occurred as a fusion partner in more than one distinct case were included. We included the entire genomic range for each gene or locus (e.g., IGH), as annotated on GRCh38, and a margin of 50Kbp on either side. Reads were base-called (and de-multiplexed, if applicable) using Dorado (v0.5.1-0.6.0) in super-accurate duplex mode.Attorney Docket No.5470.980.WO Base-called reads were aligned to the GRCh38 human reference genome using minimap2 (Li (2018) Bioinformatics 34:3094-100). Table 4. List and coordinates of genes used in 152-gene set enrichment. GRCh38 GRCh38 Chromosome Start..End Gene Chromosome Start..End GeneAttorney Docket No.5470.980.WO 92895194 36570225 NC_000009.1221917752..32393059.. 22044491 CDKN2A NC_000020.11 3263533NOL4L311Attorney Docket No.5470.980.WO 132635188 128746396 NC_000002.12218931087..36423288.. 219035657 FEV NC_000014.9 36563785SFTA321Attorney Docket No.5470.980.WO 91112159 37408149 NC_000011.108174304..831 LMO1 NC_000002.12144334375.. 8635 144570391ZEB2. GRCh38 Start.. GRCh38 Start.. Chromosome End Gene Chromosome End Gene / Attorney Docket No.5470.980.WO NC_000023.11129930302..70059768.. 130108083 BCORL1 NC_000008.11 70453NCOA28274601098523130365Attorney Docket No.5470.980.WO NC_000010.1194712706..218609656.. 94902914 CYP2C19 NC_000002.12 21871344PTD387813438713879661 1Attorney Docket No.5470.980.WO NC_000012.12132440455..36423288.. SFTA3 / 132635188 FBRSL1 NC_000014.936570225 NKX2-1 1113559172189310872 / LAKAttorney Docket No.5470.980.WO 17898071 10430572 NC_000008.1141879479..43042956.. 42101988 KAT6A NC_000021.9 43157578U2AF11

[0084] Digital karyotyping and aneuploidy inference. Relative copy number across the genome at the chromosome level (“digital karyotype”) was inferred based on relative sequencing depth by assessing genome-wide and chromosome-level depth of coverage (FIG.1A). Briefly, we infer a baseline diploid (uniform) sequencing depth equivalent to the non-blast percentage (typically low), then assess the relative read depth above baseline, where a 2:3 ratio is observed between diploid and triploid chromosomes and 1:2 between haploid and diploid, respectively. To compare with G- banding karyotypes and assess gross aneuploidy levels, we discuss only whole-chromosome and arm-level gains and losses, although smaller subchromosomal gains and losses are also clearly evident. To avoid the potentially confounding effect of adaptive sampling on relative sequencing depth assessment, we constructed this coarse-scale depth as a function of reads per million baseAttorney Docket No.5470.980.WO pairs (Mbp), where each read contributes a count of one to the bin in which the center of the read aligns. To assess minimum sensitivity to detect chromosome-level copy number changes, we consider cumulative reads at each timepoint (each minute) after sequencing is initiated, based on read timestamps. We apply a pairwise Kolmogorov-Smirnov test with a conservative threshold (p < 1e-9) and a minimum median divergence of 20% to determine the earliest timepoint at which chromosomes exhibit significantly divergent copy number meeting the clinically-relevant threshold for high hyperdiploid (≥53) or low hypodiploid / near haploid (≤39) karyotype.

[0085] Translocation and fusion detection. Putative translocations were characterized by counting reads for which multiple alignments existed to two independent genes in our enrichment set (e.g., ETV6 and RUNX1) (FIG. 1D). A read was considered “anchored” in a gene if an alignment of ≥500 nt existed within 5Kbp of the gene—this accounted for rearrangements involving the translocation of promoters outside the annotated gene boundaries. Known repetitive elements, centromere and satellite sequences, and segmental duplications were masked to avoid false positives caused by ambiguous alignments and transposable elements. Translocations between two genes in our target set with at least two independent supporting reads were subsequently validated by visualization of the supporting and putative breakpoints. Fusions detected in our samples had a minimum of 3, and as many as 104 supporting (non-duplex) reads (Table 6). Two apparently independent reads can sometimes represent both strands of a “duplex” read that were not appropriately collapsed during duplex basecalling. Duplex reads were identified conservatively and considered a single read if two reads supporting the same translocation were acquired through the same sequencing channel within 30 s of one another. Table 6. Translocation / fusion genes and read-level support. Supporting Secondary Supporting l F i R F i RAttorney Docket No.5470.980.WO 0130 MEF2D::CSF1R 210131 - -Attorney Docket No.5470.980.WO 0226 - -0227 - -g py . mal) copy number variation (e.g., iAMP21, CDKN2A, and ERG deletion, CRLF2-P2RY8 interstitial deletion) was determined by a combination of sequencing depth and split read alignment evidence (FIG.1B). Focal deletions are characterized by a haploid (or multiple thereof) drop in read depth and reads split-mapped to either side of the deletion breakpoints. We consider the full per- nucleotide sequencing depth to evaluate intragenic copy-number variation within genes in our enrichment set.

[0087] Targeted SNV and insertion / deletion calling. Due to uneven coverage resulting from adaptive sampling enrichment and to additionally capture moderate-sized insertions / deletions, existing SNV calling tools for nanopore sequencing data were found to be ineffective. We screened our samples for clinically relevant single nucleotide variants (SNVs), specifically focusing on mutations in NPM1, FLT3-ITD, and TPMT / NUTD15. We implemented a straightforward reference-guided assembly approach to identify SNVs and small-scale insertions and deletions (indels) above 0.3 minor allele frequency (MAF). This threshold of 0.3 MAF detected heterozygous and homozygous variants consistent with known molecular genetics across our pediatric leukemia cohort, the majority of which have blast percentages greater than 80%. Briefly, within each enriched gene region, we align all overlapping reads to the GRCh38 reference genome and build a consensus sequence, including SNVs and indels at or above 0.3 MAF. All overlapping reads are subsequently realigned to the consensus sequence and SNV and indel variants exceedingAttorney Docket No.5470.980.WO 0.3 MAF are reported. Small-scale indels, notably FLT3-ITD, were likewise identified by building a local consensus sequence from a high-depth sequence covering our target genes. We additionally devised a notion of in silico PCR to recapitulate commonly used capillary electrophoresis methods to characterize FLT3-ITD (Kiyoi et al. (1999) Blood 93:3074–80) where read segments bounded by FLT3 11F (GCAATTTAGGTATGAAAGCCAGC; SEQ ID NO:1) and 12R (CTTTCAGCATTTTGACGGCAACC; SEQ ID NO:2) primers were extracted and plotted by size.

[0088] Evaluation and validation. Leukemia lineage was determined by flow cytometry. The ground truth of subtypes was determined by G-banding karyotype analysis, FISH, microarray, Illumina sequencing, or some combination thereof. Illumina RNA sequencing and fusion detection and / or expression-based subtyping (e.g., DUX4r) determined a subset of cases. We evaluated the performance of our WGS-based analysis against the final consensus subtype following this multimodal characterization. Example 2: Specificity and Sensitivity of Nanopore-based WGS

[0089] Nanopore-based WGS demonstrated 100% specificity and 96% sensitivity in aggregate for determining genomic subtype across both retrospective and real-time samples. The choice of library preparation method and sequencing platform affected assay performance only through their influence on read length and sequencing coverage. Of the 57 samples analyzed, we accurately characterized gross karyotype abnormalities in all but two cases, achieving 96% sensitivity. Both occurred in specimens with very low (<30%) blast content; among specimens with >50% blasts (53 / 57), we achieved 100% accuracy. Our approach achieved 100% specificity in the sense that no gross aneuploidy was reported for cytogenetically diploid cases.

[0090] Nanopore sequencing successfully identified all fusion events previously detected by clinical testing, achieving 100% sensitivity for detection of known fusion oncogenes. Moreover, in five cases, this approach revealed additional fusion partners that were not identified through clinical karyotyping and FISH. These included cytogenetically cryptic (e.g., DUX4::IGH) fusions and rearrangements where only one gene partner was known (e.g., PDGFRB / CSF1R rearranged) that were fully resolved by nanopore sequencing. In all cases, nanopore-based translocations were consistent with available clinical cytogenetic evidence, achieving 100% specificity.Attorney Docket No.5470.980.WO

[0091] Although genome-wide SNV calling was not performed in clinical testing, our sequencing revealed the correct SNV in each of the four cases with a clinically identified SNV; clinically relevant SNVs including TMPT and NUDT15 were assessed for in each nanopore- generated dataset regardless of positive clinical findings (Tables 7-10). Across all cases, nanopore WGS with adaptive sampling confirmed all but two clinically identified genomic alterations used for risk stratification and treatment planning. Additionally, this method revealed DUX4 rearrangements in two cases that were not detected by standard clinical testing, potentially warranting revised risk assessments for these patients. Table 7. Focal copy-number variation and SNVs in enriched clinically relevant IKZF1, CDKN2A and PAX5 gene targets. Sample IKZF1 CDKN2A PAX5 0048 - - - - - -Attorney Docket No.5470.980.WO 0157 - - - - - - 0158 - - - - - - 7Table 8. Focal copy-number variation and SNVs in enriched clinically relevant ETV6, EBF1, and ERG gene targets.Attorney Docket No.5470.980.WO Sample ETV6 EBF1 ERG 0048 - - - - - -Attorney Docket No.5470.980.WO 0203 - - - - - - 0209 - - SNV S442G - -Table 9. Focal copy-number variation and SNVs in enriched clinically relevant ETV6, EBF1, and ERG gene targets. Sample RUNX1 FLT3 NPM1Attorney Docket No.5470.980.WO 0130 SNV S305A(0.42 AF)- - - -Attorney Docket No.5470.980.WO (0.45 AF) 0163 - - - - - -Attorney Docket No.5470.980.WO 0230 - - - - - - 0231 - - SNV M227T - -Table 10. Focal copy-number variation and SNVs in enriched clinically relevant TPMT and NUDT15 gene targets. Sample TPMT NUDT15Attorney Docket No.5470.980.WO 0141- - - -0145- - - -Attorney Docket No.5470.980.WO 0234 - - - - 0238 - - - -nerated an average of 12X whole-genome coverage and 86X coverage over target genes. For all samples run with adaptive sampling (including multiplexed aneuploid samples), we achieved a relative enrichment of 8.3X (range 1.46–16.4X). Samples exhibiting very poor enrichment resulted from severely fragmented input DNA (0154, 0160) or adaptive sampling failure (0229). Samples undergoing adaptive sampling from fresh samples for real-time analysis ranged from 4.2–12.5X enrichment. Example 3: Nanopore whole genome sequencing with and without adaptive sampling accurately identifies clinically relevant karyotype profiles and gene fusions in acute leukemias

[0093] The karyotype for 27 of 57 samples showed gross changes at the chromosome level (aneuploidy) (FIG. 2A-2C). Clinically relevant subtypes classified included high hyperdiploidy (>50 chromosomes; n = 9), low hypodiploidy (31–39 chromosomes; n = 4), and near-haploidy (24– 30 chromosomes; n = 4) (Table 1). Ten additional samples were aneuploid (45–50 chromosomes) with (n = 8) or without (n = 2) other known genomic drivers (Table 2). In all 27 aneuploid cases, changes in gross karyotype detected by WGS nanopore sequencing (both with (n = 25) and without (n = 2) adaptive sampling) were consistent with clinical classification.2 / 57 samples had low blast count (<30%) that prohibited precise genomic classification (samples 0131 and 0133). In cases 0154 and 0164, we estimated six copies of RUNX1 within a broader regional amplification pattern, indicating an intrachromosomal amplification of chromosome 21 (iAMP21) B-ALL genomic subtype, defined by at least four copies of RUNX1 and focal amplification.

[0094] We detected 19 unique gene fusions across all samples (Table 2 and Table 6). Identified fusion-driven genomic subtypes of B-ALL were: ETV6::RUNX1, TCF3::PBX1, Philadelphia (Ph) (BCR::ALB1), Ph-like including CRLF2r (CRLF2::IGH; PAX5::MLLT3; MEF2D::CSF1R; EBF1::PDGFRB), MEF2Dr (MEF2D::BCL9; MEF2D::HNRNPUL1), ZNF384rAttorney Docket No.5470.980.WO (ZNF384::EP300; TCF3::ZNF384; CREBBP::ZNF384), and DUX4r (DUX4::IGH). Fusion- driven genomic subtypes of AML identified (n = 3) were: RUNX1::RUNX1T1, KMT2Ar (KMT2A::MLLT1, KMT2A::MLLT10; KMT2A:USP2), and NUP98::NSD1 (Table 1 and Table 2). Gene fusions within B-ALL fell into two major categories: balanced translocations and complex rearrangements (e.g., DUX4r; CRLF2r). The primary characterization of structural rearrangements from long-read WGS consists of reads aligning to distant genes or regions. Adaptive sampling enriches for reads spanning involved genes, providing strong and consistent support for these structural variants (FIGS. 3A-3C and FIGS. 4A-4C). In several instances, nanopore-generated WGS data provided additional information about structural variation that was not detected through clinical assays. In one case, clinical classification identified one of two fusion partners, whereas nanopore WGS identified both fusion partners; for sample 0130, clinical testing classified sample 0130 as PDGFRB or CSF1R with an unknown fusion partner while sequencing data identified a precise MEF2D::PDGFRB fusion. In other cases, clinical break-apart FISH assays only identified one component of the likely fusion (NUP98 in sample 0158 and KMT2A in sample 0172), while the partner gene with prognostic significance was only identified after nanopore sequencing, (NUP98::NSD1 and KMT2A::USP2). In two cases (0157 and 0162), an IGH::DUX4 fusion clinically (by karyotyping and FISH) butwas detected in our real-time sequencing analysis. Example 4: Small-scale and intragenic copy-number variation and single nucleotide variation are sensitively characterized in enriched regions

[0095] Nanopore WGS produces support for clinically relevant small-scale structural variants affecting genes associated with ALL and AML, including CRLF2-P2RY8 interstitial deletion (0060, 0136, 0163, 0238), partial or total loss of CDKN2A (0153, 0165), and heterogeneous ERG deletion (0162). A full list of observed small variants, including CNVs and SNVs is provided in Tables 7-10. We accurately detect these deletions with a combination of sequencing depth and long reads spanning the deletion boundaries (FIG.4C).

[0096] We assessed the utility of adaptive WGS to call single-nucleotide polymorphisms (SNPs) and small insertions by examining pharmacogenomically relevant SNVs in TPMT, and FLT3 internal tandem duplication (FLT3-ITD), a driving mutation in pediatric AML (Table 6). We identified relevant SNVs in TPMT in both cases (0136, 0162) with clinically identified mutationsAttorney Docket No.5470.980.WO (A154T, Y240C). These are trivially phased in our long reads and confirmed to occur on the same haplotype. FLT3-ITD was called in one AML case (0141), characterized by a consensus tandem duplication of FLT3 CDS loci 1823–1904 (81 nt). To recapitulate commonly used PCR and capillary electrophoresis detection of FLT3-ITD (Kiyoi et al. (1999) Blood 93:3074-80), we identified 46 reads that include FLT311F and 12R primer sites. The size distribution of these ITD- spanning reads is shown in FIG. 5, producing an ITD:WT allelic ratio (AR) of 0.64, consistent with the clinically reported AR of 0.65. Example 5: Clinically relevant genomic variation and tumorigenic drivers are robustly detectable with a single rapid, low-cost assay

[0097] Real-time sequencing and analysis generates sample classifications within 9 h of sample receipt. DNA extraction and shearing using the ZYMOBIOMICS™ MagBead DNA / RNA kit took less than 2 hours, and library preparation took approximately 75 minutes. Using the high- throughput PROMETHION™ flow cells, a digital karyotype could be inferred within 15- 30 minutes of the start of sequencing. Fusion detection (defined as two or more independent reads supporting a single given fusion) typically took between 3 and 6 hours from the start of sequencing (FIG. 6). In concordance with our retrospective sampling, real-time classification using our nanopore-based WGS approach was 100% consistent with clinically-derived genomic subtype classification.

[0098] To validate detection limits, we analyzed three sample sets multiplexed on individual flow cells: two sets of seven samples with known aneuploidies and one set of four samples with known translocations. For the aneuploid and fusion-positive samples, we obtained an average of 8.6X and 23.6X coverage of enriched targets, respectively. These are broadly as expected given the level of multiplexing. In all cases, we successfully identified the clinically determined genomic subtype. We retroactively assessed the timing of translocation detection across the entire cohort. Briefly, since nanopore sequencing reads are captured asynchronously and continuously from the time sequencing is initiated, and each read is time-stamped with the time it started and finished sequencing, we can infer the total duration of sequencing necessary to call a translocation based on the time the second independent supporting read finished sequencing. Analysis of sequencing data is performed continuously in real-time and adds a negligible time for analysis after the second read is complete—we conservatively add 5 minutes. FIGS. 7A-7B show the distribution ofAttorney Docket No.5470.980.WO minimum necessary sequencing times to conclusively identify translocations using this method. When using adaptive sampling and running samples individually, 19 / 25 (76%) of translocations are identified in under 3 hours, and 23 / 25 (92%) are identified in under 6 hours (FIG. 7A). As expected, samples multiplexed on one flow cell or run without adaptive sampling take longer to detect the translocations. Two samples that were sequenced individually with adaptive sampling are apparent outliers, taking 11-14 hours to result. Sample 0136 is one, taking 11.2 hours to identify the putative driving translocation (CRLF2::P2RY8) based on the clinical diagnosis, however we also find a second putative driver, PAX5::MLLT3, in only 2 hours. For the other, Sample 0231, we detected MEF2D::BCL9 in 13.9 hours; the total sequencing efficiency and throughput were lower than typical for this sample resulting in a longer detection time.

[0099] We additionally determined the minimum amount of time needed to accurately detect and characterize chromosomal abnormalities by analyzing statistically and clinically relevant changes in chromosome copy number. We excluded the two specimens with very low blast count (<30%), as previously described. Among cases with gross karyotype abnormalities as primary genomic subtypes (high hyperdiploid, low hypodiploid, near haploid), four were sequenced individually on a high-capacity flow cell, one individually on a low-capacity Flongle flow cell, and the remainder multiplexed on a P2 flow cell. Singleplex high-throughput runs were characterized into aneuploidy classes after a median of 5 min and a maximum of 6 min (FIG.7B). Multiplexed specimens were called after a median of 16 and maximum of 39 minutes. The specimen run on a low-throughput Flongle flow cell (0048) was called after 61  minutes. Illustrative sequencing depth profiles at this minimum detection timepoint are shown in FIGS.8A-8D. Real-time analysis and assessment of karyotype abnormalities showed that these signatures are robustly detected within minutes of the start of sequencing. However, in all cases sequencing continued well past this point to identify possible concomitant translocations (e.g., BCR::ABL1). Example 6: Simplified identification of clinically useful genomic results

[0100] Advancements in genomic classification of acute leukemia over the past decades have led to improved prognostic stratification, risk-adapted therapeutic selection, and precision therapy. While some treatment centers perform comprehensive genomic profiling of all patients with acute leukemia, this is the exception. At most treatment centers in high-income countries (HICs), genomic classification continues to rely on long-established techniques in cytogenetics, supportedAttorney Docket No.5470.980.WO by targeted molecular profiling. More importantly, most cancer treatment centers globally are in low and middle-income countries (LMICs), often without access to reliable cytogenetic testing approaches. Therefore, accessible, scalable approaches to improve clinical genomic classification of acute leukemia are critically needed.

[0101] The motivation to pursue nanopore adaptive WGS is for simplified identification of well- established clinically useful genomic results, as opposed to a focus on discovering new biological insights. The results presented herein indicate that, as a single assay classification tool, nanopore- based adaptive whole-genome sequencing accurately classifies B-ALL into genomic subtypes, with the potential to identify clinically relevant AML genomic subtypes, as well as clinically actionable pharmacogenetic subtypes. We have demonstrated proof of principle that nanopore long-read WGS can provide all clinically relevant genomic information currently offered by traditional diagnostic testing (karyotype, FISH, and occasional microarray) for pediatric acute leukemia. In our cohort, there was no loss in sensitivity or throughput with a variety of sample preparations and robust results were generated from freshly collected as well as cryopreserved samples, including whole blood, isolated mononuclear cells, bone marrow aspirate, and peripheral blood. Likewise, nanopore-based WGS is less reliant on high-quality viable cells that are beneficial for karyotype and FISH analysis. Our approach comprehensively identifies known variation commonly characterized by a combination of karyotyping, FISH, and targeted molecular tests.

[0102] The long reads produced by nanopore sequencing are particularly useful for identifying complex structural rearrangements that are difficult to identify with current clinical approaches, such as DUX4 rearrangements. DUX4 rearrangements make up about 14% of B-ALL cases (Lee et al. (2021) Cancers 13:4068), yet structural variations involving DUX4 are often not well characterized due to tandem D4Z4 repeat cassettes containing DUX4 (Rehn et al. (2020) Cancers 12:2815). The identification of DUX4 rearrangements commonly relies on the detection of fusion transcripts by RNA sequencing (Barinka et al. (2022) Leukemia 36:1492-8). We detected DUX4::IGH gene fusions in three samples (one retrospective sample, 0222, and two real-time samples, 0157 and 0162) that were not identified by conventional FISH or karyotype analysis. Under standard clinical classification, these cases would have been designated as risk-neutral. However, emerging evidence suggests that DUX4 fusions may indicate a more favorable prognosis. These findings demonstrate that nanopore-adaptive WGS can reveal clinically relevant genomic alterations even in settings with sophisticated diagnostic capabilities.Attorney Docket No.5470.980.WO

[0103] This assay provides a level of characterization of aneuploid and small copy number changes that is critical for clinical decision-making. While hyperdiploid karyotype is generally a favorable prognostic factor, multiple groups have demonstrated that the prognosis is influenced by specific chromosome gains. More recently, multiple groups have demonstrated the potential prognostic value of small deletions in certain contexts, specifically IKZF1 deletion (Boer et al. (2016) Leukemia 30:32-8; Mullighan et al. (2009) N. Engl. J. Med. 360:470-80). It is important for future genomic classification approaches to B-ALL to include small deletions in the diagnostic results.

[0104] Pharmacogenomics is a rapidly expanding field with increased clinical relevance. For patients with ALL, the key pharmacogenomic information needed is TPMT and NUDT15 genotype, which has direct clinical implications. In HICs, the genotype is typically determined by a targeted molecular assay. We show the potential for nanopore adaptive WGS to detect known SNPs in TPMT with the same assay that provides genomic classification of B-ALL. This provides another example of cost savings in HIC and the expansion of clinically relevant pharmacogenomic information in LMICs. To capture additional information and expand the use of nanopore WGS in classifying pediatric leukemias, an extended 260 gene panel was developed for pediatric leukemias (Table 11). Table 11. List and coordinates of genes used in 260-gene set enrichment. GRCh38 GRCh38 Chromosome Start..End Gene Chromosome Start..End GeneAttorney Docket No.5470.980.WO DUX4L14 / DUX4L15 36278923529238Attorney Docket No.5470.980.WO NC_060926.1197823439..75961669.. 197968732 SF3B1 NC_060936.1 7610972NAP1L1692067840208666002 / 2-Attorney Docket No.5470.980.WO NC_060928.1108406807..15598199.. MIR484 / 108640746 TET2 NC_060940.115911962 NDE1 / MYH11 30418184155593699Attorney Docket No.5470.980.WO 96830392 61585104 NC_060930.1110494551..77581102.. 110610504 CD164 NC_060941.1 776843SRSF298Attorney Docket No.5470.980.WO NC_060931.1149939157..31028176.. 150116070 EZH2 NC_060944.1 3112897DUX4L35731339166153257815Attorney Docket No.5470.980.WO NC_060933.1142041913..117943452.. 14229 FNBP1 NC_060947.1SEPT67931 118121836122220973142867999p , , ces classification is constrained in cases with low blast percentages and is sometimes limited to subtypes defined by genomic structural variation. For instance, gains of chromosomes 4, 6, 14, 17, 18, and 21 were detected by microarray (but not traditional karyotyping) in sample 0133 - these gains were largely undetectable with our current analytical pipeline. Likewise, with patient sample 0131 (near haploid), gains of chromosomes 21 and X were shown through clinical testing but were undetectable in nanopore WGS. Additionally, a small portion of acute leukemia cases are defined by expression profiles without subtype defining DNA variation, such as the proposed ETV6-like, KMT2A-like, and ZNF384-like cases, limiting the utility of DNA-based classification approaches in these cases (Gu et al. (2019) Nat. Genet.51:296-307).

[0106] There are, in several cases, large-scale structural variants, including whole-chromosome gains and losses, that are detected by nanopore sequencing that are inconsistent with the reported clinical G-band karyotype (Table 2). Because clinical karyotyping involves growing cells in culture, genetic drift may occur, meaning that both the original (G-band karyotype) and observed (nanopore-derived) genetic classifications could be valid representations of the cells at different points in time. Previous digital karyotyping approaches based on genomic data have noted similar discrepancies in a minority of cases, emphasizing possible G-band karyotyping inaccuracies (Barinka et al. (2022) Leukemia 36:1492-8). However, these differences are typically minor andAttorney Docket No.5470.980.WO do not result in a different clinically relevant aneuploidy class (high hyperdiploid, low hypodiploid, near haploid). Example 7: Classification of solid tumors

[0107] In addition to B-ALL, this approach has been extended to solid malignancies, in particular pediatric solid tumors, e.g., pediatric sarcomas. For this analysis, resected tissue or tissue biopsies were used. The genomic targets are provided in the 157 for solid tumors, e.g., pediatric sarcomas (Table 12). Table 12. List and coordinates of genes used in 157-gene set enrichment. GRCh38 GRCh38 Chromosome Start..End Gene Chromosome Start..End Gene 1 3Attorney Docket No.5470.980.WO 164282389 57865097 NC_060925.1233944665..61637128.. FAM19A2 / 234050583 IRF2BP2 NC_060936.162445185 USP15Attorney Docket No.5470.980.WO 1252878 86061407 NC_060929.114030499..13982048.. 14499245 TRIO NC_060940.1 14354042MKL22Attorney Docket No.5470.980.WO 80249104 31860306 NC_060932.1122692348..36691430.. 122821670 HAS2 NC_060946.1 36MYH9898142

[0108] The foregoing examples are illustrative of the present invention, and are not to be construed as limiting thereof. Although the invention has been described in detail with reference to preferred embodiments, variations and modifications exist within the scope and spirit of the invention as described and defined in the following claims.

Claims

Attorney Docket No.5470.980.WO WHAT IS CLAIMED IS:

1. A method of classifying cancer in a subject in need thereof, comprising: extracting high molecular weight DNA from a biological sample from the subject; preparing a library from the extracted high molecular weight DNA for long-read (e.g., nanopore) sequencing; performing long-read whole genome sequencing of the library to generate sequencing data; identifying the presence or absence of a mutation or copy number variation in a panel of targets in the sequencing data; and classifying the cancer based on the presence or absence of the panel of targets.

2. The method of claim 1, wherein the biological sample is peripheral blood, bone marrow, or a mononuclear cell fraction thereof.

3. The method of claim 1 or claim 2, wherein the cancer is acute leukemia.

4. The method of claim 3, wherein the acute leukemia is pediatric acute leukemia.

5. The method of claim 1, wherein the biological sample is a resected tissue or tissue biopsy.

6. The method of claim 5, wherein the cancer is a solid tumor.

7. The method of claim 6, wherein the solid tumor is a pediatric sarcoma.

8. The method of any one of claims 1-7, wherein the biological sample is fresh or cryopreserved.

9. The method of any one of claims 1-8, wherein the long-read whole genome sequencing is performed using adaptive sampling.Attorney Docket No.5470.980.WO 10. The method of any one of claims 1-9, wherein the sequencing data is analyzed continuously in real time.

11. The method of any one of claims 1-10, wherein the method further comprises shearing the extracted high molecular weight DNA.

12. The method of any one of claims 1-11, wherein the panel of targets comprises genes with a known mutation associated with the cancer, optionally wherein the known mutation is a translocation, fusion, single nucleotide variation, and / or indel.

13. The method of any one of claims 1-12, wherein the cancer is acute leukemia and the panel of targets includes at least 50, 100, or 150 targets from any one of Tables 3, 4, 5, or 11.

14. The method of any one of claims 1-12, wherein the cancer is a solid tumor and the panel of targets includes at least 50, 100, or 150 targets from Table 12.

15. The method of any one of claims 1-14, wherein the method is completed in less than 24 hours, e.g., less than 12 hours.

16. The method of any one of claims 1-15, further comprising treating the cancer based on the classification.