Panels and methods for treatment of diffuse large b-cell lymphoma
A molecular classifier and targeted sequencing assay for DLBCL accurately identifies subtypes for personalized treatment, improving response rates by tailoring therapies like ibrutinib, lenalidomide, or decitabine based on subclass, addressing the limitations of current treatments.
Patent Information
- Application Number
- PCT/US2025/023297
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-13
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-09
AI Technical Summary
Current treatments for diffuse large B-cell lymphoma (DLBCL) have limited efficacy, with approximately 40% of patients not responding to first-line R-CHOP immunochemotherapy, highlighting the need for improved methods of characterization and targeted treatment strategies.
A molecular classifier and targeted sequencing assay are used to characterize variants in DLBCL samples, assigning them to specific subclasses (Cl, C2, C3, C4, or C5) based on mutations, structural variants, and somatic copy number alterations, followed by tailored treatments such as ibrutinib, lenalidomide, decitabine, tucidinostat, or ibrutinib, depending on the subclass.
This approach enables personalized treatment strategies that improve response rates by accurately classifying DLBCL subtypes and administering targeted therapies, enhancing treatment efficacy.
Smart Images

Figure US2025023297_09102025_PF_FP_ABST
Abstract
Description
PANELS AND METHODS FOR TREATMENT OF DIFFUSE LARGE B-CELLLYMPHOMACROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 575,541, filed April 5, 2024, and U.S. Provisional Application No. 63 / 694,603, filed September 13, 2024, which are incorporated by reference in their entireties herein.SEQUENCE LISTING
[0002] This application contains a Sequence Listing electronically submitted via EFS-Web to the United States Patent and Trademark Office as an XML file entitled “0680.003591W001. xml” having a size of 13,992,017 bytes and created on April 3, 2025. The information contained in the Sequence Listing is incorporated by reference herein.BACKGROUND OF THE INVENTION
[0003] Diffuse large B-cell lymphoma (DLBCL) is the most common lymphoid malignancy in adults and a clinically and biologically heterogeneous disease. Approximately 40% of patients fail to respond to the first line treatment, R-CHOP (rituximab, cyclophosphamide, doxorubicin, vincristine and prednisone) immunochemotherapy or develop recurrent disease. Thus, there remains a need for better methods for treatment of diffuse large B-cell lymphoma.
[0004] SUMMARY OF THE DISCLOSURE
[0005] As described below, the present disclosure features a molecular classifier and a targeted sequencing assay and related compositions and methods for use in characterization and treatment of diffuse large B-cell lymphoma.
[0006] In an aspect, the present disclosure provides a method for characterizing a diffuse large B- cell lymphoma (DLBCL) in a subject. The method involves (a) characterizing variants in a biological sample from the subject, where one or more of the variants are 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; or PIM2 and one or more additional variants are 10q23.31, l ip, l lq, l lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB,ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1 Al, EP3OO, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423, by using a targeted sequencing panel to characterize classes of the variants in the sample based upon the characterization of alterations in the variants, where the alterations are a mutation, a structural variant (SV), or a somatic copy number alteration (SCNA); (b) assigning a classification-specific weighted value to each class of variant characterized, where each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant; (c) condensing the variant classification-specific weighted values into two or more metafeatures; and (d) using the metafeatures as input variables for a computational analysis to assign the DLBCL to one of DLBCL subclasses Cl, C2, C3, C4 or C5, thereby characterizing the DLBCL.
[0007] In another aspect, the present disclosure provides a method for treating a subject having a diffuse large B-cell lymphoma (DLBCL). The method involves: (a) characterizing variants in a biological sample from the subject, where one or more of the variants are: 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; or PIM2 and one or more additional variants are: 10q23.31, l ip, l lq, l lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1 Al, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A,TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423, by characterizing classes of the variants in the sample based upon the characterization of alterations in the variants, where the alterations are a mutation, a structural variant (SV), or a somatic copy number alteration (SCNA); (b) assigning a classification-specific weighted value to each class of variant characterized, where each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant; (c) condensing the variant classification-specific weighted values into two or more metafeatures; (d) assigning the DLBCL as belonging to one of DLBCL subclasses Cl, C2, C3, C4 or C5 using a computational analysis, where the metafeatures are used as input variables for the computational analysis; and (e) administering to the subject a treatment including: ibrutinib or lenalidomide, when the DLBCL is assigned to subclass Cl ; decitabine and / or doxorubicin, when the DLBCL is assigned to subclass C2; tucidinostat, when the DLBCL is assigned to subclass C3; or ibrutinib when the DLBCL is assigned to subclass C5, thereby treating the DLBCL in the subject.
[0008] In another aspect, the present disclosure provides a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing to characterize two or more classes of variants in a biological sample based upon the characterization of alterations in the variants. The alterations are a mutation, a structural variant (SV), or a somatic copy number alteration (SCNA). One or more of the variants are 19ql 3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; orPIM2, and one or more additional variants are: 10q23.31, l ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423.
[0009] In another aspect, the present disclosure provides a method involving: instructing, by at least one processor, at least one user computing device to render at least one diffuse large B-celllymphoma (DLBCL) classification interface, the at least one DLBCL classification interface including at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receiving, by the at least one processor, via the at least one GSM array input element, the at least one GSM array input file, where the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, where one or more of the variants are 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; or PIM2, and one or more additional variants are: 10q23.31, l ip, l lq, l lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423, where the classes of variants are characterized using a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing of the variant classes; generating, by the at least one processor, at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilizing, by the at least one processor, at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, where the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and instructing, by the at least one processor, the at least one user computing device to render at least one DLBCL classification results element in the at least one DLBCL classification interface, where the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
[0010] In another aspect, the present disclosure provides a computing system including: at least one processor configured to execute computer-executable instructions which, upon execution,cause the system to perform steps to: instruct at least one user computing device to render at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface including at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receive, via the at least one GSM array input element, the at least one GSM array input file, where the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, where the variants include one or more of 19ql3.32; 5q; 6q; 9q; BCL11 A; ETS1; IRF4; METAP1D; or PIM2; and one or more additional variants are: 10q23.31, l ip, l lq, l lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423, where the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; generate at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilize at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, where the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and instruct the at least one user computing device to render at least one DLBCL classification results element in the at least one DLBCL classification interface, where the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
[0011] In another aspect, the present disclosure provides a method involving: receiving, by the at least one processor, at least one gene sample matrix (GSM) array input file associated with at leastone patient; where the at least one GSM array input file includes at least one GSM array encoding classes of variants associated with the at least one sample of the patient; where the classes of variants are characterized using a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing of the variant classes; where the variants include one or more of 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; and PIM2; and one or more additional variants are: 10q23.31, l ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423; generating, by the at least one processor, at least one biologically-relevant metafeature based at least in part on aggregating at least one grouping of the classes of variants, the at least one grouping comprising classes of variants grouped according to biological relevance mapping, co-occurrence, and / or cluster frequency; generating, by the at least one processor, at least one gene sample feature vector encoding at least one biologically-relevant metafeature; inputting, by the at least one processor, the at least one gene sample feature vector into at least one DLBCL classification machine learning model to output at least one cluster identification confidence value indicative of a probability of the at least one sample belonging to at least one cluster identification categorizing the at least one patient; where the at least one cluster identification characterizes DLBCL burden on the at least one patient; where at least one trained classification layer is configured, based at least in part on training using a cohort of training samples, the cohort of training samples including known biologically-relevant meta-features and known cluster identifications, to model at least one probability of correlation between biologically- relevant meta-features and the at least one cluster identification; determining, by the at least one processor, the at least one cluster identification based at least in part on the at least one cluster identification confidence value; and outputting, by the at least one processor, for display by at leastone display device via a DLBCL classification user interface, at least one DLBCL classification results element including a representation of the at least one cluster identification categorizing the at least one patient.
[0012] In any of the described aspects or embodiments thereof, the method further involves characterizing the DLBCL as high-risk if the DLBCL is assigned to subclass C5 or C3.
[0013] In any of the described aspects or embodiments thereof, the treatment further includes rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP). In any of the described aspects or embodiments thereof, the treatment further includes rituximab, cylcophosphamide, doxorubicin, prednisone, and polatuzumab vedotin (pola-R-CHP).
[0014] In any of the described aspects or embodiments thereof, the method further involves: i) when the DLBCL is assigned to class Cl or C5, administering to the subject a treatment including: an agent, where the agent is a NOTCH inhibitor, a BCL6 inhibitor and an activator of immune evasion, optionally an oligonucleotide inhibitor of NOTCH and / or BCL6, or a BCR / TLR signaling inhibitor and a BCL2 inhibitor, optionally oblimersen, ABT-263, Venetoclax (ABT-199), an antibody or oligonucleotide inhibitor of BCR / TLR signaling and / or an oligonucleotide inhibitor of BCL2; or (ii) when the DLBCL is assigned to DLBCL subclass C3 or C4 class, administering to the subject a treatment including: an agent, where the agent is a BCL2 inhibitor, a PI3K inhibitor and an epigenetic modifier, optionally oblimersen, ABT-263, Venetoclax (ABT-199), wortmannin, LY294002, an E2H2 inhibitor (optionally 3-deazaneplanocin A (DZNep), EPZ005687, EH, GSK126, and / or UNC1999), a CREBBP inhibitor, an oligonucleotide inhibitor of BCL2, an oligonucleotide inhibitor of PI3K and / or an oligonucleotide inhibitor of an epigenetic modifier; or a JAK / STAT inhibitor and a BRAF / MEK1 inhibitor, optionally ruxolitinib, Vemurafenib, Cobimetinib, an oligonucleotide inhibitor of JAK / STAT and / or an oligonucleotide inhibitor of BRAF / MEK1; or (iii) when the DLBCL is assigned to DLBCL subclass C2, administering to the subject a treatment including a CDK inhibitor, thereby treating the subject having a DLBCL.
[0015] In any of the described aspects or embodiments thereof, (a) further includes characterizing 10, 25, 50, 75, 100 or more variants, where the variants are: 10q23.31, l ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6,EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, or ZNF423.
[0016] In any of the described aspects or embodiments thereof, the agent includes rituximab, cyclophosphamide adriamycin, vincristine, prednisone, doxorubicin hydrochloride, or vincristine sulfate. In any of the described aspects or embodiments thereof, the agent includes rituximab, cyclophosphamide, doxorubicin hydrochloride, vincristine sulfate, and prednisone (R-CHOP).
[0017] In any of the described aspects or embodiments thereof, less than 25 metafeatures are used as the input variables. In any of the described aspects or embodiments thereof, 21 metafeatures or less are used as the input variables. In any of the described aspects or embodiments thereof, the metafeatures are selected from the group consisting of Ml, M2, M3, M4, M5, M6, M7, M8, M9, MIO, Mi l, M12, M13, M14, M15, M16, M17, M18, M19, M20, and M21.
[0018] In any of the described aspects or embodiments thereof, at least one class of variant corresponding to each metafeature is characterized. In any of the described aspects or embodiments thereof, the metafeatures include or are the metafeatures of Table 2.
[0019] In any of the described aspects or embodiments thereof, condensing the variant classification-specific weighted values involves summing the values corresponding to each metafeature. In any of the described aspects or embodiments thereof, the weighted values are proportional to a degree of variation from a reference sequence.
[0020] In any of the described aspects or embodiments thereof, in step (a), characterization involves: (i) determining for the mutation classes of variants whether there is a mutation and, if there is a mutation, whether the mutation is a silent mutation or a non-synonymous mutation; (ii) determining for the SCNA mutation class of variants whether there is an SCNA and, if there is an SCNA, whether the SCNA is a low level copy number alteration or a high level copy number alteration; and / or (iii) determining for the SV mutation class of variants whether or not an SV is present.
[0021] In any of the described aspects or embodiments thereof, the variant-specific weighted values are condensed into metafeatures as indicated in Table 2.
[0022] In any of the described aspects or embodiments thereof, the classes of variants belong to two or more of metafeatures selected from the metafeatures provided in Table 2.
[0023] In any of the described aspects or embodiments thereof, the oligonucleotides are suitable for use in targeted sequencing to characterize at least one variant class corresponding to each of the metafeatures provided in Table 2.
[0024] In any of the described aspects or embodiments thereof, the oligonucleotides are suitable for use in targeted sequencing to characterize all of the variant classes listed in Table 2.
[0025] In any of the described aspects or embodiments thereof, the targeted sequencing panel further includes oligonucleotide sequences suitable for use in targeted sequencing to measure microsatellite instability, to measure tumor mutational burden, and / or to detect Epstein Barr virus.
[0026] In any of the described aspects or embodiments thereof, the targeted sequencing panel includes polynucleotides sharing at least 85% sequence identity over a span of at least 80 nucleotides to at least one sequence listed in SEQ ID NOs: 1-9244 and Table B targeting the classes of variants.
[0027] In any of the described aspects or embodiments thereof, the method further involves: receiving, by the at least one processor, at least one DLBCL classification request from the at least one user computing device associated with the at least one patient, where the at least one DLBCL classification request includes at least one electronic request over a network; and generating, by the at least one processor, at least one rendering instruction in response to the at least one DLBCL classification request, where the at least one rendering instruction is configured to instruct the at least one user computing device to render the at least one DLBCL classification interface.
[0028] In any of the described aspects or embodiments thereof, the at least one trained classification layer includes an artificial neural network having learned weights for each of a plurality of neural network nodes.
[0029] In any of the described aspects or embodiments thereof, the method further involves: utilizing, by the at least one processor, at least one dimensionality reduction model to create a two- dimensional representation of the at least one metafeature; and generating, by the at least one processor, the at least one DLBCL classification results element including a two-dimensional visualization of the two-dimensional representation of the at least one metafeature, where the two- dimensional visualization includes at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0030] In any of the described aspects or embodiments thereof, the at least one dimensionality reduction model involves uniform manifold approximation and projection (UMAP).
[0031] In any of the described aspects or embodiments thereof, the at least one DLBCL classification results element includes a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0032] In any of the described aspects or embodiments thereof, the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; where the listing includes the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
[0033] In any of the described aspects or embodiments thereof, the at least one DLBCL classification results element includes a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, and ii) the at least one cluster identification associated with the at least one metafeature.
[0034] In any of the described aspects or embodiments thereof, the method further involves: generating, by the at least one processor, at least one treatment selection based at least in part on the at least one cluster identification; where the at least one DLBCL classification results element represents that at least one treatment selection.
[0035] In any of the described aspects or embodiments thereof, the at least one processor is remote from the at least one user computing device.
[0036] In any of the described aspects or embodiments thereof, the classes of variants are characterized based upon the characterization of alterations in the variants, wherein the alterations are a mutation, a structural variant (SV), or a somatic copy number alteration (SCNA).
[0037] In any of the described aspects or embodiments thereof, the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, wherein each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
[0038] In any of the described aspects or embodiments thereof, the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: receive at least one DLBCL classification request from the at least one user computing device associated with the at least one patient, where the at least one DLBCL classification request involves at least one electronic request over a network; and generate at least one rendering instruction in response to the at least one DLBCL classification request, where the at least one rendering instruction is configured to instruct the at least one user computing device to render the at least one DLBCL classification interface.
[0039] In any of the described aspects or embodiments thereof, the at least one trained classification layer includes an artificial neural network having learned weights for each of a plurality of neural network nodes.
[0040] In any of the described aspects or embodiments thereof, the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: utilize at least one dimensionality reduction model to create a two- dimensional representation of the at least one metafeature; and generate the at least one DLBCL classification results element including a two-dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization includes at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0041] In any of the described aspects or embodiments thereof, the at least one dimensionality reduction model includes a uniform manifold approximation and projection (UMAP).
[0042] In any of the described aspects or embodiments thereof, the at least one DLBCL classification results element includes a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0043] In any of the described aspects or embodiments thereof, the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; where the listing includes the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
[0044] In any of the described aspects or embodiments thereof, the at least one DLBCL classification results element includes a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, and ii) the at least one cluster identification associated with the at least one metafeature.
[0045] In any of the described aspects or embodiments thereof, the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: generate at least one treatment selection based at least in part on the at least one cluster identification; where the at least one DLBCL classification results element represents the at least one treatment selection.
[0046] In any of the described aspects or embodiments thereof, the at least one processor is remote from the at least one user computing device.
[0047] In any of the described aspects or embodiments thereof, the classes of variants are characterized based upon the characterization of alterations in the variants, where the alterations are a mutation, a structural variant (SV), or a somatic copy number alteration (SCNA).
[0048] In any of the described aspects or embodiments thereof, the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, where each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
[0049] In some aspects, the techniques described herein relate to a method including: causing, by a user computing device, to render, on a display device, at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface including at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receiving, by the user computing device, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein one or more of the variants are selected from the group consisting of 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; and PIM2, and one or more additional variants are selected from the group consisting of: 10q23.31, l ip, 11 q, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPL10, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing of the variant classes; generating, by the user computing device, at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilizing, by theuser computing device, at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and causing, by the user computing device, to render, on the display device, at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
[0050] In some aspects, the techniques described herein relate to a method, further including: receiving, by the user computing device, at least one DLBCL classification request associated with the at least one patient, wherein the at least one DLBCL classification request includes at least one electronic request over a network; and generating, by the user computing device, at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction is configured to instruct the display device to render the at least one DLBCL classification interface.
[0051] In some aspects, the techniques described herein relate to a method, wherein the at least one trained classification layer includes an artificial neural network having learned weights for each of a plurality of neural network nodes.
[0052] In some aspects, the techniques described herein relate to a method, further including: utilizing, by the user computing device, at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generating, by the user computing device, the at least one DLBCL classification results element including a two- dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization includes at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0053] Any of the described aspects or embodiments thereof, further including: generating, by the user computing device, at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents that at least one treatment selection.
[0054] Any of the described aspects or embodiments thereof, wherein the display device is local to the user computing device.
[0055] In some aspects, the techniques described herein relate to a user computing device including: a processor configured to execute computer-executable instructions which, uponexecution, cause the user computing device to perform steps to: render at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface including at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receive, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein the variants include one or more of 19ql3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; and PIM2; and one or more additional variants are selected from the group consisting of 10q23.31, l ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing of the variant classes; generate at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilize at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and render at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
[0056] Any of the described aspects or embodiments thereof, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause theuser computing device to perform steps to: receive at least one DLBCL classification request from the at least one user computing device associated with the at least one patient, wherein the at least one DLBCL classification request includes at least one electronic request over a network; and generate at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction is configured to instruct the user computing device to render the at least one DLBCL classification interface.
[0057] Any of the described aspects or embodiments thereof, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause the user computing device to perform steps to: utilize at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generate the at least one DLBCL classification results element including a two-dimensional visualization of the two- dimensional representation of the at least one metafeature, wherein the two-dimensional visualization includes at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
[0058] Any of the described aspects or embodiments thereof, wherein the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; wherein the listing includes the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
[0059] Any of the described aspects or embodiments thereof, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause the user computing device to perform steps to: generate at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents the at least one treatment selection.
[0060] In some aspects, the techniques described herein relate to a method including: receiving, by the user computing device, at least one gene sample matrix (GSM) array input file associated with at least one patient; wherein the at least one GSM array input file includes at least one GSM array encoding classes of variants associated with the at least one sample of the patient; wherein the classes of variants are characterized using a targeted sequencing panel including oligonucleotides suitable for use in targeted sequencing of the variant classes; wherein the variants include one or more of 19ql3.32; 5q; 6q; 9q; BCLl lA; ETS1; IRF4; METAP ID; and PIM2; and one or more additional variants are selected from the group consisting of: 10q23.31, l ip, l lq, l lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1,lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP3OO, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423; generating, by the user computing device, at least one biologically- relevant metafeature based at least in part on aggregating at least one grouping of the classes of variants, the at least one grouping including classes of variants grouped according to biological relevance mapping, co-occurrence, and / or cluster frequency; generating, by the user computing device, at least one gene sample feature vector encoding at least one biologically-relevant metafeature; inputting, by the user computing device, the at least one gene sample feature vector into at least one DLBCL classification machine learning model to output at least one cluster identification confidence value indicative of a probability of the at least one sample belonging to at least one cluster identification categorizing the at least one patient; wherein the at least one cluster identification characterizes DLBCL burden on the at least one patient; wherein at least one trained classification layer is configured, based at least in part on training using a cohort of training samples, the cohort of training samples including known biologically-relevant meta-features and known cluster identifications, to model at least one probability of correlation between biologically- relevant meta-features and the at least one cluster identification; determining, by the user computing device, the at least one cluster identification based at least in part on the at least one cluster identification confidence value; and outputting, by the user computing device, for display by at least one display device via a DLBCL classification user interface, at least one DLBCL classification results element including a representation of the at least one cluster identification categorizing the at least one patient.
[0061] The present disclosure provides a molecular classifier and a targeted sequencing assay for use in characterization and treatment of diffuse large B-cell lymphoma. Compositions and articles defined by the disclosure were isolated or otherwise manufactured in connection with the examplesprovided below. Other features and advantages of the disclosure will be apparent from the detailed description, and from the claims.Definitions
[0062] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this disclosure belongs. The following references provide one of skill with a general definition of many of the terms used in present disclosure: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless specified otherwise.
[0063] By “variant” is meant a genomic locus associated with an alteration associated with diffuse large B-cell lymphoma. The alteration can be a mutation, a somatic copy number alteration (SCNA), or a structural variant (SV; including translocations). The alteration can be a focal alteration or an arm variation. The somatic copy number alteration (SCNA) can be a copy number gain or a copy number loss. A focal alteration is an alteration affecting a small region of a chromosome; for example, the small region can be about or less than about 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 1 kbp, 5 kbp, 10 kbp, 25 kbp, 50 kbp, 100 kbp, 250 kbp, 500 kbp, or 1 Mbp in size. An arm alteration is an alteration affecting a large region of a chromosome (e.g., an arm of a chromosome). An arm alteration can affect a large region containing about or at least about 500 bp, 1 kbp, 5 kbp, 10 kbp, 25 kbp, 50 kbp, 100 kbp, 250 kbp, 500 kbp, 1 Mbp, 2 Mbp, 3 Mbp, 4 Mbp, or 5 Mbp. The alteration can be an alteration affecting a gene (i.e., a “gene” alteration). A variant can be described as a “target” for sequencing and the variant can be assigned to a “type” where the type identifies the type of genomic locus associated with the variant; for example, a gene, a large chromosomal region (“arm”), a small chromosomal region (“focal”), or microsatellite instability (“MSI” or a microsatellite).
[0064] By “variant classification” or “class of variant” is meant a particular alteration associated with a variant.
[0065] By “metafeature” is meant one variable of a reduced set of variables determined using dimensionality reduction. In embodiments, a metafeature is a weighted sum (i.e., “metafeature value”) of variant classes observed in a sample (see, e.g., Tables 1 and 2). Table 1 below provides representative definitions of metafeatures. A variant class can be assigned a higher weight proportional to the magnitude of an associated alteration; for example, a higher alteration in copynumber can be assigned a higher weight than a lower alteration in copy number (see, e.g., the weighting scheme provided in Table 1).Table 1: Weighing scale used to calculate metafeature values based on variant measurements.Type Value MeaningMutations 0 no mutation1 silent mutation2 non-synonymous mutationSCNA 0 no SCNA1 low level copy number alteration2 high level copy number alterationSV 0 no SV3 SV
[0066] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragments thereof. In some embodiments, an agent is a compound for the treatment of cancer (e.g., diffuse large B-cell lymphoma). In some embodiments, an agent is a chemotherapeutic agent. In some embodiments, the agent is a combination therapy comprising anti-CD20 antibody rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R- CHOP) that is commonly used as a primary treatment for DLBCL. In some embodiments, when a subject is characterized as having DLBCL subtype Cl, the agent ibrutinib or lenalidomide is administered to the subject. In some embodiments, when a subject is characterized as having DLBCL subtype C2, agents decitabine and / or doxorubicin are administered to the subject. In some embodiments, when a subject is characterized as having DLBCL subtype C3, the agent tucidinostat is administered to the subject. In some embodiments, when a subject is characterized as having DLBCL subtype C5, the agent ibrutinib is administered to the subject.
[0067] This additional agent “X” may be administered alongside any therapeutic regimen known in the art for treatment of DLBCL. In some embodiments, the agent is administered with R-CHOP (e.g., as R-CHOP-X). In some embodiments, polatuzumab vedotin in combination with rituximab, cyclophosphamide, doxorubicin, and prednisone (e.g., pola-R-CHP) may be administered alongside the agent (e.g., pola-R-CHP-X).
[0068] As used herein, the term “algorithm” refers to any formula, model, mathematical equation, algorithmic, analytical, or programmed process, or statistical technique or classification analysis that takes one or more inputs or parameters, whether continuous or categorical, and calculates an output value, index, index value or score. Examples of algorithms include but are not limited toratios, sums, regression operators such as exponents or coefficients, biomarker value transformations and normalizations (including, without limitation, normalization schemes that are based on clinical parameters such as age, gender, ethnicity, etc.), rules and guidelines, statistical classification models, statistical weights, and neural networks trained on populations or datasets. In some embodiments, the algorithm is a naive bayes or random forest. In some embodiments, the algorithm is a neural network.
[0069] By “ameliorate” is meant decrease, suppress, attenuate, diminish, arrest, or stabilize the development or progression of a disease.
[0070] By " alteration" is meant a change (increase or decrease) in the expression levels or activity of a gene or polypeptide as detected by standard art known methods such as those described herein. As used herein, an alteration includes a 10% change in expression levels, a 25% change, a 40% change, or a 50% or greater change in expression levels.
[0071] By "analog" is meant a molecule that is not identical, but has analogous functional or structural features. For example, a polypeptide analog retains the biological activity of a corresponding naturally-occurring polypeptide, while having certain biochemical modifications that enhance the analog's function relative to a naturally occurring polypeptide. Such biochemical modifications could increase the analog's protease resistance, membrane permeability, or half-life, without altering, for example, ligand binding. An analog may include an unnatural amino acid.
[0072] “Biological sample” as used herein refers to a sample obtained from a subject. Biological samples include samples of biological tissue or fluid origin, obtained, reached, or collected in vivo or in situ, that contains or is suspected of containing a polynucleotide. A biological sample also includes samples from a region of a biological subject containing precancerous or cancer cells or tissues. Such samples can be, but are not limited to, organs, tissues, fractions and cells isolated from mammals including, humans, mice, and rats. Biological samples also may include sections of the biological sample including tissues, for example, frozen sections taken for histologic purposes. In embodiments, a sample is a blood, plasma, or serum sample comprising circulating tumor DNA. In some embodiments, a biological sample is a sample from a subject having, or suspected of having a neoplasia (e.g., diffuse large B-cell lymphoma).
[0073] By “circulating tumor DNA (ctDNA)” is meant cell-free DNA found in the bloodstream of a subject that is derived from neoplasm cells. In embodiments, the neoplasm is a cancer (e.g., diffuse large B-cell lymphoma).
[0074] In this disclosure, "comprises," "comprising," "containing" and "having" and the like can have the meaning ascribed to them in U.S. Patent law and can mean " includes," "including," and the like; "consisting essentially of' or "consists essentially" likewise has the meaning ascribed inU.S. Patent law and the term is open-ended, allowing for the presence of more than that which is recited so long as basic or novel characteristics of that which is recited is not changed by the presence of more than that which is recited, but excludes prior art embodiments. Any embodiments specified as “comprising” a particular component(s) or element(s) are also contemplated as “consisting of’ or “consisting essentially of’ the particular component(s) or element(s) in some embodiments.
[0075] By “control” or “reference” is meant a standard of comparison. In one aspect, as used herein, “changed as compared to a control” sample or subject is understood as having a level that is statistically different than a sample from a normal, untreated, or control sample. Control samples include, for example, cells in culture, one or more laboratory test animals, one or more human subjects, or biological samples from the same (e.g., cfDNA). Methods to select and test control samples are within the ability of those in the art. Determination of statistical significance is within the ability of those skilled in the art, e.g., the number of standard deviations from the mean that constitute a positive result. In embodiments, a reference is a subject or a sample from a subject who does not have a cancer or a subject prior to a change in a treatment or administration of a drug or treatment. In embodiments, the reference is a matched normal sample or a panel of normals (PoN), where in some instances the matched normal sample is a sample from a healthy subject and / or a subject who does not have a cancer (e.g., a subject prior to being diagnosed with a diffuse large B-cell lymphoma). In some embodiments, a reference is a classification model for cancer (e.g., diffuse large B-cell lymphoma) other than a classification model of the present disclosure.
[0076] By “consist essentially” it is meant that the ingredients include only the listed components along with the normal impurities present in commercial materials and with any other additives present at levels which do not affect the operation of the disclosure, for instance at levels less than 5% by weight or less than 1% or even 0.5% by weight.
[0077] By “copy number variation (CNV),” “copy number alteration (CNA),” or “somatic copy number alteration (SCNA)” is meant an alteration that results in a gain or loss in copies of a section(s) of a genome. Non-limiting examples of SCNAs include duplications and deletions. SCNAs may be very short (focal) or almost exactly the length of a chromosome arm or whole chromosome (arm-level). Accordingly, non-limiting examples of SCNAs also include focal SCNAs and arm-level SCNAs.
[0078] As used herein, the term “coverage” refers to the number of sequence reads that align to a specific locus in a reference sequence. In embodiments, the reference sequence is a reference genome. For example, with regard to the terminal base of the following reference sequence, because there is only one sample base aligned at this locus (the bold cytosine in Read 2), thereis lx coverage of the reference sequence at this locus. At the 5’ end, there is 3x coverage of the reference sequence at the 5’ terminus guanine.Reference Sequence: 5’ GGGAAGGGCGATC 3’Read 1 GGGAAGGGCGATRead 2 GGGAAGGGCGATCRead 3 GGGAAGGGCG
[0079] When a genome is sequenced, there will be a large number of nucleotides sequenced. If an individual genome is sequenced only once, there will be a significant number of sequencing errors. To increase the sequencing accuracy, an individual genome will need to be sequenced a large number of times. The average coverage for a whole genome can be calculated from the length of the original genome (G), the number of reads (N), and the average read length (L) as N x L / G. In another example, a hypothetical genome with 2,000 base pairs reconstructed from 8 reads with an average length of 500 nucleotides will have 2* redundancy. This parameter also enables one to estimate other quantities, such as the percentage of the genome covered by reads (sometimes also called breadth of coverage). At a coverage of O.lx, only 10% of a reference sequence is covered by sequence reads. In embodiments, a sample polynucleotide is sequenced to a coverage of about, at least about, and / or no more than about le-8x, le-7x, le-6x, le-5x, le-4x, le-3x, le-2x, 0.05x, O. lx, 0.2x, 0.3x, 0.4x, 0.5x, lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, lOOx, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, lOOOx, 5000x, lOOOOx, 15000x, 20000x, 25000x, 30000x, 50000x, lOOOOOx, or more.
[0080] By “ultra-low coverage” is meant a coverage of less than at least 5x. In some instances, ultra-low coverage is a coverage of less than 0.5x, 0.2x, or O. lx.
[0081] “Detect” refers to identifying the presence, absence, or amount of the analyte to be detected.
[0082] By "detectable label" is meant a composition that when linked to a molecule of interest renders the latter detectable, via spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioactive isotopes, magnetic beads, metallic beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (for example, as commonly used in an ELISA), biotin, digoxigenin, or haptens.
[0083] By “disease” is meant any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ. Examples of diseases include cancer (e.g., Hodgkin’s lymphoma, primary mediastinal B-cell lymphoma), and related diseases or disorders. In an embodiment, the disease is diffuse large B-cell lymphoma (DLBCL).
[0084] By "effective amount" is meant the amount of an agent required to provide a therapeutic benefit in the treatment of a condition or to delay or minimize one or more symptoms associated with the condition. In some instances, the effective amount ameliorates the symptoms of a disease relative to an untreated patient. A therapeutically effective amount of an agent can mean an amount of therapeutic agent, alone or in combination with other therapies, which provides a therapeutic benefit in the treatment of the condition. The term “effective amount” can encompass an amount that improves overall therapy, reduces or avoids symptoms, signs, or causes of the condition, and / or enhances the therapeutic efficacy of another therapeutic agent. The effective amount of active compound(s) used to practice the present disclosure for therapeutic treatment of a disease varies depending upon the manner of administration, the age, body weight, and general health of the subject. Ultimately, the attending physician or veterinarian will decide the appropriate amount and dosage regimen. Such amount is referred to as an "effective" amount.
[0085] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. This portion contains, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of the reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0086] "Hybridization" means hydrogen bonding, which may be Watson-Crick, Hoogsteen or reversed Hoogsteen hydrogen bonding, between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.
[0087] By “immunotherapy” is meant a treatment that involves supplementing or stimulating the immune system. Non-limiting examples of immunotherapies include treatments involving administration of biologies, such as immune checkpoint blockades, and / or CAR T cells.
[0088] By “immune checkpoint blockade” is meant an agent that functions as an inhibitor of a polynucleotide and / or pathway that functions in inhibiting or stimulating an immune response. In embodiments, the agent is an antibody.
[0089] By “ increases” is meant a positive alteration of at least 10%, 25%, 50%, 75%, or 100%.
[0090] The terms "isolated," "purified," or "biologically pure" refer to material that is free to varying degrees from components which normally accompany it as found in its native state. "Isolate" denotes a degree of separation from an original source or surroundings. "Purify" denotes a degree of separation that is higher than isolation. A "purified" or "biologically pure" protein is sufficiently free of other materials such that any impurities do not materially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of this disclosure is purified if it is substantially free of cellular material, viral material, or culturemedium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, for example, polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can denote that a nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. For a protein that can be subjected to modifications, for example, phosphorylation or glycosylation, different modifications may give rise to different isolated proteins, which can be separately purified.
[0091] By " isolated polynucleotide" is meant a nucleic acid (e.g., a DNA) that is free of the genes which, in the naturally-occurring genome of the organism from which the nucleic acid molecule of the disclosure is derived, flank the gene. The term therefore includes, for example, a recombinant DNA that is incorporated into a vector; into an autonomously replicating plasmid or virus; or into the genomic DNA of a prokaryote or eukaryote; or that exists as a separate molecule (for example, a cDNA or a genomic or cDNA fragment produced by PCR or restriction endonuclease digestion) independent of other sequences. In addition, the term includes an RNA molecule that is transcribed from a DNA molecule, as well as a recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.
[0092] By an “isolated polypeptide” is meant a polypeptide of the disclosure that has been separated from components that naturally accompany it. Typically, the polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally-occurring organic molecules with which it is naturally associated. In some embodiments, the preparation is at least 75%, at least 90%, or at least 99%, by weight, a polypeptide of the disclosure. An isolated polypeptide of the disclosure may be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide; or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.
[0093] By “marker” is meant a protein, polynucleotide, or other analyte having an alteration in sequence, copy number, structure, expression level or activity that is associated with a disease or disorder.
[0094] By “mutation” is meant an alteration to a polynucleotide sequence. Non-limiting examples of non-synonymous mutations include single-nucleotide polymorphisms (SNPs), singlenucleotide variations (SNVs), and insertions or deletions (indel mutations). In embodiments, a non-synonymous mutation corresponds to a genomic region about or less than about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 10 bp, 50 bp, or 100 bp in size.
[0095] As used herein, the term “next-generation sequencing (NGS)” refers to a variety of high- throughput sequencing technologies that parallelize the sequencing process, producing thousands or millions of sequence reads at once. NGS parallelization of sequencing reactions can generate hundreds of megabases to gigabases of nucleotide sequence reads in a single instrument run. Unlike conventional sequencing techniques, such as Sanger sequencing, which typically report the average genotype of an aggregate collection of molecules, NGS technologies typically digitally tabulate the sequence of numerous individual DNA fragments (sequence reads discussed in detail below), such that low frequency variants (e.g., variants present at less than about 10%, 5% or 1% frequency in a heterogeneous population of nucleic acid molecules) can be detected. The term “massively parallel” can also be used to refer to the simultaneous generation of sequence information from many different template molecules by NGS. NGS sequencing platforms include, but are not limited to, the following: Massively Parallel Signature Sequencing (Lynx Therapeutics); 454 pyro-sequencing (454 Life Sciences / Roche Diagnostics); solid-phase, reversible dye-terminator sequencing (Solexa / Illumina); SOLiD technology (Applied Biosystems); Ion semiconductor sequencing (ion Torrent); and DNA nanoball sequencing (Complete Genomics). Descriptions of certain NGS platforms can be found in the following: Shendure, et al., “Next-generation DNA sequencing,” Nature, 2008, vol. 26, No. 10, 135-1 145; Mardis, “The impact of next-generation sequencing technology on genetics,” Trends in Genetics, 2007, vol. 24, No. 3, pp. 133-141 ; Su, et al., “Next-generation sequencing and its applications in molecular diagnostics” Expert Rev Mol Diagn, 2011, 11 (3):333-43; and Zhang et al., “The impact of next-generation sequencing on genomics,” J Genet Genomics, 201, 38(3): 95-109.
[0096] As used herein, “obtaining” as in “obtaining an agent” includes synthesizing, purchasing, or otherwise acquiring the agent.
[0097] By “polypeptide” or “amino acid sequence” is meant any chain of amino acids, regardless of length or post-translational modification. In various embodiments, the post-translational modification is glycosylation or phosphorylation. In various embodiments, conservative amino acid substitutions may be made to a polypeptide to provide functionally equivalent variants, or homologs of the polypeptide. In some aspects, the disclosure embraces sequence alterations that result in conservative amino acid substitutions. In some embodiments, a “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size characteristics of the protein in which the conservative amino acid substitution is made. Variants can be prepared according to methods for altering polypeptide sequences known to one of ordinary skill in the art such as are found in references that compile such methods, e.g. Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Second Edition, Cold Spring Harbor LaboratoryPress, Cold Spring Harbor, N. Y., 1989, or Current Protocols in Molecular Biology, F. M. Ausubel, et al., eds., John Wiley & Sons, Inc., New York. Non-limiting examples of conservative substitutions of amino acids include substitutions made among amino acids within the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D. In various embodiments, conservative amino acid substitutions can be made to the amino acid sequence of the proteins and polypeptides disclosed herein.
[0098] By “probe set” or “bait set” is meant a set of probes that hybridize to and characterize a target polynucleotide.
[0099] By “ reduces” is meant a negative alteration of at least 10%, 25%, 50%, 75%, or 100%.
[0100] By “ reference” is meant a standard or control condition. In embodiments, the reference is a reference sequence. A non-limiting example of a reference sequence is a genomic DNA sequence corresponding to a subject not having a DLBCL. As used herein, “changed as compared to a reference” sample or subject is understood as having a level that is statistically different than a sample from a normal, untreated, or reference sample. Reference samples include, for example, cells in culture, one or more laboratory test animals, or one or more human subjects. Methods to select and test reference samples are within the ability of those in the art. Determination of statistical significance is within the ability of those skilled in the art, e.g., the number of standard deviations from the mean that constitute a positive result.
[0101] A “reference genome” is a defined genome used as a basis for genome comparison or for alignment of sequencing reads thereto. A reference genome may be a subset of or the entirety of a specified genome; for example, a subset of a genome sequence, such as exome sequence, or the complete genome sequence.
[0102] Nucleic acid molecules useful in the methods of the disclosure include any nucleic acid molecule that encodes a polypeptide of the disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical with an endogenous nucleic acid sequence, but will typically exhibit substantial identity. Polynucleotides having “substantial identity” to an endogenous sequence are typically capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the disclosure include any nucleic acid molecule that encodes a polypeptide of the disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical with an endogenous nucleic acid sequence, but will typically exhibit substantial identity. Polynucleotides having “substantial identity” to an endogenous sequence are typically capable of hybridizing with at least one strand of a doublestranded nucleic acid molecule. By "hybridize" is meant pair to form a double-stranded molecule between complementary polynucleotide sequences (e.g., a gene described herein), or portionsthereof, under various conditions of stringency. (See, e.g., Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507).
[0103] For example, stringent salt concentration will ordinarily be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvent, e.g., formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, or at least about 50% formamide. Stringent temperature conditions will ordinarily include temperatures of at least about 30° C, at least about 37° C, or at least about 42° C. Varying additional parameters, such as hybridization time, the concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency are accomplished by combining these various conditions as needed. In an embodiment, hybridization will occur at 30° C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In an embodiment, hybridization will occur at 37° C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 pg / ml denatured salmon sperm DNA (ssDNA). In an embodiment, hybridization will occur at 42° C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 pg / ml ssDNA. Useful variations on these conditions will be readily apparent to those skilled in the art.
[0104] For most applications, washing steps that follow hybridization will also vary in stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As above, wash stringency can be increased by decreasing salt concentration or by increasing temperature. For example, stringent salt concentration for the wash steps will be less than about 30 mM NaCl and 3 mM trisodium citrate, or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash steps will ordinarily include a temperature of at least about 25° C, at least about 42° C, or at least about 68° C. In an embodiment, wash steps will occur at 25° C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In an embodiment, wash steps will occur at 42° C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In an embodiment, wash steps will occur at 68° C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations on these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196: 180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, AcademicPress, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0105] The phrase “pharmaceutically acceptable carrier” is recognized in the art and includes a pharmaceutically acceptable material, composition or vehicle, suitable for administering compounds of the present disclosure to a subject. The carriers include liquid or solid filler, diluent, excipient, solvent or encapsulating material, involved in carrying or transporting the subject agent from one organ, or portion of the body, to another organ, or portion of the body. Each carrier must be “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some non-limiting examples of materials which can serve as pharmaceutically acceptable carriers include the following: sugars, such as lactose, glucose and sucrose; starches, such as corn starch and potato starch; cellulose, and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; glycols, such as propylene glycol; polyols, such as glycerin, sorbitol, mannitol and polyethylene glycol; esters, such as ethyl oleate and ethyl laurate; agar; buffering agents, such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethyl alcohol; phosphate buffer solutions; and other non-toxic compatible substances employed in pharmaceutical formulations.
[0106] The term “salts” refers to the relatively non-toxic, inorganic and organic acid addition salts of compounds of the present disclosure. These salts can be prepared in situ during the final isolation and purification of compounds or by separately reacting a purified compound in its free base form with a suitable organic or inorganic acid and isolating the salt thus formed. Representative salts include the hydrobromide, hydrochloride, sulfate, bisulfate, nitrate, acetate, oxalate, valerate, oleate, palmitate, stearate, laurate, borate, benzoate, lactate, phosphate, tosylate, citrate, maleate, fumarate, succinate, tartrate, naphthylate mesylate, glucoheptonate, lactobionate and laurylsulphonate salts, and the like. Representative salts may further include cations based on the alkali and alkaline earth metals, such as sodium, lithium, potassium, calcium, magnesium, and the like, as well as non-toxic ammonium, tetramethylammonium, tetramethyl ammonium, methlyamine, dimethlyamine, trimethlyamine, triethlyamine, ethylamine, and the like. (See, for example, S. M. Barge et al., “Pharmaceutical Salts,” J. Pharm. Sci., 1977, 66: 1-19 which is incorporated herein by reference.).
[0107] By “structural variation (SV)” is meant a large alteration in the sequence of a genome. Non-limiting examples of structural variants include gene fusions, translocations, deletions,duplications, inversions, and translocations. In embodiments, a structural variation corresponds to a genomic region that is about or at least about 100 bp, 500 bp, 1 kb, 10 kb, 100 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb or 10 Mb in size.
[0108] By "substantially identical" is meant a polypeptide or nucleic acid molecule exhibiting at least 50% identity to a reference amino acid sequence (for example, any one of the amino acid sequences described herein) or nucleic acid sequence (for example, any one of the nucleic acid sequences described herein). In some embodiments, such a sequence is at least 60%, at least 80% or 85%, or at least 90%, 95% or even 99% identical at the amino acid level or nucleic acid to the sequence used for comparison.
[0109] Sequence identity is typically measured using sequence analysis software (for example, Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, a BLAST program may be used, with a probability score between e'3and e’100indicating a closely related sequence.
[0110] By "subject" is meant an animal. The animal can be a mammal. The mammal can be a human or non-human mammal, such as a bovine, equine, canine, ovine, rodent, or feline.
[0111] Ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0112] By “targeted sequencing” is meant a sequencing method where polynucleotide sequences of interest from a biological sample are selectively sequenced. In embodiments, targeted sequencing comprises contacting polynucleotides present in a biological sample with an oligonucleotide probe or panel of oligonucleotide probes. In embodiments, targeted sequencing involves enriching for polynucleotide sequences from a sample that hybridize to an oligonucleotide probe or panel of oligonucleotide probes. In various instances, targeted sequencing has the advantage of allowing for sequencing polynucleotide sequences of interest in a biological sample to a high sequencing coverage.
[0113] As used herein, the terms “treat,” treating,” “treatment,” and the like refer obtaining a desired pharmacologic and / or physiologic effect. The effect can involve reducing or ameliorating a disorder and / or symptoms associated therewith. The effect can be prophylactic in terms of completely or partially preventing a disease or symptom thereof and / or can be therapeutic in terms of a partial or complete cure for a disease and / or adverse effect attributable to the disease. “Treatment,” as used herein, covers any treatment of a disease or condition in a mammal, particularly in a human, and includes: (a) preventing the disease from occurring in a subject who can be predisposed to the disease but has not yet been diagnosed as having it; (b) inhibiting the disease, i.e., arresting its development; and (c) relieving the disease, i.e., causing regression of the disease. It will be appreciated that, although not precluded, treating a disorder or condition does not require that the disorder, condition or symptoms associated therewith be completely eliminated.
[0114] “ Tumor-derived DNA” means DNA that is derived from a cancer cell rather than from a healthy control cell. Tumor derived DNA often includes structural changes that are indicative of cancer.
[0115] Unless specifically stated or obvious from context, as used herein, the term "or" is understood to be inclusive. Unless specifically stated or obvious from context, as used herein, the terms "a", "an", and "the" are understood to be singular or plural.
[0116] Unless specifically stated or obvious from context, as used herein, the term “about” is understood as within a range of normal tolerance in the art, for example within 2 standard deviations of the mean. About can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise clear from context, all numerical values provided herein are modified by the term about.
[0117] The recitation of a listing of chemical groups in any definition of a variable herein includes definitions of that variable as any single group or combination of listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiments or portions thereof.
[0118] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0119] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0120] FIG. 1A-1C shows the Purity, Ploidy and Coverage Distribution of the two cohorts and SNP Filtering of Schmitz cohort. FIG. 1A includes histograms for purity (left panel), ploidy (middle panel) and mean target coverage (right panel) by cohort (Chapuy et al, blue; Schmitz, red). FIG. IB shows the distribution of SNVs visualized by normal and tumor variant allele fraction (VAF) in the 39 (before QC filtering) TCGA samples with available paired normal samples. Notably, -44% of all SNVs within the provided TCGA MAF (Schmitz et al., 2018) exhibit a normal Variant Allele Frequency (VAF) of > 0.40, indicating likely germline events (right top panel), however, these events contain only -1% driver (defined as events in genes with MutSig2CV q-value < 0.10) events (right panel). FIG. 1C shows the VAF of all mutations by purity are visualized in a color-coded heatmap as reported in Schmitz et al. (Schmitz et al., 2018) (left top panel) and after applying our tumor only pipeline (Schmitz et al., 2018) (right panel). Note the significant drop in the likely germline spill in (boxed triangle).
[0121] FIG. ID provides a cohort diagram. The dataset of this analysis that meet the inclusion criteria (purity >0.2; genetic data for mutations, SCNAs and SVs) is composed of 699 patients with newly diagnosed DLBCL from 3 cohorts (277 samples from Chapuy et al. (Chapuy et al., 2018), blue; 414 samples from Schmitz et al. (Schmitz et al., 2018), red; and 8 samples from TCGA (Schmitz et al., 2018), purple; top row). Patients treated with state-of-the-art R-CHOP -based therapies and available baseline clinical information (International Prognostic Index; IPI available yes / no) are indicated in the second and third row, respectively. Tissue preservation method of samples from which DNA was extracted is indicated in the third row (frozen [FRZN] or formalin- fixed paraffin embedded [FFPE] tissue). Patients with available patient-matched paired normal tissue are indicated in the fourth row, while patient assignment to train / test set is highlighted in the last row.
[0122] FIG. 2A-2F show genetic features and harmonization across cohorts before NMF clustering. FIGS. 2A-D are histograms showing the number of detected driver alterations in the two cohorts (Chapuy cohort, blue; Schmitz cohort, red) following harmonization of the pipelines. A, SVs; B, Mutational drivers; C, GISTIC-defined CN losses; D, GISTIC-defined CN gains. FIG 2E is a graph showing a marginal frequencies comparison of the NMF GSM that include 150 driver alterations previously reported (Chapuy et al., 2018) (minus SVs not captured in Schmitz et al.) across cohorts (Chapuy cohort, x-axis; Schmitz cohort, y-axis) after reprocessing of the Schmitz cohort and down sampling to compensate for COO skewing in Schmitz et al. cohort. Alterations are color coded (black, mutations; red, CN gains; blue, CN losses; SVs, green). Error bars reflect one standard deviation errors on the binomial distribution parameterized by the given frequency and counts for each driver. Those alterations that are significantly more frequent in one cohort(two-sided Fisher’ s Exact test with BH-correction, q<0.1) are labeled (*). FIG 2F provides an IGV screenshot of two representative samples from each cohort (top, Chapuy cohort; bottom, Schmitz cohort) indicating coverage differences due to the bait set used in Exon 1 of BTG2. These differences explain the significant more frequent BTG2 mutations found in the Schmitz cohort and visualized in panel E.
[0123] FIGS. 3A-3C provide the validation of C1-C5 DLBCL molecular substructure in an independent dataset and the combined datasets. Fig. 3A provides a comparison of the marginal frequencies of previously reported driver alterations in the Chapuy (x-axis) and Schmitz cohorts (y-axis) after reprocessing and down sampling the Schmitz dataset to compensate for COO skewing. Alterations are color coded (black, mutations; red, CN gains; blue, CN losses; SVs, green). Error bars reflect 1 -sigma errors on the binomial distribution parameterized by the given frequency and counts for each driver. Fig. 3B provides connectivity plots visualizing the distance matrices from the NMF consensus clustering for k=5 of the Chapuy, Schmitz and Combined cohorts. Cophenetic Rho values on top. Fig. 3C provides frequencies of genetic drivers in each cluster (C1-C5) of the combined series (y-axis) and the original Chapuy et al. series (x-axis), plotted as color-coded scatter plots of genetic driver alterations (black, mutations, red, CN gains; blue, CN losses; green, SVs). Pearson correlations were calculated for each plot. Top 10 most frequent alterations are labeled. Error bars reflect one standard deviation of the binomial distribution parameterized by the given frequency and counts for each driver.
[0124] FIG. 4A-4D show NMF consensus clustering of the combined cohort. A-C, Cophenetic index plot from the NMF consensus clustering for k=2 to k=8 (using *** iterations). Bootstrapping one standard deviation errors calculated using 200 iterations of resampling of the connectivity matrix, wherein each iteration rho is recalculated from the hierarchical clustering solution of the resampled set (A, Chapuy et al. set; B, Schmitz et al. set; C, Combined set). D, Visualization of co-clustering matrix from the NMF consensus clustering for k=2 to k=8 of the combined set. Cases (n) per cluster (k) are visualized below the connectivity heatmap.
[0125] Fig. 4E shows the similarity of genetic driver are plotted as color coded scatter plot (black, mutations, red, SCNA-gains; blue, SCNA-losses; green, SVs) for genetic driver alterations to cluster C1-C5 in Chapuy et al. (top row), Schmitz et al (middle row) and the combined set (lower row), respectively. Marginal frequency of significant drivers (q<0.1) for each cluster (in columns) in each set is plotted on the y-axis and compared to the frequencies previously reported in Chapuy et al. 2018 (infra) on the x-axis. Pearson correlations were calculated for each plot.
[0126] FIG. 5A-5B provides genetic features comparisons of 163 genetic discriminating drivers. FIG. 5 A provides a frequency comparison of 163 significantly discriminating driver alterationsacross cohorts (Chapuy cohort, x-axis; Schmitz cohort, y-axis). Alterations are color coded (black, mutations; red, CN gains; blue, CN losses; SVs, green) and those alterations that are significant more frequent in one cohort (two-sided Fisher’s Exact test with BH-correction, q<0.1) are indicated with a cross. FIG. 5 A provides a frequency comparison across all 163 driver alterations used for the molecular classifier are visualized across cohorts (train set, x-axis; test set, y-axis). Color coding and significant testing as in panel a.
[0127] FIG. 6 provides a workflow for developing the molecular DLBCL classifier and evaluation criteria. A, The 163 genetic driver input features were reduced by the dimensionally reduction technologies and then provided as an input vector to the respective algorithm that delivers an output vector. The output vector provides the class label and the information to compute the confidence for each class. B, The NMF cluster designations for the combined cohort were used as gold-standard C1-C5 labels. The combined cohort (n=699) was divided into serial training / validation series (black [n=440] and green [n=110]) and a separate test set (yellow [n=149]). Candidate models (box, top left) composed of different model architectures (neural network, NN; naive Bayer, NB; or random forest RF) and input features (all features, or reduced features) were evaluated in the training / validation set (grey box). The training algorithm for each single candidate model was based on 5-fold cross validation with ensemble averaging (green inner box), iterated 100 times with different random initial conditions. Individual models were evaluated based on their performance in the validation set. Only the final model, designated as DLBclass, was evaluated in an independent test set (bottom right panel).
[0128] FIGS. 7A-7E show the development of DLBclass. Fig. 7A shows model development including baseline model selection (left) and winning model optimization (middle). All candidate models were initially evaluated with our common performance metric (see FIGs. 6A-6B). The winner of a step (indicated with a black arrowhead) advanced to the subsequent step, resulting in the final model DLBclass. Fig. 7B shows a selection of optimal model and input features. Different model architecture methods, including artificial neural networks (NN), random forest (RF) and naive bayes (NB) were explored. All input features (n=163) or features reduced in their complexity by different dimensionality reduction technologies (PC A with n=2, 20, 30 or 40 dimensions or a sum of co-segregated features [CoSegF] into 21 meta-features were explored. The accuracy, the kappa calibration factor, the computed performance and the one standard deviation confidence interval were calculated for each combination and ranked by performance. The winning model had a 21-10-5 NN architecture (21 input, 10 hidden and 5 output nodes) that learned to classify patient samples from a CoSegF reduced input feature set (recipe, see FIG. 8). Fig. 7C shows a model optimization (MO). The winning model of the baseline selection method (NN-CoSegF, q<0.1,grey arrowhead) was further optimized by selecting fewer features (based thresholding their MutSig2CV significance and genomic footprint). Dimensions of the grid search are significance and feature footprints, where the grid spacing is in units of q-value and footprint ranking. Features were dropped in order of the size of the feature in the genome space (larger features were removed first). The final NN-CoSegF model use a q-value threshold of 0.1 and removed the 5 largest features (rm5 / <0.1, purple arrowhead). Fig. 7D shows a model optimization - alteration classes. The requirement for all datatypes (mutations, SCNAs and SVs) was evaluated by sequentially removing classes of alterations which decreased the performance of the model. Fig. 7E shows a model optimization - orthogonal assays. The addition of the cell-of-origin defined transcriptional subtypes (COO) and / or ploidy did not significantly improve the performance of the molecular classifier and were not added to the final model DLBclass.
[0129] FIG 8 is a table showing biological reduced features. All 163 cluster-discriminating genetic features (mutations, black; somatic copy number gains, red; somatic copy number losses, blue; and structural variants, green) were grouped into 21 meta-features (Ml -21, names to the left) that are supportive for the indicated DLBCL clusters to the left. M21-2pl6.1 is not specific (“MISC”) for a given cluster.
[0130] FIGS. 9A-9F provide DLBc / a.s.s results for the training / validation and test sets. FIG. 9A shows the classification output of the final molecular classifier, DLBc / a.s.s, as confusion matrices for the training / validation and independent test sets. The probabilistic classifier includes post-hoc confidence thresholds. Accuracy for training / validation and test cohorts without a confidence threshold (left column), or >0.7 confidence (right column). FIG. 9B shows the calibration metric for training / validation (left) and test sets (right). Kappa is the calibration factor that captures the correlation between confidence and accuracy. The winning model, DLBc / a.s.s, is visualized for the training / validation and test sets. Correctly and incorrectly classified cases are plotted (dark gray dots in the middle box, and plotted in the top box and pale gray dots in the middle box, plotted in the bottom box respectively) and grouped based on confidence bins (black points and confidence intervals). The (y=x) line represents the predictions of a theoretical optimally calibrated model, with black points representing the binned classified samples. Histograms on the top and bottom show the density of correctly and incorrectly classified samples across confidence levels (x-axis). FIG. 9C shows the relationship between accuracy, number of classified cases and confidence for both training / validation (top panel) and test sets (bottom panel). FIG. 9D shows the ranked scatterplots of the confidence for each sample by cluster (C1-C5 DLBCLs, x-axis). Correctly classified samples indicated with a X’s, incorrectly classified cases with a triangle. FIG. 9E shows a UMAP of cluster identification, C1-C5, for all cases (training / validation and test sets) withconfidence >0.7. FIG. 9F shows an alluvial plot summarizing the relationship between NMF clusters, DLBc / a.s.s classifications without and with a confidence threshold of 0.7.
[0131] FIGS. 10A and 10B provide graphs showing additional results of the accuracy, K, and performance of the Final Model versus Neural Network in A and the Neural Network versus the Winning Model in B. FIG. 10A Model optimization - alteration classes. The requirement for all datatypes (mutations, SCNAs and SVs) was evaluated by sequentially removing classes of alterations which decreased the performance of the model. FIG. 10B, Model optimization - orthogonal assays. The addition of the cell-of-origin defined transcriptional subtypes (COO) and / or ploidy did not significantly improve the performance of the molecular classifier and were not added to the final model DLBclass.
[0132] FIG. 11 shows testing of robustness of the molecular classifier for train / validation and test sets. Removal of true positive events yields linear loss of all feature classes (top), and negatively impacts classifier performance (bottom). Results shown for the train and test sets within panels a and b respectively. The false positive simulation introduces false positives to the drivers. Although linear gain of false positives is shown in the top track, the model exhibits roughly exponential decay of performance, underscoring the need for comprehensive germline filtering. Results shown for the test and train sets within panels a and b respectively. Model performance shown for the CCF threshold simulation. Results are fairly stable, but do indicate that its helpful to detect subclonal events to preserve model performance. Results shown for train and test set within panels a and b respectively.
[0133] FIGs. 12A-12F provide a visualization showing DLBclass-predicted C1-C5 tumors in the full (n= 699) series. DLBCLs are visualized with their associated landmark alterations (q-value <0.1, Fisher’s Exact test; non-synonymous mutations, black; synonymous mutations, gray; single CN loss [1.1 < CN < 1.6 copies], cyan; double CN loss [CN < 1.1], blue; low-level CN gain [3.7 copies > CN > 2.2 copies] pink; high-grade CN gain [CN > 3.7 copies], red; chromosomal rearrangements, green; no alterations, white). FIG. 12A shows features significantly associated with class Cl (purple box), with specific signatures of this set also shown to be associated with classes C2-C5. FIG. 12B shows features significantly associated with class C2 (blue box) with specific features of this set also shown to be associated with classes Cl, C3, and C5. FIG. 12C shows features significantly associated with class C3 (yellow box), with specific features of this set also shown to be associated with class C4. FIG. 12D shows features significantly associated with class C4 (green box), with specific features of this set also shown to be associated with classes C2, C3, and C5. FIG. 12E shows features significantly associated with class C5 (red box), with specific features of this set also shown to be associated with classes C1-C4. FIG. 12F shows FIGs.12A-12E as one visualization. Recurrent alterations reported in the original Chapuy et al. series are bolded; alterations driving the original C1-C5 substructure are indicated with an asterisk. Newly identified recurrent alterations in the expanded dataset are not bolded. Header shows cluster association by DLBclass and NMF clustering (Cl, purple; C2, blue; C3, orange; C4, turquoise; C5, red), LymphGen class (BN2, purple; A53, cyan; EZB, orange, MCD, Nl, brown; other and ambiguous classes, yellow, unassigned, gray), cohort membership (Chapuy, black; Schmitz, white), training / validation or test set membership (training / validation set, white; test set, black), and COO classification (ABC, red; GCB, cyan; unclassifiable, yellow; not assessed, gray). Frequency of alterations in the full cohort to the left as bar graph.
[0134] FIGs. 13A-13G provides plots and graphs showing DLBc / a.s.s C1-C5 features. FIG. 13A is an alluvial plot visualizing the comparison of DLBc / a.s.s and COO in n=646 primary DLBCLs with available COO assignment. FIGs. 13B-13F show CoMut plots of selected alterations for all 699 tumors ordered by C1-C5 subtypes: NOTCH2 pathway (FIG. 13B), TP53 / CDKN2A pathway (FIG. 13C), SVs involving BCL6, BCL2 and MYC loci (FIG. 13D), NF-KB modifier mutations (FIG. 13E), Genetic bases of immune evasion (mutations and focal copy number alterations), (FIG. 13F). FIG. 13G provides graphs showing progression-free survival (visualized as Kaplan-Meier plots) for the, ABC-enriched (Cl and C5), GCB-enriched (C3 and C4) and C2 DLBCLs (conf. > 0.7). The p-values were obtained by the log rank test.
[0135] FIGs. 14A-14C provide tables and plots showing a comparison of DLBclass to LymphGen and commonly used targeted sequencing panels and application to DLBCL cell lines. FIG. 14A is an alluvial plot visualizing the comparison of DLBclass and LymphGen subsets in the combined cohort (n=699; top panel). Summary table of DLBclass and LymphGen classes visualized for all, confidence >0.7 and confidence <0.7 (lower panel). FIG. 14B shows the simulated performance of DLBclass with commonly available targeted sequencing panels. Missing data in each panel was replaced with the mean value of the specific alteration in the cohort (n=699). Numbers of measured features indicated as bar graph to the right. See also figure legend of FIGs. 9B-9E. FIG. 14C, shows DLBclass assignment for DLBCL cell lines, ordered by C1-C5 subtype and ranked by confidence. Green line represents a confidence of 0.7.
[0136] FIGS. 15A-C illustrate contributions of mutational signature (A, B), and plot signature proportion and mutations in the indicated genes.
[0137] FIGs. 16A-16P provide stick (Lolipop) figures of driver mutations significantly associated with C1-C5 DLBCLs. For each significantly mutated gene associated with C1-C5 DLBCLs, all mutations are visualized within the functional domains of the respective protein usingMutationMapper v5.3.6 (Cerami et al., 2012; Gao et al., 2013). Genes are ordered in alphabetical order.
[0138] FIG. 17 provides a waterfall plot showing the genetic bases of immune evasion in C1-C5 DLBCLs.
[0139] FIG. 18 provides a schematic showing the overall approach of validating C1-C5 DLBCL and constructing the molecular classifier DLBclass. After assessing and reprocessing an independent dataset (top panel), the individual subsets and the harmonized combined set were used for validation of the C1-C5 DLBCL molecular substructure (left panel). Inclusion criteria were a purity above 0.2 and no missing data. The C1-C5 DLBCL labels from the combined cohort were used for the development of the molecular classifier, DLBclass (middle panel). To this end, different model architectures and input features (middle panel, to the left) were applied to the train and validation set to obtain for each candidate model accuracy and the calibration between accuracy and confidence (performance; middle panel, in the middle). The best model was defined as DLBclass (middle panel, to the right). DLBclass was compared to LymphGen classes (lower panel, to the left) and applied to independent data, including DLBCL cell lines, and made available to the community in form of a portal and a stand-alone Jupyter Notebook (lower panel, to the right).
[0140] FIGs. 19A-19D provide a visualization showing recurrently mutated genes in the combined training cohort (n=550). Number and frequency of recurrently mutated genes (left) visualized by mutation type (center) and ranked by significance (MutSig2CV q-value, right). Total mutation density is shown at the top. Asterisk indicates hypermutator cases.
[0141] FIG. 20 provides graphs showing the model training history plot. Training history is shown across training epoch number (x axis); Number of models left to be trained (y-axis, top) after early stopping criteria reached; Ensemble mean squared error (loss) for training and validation set (bottom). All (n=500) of our networks within DLBc / a.s.s reach the early stopping condition before epoch 40, indicating that training reaches a stable termination point.
[0142] FIG. 21 provides graphs showing training / validation set results from model optimization.
[0143] FIGs. 22A-22C provide graphs and visualizations showing a cluster-swap alluvial for the purity reduction simulation. High confidence cases are shown in FIG. 22B and low confidence cases are shown in FIG. 22C, respectively. Colored columns indicate the cluster size at a specific purity step, whereas colored bands indicate movement of samples across purity steps. More cluster swaps happen at higher purities in low confidence samples, but in both cases the classifier degrades severely below 20% simulated purity. In the absence of all events, the empty sample is classifiedas a C5 tumor with very low confidence of 0.24. This explains why C5 is the “dominant” cluster at the last step of purity =0.02.
[0144] FIGs. 23A-23C provide a waterfall plot showing the genetic bases of immune evasion in C1-C5 DLBCLs. Color-coded waterfall plot shows genetic alterations in indicated immune evasion molecules across C1-C5 DLBCLs. Mutations, black; copy number loss, cyan; deletion, blue; copy number gain, pink; amplification, red; structural variant, green.
[0145] FIGs. 24A-24E provide DLBc / a.s.s assignments of “Other” and “Ambiguous” LymphGen cases.
[0146] FIG. 24 A provides an exemplary Sankey plot of all 312 “Other” and “Ambiguous” LymphGen cases and their positive assignments into DLBc / a.s.s subtypes.
[0147] FIGs. 24B-24C provide an exemplary Sankey plot of “Other” and “Ambiguous” LymphGen cases and their positive assignments into DLBc / a.s.s subtypes with high confidence (conf. >0.7, B) or with low confidence (conf. <0.7, C).
[0148] FIG. 24D provides Sankey diagram that highlights the assignment of the 51 ambiguous cases, based on their LymphGen subtypes.
[0149] FIG. 24E provides an exemplary CoMut Plot of the “Other” and “Ambiguous” LymphGen cases following DLBc / a.s.s assignments, where DLBCLs are visualized with their associated landmark alterations (e.g., q-value <0.1, Fisher’s Exact test; non-synonymous mutations, black; synonymous mutations, gray; single CN loss [1.1 < CN < 1.6 copies], cyan; double CN loss [CN < 1.1], blue; low-level CN gain [3.7 copies > CN > 2.2 copies] pink; high-grade CN gain [CN > 3.7 copies], red; chromosomal rearrangements, green; no alterations, white), and where colored boxes mark features significantly associated with classes C1-C5. Recurrent alterations reported in the original Chapuy et al. series are bolded; alterations driving the original C1-C5 substructure are indicated with an asterisk. Newly identified recurrent alterations in the expanded dataset are not bolded. Header shows cluster association by DLBc / a.s.s (Cl, purple; C2, blue; C3, orange; C4, turquoise; C5, red), confidence of DLBc / a.s.s assignment as bar graph; LymphGen class (ambiguous, pink; other, yellow), cohort membership (Chapuy, light blue; Schmitz, dark blue), and COO classification (ABC, red; GCB, cyan; unclassifiable, yellow; not assessed, white).
[0150] FIGs. 25A-25D provide DLBc / a.s.s assignments of the Lacy dataset and outcome associations. FIG. 25A illustrates DLBc / a.s.s assignments in 874 Lacy cases, where confidence is shown as histogram for each DLBc / a.s.s. FIG. 25B, 25C and 25D illustrates progression-free survival for Lacy cohort patients with available PFS and IPI status visualized using Kaplan Meier Plots (FIG. 25B shows all patients; FIG. 25C shows stratified by IPI categories; FIG. 25D shows stratified by DLBclass), where the p-values are obtained using the logrank test.
[0151] FIG. 26 depicts a landing page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0152] FIG. 27 depicts an upload page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0153] FIG. 28 depicts the data selection tool of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0154] FIG. 29 depicts a classification results page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0155] FIG. 30 depicts a heatmap page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0156] FIG. 31 depicts another heatmap page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0157] FIG. 32 depicts another heatmap page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.
[0158] FIG. 33 depicts a UMAP page of the exemplary DLBCL classification software application in accordance with one or more embodiments of the present disclosure.DETAILED DESCRIPTION OF THE INVENTION
[0159] The present disclosure features panels of markers, a molecular classifier and a targeted sequencing assay for use in characterization and treatment of diffuse large B-cell lymphoma (DLBCL).
[0160] The disclosure provides additional markers (e.g., recurrent mutations of DUSP2, TET2 and SOCSP) and a molecular classifier for use in characterizing C1-C5 DLBCL. These additional markers were identified in the course of validating C1-C5 DLBCL using an enlarged dataset (FIG. 18), and constructing the molecular classifier DLBc / a.s.s. After assessing and reprocessing an independent dataset (FIG. 18, top panel), the individual subsets and the harmonized combined set were used for validation of the C1-C5 DLBCL molecular substructure (FIG. 18, left top panel). Inclusion criteria were a purity above 0.2 and no missing data. The C1-C5 DLBCL labels from the combined cohort were used for the development of the molecular classifier, DLBc / a.s.s (middle panel). To this end, different model architectures and input features (middle panel, to the left) were applied to the train and validation set to obtain for each candidate model accuracy and the calibration between accuracy and confidence (FIG. 18, performance; middle panel, in the middle). The best model was defined as DLBc / a.s.s (FIG. 18, middle panel, to the right). DLBc / a.s.s was compared to LymphGen classes (FIG. 18, lower panel, to the left) and applied to independent data,including DLBCL cell lines, and will be made available to the community in the form of a portal and a stand-alone Jupyter Notebook (FIG. 18, lower panel, to the right).Construction of the Molecular Classifier.
[0161] The disclosure provides a simple classifier that: 1) allows robust classification of C1-C5 DLBCL subtypes; 2) provides the confidence (probability) for a given sample to be in the respective class; and 3) is as simple as possible. To develop the classifier, the cohort was split into an independent test set (n=149) and a training / validation set (n=550). Next, different model architectures (naive bayes, neural networks, random forest) were explored and different dimensionality reduction technologies (all features, PCA2, PCA20, PCA30, PCA40, biologically informed) to pick a baseline model (FIG. 7A-7E, 9A . Iterative loops though the training / validation set were used to calculate the performance for each combination of model architecture and reduction technology (FIG. 7A-7E). This performance score is defined as a weighted harmonic mean between the two core metrics of accuracy and calibration, with accuracy weighted twice as much as the calibration factor, kappa (FIG. 7A-7E). The winning model in the baseline selection was a neural network with a biologically-informed reduction of markers (weighted sum, recipe see FIG. 8). Further model optimization demonstrated that removing the 5 largest features performs equally well, resulting in the final model (NN.Bioinf.5rm) (FIG. 9B). Notably, all genetic features are required and adding additional orthogonal data such as genome doubling, and cell-of-origin transcriptional subtypes did not further improve the model (FIG. 9C- 9D). Applying the winning model to the training / validation and independent test set demonstrated that the overall accuracy for the full cohort is 91% and 89%, respectively, and 97% and 98% for cases with a confidence of >0.7 (75% of all cases) (FIG. 12A-12F). Additional experiments demonstrated the robustness of the classifier (spike in, drop out and purity experiments).DLBCL
[0162] Diffuse large B-cell lymphoma is the most common type of non-Hodgkin lymphoma (NHL) accounting for about 22 percent of newly diagnosed cases of B-cell NHL in the United States. DLBCL is an aggressive (fast-growing) NHL that affects B-lymphocytes. The occurrence of DLBCL generally increases with age, and patients tend to be over the age of 60 at diagnosis. DLBCL can develop in the lymph nodes or in “extranodal sites” (areas outside the lymph nodes) such as the gastrointestinal tract, testes, thyroid, skin, breast, bone, brain, or essentially any organ of the body. It may be localized (in one spot) or generalized (spread throughout the body).
[0163] Earlier transcriptional analyses of DLBCL revealed disease heterogeneity related to putative cell s-of-ori gin (also termed “COO”), signaling and metabolic programs and tumor microenvironment. The cell s-of-ori gin framework includes: (i) activated B-cell (ABC) typeDLBCLs with genetic bases of increased NFkB activity and enhanced B-cell receptor (BCR) signaling and inferior outcomes following R-CHOP induction therapy (Rosenwald et al., N Engl J Med 346, 1937-1947; Davis et al., Nature 463, 88-92; and Ngo et al., Nature 470, 115-119); (ii) germinal center B-cell (GCB) type DLBCLs with more favorable responses to R-CHOP treatment; and (iii) additional unclassified tumors (Rosenwald, supra). A more recently defined subset of GCB-type DLBCLs with a dark zone transcriptional signature (DZsig) and frequent concurrent BCL2 and MYC translocations responds less well to standard induction therapy (Ennishi et al., J Clin Oncol 37, 190-201; Sha et al., J Clin Oncol 37, 202-212; Hilton et al., Blood 134, 1528-1532; and Alduaji et al., Blood 141, 2493-2507).
[0164] To characterize the genetic heterogeneity of primary DLBCLs, recurrent mutations were previously identified, somatic copy number alterations (SCNAs) and structural variants (SVs), and applied non-negative matrix factorization (NMF) consensus clustering to define 5 genetically distinct subtypes, Clusters 1-5 (C1-C5) (Chapuy et al., Nat Med 24, 679-690, 2018; See, also, PCT / US2022 / 020762, which is incorporated herein by reference in its entirety). C5 DLBCLs were ABC-type tumors with near-uniform 18q (BCL2) copy gain, frequent concurrent CD79B and MYD88L265Pmutations, extranodal tropism and an inferior response to R-CHOP treatment (Chapuy supra, Chapuy et al., Blood 127, 869-881, 2016). Although Cl DLBCLs were also enriched for ABC-type tumors, the defining genetic alterations resembled those of transformed marginal zone B-cell lymphomas (NOTCH2 and NF-KB pathway components and BCL6 SVs) and included bases of immune escape (Chapuy supra). In contrast to C5 DLBCLs, ABC-enriched Cl tumors had a more favorable outcome following R-CHOP therapy (Chapuy supra).
[0165] C3 DLBCLs were GCB-type tumors with frequent BCL2 translocations and mutations in chromatin modifiers, B-cell transcription factors and B-cell receptor (BCR) / PI3K signaling pathway components and an unfavorable response to R-CHOP therapy (Chapuy supra). In contrast, the GCB-enriched C4 DLBCLs had alterations in multiple linker and core histone genes, RAS / JAK / STAT pathway members and certain modulators of PI3K pathway signaling and a more favorable outcome following R-CHOP (Chapuy supra). Lastly, C2 DLBCLs were COO- independent with frequent bi-allelic inactivation of TP53, CDKN2A copy loss, increased genomic instability and a distinct outcome trajectory following R-CHOP (Chapuy supra).
[0166] In a concurrently reported study, Schmitz et al. (N Engl J Med 378, 1396-1407, 2018; the NCI team) NIH investigators characterized the recurrent genetic alterations in an independent series of newly diagnosed DLBCLs and defined molecular substructure associated with COO transcriptional categories (Schmitz supra). Three of the four identified genetic subtypes, MCD (with co-occurring MYD88L265Pand CD79B mutations), BN2 (with BCL6 translocations andN0TCH2 mutations) and EZB (with EZH2 mutations and BCL2 translocations), shared features of our C5, Cl and C3 DLBCLs, respectively (Chapuy supra, Schmitz supra). Subsequent reanalysis of the NIH DLBCL series revealed 2 additional subtypes, A53 and ST2, with genetic alterations resembling those in our C2 and C4 tumors; EZB DLBCLs with and without MYC SVs were also described (Chapuy supra, Wright et al., Cancer Cell 37, 551-568, 2020). A minor subset (2.8%) of NIH tumors termed NOTCH1 (Nl) was not detected in the DLBCL series herein (Chapuy supra, Wright et al., Cancer Cell 37, 551-568, 2020).
[0167] Shortly thereafter, Lacy et al. (Blood 135, 1759-1771, 2020; the UK team) utilized a 293 gene targeted sequencing panel and an alternative clustering algorithm to delineate DLBCLs with genetic alterations shared by C5 / MCD, C1 / BN2, C3ZEZB and C4 / ST2 DLBCLs (Lacy, supra). Additionally, the UK group detected the largely copy number-driven p53-deficient C2 DLBCLs in a publicly available dataset using independent clustering methods (Lacy, supra).
[0168] The identification of molecularly distinct subsets of COO-defined DLBCLs prompted a reanalysis of the earlier randomized phase III Phoenix trial that failed to meet its predefined therapeutic endpoint (Younes et al., J Clin Oncol 37, 1285-1295, 2019; Wilson et al., Cancer Cell 39, 1643-1653, 2021, and Mondello et al., Cancer Cell 39, 1570-1572, 2021). This clinical trial of R-CHOP plus the BTK inhibitor, ibrutinib, or placebo in newly diagnosed patients with non-GCB DLBCL was negative, in part because ibrutinib was prohibitively toxic in older patients (>60 years) (Younes, supra). When DLBCLs from the younger trial patients (<60 years) were reanalyzed for molecular subtypes, R-CHOP plus ibrutinib was selectively beneficial in patients with MCD and Nl, but not BN2, tumors (Wilson, supra).
[0169] In the more recent randomized phase II Guidance-01 trial, patients with primary DLBCL underwent prospective molecular tumor typing and received subtype-associated targeted therapy (R-CHOP plus X) or standard treatment (R-CHOP). Patients whose induction therapy was based on their DLBCL molecular subtype (R-CHOP plus X) had a significantly higher response rate and progression-free survival than those who were treated with standard R-CHOP (Zhang et al., Cancer Cell 41, 1705-1716 el705, 2023).
[0170] The increasing recognition and targeting of genetically defined DLBCLs underscores the need for robust classification algorithms. A classifier developed by Wright et al. (the NCI team), termed LymphGen, iteratively identified subtypes based on the key genetic alterations characteristic of their six subtypes, followed by a Bayesian prediction model (Wright, supra). In the original NCI series of newly diagnosed DLBCLs, LymphGen classified 57% of tumors as MCD, BN2, Nl, EZB, ST2 or A53; however, 5.7% were designated Genetically Composite and 36.9% as Other (Wright, supra).
[0171] Recently, the LymphGen algorithm was also used to retrospectively classify DLBCLs in the randomized phase III POLARIX trial of the anti-CD79b immunotoxin, polatuzumab vedotin (Pola), plus R-CHP versus R-CHOP alone in newly diagnosed patients. Again, only 53% of tumors were positively classified into the six genetic categories; 3.8% of tumors were typed as Genetically Composite and 43.2% as Other (Morschhauser et al., Blood 142, 3000-3003, 2023). Guidance-01 trial investigators utilized a different 20-gene algorithm that positively identified 60% of trial patients’ tumors as MCD-like, BN2-like, Nl-like, EZB-like or TP53mut; 39% of DLBCLs were designated NOS (not otherwise specified) (Morschhauser, supra).
[0172] These findings highlight the large numbers of patients whose newly diagnosed DLBCLs are not positively classified into genetic subtypes with current algorithms. Moreover, the concordance between recently developed genetic classifications and prospective algorithms and the necessary genetic elements remain to be defined. As reported in the examples below, this disclosure addresses these questions, validates a 5-cluster DLBCL substructure (Chapuy 2018, supra) in the independent NCI dataset and uses the combined series of primary DLBCLs to provide a robust, prospective and probabilistic molecular classifier for use in clinical trials and practice.DLBCL Classification
[0173] The methods and compositions described herein relate to identification of a clinically useful classifier for DLBCL, the development of which is based upon an assessment of a significantly powered cohort of DLBCL samples for variants across whole exome sequences, where such variant assessment included characterization and evaluation of each of the following types of variation: somatic single nucleotide variants (SNVs), small insertions and deletions (Indels), somatic copy number alterations (SCNAs) and structural variants (SVs), including identification of variation across all cancer causing genes (CCGs). The classifier can predict outcome for a subject having a DLBCL independent of current clinical prognostic classification systems (e.g., the International Prognostic Index (IPI)). Specific components of the instant DLBCL classifier include the following.Types of Samples
[0174] The present disclosure provides methods to extract and sequence a polynucleotide present in a sample. In one embodiment, the samples are biological samples generally derived from a subject (e.g., mammal, such as a human), such as a bodily fluid (such as ascites, blood, plasma, pleural fluid, serum, cerebrospinal fluid, phlegm, saliva, stool, urine, semen, prostate fluid, breast milk, or tears), or tissue sample (e.g., biopsy (e.g., needle biopsy), primary tumor sample, tissue section). In still another embodiment, the samples are biological samples from in vitro sources(e.g., cell culture medium). In an embodiment, the biological sample is plasma containing cell free DNA (cfDNA) or circulating tumor DNA (ctDNA)
[0175] In embodiments, a liquid sample (e.g., blood, plasma, serum) comprises at least about and / or less than about 1 pl, 10 pl, 100 pl, 200 pl, 300 pl, 400 pl, 500 pl, 600 pl, 700 pl, 800 pl, 900 pl, 1 ml, 2 ml, 3 ml, 4 ml, 5 ml, 6 ml, 7 ml, 8 ml, 9 ml, 10 ml, or 15 ml. In embodiments, a sample comprises at least about and / or less than about 1 mg, 10 mg, 100 mg, 200 mg, 300 mg, 400 mg, 500 mg, 600 mg, 700 mg, 800 mg, 900 mg, 1 g, 2 g, 3 g, 4 g, 5 g, 6 g, 7 g, 8 g, 9 g, 10 g, or 15 g. In various cases, the methods provided herein can be completed successfully using any of the above-listed sample volumes and / or masses.Reference Sequences
[0176] In certain aspects, the instant disclosure provides methods and kits that involve and / or allow for assessment of the presence or absence of one or more sequence variants and / or mutations (e.g., structural variants including translocations (SVs), somatic copy number alterations (SCNAs) and recurrent mutations) in a test subject, tissue, cell or sample, as compared to a corresponding reference sequence. In particular embodiments, a subject, tissue, cell and / or sample is assessed for one or more variants and / or sites of copy number variation.
[0177] Up to five alteration types (alternatively “classes”) were measured and can be used for the classifier (i.e., a prognostic classifier as exemplified herein):1.) Mutations (single nucleotide variants and / or InDeis)2.) Copy number alterations (CN gain, amplifications, CN losses, Deletions)3.) Structural variants (chromosomal translocations, inversions, tandem duplications, etc.)4.) Genome doublings5.) Mutational Signatures
[0178] In some instances, the alteration types used for the classifier include structural variants (SVs) including translocations, somatic copy number alterations (SCNAs) and mutations.
[0179] For classifier performance consistent with the exemplar described below as “Neural Network classifier” the somatic variants (1-5) should be detected with at least 90% sensitivity and less than 5% false detection rate.
[0180] Mutations in Candidate Cancer Genes (CCGs), hereafter referred to as driver mutations, were identified with MutSig2CV and CLUMPS.
[0181] Recurrent copy number alterations were identified using GISTIC2.0, as described in U.S. Patent Application Publication No. 2019 / 0292602.
[0182] Cl and C5 DLBCLs had significantly different cAID signature activity.
[0183] It is expressly contemplated that either all or a subset of these five classes of alterations may be measured as part of the methods provided herein, where any combination of the individual members of each class, or even other genes, can be used within a classifier of the instant disclosure. Detection of Alterations
[0184] In some aspects, exome sequencing or probe-hybridization is performed upon a test sample (e.g., a biological sample containing circulating tumor DNA (ctDNA), cell free DNA (cfDNA), RNA, polypeptides, mixtures thereof, and the like) for purpose of detecting variants and / or copy number variation as described herein and identifying DLBCL classification and selecting a therapy. In certain embodiments, assessment of candidate and / or test DLBCL samples can be performed using one or more amplification and / or sequencing oligonucleotides flanking the above-referenced variant sequence and / or copy number variation regions. Assessment of candidate and / or test DLBCL samples can be performed using one or more baits (described further below) for use in targeted sequencing of variants. The assessment can involve using baits to target particular sequences from a sample for subsequent sequencing.
[0185] In some embodiments, assessment of candidate and / or test DLBCL samples can be performed by sequencing a library of polynucleotides prepared from a sample (e.g., circulating tumor DNA from a subject). The library of polynucleotides can be sequenced in some embodiments as described further below and in the Examples of the present disclosure using targeted sequencing using a targeted sequencing panel comprising baits (i.e., probes or polynucleotides targeting variants). The assessment can also be performed based upon binding of a labeled bait(s) (e.g., an oligonucleotide(s)) to a target sequence in the sample. Design and use of such amplification and sequencing oligonucleotides, and / or copy number detection probes / oligonucleotides (e.g., baits), can be performed by one of ordinary skill in the art.
[0186] The methods provided herein can be used for enriching for target polynucleotides. The polynucleotides are associated with a genetic alteration of interest (e.g., SVs, SCNAs, or mutations). The polynucleotides can be enriched from a sample by about or at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100-fold.
[0187] In some instances the library is prepared using about, less than about, and / or at least about, 0.1 ng, 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 10 ng, 15 ng, 20 ng, 25 ng, 30 ng, 35 ng, 40 ng, 45 ng, 50 ng, 75 ng, 100 ng, 250 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 1,000 ng, or more of DNA. In some cases, the library is prepared using DNA fragments with an average size of about, at least about, and / or of no more than about 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, 500 bp, or 1,000 bp.
[0188] Methods for preparing libraries of polynucleotides for sequencing are known to one of skill in the art. Library preparation can include the addition of nucleotide barcodes to the library polynucleotides according to methods known in the art.
[0189] As will be appreciated by one of ordinary skill in the art, any amplification, sequencing (e.g., targeted sequencing), and / or copy number detection oligonucleotides can be modified by any of a number of art-recognized moieties and / or exogenous sequences, e.g., to enhance the processes of amplification, hybridization, sequencing reactions and / or detection. Exemplary oligonucleotide modifications that are expressly contemplated for use with the oligonucleotides of the instant disclosure include, e.g., fluorescent and / or radioactive label modifications; labeling one or more oligonucleotides with a universal amplification sequence (optionally of exogenous origin) and / or labeling one or more oligonucleotides of the instant disclosure with a unique identification sequence (e.g., a “bar-code” sequence, optionally of exogenous origin), as well as other modifications known in the art and suitable for use with oligonucleotides.
[0190] In embodiments, the polynucleotides (e.g., baits, probes, or oligonucleotides) provided herein (e.g., baits, probes, or oligonucleotides) contain one or more modifications or analogs.
[0191] For example, in some embodiments a polynucleotide contains one or more analogs (e.g., altered backbone, sugar, or nucleobase). Some non-limiting examples of analogs include 5- bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine.
[0192] In embodiments, the polynucleotide contains a modified backbone and / or linkages (e.g., between adjacent nucleosides). Non-limiting examples of modified backbones include those that contain a phosphorus atom in the backbone and those that do not contain a phosphorus atom in the backbone. Non-limiting examples of modified backbones include phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonate such as 3' -alkylene phosphonates, 5 '-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates including 3 '-amino phosphoramidate and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates having normal 3 '-5' linkages, 2'-5' linked analogs, and those having inverted polarity wherein one or more intemucleotide linkages is a 3' to 3', a 5' to 5' or a 2' to 2' linkage.
[0193] In embodiments, a polynucleotide contains short chain alkyl or cycloalkyl linkages (e.g., between adjacent nucleosides), mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatomic or heterocyclic internucleoside linkages. In embodiments, a polynucleotide includes one or more of the following: morpholino linkages (formed in part from the sugar portion of a nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methylene formacetyl and thioformacetyl backbones; riboacetyl backbones; alkene containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 component parts.
[0194] In embodiments, a polynucleotide contains a nucleic acid mimetic. The term “mimetic” can be intended to include polynucleotides wherein only the furanose ring or both the furanose ring and the intemucleotide linkage are replaced with non-furanose groups, replacement of only the furanose ring can also be referred as being a sugar surrogate. The heterocyclic base moiety or a modified heterocyclic base moiety can be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar- backbone of a polynucleotide can be replaced with an amide containing backbone, in particular an aminoethylglycine backbone. The nucleotides can be retained and are bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone. In embodiments, the backbone in PNA compounds contains two or more linked aminoethylglycine units that give PNA an amide containing backbone. Heterocyclic base moieties can be bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone.
[0195] In embodiments, a polynucleotide contains a morpholino backbone structure. For example, a nucleic acid can contain a 6-membered morpholino ring in place of a ribose ring. In some of these embodiments, a phosphorodiamidate or other non-phosphodiester internucleoside linkage can replace a phosphodiester linkage.
[0196] A polynucleotide can contain linked morpholino units having heterocyclic bases attached to the morpholino ring. Linking groups can link morpholino monomeric units. Non-ionic morpholino-based oligomeric compounds can have less undesired interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. A variety of compounds within the morpholino class can be joined using different linking groups. A further class of polynucleotide mimetic can be referred to as cyclohexenyl nucleic acids (CeNA). In some instances, the furanose ring normally present in a nucleic acid molecule is replaced with a cyclohexenyl ring. CeNA DMT protected phosphoramidite monomers can be prepared and used for oligomeric compound synthesis using phosphoramidite chemistry. In some cases,incorporation of CeNA monomers into a nucleic acid chain increases the stability of a DNA / RNA hybrid. CeNA oligoadenylates can form complexes with nucleic acid complements with similar stability to the native complexes. In embodiments, a polynucleotide contains Locked Nucleic Acids (LNAs) in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring thereby forming a 2'-C, 4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene ( — CH2), group bridging the 2' oxygen atom and the 4' carbon atom wherein n is 1 or 2. LNA and LNA analogs can display very high duplex thermal stabilities with complementary nucleic acid (Tm=+3 to +100C.), stability towards 3'-exonucleolytic degradation and good solubility properties.
[0197] In embodiments, a polynucleotide contains nucleobase modifications (often referred to simply as “base modifications”) or substitutions. In embodiments, unmodified nucleobases include one or more of the purine bases, (e.g., adenine (A) and guanine (G)), and / or the pyrimidine bases, (e.g., thymine (T), cytosine (C) and uracil (U)). Non-limiting examples of modified nucleobases include nucleobases such as 5-methylcytosine (5-me-C), 5 -hydroxymethyl cytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2- propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2- thiocytosine, 5-halouracil and cytosine, 5-propynyl ( — C=C — CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8- substituted adenines and guanines, 5-halo particularly 5-bromo, 5 -trifluoromethyl and other 5- substituted uracils and cytosines, 7-m ethyl guanine and 7-m ethyladenine, 2-F-adenine, 2- aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3- deazaguanine and 3 -deazaadenine. Further non-limiting examples of modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine(lH-pyrimido(5,4-b)(l,4)benzoxazin- 2(3H)-one), phenothiazine cytidine (lH-pyrimido(5,4-b)(l,4)benzothiazin-2(3H)-one), G-clamps such as a substituted phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b) (l,4)benzoxazin-2(3H)-one), phenothiazine cytidine (lH-pyrimido(5,4-b)(l,4)benzothiazin- 2(3H)-one), G-clamps such as a substituted phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H- pyrimido(5,4-(b) (l,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4, -b)indol-2- one), pyridoindole cytidine (H-pyrido(3',2':4, 5)pyrrolo[2,3-d]pyrimidin-2-one).
[0198] Variants can be characterized by sequencing polynucleotides. Characterization of a variant can involve sequencing all or a portion of a variant targeted by an oligonucleotide probe with an oligonucleotide sequence corresponding to any of SEQ ID NOs: 1-9244 or to one or more of the baits described further below. The polynucleotides can be DNA fragments.
[0199] Sequencing may be performed on any high-throughput platform. Methods of sequencing oligonucleotides and nucleic acids are well known in the art (see, e.g., WO 93 / 23564, WO 98 / 28440 and WO 98 / 13523; U.S. Pat. Nos. 5,525,464; 5,202,231; 5,695,940; 4,971,903; 5,902,723; 5,795,782; 5,547,839 and 5,403,708; Sanger et al., Proc. Natl. Acad. Sci. USA 74:5463 (1977); Drmanac et al., Genomics 4: 114 (1989); Koster et al., Nature Biotechnology 14: 1123 (1996); Hyman, Anal. Biochem. 174:423 (1988); Rosenthal, International Patent Application Publication 761107 (1989); Metzker et al., Nucl. Acids Res. 22:4259 (1994); Jones, Biotechniques 22:938 (1997); Ronaghi et al., Anal. Biochem. 242:84 (1996); Ronaghi et al., Science 281 :363 (1998); Nyren et al., Anal. Biochem. 151 :504 (1985); Canard and Arzumanov, Gene 11 : 1 (1994); Dyatkina and Arzumanov, Nucleic Acids Symp Ser 18: 117 (1987); Johnson et al., Anal. Biochem.136: 192 (1984); and Eigen and Rigler, Proc. Natl. Acad. Sci. USA 91(13):5740 (1994), all of which are expressly incorporated by reference). In one embodiment, the sequencing of a DNA fragment is carried out using commercially available sequencing technology SBS (sequencing by synthesis) by Illumina. In another embodiment, the sequencing of the DNA fragment is carried out using chain termination method of DNA sequencing. In yet another embodiment, the sequencing of the DNA fragment is carried out using one of the commercially available next-generation sequencing technologies, including SMRT (single-molecule real-time) sequencing from Pacific Biosciences, Ion Torrent™ sequencing from ThermoFisher Scientific, Pyrosequencing (454) from Roche, and SOLiD® technology from Applied Biosystems. Any appropriate sequencing technology may be chosen for sequencing. Further examples of sequencing methods suitable for use in the methods of the present disclosure include those described in U.S. Patent Application Publication No. 2019 / 0078232, which is incorporated herein by reference in its entirety for all purposes. A DNA sample can be pre-screened using ultra low pass sequencing to confirm that the sample contains sufficient levels of tumor-derived DNA prior to further characterization, as described in U.S. Patent Application Publication No. 2019 / 0078232.
[0200] In various aspects, the methods provided herein involve sequencing of a sample. In some embodiments, the sequencing is whole-genome sequencing (WGS) or whole-exome sequencing (WES). The sequencing is performed upon a test sample for purpose of detecting alterations, such as somatic copy number alterations, mutations (e.g., single nucleotide polymorphisms), and / or structural variations. In certain embodiments, the sequencing can be performed with or without amplification of a sample to be sequenced. In embodiments, a sample is sequenced to a coverage of about, at least about, and / or no more than about O.Olx, 0.05x, O.lx, 0.2x, 0.3x, 0.4x, 0.5x, lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, lOOx, 200x, 300x, 400x,500x, 600x, 700x, 800x, 900x, lOOOx, 5OOOx, lOOOOx, 15OOOx, 20000x, 25000x, 3OOOOx, 5OOOOx, lOOOOOx, or more.
[0201] Whole genome sequencing (also known as “WGS”, full genome sequencing, complete genome sequencing, or entire genome sequencing) is a process that involves sequencing a complete DNA sequence of an organism’s genome. A common strategy used for WGS is shotgun sequencing, in which DNA is broken up randomly into numerous small segments, which are sequenced. Sequence data obtained from one sequencing reaction is termed a “read.” The reads can be assembled together based on sequence overlap. The genome sequence is obtained by assembling the reads into a reconstructed sequence.
[0202] Whole exome sequencing (“WES”) is a technique used to sequence all the expressed genes in a cell or subject. WES includes first selecting only that portion of a polynucleotide sample that encodes proteins (e.g., cDNA, or a subset of a cfDNA sample), and then sequencing using any DNA sequencing technology well known in the art or as described herein. In a human being, there are about 180,000 exons, which constitute about 1% of the human genome, or approximately 30 million base pairs. In some embodiments, to sequence the exons of a genome, fragments of doublestranded genomic DNA are obtained (e.g., by methods such as sonication, nuclease digestion, or any other appropriate methods). Linkers or adapters are then attached to the DNA fragments, which are then hybridized to a library of polynucleotides designed to capture only the exons. The hybridized DNA fragments are then selectively isolated and subjected to sequencing using any sequencing method known in the art or described herein.
[0203] In aspects of the present disclosure, a sample is analyzed by means of a biochip (also known as a microarray) containing targeted baits (oligonucleotides specific for a target alteration). Targeted baits specific for target alterations (e.g., select SV, SCNAs, and mutations) are useful as hybridizable array elements in a biochip. Biochips generally comprise solid substrates and have a generally planar surface, to which a capture reagent (also called an adsorbent or affinity reagent) is attached. Frequently, the surface of a biochip comprises a plurality of addressable locations, each of which has the capture reagent bound there.
[0204] The array elements are organized in an ordered fashion such that each element is present at a specified location on the substrate. Useful substrate materials include membranes, composed of paper, nylon or other materials, filters, chips, glass slides, and other solid supports. The ordered arrangement of the array elements allows hybridization patterns and intensities to be interpreted as expression levels of particular genes or proteins. Methods for making nucleic acid microarrays are known to the skilled artisan and are described, for example, in U.S. Pat. No. 5,837,832, Lockhart, et al. (Nat. Biotech. 14: 1675-1680, 1996), and Schena, et al. (Proc. Natl. Acad. Sci.93: 10614-10619, 1996), herein incorporated by reference. Methods for making polypeptide microarrays are described, for example, by Ge (Nucleic Acids Res. 28: e3. i-e3. vii, 2000), MacBeath et al., (Science 289: 1760-1763, 2000), Zhu et al. (Nature Genet. 26:283-289), and in U.S. Pat. No. 6,436,665, hereby incorporated by reference.
[0205] In aspects of the present disclosure, a sample is analyzed by means of a nucleic acid biochip (also known as a nucleic acid microarray). To produce a nucleic acid biochip, oligonucleotides may be synthesized or bound to the surface of a substrate using a chemical coupling procedure and an inkjet application apparatus, as described in PCT application WO95 / 025116 (Baldeschweiler et al.). Alternatively, a gridded array may be used to arrange and link cDNA fragments or oligonucleotides to the surface of a substrate using a vacuum system, thermal, UV, mechanical or chemical bonding procedure.
[0206] Incubation conditions are adjusted such that hybridization occurs with precise complementary matches or with various degrees of less complementarity depending on the degree of stringency employed. For example, stringent salt concentration will ordinarily be less than about 750 mM NaCl and 75 mM trisodium citrate, less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvent, e.g., formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, or at least about 50% formamide. Stringent temperature conditions will ordinarily include temperatures of at least about 30 °C, of at least about 37 °C, or of at least about 42 °C. Varying additional parameters, such as hybridization time, the concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency are accomplished by combining these various conditions as needed. In an embodiment, hybridization will occur at 30 °C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In embodiments, hybridization will occur at 37 °C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 pg / ml denatured salmon sperm DNA (ssDNA). In other embodiments, hybridization will occur at 42 °C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 pg / ml ssDNA. Useful variations on these conditions will be readily apparent to those skilled in the art.
[0207] The removal of nonhybridized probes may be accomplished, for example, by washing. The washing steps that follow hybridization can also vary in stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As above, wash stringency can be increased by decreasing salt concentration or by increasing temperature. For example, stringent salt concentration for the wash steps will be less than about 30 mM NaCl and 3 mM trisodiumcitrate, or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash steps will ordinarily include a temperature of at least about 25 °C, of at least about 42 °C, or of at least about 68 °C. In embodiments, wash steps will occur at 25 °C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In an embodiment, wash steps will occur at 42 °C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In other embodiments, wash steps will occur at 68 °C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations on these conditions will be readily apparent to those skilled in the art.
[0208] Detection system for measuring the absence, presence, and amount of hybridization for all of the distinct nucleic acid sequences are well known in the art. For example, simultaneous detection is described in Heller et al., Proc. Natl. Acad. Sci. 94:2150-2155, 1997. In embodiments, a scanner is used to determine the levels and patterns of fluorescence.
[0209] For purpose of this disclosure, the term “amplification” means any method employing a primer and a polymerase for replicating a target sequence linearly or exponentially with reasonable fidelity. Amplification may be carried out by natural or recombinant DNA polymerases such as TaqGold™, T7 DNA polymerase, Klenow fragment of E.coli DNA polymerase, and reverse transcriptase. In an embodiment, the amplification method is PCR. Typically, the amplification of a sample results in an exponential increase in copy number of the amplified sequences. Amplification may involve thermocycling or isothermal amplification (such as through the methods RPA or LAMP).
[0210] Design and use of oligonucleotides for amplification and / or sequencing is within the knowledge of one of ordinary skill in the art. Oligonucleotides can be modified by any of a number of art-recognized moieties and / or exogenous sequences, e.g., to enhance the processes of amplification, sequencing reactions, and / or detection. Exemplary oligonucleotide modifications that are expressly contemplated for use with the oligonucleotides of the instant disclosure include, e.g., fluorescent and / or radioactive label modifications; labeling one or more oligonucleotides with a universal amplification sequence (optionally of exogenous origin) and / or labeling one or more oligonucleotides of the instant disclosure with a unique identification sequence (e.g., a “bar-code” sequence, optionally of exogenous origin), as well as other modifications known in the art and suitable for use with oligonucleotides.Classification
[0211] In the methods of the present disclosure, a neural network classifier (i.e., a classification model) can be employed to define DLBCL classification groups (C1-C5). As would be appreciated by one of ordinary skill in the art, other forms of classifier (e.g., nearest-neighbor, decision trees, boosting, support-vector machines, Naive Bayes, bagging, random forests, and various others) canbe applied to variant and / or copy number data, to perform such test sample classification. Methods for the design of classifiers are described, for example, in Baldi Pierre, and Brunak, Soren, Bioinformatics: The Machine Learning Approach. 2ndEd. Cambridge, MA, “A Bradford Book” 2001; Zhang, Y. and Rajapakse, J., Machine Learning in Bioinformatics, Wiley, 2009; and Srinivasa K., Siddesh G., Manisekhar S. (eds) Statistical Modelling and Machine Learning Principles for Bioinformatics Techniques, Tools, and Applications. Algorithms for Intelligent Systems. Springer, Singapore (2020), the disclosures of each of which are incorporated herein in their entirety by reference for all purposes.
[0212] In embodiments, the methods of the disclosure involve dimensionality reduction (e.g., through the use of metafeatures), where dimensionality reduction is the transformation of data from a high-dimensional space into a low-dimensional space so that the low-dimensional representation retains meaningful properties of the original data. Dimensionality reduction can involve a linear or a nonlinear approach, and / or feature selection or feature extraction. In feature selection a subset of input variables can be selected using any of a variety of methods, such as a filter strategy, a wrapper strategy, or an embedded strategy. Non-limiting examples of techniques that can be used for dimensionality reduction include feedforward neural networks, principal component analysis (PCA), non-negative matrix factorization (NMF) (see, Chapuy B, Stewart C, Dunford AJ, et al. Molecular subtypes of diffuse large B cell lymphoma are associated with distinct pathogenic mechanisms and outcomes. Nat Med. 2018;24(5):679-690, the disclosure of which is incorporated herein by reference in its entirety for all purposes), kernel PCA, graph-based kernel PCA, linear discriminant analysis (LDA), generalized discriminant analysis (GDA), autoencoder, t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation projection. In embodiments, dimensionality reduction is used to avoid overtraining a classification algorithm.
[0213] A neural network consists of units (neurons), arranged in layers, which convert an input vector (e.g., a vector comprising 21 metafeature values) into some output In embodiments, the input vector contains about, at least about, and / or no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 400, or 500 metafeature values. Each unit takes an input, applies a (often nonlinear) function to it and then passes the output on to the next layer. Generally, the networks are defined to be feed-forward: a unit feeds its output to all the units on the next layer, but there is no feedback to the previous layer.
[0214] In embodiments, the molecular classifier is probabilistic and includes a post-hoc threshold on confidence. In some cases, the confidence threshold is about, at least about, and / or no morethan about 0%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. It is advantageous in some contexts to select a lower confidence threshold (e.g., below 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%) to allow for classification of a larger number of samples, and it is advantageous in some contexts (e.g., in the clinical setting) to select a higher confidence threshold (e.g., above 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99%).
[0215] Weightings are applied to the signals passing from one unit to another, and it is these weightings which are tuned in the training phase to adapt a neural network to the particular problem at hand. This is the learning phase.
[0216] Neural networks have found application in a wide variety of problems. These range from function representation to pattern recognition, with pattern recognition being the focus of use of neural net classifiers of the instant disclosure.
[0217] In some embodiments, data derived from assays (e.g., genomic analyses) that are generated using samples such as “known samples” can then be used to “train” a classification model. A “known sample” is a sample that has been pre-classified. The data that are used to form the classification model can be referred to as a “training data set.” Once trained, the classification model can recognize patterns in data derived from unknown samples. The classification model can then be used to classify the unknown samples into classes (C1-C5). This can be useful, for example, in predicting whether or not a particular biological sample is associated with a certain biological condition (e.g., diseased versus non-diseased). The training data set that is used to form the classification model may comprise raw data or pre-processed data.
[0218] Classification models can be formed using any suitable statistical classification (or “learning”) method that attempts to segregate bodies of data into classes based on objective parameters present in the data. Classification methods may be either supervised or unsupervised. Examples of supervised and unsupervised classification processes are described in Jain, “Statistical Pattern Recognition: A Review”, IEEE Transactions on Pattern Analysis andMachine Intelligence, Vol. 22, No. 1, January 2000, the teachings of which are incorporated by reference.
[0219] In supervised classification, training data containing examples of known categories are presented to a learning mechanism, which learns one or more sets of relationships that define each of the known classes. New data may then be applied to the learning mechanism, which then classifies the new data using the learned relationships. Examples of supervised classification processes include linear regression processes (e.g., multiple linear regression (MLR), partial least squares (PLS) regression and principal components regression), binary decision trees (e.g., recursive partitioning processes such as CART - classification and regression trees), artificialneural networks such as back propagation networks, discriminant analyses (e.g., Bayesian classifier or Fischer analysis), logistic classifiers, and support vector classifiers (support vector machines).
[0220] In embodiments, a supervised classification method is a recursive partitioning process. Recursive partitioning processes use recursive partitioning trees to classify data derived from unknown samples. Further details about recursive partitioning processes are provided in U.S. Patent Application Publication No. 2002 / 0138208 Al to Paulse etal., “Method for analyzing mass spectra.”
[0221] In other embodiments, the classification models that are created can be formed using unsupervised learning methods. Unsupervised classification attempts to learn classifications based on similarities in the training data set, without pre-classifying the spectra from which the training data set was derived. Unsupervised learning methods include cluster analyses. A cluster analysis attempts to divide the data into “clusters” or groups that ideally should have members that are very similar to each other, and very dissimilar to members of other clusters. Similarity is then measured using some distance metric, which measures the distance between data items, and clusters together data items that are closer to each other. Clustering techniques include the MacQueen’s K-means algorithm and the Kohonen’s Self-Organizing Map algorithm.
[0222] Learning algorithms asserted for use in classifying biological information are described, for example, in International Publication No. WO 01 / 31580 (Barnhill etal., “Methods and devices for identifying patterns in biological systems and methods of use thereof’), U.S. Patent Application Publication No. 2002 / 0193950 Al (Gavin etal., “Method or analyzing mass spectra”), U.S. Patent Application Publication No. 2003 / 0004402 Al (Hitt et al., “Process for discriminating between biological states based on hidden patterns from biological data”), and U.S. Patent Application Publication No. 2003 / 0055615 Al (Zhang and Zhang, “Systems and methods for processing biological expression data”).
[0223] The classification models can be formed on and used on any suitable digital computer. Suitable digital computers include micro, mini, or large computers using any standard or specialized operating system, such as a Unix, Windows™ or Linux™ based operating system. The digital computer that is used may be physically separate from an instrument used to generate data of interest, or it may be coupled to the instrument.
[0224] The training data set and the classification models according to embodiments of the present disclosure can be embodied by computer code that is executed or used by a digital computer. The computer code can be stored on any suitable computer readable media including optical ormagnetic disks, sticks, tapes, etc., and can be written in any suitable computer programming language including C, C++, visual basic, etc.Clinical Classifier Scoring Algorithm
[0225] As an example, a neural net classifier was developed to prospectively identify DLBCL patients with the respective genetic signatures. The exemplified classifier utilizes 21 metafeatures (i.e., input parameters / variables), which are listed in Table 2 of the Examples below, that were selected based on biological significance for each DLBCL genetic cluster. Each metafeature contains one or more of the variants listed in Table 2. A neural network using these 21 metafeatures (each containing structural variants (SVs, which include translocations), somatic copy number alterations (SCNAs), and / or mutations) provided over 95% accuracy in assignment to classes at high confidence (e.g., > 0.7 or > 0.9) (FIG. 7A-E). The output of the classifier can include a probability associated with a DLBCL falling within one of the five classes described herein (i.e., C1-C5).
[0226] The neural net classifier may be trained using the NMF labels for C1-C5 DLBCL as “ground truth” for C1-C5 DLBCL assignment. The combined clustered dataset may be separated into training / validation and testing sets, e.g., at an 80-20 split, 90-10, 95-5, or other split for training and testing. The neural net classifier may be constructed from an ensemble of predictors, with the final prediction calculated as the average output vector across all predicting models. The final neural net classifier may be produced according to a model curation to further optimize the classifier. Further optimization may be performed (e.g., by trimming of features by q value and feature genomic territory, determining if all data types were needed, and / or assessing potential improvement by adding of COO and ploidy as input features).
[0227] Each metafeature can be calculated as a sum of variant class-specific weighted values (see Table 1) determined based upon the measurement of each variant included in the metafeature. For example, the metafeature BCL6.C1 listed in Table 2 could range from 0 (no mutation and no SV) to 5 (non-synonymous mutation and SV). The particular values assigned to each mutation, SCNA, or SV measured need not be limiting. For example, in some instances the relative magnitude of the various types of alterations listed in Table 1 (e.g., small, medium, and large) can be preserved while altering the absolute value assigned. For example, in the case of mutations, “no mutation” may be assigned a value of 100, “silent mutation” may be assigned a value of “1000” and “non- synonymous mutation” may be assigned a value of 2000. An intention of the weights / values listed in Table 1 can be to provide a means for quantifying the relative magnitude or severity of a particular observed alteration relative to a reference sequence (e.g., a wild-type genome). For example, a non-synonymous mutation may be considered as being more severe than a silentmutation, which may be considered more severe than no mutation. Therefore, a non-synonymous mutation may be assigned a greater weighting value than a silent mutation and both of these mutations may be assigned greater weighing values than no mutation. Similarly, an SV is considered as more severe than no SV and greater alterations in copy number are more severe than lower copy number alterations. Relative magnitudes of values assigned to mutations can be preserved across types of mutations. For example, a high level copy number alteration may be assigned a value equivalent to a non-synonymous mutation, a silent mutation can be assigned a value lower than (e.g., 3-fold lower than) that assigned to an SV, etc. The particular values assigned to each alteration observed for each variant are not intended to be limiting.
[0228] A somatic copy number alteration can be a low level copy number alteration or a high level copy number alteration. Non-limiting examples of low-level copy number alterations include an increase or decrease in copy number of not more than about 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, or lOx. Non-limiting examples of high-level copy number alterations include an increase or decrease in copy number of about or at least about 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, l lx, 12x, 13x, 14x, 15x, 20x, 25x, 50x, 75x, lOOx, 200x, 250x, 500x, 750x, l,000x, 10,000x, or 100,000x.
[0229] It is expressly contemplated that a classifier of the instant disclosure can be used to link discrete genetic signatures, clinical outcome and specific targeted therapy in clinical trials and in practice. Specifically, it is contemplated that tumors of patients with DLBCL can be analyzed prospectively with an exemplified classifier or other classifier within the scope of the instant disclosure. The resulting cluster identifications are predictive of the likelihood of response to standard combination chemotherapy and suggest rational targeted therapies based on clusterspecific biology. It is further expressly contemplated that a classifier of the instant disclosure can also be applied retrospectively to archival tissue from patients on specific clinical trials or therapies.Clinical Classifier Computer Interface
[0230] In some embodiments, one or more user computing devices may be configured to present a DLBCL classification interface that represents the cluster identifications of the metafeatures of a gene sample matrix (GSM) array representing classes of variants of a sample. The user computing device(s) each at least include a computer-readable medium, such as a random-access memory (RAM) coupled to a processor or FLASH memory. In some embodiments, the processor may execute computer-executable program instructions stored in memory. In some embodiments, the processor may include a microprocessor, an ASIC, and / or a state machine. In some embodiments, the processor may include, or may be in communication with, media, for example computer-readable media, which stores instructions that, when executed by one or more processor,may cause the processor to perform one or more steps described herein. In some embodiments, examples of computer-readable media may include, but are not limited to, an electronic, optical, magnetic, or other storage or transmission device capable of providing a processor, such as the processor of a computing device, with computer-readable instructions. In some embodiments, other examples of suitable media may include, but are not limited to, a floppy disk, CD-ROM, DVD, magnetic disk, memory chip, ROM, RAM, an ASIC, a configured processor, all optical media, all magnetic tape or other magnetic media, or any other medium from which a computer processor can read instructions. Also, various other forms of computer-readable media may transmit or carry instructions to a computer, including a router, private or public network, or other transmission device or channel, both wired and wireless. In some embodiments, the instructions may comprise code from any computer-programming language, including, for example, C, C++, Visual Basic, Java, Python, Perl, JavaScript, and etc.
[0231] In some embodiments, the user computing devices may include one or more remote computing devices located remotely from a user computing device, where the one or more remote computing devices is configured to perform DLBCL classification for presentation via the DLBCL classification interface at the user computing device. Accordingly, the one or more remote computing devices may include any one or more of, without limitation, a server, a content delivery network, a mainframe, a cloud platform, a remotely located user computing device running a remote desktop on the user computing device, among others or any combination thereof.
[0232] In some embodiments, user computing devices may also comprise a number of external or internal devices such as a mouse, a CD-ROM, DVD, a physical or virtual keyboard, a display, or other input or output devices.
[0233] In some embodiments, the computing devices may include the user computing device, where the user computing device is configured to perform DLBCL classification locally on the user computing device for presentation via the DLBCL classification interface on a display of the user computing device. In some embodiments, examples of user computing devices may be any type of processor-based platforms that are connected to a network such as, without limitation, personal computers, digital assistants, personal digital assistants, smart phones, pagers, digital tablets, laptop computers, Internet appliances, and other processor-based devices. In some embodiments, user computing devices may be specifically programmed with one or more application programs in accordance with one or more principles / methodologies detailed herein. In some embodiments, user computing devices may operate on any operating system capable of supporting a browser or browser-enabled application, such as Microsoft™, Windows™, and / or Linux. In some embodiments, user computing devices shown may include, for example, personalcomputers executing a browser application program such as Microsoft Corporation's Internet Explorer™, Apple Computer, Inc.'s Safari™, Mozilla Firefox, and / or Opera.
[0234] In some embodiments, through the user computing devices, users may communicate over the exemplary network with each other and / or with other systems and / or devices coupled to the network. Server devices may include a processor as well as memory. In some embodiments, the server devices may be also coupled to the network. In some embodiments, one or more user computing devices may be mobile clients.
[0235] In some embodiments, at least one database may be in communication with the server devices. The database may be any type of database, including a database managed by a database management system (DBMS). In some embodiments, an exemplary DBMS-managed database may be specifically programmed as an engine that controls organization, storage, management, and / or retrieval of data in the respective database. In some embodiments, the exemplary DBMS- managed database may be specifically programmed to provide the ability to query, backup and replicate, enforce rules, provide security, compute, perform change and access logging, and / or automate optimization. In some embodiments, the exemplary DBMS-managed database may be chosen from Oracle database, IBM DB2, Adaptive Server Enterprise, FileMaker, Microsoft Access, Microsoft SQL Server, MySQL, PostgreSQL, and a NoSQL implementation. In some embodiments, the exemplary DBMS-managed database may be specifically programmed to define each respective schema of each database in the exemplary DBMS, according to a particular database model of the present disclosure which may include a hierarchical model, network model, relational model, object model, or some other suitable organization that may result in one or more applicable data structures that may include fields, records, files, and / or objects. In some embodiments, the exemplary DBMS-managed database may be specifically programmed to include metadata about the data that is stored.
[0236] In some embodiments, the server(s) may be configured to use the DLBCL classification model to produce the cluster identifications of the metafeatures. The server may provide instructions across the network to cause the user computing device(s) to present the DLBCL classification interface that represents the cluster identifications of the metafeatures of a gene sample matrix (GSM) array representing classes of variants of a sample. In some embodiments, to provide the DLBCL classification, the server may be specifically configured to operate in a cloud computing / architecture such as, but not limiting to: infrastructure as a service (laaS), platform as a service (PaaS), and / or software as a service (SaaS) using a web browser, mobile app, thin client, terminal emulator or other endpoint including the user computing device.
[0237] In some embodiments, the cluster identifications may be presented to the user via a DLBCL classification interface rendered on a display of the user computing device, e.g., via direct instruction of the processor of the user computing device or by instruction over the network by the server, or a combination thereof. The DLBCL classification interface may present, e.g., a heatmap including a visualization of variants and associated cluster identifications, confidence levels of each cluster identification, and gene mutations, a list / table of each sample / metafeature and the associated confidence levels for each possible cluster identification and / or a predicted cluster with the associated confidence level, a two-dimensional or three-dimensional representation of the metafeatures, among other visualizations. In some embodiments, the two-dimensional or three- dimensional representation of the metafeatures may include a label (e.g., via color coding, shape coding, text-based labeling, etc.) of each metafeature, where the label represents that associated cluster identification. Moreover, the two-dimensional or three-dimensional representation of the metafeatures may be produced via dimensionality reduction, such as by using a suitable dimensionality reduction model including uniform manifold approximation and projection (UMAP), principal component analysis (PCA), or other dimensionality reduction model or any combination thereof.
[0238] Accordingly, the DLBCL classification model may be implemented remotely from the user computing device(s), e.g., as portal to a website or cloud-based service, or may be implemented locally to each computing device(s).
[0239] In some embodiments, DLBCL classification, including the detection of alterations, classification, and / or clinical scoring, among other operations as detailed herein, may be performed via one or more software applications installed or loaded on the user computing device to use the DLBCL classification model to produce the cluster identifications of the metafeatures. The software application(s) may be one or more native software applications programmed in a programming language configured to run natively on the user computing device, such as, e.g., C™, C++™, C#™, Swift™, Kotlin™, Java™, Python™, among others or any combination thereof. The software application(s) may be one or more web-based software applications programmed in a programming language configured to run as a web application on the user computing device, such as, e.g., HTML, CCS, Javascript™, Rust™, among others or any combination thereof.
[0240] In some embodiments, running the DLBCL classification models in a data center, cloud platform, server, or other remote computing device(s) may provide speed and processing power at the expense of compute resources of the remote computing device(s) as well as potential leakage of data between systems and / or devices. Thus, to enable users to run the DLBCL classificationmodels locally, thereby protecting data security and confidentiality while freeing resources of the remote computing device(s), an exemplary software program may be distributed, e.g., via physical media and / or by download over a network such as the Internet, and may be provided to users for local installation. The user may be permissioned or otherwise provided with a copy of the installable software program based on a type of user (e.g., education, academic, student, non-profit, personal use, etc.).
[0241] The software program may include a stand-alone program including an underlying codebase for performing some or all functions of the DLBCL classification within a simple-to-use graphical user interface (GUI). The GUI allows the user to load their own data in suitable gene sample data representations, such as in the form of a gene sample matrix (GSM). The DLBCL classification model may ingest the gene sample data and calculate probabilities for each possible DLBCL class. As detailed above, the DLBCL classes may include 1, 2, 3, 4, 5 or more classes, such as the five classes of C1-C5.
[0242] In some embodiments, to prepare the user’s gene sample data for classification without the benefit of the large datasets resident in a data center, the software program may preprocess the gene sample data. For example, the software program may place blank or “not applicable (“N / A”)” fields with pre-established mean values for the associated gene(s). Thus, the software program may be installed with the pre-established mean values. Moreover, such mean values may be updated via software updates in later versions and / or pushed via a server and / or web call, e.g., upon user request and / or with a notification notifying the user, via the GUI, of the updated mean values.
[0243] Alternatively, or additionally, pre-analyzed datasets may be provided with the software application as part of an installation package including in software version updates, and / or downloaded after installation. The datasets may enable analysis and visualization of the user’s gene sample data in the context of existing datasets, which may be presented alongside the user’s gene sample data in interface elements such as the table(s), heatmap(s) and / or dimensionality reduction plot(s).
[0244] The software program may output, via the GUI, the input GSM and corresponding DLBCL classes via one or more data visualizations, such as, without limitation, as one or more tables, as one or more heatmaps, and / or as one or more dimensionality reduction plots (e.g., a UMAP plot). In some embodiments, each data visualization may be on separate tabs within the GUI or other form of section within the GUI (e.g., pages, screens, windows, or others or any combination thereof). Thus, the user may compare DLBCL classifications results from their own gene sample data to DLBCL classification results from the publication or other samples.
[0245] In some embodiments, the software program may store the DLBCL classifications and / or data visualizations in memory, and / or enable the user to download the DLBCL classifications and / or data visualizations. For example, the software program may store and / or download the table(s) as a tab delimited file, comma delimited file, Excel™ spreadsheet, or other file or represented tabulated data or any combination thereof. In the same or different example, the software program may store and / or download the heatmap(s), e.g., as an image (including, but not limited to, a vector graphics format, raster graphics format, PDF, or other graphical format file or any combination thereof). Similarly, in a same or different example, the software program may store and / or download the dimensionality reduction plot(s) as an image (including, but not limited to, a vector graphics format, raster graphics format, PDF, or other graphical format file or any combination thereof). Thus, the locally installable software program easily allows non- computational researchers to explore DLBCL classifications without a technical learning curve.Bait Sets
[0246] Provided herein are bait sets (e.g., sets of oligonucleotide probes) for characterization of variants in a biological sample (e.g., a biological sample containing circulating tumor DNA). The bait sets can be used in combination with the method for DLBCL classification described herein to provide a targeted sequencing assay for use in the clinic. The bait sets can comprise oligonucleotide sequences targeting structural variants (SVs) including translocations, somatic copy number alterations (SCNAs), and mutations. The bait sets can contain primer sequences allowing for targeted sequencing of a sample or for preparation of an amplicon(s) from a sample. In embodiments, the bait sets make up part of a targeted sequencing panel. Methods for design and manufacture of a targeted sequencing panel are known in the art (see, e.g., Moorthie, etal. “Review of massively parallel DNA sequencing technologies”, The HUGO Journal, 5: 1-12 (2011)). The targeted sequencing panel can be hybridization capture-based, circularization-based, or amplicon sequencing-based. The bait sets can be used to prepare one of the biochips described above.
[0247] The set of polynucleotides (i.e., bait set) can include all or a sub-set of polynucleotides identified as targeting a particular variant. The set of polynucleotides can include all or a sub-set of polynucleotides targeting a particular variant class and / or variant (i.e., target) type. The set of polynucleotides can include all or a sub-set of polynucleotides targeting all of the variants listed in Table 2.
[0248] The bait sets can be used to determine tumor mutational burden in a subject or for quantifying levels of circulating tumor DNA in a subject. The bait sets capture alterations necessary and sufficient to classify DLBCLs as belonging to C1-C5. The bait sets can also be used to measure microsatellite instability and / or tumor mutational burden.List of Sequence Descriptions:
[0249] The following list provides descriptions of the sequences provided in the Sequence Listing, where in the list each SEQ ID NO is followed by the name of the targeted genetic locus, which is followed by a description of the target type in parentheses: SEQ ID NO: 1, lp36.32_DLBCL (FOCAL); SEQ ID NO: 2, lp36.32_DLBCL (FOCAL); SEQ ID NO: 3, lp36.32_DLBCL (FOCAL); SEQ ID NO: 4, lp36.32_DLBCL (FOCAL); SEQ ID NO: 5, lp36.32_DLBCL (FOCAL); SEQ ID NO: 6, GNB1 (Gene); SEQ ID NO: 7, GNB1 (Gene); SEQ ID NO: 8, GNB1 (Gene); SEQ ID NO: 9, GNB1 (Gene); SEQ ID NO: 10, GNB1 (Gene); SEQ ID NO: 11, GNB1 (Gene); SEQ ID NO: 12, GNB1 (Gene); SEQ ID NO: 13, GNB1 (Gene); SEQ ID NO: 14, GNB1 (Gene); SEQ ID NO: 15, lp36.32_DLBCL (FOCAL); SEQ ID NO: 16, lp36.32_DLBCL (FOCAL); SEQ ID NO: 17, lp36.32_DLBCL (FOCAL); SEQ ID NO: 18, TNFRSF14 (Gene); SEQ ID NO: 19, TNFRSF14 (Gene); SEQ ID NO: 20, TNFRSF14 (Gene); SEQ ID NO: 21, TNFRSF14 (Gene); SEQ ID NO: 22, TNFRSF14 (Gene); SEQ ID NO: 23, TNFRSF14 (Gene); SEQ ID NO: 24, TNFRSF14 (Gene); SEQ ID NO: 25, TNFRSF14 (Gene); SEQ ID NO: 26, lp36.32_DLBCL (FOCAL); SEQ ID NO: 27, lp36.32_DLBCL (FOCAL); SEQ ID NO: 28, lp36.32_DLBCL (FOCAL); SEQ ID NO: 29, lp36.32_DLBCL (FOCAL); SEQ ID NO: 30, lp36.32_DLBCL (FOCAL); SEQ ID NO: 31, lp36.32_DLBCL (FOCAL); SEQ ID NO: 32, FP (FP); SEQ ID NO: 33, IP (ARM); SEQ ID NO: 34, IP (ARM); SEQ ID NO: 35, MSI (MSI); SEQ ID NO: 36, MSI (MSI); SEQ ID NO: 37, IP (ARM); SEQ ID NO: 38, IP (ARM); SEQ ID NO: 39, PIK3CD (Gene); SEQ ID NO: 40, PIK3CD (Gene); SEQ ID NO: 41, PIK3CD (Gene); SEQ ID NO: 42, PIK3CD (Gene); SEQ ID NO: 43, PIK3CD (Gene); SEQ ID NO: 44, PIK3CD (Gene); SEQ ID NO: 45, PIK3CD (Gene); SEQ ID NO: 46, PIK3CD (Gene); SEQ ID NO: 47, PIK3CD (Gene); SEQ ID NO: 48, PIK3CD (Gene); SEQ ID NO: 49, PIK3CD (Gene); SEQ ID NO: 50, PIK3CD (Gene); SEQ ID NO: 51, PIK3CD (Gene); SEQ ID NO: 52, PIK3CD (Gene); SEQ ID NO: 53, PIK3CD (Gene); SEQ ID NO: 54, PIK3CD (Gene); SEQ ID NO: 55, PIK3CD (Gene); SEQ ID NO: 56, PIK3CD (Gene); SEQ ID NO: 57, PIK3CD (Gene); SEQ ID NO: 58, PIK3CD (Gene); SEQ ID NO: 59, PIK3CD (Gene); SEQ ID NO: 60, PIK3CD (Gene); SEQ ID NO: 61, IP (ARM); SEQ ID NO: 62, MTOR (Gene); SEQ ID NO: 63, MTOR (Gene); SEQ ID NO: 64, MTOR (Gene); SEQ ID NO: 65, MTOR (Gene); SEQ ID NO: 66, MTOR (Gene); SEQ ID NO: 67, MTOR (Gene); SEQ ID NO: 68, MTOR (Gene); SEQ ID NO: 69, MTOR (Gene); SEQ ID NO: 70, MTOR (Gene); SEQ ID NO: 71, MTOR (Gene); SEQ ID NO: 72, MTOR (Gene); SEQ ID NO: 73, MTOR (Gene); SEQ ID NO: 74, MTOR (Gene); SEQ ID NO: 75, MTOR (Gene); SEQ ID NO: 76, MTOR (Gene); SEQ ID NO: 77, MTOR (Gene); SEQ ID NO: 78, MTOR (Gene); SEQ ID NO: 79, MTOR (Gene); SEQ ID NO: 80, MTOR (Gene); SEQ ID NO: 81, MTOR (Gene); SEQ ID NO: 82, MTOR(Gene); SEQ ID NO: 83, MTOR (Gene); SEQ ID NO: 84, MTOR (Gene); SEQ ID NO: 85, MTOR (Gene); SEQ ID NO: 86, MTOR (Gene); SEQ ID NO: 87, MTOR (Gene); SEQ ID NO: 88, MTOR (Gene); SEQ ID NO: 89, MTOR (Gene); SEQ ID NO: 90, MTOR (Gene); SEQ ID NO: 91, MTOR (Gene); SEQ ID NO: 92, MTOR (Gene); SEQ ID NO: 93, MTOR (Gene); SEQ ID NO: 94, MTOR (Gene); SEQ ID NO: 95, MTOR (Gene); SEQ ID NO: 96, MTOR (Gene); SEQ ID NO: 97, MTOR (Gene); SEQ ID NO: 98, MTOR (Gene); SEQ ID NO: 99, MTOR (Gene); SEQ ID NO: 100, MTOR (Gene); SEQ ID NO: 101, MTOR (Gene); SEQ ID NO: 102, MTOR (Gene); SEQ ID NO: 103, MTOR (Gene); SEQ ID NO: 104, MTOR (Gene); SEQ ID NO: 105, MTOR (Gene); SEQ ID NO: 106, MTOR (Gene); SEQ ID NO: 107, MTOR (Gene); SEQ ID NO: 108, MTOR (Gene); SEQ ID NO: 109, MTOR (Gene); SEQ ID NO: 110, MTOR (Gene); SEQ ID NO: 111, MTOR (Gene); SEQ ID NO: 112, MTOR (Gene); SEQ ID NO: 113, MTOR (Gene); SEQ ID NO: 114, MTOR (Gene); SEQ ID NO: 115, MTOR (Gene); SEQ ID NO: 116, MTOR (Gene); SEQ ID NO: 117, MTOR (Gene); SEQ ID NO: 118, MTOR (Gene); SEQ ID NO: 119, MTOR (Gene); SEQ ID NO: 120, MTOR (Gene); SEQ ID NO: 121, MTOR (Gene); SEQ ID NO: 122, IP (ARM); SEQ ID NO: 123, TNFRSF8 (Gene); SEQ ID NO: 124, TNFRSF8 (Gene); SEQ ID NO: 125, TNFRSF8 (Gene); SEQ ID NO: 126, TNFRSF8 (Gene); SEQ ID NO: 127, TNFRSF8 (Gene); SEQ ID NO: 128, TNFRSF8 (Gene); SEQ ID NO: 129, TNFRSF8 (Gene); SEQ ID NO: 130, TNFRSF8 (Gene); SEQ ID NO: 131, TNFRSF8 (Gene); SEQ ID NO: 132, TNFRSF8 (Gene); SEQ ID NO: 133, TNFRSF8 (Gene); SEQ ID NO: 134, TNFRSF8 (Gene); SEQ ID NO: 135, TNFRSF8 (Gene); SEQ ID NO: 136, TNFRSF8 (Gene); SEQ ID NO: 137, TNFRSF8 (Gene); SEQ ID NO: 138, TNFRSF8 (Gene); SEQ ID NO: 139, FP (FP); SEQ ID NO: 140, IP (ARM); SEQ ID NO: 141, IP (ARM); SEQ ID NO: 142, FP (FP); SEQ ID NO: 143, IP (ARM); SEQ ID NO: 144, SPEN (Gene); SEQ ID NO: 145, SPEN (Gene); SEQ ID NO: 146, SPEN (Gene); SEQ ID NO: 147, SPEN (Gene); SEQ ID NO: 148, SPEN (Gene); SEQ ID NO: 149, SPEN (Gene); SEQ ID NO: 150, SPEN (Gene); SEQ ID NO: 151, SPEN (Gene); SEQ ID NO: 152, SPEN (Gene); SEQ ID NO: 153, SPEN (Gene); SEQ ID NO: 154, SPEN (Gene); SEQ ID NO: 155, SPEN (Gene); SEQ ID NO: 156, SPEN (Gene); SEQ ID NO: 157, SPEN (Gene); SEQ ID NO: 158, SPEN (Gene); SEQ ID NO: 159, SPEN (Gene); SEQ ID NO: 160, SPEN (Gene); SEQ ID NO: 161, SPEN (Gene); SEQ ID NO: 162, SPEN (Gene); SEQ ID NO: 163, SPEN (Gene); SEQ ID NO: 164, SPEN (Gene); SEQ ID NO: 165, SPEN (Gene); SEQ ID NO: 166, SPEN (Gene); SEQ ID NO: 167, SPEN (Gene); SEQ ID NO: 168, SPEN (Gene); SEQ ID NO: 169, SPEN (Gene); SEQ ID NO: 170, SPEN (Gene); SEQ ID NO: 171, SPEN (Gene); SEQ ID NO: 172, SPEN (Gene); SEQ ID NO: 173, SPEN (Gene); SEQ ID NO: 174, SPEN (Gene); SEQ ID NO: 175, SPEN (Gene); SEQ ID NO: 176, SPEN (Gene); SEQ ID NO: 177, SPEN (Gene); SEQ ID NO: 178, IP (ARM); SEQ ID NO: 179, IP (ARM); SEQID NO: 180, IP (ARM); SEQ ID NO: 181, IP (ARM); SEQ ID NO: 182, MSI (MSI); SEQ ID NO: 183, MSI (MSI); SEQ ID NO: 184, IP (ARM); SEQ ID NO: 185, lp36.11 DLBCL (FOCAL); SEQ ID NO: 186, lp36.11_DLBCL (FOCAL); SEQ ID NO: 187, lp36.11 DLBCL (FOCAL); SEQ ID NO: 188, lp36.11_DLBCL (FOCAL); SEQ ID NO: 189, lp36.11 DLBCL (FOCAL); SEQ ID NO: 190, ARID 1 A (Gene); SEQ ID NO: 191, ARID 1 A (Gene); SEQ ID NO: 192, ARID 1 A (Gene); SEQ ID NO: 193, ARID 1 A (Gene); SEQ ID NO: 194, ARID 1 A (Gene); SEQ ID NO: 195, ARID1A (Gene); SEQ ID NO: 196, ARID1A (Gene); SEQ ID NO: 197, ARID 1 A (Gene); SEQ ID NO: 198, ARID 1 A (Gene); SEQ ID NO: 199, ARID 1 A (Gene); SEQ ID NO: 200, ARID 1 A (Gene); SEQ ID NO: 201, ARID 1 A (Gene); SEQ ID NO: 202, ARID 1 A (Gene); SEQ ID NO: 203, ARID 1 A (Gene); SEQ ID NO: 204, ARID 1 A (Gene); SEQ ID NO: 205, ARID1A (Gene); SEQ ID NO: 206, ARID1A (Gene); SEQ ID NO: 207, ARID1A (Gene); SEQ ID NO: 208, ARID 1 A (Gene); SEQ ID NO: 209, ARID 1 A (Gene); SEQ ID NO: 210, ARID1A (Gene); SEQ ID NO: 211, ARID1A (Gene); SEQ ID NO: 212, ARID1A (Gene); SEQ ID NO: 213, ARID1A (Gene); SEQ ID NO: 214, ARID1A (Gene); SEQ ID NO: 215, ARID1A (Gene); SEQ ID NO: 216, ARID1A (Gene); SEQ ID NO: 217, lp36.11 DLBCL (FOCAL); SEQ ID NO: 218, lp36.11 DLBCL (FOCAL); SEQ ID NO: 219, MSI (MSI); SEQ ID NO: 220, MSI (MSI); SEQ ID NO: 221, lp36.11_DLBCL (FOCAL); SEQ ID NO: 222, lp36.11 DLBCL (FOCAL); SEQ ID NO: 223, lp36.11 DLBCL (FOCAL); SEQ ID NO: 224, IP (ARM); SEQ ID NO: 225, IP (ARM); SEQ ID NO: 226, IP (ARM); SEQ ID NO: 227, IP (ARM); SEQ ID NO: 228, IP (ARM); SEQ ID NO: 229, IP (ARM); SEQ ID NO: 230, IP (ARM); SEQ ID NO: 231, IP (ARM); SEQ ID NO: 232, IP (ARM); SEQ ID NO: 233, MSI (MSI); SEQ ID NO: 234, MSI (MSI); SEQ ID NO: 235, IP (ARM); SEQ ID NO: 236, ZC3H12A (Gene); SEQ ID NO: 237, ZC3H12A (Gene); SEQ ID NO: 238, ZC3H12A (Gene); SEQ ID NO: 239, ZC3H12A (Gene); SEQ ID NO: 240, ZC3H12A (Gene); SEQ ID NO: 241, ZC3H12A (Gene); SEQ ID NO: 242, IP (ARM); SEQ ID NO: 243, IP (ARM); SEQ ID NO: 244, IP (ARM); SEQ ID NO: 245, IP (ARM); SEQ ID NO: 246, IP (ARM); SEQ ID NO: 247, IP (ARM); SEQ ID NO: 248, IP (ARM); SEQ ID NO: 249, IP (ARM); SEQ ID NO: 250, IP (ARM); SEQ ID NO: 251, IP (ARM); SEQ ID NO: 252, IP (ARM); SEQ ID NO: 253, IP (ARM); SEQ ID NO: 254, IP (ARM); SEQ ID NO: 255, IP (ARM); SEQ ID NO: 256, IP (ARM); SEQ ID NO: 257, IP (ARM); SEQ ID NO: 258, IP (ARM); SEQ ID NO: 259, IP (ARM); SEQ ID NO: 260, IP (ARM); SEQ ID NO: 261, IP (ARM); SEQ ID NO: 262, IP (ARM); SEQ ID NO: 263, JAK1 (Gene); SEQ ID NO: 264, JAK1 (Gene);SEQ ID NO: 265, JAK1 (Gene); SEQ ID NO: 266, JAK1 (Gene); SEQ ID NO: 267, JAK1 (Gene);SEQ ID NO: 268, JAK1 (Gene); SEQ ID NO: 269, JAK1 (Gene); SEQ ID NO: 270, JAK1 (Gene);SEQ ID NO: 271, JAK1 (Gene); SEQ ID NO: 272, JAK1 (Gene); SEQ ID NO: 273, JAK1 (Gene);SEQ ID NO: 274, JAK1 (Gene); SEQ ID NO: 275, JAK1 (Gene); SEQ ID NO: 276, JAK1 (Gene);SEQ ID NO: 277, JAK1 (Gene); SEQ ID NO: 278, JAK1 (Gene); SEQ ID NO: 279, JAK1 (Gene);SEQ ID NO: 280, JAK1 (Gene); SEQ ID NO: 281, JAK1 (Gene); SEQ ID NO: 282, JAK1 (Gene);SEQ ID NO: 283, JAK1 (Gene); SEQ ID NO: 284, JAK1 (Gene); SEQ ID NO: 285, JAK1 (Gene);SEQ ID NO: 286, JAK1 (Gene); SEQ ID NO: 287, IP (ARM); SEQ ID NO: 288, IP (ARM); SEQ ID NO: 289, IP (ARM); SEQ ID NO: 290, IP (ARM); SEQ ID NO: 291, IP (ARM); SEQ ID NO: 292, IP (ARM); SEQ ID NO: 293, IP (ARM); SEQ ID NO: 294, IP (ARM); SEQ ID NO: 295, lp31.1_DLBCL (FOCAL); SEQ ID NO: 296, lp31.1_DLBCL (FOCAL); SEQ ID NO: 297, lp31.1_DLBCL (FOCAL); SEQ ID NO: 298, lp31.1_DLBCL (FOCAL); SEQ ID NO: 299, lp31.1_DLBCL (FOCAL); SEQ ID NO: 300, lp31.1_DLBCL (FOCAL); SEQ ID NO: 301, lp31.1_DLBCL (FOCAL); SEQ ID NO: 302, lp31.1_DLBCL (FOCAL); SEQ ID NO: 303, lp31.1_DLBCL (FOCAL); SEQ ID NO: 304, lp31.1_DLBCL (FOCAL); SEQ ID NO: 305, lp31.1_DLBCL (FOCAL); SEQ ID NO: 306, lp31.1_DLBCL (FOCAL); SEQ ID NO: 307, lp31.1_DLBCL (FOCAL); SEQ ID NO: 308, lp31.1_DLBCL (FOCAL); SEQ ID NO: 309, lp31.1_DLBCL (FOCAL); SEQ ID NO: 310, lp31.1_DLBCL (FOCAL); SEQ ID NO: 311, lp31.1_DLBCL (FOCAL); SEQ ID NO: 312, lp31.1_DLBCL (FOCAL); SEQ ID NO: 313, lp31.1_DLBCL (FOCAL); SEQ ID NO: 314, lp31.1_DLBCL (FOCAL); SEQ ID NO: 315, lp31.1_DLBCL (FOCAL); SEQ ID NO: 316, lp31.1_DLBCL (FOCAL); SEQ ID NO: 317, lp31.1_DLBCL (FOCAL); SEQ ID NO: 318, lp31.1_DLBCL (FOCAL); SEQ ID NO: 319, lp31.1_DLBCL (FOCAL); SEQ ID NO: 320, lp31.1_DLBCL (FOCAL); SEQ ID NO: 321, lp31.1_DLBCL (FOCAL); SEQ ID NO: 322, lp31.1 DLBCL (FOCAL); SEQ ID NO: 323, lp21- p22_MCL (FOCAL); SEQ ID NO: 324, lp21-p22_MCL (FOCAL); SEQ ID NO: 325, lp21- p22_MCL (FOCAL); SEQ ID NO: 326, BCL10 (Gene); SEQ ID NO: 327, BCL10 (Gene); SEQ ID NO: 328, BCL10 (Gene); SEQ ID NO: 329, lp21-p22_MCL (FOCAL); SEQ ID NO: 330, lp21-p22_MCL (FOCAL); SEQ ID NO: 331, lp21-p22_MCL (FOCAL); SEQ ID NO: 332, lp21- p22_MCL (FOCAL); SEQ ID NO: 333, lp21-p22_MCL (FOCAL); SEQ ID NO: 334, lp21- p22_MCL (FOCAL); SEQ ID NO: 335, lp21-p22_MCL (FOCAL); SEQ ID NO: 336, lp21- p22_MCL (FOCAL); SEQ ID NO: 337, lp21-p22_MCL (FOCAL); SEQ ID NO: 338, lp21- p22_MCL (FOCAL); SEQ ID NO: 339, lp21-p22_MCL (FOCAL); SEQ ID NO: 340, lp21- p22_MCL (FOCAL); SEQ ID NO: 341, lp21-p22_MCL (FOCAL); SEQ ID NO: 342, lp21- p22_MCL (FOCAL); SEQ ID NO: 343, lp21-p22_MCL (FOCAL); SEQ ID NO: 344, lp21- p22_MCL (FOCAL); SEQ ID NO: 345, lp21-p22_MCL (FOCAL); SEQ ID NO: 346, lp21- p22_MCL (FOCAL); SEQ ID NO: 347, lp21-p22_MCL (FOCAL); SEQ ID NO: 348, lp21- p22_MCL (FOCAL); SEQ ID NO: 349, lp21-p22_MCL (FOCAL); SEQ ID NO: 350, lp21-p22_MCL (FOCAL); SEQ ID NO: 351, lp21-p22_MCL (FOCAL); SEQ ID NO: 352, lp21- p22_MCL (FOCAL); SEQ ID NO: 353, lp21-p22_MCL (FOCAL); SEQ ID NO: 354, lp21- p22_MCL (FOCAL); SEQ ID NO: 355, lp21-p22_MCL (FOCAL); SEQ ID NO: 356, lp21- p22_MCL (FOCAL); SEQ ID NO: 357, lp21-p22_MCL (FOCAL); SEQ ID NO: 358, lp21- p22_MCL (FOCAL); SEQ ID NO: 359, lp21-p22_MCL (FOCAL); SEQ ID NO: 360, lp21- p22_MCL (FOCAL); SEQ ID NO: 361, lp21-p22_MCL (FOCAL); SEQ ID NO: 362, lp21- p22_MCL (FOCAL); SEQ ID NO: 363, lp21-p22_MCL (FOCAL); SEQ ID NO: 364, lp21- p22_MCL (FOCAL); SEQ ID NO: 365, lp21-p22_MCL (FOCAL); SEQ ID NO: 366, lp21- p22_MCL (FOCAL); SEQ ID NO: 367, MSI (MSI); SEQ ID NO: 368, MSI (MSI); SEQ ID NO: 369, lp21-p22_MCL (FOCAL); SEQ ID NO: 370, lp21-p22_MCL (FOCAL); SEQ ID NO: 371, lp21-p22_MCL (FOCAL); SEQ ID NO: 372, lp21-p22_MCL (FOCAL); SEQ ID NO: 373, lp21- p22_MCL (FOCAL); SEQ ID NO: 374, lp21-p22_MCL (FOCAL); SEQ ID NO: 375, lp21- p22_MCL (FOCAL); SEQ ID NO: 376, lp21-p22_MCL (FOCAL); SEQ ID NO: 377, lp21- p22_MCL (FOCAL); SEQ ID NO: 378, lp21-p22_MCL (FOCAL); SEQ ID NO: 379, lp21- p22_MCL (FOCAL); SEQ ID NO: 380, lp21-p22_MCL (FOCAL); SEQ ID NO: 381, lp21- p22_MCL (FOCAL); SEQ ID NO: 382, lp21-p22_MCL (FOCAL); SEQ ID NO: 383, lp21- p22_MCL (FOCAL); SEQ ID NO: 384, lp21-p22_MCL (FOCAL); SEQ ID NO: 385, lp21- p22_MCL (FOCAL); SEQ ID NO: 386, lp21-p22_MCL (FOCAL); SEQ ID NO: 387, lp21- p22_MCL (FOCAL); SEQ ID NO: 388, lp21-p22_MCL (FOCAL); SEQ ID NO: 389, lp21- p22_MCL (FOCAL); SEQ ID NO: 390, lp21-p22_MCL (FOCAL); SEQ ID NO: 391, lp21- p22_MCL (FOCAL); SEQ ID NO: 392, lp21-p22_MCL (FOCAL); SEQ ID NO: 393, lp21- p22_MCL (FOCAL); SEQ ID NO: 394, lp21-p22_MCL (FOCAL); SEQ ID NO: 395, lp21- p22_MCL (FOCAL); SEQ ID NO: 396, lp21-p22_MCL (FOCAL); SEQ ID NO: 397, lp21- p22_MCL (FOCAL); SEQ ID NO: 398, lp21-p22_MCL (FOCAL); SEQ ID NO: 399, lp21- p22_MCL (FOCAL); SEQ ID NO: 400, lp21-p22_MCL (FOCAL); SEQ ID NO: 401, lp21- p22_MCL (FOCAL); SEQ ID NO: 402, lp21-p22_MCL (FOCAL); SEQ ID NO: 403, lp21- p22_MCL (FOCAL); SEQ ID NO: 404, lp21-p22_MCL (FOCAL); SEQ ID NO: 405, lp21- p22_MCL (FOCAL); SEQ ID NO: 406, lp21-p22_MCL (FOCAL); SEQ ID NO: 407, lp21- p22_MCL (FOCAL); SEQ ID NO: 408, lp21-p22_MCL (FOCAL); SEQ ID NO: 409, lp21- p22_MCL (FOCAL); SEQ ID NO: 410, lp21-p22_MCL (FOCAL); SEQ ID NO: 411, lp21- p22_MCL (FOCAL); SEQ ID NO: 412, lp21-p22_MCL (FOCAL); SEQ ID NO: 413, lp21- p22_MCL (FOCAL); SEQ ID NO: 414, lp21-p22_MCL (FOCAL); SEQ ID NO: 415, lp21- p22_MCL (FOCAL); SEQ ID NO: 416, lp21-p22_MCL (FOCAL); SEQ ID NO: 417, lp21- p22_MCL (FOCAL); SEQ ID NO: 418, lp21-p22_MCL (FOCAL); SEQ ID NO: 419, lp21-p22_MCL (FOCAL); SEQ ID NO: 420, lp21-p22_MCL (FOCAL); SEQ ID NO: 421, IP (ARM); SEQ ID NO: 422, IP (ARM); SEQ ID NO: 423, IP (ARM); SEQ ID NO: 424, IP (ARM); SEQ ID NO: 425, IP (ARM); SEQ ID NO: 426, IP (ARM); SEQ ID NO: 427, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 428, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 429, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 430, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 431, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 432, NRAS (Gene); SEQ ID NO: 433, NRAS (Gene); SEQ ID NO: 434, NRAS (Gene); SEQ ID NO: 435, NRAS (Gene); SEQ ID NO: 436, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 437, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 438, lpl3.1_DLB CL (FOCAL); SEQ ID NO: 439, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 440, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 441, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 442, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 443, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 444, CD58 (Gene); SEQ ID NO: 445, CD58 (Gene); SEQ ID NO: 446, CD58 (Gene); SEQ ID NO: 447, CD58 (Gene); SEQ ID NO: 448, CD58 (Gene); SEQ ID NO: 449, CD58 (Gene); SEQ ID NO: 450, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 451, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 452, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 453, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 454, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 455, lpl3.1_DLBCL (FOCAL); SEQ ID NO: 456, IP (ARM); SEQ ID NO: 457, NOTCH2 (Gene); SEQ ID NO: 458, NOTCH2 (Gene); SEQ ID NO: 459, NOTCH2 (Gene); SEQ ID NO: 460, NOTCH2 (Gene); SEQ ID NO: 461, NOTCH2 (Gene); SEQ ID NO: 462, NOTCH2 (Gene); SEQ ID NO: 463, NOTCH2 (Gene); SEQ ID NO: 464, NOTCH2 (Gene); SEQ ID NO: 465, NOTCH2 (Gene); SEQ ID NO: 466, NOTCH2 (Gene); SEQ ID NO: 467, NOTCH2 (Gene); SEQ ID NO: 468, NOTCH2 (Gene); SEQ ID NO: 469, NOTCH2 (Gene); SEQ ID NO: 470, NOTCH2 (Gene); SEQ ID NO: 471, NOTCH2 (Gene); SEQ ID NO: 472, NOTCH2 (Gene); SEQ ID NO: 473, NOTCH2 (Gene); SEQ ID NO: 474, NOTCH2 (Gene); SEQ ID NO: 475, NOTCH2 (Gene); SEQ ID NO: 476, NOTCH2 (Gene); SEQ ID NO: 477, NOTCH2 (Gene); SEQ ID NO: 478, NOTCH2 (Gene); SEQ ID NO: 479, NOTCH2 (Gene); SEQ ID NO: 480, NOTCH2 (Gene); SEQ ID NO: 481, NOTCH2 (Gene); SEQ ID NO: 482, NOTCH2 (Gene); SEQ ID NO: 483, NOTCH2 (Gene); SEQ ID NO: 484, NOTCH2 (Gene); SEQ ID NO: 485, NOTCH2 (Gene); SEQ ID NO: 486, NOTCH2 (Gene); SEQ ID NO: 487, NOTCH2 (Gene); SEQ ID NO: 488, NOTCH2 (Gene); SEQ ID NO: 489, NOTCH2 (Gene); SEQ ID NO: 490, NOTCH2 (Gene); SEQ ID NO: 491, NOTCH2 (Gene); SEQ ID NO: 492, NOTCH2 (Gene); SEQ ID NO: 493, NOTCH2 (Gene); SEQ ID NO: 494, lq23.3_DLBCL (FOCAL); SEQ ID NO: 495, PDE4DIP (Gene); SEQ ID NO: 496, PDE4DIP (Gene); SEQ ID NO: 497, PDE4DIP (Gene); SEQ ID NO: 498, PDE4DIP (Gene); SEQ ID NO: 499, PDE4DIP (Gene); SEQ ID NO: 500, PDE4DIP (Gene); SEQ ID NO: 501, PDE4DIP (Gene); SEQ ID NO: 502, PDE4DIP (Gene); SEQ ID NO: 503, PDE4DIP (Gene); SEQ ID NO: 504,PDE4DIP (Gene); SEQ ID NO: 505, PDE4DIP (Gene); SEQ ID NO: 506, PDE4DIP (Gene); SEQ ID NO: 507, PDE4DIP (Gene); SEQ ID NO: 508, PDE4DIP (Gene); SEQ ID NO: 509, PDE4DIP (Gene); SEQ ID NO: 510, PDE4DIP (Gene); SEQ ID NO: 511, PDE4DIP (Gene); SEQ ID NO: 512, PDE4DIP (Gene); SEQ ID NO: 513, PDE4DIP (Gene); SEQ ID NO: 514, PDE4DIP (Gene); SEQ ID NO: 515, PDE4DIP (Gene); SEQ ID NO: 516, PDE4DIP (Gene); SEQ ID NO: 517, PDE4DIP (Gene); SEQ ID NO: 518, PDE4DIP (Gene); SEQ ID NO: 519, PDE4DIP (Gene); SEQ ID NO: 520, PDE4DIP (Gene); SEQ ID NO: 521, PDE4DIP (Gene); SEQ ID NO: 522, PDE4DIP (Gene); SEQ ID NO: 523, PDE4DIP (Gene); SEQ ID NO: 524, PDE4DIP (Gene); SEQ ID NO: 525, PDE4DIP (Gene); SEQ ID NO: 526, PDE4DIP (Gene); SEQ ID NO: 527, PDE4DIP (Gene); SEQ ID NO: 528, PDE4DIP (Gene); SEQ ID NO: 529, PDE4DIP (Gene); SEQ ID NO: 530, PDE4DIP (Gene); SEQ ID NO: 531, PDE4DIP (Gene); SEQ ID NO: 532, PDE4DIP (Gene); SEQ ID NO: 533, PDE4DIP (Gene); SEQ ID NO: 534, PDE4DIP (Gene); SEQ ID NO: 535, PDE4DIP (Gene); SEQ ID NO: 536, PDE4DIP (Gene); SEQ ID NO: 537, PDE4DIP (Gene); SEQ ID NO: 538, PDE4DIP (Gene); SEQ ID NO: 539, PDE4DIP (Gene); SEQ ID NO: 540, PDE4DIP (Gene); SEQ ID NO: 541, PDE4DIP (Gene); SEQ ID NO: 542, PDE4DIP (Gene); SEQ ID NO: 543, PDE4DIP (Gene); SEQ ID NO: 544, PDE4DIP (Gene); SEQ ID NO: 545, PDE4DIP (Gene); SEQ ID NO: 546, PDE4DIP (Gene); SEQ ID NO: 547, PDE4DIP (Gene); SEQ ID NO: 548, PDE4DIP (Gene); SEQ ID NO: 549, lq23.3_DLBCL (FOCAL); SEQ ID NO: 550, PDE4DIP (Gene); SEQ ID NO: 551, lq23.3_DLBCL (FOCAL); SEQ ID NO: 552, lq23.3_DLBCL (FOCAL); SEQ ID NO: 553, lq23.3_DLBCL (FOCAL); SEQ ID NO: 554, lq23.3_DLBCL (FOCAL); SEQ ID NO: 555, lq23.3_DLBCL (FOCAL); SEQ ID NO: 556, lq23.3_DLBCL (FOCAL); SEQ ID NO: 557, lq23.3_DLBCL (FOCAL); SEQ ID NO: 558, lq23.3_DLBCL (FOCAL); SEQ ID NO: 559, lq23.3_DLBCL (FOCAL); SEQ ID NO: 560, HIST2H2BE (Gene); SEQ ID NO: 561, lq23.3_DLBCL (FOCAL); SEQ ID NO: 562, lq23.3_DLBCL (FOCAL); SEQ ID NO: 563, lq23.3_DLBCL (FOCAL); SEQ ID NO: 564, MCL1 (Gene); SEQ ID NO: 565, MCL1 (Gene); SEQ ID NO: 566, MCL1 (Gene); SEQ ID NO: 567, MCL1 (Gene); SEQ ID NO: 568, MCL1 (Gene); SEQ ID NO: 569, lq23.3_DLBCL (FOCAL); SEQ ID NO: 570, SETDB1 (Gene); SEQ ID NO: 571, SETDB1 (Gene); SEQ ID NO: 572, SETDB1 (Gene); SEQ ID NO: 573, SETDB1 (Gene); SEQ ID NO: 574, SETDB1 (Gene); SEQ ID NO: 575, SETDB1 (Gene); SEQ ID NO: 576, SETDB1 (Gene); SEQ ID NO: 577, SETDB1 (Gene); SEQ ID NO: 578, SETDB1 (Gene); SEQ ID NO: 579, SETDB1 (Gene); SEQ ID NO: 580, SETDB1 (Gene); SEQ ID NO: 581, SETDB1 (Gene); SEQ ID NO: 582, SETDB1 (Gene); SEQ ID NO: 583, SETDB1 (Gene); SEQ ID NO: 584, SETDB1 (Gene); SEQ ID NO: 585, SETDB1 (Gene); SEQ ID NO: 586, SETDB1 (Gene); SEQ ID NO: 587, SETDB1 (Gene); SEQ ID NO: 588, SETDB1 (Gene); SEQ ID NO:589, SETDB1 (Gene); SEQ ID NO: 590, SETDB1 (Gene); SEQ ID NO: 591, SETDB1 (Gene)SEQ ID NO: 592, SETDB1 (Gene); SEQ ID NO: 593, SETDB1 (Gene); SEQ ID NO: 594, lq23.3_DLBCL (FOCAL); SEQ ID NO: 595, lq23.3_DLBCL (FOCAL); SEQ ID NO: 596, lq23.3_DLBCL (FOCAL); SEQ ID NO: 597, lq23.3_DLBCL (FOCAL); SEQ ID NO: 598, lq23.3_DLBCL (FOCAL); SEQ ID NO: 599, lq23.3_DLBCL (FOCAL); SEQ ID NO: 600, lq23.3_DLBCL (FOCAL); SEQ ID NO: 601, lq23.3_DLBCL (FOCAL); SEQ ID NO: 602, lq23.3_DLBCL (FOCAL); SEQ ID NO: 603, lq23.3_DLBCL (FOCAL); SEQ ID NO: 604, lq23.3_DLBCL (FOCAL); SEQ ID NO: 605, lq23.3_DLBCL (FOCAL); SEQ ID NO: 606, lq23.3_DLBCL (FOCAL); SEQ ID NO: 607, lq23.3_DLBCL (FOCAL); SEQ ID NO: 608, lq23.3_DLBCL (FOCAL); SEQ ID NO: 609, lq23.3_DLBCL (FOCAL); SEQ ID NO: 610, lq23.3_DLBCL (FOCAL); SEQ ID NO: 611, lq23.3_DLBCL (FOCAL); SEQ ID NO: 612, lq23.3_DLBCL (FOCAL); SEQ ID NO: 613, lq23.3_DLBCL (FOCAL); SEQ ID NO: 614, lq23.3_DLBCL (FOCAL); SEQ ID NO: 615, MSI (MSI); SEQ ID NO: 616, MSI (MSI); SEQ ID NO: 617, lq23.3_DLBCL (FOCAL); SEQ ID NO: 618, lq23.3_DLBCL (FOCAL); SEQ ID NO: 619, lq23.3_DLBCL (FOCAL); SEQ ID NO: 620, lq23.3_DLBCL (FOCAL); SEQ ID NO: 621, lq23.3_DLBCL (FOCAL); SEQ ID NO: 622, lq23.3_DLBCL (FOCAL); SEQ ID NO: 623,NTRK1 (Gene); SEQ ID NO: 624, NTRK1 (Gene); SEQ ID NO: 625, NTRK1 (Gene); SEQ ID NO: 626, NTRK1 (Gene); SEQ ID NO: 627, NTRK1 (Gene); SEQ ID NO: 628, NTRK1 (Gene); SEQ ID NO: 629, NTRK1 (Gene); SEQ ID NO: 630, NTRK1 (Gene); SEQ ID NO: 631, NTRK1 (Gene); SEQ ID NO: 632, NTRK1 (Gene); SEQ ID NO: 633, NTRK1 (Gene); SEQ ID NO: 634,NTRK1 (Gene); SEQ ID NO: 635, NTRK1 (Gene); SEQ ID NO: 636, NTRK1 (Gene); SEQ IDNO: 637, NTRK1 (Gene); SEQ ID NO: 638, NTRK1 (Gene); SEQ ID NO: 639, NTRK1 (Gene);SEQ ID NO: 640, NTRK1 (Gene); SEQ ID NO: 641, NTRK1 (Gene); SEQ ID NO: 642, NTRK1(Gene); SEQ ID NO: 643, lq23.3_DLBCL (FOCAL); SEQ ID NO: 644, lq23.3_DLBCL (FOCAL); SEQ ID NO: 645, lq23.3_DLBCL (FOCAL); SEQ ID NO: 646, lq23.3_DLBCL (FOCAL); SEQ ID NO: 647, lq23.3_DLBCL (FOCAL); SEQ ID NO: 648, lq23.3_DLBCL (FOCAL); SEQ ID NO: 649, lq23.3_DLBCL (FOCAL); SEQ ID NO: 650, FP (FP); SEQ ID NO: 651, lq23.3_DLBCL (FOCAL); SEQ ID NO: 652, lq23.3_DLBCL (FOCAL); SEQ ID NO: 653, lq23.3_DLBCL (FOCAL); SEQ ID NO: 654, lq23.3_DLBCL (FOCAL); SEQ ID NO: 655, lq23.3_DLBCL (FOCAL); SEQ ID NO: 656, lq23.3_DLBCL (FOCAL); SEQ ID NO: 657, lq23.3_DLBCL (FOCAL); SEQ ID NO: 658, lq23.3_DLBCL (FOCAL); SEQ ID NO: 659, lq23.3_DLBCL (FOCAL); SEQ ID NO: 660, lq23.3_DLBCL (FOCAL); SEQ ID NO: 661, lq23.3_DLBCL (FOCAL); SEQ ID NO: 662, lq23.3_DLBCL (FOCAL); SEQ ID NO: 663, lq23.3_DLBCL (FOCAL); SEQ ID NO: 664, lq23.3_DLBCL (FOCAL); SEQ ID NO: 665,lq23.3_DLBCL (FOCAL); SEQ ID NO: 666, lq23.3_DLBCL (FOCAL); SEQ ID NO: 667, lq23.3_DLBCL (FOCAL); SEQ ID NO: 668, lq23.3_DLBCL (FOCAL); SEQ ID NO: 669, lq23.3_DLBCL (FOCAL); SEQ ID NO: 670, lq23.3_DLBCL (FOCAL); SEQ ID NO: 671, lq23.3_DLBCL (FOCAL); SEQ ID NO: 672, lq23.3_DLBCL (FOCAL); SEQ ID NO: 673, lq23.3_DLBCL (FOCAL); SEQ ID NO: 674, lq23.3_DLBCL (FOCAL); SEQ ID NO: 675, lq23.3_DLBCL (FOCAL); SEQ ID NO: 676, lq23.3_DLBCL (FOCAL); SEQ ID NO: 677, lq23.3_DLBCL (FOCAL); SEQ ID NO: 678, lq23.3_DLBCL (FOCAL); SEQ ID NO: 679, lq23.3_DLBCL (FOCAL); SEQ ID NO: 680, lq23.3_DLBCL (FOCAL); SEQ ID NO: 681, lq23.3_DLBCL (FOCAL); SEQ ID NO: 682, lq23.3_DLBCL (FOCAL); SEQ ID NO: 683, lq23.3_DLBCL (FOCAL); SEQ ID NO: 684, lq23.3_DLBCL (FOCAL); SEQ ID NO: 685, lq23.3_DLBCL (FOCAL); SEQ ID NO: 686, lq23.3_DLBCL (FOCAL); SEQ ID NO: 687, lq23.3_DLBCL (FOCAL); SEQ ID NO: 688, lq23.3_DLBCL (FOCAL); SEQ ID NO: 689, lq23.3_DLBCL (FOCAL); SEQ ID NO: 690, lq23.3_DLBCL (FOCAL); SEQ ID NO: 691, lq23.3_DLBCL (FOCAL); SEQ ID NO: 692, lq23.3_DLBCL (FOCAL); SEQ ID NO: 693, lq23.3_DLBCL (FOCAL); SEQ ID NO: 694, lq23.3_DLBCL (FOCAL); SEQ ID NO: 695, lq23.3_DLBCL (FOCAL); SEQ ID NO: 696, lq23.3_DLBCL (FOCAL); SEQ ID NO: 697, lq23.3_DLBCL (FOCAL); SEQ ID NO: 698, lq23.3_DLBCL (FOCAL); SEQ ID NO: 699, lq23.3_DLBCL (FOCAL); SEQ ID NO: 700, lq23.3_DLBCL (FOCAL); SEQ ID NO: 701, lq23.3_DLBCL (FOCAL); SEQ ID NO: 702, lq23.3_DLBCL (FOCAL); SEQ ID NO: 703, lq23.3_DLBCL (FOCAL); SEQ ID NO: 704, lq23.3_DLBCL (FOCAL); SEQ ID NO: 705, lq23.3_DLBCL (FOCAL); SEQ ID NO: 706, lq23.3_DLBCL (FOCAL); SEQ ID NO: 707, lq23.3_DLBCL (FOCAL); SEQ ID NO: 708, lq23.3_DLBCL (FOCAL); SEQ ID NO: 709, lq23.3_DLBCL (FOCAL); SEQ ID NO: 710, lq23.3_DLBCL (FOCAL); SEQ ID NO: 711, lq23.3_DLBCL (FOCAL); SEQ ID NO: 712, lq23.3_DLBCL (FOCAL); SEQ ID NO: 713, lq23.3_DLBCL (FOCAL); SEQ ID NO: 714, lq23.3_DLBCL (FOCAL); SEQ ID NO: 715, lq23.3_DLBCL (FOCAL); SEQ ID NO: 716, lq23.3_DLBCL (FOCAL); SEQ ID NO: 717, lq23.3_DLBCL (FOCAL); SEQ ID NO: 718, lq23.3_DLBCL (FOCAL); SEQ ID NO: 719, lq23.3_DLBCL (FOCAL); SEQ ID NO: 720, lq23.3_DLBCL (FOCAL); SEQ ID NO: 721, lq23.3_DLBCL (FOCAL); SEQ ID NO: 722, lq23.3_DLBCL (FOCAL); SEQ ID NO: 723, lq23.3_DLBCL (FOCAL); SEQ ID NO: 724, lq23.3_DLBCL (FOCAL); SEQ ID NO: 725, lq23.3_DLBCL (FOCAL); SEQ ID NO: 726, lq23.3_DLBCL (FOCAL); SEQ ID NO: 727, lq23.3_DLBCL (FOCAL); SEQ ID NO: 728, lq23.3_DLBCL (FOCAL); SEQ ID NO: 729, 1Q DLBCL (ARM); SEQ ID NO: 730, 1Q_D ,BCL (ARM); SEQ ID NO: 731, 1Q DLBCL(ARM); SEQ ID NO: 732, 1Q DLBCL (ARM); SEQ ID NO: 733, 1Q DLBCL (ARM); SEQ IDNO: 734, 1Q DLBCL (ARM); SEQ ID NO: 735, 1Q DLBCL (ARM); SEQ ID NO: 736, 1Q DLBCL (ARM); SEQ ID NO: 737, 1Q DLBCL (ARM); SEQ ID NO: 738, 1Q DLBCL (ARM); SEQ ID NO: 739, 1Q DLBCL (ARM); SEQ ID NO: 740, 1Q DLBCL (ARM); SEQ ID NO: 741, 1Q DLBCL (ARM); SEQ ID NO: 742, 1Q DLBCL (ARM); SEQ ID NO: 743, 1Q DLBCL (ARM); SEQ ID NO: 744, 1Q DLBCL (ARM); SEQ ID NO: 745, 1Q DLBCL (ARM); SEQ ID NO: 746, 1Q DLBCL (ARM); SEQ ID NO: 747, 1Q DLBCL (ARM); SEQ ID NO: 748, 1Q DLBCL (ARM); SEQ ID NO: 749, 1Q DLBCL (ARM); SEQ ID NO: 750, 1Q DLBCL (ARM); SEQ ID NO: 751, 1Q DLBCL (ARM); SEQ ID NO: 752, 1Q DLBCL (ARM); SEQ ID NO: 753, 1Q DLBCL (ARM); SEQ ID NO: 754, 1Q DLBCL (ARM); SEQ ID NO: 755, 1Q DLBCL (ARM); SEQ ID NO: 756, BTG2 (Gene); SEQ ID NO: 757, BTG2 (Gene); SEQ ID NO: 758, 1Q DLBCL (ARM); SEQ ID NO: 759, 1Q DLBCL (ARM); SEQ ID NO: 760, 1Q DLBCL (ARM); SEQ ID NO: 761, 1Q DLBCL (ARM); SEQ ID NO: 762, 1Q DLBCL (ARM); SEQ ID NO: 763, 1Q DLBCL (ARM); SEQ ID NO: 764, 1Q DLBCL (ARM); SEQ ID NO: 765, 1Q DLBCL (ARM); SEQ ID NO: 766, 1Q DLBCL (ARM); SEQ ID NO: 767, 1Q DLBCL (ARM); SEQ ID NO: 768, PTPN14 (Gene); SEQ ID NO: 769, PTPN14 (Gene); SEQ ID NO: 770, PTPN14 (Gene); SEQ ID NO: 771, PTPN14 (Gene); SEQ ID NO: 772, PTPN14 (Gene); SEQ ID NO: 773, PTPN14 (Gene); SEQ ID NO: 774, PTPN14 (Gene); SEQ ID NO: 775, PTPN14 (Gene); SEQ ID NO: 776, PTPN14 (Gene); SEQ ID NO: 777, PTPN14 (Gene); SEQ ID NO: 778, PTPN14 (Gene); SEQ ID NO: 779, PTPN14 (Gene); SEQ ID NO: 780, PTPN14 (Gene); SEQ ID NO: 781, PTPN14 (Gene); SEQ ID NO: 782, PTPN14 (Gene); SEQ ID NO: 783, PTPN14 (Gene); SEQ ID NO: 784, PTPN14 (Gene); SEQ ID NO: 785, PTPN14 (Gene); SEQ ID NO: 786, PTPN14 (Gene); SEQ ID NO: 787, PTPN14 (Gene); SEQ ID NO: 788, PTPN14 (Gene); SEQ ID NO: 789, 1Q DLBCL (ARM); SEQ ID NO: 790, 1Q DLBCL (ARM); SEQ ID NO: 791, 1Q DLBCL (ARM); SEQ ID NO: 792, lq42.12_DLBCL (FOCAL); SEQ ID NO: 793, lq42.12_DLBCL (FOCAL); SEQ ID NO: 794, lq42.12_DLBCL (FOCAL); SEQ ID NO: 795, lq42.12_DLBCL (FOCAL); SEQ ID NO: 796, lq42.12_DLBCL (FOCAL); SEQ ID NO: 797, lq42.12_DLBCL (FOCAL); SEQ ID NO: 798, lq42.12_DLBCL (FOCAL); SEQ ID NO: 799, lq42.12_DLBCL (FOCAL); SEQ ID NO: 800, lq42.12_DLBCL (FOCAL); SEQ ID NO: 801, lq42.12_DLBCL (FOCAL); SEQ ID NO: 802, lq42.12_DLBCL (FOCAL); SEQ ID NO: 803, lq42.12_DLBCL (FOCAL); SEQ ID NO: 804, lq42.12_DLBCL (FOCAL); SEQ ID NO: 805, lq42.12_DLBCL (FOCAL); SEQ ID NO: 806, FP (FP); SEQ ID NO: 807, lq42.12_DLBCL(FOCAL); SEQ ID NO: 808, lq42.12_DLBCL (FOCAL); SEQ ID NO: 809, lq42.12_DLBCL(FOCAL); SEQ ID NO: 810, lq42.12_DLBCL (FOCAL); SEQ ID NO: 811, lq42.12_DLBCL(FOCAL); SEQ ID NO: 812, lq42.12_DLBCL (FOCAL); SEQ ID NO: 813, lq42.12_DLBCL(FOCAL); SEQ ID NO: 814, lq42.12_DLBCL (FOCAL); SEQ ID NO: 815, lq42.12_DLBCL(FOCAL); SEQ ID NO: 816, lq42.12_DLBCL (FOCAL); SEQ ID NO: 817, lq42.12_DLBCL(FOCAL); SEQ ID NO: 818, lq42.12_DLBCL (FOCAL); SEQ ID NO: 819, lq42.12_DLBCL(FOCAL); SEQ ID NO: 820, lq42.12_DLBCL (FOCAL); SEQ ID NO: 821, lq42.12_DLBCL(FOCAL); SEQ ID NO: 822, lq42.12_DLBCL (FOCAL); SEQ ID NO: 823, lq42.12_DLBCL(FOCAL); SEQ ID NO: 824, lq42.12_DLBCL (FOCAL); SEQ ID NO: 825, lq42.12_DLBCL(FOCAL); SEQ ID NO: 826, lq42.12_DLBCL (FOCAL); SEQ ID NO: 827, lq42.12_DLBCL(FOCAL); SEQ ID NO: 828, lq42.12_DLBCL (FOCAL); SEQ ID NO: 829, lq42.12_DLBCL(FOCAL); SEQ ID NO: 830, lq42.12_DLBCL (FOCAL); SEQ ID NO: 831, lq42.12_DLBCL(FOCAL); SEQ ID NO: 832, lq42.12_DLBCL (FOCAL); SEQ ID NO: 833, lq42.12_DLBCL(FOCAL); SEQ ID NO: 834, lq42.12_DLBCL (FOCAL); SEQ ID NO: 835, lq42.12_DLBCL(FOCAL); SEQ ID NO: 836, lq42.12_DLBCL (FOCAL); SEQ ID NO: 837, lq42.12_DLBCL(FOCAL); SEQ ID NO: 838, lq42.12_DLBCL (FOCAL); SEQ ID NO: 839, lq42.12_DLBCL(FOCAL); SEQ ID NO: 840, lq42.12_DLBCL (FOCAL); SEQ ID NO: 841, lq42.12_DLBCL(FOCAL); SEQ ID NO: 842, lq42.12_DLBCL (FOCAL); SEQ ID NO: 843, lq42.12_DLBCL(FOCAL); SEQ ID NO: 844, lq42.12_DLBCL (FOCAL); SEQ ID NO: 845, lq42.12_DLBCL(FOCAL); SEQ ID NO: 846, lq42.12_DLBCL (FOCAL); SEQ ID NO: 847, lq42.12_DLBCL(FOCAL); SEQ ID NO: 848, lq42.12_DLBCL (FOCAL); SEQ ID NO: 849, lq42.12_DLBCL(FOCAL); SEQ ID NO: 850, lq42.12_DLBCL (FOCAL); SEQ ID NO: 851, lq42.12_DLBCL(FOCAL); SEQ ID NO: 852, lq42.12_DLBCL (FOCAL); SEQ ID NO: 853, lq42.12_DLBCL(FOCAL); SEQ ID NO: 854, lq42.12_DLBCL (FOCAL); SEQ ID NO: 855, lq42.12_DLBCL(FOCAL); SEQ ID NO: 856, lq42.12_DLBCL (FOCAL); SEQ ID NO: 857, lq42.12_DLBCL(FOCAL); SEQ ID NO: 858, ITPKB (Gene); SEQ ID NO: 859, ITPKB (Gene); SEQ ID NO: 860, ITPKB (Gene); SEQ ID NO: 861, ITPKB (Gene); SEQ ID NO: 862, ITPKB (Gene); SEQ ID NO: 863, ITPKB (Gene); SEQ ID NO: 864, lq42.12_DLBCL (FOCAL); SEQ ID NO: 865, ITPKB (Gene); SEQ ID NO: 866, ITPKB (Gene); SEQ ID NO: 867, ITPKB (Gene); SEQ ID NO: 868, ITPKB (Gene); SEQ ID NO: 869, ITPKB (Gene); SEQ ID NO: 870, lq42.12_DLBCL (FOCAL); SEQ ID NO: 871, lq42.12_DLBCL (FOCAL); SEQ ID NO: 872, lq42.12_DLBCL (FOCAL);SEQ ID NO: 873, lq42.12_DLBCL (FOCAL); SEQ ID NO: 874, lq42.12_DLBCL (FOCAL);SEQ ID NO: 875, lq42.12_DLBCL (FOCAL); SEQ ID NO: 876, lq42.12_DLBCL (FOCAL);SEQ ID NO: 877, lq42.12_DLBCL (FOCAL); SEQ ID NO: 878, lq42.12_DLBCL (FOCAL);SEQ ID NO: 879, lq42.12_DLBCL (FOCAL); SEQ ID NO: 880, lq42.12_DLBCL (FOCAL);SEQ ID NO: 881, lq42.12_DLBCL (FOCAL); SEQ ID NO: 882, lq42.12_DLBCL (FOCAL);SEQ ID NO: 883, lq42.12_DLBCL (FOCAL); SEQ ID NO: 884, lq42.12_DLBCL (FOCAL);SEQ ID NO: 885, lq42.12_DLBCL (FOCAL); SEQ ID NO: 886, lq42.12_DLBCL (FOCAL);SEQ ID NO: 887, lq42.12_DLBCL (FOCAL); SEQ ID NO: 888, lq42.12_DLBCL (FOCAL);SEQ ID NO: 889, lq42.12_DLBCL (FOCAL); SEQ ID NO: 890, lq42.12_DLBCL (FOCAL);SEQ ID NO: 891, lq42.12_DLBCL (FOCAL); SEQ ID NO: 892, lq42.12_DLBCL (FOCAL);SEQ ID NO: 893, lq42.12_DLBCL (FOCAL); SEQ ID NO: 894, lq42.12_DLBCL (FOCAL);SEQ ID NO: 895, lq42.12_DLBCL (FOCAL); SEQ ID NO: 896, lq42.12_DLBCL (FOCAL);SEQ ID NO: 897, lq42.12_DLBCL (FOCAL); SEQ ID NO: 898, lq42.12_DLBCL (FOCAL);SEQ ID NO: 899, lq42.12_DLBCL (FOCAL); SEQ ID NO: 900, lq42.12_DLBCL (FOCAL);SEQ ID NO: 901, lq42.12_DLBCL (FOCAL); SEQ ID NO: 902, lq42.12_DLBCL (FOCAL);SEQ ID NO: 903, lq42.12_DLBCL (FOCAL); SEQ ID NO: 904, lq42.12_DLBCL (FOCAL);SEQ ID NO: 905, lq42.12_DLBCL (FOCAL); SEQ ID NO: 906, lq42.12_DLBCL (FOCAL);SEQ ID NO: 907, lq42.12_DLBCL (FOCAL); SEQ ID NO: 908, lq42.12_DLBCL (FOCAL);SEQ ID NO: 909, lq42.12_DLBCL (FOCAL); SEQ ID NO: 910, MSI (MSI); SEQ ID NO: 911, MSI (MSI); SEQ ID NO: 912, lq42.12_DLBCL (FOCAL); SEQ ID NO: 913, lq42.12_DLBCL (FOCAL); SEQ ID NO: 914, lq42.12_DLBCL (FOCAL); SEQ ID NO: 915, lq42.12_DLBCL(FOCAL); SEQ ID NO: 916, lq42.12_DLBCL (FOCAL); SEQ ID NO: 917, lq42.12_DLBCL(FOCAL); SEQ ID NO: 918, lq42.12_DLBCL (FOCAL); SEQ ID NO: 919, lq42.12_DLBCL(FOCAL); SEQ ID NO: 920, lq42.12_DLBCL (FOCAL); SEQ ID NO: 921, lq42.12_DLBCL(FOCAL); SEQ ID NO: 922, lq42.12_DLBCL (FOCAL); SEQ ID NO: 923, lq42.12_DLBCL(FOCAL); SEQ ID NO: 924, lq42.12_DLBCL (FOCAL); SEQ ID NO: 925, lq42.12_DLBCL(FOCAL); SEQ ID NO: 926, lq42.12_DLBCL (FOCAL); SEQ ID NO: 927, lq42.12_DLBCL(FOCAL); SEQ ID NO: 928, lq42.12_DLBCL (FOCAL); SEQ ID NO: 929, lq42.12_DLBCL(FOCAL); SEQ ID NO: 930, lq42.12_DLBCL (FOCAL); SEQ ID NO: 931, lq42.12_DLBCL(FOCAL); SEQ ID NO: 932, lq42.12_DLBCL (FOCAL); SEQ ID NO: 933, lq42.12_DLBCL(FOCAL); SEQ ID NO: 934, lq42.12_DLBCL (FOCAL); SEQ ID NO: 935, lq42.12_DLBCL(FOCAL); SEQ ID NO: 936, lq42.12_DLBCL (FOCAL); SEQ ID NO: 937, lq42.12_DLBCL(FOCAL); SEQ ID NO: 938, lq42.12_DLBCL (FOCAL); SEQ ID NO: 939, lq42.12_DLBCL(FOCAL); SEQ ID NO: 940, lq42.12_DLBCL (FOCAL); SEQ ID NO: 941, lq42.12_DLBCL(FOCAL); SEQ ID NO: 942, IRF2BP2 (Gene); SEQ ID NO: 943, IRF2BP2 (Gene); SEQ ID NO: 944, IRF2BP2 (Gene); SEQ ID NO: 945, IRF2BP2 (Gene); SEQ ID NO: 946, IRF2BP2 (Gene); SEQ ID NO: 947, lq42.12_DLBCL (FOCAL); SEQ ID NO: 948, lq42.12_DLBCL (FOCAL);SEQ ID NO: 949, lq42.12_DLBCL (FOCAL); SEQ ID NO: 950, lq42.12_DLBCL (FOCAL);SEQ ID NO: 951, lq42.12_DLBCL (FOCAL); SEQ ID NO: 952, lq42.12_DLBCL (FOCAL);SEQ ID NO: 953, lq42.12_DLBCL (FOCAL); SEQ ID NO: 954, lq42.12_DLBCL (FOCAL);SEQ ID NO: 955, lq42.12_DLBCL (FOCAL); SEQ ID NO: 956, lq42.12_DLBCL (FOCAL);SEQ ID NO: 957, lq42.12_DLBCL (FOCAL); SEQ ID NO: 958, lq42.12_DLBCL (FOCAL);SEQ ID NO: 959, lq42.12_DLBCL (FOCAL); SEQ ID NO: 960, lq42.12_DLBCL (FOCAL);SEQ ID NO: 961, lq42.12_DLBCL (FOCAL); SEQ ID NO: 962, lq42.12_DLBCL (FOCAL);SEQ ID NO: 963, lq42.12_DLBCL (FOCAL); SEQ ID NO: 964, lq42.12_DLBCL (FOCAL);SEQ ID NO: 965, lq42.12_DLBCL (FOCAL); SEQ ID NO: 966, lq42.12_DLBCL (FOCAL);SEQ ID NO: 967, lq42.12_DLBCL (FOCAL); SEQ ID NO: 968, lq42.12_DLBCL (FOCAL);SEQ ID NO: 969, lq42.12_DLBCL (FOCAL); SEQ ID NO: 970, lq42.12_DLBCL (FOCAL);SEQ ID NO: 971, lq42.12_DLBCL (FOCAL); SEQ ID NO: 972, lq42.12_DLBCL (FOCAL);SEQ ID NO: 973, lq42.12_DLBCL (FOCAL); SEQ ID NO: 974, lq42.12_DLBCL (FOCAL);SEQ ID NO: 975, lq42.12_DLBCL (FOCAL); SEQ ID NO: 976, lq42.12_DLBCL (FOCAL);SEQ ID NO: 977, lq42.12_DLBCL (FOCAL); SEQ ID NO: 978, FP (FP); SEQ ID NO: 979, lq42.12_DLBCL (FOCAL); SEQ ID NO: 980, lq42.12_DLBCL (FOCAL); SEQ ID NO: 981, lq42.12_DLBCL (FOCAL); SEQ ID NO: 982, lq42.12_DLBCL (FOCAL); SEQ ID NO: 983, lq42.12_DLBCL (FOCAL); SEQ ID NO: 984, lq42.12_DLBCL (FOCAL); SEQ ID NO: 985, lq42.12_DLBCL (FOCAL); SEQ ID NO: 986, lq42.12_DLBCL (FOCAL); SEQ ID NO: 987, FP (FP); SEQ ID NO: 988, lq42.12_DLBCL (FOCAL); SEQ ID NO: 989, lq42.12_DLBCL (FOCAL); SEQ ID NO: 990, lq42.12_DLBCL (FOCAL); SEQ ID NO: 991, lq42.12_DLBCL(FOCAL); SEQ ID NO: 992, lq42.12_DLBCL (FOCAL); SEQ ID NO: 993, lq42.12_DLBCL(FOCAL); SEQ ID NO: 994, lq42.12_DLBCL (FOCAL); SEQ ID NO: 995, lq42.12_DLBCL(FOCAL); SEQ ID NO: 996, lq42.12_DLBCL (FOCAL); SEQ ID NO: 997, lq42.12_DLBCL(FOCAL); SEQ ID NO: 998, lq42.12_DLBCL (FOCAL); SEQ ID NO: 999, lq42.12_DLBCL(FOCAL); SEQ ID NO: 1000, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1001, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1002, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1003, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1004, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1005, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1006, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1007, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1008, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1009, FP (FP); SEQ ID NO: 1010, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1011, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1012, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1013, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1014, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1015, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1016, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1017, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1018, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1019, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1020, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1021, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1022, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1023, lq42.12_DLBCL (FOCAL); SEQ IDNO: 1024, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1025, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1026, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1027, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1028, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1029, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1030, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1031, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1032, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1033, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1034, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1035, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1036, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1037, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1038, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1039, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1040, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1041, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1042, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1043, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1044, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1045, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1046, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1047, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1048, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1049, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1050, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1051, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1052, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1053, lq42.12_DLBCL (FOCAL); SEQ ID NO: 1054, 10P (ARM); SEQ ID NO: 1055, 10P (ARM); SEQ ID NO: 1056, 10P (ARM); SEQ ID NO: 1057, 10P (ARM); SEQ ID NO: 1058, FP (FP); SEQ ID NO: 1059, 10P (ARM); SEQ ID NO: 1060, 10P (ARM); SEQ ID NO: 1061, 1 OP (ARM); SEQ ID NO: 1062, 10P (ARM); SEQ ID NO: 1063, 10P (ARM); SEQ ID NO: 1064, 10P (ARM); SEQ ID NO: 1065, 10P (ARM); SEQ ID NO: 1066, 10P (ARM); SEQ ID NO: 1067, 10P (ARM); SEQ ID NO: 1068, 10P (ARM); SEQ ID NO: 1069, 10P (ARM); SEQ ID NO: 1070, 10P (ARM); SEQ ID NO: 1071, 1 OP (ARM); SEQ ID NO: 1072, 10P (ARM); SEQ ID NO: 1073, 10P (ARM); SEQ ID NO: 1074, 10P (ARM); SEQ ID NO: 1075, 10P (ARM); SEQ ID NO: 1076, 10P (ARM); SEQ ID NO: 1077, 10P (ARM); SEQ ID NO: 1078, 10P (ARM); SEQ ID NO: 1079, 10P (ARM); SEQ ID NO: 1080, 10P (ARM); SEQ ID NO: 1081, 1 OP (ARM); SEQ ID NO: 1082, 10P (ARM); SEQ ID NO: 1083, WAC (Gene); SEQ ID NO: 1084, WAC (Gene); SEQ ID NO: 1085, WAC (Gene); SEQ ID NO: 1086, WAC (Gene); SEQ ID NO: 1087, WAC (Gene); SEQ ID NO: 1088, WAC (Gene); SEQ ID NO: 1089, WAC (Gene); SEQ ID NO: 1090, WAC (Gene); SEQ ID NO: 1091, WAC (Gene); SEQ ID NO: 1092, WAC (Gene); SEQ ID NO: 1093, WAC (Gene); SEQ ID NO: 1094, WAC (Gene); SEQ ID NO: 1095, WAC (Gene); SEQ ID NO: 1096, WAC (Gene); SEQ ID NO: 1097, WAC (Gene); SEQ ID NO: 1098, 10P (ARM); SEQ ID NO: 1099, 10P (ARM); SEQ ID NO: 1100, MSI (MSI); SEQ ID NO: 1101, MSI (MSI); SEQ ID NO: 1102, 10P (ARM); SEQ ID NO: 1103, 10P (ARM); SEQ ID NO: 1104, 10P (ARM); SEQ ID NO: 1105, 10P (ARM); SEQ ID NO: 1106, 10P (ARM); SEQ ID NO: 1107, RET (Gene); SEQ ID NO: 1108, RET (Gene); SEQ ID NO: 1109, RET (Gene); SEQID NO: 1110, RET (Gene); SEQ ID NO: 1111, RET (Gene); SEQ ID NO: 1112, RET (Gene); SEQ ID NO: 1113, RET (Gene); SEQ ID NO: 1114, RET (Gene); SEQ ID NO: 1115, RET (Gene); SEQ ID NO: 1116, RET (Gene); SEQ ID NO: 1117, RET (Gene); SEQ ID NO: 1118, RET (Gene); SEQ ID NO: 1119, RET (Gene); SEQ ID NO: 1120, RET (Gene); SEQ ID NO: 1121, RET (Gene); SEQ ID NO: 1122, RET (Gene); SEQ ID NO: 1123, RET (Gene); SEQ ID NO: 1124, RET (Gene); SEQ ID NO: 1125, RET (Gene); SEQ ID NO: 1126, RET (Gene); SEQ ID NO: 1127, RET (Gene); SEQ ID NO: 1128, 10P (ARM); SEQ ID NO: 1129, 10P (ARM); SEQ ID NO: 1130, 10P (ARM); SEQ ID NO: 1131, lOP (ARM); SEQ ID NO: 1132, lOP (ARM); SEQ ID NO: 1133, lOP (ARM); SEQ ID NO: 1134, 10P (ARM); SEQ ID NO: 1135, 10P (ARM); SEQ ID NO: 1136, 10P (ARM); SEQ ID NO: 1137, 10P (ARM); SEQ ID NO: 1138, 10P (ARM); SEQ ID NO: 1139, 10P (ARM); SEQ ID NO: 1140, FP (FP); SEQ ID NO: 1141, 10P (ARM); SEQ ID NO: 1142, 10P (ARM); SEQ ID NO: 1143, 10P (ARM); SEQ ID NO: 1144, 10P (ARM); SEQ ID NO: 1145, 10P (ARM); SEQ ID NO: 1146, ARID5B (Gene); SEQ ID NO: 1147, ARID5B (Gene); SEQ ID NO: 1148, ARID5B (Gene); SEQ ID NO: 1149, ARID5B (Gene); SEQ ID NO: 1150, ARID5B (Gene); SEQ ID NO: 1151, ARID5B (Gene); SEQ ID NO: 1152, ARID5B (Gene); SEQ ID NO: 1153, ARID5B (Gene); SEQ ID NO: 1154, ARID5B (Gene); SEQ ID NO: 1155, ARID5B (Gene); SEQ ID NO: 1156, ARID5B (Gene); SEQ ID NO: 1157, ARID5B (Gene); SEQ ID NO: 1158, ARID5B (Gene); SEQ ID NO: 1159, ARID5B (Gene); SEQ ID NO: 1160, ARID5B (Gene); SEQ ID NO: 1161, ARID5B (Gene); SEQ ID NO: 1162, ARID5B (Gene); SEQ ID NO: 1163, ARID5B (Gene); SEQ ID NO: 1164, ARID5B (Gene); SEQ ID NO: 1165, ARID5B (Gene); SEQ ID NO: 1166, 10P (ARM); SEQ ID NO: 1167, 1 OP (ARM); SEQ ID NO: 1168, 10P (ARM); SEQ ID NO: 1169, 10P(ARM); SEQ ID NO: 1170, 10P (ARM); SEQ ID NO: 1171, MSI (MSI); SEQ ID NO: 1172, MSI(MSI); SEQ ID NO: 1173, 10P (ARM); SEQ ID NO: 1174, 10P (ARM); SEQ ID NO: 1175, 10P (ARM); SEQ ID NO: 1176, 1 OP (ARM); SEQ ID NO: 1177, 10P (ARM); SEQ ID NO: 1178, 10P(ARM); SEQ ID NO: 1179, 1 OP (ARM); SEQ ID NO: 1180, 10P (ARM); SEQ ID NO: 1181, 10P(ARM); SEQ ID NO: 1182, 10P (ARM); SEQ ID NO: 1183, 10P (ARM); SEQ ID NO: 1184, 10P(ARM); SEQ ID NO: 1185, 1 OP (ARM); SEQ ID NO: 1186, 10P (ARM); SEQ ID NO: 1187, 10P(ARM); SEQ ID NO: 1188, 10P (ARM); SEQ ID NO: 1189, 10P (ARM); SEQ ID NO: 1190, 10P(ARM); SEQ ID NO: 1191, 10P (ARM); SEQ ID NO: 1192, 10P (ARM); SEQ ID NO: 1193, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1194, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1195, 10q23.31 DLBCL (FOCAL); SEQ ID NO: 1196, PTEN (Gene); SEQ ID NO: 1197, PTEN (Gene); SEQ ID NO: 1198, PTEN (Gene); SEQ ID NO: 1199, PTEN (Gene); SEQ ID NO: 1200, PTEN (Gene); SEQ ID NO: 1201, PTEN (Gene); SEQ ID NO: 1202, PTEN (Gene); SEQ ID NO: 1203, PTEN (Gene); SEQ ID NO: 1204, PTEN (Gene); SEQ ID NO: 1205, PTEN (Gene); SEQID NO: 1206, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1207, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1208, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1209, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1210, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1211, FAS (Gene); SEQ ID NO: 1212, FAS (Gene); SEQ ID NO: 1213, FAS (Gene); SEQ ID NO: 1214, FAS (Gene); SEQ ID NO: 1215, FAS (Gene); SEQ ID NO: 1216, FAS (Gene); SEQ ID NO: 1217, FAS (Gene); SEQ ID NO: 1218, FAS (Gene); SEQ ID NO: 1219, FAS (Gene); SEQ ID NO: 1220, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1221, 10q23.31_DLBCL (FOCAL); SEQ ID NO: 1222, 10P (ARM); SEQ ID NO: 1223, 10P (ARM); SEQ ID NO: 1224, 10P (ARM); SEQ ID NO: 1225, 10P (ARM); SEQ ID NO: 1226, 10P (ARM); SEQ ID NO: 1227, 10P (ARM); SEQ ID NO: 1228, MSI (MSI); SEQ ID NO: 1229, MSI (MSI); SEQ ID NO: 1230, 10P (ARM); SEQ ID NO: 1231, MSI (MSI); SEQ ID NO: 1232, MSI (MSI); SEQ ID NO: 1233, 10P (ARM); SEQ ID NO: 1234, 10Q (ARM); SEQ ID NO: 1235, 10Q (ARM); SEQ ID NO: 1236, 10Q (ARM); SEQ ID NO: 1237, 10Q (ARM); SEQ ID NO: 1238, 10Q (ARM); SEQ ID NO: 1239, 10Q (ARM); SEQ ID NO: 1240, 10Q (ARM); SEQ ID NO: 1241, 10Q (ARM); SEQ ID NO: 1242, MSI (MSI); SEQ ID NO: 1243, MSI (MSI); SEQ ID NO: 1244, 10Q (ARM); SEQ ID NO: 1245, 10Q (ARM); SEQ ID NO: 1246, 10Q (ARM); SEQ ID NO: 1247, 10Q (ARM); SEQ ID NO: 1248, 10Q (ARM); SEQ ID NO: 1249, 10Q (ARM); SEQ ID NO: 1250, 10Q (ARM); SEQ ID NO: 1251, 10Q (ARM); SEQ ID NO: 1252, 10Q (ARM); SEQ ID NO: 1253, 10Q (ARM); SEQ ID NO: 1254, 10Q (ARM); SEQ ID NO: 1255, 10Q (ARM); SEQ ID NO: 1256, 10Q (ARM); SEQ ID NO: 1257, 10Q (ARM); SEQ ID NO: 1258, 10Q (ARM); SEQ ID NO: 1259, 10Q (ARM); SEQ ID NO: 1260, 10Q (ARM); SEQ ID NO: 1261, 10Q (ARM); SEQ ID NO: 1262, 10Q (ARM); SEQ ID NO: 1263, 10Q (ARM); SEQ ID NO: 1264, FP (FP); SEQ ID NO: 1265, 10Q (ARM); SEQ ID NO: 1266, 10Q (ARM); SEQ ID NO: 1267, 10Q (ARM); SEQ ID NO: 1268, 10Q (ARM); SEQ ID NO: 1269, 10Q (ARM); SEQ ID NO: 1270, FP (FP); SEQ ID NO: 1271, 10Q (ARM); SEQ ID NO: 1272, 10Q (ARM); SEQ ID NO: 1273, 10Q (ARM); SEQ ID NO: 1274, 11P DLBCL (ARM); SEQ ID NO: 1275, 11P DLBCL (ARM); SEQ ID NO: 1276, 11P DLBCL (ARM); SEQ ID NO: 1277, 11P DLBCL (ARM); SEQ ID NO: 1278, 11P DLBCL (ARM); SEQ ID NO: 1279, 11P DLBCL (ARM); SEQ ID NO: 1280, 11P DLBCL (ARM); SEQ ID NO: 1281, 11P DLBCL (ARM); SEQ ID NO: 1282, 11P DLBCL (ARM); SEQ ID NO: 1283, 11P DLBCL (ARM); SEQ ID NO: 1284, FP (FP); SEQ ID NO: 1285, 11P DLBCL (ARM); SEQ ID NO: 1286, 11P DLBCL (ARM); SEQ ID NO: 1287, 11P DLBCL (ARM); SEQ ID NO: 1288, 11P DLBCL (ARM); SEQ ID NO: 1289, 11P DLBCL (ARM); SEQ ID NO: 1290, 11P DLBCL (ARM); SEQ ID NO: 1291, 11P DLBCL (ARM); SEQ ID NO: 1292, 11P DLBCL (ARM); SEQ ID NO: 1293, 11P DLBCL (ARM); SEQ ID NO: 1294, 11P DLBCL (ARM); SEQ ID NO: 1295, 11P DLBCL (ARM); SEQ ID NO: 1296, 11P DLBCL(ARM); SEQ ID NO: 1297, 11P DLBCL (ARM); SEQ ID NO: 1298, 11P DLBCL (ARM); SEQ ID NO: 1299, 11P DLBCL (ARM); SEQ ID NO: 1300, 11P DLBCL (ARM); SEQ ID NO: 1301, 11P DLBCL (ARM); SEQ ID NO: 1302, 11P DLBCL (ARM); SEQ ID NO: 1303, 11P DLBCL (ARM); SEQ ID NO: 1304, 11P DLBCL (ARM); SEQ ID NO: 1305, 11P DLBCL (ARM); SEQ ID NO: 1306, 11P DLBCL (ARM); SEQ ID NO: 1307, 11P DLBCL (ARM); SEQ ID NO: 1308, 11P DLBCL (ARM); SEQ ID NO: 1309, 11P DLBCL (ARM); SEQ ID NO: 1310, 11P DLBCL (ARM); SEQ ID NO: 1311, 11P DLBCL (ARM); SEQ ID NO: 1312, 11P DLBCL (ARM); SEQ ID NO: 1313, 11P DLBCL (ARM); SEQ ID NO: 1314, 11P DLBCL (ARM); SEQ ID NO: 1315, 11P DLBCL (ARM); SEQ ID NO: 1316, 11P DLBCL (ARM); SEQ ID NO: 1317, 11P DLBCL (ARM); SEQ ID NO: 1318, 11P DLBCL (ARM); SEQ ID NO: 1319, 11P DLBCL (ARM); SEQ ID NO: 1320, 11P DLBCL (ARM); SEQ ID NO: 1321, 11P DLBCL (ARM); SEQ ID NO: 1322, 11Q DLBCL (ARM); SEQ ID NO: 1323, MS4A1 (Gene); SEQ ID NO: 1324, MS4A1 (Gene); SEQ ID NO: 1325, MS4A1 (Gene); SEQ ID NO: 1326, 11Q DLBCL (ARM); SEQ ID NO: 1327, MS4A1 (Gene); SEQ ID NO: 1328, MS4A1 (Gene); SEQ ID NO: 1329, MS4A1 (Gene); SEQ ID NO: 1330, 11Q DLBCL (ARM); SEQ ID NO: 1331, 11Q DLBCL (ARM); SEQ ID NO: 1332, MSI (MSI); SEQ ID NO: 1333, MSI (MSI); SEQ ID NO: 1334, 11Q DLBCL (ARM); SEQ ID NO: 1335, 11Q DLBCL (ARM); SEQ ID NO: 1336, CCND1 MTC (SV); SEQ ID NO: 1337, CCND1 MTC (SV); SEQ ID NO: 1338, CCND1 MTC (SV); SEQ ID NO: 1339, CCND1 MTC (SV); SEQ ID NO: 1340, CCND1 MTC (SV); SEQ ID NO: 1341, CCND1 MTC (SV); SEQ ID NO: 1342, CCND1 MTC (SV); SEQ ID NO: 1343, CCND1 MTC (SV); SEQ ID NO: 1344, CCND1 MTC (SV); SEQ ID NO: 1345, CCND1 MTC (SV); SEQ ID NO: 1346, CCND1 MTC (SV); SEQ ID NO: 1347, CCND1 MTC (SV); SEQ ID NO: 1348, CCND1 MTC (SV); SEQ ID NO: 1349, CCND1 MTC (SV); SEQ ID NO: 1350, CCND1 MTC (SV); SEQ ID NO: 1351, CCND1 MTC (SV); SEQ ID NO: 1352, CCND1 MTC (SV); SEQ ID NO: 1353, CCND1 MTC (SV); SEQ ID NO: 1354, CCND1 MTC (SV); SEQ ID NO: 1355, CCND1 MTC (SV); SEQ ID NO: 1356, CCND1 MTC (SV); SEQ ID NO: 1357, CCND1 MTC (SV); SEQ ID NO: 1358, CCND1 MTC (SV); SEQ ID NO: 1359, CCND1 MTC (SV); SEQ ID NO: 1360, CCND1 MTC (SV); SEQ ID NO: 1361, CCND1 MTC (SV); SEQ ID NO: 1362, CCND1 MTC (SV); SEQ ID NO: 1363, CCND1 MTC (SV); SEQ ID NO: 1364, CCND1 MTC (SV); SEQ ID NO: 1365, CCND1 MTC (SV); SEQ ID NO: 1366, CCND1 MTC (SV); SEQ ID NO: 1367, CCND1 MTC (SV); SEQ ID NO: 1368, CCND1 MTC (SV); SEQ ID NO: 1369, CCND1 MTC (SV); SEQ ID NO: 1370, CCND1 MTC (SV); SEQ ID NO: 1371, CCND1 MTC (SV); SEQ ID NO: 1372, CCND1 MTC (SV); SEQ ID NO: 1373, CCND1 MTC (SV); SEQ ID NO: 1374, CCND1 MTC (SV); SEQ ID NO: 1375, CCND1 MTC (SV); SEQ ID NO: 1376, CCND1 MTC (SV); SEQ IDNO: 1377, CCND1 MTC (SV); SEQ ID NO: 1378, CCND1 MTC (SV); SEQ ID NO: 1379, CCND1 MTC (SV); SEQ ID NO: 1380, CCND1 MTC (SV); SEQ ID NO: 1381, CCND1 MTC (SV); SEQ ID NO: 1382, CCND1 MTC (SV); SEQ ID NO: 1383, CCND1 MTC (SV); SEQ ID NO: 1384, CCND1 MTC (SV); SEQ ID NO: 1385, CCND1 MTC (SV); SEQ ID NO: 1386,CCND1 MTC (SV); SEQ ID NO: 1387, CCND1 MTC (SV); SEQ ID NO: 1388, CCND1 MTC(SV); SEQ ID NO: 1389, CCND1 MTC (SV); SEQ ID NO: 1390, CCND1 MTC (SV); SEQ IDNO: 1391, CCND1 MTC (SV); SEQ ID NO: 1392, CCND1 MTC (SV); SEQ ID NO: 1393,CCNDl_promoter (SV); SEQ ID NO: 1394, CCND1 promoter (SV); SEQ ID NO: 1395,CCNDl_promoter (SV); SEQ ID NO: 1396, CCND1 promoter (SV); SEQ ID NO: 1397,CCNDl_promoter (SV); SEQ ID NO: 1398, CCND1 promoter (SV); SEQ ID NO: 1399,CCNDl_promoter (SV); SEQ ID NO: 1400, CCND1 promoter (SV); SEQ ID NO: 1401,CCNDl_promoter (SV); SEQ ID NO: 1402, CCND1 promoter (SV); SEQ ID NO: 1403,CCNDl_promoter (SV); SEQ ID NO: 1404, CCND1 promoter (SV); SEQ ID NO: 1405,CCNDl_promoter (SV); SEQ ID NO: 1406, CCND1 promoter (SV); SEQ ID NO: 1407,CCNDl_promoter (SV); SEQ ID NO: 1408, CCND1 promoter (SV); SEQ ID NO: 1409,CCNDl_promoter (SV); SEQ ID NO: 1410, CCND1 promoter (SV); SEQ ID NO: 1411,CCNDl_promoter (SV); SEQ ID NO: 1412, CCND1 promoter (SV); SEQ ID NO: 1413,CCNDl_promoter (SV); SEQ ID NO: 1414, CCND1 promoter (SV); SEQ ID NO: 1415,CCNDl_promoter (SV); SEQ ID NO: 1416, CCND1 promoter (SV); SEQ ID NO: 1417,CCNDl_promoter (SV); SEQ ID NO: 1418, CCND1 promoter (SV); SEQ ID NO: 1419,CCNDl_promoter (SV); SEQ ID NO: 1420, CCND1 promoter (SV); SEQ ID NO: 1421,CCNDl_promoter (SV); SEQ ID NO: 1422, CCND1 promoter (SV); SEQ ID NO: 1423,CCNDl_promoter (SV); SEQ ID NO: 1424, CCND1 promoter (SV); SEQ ID NO: 1425,CCNDl_promoter (SV); SEQ ID NO: 1426, CCND1 promoter (SV); SEQ ID NO: 1427,CCNDl_promoter (SV); SEQ ID NO: 1428, CCND1 promoter (SV); SEQ ID NO: 1429,CCNDl_promoter (SV); SEQ ID NO: 1430, CCND1 promoter (SV); SEQ ID NO: 1431,CCNDl_promoter (SV); SEQ ID NO: 1432, CCND1 promoter (SV); SEQ ID NO: 1433,CCNDl_promoter (SV); SEQ ID NO: 1434, CCND1 promoter (SV); SEQ ID NO: 1435,CCNDl_promoter (SV); SEQ ID NO: 1436, CCND1 promoter (SV); SEQ ID NO: 1437,CCNDl_promoter (SV); SEQ ID NO: 1438, CCND1 promoter (SV); SEQ ID NO: 1439,CCNDl_promoter (SV); SEQ ID NO: 1440, CCND1 promoter (SV); SEQ ID NO: 1441,CCNDl_promoter (SV); SEQ ID NO: 1442, CCND1 promoter (SV); SEQ ID NO: 1443,CCNDl_promoter (SV); SEQ ID NO: 1444, CCND1_SV (SV); SEQ ID NO: 1445,CCNDl_promoter (SV); SEQ ID NO: 1446, CCND1 promoter (SV); SEQ ID NO: 1447,CCND1 SV (SV); SEQ ID NO: 1448, CCND1 SV (SV); SEQ ID NO: 1449, CCND1 (Gene);SEQ ID NO: 1450, CCND1 (Gene); SEQ ID NO: 1451, CCND1 (Gene); SEQ ID NO: 1452, 11Q DLBCL (ARM); SEQ ID NO: 1453, 11Q DLBCL (ARM); SEQ ID NO: 1454,11Q DLBCL (ARM); SEQ ID NO: 1455, 11Q DLBCL (ARM); SEQ ID NO: 1456,11Q DLBCL (ARM); SEQ ID NO: 1457, 11Q DLBCL (ARM); SEQ ID NO: 1458,11Q DLBCL (ARM); SEQ ID NO: 1459, 11Q DLBCL (ARM); SEQ ID NO: 1460,11Q DLBCL (ARM); SEQ ID NO: 1461, 11Q DLBCL (ARM); SEQ ID NO: 1462,11Q DLBCL (ARM); SEQ ID NO: 1463, 11Q DLBCL (ARM); SEQ ID NO: 1464,11Q DLBCL (ARM); SEQ ID NO: 1465, 11Q DLBCL (ARM); SEQ ID NO: 1466,11Q DLBCL (ARM); SEQ ID NO: 1467, 11Q DLBCL (ARM); SEQ ID NO: 1468,11Q DLBCL (ARM); SEQ ID NO: 1469, 11Q DLBCL (ARM); SEQ ID NO: 1470,11Q DLBCL (ARM); SEQ ID NO: 1471, 11Q DLBCL (ARM); SEQ ID NO: 1472,11Q DLBCL (ARM); SEQ ID NO: 1473, 11Q DLBCL (ARM); SEQ ID NO: 1474,11Q DLBCL (ARM); SEQ ID NO: 1475, 11Q DLBCL (ARM); SEQ ID NO: 1476,11Q DLBCL (ARM); SEQ ID NO: 1477, 11Q DLBCL (ARM); SEQ ID NO: 1478,11Q DLBCL (ARM); SEQ ID NO: 1479, 11Q DLBCL (ARM); SEQ ID NO: 1480,11Q DLBCL (ARM); SEQ ID NO: 1481, 11Q DLBCL (ARM); SEQ ID NO: 1482,11Q DLBCL (ARM); SEQ ID NO: 1483, BIRC3 (Gene); SEQ ID NO: 1484, BIRC3 (Gene); SEQ ID NO: 1485, BIRC3 (Gene); SEQ ID NO: 1486, BIRC3 (Gene); SEQ ID NO: 1487, BIRC3 (Gene); SEQ ID NO: 1488, BIRC3 (Gene); SEQ ID NO: 1489, BIRC3 (Gene); SEQ ID NO: 1490,BIRC3 (Gene); SEQ ID NO: 1491, BIRC3 (Gene); SEQ ID NO: 1492, BIRC3 (Gene); SEQ ID NO: 1493, 11Q DLBCL (ARM); SEQ ID NO: 1494, 11Q DLBCL (ARM); SEQ ID NO: 1495, 11Q DLBCL (ARM); SEQ ID NO: 1496, 11Q DLBCL (ARM); SEQ ID NO: 1497, ATM (Gene); SEQ ID NO: 1498, ATM (Gene); SEQ ID NO: 1499, ATM (Gene); SEQ ID NO: 1500, ATM (Gene); SEQ ID NO: 1501, ATM (Gene); SEQ ID NO: 1502, ATM (Gene); SEQ ID NO: 1503, ATM (Gene); SEQ ID NO: 1504, ATM (Gene); SEQ ID NO: 1505, ATM (Gene); SEQ ID NO: 1506, ATM (Gene); SEQ ID NO: 1507, ATM (Gene); SEQ ID NO: 1508, ATM (Gene); SEQ ID NO: 1509, ATM (Gene); SEQ ID NO: 1510, ATM (Gene); SEQ ID NO: 1511, ATM (Gene); SEQ ID NO: 1512, ATM (Gene); SEQ ID NO: 1513, ATM (Gene); SEQ ID NO: 1514, ATM (Gene); SEQ ID NO: 1515, ATM (Gene); SEQ ID NO: 1516, ATM (Gene); SEQ ID NO: 1517, ATM (Gene); SEQ ID NO: 1518, ATM (Gene); SEQ ID NO: 1519, ATM (Gene); SEQ ID NO: 1520, ATM (Gene); SEQ ID NO: 1521, ATM (Gene); SEQ ID NO: 1522, ATM (Gene); SEQ ID NO: 1523, ATM (Gene); SEQ ID NO: 1524, ATM (Gene); SEQ ID NO: 1525, ATM (Gene); SEQ ID NO: 1526, ATM (Gene); SEQ ID NO: 1527, ATM (Gene); SEQ ID NO: 1528, ATM (Gene);SEQ ID NO: 1529, ATM (Gene); SEQ ID NO: 1530, ATM (Gene); SEQ ID NO: 1531, ATM (Gene); SEQ ID NO: 1532, ATM (Gene); SEQ ID NO: 1533, ATM (Gene); SEQ ID NO: 1534, ATM (Gene); SEQ ID NO: 1535, ATM (Gene); SEQ ID NO: 1536, ATM (Gene); SEQ ID NO: 1537, ATM (Gene); SEQ ID NO: 1538, ATM (Gene); SEQ ID NO: 1539, ATM (Gene); SEQ ID NO: 1540, ATM (Gene); SEQ ID NO: 1541, ATM (Gene); SEQ ID NO: 1542, ATM (Gene); SEQ ID NO: 1543, ATM (Gene); SEQ ID NO: 1544, ATM (Gene); SEQ ID NO: 1545, ATM (Gene); SEQ ID NO: 1546, ATM (Gene); SEQ ID NO: 1547, ATM (Gene); SEQ ID NO: 1548, ATM (Gene); SEQ ID NO: 1549, ATM (Gene); SEQ ID NO: 1550, ATM (Gene); SEQ ID NO: 1551, ATM (Gene); SEQ ID NO: 1552, ATM (Gene); SEQ ID NO: 1553, ATM (Gene); SEQ ID NO: 1554, ATM (Gene); SEQ ID NO: 1555, ATM (Gene); SEQ ID NO: 1556, ATM (Gene); SEQ ID NO: 1557, ATM (Gene); SEQ ID NO: 1558, ATM (Gene); SEQ ID NO: 1559, ATM (Gene); SEQ ID NO: 1560, 11Q DLBCL (ARM); SEQ ID NO: 1561, FP (FP); SEQ ID NO: 1562, POU2AF1 (Gene); SEQ ID NO: 1563, POU2AF1 (Gene); SEQ ID NO: 1564, POU2AF1 (Gene); SEQ ID NO: 1565, POU2AF1 (Gene); SEQ ID NO: 1566, POU2AF1 (Gene); SEQ ID NO: 1567, 11Q DLBCL (ARM); SEQ ID NO: 1568, 11Q DLBCL (ARM); SEQ ID NO: 1569,11Q DLBCL (ARM); SEQ ID NO: 1570, 11Q DLBCL (ARM); SEQ ID NO: 1571,11Q DLBCL (ARM); SEQ ID NO: 1572, 11Q DLBCL (ARM); SEQ ID NO: 1573,11Q DLBCL (ARM); SEQ ID NO: 1574, 11Q DLBCL (ARM); SEQ ID NO: 1575,11Q DLBCL (ARM); SEQ ID NO: 1576, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1577, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1578, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1579, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1580, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1581, IL 1 ORA (Gene); SEQ ID NO: 1582, IL 1 ORA (Gene); SEQ ID NO: 1583, IL 1 ORA (Gene); SEQ ID NO: 1584, IL10RA (Gene); SEQ ID NO: 1585, IL10RA (Gene); SEQ ID NO: 1586, IL10RA (Gene); SEQ ID NO: 1587, IL10RA (Gene); SEQ ID NO: 1588, IL10RA (Gene); SEQ ID NO: 1589, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1590, 1 lq23.3_DLBCL (FOCAL); SEQ ID NO: 1591, MSI (MSI); SEQ ID NO: 1592, MSI (MSI); SEQ ID NO: 1593, 11Q DLBCL (ARM); SEQ ID NO: 1594, 11Q DLBCL (ARM); SEQ ID NO: 1595, 11Q DLBCL (ARM); SEQ ID NO: 1596, 11Q DLBCL (ARM); SEQ ID NO: 1597, 11Q DLBCL (ARM); SEQ ID NO: 1598,11Q DLBCL (ARM); SEQ ID NO: 1599, 11Q DLBCL (ARM); SEQ ID NO: 1600,11Q DLBCL (ARM); SEQ ID NO: 1601, 11Q DLBCL (ARM); SEQ ID NO: 1602,11Q DLBCL (ARM); SEQ ID NO: 1603, 11Q DLBCL (ARM); SEQ ID NO: 1604, MSI (MSI);SEQ ID NO: 1605, MSI (MSI); SEQ ID NO: 1606, 11Q DLBCL (ARM); SEQ ID NO: 1607, 11Q DLBCL (ARM); SEQ ID NO: 1608, 11Q DLBCL (ARM); SEQ ID NO: 1609, 11Q DLBCL (ARM); SEQ ID NO: 1610, ETS1 (Gene); SEQ ID NO: 1611, ETS1 (Gene); SEQID NO: 1612, ETS1 (Gene); SEQ ID NO: 1613, ETS1 (Gene); SEQ ID NO: 1614, ETS1 (Gene); SEQ ID NO: 1615, ETS1 (Gene); SEQ ID NO: 1616, ETS1 (Gene); SEQ ID NO: 1617, ETS1 (Gene); SEQ ID NO: 1618, ETS1 (Gene); SEQ ID NO: 1619, ETS1 (Gene); SEQ ID NO: 1620, ETS1 (Gene); SEQ ID NO: 1621, 11Q DLBCL (ARM); SEQ ID NO: 1622, 11Q DLBCL (ARM); SEQ ID NO: 1623, 11Q DLBCL (ARM); SEQ ID NO: 1624, 11Q DLBCL (ARM); SEQ ID NO: 1625, 11Q DLBCL (ARM); SEQ ID NO: 1626, 11Q DLBCL (ARM); SEQ ID NO: 1627, 11Q DLBCL (ARM); SEQ ID NO: 1628, 11Q DLBCL (ARM); SEQ ID NO: 1629, 11Q DLBCL (ARM); SEQ ID NO: 1630, 11Q DLBCL (ARM); SEQ ID NO: 1631, 11Q DLBCL (ARM); SEQ ID NO: 1632, 11Q DLBCL (ARM); SEQ ID NO: 1633, FP (FP); SEQ ID NO: 1634, 11Q DLBCL (ARM); SEQ ID NO: 1635, MSI (MSI); SEQ ID NO: 1636, MSI (MSI); SEQ ID NO: 1637, FP (FP); SEQ ID NO: 1638, 12P DLBCL (ARM); SEQ ID NO: 1639, 12P DLBCL (ARM); SEQ ID NO: 1640, 12P DLBCL (ARM); SEQ ID NO: 1641, 12P DLBCL (ARM); SEQ ID NO: 1642, 12P DLBCL (ARM); SEQ ID NO: 1643, 12P DLBCL (ARM); SEQ ID NO: 1644, 12P DLBCL (ARM); SEQ ID NO: 1645, 12P DLBCL (ARM); SEQ ID NO: 1646, PTPN6 (Gene); SEQ ID NO: 1647, PTPN6 (Gene); SEQ ID NO: 1648, PTPN6 (Gene); SEQ ID NO: 1649, PTPN6 (Gene); SEQ ID NO: 1650, PTPN6 (Gene); SEQ ID NO: 1651, PTPN6 (Gene); SEQ ID NO: 1652, PTPN6 (Gene); SEQ ID NO: 1653, PTPN6 (Gene); SEQ ID NO: 1654, PTPN6 (Gene); SEQ ID NO: 1655, PTPN6 (Gene); SEQ ID NO: 1656, PTPN6 (Gene); SEQ ID NO: 1657, PTPN6 (Gene); SEQ ID NO: 1658, PTPN6 (Gene); SEQ ID NO: 1659, PTPN6 (Gene); SEQ ID NO: 1660, PTPN6 (Gene); SEQ ID NO: 1661, PTPN6 (Gene); SEQ ID NO: 1662, 12P DLBCL (ARM); SEQ ID NO: 1663, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1664, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1665, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1666, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1667, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1668, ETV6 (Gene); SEQ ID NO: 1669, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1670, ETV6 SV (SV); SEQ ID NO: 1671, ETV6 SV (SV); SEQ ID NO: 1672, ETV6 SV (SV); SEQ ID NO: 1673, ETV6 SV (SV); SEQ ID NO: 1674, ETV6 SV (SV); SEQ ID NO: 1675, ETV6 SV (SV); SEQ ID NO: 1676, ETV6 SV (SV); SEQ ID NO: 1677, ETV6 SV (SV); SEQ ID NO: 1678, ETV6 SV (SV); SEQ ID NO: 1679, ETV6 SV (SV); SEQ ID NO: 1680, ETV6 SV (SV); SEQ ID NO: 1681, ETV6_SV (SV); SEQ ID NO: 1682, ETV6 SV (SV); SEQ ID NO: 1683, ETV6_SV (SV); SEQ ID NO: 1684, ETV6 SV (SV); SEQ ID NO: 1685, ETV6 SV (SV); SEQ ID NO: 1686, ETV6 SV (SV); SEQ ID NO: 1687, ETV6 SV (SV); SEQ ID NO: 1688, ETV6 SV (SV); SEQ ID NO: 1689, ETV6 SV (SV); SEQ ID NO: 1690, ETV6 SV (SV); SEQ ID NO: 1691, ETV6 SV (SV); SEQ ID NO: 1692, ETV6 SV (SV); SEQ ID NO: 1693, ETV6 SV (SV); SEQ ID NO: 1694, ETV6_SV (SV); SEQ ID NO: 1695, ETV6 SV (SV); SEQ ID NO: 1696, ETV6_SV(SV); SEQ ID NO: 1697, ETV6 SV (SV); SEQ ID NO: 1698, ETV6 SV (SV); SEQ ID NO: 1699, ETV6 SV (SV); SEQ ID NO: 1700, ETV6_SV (SV); SEQ ID NO: 1701, ETV6_SV (SV); SEQ ID NO: 1702, ETV6 SV (SV); SEQ ID NO: 1703, ETV6 SV (SV); SEQ ID NO: 1704, ETV6 SV (SV); SEQ ID NO: 1705, ETV6 SV (SV); SEQ ID NO: 1706, ETV6 SV (SV); SEQ ID NO: 1707, ETV6_SV (SV); SEQ ID NO: 1708, ETV6 SV (SV); SEQ ID NO: 1709,ETV6_SV (SV); SEQ ID NO: 1710, ETV6 SV (SV); SEQ ID NO: 1711, ETV6 SV (SV); SEQ ID NO: 1712, ETV6 SV (SV); SEQ ID NO: 1713, ETV6 SV (SV); SEQ ID NO: 1714, ETV6 SV (SV); SEQ ID NO: 1715, ETV6 SV (SV); SEQ ID NO: 1716, ETV6 SV (SV); SEQ ID NO: 1717, ETV6 SV (SV); SEQ ID NO: 1718, ETV6 SV (SV); SEQ ID NO: 1719, ETV6 SV (SV); SEQ ID NO: 1720, ETV6_SV (SV); SEQ ID NO: 1721, ETV6 SV (SV); SEQ ID NO: 1722,ETV6_SV (SV); SEQ ID NO: 1723, ETV6 SV (SV); SEQ ID NO: 1724, ETV6 SV (SV); SEQ ID NO: 1725, ETV6 SV (SV); SEQ ID NO: 1726, ETV6 SV (SV); SEQ ID NO: 1727, ETV6 SV (SV); SEQ ID NO: 1728, ETV6 SV (SV); SEQ ID NO: 1729, ETV6 SV (SV); SEQ ID NO: 1730, ETV6 SV (SV); SEQ ID NO: 1731, ETV6 SV (SV); SEQ ID NO: 1732, ETV6 SV (SV); SEQ ID NO: 1733, ETV6_SV (SV); SEQ ID NO: 1734, ETV6 SV (SV); SEQ ID NO: 1735,ETV6_SV (SV); SEQ ID NO: 1736, ETV6 SV (SV); SEQ ID NO: 1737, ETV6 SV (SV); SEQ ID NO: 1738, ETV6 SV (SV); SEQ ID NO: 1739, ETV6 SV (SV); SEQ ID NO: 1740, ETV6 SV (SV); SEQ ID NO: 1741, ETV6 SV (SV); SEQ ID NO: 1742, ETV6 SV (SV); SEQ ID NO: 1743, ETV6 SV (SV); SEQ ID NO: 1744, ETV6 SV (SV); SEQ ID NO: 1745, ETV6 SV (SV); SEQ ID NO: 1746, ETV6_SV (SV); SEQ ID NO: 1747, ETV6 SV (SV); SEQ ID NO: 1748,ETV6_SV (SV); SEQ ID NO: 1749, ETV6 SV (SV); SEQ ID NO: 1750, ETV6 SV (SV); SEQ ID NO: 1751, ETV6 SV (SV); SEQ ID NO: 1752, ETV6 SV (SV); SEQ ID NO: 1753, ETV6 SV (SV); SEQ ID NO: 1754, ETV6 SV (SV); SEQ ID NO: 1755, ETV6 SV (SV); SEQ ID NO: 1756, ETV6 SV (SV); SEQ ID NO: 1757, ETV6 SV (SV); SEQ ID NO: 1758, ETV6 SV (SV); SEQ ID NO: 1759, ETV6_SV (SV); SEQ ID NO: 1760, ETV6 SV (SV); SEQ ID NO: 1761,ETV6_SV (SV); SEQ ID NO: 1762, ETV6 SV (SV); SEQ ID NO: 1763, ETV6 SV (SV); SEQ ID NO: 1764, ETV6 SV (SV); SEQ ID NO: 1765, ETV6 SV (SV); SEQ ID NO: 1766, ETV6 SV (SV); SEQ ID NO: 1767, ETV6 SV (SV); SEQ ID NO: 1768, ETV6 SV (SV); SEQ ID NO: 1769, ETV6 SV (SV); SEQ ID NO: 1770, ETV6 SV (SV); SEQ ID NO: 1771, ETV6 SV (SV); SEQ ID NO: 1772, ETV6_SV (SV); SEQ ID NO: 1773, ETV6 SV (SV); SEQ ID NO: 1774,ETV6_SV (SV); SEQ ID NO: 1775, ETV6 SV (SV); SEQ ID NO: 1776, ETV6 SV (SV); SEQ ID NO: 1777, ETV6 SV (SV); SEQ ID NO: 1778, ETV6 SV (SV); SEQ ID NO: 1779, ETV6 SV (SV); SEQ ID NO: 1780, ETV6 SV (SV); SEQ ID NO: 1781, ETV6 SV (SV); SEQ ID NO: 1782, ETV6 SV (SV); SEQ ID NO: 1783, ETV6 SV (SV); SEQ ID NO: 1784, ETV6 SV (SV); SEQID NO: 1785, ETV6_SV (SV); SEQIDNO: 1786, ETV6 SV (SV); SEQIDNO: 1787,ETV6_SV (SV); SEQ ID NO: 1788, ETV6 SV (SV); SEQ ID NO: 1789, ETV6 SV (SV); SEQ ID NO: 1790, ETV6 SV (SV); SEQ ID NO: 1791, ETV6 SV (SV); SEQ ID NO: 1792, ETV6 SV (SV); SEQ ID NO: 1793, ETV6 SV (SV); SEQ ID NO: 1794, ETV6 SV (SV); SEQ ID NO: 1795, ETV6 SV (SV); SEQ ID NO: 1796, ETV6 SV (SV); SEQ ID NO: 1797, ETV6 SV (SV); SEQ ID NO: 1798, ETV6_SV (SV); SEQIDNO: 1799, ETV6 SV (SV); SEQIDNO: 1800,ETV6_SV (SV); SEQ ID NO: 1801, ETV6 SV (SV); SEQ ID NO: 1802, ETV6 SV (SV); SEQ ID NO: 1803, ETV6 SV (SV); SEQ ID NO: 1804, ETV6_SV (SV); SEQ ID NO: 1805, ETV6_SV (SV); SEQ ID NO: 1806, ETV6 SV (SV); SEQ ID NO: 1807, ETV6 SV (SV); SEQ ID NO: 1808, ETV6 SV (SV); SEQ ID NO: 1809, ETV6 SV (SV); SEQ ID NO: 1810, ETV6 SV (SV); SEQ ID NO: 1811, ETV6_SV(SV); SEQIDNO: 1812, ETV6 SV (SV); SEQIDNO: 1813,ETV6_SV (SV); SEQ ID NO: 1814, ETV6 SV (SV); SEQ ID NO: 1815, ETV6 SV (SV); SEQ ID NO: 1816, ETV6 SV (SV); SEQ ID NO: 1817, ETV6 SV (SV); SEQ ID NO: 1818, ETV6 SV (SV); SEQ ID NO: 1819, ETV6 SV (SV); SEQ ID NO: 1820, ETV6 SV (SV); SEQ ID NO: 1821, ETV6 SV (SV); SEQ ID NO: 1822, ETV6 SV (SV); SEQ ID NO: 1823, ETV6 SV (SV); SEQ ID NO: 1824, ETV6_SV (SV); SEQIDNO: 1825, ETV6 SV (SV); SEQIDNO: 1826,ETV6_SV (SV); SEQ ID NO: 1827, ETV6 SV (SV); SEQ ID NO: 1828, ETV6 SV (SV); SEQ ID NO: 1829, ETV6 SV (SV); SEQ ID NO: 1830, ETV6 SV (SV); SEQ ID NO: 1831, ETV6 SV (SV); SEQ ID NO: 1832, ETV6 SV (SV); SEQ ID NO: 1833, ETV6 SV (SV); SEQ ID NO: 1834, ETV6 SV (SV); SEQ ID NO: 1835, ETV6 SV (SV); SEQ ID NO: 1836, ETV6 SV (SV); SEQ ID NO: 1837, ETV6_SV (SV); SEQIDNO: 1838, ETV6 SV (SV); SEQIDNO: 1839,ETV6_SV (SV); SEQ ID NO: 1840, ETV6 SV (SV); SEQ ID NO: 1841, ETV6 SV (SV); SEQ ID NO: 1842, ETV6 SV (SV); SEQ ID NO: 1843, ETV6 SV (SV); SEQ ID NO: 1844, ETV6 SV (SV); SEQ ID NO: 1845, ETV6 SV (SV); SEQ ID NO: 1846, ETV6 SV (SV); SEQ ID NO: 1847, ETV6 SV (SV); SEQ ID NO: 1848, ETV6 SV (SV); SEQ ID NO: 1849, ETV6 SV (SV); SEQ ID NO: 1850, ETV6_SV (SV); SEQIDNO: 1851, ETV6 SV (SV); SEQIDNO: 1852,ETV6_SV (SV); SEQ ID NO: 1853, ETV6 SV (SV); SEQ ID NO: 1854, ETV6 SV (SV); SEQ ID NO: 1855, ETV6 SV (SV); SEQ ID NO: 1856, ETV6 SV (SV); SEQ ID NO: 1857, ETV6 SV (SV); SEQ ID NO: 1858, ETV6 SV (SV); SEQ ID NO: 1859, ETV6 SV (SV); SEQ ID NO: 1860, ETV6 SV (SV); SEQ ID NO: 1861, ETV6 SV (SV); SEQ ID NO: 1862, ETV6 SV (SV); SEQ ID NO: 1863, ETV6_SV (SV); SEQIDNO: 1864, ETV6 SV (SV); SEQIDNO: 1865,ETV6_SV (SV); SEQ ID NO: 1866, ETV6 SV (SV); SEQ ID NO: 1867, ETV6 SV (SV); SEQ ID NO: 1868, ETV6 SV (SV); SEQ ID NO: 1869, ETV6 SV (SV); SEQ ID NO: 1870, ETV6 SV (SV); SEQ ID NO: 1871, ETV6 SV (SV); SEQ ID NO: 1872, ETV6 SV (SV); SEQ ID NO: 1873,ETV6 SV (SV); SEQ ID NO: 1874, ETV6 SV (SV); SEQ ID NO: 1875, ETV6 SV (SV); SEQ ID NO: 1876, ETV6_SV (SV); SEQ ID NO: 1877, ETV6 SV (SV); SEQ ID NO: 1878, ETV6_SV (SV); SEQ ID NO: 1879, ETV6 SV (SV); SEQ ID NO: 1880, ETV6 SV (SV); SEQ ID NO: 1881, ETV6 SV (SV); SEQ ID NO: 1882, ETV6 SV (SV); SEQ ID NO: 1883, ETV6 SV (SV); SEQ ID NO: 1884, ETV6 SV (SV); SEQ ID NO: 1885, ETV6 SV (SV); SEQ ID NO: 1886, ETV6 SV (SV); SEQ ID NO: 1887, ETV6 SV (SV); SEQ ID NO: 1888, ETV6 SV (SV); SEQ ID NO: 1889, ETV6_SV (SV); SEQ ID NO: 1890, ETV6 SV (SV); SEQ ID NO: 1891, ETV6_SV (SV); SEQ ID NO: 1892, ETV6 SV (SV); SEQ ID NO: 1893, ETV6 SV (SV); SEQ ID NO: 1894, ETV6 SV (SV); SEQ ID NO: 1895, ETV6 SV (SV); SEQ ID NO: 1896, ETV6 SV (SV); SEQ ID NO: 1897, ETV6 SV (SV); SEQ ID NO: 1898, ETV6 SV (SV); SEQ ID NO: 1899, ETV6 SV (SV); SEQ ID NO: 1900, ETV6 SV (SV); SEQ ID NO: 1901, ETV6 SV (SV); SEQ ID NO: 1902, ETV6_SV (SV); SEQ ID NO: 1903, ETV6 SV (SV); SEQ ID NO: 1904, ETV6_SV (SV); SEQ ID NO: 1905, ETV6 SV (SV); SEQ ID NO: 1906, ETV6 SV (SV); SEQ ID NO: 1907, ETV6 SV (SV); SEQ ID NO: 1908, ETV6_SV (SV); SEQ ID NO: 1909, ETV6_SV (SV); SEQ ID NO: 1910, ETV6 SV (SV); SEQ ID NO: 1911, ETV6 SV (SV); SEQ ID NO: 1912, ETV6 SV (SV); SEQ ID NO: 1913, ETV6 SV (SV); SEQ ID NO: 1914, ETV6 SV (SV); SEQ ID NO: 1915, ETV6 (Gene); SEQ ID NO: 1916, ETV6 (Gene); SEQ ID NO: 1917, ETV6 (Gene); SEQ ID NO: 1918, ETV6 (Gene); SEQ ID NO: 1919, ETV6 (Gene); SEQ ID NO: 1920, ETV6 (Gene); SEQ ID NO: 1921, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1922, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1923, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1924, CDKN1B (Gene); SEQ ID NO: 1925, CDKN1B (Gene); SEQ ID NO: 1926, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1927, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1928, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1929, 12pl3.2_DLBCL (FOCAL); SEQ ID NO: 1930, 12P DLBCL (ARM); SEQ ID NO: 1931, 12P DLBCL (ARM); SEQ ID NO: 1932, 12P DLBCL (ARM); SEQ ID NO: 1933, 12P DLBCL (ARM); SEQ ID NO: 1934, 12P DLBCL (ARM); SEQ ID NO: 1935, 12P DLBCL (ARM); SEQ ID NO: 1936, FP (FP); SEQ ID NO: 1937, 12P DLBCL (ARM); SEQ ID NO: 1938, KRAS (Gene); SEQ ID NO: 1939, KRAS (Gene); SEQ ID NO: 1940, KRAS (Gene); SEQ ID NO: 1941, KRAS (Gene); SEQ ID NO: 1942, KRAS (Gene); SEQ ID NO: 1943, 12P DLBCL (ARM); SEQ ID NO: 1944, FP (FP); SEQ ID NO: 1945, 12P DLBCL (ARM); SEQ ID NO: 1946, 12P DLBCL (ARM); SEQ ID NO: 1947, 12P DLBCL (ARM); SEQ ID NO: 1948, 12P DLBCL (ARM); SEQ ID NO: 1949, 12P DLBCL (ARM); SEQ ID NO: 1950, 12P DLBCL (ARM); SEQ ID NO: 1951, 12Q (ARM); SEQ ID NO: 1952, 12Q (ARM); SEQ ID NO: 1953, 12Q (ARM); SEQ ID NO: 1954, 12Q (ARM); SEQ ID NO: 1955, 12Q (ARM); SEQ ID NO: 1956, 12Q (ARM); SEQ ID NO: 1957, 12Q (ARM); SEQ ID NO: 1958, 12Q (ARM); SEQ ID NO: 1959, FP (FP);SEQ ID NO: 1960, 12Q (ARM); SEQ ID NO: 1961, 12Q (ARM); SEQ ID NO: 1962, KMT2D (Gene); SEQ ID NO: 1963, KMT2D (Gene); SEQ ID NO: 1964, KMT2D (Gene); SEQ ID NO: 1965, KMT2D (Gene); SEQ ID NO: 1966, KMT2D (Gene); SEQ ID NO: 1967, KMT2D (Gene); SEQ ID NO: 1968, KMT2D (Gene); SEQ ID NO: 1969, KMT2D (Gene); SEQ ID NO: 1970, KMT2D (Gene); SEQ ID NO: 1971, KMT2D (Gene); SEQ ID NO: 1972, KMT2D (Gene); SEQ ID NO: 1973, KMT2D (Gene); SEQ ID NO: 1974, KMT2D (Gene); SEQ ID NO: 1975, KMT2D (Gene); SEQ ID NO: 1976, KMT2D (Gene); SEQ ID NO: 1977, KMT2D (Gene); SEQ ID NO: 1978, KMT2D (Gene); SEQ ID NO: 1979, KMT2D (Gene); SEQ ID NO: 1980, KMT2D (Gene); SEQ ID NO: 1981, KMT2D (Gene); SEQ ID NO: 1982, KMT2D (Gene); SEQ ID NO: 1983, KMT2D (Gene); SEQ ID NO: 1984, KMT2D (Gene); SEQ ID NO: 1985, KMT2D (Gene); SEQ ID NO: 1986, KMT2D (Gene); SEQ ID NO: 1987, KMT2D (Gene); SEQ ID NO: 1988, KMT2D (Gene); SEQ ID NO: 1989, KMT2D (Gene); SEQ ID NO: 1990, KMT2D (Gene); SEQ ID NO: 1991, KMT2D (Gene); SEQ ID NO: 1992, KMT2D (Gene); SEQ ID NO: 1993, KMT2D (Gene); SEQ ID NO: 1994, KMT2D (Gene); SEQ ID NO: 1995, KMT2D (Gene); SEQ ID NO: 1996, KMT2D (Gene); SEQ ID NO: 1997, KMT2D (Gene); SEQ ID NO: 1998, KMT2D (Gene); SEQ ID NO: 1999, KMT2D (Gene); SEQ ID NO: 2000, KMT2D (Gene); SEQ ID NO: 2001, KMT2D (Gene); SEQ ID NO: 2002, KMT2D (Gene); SEQ ID NO: 2003, KMT2D (Gene); SEQ ID NO: 2004, KMT2D (Gene); SEQ ID NO: 2005, KMT2D (Gene); SEQ ID NO: 2006, KMT2D (Gene); SEQ ID NO: 2007, KMT2D (Gene); SEQ ID NO: 2008, KMT2D (Gene); SEQ ID NO: 2009, KMT2D (Gene); SEQ ID NO: 2010, KMT2D (Gene); SEQ ID NO: 2011, KMT2D (Gene); SEQ ID NO: 2012, KMT2D (Gene); SEQ ID NO: 2013, KMT2D (Gene); SEQ ID NO: 2014, KMT2D (Gene); SEQ ID NO: 2015, KMT2D (Gene); SEQ ID NO: 2016, KMT2D (Gene); SEQ ID NO: 2017, KMT2D (Gene); SEQ ID NO: 2018, KMT2D (Gene); SEQ ID NO: 2019, KMT2D (Gene); SEQ ID NO: 2020, KMT2D (Gene); SEQ ID NO: 2021, KMT2D (Gene); SEQ ID NO: 2022, KMT2D (Gene); SEQ ID NO: 2023, KMT2D (Gene); SEQ ID NO: 2024, KMT2D (Gene); SEQ ID NO: 2025, KMT2D (Gene); SEQ ID NO: 2026, KMT2D (Gene); SEQ ID NO: 2027, KMT2D (Gene); SEQ ID NO: 2028, KMT2D (Gene); SEQ ID NO: 2029, KMT2D (Gene); SEQ ID NO: 2030, KMT2D (Gene); SEQ ID NO: 2031, KMT2D (Gene); SEQ ID NO: 2032, KMT2D (Gene); SEQ ID NO: 2033, KMT2D (Gene); SEQ ID NO: 2034, KMT2D (Gene); SEQ ID NO: 2035, KMT2D (Gene); SEQ ID NO: 2036, KMT2D (Gene); SEQ ID NO: 2037, KMT2D (Gene); SEQ ID NO: 2038, KMT2D (Gene); SEQ ID NO: 2039, KMT2D (Gene); SEQ ID NO: 2040, KMT2D (Gene); SEQ ID NO: 2041, 12Q (ARM); SEQ ID NO: 2042, 12Q (ARM); SEQ ID NO: 2043, 12Q (ARM); SEQ ID NO: 2044, MSI (MSI); SEQ ID NO: 2045, MSI (MSI); SEQ ID NO: 2046, MSI (MSI); SEQ ID NO: 2047, MSI (MSI); SEQ ID NO: 2048, STAT6 (Gene); SEQ ID NO:2049, STAT6 (Gene); SEQ ID NO: 2050, STAT6 (Gene); SEQ ID NO: 2051, STAT6 (Gene); SEQ ID NO: 2052, STAT6 (Gene); SEQ ID NO: 2053, STAT6 (Gene); SEQ ID NO: 2054, STAT6 (Gene); SEQ ID NO: 2055, STAT6 (Gene); SEQ ID NO: 2056, STAT6 (Gene); SEQ ID NO: 2057, STAT6 (Gene); SEQ ID NO: 2058, STAT6 (Gene); SEQ ID NO: 2059, STAT6 (Gene); SEQ ID NO: 2060, STAT6 (Gene); SEQ ID NO: 2061, STAT6 (Gene); SEQ ID NO: 2062, STAT6 (Gene); SEQ ID NO: 2063, STAT6 (Gene); SEQ ID NO: 2064, STAT6 (Gene); SEQ ID NO: 2065, STAT6 (Gene); SEQ ID NO: 2066, STAT6 (Gene); SEQ ID NO: 2067, STAT6 (Gene); SEQ ID NO: 2068, STAT6 (Gene); SEQ ID NO: 2069, 12Q (ARM); SEQ ID NO: 2070, 12Q (ARM); SEQ ID NO: 2071, 12Q (ARM); SEQ ID NO: 2072, 12Q (ARM); SEQ ID NO: 2073, 12Q (ARM); SEQ ID NO: 2074, 12Q (ARM); SEQ ID NO: 2075, 12Q (ARM); SEQ ID NO: 2076, 12Q (ARM); SEQ ID NO: 2077, 12Q (ARM); SEQ ID NO: 2078, 12Q (ARM); SEQ ID NO: 2079, 12Q (ARM); SEQ ID NO: 2080, 12Q (ARM); SEQ ID NO: 2081, 12Q (ARM); SEQ ID NO: 2082, 12Q (ARM); SEQ ID NO: 2083, 12Q (ARM); SEQ ID NO: 2084, 12Q (ARM); SEQ ID NO: 2085, 12Q (ARM); SEQ ID NO: 2086, 12Q (ARM); SEQ ID NO: 2087, 12Q (ARM); SEQ ID NO: 2088, 12Q (ARM); SEQ ID NO: 2089, 12Q (ARM); SEQ ID NO: 2090, 12Q (ARM); SEQ ID NO: 2091, 12Q (ARM); SEQ ID NO: 2092, BTG1 (Gene); SEQ ID NO: 2093, BTG1 (Gene); SEQ ID NO: 2094, BTG1 (Gene); SEQ ID NO: 2095, BTG1 (Gene); SEQ ID NO: 2096, 12Q (ARM); SEQ ID NO: 2097, 12Q (ARM); SEQ ID NO: 2098, 12Q (ARM); SEQ ID NO: 2099, 12Q (ARM); SEQ ID NO: 2100, 12Q (ARM); SEQ ID NO: 2101, 12Q (ARM); SEQ ID NO: 2102, 12Q (ARM); SEQ ID NO : 2103 , 12Q (ARM); SEQ ID NO : 2104, 12Q (ARM); SEQ ID NO : 2105, 12Q (ARM); SEQ ID NO: 2106, FP (FP); SEQ ID NO: 2107, 12Q (ARM); SEQ ID NO: 2108, 12Q (ARM); SEQ ID NO : 2109, 12Q (ARM); SEQ ID NO : 2110, 12Q (ARM); SEQ ID NO : 2111 , 12Q (ARM); SEQ ID NO: 2112, 12Q (ARM); SEQ ID NO: 2113, 12Q (ARM); SEQ ID NO: 2114, HVCN1 (Gene); SEQ ID NO: 2115, HVCN1 (Gene); SEQ ID NO: 2116, HVCN1 (Gene); SEQ ID NO: 2117, HVCN1 (Gene); SEQ ID NO: 2118, HVCN1 (Gene); SEQ ID NO: 2119, HVCN1 (Gene); SEQ ID NO : 2120, 12Q (ARM); SEQ ID NO : 2121 , 12Q (ARM); SEQ ID NO : 2122, 12Q (ARM); SEQ ID NO: 2123, DTX1 (Gene); SEQ ID NO: 2124, DTX1 (Gene); SEQ ID NO: 2125, DTX1 (Gene); SEQ ID NO: 2126, DTX1 (Gene); SEQ ID NO: 2127, DTX1 (Gene); SEQ ID NO: 2128, DTX1 (Gene); SEQ ID NO: 2129, DTX1 (Gene); SEQ ID NO: 2130, DTX1 (Gene); SEQ ID NO: 2131, DTX1 (Gene); SEQ ID NO: 2132, DTX1 (Gene); SEQ ID NO: 2133, 12Q (ARM); SEQ ID NO: 2134, 12Q (ARM); SEQ ID NO: 2135, 12Q (ARM); SEQ ID NO: 2136, 12Q (ARM); SEQ ID NO: 2137, 12Q (ARM); SEQ ID NO: 2138, 12Q (ARM); SEQ ID NO: 2139, 12Q (ARM); SEQ ID NO : 2140, 12Q (ARM); SEQ ID NO : 2141 , 12Q (ARM); SEQ ID NO : 2142, 12Q (ARM); SEQ ID NO: 2143, 12Q (ARM); SEQ ID NO: 2144, SETD1B (Gene); SEQ ID NO: 2145,SETD1B (Gene); SEQ ID NO: 2146, SETD1B (Gene); SEQ ID NO: 2147, SETD1B (Gene); SEQ ID NO: 2148, SETD1B (Gene); SEQ ID NO: 2149, SETD1B (Gene); SEQ ID NO: 2150, SETD1B (Gene); SEQ ID NO: 2151, SETD1B (Gene); SEQ ID NO: 2152, SETD1B (Gene); SEQ ID NO: 2153, SETD1B (Gene); SEQ ID NO: 2154, SETD1B (Gene); SEQ ID NO: 2155, SETD1B (Gene); SEQ ID NO: 2156, SETD1B (Gene); SEQ ID NO: 2157, SETD1B (Gene); SEQ ID NO: 2158, SETD1B (Gene); SEQ ID NO: 2159, SETD1B (Gene); SEQ ID NO: 2160, SETD1B (Gene); SEQ ID NO: 2161, SETD1B (Gene); SEQ ID NO: 2162, SETD1B (Gene); SEQ ID NO: 2163, SETD1B (Gene); SEQ ID NO: 2164, SETD1B (Gene); SEQ ID NO: 2165, SETD1B (Gene); SEQ ID NO: 2166, SETD1B (Gene); SEQ ID NO: 2167, SETD1B (Gene); SEQ ID NO: 2168, BCL7A (Gene); SEQ ID NO: 2169, BCL7A (Gene); SEQ ID NO: 2170, BCL7A (Gene); SEQ ID NO: 2171, BCL7A (Gene); SEQ ID NO: 2172, BCL7A (Gene); SEQ ID NO: 2173, BCL7A (Gene); SEQ ID NO: 2174, 12Q (ARM); SEQ ID NO: 2175, 12Q (ARM); SEQ ID NO: 2176, 12Q (ARM); SEQ ID NO: 2177, 12Q (ARM); SEQ ID NO: 2178, 12Q (ARM); SEQ ID NO: 2179, 12Q (ARM); SEQ ID NO : 2180, 12Q (ARM); SEQ ID NO : 2181 , 12Q (ARM); SEQ ID NO : 2182, 12Q (ARM); SEQ ID NO : 2183 , 12Q (ARM); SEQ ID NO : 2184, 12Q (ARM); SEQ ID NO : 2185, 12Q (ARM); SEQ ID NO: 2186, 13q-ARM_CLL (ARM); SEQ ID NO: 2187, 13q-ARM_CLL (ARM); SEQ ID NO: 2188, FP (FP); SEQ ID NO: 2189, 13q-ARM_CLL (ARM); SEQ ID NO: 2190, 13q- ARM CLL (ARM); SEQ ID NO: 2191, 13q-ARM_CLL (ARM); SEQ ID NO: 2192, 13q- ARM CLL (ARM); SEQ ID NO: 2193, 13q-ARM_CLL (ARM); SEQ ID NO: 2194, 13q- ARM CLL (ARM); SEQ ID NO: 2195, PABPC3 (Gene); SEQ ID NO: 2196, PABPC3 (Gene); SEQ ID NO: 2197, PABPC3 (Gene); SEQ ID NO: 2198, PABPC3 (Gene); SEQ ID NO: 2199, 13q-ARM_CLL (ARM); SEQ ID NO: 2200, 13q-ARM_CLL (ARM); SEQ ID NO: 2201, 13 q- ARM CLL (ARM); SEQ ID NO: 2202, 13q-ARM_CLL (ARM); SEQ ID NO: 2203, 13ql2.3- ql3.1_MCL (FOCAL); SEQ ID NO: 2204, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2205, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2206, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2207, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2208, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2209, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2210, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2211, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2212, 13ql2.3- ql3.1_MCL (FOCAL); SEQ ID NO: 2213, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2214, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2215, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2216, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2217, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2218, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2219, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2220, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2221, 13ql2.3- ql3.1_MCL (FOCAL); SEQ ID NO: 2222, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2223,13ql2.3-ql3.1_MCL (FOCAL); SEQ IDNO: 2224, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2225, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2226, BRCA2 (Gene); SEQ ID NO: 2227, BRCA2 (Gene); SEQ ID NO: 2228, BRCA2 (Gene); SEQ ID NO: 2229, BRCA2 (Gene); SEQ ID NO: 2230, BRCA2 (Gene); SEQ ID NO: 2231, BRCA2 (Gene); SEQ ID NO: 2232, BRCA2 (Gene); SEQ ID NO: 2233, BRCA2 (Gene); SEQ ID NO: 2234, BRCA2 (Gene); SEQ ID NO: 2235, BRCA2 (Gene); SEQ ID NO: 2236, BRCA2 (Gene); SEQ ID NO: 2237, BRCA2 (Gene); SEQ ID NO: 2238, BRCA2 (Gene); SEQ ID NO: 2239, BRCA2 (Gene); SEQ ID NO: 2240, BRCA2 (Gene); SEQ ID NO: 2241, BRCA2 (Gene); SEQ ID NO: 2242, BRCA2 (Gene); SEQ ID NO: 2243, BRCA2 (Gene); SEQ ID NO: 2244, BRCA2 (Gene); SEQ ID NO: 2245, BRCA2 (Gene); SEQ ID NO: 2246, BRCA2 (Gene); SEQ ID NO: 2247, BRCA2 (Gene); SEQ ID NO: 2248, BRCA2 (Gene); SEQ ID NO: 2249, BRCA2 (Gene); SEQ ID NO: 2250, BRCA2 (Gene); SEQ ID NO: 2251, BRCA2 (Gene); SEQ ID NO: 2252, BRCA2 (Gene); SEQ ID NO: 2253, BRCA2 (Gene); SEQ ID NO: 2254, BRCA2 (Gene); SEQ ID NO: 2255, BRCA2 (Gene); SEQ ID NO: 2256, BRCA2 (Gene); SEQ ID NO: 2257, BRCA2 (Gene); SEQ ID NO: 2258, BRCA2 (Gene); SEQ ID NO: 2259, BRCA2 (Gene); SEQ ID NO: 2260, BRCA2 (Gene); SEQ ID NO: 2261, BRCA2 (Gene); SEQ ID NO: 2262, BRCA2 (Gene); SEQ ID NO: 2263, BRCA2 (Gene); SEQ ID NO: 2264, BRCA2 (Gene); SEQ ID NO: 2265, BRCA2 (Gene); SEQ ID NO: 2266, BRCA2 (Gene); SEQ ID NO: 2267, BRCA2 (Gene); SEQ ID NO: 2268, BRCA2 (Gene); SEQ ID NO: 2269, BRCA2 (Gene); SEQ ID NO: 2270, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2271, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2272, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2273, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2274, 13ql2.3-ql3.1_MCL (FOCAL); SEQ ID NO: 2275, 13q-ARM_CLL (ARM); SEQ ID NO: 2276, 13q-ARM_CLL (ARM); SEQ ID NO: 2277, FP (FP); SEQ ID NO: 2278, 13q-ARM_CLL (ARM); SEQ ID NO: 2279, 13q-ARM_CLL (ARM); SEQ ID NO: 2280, 13q-ARM_CLL (ARM); SEQ ID NO: 2281, 13q-ARM_CLL (ARM); SEQ ID NO: 2282, 13q-ARM_CLL (ARM); SEQ ID NO: 2283, FOXO1 (Gene); SEQ ID NO: 2284, FOXO1 (Gene); SEQ ID NO: 2285, FOXO1 (Gene); SEQ ID NO: 2286, FOXO1 (Gene); SEQ ID NO: 2287, FOXO1 (Gene); SEQ ID NO: 2288, 13q-ARM_CLL (ARM); SEQ ID NO: 2289, 13q-ARM_CLL (ARM); SEQ ID NO: 2290, 13q-ARM_CLL (ARM); SEQ ID NO: 2291, 13q-ARM_CLL (ARM); SEQ ID NO: 2292, 13q-ARM_CLL (ARM); SEQ ID NO: 2293, 13q-ARM_CLL (ARM); SEQ ID NO: 2294, 13q-ARM_CLL (ARM); SEQ ID NO: 2295, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2296, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2297, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2298, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2299, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2300, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2301, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2302, FP (FP); SEQ ID NO: 2303,13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2304, RBI (Gene); SEQ ID NO: 2305, RBI (Gene); SEQ ID NO: 2306, RBI (Gene); SEQ ID NO: 2307, RBI (Gene); SEQ ID NO: 2308, RBI (Gene); SEQ ID NO: 2309, RBI (Gene); SEQ ID NO: 2310, RBI (Gene); SEQ ID NO: 2311, RBI (Gene); SEQ ID NO: 2312, RBI (Gene); SEQ ID NO: 2313, RBI (Gene); SEQ ID NO: 2314, RBI (Gene); SEQ ID NO: 2315, RBI (Gene); SEQ ID NO: 2316, RBI (Gene); SEQ ID NO: 2317, RBI (Gene); SEQ ID NO: 2318, RBI (Gene); SEQ ID NO: 2319, RBI (Gene); SEQ ID NO: 2320, RBI (Gene); SEQ ID NO: 2321, RBI (Gene); SEQ ID NO: 2322, RBI (Gene); SEQ ID NO: 2323, RBI (Gene); SEQ ID NO: 2324, RBI (Gene); SEQ ID NO: 2325, RBI (Gene); SEQ ID NO: 2326, RBI (Gene); SEQ ID NO: 2327, RBI (Gene); SEQ ID NO: 2328, RBI (Gene); SEQ ID NO: 2329, RBI (Gene); SEQ ID NO: 2330, RBI (Gene); SEQ ID NO: 2331, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2332, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2333, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2334, SETDB2 (Gene); SEQ ID NO: 2335, SETDB2 (Gene); SEQ ID NO: 2336, SETDB2 (Gene); SEQ ID NO: 2337, SETDB2 (Gene); SEQ ID NO: 2338, SETDB2 (Gene); SEQ ID NO: 2339, SETDB2 (Gene); SEQ ID NO: 2340, SETDB2 (Gene); SEQ ID NO: 2341, SETDB2 (Gene); SEQ ID NO: 2342, SETDB2 (Gene); SEQ ID NO: 2343, SETDB2 (Gene); SEQ ID NO: 2344, SETDB2 (Gene); SEQ ID NO: 2345, SETDB2 (Gene); SEQ ID NO: 2346, SETDB2 (Gene); SEQ ID NO: 2347, SETDB2 (Gene); SEQ ID NO: 2348, SETDB2 (Gene); SEQ ID NO: 2349, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2350, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2351, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2352, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2353, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2354, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2355, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2356, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2357, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2358, 13ql4.2_DLBCL (FOCAL); SEQ ID NO: 2359, 13q-ARM_CLL (ARM); SEQ ID NO: 2360, 13q-ARM_CLL (ARM); SEQ ID NO: 2361, 13 q- ARM CLL (ARM); SEQ ID NO: 2362, 13q-ARM_CLL (ARM); SEQ ID NO: 2363, 13q-ARM CLL (ARM); SEQ ID NO: 2364, 13q-ARM_CLL (ARM); SEQ ID NO: 2365, 13q-ARM CLL (ARM); SEQ ID NO: 2366, 13q-ARM_CLL (ARM); SEQ ID NO: 2367, 13q-ARM CLL (ARM); SEQ ID NO: 2368, FP (FP); SEQ ID NO: 2369, 13q-ARM_CLL (ARM);SEQ ID NO: 2370, 13q-ARM_CLL (ARM); SEQ ID NO: 2371, 13q-ARM_CLL (ARM); SEQ ID NO: 2372, 13q-ARM_CLL (ARM); SEQ ID NO: 2373, 13q-ARM_CLL (ARM); SEQ ID NO: 2374, 13q-ARM_CLL (ARM); SEQ ID NO: 2375, 13q-ARM_CLL (ARM); SEQ ID NO: 2376, 13q-ARM_CLL (ARM); SEQ ID NO: 2377, 13q-ARM_CLL (ARM); SEQ ID NO: 2378, MYCBP2 (Gene); SEQ ID NO: 2379, MYCBP2 (Gene); SEQ ID NO: 2380, MYCBP2 (Gene); SEQ ID NO: 2381, MYCBP2 (Gene); SEQ ID NO: 2382, MYCBP2 (Gene); SEQ ID NO: 2383, MYCBP2 (Gene); SEQ ID NO: 2384, MYCBP2 (Gene); SEQ ID NO: 2385, MYCBP2 (Gene);SEQ ID NO: 2386, MYCBP2 (Gene); SEQ ID NO: 2387, MYCBP2 (Gene); SEQ ID NO: 2388, MYCBP2 (Gene); SEQ ID NO: 2389, MYCBP2 (Gene); SEQ ID NO: 2390, MYCBP2 (Gene); SEQ ID NO: 2391, MYCBP2 (Gene); SEQ ID NO: 2392, MYCBP2 (Gene); SEQ ID NO: 2393, MYCBP2 (Gene); SEQ ID NO: 2394, MYCBP2 (Gene); SEQ ID NO: 2395, MYCBP2 (Gene); SEQ ID NO: 2396, MYCBP2 (Gene); SEQ ID NO: 2397, MYCBP2 (Gene); SEQ ID NO: 2398, MYCBP2 (Gene); SEQ ID NO: 2399, MYCBP2 (Gene); SEQ ID NO: 2400, MYCBP2 (Gene); SEQ ID NO: 2401, MYCBP2 (Gene); SEQ ID NO: 2402, MYCBP2 (Gene); SEQ ID NO: 2403, MYCBP2 (Gene); SEQ ID NO: 2404, MYCBP2 (Gene); SEQ ID NO: 2405, MYCBP2 (Gene); SEQ ID NO: 2406, MYCBP2 (Gene); SEQ ID NO: 2407, MYCBP2 (Gene); SEQ ID NO: 2408, MYCBP2 (Gene); SEQ ID NO: 2409, MYCBP2 (Gene); SEQ ID NO: 2410, MYCBP2 (Gene); SEQ ID NO: 2411, MYCBP2 (Gene); SEQ ID NO: 2412, MYCBP2 (Gene); SEQ ID NO: 2413, MYCBP2 (Gene); SEQ ID NO: 2414, MYCBP2 (Gene); SEQ ID NO: 2415, MYCBP2 (Gene); SEQ ID NO: 2416, MYCBP2 (Gene); SEQ ID NO: 2417, MYCBP2 (Gene); SEQ ID NO: 2418, MYCBP2 (Gene); SEQ ID NO: 2419, MYCBP2 (Gene); SEQ ID NO: 2420, MYCBP2 (Gene); SEQ ID NO: 2421, MYCBP2 (Gene); SEQ ID NO: 2422, MYCBP2 (Gene); SEQ ID NO: 2423, MYCBP2 (Gene); SEQ ID NO: 2424, MYCBP2 (Gene); SEQ ID NO: 2425, MYCBP2 (Gene); SEQ ID NO: 2426, MYCBP2 (Gene); SEQ ID NO: 2427, MYCBP2 (Gene); SEQ ID NO: 2428, MYCBP2 (Gene); SEQ ID NO: 2429, MYCBP2 (Gene); SEQ ID NO: 2430, MYCBP2 (Gene); SEQ ID NO: 2431, MYCBP2 (Gene); SEQ ID NO: 2432, MYCBP2 (Gene); SEQ ID NO: 2433, MYCBP2 (Gene); SEQ ID NO: 2434, MYCBP2 (Gene); SEQ ID NO: 2435, MYCBP2 (Gene); SEQ ID NO: 2436, MYCBP2 (Gene); SEQ ID NO: 2437, MYCBP2 (Gene); SEQ ID NO: 2438, MYCBP2 (Gene); SEQ ID NO: 2439, MYCBP2 (Gene); SEQ ID NO: 2440, MYCBP2 (Gene); SEQ ID NO: 2441, MYCBP2 (Gene); SEQ ID NO: 2442, MYCBP2 (Gene); SEQ ID NO: 2443, MYCBP2 (Gene); SEQ ID NO: 2444, MYCBP2 (Gene); SEQ ID NO: 2445, MYCBP2 (Gene); SEQ ID NO: 2446, MYCBP2 (Gene); SEQ ID NO: 2447, MYCBP2 (Gene); SEQ ID NO: 2448, MYCBP2 (Gene); SEQ ID NO: 2449, MYCBP2 (Gene); SEQ ID NO: 2450, MYCBP2 (Gene); SEQ ID NO: 2451, MYCBP2 (Gene); SEQ ID NO: 2452, MYCBP2 (Gene); SEQ ID NO: 2453, MYCBP2 (Gene); SEQ ID NO: 2454, MYCBP2 (Gene); SEQ ID NO: 2455, MYCBP2 (Gene); SEQ ID NO: 2456, MYCBP2 (Gene); SEQ ID NO: 2457, MYCBP2 (Gene); SEQ ID NO: 2458, MYCBP2 (Gene); SEQ ID NO: 2459, MYCBP2 (Gene); SEQ ID NO: 2460, MYCBP2 (Gene); SEQ ID NO: 2461, MYCBP2 (Gene); SEQ ID NO: 2462, MYCBP2 (Gene); SEQ ID NO: 2463, MYCBP2 (Gene); SEQ ID NO: 2464, MYCBP2 (Gene); SEQ ID NO: 2465, MYCBP2 (Gene); SEQ ID NO: 2466, 13q-ARM_CLL (ARM); SEQ ID NO: 2467, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2468, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2469, 13q31.3_DLBCL (FOCAL); SEQID NO: 2470, 13q31.3 DLBCL (FOCAL); SEQ ID NO: 2471, 13q31.3_DLBCL (FOCAL); SEQID NO: 2472, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2473, 13q31.3_DLBCL (FOCAL); SEQID NO: 2474, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2475, 13q31.3_DLBCL (FOCAL); SEQID NO: 2476, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2477, 13q31.3_DLBCL (FOCAL); SEQID NO: 2478, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2479, 13q31.3_DLBCL (FOCAL); SEQID NO: 2480, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2481, 13q31.3_DLBCL (FOCAL); SEQID NO: 2482, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2483, 13q31.3_DLBCL (FOCAL); SEQID NO: 2484, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2485, 13q31.3_DLBCL (FOCAL); SEQID NO: 2486, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2487, 13q31.3_DLBCL (FOCAL); SEQID NO: 2488, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2489, 13q31.3_DLBCL (FOCAL); SEQID NO: 2490, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2491, 13q31.3_DLBCL (FOCAL); SEQID NO: 2492, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2493, 13q31.3_DLBCL (FOCAL); SEQID NO: 2494, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2495, 13q31.3_DLBCL (FOCAL); SEQID NO: 2496, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2497, 13q31.3_DLBCL (FOCAL); SEQID NO: 2498, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2499, 13q31.3_DLBCL (FOCAL); SEQID NO: 2500, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2501, 13q31.3_DLBCL (FOCAL); SEQID NO: 2502, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2503, 13q31.3_DLBCL (FOCAL); SEQID NO: 2504, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2505, 13q31.3_DLBCL (FOCAL); SEQID NO: 2506, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2507, 13q31.3_DLBCL (FOCAL); SEQID NO: 2508, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2509, 13q31.3_DLBCL (FOCAL); SEQID NO: 2510, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2511, 13q31.3_DLBCL (FOCAL); SEQID NO: 2512, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2513, 13q31.3_DLBCL (FOCAL); SEQID NO: 2514, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2515, 13q31.3_DLBCL (FOCAL); SEQID NO: 2516, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2517, 13q31.3_DLBCL (FOCAL); SEQID NO: 2518, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2519, 13q31.3_DLBCL (FOCAL); SEQID NO: 2520, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2521, 13q31.3_DLBCL (FOCAL); SEQID NO: 2522, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2523, 13q31.3_DLBCL (FOCAL); SEQID NO: 2524, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2525, 13q31.3_DLBCL (FOCAL); SEQID NO: 2526, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2527, 13q31.3_DLBCL (FOCAL); SEQID NO: 2528, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2529, 13q31.3_DLBCL (FOCAL); SEQID NO: 2530, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2531, 13q31.3_DLBCL (FOCAL); SEQID NO: 2532, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2533, 13q31.3_DLBCL (FOCAL); SEQID NO: 2534, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2535, 13q31.3_DLBCL (FOCAL); SEQID NO: 2536, 13q31. 3 DLBCL (FOCAL); SEQ ID NO: 2537, 13q31.3_DLBCL (FOCAL); SEQID NO: 2538, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2539, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2540, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2541, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2542, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2543, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2544, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2545, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2546, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2547, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2548, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2549, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2550, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2551, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2552, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2553, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2554, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2555, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2556, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2557, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2558, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2559, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2560, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2561, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2562, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2563, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2564, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2565, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2566, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2567, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2568, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2569, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2570, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2571, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2572, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2573, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2574, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2575, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2576, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2577, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2578, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2579, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2580, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2581, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2582, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2583, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2584, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2585, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2586, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2587, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2588, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2589, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2590, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2591, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2592, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2593, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2594, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2595, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2596, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2597, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2598, 13q31.3_DLBCL (FOCAL); SEQIDNO: 2599, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2600, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2601, FP (FP); SEQ ID NO: 2602, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2603, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2604, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2605, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2606,13q31.3_DLBCL (FOCAL); SEQ ID NO: 2607, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2608, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2609, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2610, 13q31.3_DLBCL (FOCAL); SEQ ID NO: 2611 , 13q34_DLBCL (FOCAL); SEQ ID NO: 2612, 13q34_DLBCL (FOCAL); SEQ ID NO: 2613, 13q34_DLBCL (FOCAL); SEQ ID NO: 2614, 13q34_DLBCL (FOCAL); SEQ ID NO: 2615, 13q34_DLBCL (FOCAL); SEQ ID NO: 2616, 13q34_DLBCL (FOCAL); SEQ ID NO: 2617, 13q34_DLBCL (FOCAL); SEQ ID NO: 2618, 13q34_DLBCL (FOCAL); SEQ ID NO: 2619, 13q34_DLBCL (FOCAL); SEQ ID NO: 2620, 13q34_DLBCL (FOCAL); SEQ ID NO: 2621, 13q34_DLBCL (FOCAL); SEQ ID NO: 2622, 13q34_DLBCL (FOCAL); SEQ ID NO: 2623, 13q34_DLBCL (FOCAL); SEQ ID NO: 2624, 13q34_DLBCL (FOCAL); SEQ ID NO: 2625, 13q34_DLBCL (FOCAL); SEQ ID NO: 2626, 13q34_DLBCL (FOCAL); SEQ ID NO: 2627, 13q34_DLBCL (FOCAL); SEQ ID NO: 2628, 13q34_DLBCL (FOCAL); SEQ ID NO: 2629, 13q34_DLBCL (FOCAL); SEQ ID NO: 2630, 13q34_DLBCL (FOCAL); SEQ ID NO: 2631, 13q34_DLBCL (FOCAL); SEQ ID NO: 2632, 13q34_DLBCL (FOCAL); SEQ ID NO: 2633, 13q34_DLBCL (FOCAL); SEQ ID NO: 2634, CHD8 (Gene); SEQ ID NO: 2635, CHD8 (Gene) ; SEQ ID NO: 2636, CHD8 (Gene); SEQ ID NO: 2637, CHD8 (Gene); SEQ ID NO: 2638, CHD8 (Gene); SEQ ID NO: 2639, CHD8 (Gene); SEQID NO: 2640, CHD8 (Gene); SEQ ID NO: 2641, CHD8 (Gene); SEQ ID NO: 2642, CHD8 (Gene); SEQ ID NO: 2643, CHD8 (Gene); SEQ ID NO: 2644, CHD8 (Gene); SEQ ID NO: 2645, CHD8 (Gene); SEQ ID NO: 2646, CHD8 (Gene); SEQ ID NO: 2647, CHD8 (Gene); SEQ ID NO: 2648, CHD8 (Gene); SEQ ID NO: 2649, CHD8 (Gene); SEQ ID NO: 2650, CHD8 (Gene); SEQ ID NO: 2651, CHD8 (Gene); SEQ ID NO: 2652, CHD8 (Gene); SEQ ID NO: 2653, CHD8 (Gene); SEQ ID NO: 2654, CHD8 (Gene); SEQ ID NO: 2655, CHD8 (Gene); SEQ ID NO: 2656, CHD8 (Gene); SEQ ID NO: 2657, CHD8 (Gene); SEQ ID NO: 2658, CHD8 (Gene); SEQ ID NO: 2659, CHD8 (Gene); SEQ ID NO: 2660, CHD8 (Gene); SEQ ID NO: 2661, CHD8 (Gene); SEQ ID NO: 2662, CHD8 (Gene); SEQ ID NO: 2663, CHD8 (Gene); SEQ ID NO: 2664, CHD8 (Gene); SEQ ID NO: 2665, CHD8 (Gene); SEQ ID NO: 2666, CHD8 (Gene); SEQ ID NO: 2667, CHD8 (Gene); SEQ ID NO: 2668, CHD8 (Gene); SEQ ID NO: 2669, CHD8 (Gene); SEQ ID NO: 2670, CHD8 (Gene); SEQ ID NO: 2671, CHD8 (Gene); SEQ ID NO: 2672, CHD8 (Gene); SEQ ID NO: 2673, CHD8 (Gene); SEQ ID NO: 2674, CHD8 (Gene); SEQ ID NO: 2675, CHD8 (Gene); SEQ ID NO: 2676, CHD8 (Gene); SEQ ID NO: 2677, CHD8 (Gene); SEQ ID NO: 2678, 14Q (ARM); SEQ ID NO: 2679, 14Q (ARM); SEQ ID NO: 2680, 14Q (ARM); SEQ ID NO: 2681, FP (FP); SEQ ID NO: 2682, 14Q (ARM); SEQ ID NO: 2683, 14Q (ARM); SEQ ID NO: 2684, FP (FP); SEQ ID NO: 2685, 14Q (ARM); SEQ ID NO: 2686, 14Q (ARM); SEQ ID NO: 2687, 14Q (ARM); SEQ ID NO: 2688, 14Q (ARM); SEQ ID NO: 2689, FP (FP); SEQ ID NO: 2690, 14Q (ARM); SEQ IDNO: 2691, 14Q (ARM); SEQ ID NO: 2692, NFKBIA (Gene); SEQ ID NO: 2693, NFKBIA (Gene); SEQ ID NO: 2694, NFKBIA (Gene); SEQ ID NO: 2695, NFKBIA (Gene); SEQ ID NO: 2696, NFKBIA (Gene); SEQ ID NO: 2697, NFKBIA (Gene); SEQ ID NO: 2698, 14Q (ARM); SEQ ID NO: 2699, 14Q (ARM); SEQ ID NO: 2700, 14Q (ARM); SEQ ID NO: 2701, 14Q (ARM); SEQ ID NO: 2702, 14Q (ARM); SEQ ID NO: 2703, 14Q (ARM); SEQ ID NO: 2704, 14Q (ARM); SEQ ID NO: 2705, 14Q (ARM); SEQ ID NO: 2706, 14Q (ARM); SEQ ID NO: 2707, 14Q (ARM); SEQ ID NO: 2708, 14Q (ARM); SEQ ID NO: 2709, 14Q (ARM); SEQ ID NO: 2710, 14Q (ARM); SEQ ID NO : 2711 , 14Q (ARM); SEQ ID NO : 2712, 14Q (ARM); SEQ ID NO : 2713 , 14Q (ARM); SEQ ID NO : 2714, 14Q (ARM); SEQ ID NO : 2715 , 14Q (ARM); SEQ ID NO : 2716, 14Q (ARM); SEQ ID NO : 2717, 14Q (ARM); SEQ ID NO : 2718, 14Q (ARM); SEQ ID NO : 2719, 14Q (ARM); SEQ ID NO: 2720, 14Q (ARM); SEQ ID NO: 2721, 14Q (ARM); SEQ ID NO: 2722, 14Q (ARM); SEQ ID NO: 2723, 14Q (ARM); SEQ ID NO: 2724, FP (FP); SEQ ID NO: 2725, 14Q (ARM); SEQ ID NO: 2726, 14Q (ARM); SEQ ID NO: 2727, ZFP36L1 (Gene); SEQ ID NO: 2728, ZFP36L1 (Gene); SEQ ID NO: 2729, ZFP36L1 (Gene); SEQ ID NO: 2730, ZFP36L1 (Gene); SEQ ID NO: 2731, 14Q (ARM); SEQ ID NO: 2732, 14Q (ARM); SEQ ID NO: 2733, 14Q (ARM); SEQ ID NO: 2734, 14Q (ARM); SEQ ID NO: 2735, 14Q (ARM); SEQ ID NO: 2736, 14Q (ARM); SEQ ID NO: 2737, 14Q (ARM); SEQ ID NO: 2738, 14Q (ARM); SEQ ID NO: 2739, 14Q (ARM); SEQ ID NO: 2740, 14Q (ARM); SEQ ID NO: 2741, 14Q (ARM); SEQ ID NO: 2742, 14Q (ARM); SEQ ID NO: 2743, 14Q (ARM); SEQ ID NO: 2744, 14Q (ARM); SEQ ID NO: 2745, 14Q (ARM); SEQ ID NO: 2746, 14Q (ARM); SEQ ID NO: 2747, 14Q (ARM); SEQ ID NO: 2748, 14Q (ARM); SEQ ID NO: 2749, 14Q (ARM); SEQ ID NO: 2750, 14Q (ARM); SEQ ID NO: 2751, PPP4R3A (Gene); SEQ ID NO: 2752, PPP4R3 A (Gene); SEQ ID NO: 2753, PPP4R3 A (Gene); SEQ ID NO: 2754, PPP4R3A (Gene); SEQ ID NO: 2755, PPP4R3A (Gene); SEQ ID NO: 2756, PPP4R3A (Gene); SEQ ID NO: 2757, PPP4R3 A (Gene); SEQ ID NO: 2758, PPP4R3 A (Gene); SEQ ID NO: 2759, PPP4R3A (Gene); SEQ ID NO: 2760, PPP4R3A (Gene); SEQ ID NO: 2761, PPP4R3A (Gene); SEQ ID NO: 2762, PPP4R3 A (Gene); SEQ ID NO: 2763, PPP4R3 A (Gene); SEQ ID NO: 2764, PPP4R3A (Gene); SEQ ID NO: 2765, PPP4R3A (Gene); SEQ ID NO: 2766, PPP4R3A (Gene); SEQ ID NO: 2767, 14Q (ARM); SEQ ID NO: 2768, 14Q (ARM); SEQ ID NO: 2769, 14Q (ARM); SEQ ID NO: 2770, 14Q (ARM); SEQ ID NO: 2771, FP (FP); SEQ ID NO: 2772, TCL1 A (Gene); SEQ ID NO: 2773, TCL1 A (Gene); SEQ ID NO: 2774, TCL1 A (Gene); SEQ ID NO: 2775, 14Q (ARM); SEQ ID NO: 2776, 14Q (ARM); SEQ ID NO: 2777, 14Q (ARM); SEQ ID NO: 2778, FP (FP); SEQ ID NO: 2779, 14Q (ARM); SEQ ID NO: 2780, 14Q (ARM); SEQ ID NO: 2781, 14Q (ARM); SEQ ID NO: 2782, YY1 (Gene); SEQ ID NO: 2783, YY1 (Gene); SEQ ID NO: 2784, YY1 (Gene); SEQ ID NO: 2785, YY1 (Gene); SEQ ID NO: 2786, YY1 (Gene);SEQ ID NO: 2787, YY1 (Gene); SEQ ID NO: 2788, 14Q (ARM); SEQ ID NO: 2789, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2790, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2791, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2792, RCOR1 (Gene); SEQ ID NO: 2793, RCOR1 (Gene); SEQ ID NO: 2794, RCOR1 (Gene); SEQ ID NO: 2795, RCOR1 (Gene); SEQ ID NO: 2796, RCOR1 (Gene); SEQ ID NO: 2797, RCOR1 (Gene); SEQ ID NO: 2798, RCOR1 (Gene); SEQ ID NO: 2799, RCOR1 (Gene); SEQ ID NO: 2800, RCOR1 (Gene); SEQ ID NO: 2801, RCOR1 (Gene); SEQ ID NO: 2802, RCOR1 (Gene); SEQ ID NO: 2803, RCOR1 (Gene); SEQ ID NO: 2804, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2805, TRAF3 (Gene); SEQ ID NO: 2806, TRAF3 (Gene); SEQ ID NO: 2807, TRAF3 (Gene); SEQ ID NO: 2808, TRAF3 (Gene); SEQ ID NO: 2809, TRAF3 (Gene); SEQ ID NO: 2810, TRAF3 (Gene); SEQ ID NO: 2811, TRAF3 (Gene); SEQ ID NO: 2812, TRAF3 (Gene); SEQ ID NO: 2813, TRAF3 (Gene); SEQ ID NO: 2814, TRAF3 (Gene); SEQ ID NO: 2815, TRAF3 (Gene); SEQ ID NO: 2816, TRAF3 (Gene); SEQ ID NO: 2817, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2818, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2819, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2820, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2821, 14q32.31_DLBCL (FOCAL); SEQ ID NO: 2822, 14Q (ARM); SEQ ID NO: 2823, CRIP1 (Gene); SEQ ID NO: 2824, CRIP1 (Gene); SEQ ID NO: 2825, CRIP1 (Gene); SEQ ID NO: 2826, CRIP1 (Gene); SEQ ID NO: 2827, IGH SV (SV); SEQ ID NO: 2828, IGH SV (SV); SEQ ID NO: 2829, IGH SV (SV); SEQ ID NO: 2830, IGH SV (SV); SEQ ID NO: 2831, IGH SV (SV); SEQ ID NO: 2832, IGH SV (SV); SEQ ID NO: 2833, IGH SV (SV); SEQ ID NO: 2834, IGH SV (SV); SEQ ID NO: 2835, IGH SV (SV); SEQ ID NO: 2836, IGH SV (SV); SEQ ID NO: 2837, IGH SV (SV); SEQ ID NO: 2838, IGH SV (SV); SEQ ID NO: 2839, IGH SV (SV); SEQ ID NO: 2840, IGH SV (SV); SEQ ID NO: 2841, IGH SV (SV); SEQ ID NO: 2842, IGH SV (SV); SEQ ID NO: 2843, IGH SV (SV); SEQ ID NO: 2844, IGH SV (SV); SEQ ID NO: 2845, IGH SV (SV); SEQ ID NO: 2846, IGH SV (SV); SEQ ID NO: 2847, IGH SV (SV); SEQ ID NO: 2848, IGH SV (SV); SEQ ID NO: 2849, IGH SV (SV); SEQ ID NO: 2850, IGH SV (SV); SEQ ID NO: 2851, IGH SV (SV); SEQ ID NO: 2852, IGH SV (SV); SEQ ID NO: 2853, IGH SV (SV); SEQ ID NO: 2854, IGH SV (SV); SEQ ID NO: 2855, IGH SV (SV); SEQ ID NO: 2856, IGH SV (SV); SEQ ID NO: 2857, IGH SV (SV); SEQ ID NO: 2858, IGH SV (SV); SEQ ID NO: 2859, IGH SV (SV); SEQ ID NO: 2860, IGH SV (SV); SEQ ID NO: 2861, IGH SV (SV); SEQ ID NO: 2862, IGH SV (SV); SEQ ID NO: 2863, IGH SV (SV); SEQ ID NO: 2864, IGH SV (SV); SEQ ID NO: 2865, IGH SV (SV); SEQ ID NO: 2866, IGH SV (SV); SEQ ID NO: 2867, IGH SV (SV); SEQ ID NO: 2868, IGH SV (SV); SEQ ID NO: 2869, IGH SV (SV); SEQ ID NO: 2870, IGH SV (SV); SEQ ID NO: 2871, IGH SV (SV); SEQ ID NO: 2872, IGH SV (SV); SEQ ID NO: 2873, IGH SV (SV); SEQ ID NO: 2874,IGH SV (SV); SEQ ID NO: 2875, IGH SV (SV); SEQ ID NO: 2876, IGH SV (SV); SEQ ID NO: 2877, IGH SV (SV); SEQ ID NO: 2878, IGH SV (SV); SEQ ID NO: 2879, IGH SV (SV); SEQ ID NO: 2880, IGH SV (SV); SEQ ID NO: 2881, IGH SV (SV); SEQ ID NO: 2882, IGH SV (SV); SEQ ID NO: 2883, IGH SV (SV); SEQ ID NO: 2884, IGH SV (SV); SEQ ID NO: 2885, IGH SV (SV); SEQ ID NO: 2886, IGH SV (SV); SEQ ID NO: 2887, IGH SV (SV); SEQ ID NO: 2888, IGH SV (SV); SEQ ID NO: 2889, IGH SV (SV); SEQ ID NO: 2890, IGH SV (SV); SEQ ID NO: 2891, IGH SV (SV); SEQ ID NO: 2892, IGH SV (SV); SEQ ID NO: 2893, IGH SV (SV); SEQ ID NO: 2894, IGH SV (SV); SEQ ID NO: 2895, IGH SV (SV); SEQ ID NO: 2896, IGH SV (SV); SEQ ID NO: 2897, IGH SV (SV); SEQ ID NO: 2898, IGH SV (SV); SEQ ID NO: 2899, IGH SV (SV); SEQ ID NO: 2900, IGH SV (SV); SEQ ID NO: 2901, IGH SV (SV); SEQ ID NO: 2902, IGH SV (SV); SEQ ID NO: 2903, IGH SV (SV); SEQ ID NO: 2904, IGH SV (SV); SEQ ID NO: 2905, IGH SV (SV); SEQ ID NO: 2906, IGH SV (SV); SEQ ID NO: 2907, IGH SV (SV); SEQ ID NO: 2908, IGH SV (SV); SEQ ID NO: 2909, IGH SV (SV); SEQ ID NO: 2910, IGH SV (SV); SEQ ID NO: 2911, IGH SV (SV); SEQ ID NO: 2912, IGH SV (SV); SEQ ID NO: 2913, IGH SV (SV); SEQ ID NO: 2914, IGH SV (SV); SEQ ID NO: 2915, IGH SV (SV); SEQ ID NO: 2916, IGH SV (SV); SEQ ID NO: 2917, IGH SV (SV); SEQ ID NO: 2918, IGH SV (SV); SEQ ID NO: 2919, IGH SV (SV); SEQ ID NO: 2920, IGH SV (SV); SEQ ID NO: 2921, IGH SV (SV); SEQ ID NO: 2922, IGH SV (SV); SEQ ID NO: 2923, IGH SV (SV); SEQ ID NO: 2924, IGH SV (SV); SEQ ID NO: 2925, IGH SV (SV); SEQ ID NO: 2926, IGH SV (SV); SEQ ID NO: 2927, IGH SV (SV); SEQ ID NO: 2928, IGH SV (SV); SEQ ID NO: 2929, IGH SV (SV); SEQ ID NO: 2930, IGH SV (SV); SEQ ID NO: 2931, IGH SV (SV); SEQ ID NO: 2932, IGH SV (SV); SEQ ID NO: 2933, IGH SV (SV); SEQ ID NO: 2934, IGH SV (SV); SEQ ID NO: 2935, IGH SV (SV); SEQ ID NO: 2936, IGH SV (SV); SEQ ID NO: 2937, IGH SV (SV); SEQ ID NO: 2938, IGH SV (SV); SEQ ID NO: 2939, IGH SV (SV); SEQ ID NO: 2940, IGH SV (SV); SEQ ID NO: 2941, IGH SV (SV); SEQ ID NO: 2942, IGH SV (SV); SEQ ID NO: 2943, IGH SV (SV); SEQ ID NO: 2944, IGH SV (SV); SEQ ID NO: 2945, IGH SV (SV); SEQ ID NO: 2946, IGH SV (SV); SEQ ID NO: 2947, IGH SV (SV); SEQ ID NO: 2948, IGH SV (SV); SEQ ID NO: 2949, IGH SV (SV); SEQ ID NO: 2950, IGH SV (SV); SEQ ID NO: 2951, IGH SV (SV); SEQ ID NO: 2952, IGH SV (SV); SEQ ID NO: 2953, IGH SV (SV); SEQ ID NO: 2954, IGH SV (SV); SEQ ID NO: 2955, IGH SV (SV); SEQ ID NO: 2956, IGH SV (SV); SEQ ID NO: 2957, IGH SV (SV); SEQ ID NO: 2958, IGH SV (SV); SEQ ID NO: 2959, IGH SV (SV); SEQ ID NO: 2960, 15Q (ARM); SEQ ID NO: 2961, FP (FP); SEQ ID NO: 2962, 15Q (ARM); SEQ ID NO: 2963, 15Q (ARM); SEQ ID NO: 2964, 15Q (ARM); SEQ ID NO: 2965, 15Q (ARM); SEQ ID NO: 2966, 15Q (ARM); SEQ ID NO: 2967, 15Q (ARM); SEQ ID NO: 2968,15Q (ARM); SEQ ID NO: 2969, 15Q (ARM); SEQ ID NO: 2970, 15Q (ARM); SEQ ID NO: 2971, 15Q (ARM); SEQ ID NO: 2972, 15Q (ARM); SEQ ID NO: 2973, 15Q (ARM); SEQ ID NO: 2974, 15Q (ARM); SEQ ID NO: 2975, 15Q (ARM); SEQ ID NO: 2976, 15Q (ARM); SEQ ID NO: 2977, 15Q (ARM); SEQ ID NO: 2978, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 2979, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 2980, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 2981, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 2982, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 2983, MGA (Gene); SEQ ID NO: 2984, MGA (Gene); SEQ ID NO: 2985, MGA (Gene); SEQ ID NO: 2986, MGA (Gene); SEQ ID NO: 2987, MGA (Gene); SEQ ID NO: 2988, MGA (Gene); SEQ ID NO: 2989, MGA (Gene); SEQ ID NO: 2990, MGA (Gene); SEQ ID NO: 2991, MGA (Gene); SEQ ID NO: 2992, MGA (Gene); SEQ ID NO: 2993, MGA (Gene); SEQ ID NO: 2994, MGA (Gene); SEQ ID NO: 2995, MGA (Gene); SEQ ID NO: 2996, MGA (Gene); SEQ ID NO: 2997, MGA (Gene); SEQ ID NO: 2998, MGA (Gene); SEQ ID NO: 2999, MGA (Gene); SEQ ID NO: 3000, MGA (Gene); SEQ ID NO: 3001, MGA (Gene); SEQ ID NO: 3002, MGA (Gene); SEQ ID NO: 3003, MGA (Gene); SEQ ID NO: 3004, MGA (Gene); SEQ ID NO: 3005, MGA (Gene); SEQ ID NO: 3006, MGA (Gene); SEQ ID NO: 3007, MGA (Gene); SEQ ID NO: 3008, MGA (Gene); SEQ ID NO: 3009, MGA (Gene); SEQ ID NO: 3010, MGA (Gene); SEQ ID NO: 3011, MGA (Gene); SEQ ID NO: 3012, MGA (Gene); SEQ ID NO: 3013, MGA (Gene); SEQ ID NO: 3014, MGA (Gene); SEQ ID NO: 3015, MGA (Gene); SEQ ID NO: 3016, MGA (Gene); SEQ ID NO: 3017, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3018, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3019, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3020, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3021, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3022, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3023, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3024, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3025, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3026, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3027, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3028, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3029, B2M (Gene); SEQ ID NO: 3030, B2M (Gene); SEQ ID NO: 3031, B2M (Gene); SEQ ID NO: 3032, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3033, 15ql5.3_DLBCL (FOCAL); SEQ ID NO: 3034, 15Q (ARM); SEQ ID NO: 3035, 15Q (ARM); SEQ ID NO: 3036, FP (FP); SEQ ID NO: 3037, 15Q (ARM); SEQ ID NO: 3038, 15Q (ARM); SEQ ID NO: 3039, 15Q (ARM); SEQ ID NO: 3040, 15Q (ARM); SEQ ID NO: 3041, SCG3 (Gene); SEQ ID NO: 3042, SCG3 (Gene); SEQ ID NO: 3043, SCG3 (Gene); SEQ ID NO: 3044, SCG3 (Gene); SEQ ID NO: 3045, SCG3 (Gene); SEQ ID NO: 3046, SCG3 (Gene); SEQ ID NO: 3047, SCG3 (Gene); SEQ ID NO: 3048, SCG3 (Gene); SEQ ID NO: 3049, SCG3 (Gene); SEQ ID NO: 3050, SCG3 (Gene); SEQ ID NO: 3051, SCG3 (Gene); SEQ ID NO: 3052, SCG3 (Gene); SEQ ID NO: 3053, 15Q (ARM); SEQ ID NO: 3054, 15Q (ARM); SEQ ID NO: 3055, FP (FP); SEQ ID NO: 3056, 15Q (ARM); SEQ IDNO: 3057, 15Q (ARM); SEQ ID NO: 3058, 15Q (ARM); SEQ ID NO: 3059, FP (FP); SEQ ID NO: 3060, 15Q (ARM); SEQ ID NO: 3061, 15Q (ARM); SEQ ID NO: 3062, 15Q (ARM); SEQ ID NO: 3063, 15Q (ARM); SEQ ID NO: 3064, 15Q (ARM); SEQ ID NO: 3065, 15Q (ARM); SEQ ID NO: 3066, 15Q (ARM); SEQ ID NO: 3067, 15Q (ARM); SEQ ID NO: 3068, 15Q (ARM); SEQ ID NO: 3069, 15Q (ARM); SEQ ID NO: 3070, 15Q (ARM); SEQ ID NO: 3071, MSI (MSI); SEQ ID NO: 3072, MSI (MSI); SEQ ID NO: 3073, MAP2K1 (Gene); SEQ ID NO: 3074, MAP2K1 (Gene); SEQ ID NO: 3075, MAP2K1 (Gene); SEQ ID NO: 3076, MAP2K1 (Gene); SEQ ID NO: 3077, MAP2K1 (Gene); SEQ ID NO: 3078, MAP2K1 (Gene); SEQ ID NO: 3079, MAP2K1 (Gene); SEQ ID NO: 3080, MAP2K1 (Gene); SEQ ID NO: 3081, MAP2K1 (Gene); SEQ ID NO: 3082, MAP2K1 (Gene); SEQ ID NO: 3083, MAP2K1 (Gene); SEQ ID NO: 3084, 15Q (ARM); SEQ ID NO: 3085, 15Q (ARM); SEQ ID NO: 3086, 15Q (ARM); SEQ ID NO: 3087, 15Q (ARM); SEQ ID NO: 3088, 15Q (ARM); SEQ ID NO: 3089, 15Q (ARM); SEQ ID NO: 3090, 15Q (ARM); SEQ ID NO: 3091, 15Q (ARM); SEQ ID NO: 3092, SIN3A (Gene); SEQ ID NO: 3093, SIN3 A (Gene); SEQ ID NO: 3094, SIN3 A (Gene); SEQ ID NO: 3095, SIN3 A (Gene); SEQ ID NO: 3096, SIN3A (Gene); SEQ ID NO: 3097, SIN3A (Gene); SEQ ID NO: 3098, SIN3A (Gene); SEQ ID NO: 3099, SIN3 A (Gene); SEQ ID NO: 3100, SIN3 A (Gene); SEQ ID NO: 3101, SIN3A (Gene); SEQ ID NO: 3102, SIN3A (Gene); SEQ ID NO: 3103, SIN3A (Gene); SEQ ID NO: 3104, SIN3A (Gene); SEQ ID NO: 3105, SIN3A (Gene); SEQ ID NO: 3106, SIN3A (Gene); SEQ ID NO: 3107, SIN3 A (Gene); SEQ ID NO: 3108, SIN3 A (Gene); SEQ ID NO: 3109, SIN3 A (Gene); SEQ ID NO: 3110, SIN3 A (Gene); SEQ ID NO: 3111, SIN3 A (Gene); SEQ ID NO: 3112, SIN3A (Gene); SEQ ID NO: 3113, 15Q (ARM); SEQ ID NO: 3114, 15Q (ARM); SEQ ID NO: 3115, MSI (MSI); SEQ ID NO: 3116, MSI (MSI); SEQ ID NO: 3117, 15Q (ARM); SEQ ID NO: 3118, 15Q (ARM); SEQ ID NO: 3119, 15Q (ARM); SEQ ID NO: 3120, 15Q (ARM); SEQ ID NO: 3121, 15Q (ARM); SEQ ID NO: 3122, 15Q (ARM); SEQ ID NO: 3123, 15Q (ARM); SEQ ID NO: 3124, 15Q (ARM); SEQ ID NO: 3125, 15Q (ARM); SEQ ID NO: 3126, 15Q (ARM); SEQ ID NO: 3127, IDH2 (Gene); SEQ ID NO: 3128, IDH2 (Gene); SEQ ID NO: 3129, IDH2 (Gene); SEQ ID NO: 3130, IDH2 (Gene); SEQ ID NO: 3131, IDH2 (Gene); SEQ ID NO: 3132, IDH2 (Gene); SEQ ID NO: 3133, IDH2 (Gene); SEQ ID NO: 3134, IDH2 (Gene); SEQ ID NO: 3135, IDH2 (Gene); SEQ ID NO: 3136, IDH2 (Gene); SEQ ID NO: 3137, IDH2 (Gene); SEQ ID NO: 3138, 15Q (ARM); SEQ ID NO: 3139, MSI (MSI); SEQ ID NO: 3140, MSI (MSI); SEQ ID NO: 3141, FP (FP); SEQ ID NO: 3142, 15Q (ARM); SEQ ID NO: 3143, 15Q (ARM); SEQ ID NO: 3144, 15Q (ARM); SEQ ID NO: 3145, 15Q (ARM); SEQ ID NO: 3146, 15Q (ARM); SEQ ID NO: 3147, 15Q (ARM); SEQ ID NO: 3148, 15Q (ARM); SEQ ID NO: 3149, 15Q (ARM); SEQ ID NO: 3150, 15Q (ARM); SEQ ID NO: 3151, 15Q (ARM); SEQ ID NO: 3152, 15Q (ARM);SEQ ID NO: 3153, 15Q (ARM); SEQ ID NO: 3154, 15Q (ARM); SEQ ID NO: 3155, TSC2 (Gene); SEQ ID NO: 3156, TSC2 (Gene); SEQ ID NO: 3157, TSC2 (Gene); SEQ ID NO: 3158, TSC2 (Gene); SEQ ID NO: 3159, TSC2 (Gene); SEQ ID NO: 3160, TSC2 (Gene); SEQ ID NO: 3161, TSC2 (Gene); SEQ ID NO: 3162, TSC2 (Gene); SEQ ID NO: 3163, TSC2 (Gene); SEQ ID NO: 3164, TSC2 (Gene); SEQ ID NO: 3165, TSC2 (Gene); SEQ ID NO: 3166, TSC2 (Gene); SEQ ID NO: 3167, TSC2 (Gene); SEQ ID NO: 3168, TSC2 (Gene); SEQ ID NO: 3169, TSC2 (Gene); SEQ ID NO: 3170, TSC2 (Gene); SEQ ID NO: 3171, TSC2 (Gene); SEQ ID NO: 3172, TSC2 (Gene); SEQ ID NO: 3173, TSC2 (Gene); SEQ ID NO: 3174, TSC2 (Gene); SEQ ID NO: 3175, TSC2 (Gene); SEQ ID NO: 3176, TSC2 (Gene); SEQ ID NO: 3177, TSC2 (Gene); SEQ ID NO: 3178, TSC2 (Gene); SEQ ID NO: 3179, TSC2 (Gene); SEQ ID NO: 3180, TSC2 (Gene); SEQ ID NO: 3181, TSC2 (Gene); SEQ ID NO: 3182, TSC2 (Gene); SEQ ID NO: 3183, TSC2 (Gene); SEQ ID NO: 3184, TSC2 (Gene); SEQ ID NO: 3185, TSC2 (Gene); SEQ ID NO: 3186, TSC2 (Gene); SEQ ID NO: 3187, TSC2 (Gene); SEQ ID NO: 3188, TSC2 (Gene); SEQ ID NO: 3189, TSC2 (Gene); SEQ ID NO: 3190, TSC2 (Gene); SEQ ID NO: 3191, TSC2 (Gene); SEQ ID NO: 3192, TSC2 (Gene); SEQ ID NO: 3193, TSC2 (Gene); SEQ ID NO: 3194, TSC2 (Gene); SEQ ID NO: 3195, TSC2 (Gene); SEQ ID NO: 3196, TSC2 (Gene); SEQ ID NO: 3197, TSC2 (Gene); SEQ ID NO: 3198, TSC2 (Gene); SEQ ID NO: 3199, TSC2 (Gene); SEQ ID NO: 3200, 16P (ARM); SEQ ID NO: 3201, CREBBP (Gene); SEQ ID NO: 3202, CREBBP (Gene); SEQ ID NO: 3203, CREBBP (Gene); SEQ ID NO: 3204, CREBBP (Gene); SEQ ID NO: 3205, CREBBP (Gene); SEQ ID NO: 3206, CREBBP (Gene); SEQ ID NO: 3207, CREBBP (Gene); SEQ ID NO: 3208, CREBBP (Gene); SEQ ID NO: 3209, CREBBP (Gene); SEQ ID NO: 3210, CREBBP (Gene); SEQ ID NO: 3211, CREBBP (Gene); SEQ ID NO: 3212, CREBBP (Gene); SEQ ID NO: 3213, CREBBP (Gene); SEQ ID NO: 3214, CREBBP (Gene); SEQ ID NO: 3215, CREBBP (Gene); SEQ ID NO: 3216, CREBBP (Gene); SEQ ID NO: 3217, CREBBP (Gene); SEQ ID NO: 3218, CREBBP (Gene); SEQ ID NO: 3219, CREBBP (Gene); SEQ ID NO: 3220, CREBBP (Gene); SEQ ID NO: 3221, CREBBP (Gene); SEQ ID NO: 3222, CREBBP (Gene); SEQ ID NO: 3223, CREBBP (Gene); SEQ ID NO: 3224, CREBBP (Gene); SEQ ID NO: 3225, CREBBP (Gene); SEQ ID NO: 3226, CREBBP (Gene); SEQ ID NO: 3227, CREBBP (Gene); SEQ ID NO: 3228, CREBBP (Gene); SEQ ID NO: 3229, CREBBP (Gene); SEQ ID NO: 3230, CREBBP (Gene); SEQ ID NO: 3231, CREBBP (Gene); SEQ ID NO: 3232, CREBBP (Gene); SEQ ID NO: 3233, CREBBP (Gene); SEQ ID NO: 3234, CREBBP (Gene); SEQ ID NO: 3235, CREBBP (Gene); SEQ ID NO: 3236, CREBBP (Gene); SEQ ID NO: 3237, CREBBP (Gene); SEQ ID NO: 3238, CREBBP (Gene); SEQ ID NO: 3239, CREBBP (Gene); SEQ ID NO: 3240, CREBBP (Gene); SEQ ID NO: 3241, FP (FP); SEQ ID NO: 3242, 16P (ARM); SEQ ID NO: 3243, 16P(ARM); SEQ ID NO: 3244, 16P (ARM); SEQ ID NO: 3245, 16P (ARM); SEQ ID NO: 3246, 16P (ARM); SEQ ID NO: 3247, 16P (ARM); SEQ ID NO: 3248, 16P (ARM); SEQ ID NO: 3249, 16P (ARM); SEQ ID NO: 3250, 16P (ARM); SEQ ID NO: 3251, MSI (MSI); SEQ ID NO: 3252, MSI (MSI); SEQ ID NO: 3253, CIITA SV (SV); SEQ ID NO: 3254, CIITA SV (SV); SEQ ID NO: 3255, CIITA SV (SV); SEQ ID NO: 3256, CIITA SV (SV); SEQ ID NO: 3257, CIITA SV (SV); SEQ ID NO: 3258, CIITA SV (SV); SEQ ID NO: 3259, CIITA SV (SV); SEQ ID NO: 3260, CIITA SV (SV); SEQ ID NO: 3261, CIITA SV (SV); SEQ ID NO: 3262, CIITA SV (SV); SEQ ID NO: 3263, CIITA SV (SV); SEQ ID NO: 3264, CIITA SV (SV); SEQ ID NO: 3265, CIITA SV (SV); SEQ ID NO: 3266, CIITA SV (SV); SEQ ID NO: 3267, CIITA SV (SV); SEQ ID NO: 3268, CIITA SV (SV); SEQ ID NO: 3269, CIITA SV (SV); SEQ ID NO: 3270, CIITA SV (SV); SEQ ID NO: 3271, CIITA SV (SV); SEQ ID NO: 3272, CIITA SV (SV); SEQ ID NO: 3273, CIITA SV (SV); SEQ ID NO: 3274, CIITA SV (SV); SEQ ID NO: 3275, CIITA SV (SV); SEQ ID NO: 3276, CIITA SV (SV); SEQ ID NO: 3277, CIITA SV (SV); SEQ ID NO: 3278, CIITA SV (SV); SEQ ID NO: 3279, CIITA SV (SV); SEQ ID NO: 3280, CIITA SV (SV); SEQ ID NO: 3281, CIITA SV (SV); SEQ ID NO: 3282, CIITA SV (SV); SEQ ID NO: 3283, CIITA SV (SV); SEQ ID NO: 3284, CIITA SV (SV); SEQ ID NO: 3285, CIITA SV (SV); SEQ ID NO: 3286, CIITA SV (SV); SEQ ID NO: 3287, CIITA SV (SV); SEQ ID NO: 3288, CIITA SV (SV); SEQ ID NO: 3289, CIITA SV (SV); SEQ ID NO: 3290, CIITA (Gene); SEQ ID NO: 3291, CIITA (Gene); SEQ ID NO: 3292, CIITA (Gene); SEQ ID NO: 3293, CIITA (Gene); SEQ ID NO: 3294, CIITA (Gene); SEQ ID NO: 3295, CIITA (Gene); SEQ ID NO: 3296, CIITA (Gene); SEQ ID NO: 3297, CIITA (Gene); SEQ ID NO: 3298, CIITA (Gene); SEQ ID NO: 3299, CIITA (Gene); SEQ ID NO: 3300, CIITA (Gene); SEQ ID NO: 3301, CIITA (Gene); SEQ ID NO: 3302, CIITA (Gene); SEQ ID NO: 3303, CIITA (Gene); SEQ ID NO: 3304, CIITA (Gene); SEQ ID NO: 3305, CIITA (Gene); SEQ ID NO: 3306, CIITA (Gene); SEQ ID NO: 3307, CIITA (Gene); SEQ ID NO: 3308, CIITA (Gene); SEQ ID NO: 3309, CIITA (Gene); SEQ ID NO: 3310, CIITA (Gene); SEQ ID NO: 3311, CIITA (Gene); SEQ ID NO: 3312, CIITA (Gene); SEQ ID NO: 3313, S0CS1 (Gene); SEQ ID NO: 3314, S0CS1 (Gene); SEQ ID NO: 3315, 16P (ARM); SEQ ID NO: 3316, 16P (ARM); SEQ ID NO: 3317, 16P (ARM); SEQ ID NO: 3318, 16P (ARM); SEQ ID NO: 3319, 16P (ARM); SEQ ID NO: 3320, 16P (ARM); SEQ ID NO: 3321, 16P (ARM); SEQ ID NO: 3322, 16P (ARM); SEQ ID NO: 3323, PRKCB (Gene); SEQ ID NO: 3324, PRKCB (Gene); SEQ ID NO: 3325, PRKCB (Gene); SEQ ID NO: 3326, PRKCB (Gene); SEQ ID NO: 3327, PRKCB (Gene); SEQ ID NO: 3328, PRKCB (Gene); SEQ ID NO: 3329, PRKCB (Gene); SEQ ID NO: 3330, PRKCB (Gene); SEQ ID NO: 3331, PRKCB (Gene); SEQ ID NO: 3332, 16P (ARM); SEQ ID NO: 3333, PRKCB (Gene); SEQ ID NO: 3334, PRKCB(Gene); SEQ ID NO: 3335, PRKCB (Gene); SEQ ID NO: 3336, PRKCB (Gene); SEQ ID NO: 3337, PRKCB (Gene); SEQ ID NO: 3338, PRKCB (Gene); SEQ ID NO: 3339, PRKCB (Gene); SEQ ID NO: 3340, PRKCB (Gene); SEQ ID NO: 3341, PRKCB (Gene); SEQ ID NO: 3342, 16P (ARM); SEQ ID NO: 3343, 16P (ARM); SEQ ID NO: 3344, 16P (ARM); SEQ ID NO: 3345, 16P (ARM); SEQ ID NO: 3346, IL4R (Gene); SEQ ID NO: 3347, IL4R (Gene); SEQ ID NO: 3348, IL4R (Gene); SEQ ID NO: 3349, IL4R (Gene); SEQ ID NO: 3350, IL4R (Gene); SEQ ID NO: 3351, IL4R (Gene); SEQ ID NO: 3352, IL4R (Gene); SEQ ID NO: 3353, IL4R (Gene); SEQ ID NO: 3354, IL4R (Gene); SEQ ID NO: 3355, IL4R (Gene); SEQ ID NO: 3356, IL4R (Gene); SEQ ID NO: 3357, IL4R (Gene); SEQ ID NO: 3358, IL4R (Gene); SEQ ID NO: 3359, 16P (ARM); SEQ ID NO: 3360, CD19 (Gene); SEQ ID NO: 3361, CD19 (Gene); SEQ ID NO: 3362, CD19 (Gene); SEQ ID NO: 3363, CD19 (Gene); SEQ ID NO: 3364, CD19 (Gene); SEQ ID NO: 3365, CD 19 (Gene); SEQ ID NO: 3366, CD 19 (Gene); SEQ ID NO: 3367, CD 19 (Gene); SEQ ID NO: 3368, CD 19 (Gene); SEQ ID NO: 3369, CD 19 (Gene); SEQ ID NO: 3370, CD 19 (Gene); SEQ ID NO: 3371, CD19 (Gene); SEQ ID NO: 3372, CD19 (Gene); SEQ ID NO: 3373, CD19 (Gene); SEQ ID NO: 3374, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3375, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3376, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3377, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3378, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3379, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3380, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3381, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3382, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3383, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3384, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3385, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3386, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3387, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3388, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3389, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3390, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3391, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3392, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3393, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3394, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3395, FP (FP); SEQ ID NO: 3396, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3397, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3398, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3399, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3400, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3401, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3402, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3403, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3404, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3405, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3406, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3407, 16ql2.1_DLBCL_trim (FOCAL); SEQ IDNO: 3408, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3409, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3410, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3411, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3412, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3413, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3414, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3415, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3416, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3417, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3418, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3419, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3420, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3421, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3422, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3423, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3424, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3425, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3426, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3427, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3428, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3429, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3430, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3431, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3432, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3433, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3434, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3435, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3436, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3437, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3438, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3439, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3440, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3441, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3442, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3443, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3444, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3445, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3446, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3447, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3448, 16ql2.1_DLBCL_trim (FOCAL); SEQ ID NO: 3449, 16Q (ARM); SEQ ID NO: 3450, 16Q (ARM); SEQ ID NO: 3451, 16Q (ARM); SEQ ID NO: 3452, 16Q (ARM); SEQ ID NO: 3453, 16Q (ARM); SEQ ID NO: 3454, 16Q (ARM); SEQ ID NO: 3455, MSI (MSI); SEQ ID NO: 3456, MSI (MSI); SEQ ID NO: 3457, 16Q (ARM); SEQ ID NO: 3458, 16Q (ARM); SEQ ID NO: 3459, 16Q (ARM); SEQ ID NO: 3460, 16Q (ARM); SEQ ID NO: 3461, 16Q (ARM); SEQ ID NO: 3462,16Q (ARM); SEQ ID NO: 3463, 16Q (ARM); SEQ ID NO: 3464, 16Q (ARM); SEQ ID NO: 3465,16Q (ARM); SEQ ID NO: 3466, 16Q (ARM); SEQ ID NO: 3467, 16Q (ARM); SEQ ID NO: 3468,16Q (ARM); SEQ ID NO: 3469, 16Q (ARM); SEQ ID NO: 3470, 16Q (ARM); SEQ ID NO: 3471,16Q (ARM); SEQ ID NO: 3472, FP (FP); SEQ ID NO: 3473, 16Q (ARM); SEQ ID NO: 3474, 16Q (ARM); SEQ ID NO: 3475, PLCG2 (Gene); SEQ ID NO: 3476, PLCG2 (Gene); SEQ IDNO: 3477, PLCG2 (Gene); SEQ ID NO: 3478, PLCG2 (Gene); SEQ ID NO: 3479, PLCG2 (Gene); SEQ ID NO: 3480, PLCG2 (Gene); SEQ ID NO: 3481, PLCG2 (Gene); SEQ ID NO: 3482, PLCG2 (Gene); SEQ ID NO: 3483, PLCG2 (Gene); SEQ ID NO: 3484, PLCG2 (Gene); SEQ ID NO: 3485, PLCG2 (Gene); SEQ ID NO: 3486, PLCG2 (Gene); SEQ ID NO: 3487, PLCG2 (Gene); SEQ ID NO: 3488, PLCG2 (Gene); SEQ ID NO: 3489, PLCG2 (Gene); SEQ ID NO: 3490, PLCG2 (Gene); SEQ ID NO: 3491, PLCG2 (Gene); SEQ ID NO: 3492, PLCG2 (Gene); SEQ ID NO: 3493, PLCG2 (Gene); SEQ ID NO: 3494, PLCG2 (Gene); SEQ ID NO: 3495, PLCG2 (Gene); SEQ ID NO: 3496, 16Q (ARM); SEQ ID NO: 3497, PLCG2 (Gene); SEQ ID NO: 3498, PLCG2 (Gene); SEQ ID NO: 3499, PLCG2 (Gene); SEQ ID NO: 3500, PLCG2 (Gene); SEQ ID NO: 3501, PLCG2 (Gene); SEQ ID NO: 3502, PLCG2 (Gene); SEQ ID NO: 3503, PLCG2 (Gene); SEQ ID NO: 3504, PLCG2 (Gene); SEQ ID NO: 3505, PLCG2 (Gene); SEQ ID NO: 3506, PLCG2 (Gene); SEQ ID NO: 3507, PLCG2 (Gene); SEQ ID NO: 3508, 16Q (ARM); SEQ ID NO: 3509, 16Q (ARM); SEQ ID NO: 3510, 16Q (ARM); SEQ ID NO: 3511, 16Q (ARM); SEQ ID NO: 3512, 16Q (ARM); SEQ ID NO: 3513, 16Q (ARM); SEQ ID NO: 3514, MSI (MSI); SEQ ID NO: 3515, MSI (MSI); SEQ ID NO: 3516, IRF8 (Gene); SEQ ID NO: 3517, IRF8 (Gene); SEQ ID NO: 3518, IRF8 (Gene); SEQ ID NO: 3519, IRF8 (Gene); SEQ ID NO: 3520, IRF8 (Gene); SEQ ID NO: 3521, IRF8 (Gene); SEQ ID NO: 3522, IRF8 (Gene); SEQ ID NO: 3523, IRF8 (Gene); SEQ ID NO: 3524, IRF8 (Gene); SEQ ID NO: 3525, 16Q (ARM); SEQ ID NO: 3526, 16Q (ARM); SEQ ID NO: 3527, 16Q (ARM); SEQ ID NO: 3528, ANKRD11 (Gene); SEQ ID NO: 3529, ANKRD11 (Gene); SEQ ID NO: 3530, ANKRD11 (Gene); SEQ ID NO: 3531, ANKRD11 (Gene); SEQ ID NO: 3532, ANKRD11 (Gene); SEQ ID NO: 3533, ANKRD11 (Gene); SEQ ID NO: 3534, ANKRD11 (Gene); SEQ ID NO: 3535, ANKRD11 (Gene); SEQ ID NO: 3536, ANKRD11 (Gene); SEQ ID NO: 3537, ANKRD11 (Gene); SEQ ID NO: 3538, ANKRD11 (Gene); SEQ ID NO: 3539, ANKRD11 (Gene); SEQ ID NO: 3540, ANKRD11 (Gene); SEQ ID NO: 3541, ANKRD11 (Gene); SEQ ID NO: 3542, ANKRD11 (Gene); SEQ ID NO: 3543, ANKRD11 (Gene); SEQ ID NO: 3544, ANKRD11 (Gene); SEQ ID NO: 3545, ANKRD11 (Gene); SEQ ID NO: 3546, ANKRD11 (Gene); SEQ ID NO: 3547, ANKRD11 (Gene); SEQ ID NO: 3548, ANKRD11 (Gene); SEQ ID NO: 3549, ANKRD11 (Gene); SEQ ID NO: 3550, ANKRD11 (Gene); SEQ ID NO: 3551, ANKRD11 (Gene); SEQ ID NO: 3552, ANKRD11 (Gene); SEQ ID NO: 3553, 16Q (ARM); SEQ ID NO: 3554, 16Q (ARM); SEQ ID NO: 3555, 17p_CLL (TP53) (ARM); SEQ ID NO: 3556, 17p_CLL (TP53) (ARM); SEQ ID NO: 3557, 17p_CLL (TP53) (ARM); SEQ ID NO: 3558, 17p_CLL (TP53) (ARM); SEQ ID NO: 3559, 17p_CLL (TP53) (ARM); SEQ ID NO: 3560, FP (FP); SEQ ID NO: 3561, 17p_CLL (TP53) (ARM); SEQ ID NO: 3562, 17p_CLL (TP53) (ARM); SEQ ID NO: 3563, GPS2 (Gene); SEQ IDNO: 3564, GPS2 (Gene); SEQ ID NO: 3565, GPS2 (Gene); SEQ ID NO: 3566, GPS2 (Gene); SEQ ID NO: 3567, GPS2 (Gene); SEQ ID NO: 3568, GPS2 (Gene); SEQ ID NO: 3569, GPS2 (Gene); SEQ ID NO: 3570, GPS2 (Gene); SEQ ID NO: 3571, GPS2 (Gene); SEQ ID NO: 3572, GPS2 (Gene); SEQ ID NO: 3573, TP53 (Gene); SEQ ID NO: 3574, TP53 (Gene); SEQ ID NO: 3575, TP53 (Gene); SEQ ID NO: 3576, TP53 (Gene); SEQ ID NO: 3577, TP53 (Gene); SEQ ID NO: 3578, TP53 (Gene); SEQ ID NO: 3579, TP53 (Gene); SEQ ID NO: 3580, TP53 (Gene); SEQ ID NO: 3581, TP53 (Gene); SEQ ID NO: 3582, TP53 (Gene); SEQ ID NO: 3583, TP53 (Gene); SEQ ID NO: 3584, DNAH2 (Gene); SEQ ID NO: 3585, DNAH2 (Gene); SEQ ID NO: 3586, DNAH2 (Gene); SEQ ID NO: 3587, DNAH2 (Gene); SEQ ID NO: 3588, DNAH2 (Gene); SEQ ID NO: 3589, DNAH2 (Gene); SEQ ID NO: 3590, DNAH2 (Gene); SEQ ID NO: 3591, DNAH2 (Gene); SEQ ID NO: 3592, DNAH2 (Gene); SEQ ID NO: 3593, DNAH2 (Gene); SEQ ID NO: 3594, DNAH2 (Gene); SEQ ID NO: 3595, DNAH2 (Gene); SEQ ID NO: 3596, DNAH2 (Gene); SEQ ID NO: 3597, DNAH2 (Gene); SEQ ID NO: 3598, DNAH2 (Gene); SEQ ID NO: 3599, DNAH2 (Gene); SEQ ID NO: 3600, DNAH2 (Gene); SEQ ID NO: 3601, DNAH2 (Gene); SEQ ID NO: 3602, DNAH2 (Gene); SEQ ID NO: 3603, DNAH2 (Gene); SEQ ID NO: 3604, DNAH2 (Gene); SEQ ID NO: 3605, DNAH2 (Gene); SEQ ID NO: 3606, DNAH2 (Gene); SEQ ID NO: 3607, DNAH2 (Gene); SEQ ID NO: 3608, DNAH2 (Gene); SEQ ID NO: 3609, DNAH2 (Gene); SEQ ID NO: 3610, DNAH2 (Gene); SEQ ID NO: 3611, DNAH2 (Gene); SEQ ID NO: 3612, DNAH2 (Gene); SEQ ID NO: 3613, DNAH2 (Gene); SEQ ID NO: 3614, DNAH2 (Gene); SEQ ID NO: 3615, DNAH2 (Gene); SEQ ID NO: 3616, DNAH2 (Gene); SEQ ID NO: 3617, DNAH2 (Gene); SEQ ID NO: 3618, DNAH2 (Gene); SEQ ID NO: 3619, DNAH2 (Gene); SEQ ID NO: 3620, DNAH2 (Gene); SEQ ID NO: 3621, DNAH2 (Gene); SEQ ID NO: 3622, DNAH2 (Gene); SEQ ID NO: 3623, DNAH2 (Gene); SEQ ID NO: 3624, DNAH2 (Gene); SEQ ID NO: 3625, DNAH2 (Gene); SEQ ID NO: 3626, DNAH2 (Gene); SEQ ID NO: 3627, DNAH2 (Gene); SEQ ID NO: 3628, DNAH2 (Gene); SEQ ID NO: 3629, DNAH2 (Gene); SEQ ID NO: 3630, DNAH2 (Gene); SEQ ID NO: 3631, DNAH2 (Gene); SEQ ID NO: 3632, DNAH2 (Gene); SEQ ID NO: 3633, DNAH2 (Gene); SEQ ID NO: 3634, DNAH2 (Gene); SEQ ID NO: 3635, DNAH2 (Gene); SEQ ID NO: 3636, DNAH2 (Gene); SEQ ID NO: 3637, DNAH2 (Gene); SEQ ID NO: 3638, DNAH2 (Gene); SEQ ID NO: 3639, DNAH2 (Gene); SEQ ID NO: 3640, DNAH2 (Gene); SEQ ID NO: 3641, DNAH2 (Gene); SEQ ID NO: 3642, DNAH2 (Gene); SEQ ID NO: 3643, DNAH2 (Gene); SEQ ID NO: 3644, DNAH2 (Gene); SEQ ID NO: 3645, DNAH2 (Gene); SEQ ID NO: 3646, DNAH2 (Gene); SEQ ID NO: 3647, DNAH2 (Gene); SEQ ID NO: 3648, DNAH2 (Gene); SEQ ID NO: 3649, DNAH2 (Gene); SEQ ID NO: 3650, DNAH2 (Gene); SEQ ID NO: 3651, DNAH2 (Gene); SEQ ID NO: 3652, DNAH2 (Gene); SEQ ID NO: 3653, DNAH2 (Gene); SEQID NO: 3654, DNAH2 (Gene); SEQ ID NO: 3655, DNAH2 (Gene); SEQ ID NO: 3656, DNAH2 (Gene); SEQ ID NO: 3657, DNAH2 (Gene); SEQ ID NO: 3658, DNAH2 (Gene); SEQ ID NO: 3659, DNAH2 (Gene); SEQ ID NO: 3660, DNAH2 (Gene); SEQ ID NO: 3661, DNAH2 (Gene); SEQ ID NO: 3662, DNAH2 (Gene); SEQ ID NO: 3663, DNAH2 (Gene); SEQ ID NO: 3664, DNAH2 (Gene); SEQ ID NO: 3665, DNAH2 (Gene); SEQ ID NO: 3666, DNAH2 (Gene); SEQ ID NO: 3667, DNAH2 (Gene); SEQ ID NO: 3668, DNAH2 (Gene); SEQ ID NO: 3669, DNAH2 (Gene); SEQ ID NO: 3670, DNAH2 (Gene); SEQ ID NO: 3671, MSI (MSI); SEQ ID NO: 3672, MSI (MSI); SEQ ID NO: 3673, 17p_CLL (...
Claims
CLAIMSWhat is claimed is:
1. A method comprising: causing, by at least one user computing device, to render, on a display device, at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface comprising at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receiving, by the at least one user computing device, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein one or more of the variants are selected from the group consisting of 19ql 3.32; 5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; and PIM2, and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes;generating, by the at least one user computing device, at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilizing, by the at least one user computing device, at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and causing, by the at least one user computing device, to render, on the display device, at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
2. The method of claim 1, further comprising: receiving, by the at least one user computing device, at least one DLBCL classification request from the at least one user computing device associated with the at least one patient, wherein the at least one DLBCL classification request comprises at least one electronic request over a network; and generating, by the at least one user computing device, at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction is configured to instruct the display device to render the at least one DLBCL classification interface.
3. The method of claim 1, wherein the at least one trained classification layer comprises an artificial neural network having learned weights for each of a plurality of neural network nodes.
4. The method of claim 1, further comprising: utilizing, by the at least one user computing device, at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generating, by the at least one user computing device, the at least one DLBCL classification results element comprising a two-dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization comprises at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
5. The method of claim 1, wherein the at least one dimensionality reduction model comprises uniform manifold approximation and projection (UMAP).
6. The method of claim 1, wherein the at least one DLBCL classification results element comprises a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
7. The method of claim 6, wherein the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; wherein the listing comprises the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
8. The method of claim 1, wherein the at least one DLBCL classification results element comprises a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, and ii) the at least one cluster identification associated with the at least one metafeature.
9. The method of claim 1, further comprising: generating, by the at least one user computing device, at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents that at least one treatment selection.
10. The method of claim 1, wherein the display device is local to the at least one user computing device.
11. The method of any of claims 1-10, wherein the classes of variants are characterized based upon the characterization of alterations in the variants, wherein the alterations are selected from the group consisting of a mutation, a structural variant (SV), and a somatic copy number alteration (SCNA);12. The method of claim 11, wherein the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, wherein eachclassification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
13. A user computing device comprising: a processor configured to execute computer-executable instructions which, upon execution, cause the user computing device to perform steps to: render at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface comprising at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receive, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein the variants comprise one or more of 19q 13.32; 5q; 6q; 9q; BCL11 A; ETS1; IRF4; METAP ID; and PIM2; and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; generate at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array;utilize at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and render at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
14. The user computing device of claim 13, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause the user computing device to perform steps to: receive at least one DLBCL classification request from the user computing device associated with the at least one patient, wherein the at least one DLBCL classification request comprises at least one electronic request over a network; and generate at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction is configured to instruct the user computing device to render the at least one DLBCL classification interface.
15. The user computing device of claim 13, wherein the at least one trained classification layer comprises an artificial neural network having learned weights for each of a plurality of neural network nodes.
16. The user computing device of claim 13, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause the user computing device to perform steps to: utilize at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generate the at least one DLBCL classification results element comprising a two- dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization comprises at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
17. The user computing device of claim 13, wherein the at least one dimensionality reduction model comprises a uniform manifold approximation and projection (UMAP).
18. The user computing device of claim 13, wherein the at least one DLBCL classification results element comprises a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
19. The user computing device of claim 18, wherein the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; wherein the listing comprises the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
20. The user computing device of claim 18, wherein the at least one DLBCL classification results element comprises a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, and ii) the at least one cluster identification associated with the at least one metafeature.
21. The user computing device of claim 13, wherein the processor is further configured to execute computer-executable instructions which, upon execution, further cause the user computing device to perform steps to: generate at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents the at least one treatment selection.
22. The user computing device of claim 13, wherein the processor is remote from the user computing device.
23. The user computing device of any of claims 13-22, wherein the classes of variants are characterized based upon the characterization of alterations in the variants, wherein the alterations are selected from the group consisting of a mutation, a structural variant (SV), and a somatic copy number alteration (SCNA).
24. The user computing device of claim 13, wherein the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, wherein each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
25. A method comprising: receiving, by the user computing device, at least one gene sample matrix (GSM) array input file associated with at least one patient; wherein the at least one GSM array input file comprises at least one GSM array encoding classes of variants associated with the at least one sample of the patient; wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; wherein the variants comprise one or more of 19q 13.32; 5q; 6q; 9q; BCL11 A; ETS1; IRF4; METAP ID; and PIM2; and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p, 17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP300, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, N0TCH2, OSBPLIO, PABPC1, PIM1, P0U2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423; generating, by the user computing device, at least one biologically-relevant metafeature based at least in part on aggregating at least one grouping of the classes of variants, the at least one grouping comprising classes of variants grouped according to biological relevance mapping, co-occurrence, and / or cluster frequency;generating, by the user computing device, at least one gene sample feature vector encoding at least one biologically-relevant metafeature; inputting, by the user computing device, the at least one gene sample feature vector into at least one DLBCL classification machine learning model to output at least one cluster identification confidence value indicative of a probability of the at least one sample belonging to at least one cluster identification categorizing the at least one patient; wherein the at least one cluster identification characterizes DLBCL burden on the at least one patient; wherein at least one trained classification layer is configured, based at least in part on training using a cohort of training samples, the cohort of training samples comprising known biologically-relevant meta-features and known cluster identifications, to model at least one probability of correlation between biologically-relevant meta-features and the at least one cluster identification; determining, by the user computing device, the at least one cluster identification based at least in part on the at least one cluster identification confidence value; and outputting, by the user computing device, for display by at least one display device via a DLBCL classification user interface, at least one DLBCL classification results element comprising a representation of the at least one cluster identification categorizing the at least one patient.
26. A method comprising: instructing, by at least one processor, at least one computing device to render at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface comprising at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receiving, by the at least one processor, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein one or more of the variants are selected from the group consisting of 19ql 3.32;5q; 6q; 9q; BCL11A; ETS1; IRF4; METAP1D; and PIM2, and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p,17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP3OO, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; generating, by the at least one processor, at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilizing, by the at least one processor, at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and instructing, by the at least one processor, the at least one computing device to render at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
27. The method of claim 26, further comprising: receiving, by the at least one processor, at least one DLBCL classification request from the at least one computing device associated with the at least one patient, wherein the at least one DLBCL classification request comprises at least one electronic request over a network; and generating, by the at least one processor, at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction isconfigured to instruct the at least one computing device to render the at least one DLBCL classification interface.
28. The method of claim 26, wherein the at least one trained classification layer comprises an artificial neural network having learned weights for each of a plurality of neural network nodes.
29. The method of claim 26, further comprising: utilizing, by the at least one processor, at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generating, by the at least one processor, the at least one DLBCL classification results element comprising a two-dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization comprises at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
30. The method of claim 26, wherein the at least one dimensionality reduction model comprises uniform manifold approximation and projection (UMAP).
31. The method of claim 26, wherein the at least one DLBCL classification results element comprises a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
32. The method of claim 31, wherein the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; wherein the listing comprises the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
33. The method of claim 26, wherein the at least one DLBCL classification results element comprises a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, and ii) the at least one cluster identification associated with the at least one metafeature.
34. The method of claim 26, further comprising: generating, by the at least one processor, at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents that at least one treatment selection.
35. The method of claim 26, wherein the at least one processor is remote from the at least one computing device.
36. The method of any of claims 26-35, wherein the classes of variants are characterized based upon the characterization of alterations in the variants, wherein the alterations are selected from the group consisting of a mutation, a structural variant (SV), and a somatic copy number alteration (SCNA);37. The method of claim 36, wherein the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, wherein each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
38. A computing system comprising: at least one processor configured to execute computer-executable instructions which, upon execution, cause the system to perform steps to: instruct at least one computing device to render at least one diffuse large B-cell lymphoma (DLBCL) classification interface, the at least one DLBCL classification interface comprising at least one gene sample matrix (GSM) array input element configured to accept at least one GSM array input file storing at least one GSM array associated with at least one patient; receive, via the at least one GSM array input element, the at least one GSM array input file, wherein the at least one GSM array input file represents at least one GSM array that characterizes classes of variants, wherein the variants comprise one or more of 19q 13.32; 5q; 6q; 9q; BCL11 A; ETS1; IRF4; METAP ID; and PIM2; and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p,17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP3OO, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423, wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; generate at least one metafeature based at least in part on a weighted sum of the classes of the variants in the at least one GSM array; utilize at least one DLBCL classification machine learning model to generate at least one cluster identification categorizing the at least one patient based at least in part on the at least one metafeature and at least one trained classification layer, wherein the at least one cluster identification characterizes a DLBCL burden on a subject associated with the at least one patient; and instruct the at least one computing device to render at least one DLBCL classification results element in the at least one DLBCL classification interface, wherein the at least one DLBCL classification results element depicts a representation of the at least one cluster identification categorizing the at least one patient.
39. The system of claim 38, wherein the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: receive at least one DLBCL classification request from the at least one computing device associated with the at least one patient, wherein the at least one DLBCL classification request comprises at least one electronic request over a network; andgenerate at least one rendering instruction in response to the at least one DLBCL classification request, wherein the at least one rendering instruction is configured to instruct the at least one computing device to render the at least one DLBCL classification interface.
40. The system of claim 38, wherein the at least one trained classification layer comprises an artificial neural network having learned weights for each of a plurality of neural network nodes.
41. The system of claim 38, wherein the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: utilize at least one dimensionality reduction model to create a two-dimensional representation of the at least one metafeature; and generate the at least one DLBCL classification results element comprising a two- dimensional visualization of the two-dimensional representation of the at least one metafeature, wherein the two-dimensional visualization comprises at least one labelled data point representing the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
42. The system of claim 38, wherein the at least one dimensionality reduction model comprises a uniform manifold approximation and projection (UMAP).
43. The system of claim 38, wherein the at least one DLBCL classification results element comprises a listing of the at least one metafeature and the at least one cluster identification associated with the at least one metafeature.
44. The system of claim 43, wherein the at least one trained classification layer is configured to produce at least one confidence score associated with the at least one cluster identification for the at least one metafeature; wherein the listing comprises the at least one confidence score associated with the at least one cluster identification for the at least one metafeature.
45. The system of claim 43, wherein the at least one DLBCL classification results element comprises a heatmap depicting: i) at least one gene mutation associated with the at least one metafeature, andii) the at least one cluster identification associated with the at least one metafeature.
46. The system of claim 38, wherein the at least one processor is further configured to execute computer-executable instructions which, upon execution, further cause the system to perform steps to: generate at least one treatment selection based at least in part on the at least one cluster identification; wherein the at least one DLBCL classification results element represents the at least one treatment selection.
47. The system of claim 38, wherein the at least one processor is remote from the at least one computing device.
48. The system of any of claims 38-47, wherein the classes of variants are characterized based upon the characterization of alterations in the variants, wherein the alterations are selected from the group consisting of a mutation, a structural variant (SV), and a somatic copy number alteration (SCNA);49. The system of claim 38, wherein the weighted sum is generated by assigning a classification-specific weighted value to each class of variant characterized, wherein each classification-specific weighted value reflects the magnitude of the characterized alteration in each class of variant.
50. A method comprising: receiving, by the at least one processor, at least one gene sample matrix (GSM) array input file associated with at least one patient; wherein the at least one GSM array input file comprises at least one GSM array encoding classes of variants associated with the at least one sample of the patient; wherein the classes of variants are characterized using a targeted sequencing panel comprising oligonucleotides suitable for use in targeted sequencing of the variant classes; wherein the variants comprise one or more of 19q 13.32; 5q; 6q; 9q; BCL11 A; ETS1;IRF4; METAP ID; and PIM2; and one or more additional variants are selected from the group consisting of: 10q23.31, 1 Ip, l lq, 1 lq23.3, 12p, 12pl3.2, 12q, 13q, 13ql4.2, 13q34, 14q32.31, 15ql5.3, 16ql2.1, 17p,17q24.3, 17q25.1, 18p, 18q, 18q21.32, 18q22.2, 18q23, 19pl3.2, 19pl3.3, 19q, 19ql3.42, lpl3.1, lp31.1, lp36.11, lp36.32, Iq, lq32.1, lq42.12, 21q, 2pl6.1, 2q22.2, 3p, 3p21.31, 3q28, 4q21.22, 5p, 6p, 6p21.1, 6p21.33, 6ql4.1, 6q21, 7p, 7q, 7q22.1, 8ql2.1, 8q24.22, 9p21.3, 9p24.1, ACTB, ATP2A2, B2M, BCL10, BCL2, BCL6, BCL7A, BRAF, BTG1, BTG2, CARD11, CCDC27, CD274, CD58, CD70, CD79Bmut, CD83, CREBBP, CRIP1, CXCR4, DTX1, DUSP2, EBF1, EEF1A1, EP3OO, ETV6, EZH2, FADD, FAS, GNA13, GNAI2, GRHPR, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H2AC, HIST1H2AM, HIST1H2BC, HLA-A, HLA-B, HLA-C, HVCN1, IGLL5, IKZF3, IRF2BP2, IRF8, KLHL6, KMT2D, KRAS, LTB, LYN, MAP2K1, MEF2B, MEF2C, MYC, MYD88, MYD88L265P, MYD88OTHER, NFKBIA, NFKBIE, NOTCH2, OSBPLIO, PABPC1, PIM1, POU2AF1, POU2F2, PRDM1, PTEN, PTPN6, RAC2, RHOA, SESN3, SF3B1, SGK1, SMG7, SOCS1, SPEN, STAT3, TBL1XR1, TET2, TMEM30A, TMSB4X, TNFAIP3, TNFRSF14, TNIP1, TOX, TP53, TUBGCP5, UBE2A, YY1, ZC3H12A, ZEB2, ZFP36L1, and ZNF423; generating, by the at least one processor, at least one biologically-relevant metafeature based at least in part on aggregating at least one grouping of the classes of variants, the at least one grouping comprising classes of variants grouped according to biological relevance mapping, co-occurrence, and / or cluster frequency; generating, by the at least one processor, at least one gene sample feature vector encoding at least one biologically-relevant metafeature; inputting, by the at least one processor, the at least one gene sample feature vector into at least one DLBCL classification machine learning model to output at least one cluster identification confidence value indicative of a probability of the at least one sample belonging to at least one cluster identification categorizing the at least one patient; wherein the at least one cluster identification characterizes DLBCL burden on the at least one patient; wherein at least one trained classification layer is configured, based at least in part on training using a cohort of training samples, the cohort of training samples comprising known biologically-relevant meta-features and known cluster identifications, to model at least one probability of correlation between biologically-relevant meta-features and the at least one cluster identification; determining, by the at least one processor, the at least one cluster identification based at least in part on the at least one cluster identification confidence value; and outputting, by the at least one processor, for display by at least one display device via a DLBCL classification user interface, at least one DLBCL classification results elementcomprising a representation of the at least one cluster identification categorizing the at least one patient.
Citation Information
Patent Citations
Therapeutic treatment of select diffuse large b cell lymphomas exhibiting distinct pathogenic mechanisms and outcomes
US20190292602A1
Method and kit for predicting therapeutic effectiveness of chemotherapy for diffuse large b-cell lymphoma patients
US20200270703A1
Panels and methods for treatment of diffuse large b-cell lymphoma
WO2022197930A2