Molecular subtyping of tumors from archival tissue

By solubilizing, digesting, and enriching RNA from aged FFPE samples, combined with digital droplet PCR, the challenges of molecular subtype analysis in aged samples were addressed, enabling efficient molecular subtype identification and diagnosis, and supporting personalized treatment decisions.

JP2025530330APending Publication Date: 2025-09-11BIOVENTURES LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515364
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-21
Filing Date
2023-06-21
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively use archived, degraded samples for molecular subtyping analysis, especially for cancer diagnosis. Common molecular tests such as Oncotype Dx Recurrence Score do not work well in aging samples.

Method used

By solubilizing aged FFPE samples with mineral oil, RNA was extracted and amplified for molecular subtype analysis, combining protease digestion, DNase treatment, incubation with salt-based buffer, RNA enrichment, and digital droplet PCR.

Benefits of technology

Effective molecular subtyping of aged FFPE samples is achieved, improving the diagnostic value of archived samples and supporting retrospective clinical trials and personalized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025530330000001_ABST
    Figure 2025530330000001_ABST
Patent Text Reader

Abstract

The present disclosure encompasses methods for molecular subtyping of formalin-fixed, paraffin-embedded tumor samples. The disclosure is particularly useful for older, degraded (archived) samples where standard methods are not feasible. Furthermore, the disclosed methods provide a low correlation between patient outcome data and tumor molecular subtype, providing a wealth of information to guide treatment decisions and / or therapeutic drug selection.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 354,128, entitled "METHODS FOR THE MOLECULAR SUBTYPING OF TUMORS FROM ARCHIVAL TISSUE," filed June 21, 2022. The contents of the aforementioned application are incorporated herein by reference in their entirety.

[0002] Technical Field The present disclosure encompasses methods for identifying molecular subtypes of cancer or tumors, which are also useful for archival samples that are old and generally in a deteriorated state. [Background technology]

[0003] background Molecular biology research is now routinely performed in pathology laboratories for the diagnosis of various pathologies, including cancer, infectious diseases, and genetic disorders. One of the most commonly performed molecular tests is polymerase chain reaction (PCR), which can amplify specific nucleic acid sequences from extremely small amounts of genetic starting material. In most laboratories, PCR is typically performed on a variety of fresh specimens, including blood, body fluids, and tissues. Fresh material may not be available, and biomedical needs for PCR on archived material may be unmet. This principle is crucial when a diagnosis is unconfirmed, when fresh tissue is no longer available, or when testing remotely obtained samples. However, the suitability of archived material may be questionable, as probes used in assays such as i) Oncotype Dx Recurrence Score, ii) Prosigna Predicition Analysis of Microarray 50 (PAM50) Recurrence Risk, iii) EndoPredict, iv) MamaPrint, and v) Breast Cancer Index may not function well, especially on older, degraded specimens.

[0004] In contrast to fresh samples, archived materials such as cytological specimens, paraffin-embedded tissues, and frozen tissues offer the opportunity for careful morphological review and interpretation before molecular analysis. This allows trained cytologists to preselect cases and slides for subsequent molecular analysis, resulting in optimal utilization and cost control in the molecular laboratory. As molecular techniques are improved and refined, diagnostic possibilities will become a reality.

[0005] Clinicians and pathologists alike have operated under the assumption that fresh specimens offer the highest diagnostic yield in molecular biology studies. However, fresh material for performing molecular biology tests is not always available, and careful comparative studies have been few and far between. Therefore, there is a need in the art for methods to evaluate archival tissues used as diagnostic samples. This becomes even more important as clinical outcome data become available and can be linked to these samples, allowing for the conduct and analysis of retrospective clinical trials using molecular profiling data obtained from these "currently available" samples. Summary of the Invention

[0006] overview In some aspects, the present disclosure encompasses a method for molecular subtyping of a cancer sample obtained from a subject, the method comprising: (a) enhancing solubilization of old, degraded FFPE sample material through the non-discretionary use of mineral oil; (b) digesting the sample with a proteinase; (c) incubating the digested sample from step a) with DNase; (d) incubating the mixture from step b) in a guanidine salt-based buffer; (e) enriching and isolating RNA; (f) pre-amplifying; and (g) performing digital droplet PCR.

[0007] In certain embodiments, the sample is a formalin-fixed, paraffin-embedded sample. In some embodiments, the sample is at least 5 years old.

[0008] In one embodiment, the sample is digested with proteinase K. In one embodiment, the sample is digested at about 65° C. to about 70° C. for about 75 minutes.

[0009] In some embodiments, step a) further comprises separating the sample into an aqueous phase by centrifugation and separating the aqueous phase from the residual lysate, wherein the aqueous phase is used in step b).

[0010] In some embodiments of this method, spin columns are used to concentrate and isolate RNA, and in some embodiments, the remaining lysate is used for DNA extraction.

[0011] In some embodiments of this method, a proteinase is incubated with the remaining lysate at about 65°C to about 70°C for about 13 hours to about 18 hours to produce a digested tissue lysate. In some aspects, the digested tissue lysate is incubated with RNase.

[0012] Another embodiment of this method further comprises concentrating and isolating the DNA using a spin column. In some embodiments, the method further comprises pre-amplifying the isolated DNA and performing ddPCR.

[0013] In some embodiments of the method, the target nucleic acid and optionally the reference nucleic acid are quantified.In some aspects, the target nucleic acid or its fragment encodes estrogen receptor 1 (ESR1), progesterone receptor (PGR), B-cell lymphoma 2 (BCL2), signal peptide, CUB domain and EGF-like domain containing 2 (SCUBE2), human epidermal growth factor receptor 2 (HER2), growth factor receptor-bound protein 7 (GRB7), proliferation marker Ki-67 (MKI67), aurora kinase A (AURKA), Baculoviral IAP Repeat Containing 5 (BIRC5), cyclin B1 (CCNB1), MYB Proto-Oncogene Like 2 (MYBL2), thymidine kinase 1 (TK1), or a combination thereof.

[0014] In some embodiments, the cancer sample is a breast cancer sample. In some embodiments, the subtype consists of a luminal A subtype (Lum A), a luminal B subtype (Lum B), a HER2 subtype (HER2), or a triple-negative subtype (TN).

[0015] In some embodiments, a sample is determined to be Lum A if the level of one or more of the target nucleic acids or fragments thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated in the sample.

[0016] In some embodiments, a sample is determined to be Lum B if the level of a nucleic acid or fragment thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated, and if the level of a nucleic acid or fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated in the sample.

[0017] In some embodiments, a sample is determined to be HER2 if the level of a nucleic acid or fragment thereof encoding HER2, GRB7, or a combination thereof is elevated in the sample.

[0018] In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated in the sample.

[0019] In a further embodiment, the amount of target nucleic acid is compared to the subject's outcome or treatment response. [Brief explanation of the drawings]

[0020] Brief description of the diagram The patent or application file contains at least one drawing in color. Copies of this patent or patent application publication (with color drawing(s)) will be provided by the Office upon request and payment of the necessary fee.

[0021] [Figure 1A] Figure 1A shows the S8897 specimen repository used in the examples. The S8897 specimen archive contains the largest number of patient specimens who did not receive treatment after surgery for small breast tumors. It is estimated that most patient specimens contain 4-6 unstained formalin-fixed, paraffin-embedded (FFPE) slides that are more than 20-30 years old. The criteria for the low-risk group are as follows: patients with a T less than 1 cm and no HR status or S-phase flow cytometry assessment (also known as the initial low-risk group). Specimens include unstained FFPE slides.

[0022] [Figure 1B] Figure 1B is a flow diagram of the archived specimen and its molecular profiling history.

[0023] [Figure 2] Figures 2A-2C show representative fragment analyzer (FA) traces of samples (extracted nucleic acids). Figure 2A shows representative RNA Fragment Analyzer traces of a breast tumor specimen. Figure 2B shows representative RNA Fragment Analyzer traces of a normal LN. Figure 2C shows representative genomic DNA fragment analysis traces of a breast tumor specimen. Similar RNA Fragment Analyzer traces were performed on all tumor specimens T1-T33 (data not shown).

[0024] [Figure 3] Figures 3A-3D show breast cancer tumor subtyping for samples T15, T19, T23, and T21. The x-axis represents molecular targets, and the y-axis represents the z-score transformation of ddPCR counts. Figure 3A shows specimen T15 (Lum A). Figure 3B shows specimen T19 (Lum B). Figure 3C shows specimen T23 (Her 2). Figure 3D shows specimen T21 (TN).

[0025] [Figure 4] Figures 4A–4U show breast cancer tumor subtyping using ddPCR for samples. The x-axis represents molecular targets, and the y-axis represents the z-score transformation of ddPCR counts. Figure 4A shows specimen T1. Figure 4B shows specimen T2. ​​Figure 4C shows specimen T3 (Her 2). Figure 4D shows specimen T4. Figure 4E shows specimen T5. Figure 4F shows specimen T6. Figure 4G shows specimen T7. Figure 4H shows specimen T8. Figure 4I shows specimen T9. Figure 4J shows specimen T10. Figure 4K shows specimen T11. Figure 4L shows specimen T12. Figure 4M shows specimen T13. Figure 4N shows specimen T14. Figure 4O shows specimen T16. Figure 4P shows specimen T17. Figure 4Q shows specimen T18. Figure 4R shows specimen T20. Figure 4S shows specimen T22. Figure 4T shows specimen T24. Figure 4U shows specimen T25.

[0026] [Figure 5] Figures 5A-5F show the distance geometry. Figure 5A shows the specimen-based HCA of 25 breast cancer specimens. Figure 5B shows the same 25 specimens after PCA. Figure 5C shows the HCA of the same 25 breast cancer specimens along the x-axis, along with the 12 molecular markers used in the ddPCR assay. Figures 5D-5F show the results for 25 breast tumor specimens (T1-T25) and 13 normal human lymph node specimens (LN1-LN13) from the S8897 cohort after processing and analysis by the ddPCR assay.

[0027] [Figure 6]Figures 6A-6F show statistical simulations of T1-T25 and LN1-LN13. Molecular profiling was performed based on breast tumor RNA. Figure 6A shows specimen-based HCA of T1-T25 and synthetic breast cancer specimens. Figure 6B shows the results of PCA including T1-T25 along with 500 synthetic specimens for each PAM50 subtype. Figure 6C shows HCA of T1-T25 specimens using 500 synthetic specimens for each PAM50 subtype. Figure 6D displays HCA of T1-T25, LN1-LN13, and synthetic specimens. Figure 6E shows the results of PCA including T1-T25 and LN1-LN13 along with 500 synthetic specimens for each PAM50 subtype and lymph node. Figure 6F shows the HCA of T1-T25 and LN1-LN13 specimens, with the x-axis comparing 500 synthetic specimens for each PAM50 subtype and lymph node with the 12 molecular markers used in the ddPCR assay.

[0028] [Figure 7] Figures 7A-7B show the clinical and genomic characteristics. Figure 7A shows the clinical and genomic characteristics. Figure 7B shows the IntClust-associated chromosomal bands and associated copy numbers.

[0029] [Figure 8] Figures 8A-8C show DNA specimen analysis for sample T6. Figure 8A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S98-9, the final specimen ID is T6, HistoPath is ILC, Grade is 1, tumor percentage is 60%, and molecular subtype (ddPCR) is Lum A. Figure 8B shows IntClust-associated chromosomal bands and associated copy numbers (basal via InClust 10). Figure 8C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0030] [Figure 9]Figures 9A-9C show DNA specimen analysis for sample T1. Figure 9A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). Initial specimen ID is S98-3, final specimen ID is T1, HistoPath is IDC, Grade is 3, tumor percentage is 30%, and molecular subtype (ddPCR) is Lum B. Figure 9B shows the chromosomal bands associated with IntClust and the associated copy number (LumA via InClust 8). Figure 9C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0031] [Figure 10] Figure 10A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). Initial specimen ID is S98-4, final specimen ID is T2, HistoPath is IDC, grade is 3, tumor percentage is 40%, and molecular subtype (ddPCR) is TN. Figure 10B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 10C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0032] [Figure 11] Figures 11A-C show DNA specimen analysis for sample T3. Figure 11A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S98-5, the final specimen ID is T3, HistoPath is IDC, grade is 3, tumor percentage is 40%, and molecular subtype (ddPCR) is TN. Figure 11B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 11C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0033] [Figure 12]Figures 12A-12C show DNA specimen analysis for sample T4. Figure 12A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S98-6, the final specimen ID is T4, HistoPath is IDC, grade is 3, tumor percentage is 40%, and molecular subtype (ddPCR) is Lum B. Figure 12B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10, Lum B via InClust 1 and 2). Figure 12C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0034] [Figure 13] Figures 13A-13C show DNA specimen analysis for sample T5. Figure 13A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). Initial specimen ID is S98-8, final specimen ID is T5, HistoPath is IDC, grade is 1, tumor percentage is 70%, and molecular subtype (ddPCR) is Lum A. Figure 13B shows the chromosomal bands associated with IntClust and the associated copy number (Lum A via InClust 8). Figure 13C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0035] [Figure 14] Figures 14A-14C show DNA specimen analysis for sample T7. Figure 14A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial sample ID was S98-11, the final sample ID was T7, the HistoPath was LN Met, the Grade was 2, the tumor percentage was 50%, and the molecular subtype (ddPCR) was Lum A;B. Figure 14B shows the chromosomal bands associated with IntClust and the associated copy number (Lum A via InClust 8;9). Figure 14C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of the mutation.

[0036] [Figure 15] Figures 15A-15C show DNA specimen analysis for sample T8. Figure 15A shows a 20-year-old tumor from UAMS, and a normal blood pool was used due to CNV (no matched normal). Initial specimen ID: S98-12; Final specimen ID: T8; HistoPath: ILC; Grade: 2; Tumor %=50%; Molecular subtype (ddPCR): Lum A. Figure 15B shows IntClust-associated chromosomal bands and associated copy numbers (Lum A; B via InClust 7; 1, 2, 6). Figure 15C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0037] [Figure 16] Figures 16A-16C show DNA specimen analysis for sample T9. Figure 16A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S99-7, the final specimen ID is T9, HistoPath is DCIS, grade is 2, tumor percentage is 30%, and molecular subtype (ddPCR) is Lum A. Figure 16B shows the chromosomal bands associated with IntClust and the associated copy number (Lum A; Her2 via InClust 7; 8, 5). Figure 16C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0038] [Figure 17] Figures 17A-17C show DNA specimen analysis for sample T10. Figure 17A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). Initial specimen ID is S99-16, final specimen ID is T10, HistoPath is IDS, grade is 3, tumor percentage is 60%, and molecular subtype (ddPCR) is TN. Figure 17B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 17C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0039] [Figure 18] Figures 18A-18C show DNA specimen analysis for sample T11. Figure 18A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S99-17, the final specimen ID is T11, HistoPath is IDC, grade is 3, tumor percentage is 80%, and molecular subtype (ddPCR) is TN. Figure 18B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 18C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0040] [Figure 19] Figures 19A-19C show DNA specimen analysis for sample T12. Figure 19A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). Initial specimen ID is S99-18, final specimen ID is T12, HistoPath is IDC, grade is 3, tumor percentage is 50%, and molecular subtype (ddPCR) is TN. Figure 19B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 19C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0041] [Figure 20] Figures 20A-20C show DNA specimen analysis for sample T13. Figure 20A shows a 20-year-old tumor from UAMS, using a matched pooled normal for CNVs (no matched normal). The initial specimen ID is S99-19, the final specimen ID is T13, HistoPath is IDC, grade is 3, tumor percentage is 70%, and molecular subtype (ddPCR) is TN. Figure 20B shows the chromosomal bands associated with IntClust and the associated copy number (basal via InClust 10). Figure 20C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0042] [Figure 21] Figures 21A-21C show DNA specimen analysis for sample T14. Figure 21A shows the S8897 tumor and matched LN. Initial sample ID, Tu2; Final sample ID, T14; HistoPath, IDC; Grade 2; Tumor %=50%; Molecular subtype (ddPCR), Lum A. Figure 21B shows the chromosomal bands associated with IntClust and associated copy number (LumA via InClust 7). Figure 21C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0043] [Figure 22] Figures 22A-22C show DNA specimen analysis for sample T15. Figure 22A shows the S8897 tumor and matched LN. Initial specimen ID, Tu3; Final specimen ID, T15; HistoPath, IDC; Grade 2; Tumor %=95%; Molecular subtype (ddPCR), Lum A. Figure 22B shows the chromosomal bands associated with IntClust and associated copy number (LumA via InClust 7). Figure 22C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0044] [Figure 23] Figures 23A-C show DNA specimen analysis for sample T16. Figure 23A shows the S8897 tumor and matched LN. Initial sample ID, Tu11; Final sample ID, T16; HistoPath, IDC; Grade 2; Tumor %=20%; Molecular subtype (ddPCR), Lum A. Figure 23B shows IntClust-associated chromosomal bands and associated copy numbers (Lum A, B via InClust 8; 2). Figure 23C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutational impact.

[0045] [Figure 24]Figures 24A-24B show DNA specimen analysis for sample T17. Figure 24A shows the S8897 tumor and matched LN. Initial sample ID, Tu12; Final sample ID, T17; HistoPath, ADH, ALH; Grade, NA; Tumor %=0%; Molecular subtype (ddPCR), NA. Figure 24B shows the chromosomal bands associated with IntClust and the associated copy number.

[0046] [Figure 25] Figures 25A-25C show DNA specimen analysis for sample T18. Figure 25A shows the S8897 tumor and matched LN. Initial specimen ID, Tu13; Final specimen ID, T18; HistoPath, DCIS & IDC; Grade 1; Tumor %=20%; Molecular subtype (ddPCR), Lum A. Figure 25B shows the chromosomal bands associated with IntClust and associated copy number (Lum A, via InClust 7;8). Figure 25C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0047] [Figure 26] Figures 26A-26C show DNA specimen analysis for sample T19. Figure 26A shows the S8897 tumor and matched LN. Initial sample ID, Tu15; Final sample ID, T19; HistoPath, IDC; Grade 2; Tumor %=20%; Molecular subtype (ddPCR), Lum B. Figure 26B shows IntClust-associated chromosomal bands and associated copy number (Lum B via InClust 8). Figure 26C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0048] [Figure 27]Figures 27A-27C show DNA specimen analysis for sample T20. Figure 27A shows the S8897 tumor and matched LN. Initial specimen ID, Tu16; Final specimen ID, T20; HistoPath, DCIS & IDC (4:1); Grade 1; Tumor %=25%; Molecular subtype (ddPCR), Lum A. Figure 27B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 8). Figure 27C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0049] [Figure 28] Figures 28A-C show DNA specimen analysis for sample T21. Figure 28A shows the S8897 tumor and matched LN. Initial specimen ID, Tu18; Final specimen ID, T21; HistoPath, IDC; Grade 2; <5% tumor; Molecular subtype (ddPCR), TN. Figure 28B shows the chromosomal bands associated with IntClust and associated copy number (basal vs Lum B, via InClust 10, 9). Figure 28C shows the top 27 targets of the UMI-based gene panel, color-coded by mutation impact.

[0050] [Figure 29] Figures 29A-C show DNA specimen analysis for sample T22. Figure 29A shows the S8897 tumor and matched LN. Initial specimen ID, Tu19; Final specimen ID, T22; HistoPath, IDC; Grade 1; Tumor %=100%; Molecular subtype (ddPCR), Lum A. Figure 29B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 7). Figure 29C shows the top 27 targets of the UMI-based gene panel, color-coded by mutation impact.

[0051] [Figure 30]Figures 30A-30B show DNA specimen analysis for sample T23. Figure 30A shows the S8897 tumor and matched LN. Initial sample ID, Tu20; Final sample ID, T23; HistoPath, DCIS & IDC (4:1); Grade 2; Tumor rate = 50%; Molecular subtype (ddPCR); HER2. Figure 30B shows the chromosomal bands associated with IntClust and the associated copy number (HER2 via InClust 5).

[0052] [Figure 31] Figures 31A-C show DNA specimen analysis for sample T26. Figure 31A shows the S8897 tumor and matched LN. Initial sample ID, Tu1; Final sample ID, T26; HistoPath, IDC; Grade 1; Tumor rate 90%. Figure 31B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 8). Figure 31C shows the top 27 targets of the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0053] [Figure 32] Figures 32A-32C show DNA specimen analysis for sample T27. Figure 32A shows the S8897 tumor and matched LN. Initial sample ID, Tu4; Final sample ID, T27; HistoPath, IDC; Grade 2; Tumor rate = 100%. Figure 32B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 7 and 8). Figure 32C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0054] [Figure 33]Figures 33A-33C show DNA specimen analysis for sample T28. Figure 33A shows the S8897 tumor and matched LN. Initial sample ID, Tu5; Final sample ID, T28; HistoPath, IDC; Grade 1; Tumor rate = 60%. Figure 33B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 8). Figure 33C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate mutation impact.

[0055] [Figure 34] Figures 34A-34C show DNA specimen analysis for sample T29. Figure 34A shows the S8897 tumor and matched LN. Initial sample ID, Tu6; Final sample ID, T29; HistoPath, IDC & focal DCIS; Grade 2; Tumor rate = 95%. Figure 34B shows the chromosomal bands associated with IntClust and associated copy number (Lum A via InClust 7). Figure 34C shows the top 27 targets of the UMI-based gene panel, color-coded by mutation impact.

[0056] [Figure 35] Figures 35A-C show DNA specimen analysis for sample T30. Figure 35A shows the S8897 tumor and matched LN. Initial sample ID: Tu7, Final sample ID: T30, HistoPath: IDC & DCIS, Grade 3, Tumor rate: 40%. Figure 35B shows the chromosomal bands associated with IntClust and the associated copy number (Lum A, B via InClust 7; 8, 9). Figure 35C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0057] [Figure 36]Figures 36A-36C show DNA specimen analysis for sample T31. Figure 36A shows the S8897 tumor and matched LN. Initial sample ID: Tu8, Final sample ID: T31, HistoPath: IDC, Grade 3, Tumor rate: 30%. Figure 36B shows the chromosomal bands associated with IntClust and the associated copy number (Lum B via InClust 2). Figure 36C shows the top 27 targets from the UMI-based gene panel, color-coded to indicate the impact of mutations.

[0058] [Figure 37] Figures 37A-37C show DNA specimen analysis for sample T32. Figure 37A shows the S8897 tumor and matched LN. Initial specimen ID, Tu10; Final specimen ID, T32; HistoPath, IDC & DCIS (9:1); Grade 2; Tumor rate = 100%. Figure 37B shows the chromosomal bands associated with IntClust and associated copy number (Lum B via InClust 1). Figure 37C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0059] [Figure 38] Figures 38A-C show DNA specimen analysis for sample T33. Figure 38A shows the S8897 tumor and matched LN. Initial sample ID, Tu14; Final sample ID, T33; HistoPath, DCIS; Grade 2; Tumor rate = 15%. Figure 38B shows the chromosomal bands associated with IntClust and the associated copy number (Lum A / B). Figure 38C shows the top 27 targets from the UMI-based gene panel, color-coded by mutation impact.

[0060] [Figure 39] Figure 39 shows the copy number mapping criteria to PAM50 subtypes.

[0061] [Figure 40] Figure 40 shows specimen T1, S98-3, ER → ER IHC was negative.

[0062] [Figure 41] Figure 41 shows specimens T2, S98-4, and S98-3292 A3 H&E stained.

[0063] [Figure 42] Figure 42 shows specimen T2, S98-4, ER → ER negative.

[0064] [Figure 43] Figure 43 shows specimens T3, S98-5, and S98-3292 A4 H&E stained.

[0065] [Figure 44] Figure 44 shows specimen T3, S98-5, ER → ER IHC negative.

[0066] [Figure 45] Figure 45 shows H&E staining of specimens T4, S98-6, and S98-3292 A5.

[0067] [Figure 46] Figure 46 shows specimen T4, S98-6, ER → ER IHC negative.

[0068] [Figure 47] Figure 47 shows H&E staining of specimens T5, S98-8, and S98-8029 A3.

[0069] [Figure 48] Figure 48 shows specimen T5, S98-8, ER → ER IHC positive.

[0070] [Figure 49] Figure 49 shows specimens T6, S98-9, and S98-8029 A4 H&E stained.

[0071] [Figure 50] Figure 50 shows specimen T6, S98-9, ER → ER IHC positive.

[0072] [Figure 51] Figure 51 shows specimens T7, S98-11, and S98-8045 A1 H&E stained.

[0073] [Figure 52] Figure 52 shows specimen T6, S98-11, ER → ER IHC positive.

[0074] [Figure 53] Figure 53 shows H&E staining of specimens T8, S98-12, and S98-11141 C1.

[0075] [Figure 54] Figure 54 shows the IHC of specimen T8, S98-12, ER. The IHC was inconclusive due to extensive necrosis and crush artifact.

[0076] [Figure 55] Figure 55 shows H&E staining of specimens T1, S98-3, and S98-3292 A1.

[0077] [Figure 56] Figure 56 shows specimens T9, S99-7, and S99-4808 A3 H&E stained.

[0078] [Figure 57] Figure 57 shows specimen T9, S99-7, ER → ER IHC negative.

[0079] [Figure 58] Figure 58 shows specimen T9, S99-7, HER2 positive by IHC → HER2 positive.

[0080] [Figure 59] Figure 59 shows H&E staining of specimens T10, S99-16, and S99-14088 C10.

[0081] [Figure 60] Figure 60 shows specimen T10, S99-16, ER → ER IHC negative.

[0082] [Figure 61] Figure 61 shows H&E staining of specimens T11, S99-17, and S99-14088 C11.

[0083] [Figure 62] Figure 62 shows specimen T11, S99-17, ER → ER IHC negative.

[0084] [Figure 63] Figure 63 shows specimens T12, S99-18, and S99-14088 C12 H&E stained.

[0085] [Figure 64] Figure 64 shows specimen T12, S99-18, ER → ER IHC negative.

[0086] [Figure 65] Figure 65 shows specimens T13, S99-19, and S99-14088 C13 H&E stained.

[0087] [Figure 66] Figure 66 shows specimen T13, S99-19, ER → ER IHC negative. DETAILED DESCRIPTION OF THE INVENTION

[0088] Detailed explanation The disclosure provided herein is based in part on the discovery and development of methods that can be used to assess tumor subtypes and specific marker-based signatures (e.g., risk of recurrence) in archival specimens (e.g., 20-30 years old). This disclosure provides an in vitro, multiparameter, RNA-based molecular diagnostic method for identifying the molecular subtype of a tumor or cancer sampled from a subject. As clinical cancer diagnoses become more clearly defined, prognosis can be better determined, and by identifying a subject's cancer or tumor subtype and correlating that information with treatment outcomes, the predictability of treatment response can be better established. Active subject tissues and archival samples can be used for subtyping. Archival samples are a rich source of discovery when outcome data are available, but current approaches currently do not work well with older, degraded archival samples. With the right molecular diagnostic methods for effective subtyping, they could be used as a retrospective discovery source. Furthermore, such samples can be used to identify situations in which toxic adjuvant therapy can be avoided.

[0089] I. Definition In order to make the present disclosure more readily understandable, certain terms are first defined. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the embodiments of the present disclosure pertain. Although many methods and materials similar, modified, or equivalent to those described herein can be used to implement various embodiments of the present disclosure without undue experimentation, the preferred materials and methods are described herein. In describing and claiming the embodiments of the present disclosure, the following terms will be used in accordance with the definitions set forth below.

[0090] In this specification, concentrations, amounts, and other numerical data may be expressed or presented in range format. It should be understood that such range format is used merely for convenience and brevity and should be interpreted flexibly to include not only the numerical values ​​explicitly stated as range limits, but also all individual numerical values ​​or subranges subsumed within the range, as if each numerical value and subrange were expressly recited. As an example, a numerical range of "about 2 to about 50" should be interpreted to include not only the explicitly recited values ​​of 2 to 50, but also all individual values ​​and subranges within the indicated range. Thus, this numerical range includes the following: 2, 2.4, 3, 3.7, 4, 5.5, 10, 10.1, 14, 15, 15.98, 20, 20.13, 23, 25.06, 30, 35.1, 38.0, 40, 44, 44.6, 45, 48, and subranges such as from 1-3, from 2-4, from 5-10, from 5-20, from 5-25, from 5-30, from 5-35, from 5-40, from 5-50, from 2-10, from 2-20, from 2-30, from 2-40, from 2-50, etc. This same principle applies to ranges that list only a single number as the minimum or maximum value. Furthermore, this interpretation should apply regardless of the broadness of the range or the characteristics being described.

[0091] As used herein, the term "about" refers to the variation of a numerical quantity that may occur through typical measurement techniques and equipment with respect to any quantifiable variable, including, but not limited to, mass, volume, time, distance, or amount. Furthermore, given the solid and liquid handling procedures used in the real world, there are inadvertent errors and variations that may occur due to differences in the manufacture, source, or purity of ingredients used to prepare a composition or carry out a method. The term "about" also encompasses variations of up to ±5%, but may also include ±4%, ±3%, ±2%, ±1%, etc. Whether or not modified by the term "about," the claims include equivalents to the quantities.

[0092] In this disclosure, terms such as "comprises," "comprising," "containing," and "having" can have the meanings given them in U.S. patent law, can mean "includes," "including," and the like, and are generally construed as open-ended terms. The terms "consisting of" or "consisting of" are closed terms, including those specifically recited with such terms, as well as those pursuant to U.S. patent law. "Consisting essentially of" or "consisting essentially of" have the meaning generally accepted in U.S. patent law. Notably, such terms are generally closed terms, allowing for the inclusion of additional items, materials, components, steps, or elements that do not materially affect the basic and novel characteristics or function of the item used in connection therewith. For example, trace elements present in a composition but that do not affect the nature or properties of the composition are permissible under the language "consisting essentially of," even if they are not explicitly recited in the list of items following such terms. Where open-ended terms such as "consisting of" or "including" are used herein, it is understood that direct support should also be given to the phrase "consisting essentially of," as if expressly stated, and vice versa.

[0093] As used herein, "individual," "subject," "host," and "patient" can be used interchangeably and can refer to a human or non-human mammalian subject, e.g., a human, pet, livestock, horse, or other animal, for whom diagnosis, treatment, prevention, or therapy is desired. In some aspects, the subject is a human. In other embodiments, the subject is a human with cancer (e.g., breast cancer), including humans who have undergone or are candidates for resection (surgery) to remove cancerous tissue.

[0094] In one embodiment, the subject may be a rodent, such as a mouse, rat, guinea pig, etc. In another embodiment, the subject may be a livestock animal. Non-limiting examples of suitable livestock animals include pigs, cows, horses, goats, sheep, llamas, alpacas, etc. In yet another embodiment, the subject may be a companion animal. Non-limiting examples of companion animals include pets such as dogs, cats, rabbits, etc. In yet another embodiment, the subject may be a zoological animal. As used herein, "zoo animal" refers to an animal found in a zoo. Such animals include primates, big cats, wolves, bears, etc. In a preferred embodiment, the subject is a human.

[0095] As used herein, the term "surgery" refers to surgical procedures performed to remove cancerous tissue, such as mastectomy, lumpectomy, lymphadenectomy, sentinel lymph node dissection, prophylactic mastectomy, prophylactic oophorectomy, cryotherapy, tumor biopsy, etc. The tumor sample used in the methods of the present invention may be obtained from any of these procedures.

[0096] As used herein, the terms "treat," "treat," or "treatment" refer to both therapeutic treatment and prophylactic or preventative measures, the purpose of which is to prevent or slow (alleviate) an undesirable physiological change or disease / disorder. Beneficial or desired clinical results include, but are not limited to, alleviation of symptoms, reduction in the extent of disease, stabilization of the disease state (i.e., not worsening), delay or slowing of disease progression, improvement or palliation of the disease state, and remission (partial or total), whether detectable or undetectable. "Treatment" can also mean prolonging survival compared to expected survival if not treated. Those in need of treatment include those already suffering from the disease, condition, or disorder, as well as those susceptible to the disease, condition, or disorder, and those in whom the disease, condition, or disorder is desired to be prevented.

[0097] As used herein, "cancer," "tumor," or "malignant tumor" may refer to one or more neoplasms or cancers. Neoplasms can be malignant or benign, cancers can be primary or metastatic, and neoplasms or cancers can be early-stage or late-stage. Non-limiting examples of neoplasms or cancers include acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendix cancer, astrocytoma (pediatric cerebellar or cerebral), basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, brain stem glioma, brain tumors (cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma), breast cancer, bronchial adenoma / carcinoid, Burkitt's lymphoma, carcinoid tumor (small Children, gastrointestinal), cancer of unknown primary, central nervous system lymphoma (primary), cerebellar astrocytoma, brain astrocytoma / malignant glioma, cervical cancer, childhood cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorder, colon cancer, cutaneous T-cell lymphoma, deneoplastic small round cell tumor, endometrial cancer, ependymoma, esophageal cancer, Ewing's sarcoma of the Ewing's tumor family, extracranial germ cell tumor (childhood), extramaxillary germ cell tumor, extrahepatic bile duct cancer, eye cancer (intraocular melanoma, retinoblastoma), gallbladder cancer, stomach cancer, gastrointestinal carcinoid tumor, gastrointestinal Stromal tumors, germ cell tumors (pediatric extracranial, extraovarian), gestational trophoblastic tumors, gliomas (adult, pediatric brainstem, pediatric brain astrocytoma, pediatric visual pathway and hypothalamic), gastric carcinoid, hairy cell leukemia, head and neck cancer, hepatocellular (liver) cancer, Hodgkin's lymphoma, hypopharyngeal cancer, hypothalamic and visual pathway gliomas (childhood), intraocular melanoma, pancreatic islet cell carcinoma, Kaposi's sarcoma, kidney cancer (renal cell carcinoma), laryngeal cancer, leukemia (acute lymphoblastic, acute myeloid, chronic lymphocytic, chronic myeloid, hairy cell), lip and oral cavity cancer, liver Cancer (primary), lung cancer (non-small cell, small cell), lymphoma (AIDS-related, Burkitt's disease, cutaneous T-cell, Hodgkin's disease, non-Hodgkin's disease, central nervous system primary), macroglobulinemia (Waldenstrom's), bone malignant fibrous histiocytoma / osteosarcoma, medulloblastoma (childhood), melanoma, intraocular melanoma, Merkel cell carcinoma, mesothelioma (cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma), metastatic squamous cell carcinoma of the neck of unknown primary origin, oral cancer,Multiple endocrine neoplasia syndrome (childhood), multiple myeloma / plasma cell neoplasm, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative disease, myeloid leukemia (chronic), myeloid leukemia (cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma)1), multiple myeloma, myeloproliferative disease (chronic), nasal cavity and paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, oral cancer, oropharyngeal cancer, osteosarcoma / Malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial cancer (cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma)2), ovarian germ cell tumor, ovarian low-grade malignant tumor, pancreatic cancer, pancreatic cancer (islet cell), paranasal sinus cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal embryonal tumor, pineoblastoma and supraventricular primitive neuroectodermal tumor (childhood), pituitary adenoma, plasma cell neoplasm, pleuropulmonary blastoma, primary central nervous system tumor Nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma (kidney cancer), renal pelvis / ureteral transitional cell carcinoma, retinoblastoma, rhabdomyosarcoma (childhood), salivary gland cancer, sarcoma (cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma)3), Sézary syndrome, skin cancer (non-melanoma, melanoma), skin cancer (Merkel cell), small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, cervical squamous cell carcinoma of unknown primary (metastatic), stomach cancer, brain Epithelial neuroectodermal tumor (childhood), T-cell lymphoma (skin), testicular cancer, pharyngeal cancer, thymoma (childhood), thymoma and thymic carcinoma, thyroid cancer, thyroid cancer (childhood), transitional cell carcinoma of the renal pelvis and ureter, choriocarcinoma (gestational), primary site unknown (adult, childhood), transitional cell carcinoma of the ureter and renal pelvis, urethral cancer, endometrial cancer, uterine sarcoma, vaginal cancer, visual pathway and hypothalamic glioma (childhood), vulvar cancer, Waldenstrom's macroglobulinemia, Wilms' tumor (childhood). (adult malignancies, children) (adult acute phase, children acute phase) (surface epithelial-stromal tumors) (Ewing family tumors, Kaposi's tumor, soft tissue tumors, uterine tumors)

[0098] As used herein, a "sample," "biological sample," or "specimen" may refer to any biological tissue, fluid, or cell from a subject. The term "tumor" as used herein refers to all neoplastic cell proliferation and growth, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues. The term "tumor sample" as used herein refers to a sample containing tumor material obtained from a cancer subject. The sample may be solid or liquid. The sample may be a heterogeneous population of cells. Non-limiting examples of suitable biological samples include sputum, serum, blood, blood cells (e.g., white blood cells), biopsy, urine, peritoneal fluid, pleural effusion, or cells derived therefrom. The biopsy may be a fine-needle aspiration biopsy, core needle biopsy, vacuum-assisted biopsy, open biopsy, shave biopsy, punch biopsy, incisional biopsy, scraping biopsy, or deep shave biopsy. Biological samples also include tissue sections, such as frozen sections and formalin-fixed sections, taken for histological purposes. The sample may be tumor tissue, peritumor tissue, or non-tumor tissue. A tumor sample can be a fixed, wax-embedded tissue sample, such as a formalin-fixed, paraffin-embedded tissue sample. Furthermore, the term "tumor sample" encompasses samples comprised of tumor cells obtained from a site other than the primary tumor, e.g., circulating tumor cells. The term also encompasses cells that are the progeny of tumor cells in a subject, e.g., cell culture samples derived from primary tumor cells or circulating tumor cells. The term further encompasses samples that may contain proteins or nucleic acid material shed by tumor cells in vivo, e.g., bone marrow, blood, plasma, serum, etc. The term also encompasses samples enriched for tumor cells, samples that have been further manipulated after collection, and samples containing polynucleotides and / or polypeptides obtained from tumor material in a subject. Methods for obtaining biological samples from a subject are well known in the art.

[0099] A specimen from a subject can be collected one or more times before, during, or after diagnosis. In some embodiments, a sample can be collected from a subject before, during, and / or after treatment for cancer. In some embodiments, a control sample can be prepared from a healthy subject. In some embodiments, the control sample consists of non-cancerous cells. In some embodiments, the non-cancerous cells can be from the same tissue type as the cancer cells. For example, if the cancer cells are from breast cancer, the non-cancerous cells can be from healthy breast tissue. In some embodiments, the control can comprise the average level of a molecular profile in samples from a subject before the onset of cancer. In some embodiments, the control sample can be a sample from a subject before diagnosis or treatment. In certain embodiments, the molecular profile can be measured in a person other than the subject with cancer. In some embodiments, the control is a person or persons with similar characteristics to the subject with cancer. In some embodiments, the control can be an average combination of the disclosed molecular profile levels from different healthy sources (e.g., multiple healthy control subjects). In some embodiments, the control sample can be a pooled sample.

[0100] As used herein, "nucleic acid" refers to a polymer of ribonucleosides and / or deoxyribonucleosides, typically covalently linked by intersubunit phosphodiester bonds, but sometimes by phosphorothioates, methylphosphonates, etc. Examples of nucleic acids include genomic DNA, circular DNA, low molecular weight DNA, plasmid DNA, circulating DNA, circulating tumor DNA (ctDNA), hnRNA, mRNA, non-coding RNA, including rRNA, tRNA, microRNA, small interfering RNA, small nucleolar RNA, small nuclear RNA, and small temporal RNA; fragmented or degraded nucleic acids; PNA; nucleic acids obtained from intracellular organelles such as mitochondria and chloroplasts; and nucleic acids obtained from microorganisms, parasites, DNA or RNA viruses that may be present in a biological sample.

[0101] As used herein, the terms "level," "amount," or "abundance" refer to the qualitative or quantitative determination of the number of copies of a coding or non-coding RNA transcript, polypeptide / protein, or analyte. An RNA transcript or polypeptide / protein exhibits an "increased level" if the level of the RNA transcript or polypeptide / protein is higher in a first sample, such as a clinically relevant subpopulation of patients (e.g., patients who have experienced a cancer recurrence), than in a second sample, such as a related subpopulation (e.g., patients who have not experienced a cancer recurrence). In the context of analyzing RNA transcript or polypeptide / protein levels in tumor samples obtained from individual patients, an RNA transcript or polypeptide / protein exhibits an "increased level" if the level of the RNA transcript or polypeptide / protein in the subject shifts toward or more closely resembles the level characteristic of a clinically relevant subpopulation of patients. The amount may be the concentration, number, ratio, proportion, or percentage of the analyte compared to a control sample, or determined using a standard curve. The amount may be an absolute or relative amount.

[0102] Thus, for example, if the analyzed RNA transcript is one that exhibits increased levels in subjects who have experienced long-term cancer recurrence-free survival compared to subjects who have not, the "increased" level of the given RNA transcript can be expressed as being positively correlated with the likelihood of long-term cancer recurrence-free survival. If the level of the RNA transcript in the individual patient being evaluated trends toward a level characteristic of subjects who have experienced long-term cancer recurrence-free survival, the level of the RNA transcript supports the determination that the individual patient is more likely to experience long-term cancer recurrence-free survival. If the level of the RNA transcript in the individual patient trends toward a level characteristic of subjects who have experienced cancer recurrence, the level of the RNA transcript supports the determination that the individual patient is more likely to experience cancer recurrence.

[0103] As used herein, " degraded sample " refers to the sample that may be more degraded than the normal sample used for expression analysis, and has some degree of degradation.After surgical treatment, the biopsy sample from tumor is routinely preserved by FFPE, which may damage the integrity of DNA and RNA.In some embodiments, the degraded sample is comprised of FFPE sample.

[0104] As used herein, a "preserved" or "preserved" sample is a paraffin-embedded and / or fixed tissue biopsy, or a paraffin-embedded and / or fixed tissue section or portion thereof, e.g., a microdissected sample.

[0105] As used herein, "biopsy" refers to any type of needle biopsy or any type of tissue sample taken during surgery.

[0106] As used herein, the term "tissue section" refers to any portion of a biopsy obtained, for example, by sectioning the biopsy with a microtome.

[0107] The term "likelihood score" refers to an arithmetic or mathematically calculated numerical value that simplifies or reveals or helps inform the analysis of more complex quantitative information, such as the correlation between a particular level of a disclosed RNA transcript, its expression product, or gene network and the likelihood of a particular clinical outcome in a breast cancer patient, such as the likelihood of long-term survival without breast cancer recurrence. A likelihood score may be determined by applying a particular algorithm. The algorithm used to calculate the likelihood score can group RNA transcripts or their expression products into a gene network. A likelihood score can be determined for a gene network by determining the levels of one or more RNA transcripts or their expression products and weighting their contribution to a particular clinical outcome, such as recurrence. A patient's likelihood score may also be determined. In some embodiments, the likelihood score is a recurrence score, and an increase in the recurrence score negatively correlates with an increase in the likelihood of long-term survival without breast cancer recurrence. In other words, an increase in the recurrence score correlates with a poor prognosis. An example of a method for determining a likelihood score or recurrence score is disclosed in U.S. Pat. No. 7,526,387. No. 7,526,387.

[0108] As used herein, the term "long-term" survival means survival for at least three years. In other aspects, it can refer to survival for at least five years or at least ten years after surgery or other treatment.

[0109] As used herein, the term "normalized" with respect to a coding or non-coding RNA transcript, or an expression product of a coding RNA transcript, refers to the level of the RNA transcript or its expression product relative to the average level of the transcript / product of a set of reference RNA transcripts or their expression products. The reference RNA transcripts or their expression products are based on minimal variation across patients, tissues, or treatments. Alternatively, the coding or non-coding RNA transcripts or their expression products can be normalized to the totality of the tested RNA transcripts, or a subset of such tested RNA transcripts.

[0110] As used herein, the term "pathology" of cancer includes all phenomena that constitute the well-being of the patient, including, but not limited to, abnormal or uncontrolled cell proliferation, metastasis, interference with the normal function of neighboring cells, release of abnormal levels of cytokines or other secretory products, suppressed or exacerbated inflammatory or immunological responses, neoplasia, premalignancy, malignancy, invasion of surrounding or distant tissues or organs such as lymph nodes, etc.

[0111] "Subject response" includes, but is not limited to, (1) inhibition of tumor growth (including slowing and complete cessation of growth), (2) reduction in tumor cell count, (3) reduction in tumor size, (4) inhibition of tumor cell infiltration into adjacent peripheral organs and / or tissues (i.e., reduction in size, slowing down, complete cessation, etc.), (5) inhibition of metastasis (i.e., reduction in size, slowing down, complete cessation, etc.), (6) enhancement of an anti-tumor immune response (which may, but need not, result in tumor regression or rejection), (7) inhibition of metastasis (i.e., reduction in size, slowing down, complete cessation, etc.); (6) enhancement of an anti-tumor immune response (which may, but need not, result in tumor regression or rejection); (7) alleviation to some extent of one or more symptoms associated with cancer; (8) increased survival after treatment; and / or (9) reduced mortality at a given time point after treatment.

[0112] The term "prognosis" as used herein refers to the prediction of the likelihood of cancer-related death or progression, including recurrence, metastasis, and drug resistance, of a neoplastic disease such as breast cancer. The term "prediction" as used herein refers to the likelihood that a patient will respond favorably or unfavorably to a drug or series of drugs, as well as the degree of such response, or the likelihood that a patient will survive for a certain period of time without cancer recurrence after surgical resection of the primary tumor and / or chemotherapy. The methods of the present invention can be used clinically to determine treatment strategies by selecting the most appropriate treatment for a particular patient. The methods of the present invention are tools for predicting whether a patient is likely to respond favorably to a treatment regimen, such as surgical intervention, chemotherapy with a specific drug or drug combination, and / or radiation therapy, or whether a patient will survive long-term without cancer recurrence after completion of surgery and / or chemotherapy or other treatments.

[0113] The term "breast cancer prognostic biomarker" refers to an RNA transcript, or its expression product, intronic RNA, lincRNA, intergenic sequence, and / or intergenic region, that has been found to be associated with long-term breast cancer recurrence-free survival, as disclosed herein.

[0114] The term "reference" RNA transcript or expression product thereof, as used herein, refers to an RNA transcript or expression product thereof that can be used to compare the level of the RNA transcript or expression product thereof in a test sample. In one embodiment of the invention, reference RNA transcripts include housekeeping genes such as β-globin, alcohol dehydrogenase, and other RNA transcripts whose levels or expression do not change due to the disease state of the cell containing the RNA transcript or its expression product. In another embodiment, all of the assayed RNA transcripts, or their expression products, or a subset thereof, may serve as reference RNA transcripts or reference RNA expression products.

[0115] As used herein, the term "RefSeq RNA" refers to RNA that can be found in the Reference Sequence (RefSeq) database, a collection of publicly available nucleotide sequences and their protein products compiled by the National Center for Biotechnology Information (NCBI). The RefSeq database provides an annotated, non-redundant record for each naturally occurring biological molecule (DNA, RNA, or protein) included in the database. Therefore, the sequences of RefSeq RNA are well known and can be found at the RefSeq database on the internet site: www (dot) ncbi (dot) nlm (dot) nih (dot) gov (slash) RefSeq (slash). See also Pruitt et al., Nucl. Acids Res: D501-D504 (2005). The accession numbers for each RefSeq (including accession numbers for alternative splice forms) are listed in Tables 1 and 2, Table B. Nevertheless, the coordinates for each intronic sequence listed in Table 3 are listed in Table A. Therefore, the base sequences of each RNA sequence in Tables 1 to 3 and Table 15 can be readily obtained from publicly available sources.

[0116] As used herein, the term "RNA transcript" refers to the RNA transcript of DNA, including coding RNA transcript and non-coding RNA transcript.RNA transcript includes, for example, mRNA, unspliced ​​RNA, splice variant mRNA, microRNA, fragmented RNA, long intergenic non-coding RNA (lincRNA), intergenic RNA sequence or region, and intron RNA.

[0117] The term "RNA-Seq" or "transcriptome sequencing" refers to sequencing performed on RNA (or cDNA) instead of DNA, typically with the primary goal of measuring expression levels, detecting fusion transcripts, alternative splicing, and other genomic changes that can be better assessed from RNA. RNA-Seq includes whole-transcriptome sequencing and target-specific sequencing.

[0118] As used herein, the term "computer-based system" refers to the hardware, software, and data storage means used to analyze information. The minimum hardware for a patient computer-based system consists of a central processing unit (CPU), input means, output means, and data storage means. Those skilled in the art will readily appreciate that many currently available computer-based systems are suitable for use with the present invention and can be programmed to perform the specific measurement and / or calculation functions of the present invention.

[0119] To "record" data, programming, or other information on a computer-readable medium refers to a process for storing the information using such methods as are known in the art. Any convenient data storage structure can be selected based on the means used to access the stored information. A variety of data processing programs and formats can be used for storage, such as word processing text files, database formats, and the like.

[0120] A "processor" or "computing means" refers to a combination of hardware and / or software that performs the required functions. For example, any processor herein may be a programmable digital microprocessor, such as those available in the form of an electronic controller, mainframe, server, or personal computer (desktop or portable). If the processor is programmable, the appropriate programming may be transmitted to the processor remotely or pre-stored on a computer program product (such as a portable or fixed computer-readable storage medium, whether magnetic, optical, or solid-state device-based). For example, a magnetic medium or optical disk may carry the programming, which may be read by an appropriate reader that communicates with each processor at the corresponding station.

[0121] II. Nucleic Acid Extraction Methods In some embodiments, the present disclosure provides DNA and RNA extraction from unstained FFPE material mounted on glass microscope slides that has been stored for over 30 years without special handling or storage conditions. In some aspects, the samples used for extraction can be excised from 20-year-old FFPE blocks and stored without special handling or storage conditions. Rigorous QA and QC assessments can be performed along with advanced molecular profiling consisting of next-generation sequencing (NGS) approaches to profile DNA mutations and copy number alterations. RNA-seq can be performed, and custom droplet digital PCR (ddPCR) assays have been developed and optimized for archival samples to determine tumor molecular subtypes.

[0122] In some embodiments, a method for molecular subtyping a tumor sample obtained from a subject includes the steps of: a) enhancing solubilization of old, degraded FFPE sample material through the non-discretionary use of mineral oil; b) digesting the sample with a proteinase; incubating the digested sample from step a) with DNase; incubating the mixture from step b) in a guanidine salt-based buffer; concentrating and isolating RNA; pre-amplifying; and performing digital droplet PCR.

[0123] The methodologies disclosed herein are described in more detail below.

[0124] (a) Sample The present disclosure encompasses methods for molecular subtyping of tumors in nucleic acid-containing samples. The term encompasses tumor tissue samples, such as tissue obtained by surgical resection or by biopsy, such as core biopsy or fine needle biopsy. In certain embodiments, the tumor sample is a fixed, wax-embedded tissue sample, such as a formalin-fixed, paraffin-embedded tissue sample.

[0125] Methods for fixing and embedding tissue samples are known in the art. In some embodiments, the method involves fixing, dehydrating, clearing, infiltrating or impregnating the specimen with paraffin, blocking or embedding the block and specimen in a paraffin block, slicing the block and specimen into thin sections, mounting the sections on microscope slides, or any combination thereof. In some embodiments, the tissue sample can be fixed in a formalin solution (e.g., a 10% formalin solution contains 3.7% formaldehyde and 1.0-1.5% methanol). The sample can then be dehydrated, such as by placing it in alcohol, and the alcohol can be "removed" by exposure to a solvent such as xylene. In this case, the sample is embedded in paraffin, and paraffin replaces the xylene in the sample and surrounds the sample. In some embodiments, the fixed and embedded tissue sample is further stored. In some embodiments, storage is at room temperature. In some aspects, the sample is old or deteriorated.

[0126] In one aspect, the present disclosure encompasses a method for identifying molecular subtypes in old / aged or fresh samples containing nucleic acids. Generally speaking, uncertainty about the fidelity of aged RNA samples remains a serious limitation. Aged tissue processing and sample storage are known to result in highly degraded RNA, which limits detection and leads to sequencing artifacts. However, the present disclosure provides a method to overcome these limitations. The age of a sample can be determined based on the time elapsed since the sample was obtained from a subject. In some embodiments, an old sample for use within the present methods can be at least 1 month old, at least 2 months old, at least 3 months old, at least 4 months old, at least 5 months old, at least 6 months old, at least 7 months old, at least 8 months old, at least 9 months old, at least 10 months old, at least 11 months old, at least 1 year old, at least 2 years old, at least 3 years old, at least 4 years old, at least 5 years old, at least 6 years old, at least 7 years old, at least 8 years old, at least 9 years old, at least 10 years old, at least 11 years old, at least 12 years old, at least 13 years old, at least 14 years old, at least 15 years old, at least 16 years old, at least 17 years old, at least 18 years old, at least 19 years ... 0 years ago, at least 21 years ago, at least 22 years ago, at least 23 years ago, at least 24 years ago, at least 25 years ago, at least 26 years ago, at least 27 years ago, at least 28 years ago, at least 29 years ago, at least 30 years ago, at least 31 years ago, at least 32 years ago, at least 33 years ago, at least 34 years ago, at least 35 years ago, at least 36 or more years ago, at least 37 or more years ago, at least 38 or more years ago, at least 39 or more years ago, at least 40 or more years ago, at least 41 or more years ago, at least 42 or more years ago, at least 43 or more years ago, at least 44 or more years ago, at least 45 or more years ago, at least 46 or more years ago, at least 47 or more years ago, at least 48 or more years ago, at least 49 or more years ago, at least 50 or more years ago.

[0127] In some embodiments, the present disclosure encompasses a method for identifying molecular subtypes in a degraded sample. In some embodiments, the degraded sample is an FFPE sample. In some embodiments, the method comprises identifying molecular subtypes in an old, degraded sample. In some embodiments, the old, degraded sample is an FFPE. In some embodiments, the FFPE sample for use in the method can be at least 1 month old, at least 2 months old, at least 3 months old, at least 4 months old, at least 5 months old, at least 6 months old, at least 7 months old, at least 8 months old, at least 9 months old, at least 10 months old, at least 11 months old, at least 1 year old, at least 2 years old, at least 3 years old, at least 4 years old, at least 5 years old, at least 6 years old, at least 7 years old, at least 8 years old, at least 9 years old, at least 10 years old, at least 11 years old, at least 12 years old, at least 13 years old, at least 14 years old, at least 15 years old, at least 16 years old, at least 17 years old, at least 18 years old, at least 19 years old, at least 20 years ago, at least 21 years ago, at least 22 years ago, at least 23 years ago, at least 24 years ago, at least 25 years ago, at least 26 years ago, at least 27 years ago, at least 28 years ago, at least 29 years ago, at least 30 years ago, at least 31 years ago, at least 32 years ago, at least 33 years ago, at least 34 years ago, at least 35 years ago, at least 36 or more years ago, at least 37 or more years ago, at least 38 or more years ago, at least 39 or more years ago, at least 40 or more years ago, at least 41 or more years ago, at least 42 or more years ago, at least 43 or more years ago, at least 44 or more years ago, at least 45 or more years ago, at least 46 or more years ago, at least 47 or more years ago, at least 48 or more years ago, at least 49 or more years ago, at least 50 or more years ago.

[0128] In some embodiments, the sample is a tissue sample, a cell sample, a whole blood sample, a plasma sample, a serum sample, or a combination thereof. In some aspects, the tissue sample is a cancer or tumor tissue sample. In some embodiments, the cancer or tumor tissue sample consists of cancer or tumor cells, tumor-infiltrating immune cells, stromal cells, or a combination thereof. In some embodiments, the sample is a fixed sample. In some embodiments, the fixed sample is fixed with a compound selected from the group consisting of formalin, glutaraldehyde, alcohol, osmic acid, and paraformaldehyde. In some embodiments, the sample is paraffin-embedded. In some embodiments, the sample is a formalin-fixed paraffin-embedded sample selected from the group consisting of fine needle aspirate (FNA), core biopsy, and needle biopsy. In some embodiments, the cancer or tumor tissue sample is a formalin-fixed paraffin-embedded (FFPE) sample, an archived sample, a fresh sample, or a frozen sample. In some embodiments, the cancer or tumor tissue sample is an FFPE sample.

[0129] In some embodiments, the cancer is selected from the group consisting of lung cancer, kidney cancer, bladder cancer, breast cancer, colorectal cancer, ovarian cancer, pancreatic cancer, gastric cancer, esophageal cancer, mesothelioma, melanoma, head and neck cancer, thyroid cancer, sarcoma, prostate cancer, glioblastoma, cervical cancer, thymic cancer, leukemia, lymphoma, myeloma, mycosis fungoides, Merkel cell carcinoma, or hematological malignancies. In some aspects, the cancer is lung cancer, kidney cancer, bladder cancer, or breast cancer. In some aspects, the lung cancer is non-small cell lung cancer (NSCLC). In some aspects, the kidney cancer is renal cell carcinoma (RCC). In some aspects, the bladder cancer is urothelial bladder cancer (UBC). In some aspects, the breast cancer is triple-negative breast cancer (TNBC).

[0130] (b) Nucleic acid extraction In one embodiment, the method of the present disclosure includes, in part, a method for extracting nucleic acids from old and potentially deteriorated formalin-fixed, paraffin-embedded (FFPE) samples. For example, in the case of archival FFPE samples, tissues are often mounted on glass slides. Before further processing of slide-mounted FFPE samples, the use of an organic solvent is essential, which promotes sample solubilization, ultimately improving the release of nucleic acids and overall yield. In one embodiment, the organic solvent is xylene, CitriSolv, or mineral oil. In one embodiment, the organic solvent is mineral oil. In one embodiment, the mineral oil is light mineral oil.

[0131] In some embodiments, the method comprises extracting nucleic acids from an FFPE sample comprising a first heating step in which the sample is heated in the presence of mineral oil. In one embodiment, the sample is heated to a temperature between about 65°C and about 95°C. In some embodiments, the sample is heated to a temperature of about 80°C. In one embodiment, the first heating step can be from about 5 minutes to about 15 minutes. In some embodiments, the first heating step is about 10 minutes.

[0132] In some embodiments, the method for extracting nucleic acids from an FFPE sample further comprises digesting the sample with a proteinase. In certain embodiments, the protease is proteinase K or trypsin. In some embodiments, the protease is proteinase K. In some embodiments, the sample and proteinase can be incubated for a period of time (e.g., between 15 minutes and 1 hour, or for 1 hour or more). In certain embodiments, the sample is incubated with the proteinase for about 1 hour. In certain embodiments, the sample is incubated with the proteinase for about 15 minutes. In certain embodiments, the incubation of the sample with the proteinase is for about 30 minutes.

[0133] After the initial heating step, the sample can be digested in an appropriate digestion buffer, such as, but not limited to, Proteinase K Digestion Buffer (e.g., commercially available from Qiagen). The sample can be incubated in the digestion buffer for a suitable period of time to allow for solubilization and release of DNA. Because much of the DNA is wrapped around histone proteins, enhanced solubilization improves sample purification and enhances polymerase activity in downstream processes. In one embodiment, the step of digesting the sample involves incubating the sample in the digestion buffer at a temperature of about 55°C to about 80°C, optionally with constant agitation of the buffer and sample. In a preferred embodiment, the sample is digested between about 65°C to about 70°C for about 45 minutes. After about 45 minutes, the digestion buffer and sample are heated between about 60°C to about 100°C for about 15 minutes and then centrifuged to separate the upper and lower phases. In some embodiments, additional Proteinase K is added to the lower phase, and the Proteinase K and sample are incubated at a temperature of about 55°C to about 80°C for about 30 minutes, optionally with agitation.

[0134] After digesting the sample with PKD buffer, the sample can be centrifuged to separate the aqueous phase (lower layer) from the remaining lysate. In some embodiments, in addition to protease treatment, the sample can be subjected to RNase digestion to remove residual RNA. Conversely, when treating an RNA sample, the sample can be subjected to DNase digestion to remove residual DNA. In some embodiments, DNase is added to the aqueous phase and incubated for a period of time. In some embodiments, the sample incubation can be extended for a suitable period of time, for example, more than one hour, or less than one hour (e.g., 15 minutes). In some embodiments, the sample is incubated with DNAse for approximately 15 minutes.

[0135] In one embodiment, buffer RBC is added to the aqueous phase-DNase mixture and mixed before adding ethanol. In a further embodiment, RNA is purified and concentrated, for example, using an RNeasy MinElute spin column. The concentrated and purified RNA is eluted with RNase-free water. Those skilled in the art will appreciate that suitable equivalent buffers are useful in the disclosed methods. For example, a guanidine hydrochloride-based RBC buffer (a guanidine salt-based buffer) is used to adjust RNA binding conditions, although alternative guanidine salt-based buffers can also be "home-made." These types of approaches (e.g., nucleic acid / RNA binding conditions) are described in the Molecular Cloning - Lab Manual (also known as the Maniatis manual), which is incorporated herein by reference. As another example, other alcohols can be used, such as EtOH versus isopropanol versus a combination-based approach. For example, the Qiagen All-Prep kit first uses isopropanol to aid in RNA isolation, and then EtOH is used for DNA in a later step. Therefore, other alcohols can be used and "stepped in" depending on the goals and objectives of the protocol.

[0136] In a specific embodiment, the protocol for extracting RNA from FFPE samples is as follows: Preparation stage: 1. Set the rotisserie oven to 65-70°C. 2. Set the heating block to 80°C. 3. Clean the razor blades with EtOH. One razor blade is needed per sample. Day 1 4. Apply 10 μl of light mineral oil to wet the tissue. Note: In this step and throughout this protocol, the use of mineral oil or equivalent is not optional. 5. Scrape the tissue of interest from the slide into a 2 ml screw-cap tube (note: there may be a lot of paraffin around the tissue, so be careful not to scrape off any unwanted paraffin). 6. Save the slides for future reference. RNA extraction: 7. Add 0.8 ml of mineral oil to the sample tube. 8. Heat on a heating block (80°C) for 10 minutes. 9. Remove the tube from the heating block and give it a quick spin. 10. Add 360 μl of Buffer PKD. 11. Add 40 μl of Qiagen proteinase K. 12. Incubate in a rotisserie oven set at 65-70°C for 45 minutes. 13. Incubate in a heating block (80°C) for 15 minutes. 14. Quickspin. 15. Add 25 μl Qiagen proteinase K to the lower phase. 16. Incubate in a rotisserie oven set at 65-70°C for 30 minutes. 17. Thaw DNase I on ice. 18. Centrifuge the sample at maximum speed (13 Krpm) for 15 minutes. 19. For each sample, set and label: two 2ml tubes, one Qiagen RNeasy mini elution tube from the FFPE kit (store at 4°C), and one 1.5ml elution tube. 20. After centrifugation, carefully transfer 250 μl of the aqueous (bottom) phase to a new 2 ml labeled tube, without disturbing the pellet or aspirating the mineral oil. Save the remaining lysate tube for later DNA extraction. 21. Add 25 μl of DNase Booster buffer and 10 μl of DNase I solution to the 250 μl aqueous phase. Mix by inverting the tube. 22. Quickspin. 23. Incubate at RT for 15 minutes (if DNA extraction is to be performed the next day, incubate the lysate remaining from step 13 overnight). Add 120 μl of ATL and 30 μl of Proteinase K to each sample, ensuring the tube caps are screwed on tightly. Place in a rotisserie oven set to 65-70°C and incubate overnight). 24. Quickspin. 25. Add 500 μl Buffer RBC (from the RNeasy FFPE kit) and vortex to mix. 26. Divide the sample into two 2 ml tubes, adding 392.5 μl to each tube. 27. Add 875 μl of 100% EtOH to each tube and mix well by pipetting up and down. Proceed immediately to the next step. 28. Transfer 700 μl of sample to a labeled RNeasy MiniElute spin column and centrifuge for 15 seconds. 29. Discard the flow-through and repeat the previous step until the entire sample has passed through the column. 30. Add 500 μl Buffer RPE to the column and centrifuge at maximum speed for 15 seconds. 31. Add 500 μl Buffer RPE to the column and centrifuge at maximum speed for 1 minute. 32. Discard the flow-through, open the spin column, and centrifuge at full speed for 5 minutes. 33. Discard the flow-through and place the spin column into a labeled 1.5 ml elution tube. 34. Add 30 μl RNase-free water directly to the column membrane and incubate at RT for 1 minute. 35. Centrifuge at maximum speed for 1 minute to elute the RNA in a final volume of 30 μl. 36. Check RNA concentration using Qubit.

[0137] In another embodiment, the method involves extracting DNA from FFPE samples by incubating the residual lysate, as described above, with ATL and proteinase K buffer at temperatures between about 55°C and about 80°C, optionally with constant agitation, for about 11 to about 15 hours. ATL buffer is a tissue lysis buffer used for nucleic acid purification. It contains Buffer ATL·EDTA and SDS sodium dodecyl sulfate (SDS), an anionic surfactant (detergent) that aids in tissue lysis by disrupting noncovalent protein bonds, aiding in the overall denaturation process critical for nucleic acid release. It should also be noted that this protocol is not dependent on a specific column (e.g., Qiagen), as modifications using "bead-based" methods for the molecular separation step have been successfully implemented.

[0138] In some embodiments, proteinase K buffer can be added during the incubation period to ensure complete digestion of the tissue. After complete digestion of the tissue, RNase A can be added along with binding buffer PM and sodium acetate. Buffer PM is a molecular biology binding buffer consisting of guanidinium chloride and 2-propanol. The buffer PM replacement solution of the present invention is 64% tetraethylene glycol (v / v); 24% ethanol (v / v); 100 mM NaCl; 10 mM Tris pH 7.5. The DNA in the mixture can then be isolated and concentrated, for example, using a spin column. The DNA can then be eluted with a heated elution buffer.

[0139] In a specific embodiment, the protocol for extracting DNA from FFPE samples is as follows: 1. Add 120 μl of ATL and 30 μl of Proteinase K to each sample of residual lysate (obtained from the RNA extraction protocol), ensure the tube caps are screwed on tightly, and place in a rotisserie oven set to 65-70°C and incubate overnight. 2. Add 20 μl proteinase K to the aqueous phase and pipette up and down. 3. Place back in the rotisserie oven at 65-70°C for an additional 1-2 hours to allow the tissue to fully digest. 4. If any unlysed tissue remains, repeat steps 2-3. 5. While waiting, label one 1.5 ml Eppendorf tube for binding buffer mixing, one Qiaquick spin column, one collection tube, and one 1.5 ml elution tube for each sample. Prepare an 80% EtOH solution. Set the heat block to 65°C. 6. Quickspin. 7. Add 1.5 μl RNase A, vortex for 3-5 seconds, quick spin, and incubate at room temperature for 5 minutes. 8. Add 490 μl Binding Buffer PM and 10 μl 3 M Sodium Acetate to a labeled 1.5 ml Eppendorf tube. 9. Add approximately 250 μl of the aqueous phase (bottom layer) to the tube (check the volume of the aqueous phase first). Then, pipette in 200 μl first, then slowly pipette in the remaining amount. Mix the sample and buffer by pipetting up and down. Store the remaining eluted tissue in the 2 ml tube indefinitely until the presence of DNA in the eluted sample has been confirmed. 10. Add 700 μl of sample to a labeled Qiaquick spin column and centrifuge at 2,000 rpm for 90 seconds. 11. Not all DNA will bind to the column, so you may need to reapply the flow-through; save the flow-through and the labeled collection tube. 12. Place the spin column in a new collection tube, add the remaining sample from step 9, and centrifuge at 2,000 rpm for 90 seconds. 13. Place the spin column in a new collection tube, apply the flow-through from step 11, and centrifuge at 2,000 rpm for 90 seconds. Repeat with the remaining flow-through from step 12. 14. Add 700 μl of Buffer PE (this is a wash buffer containing a weak organic base). 15. Add 10 nM Tris-HCl pH 7.5, 80% EtOH and centrifuge at 10,000 rpm for 15 seconds. Discard the flow-through. 16. Add 700 μl of 80% EtOH and centrifuge at maximum speed for 1 minute. 17. Discard the flow-through and centrifuge at maximum speed for 5 minutes. (c) Preparation of preamplification cDNA products and ddPCR

[0140] PCR-based preamplification is a method used to increase the concentration of a specific target panel in a sample prior to qPCR analysis, thereby reducing the sample input required for multi-target qPCR experiments. Preamplification is essentially a highly multiplexed PCR reaction performed for a limited number of cycles using the same primer set used in the downstream qPCR reaction. Using reagents designed for preamplification with limited PCR cycles maintains optimal amplification efficiency for each target, which is essential to prevent bias in the qPCR analysis. As few as 10–14 cycles of preamplification can increase the concentration of each target by more than 1000-fold, providing sufficient preamplified sample to analyze all targets by qPCR without compromising the sensitivity of the qPCR analysis. As a rule of thumb, preamplification is effective when the amount of sample available limits the number of targets that can be efficiently analyzed. Preamplification can be performed using methods well known in the art. In an exemplary embodiment, TaqMan preamp Supermix (ThermoFisher) can be used. In another exemplary embodiment, SsoAdvanced PreAmp Supermix (Bio-Rad) can be used.

[0141] Typically, amplification of a region of interest is performed using polymerase chain reaction (PCR). A PCR reaction can contain a sample containing nucleic acid, one or more primer pairs, polymerase, water, a buffer, and deoxynucleotide triphosphates (dNTPs) in a single reaction vial. PCR can be performed according to standard methods in the art. As a non-limiting example, a PCR reaction can consist of denaturation, followed by about 15 to about 30 cycles of denaturation, annealing, and extension, followed by a final extension. In an exemplary embodiment, the PCR reaction consists of denaturation at about 98°C for about 30 seconds, followed by about 15 to about 30 cycles (about 98°C for about 10 seconds, about 62-72°C for about 30 seconds, and about 72°C for about 30 seconds), followed by a final extension at about 72°C for about 2 minutes.

[0142] In certain cases, Droplet Digital PCR (ddPCR) is used. ddPCR is a digital PCR implementation based on water-oil emulsion droplet technology. The sample is partitioned into 20,000 droplets, and PCR amplification of template molecules occurs in each droplet. ddPCR technology uses reagents and a workflow similar to those used in most standard TaqMan probe-based assays. Large-scale sample partitioning is a key aspect of ddPCR technology. Droplets are formed in a water-oil emulsion, creating partitions that separate template DNA molecules. Droplets essentially perform the same function as individual test tubes or wells in a plate where PCR reactions occur, but in a much smaller format. Large-scale sample partitioning is a key aspect of ddPCR technology. The Droplet Digital PCR system partitions a nucleic acid sample into thousands of nanoliter-sized droplets, and PCR amplification occurs within each droplet. This technology requires less sample than other commercially available digital PCR systems, thereby reducing costs and preserving valuable sample.

[0143] In one embodiment, one or more droplets are formed, each containing a heterogeneous mixture of nucleic acid and primer pairs and probes specific to multiple target sites on the template. For example, a first fluid (continuous or discontinuous as in droplets) containing a single nucleic acid template (DNA or RNA) is merged with a second fluid (also either continuous or discontinuous as in droplets) containing multiple primer pairs and multiple probes, each specific to multiple target sites on the nucleic acid template, to form droplets containing a heterogeneous mixture of the single nucleic acid template and the primer pairs and probes. The second fluid can also contain reagents for performing a PCR reaction, such as polymerase and dNTPs.

[0144] In some embodiments, certain members of the plurality of probes contain a detectable label. Each member of the plurality of probes can contain the same detectable label or different detectable labels. The detectable label is preferably a fluorescent label. The plurality of probes can contain one or more probe groups with different concentrations. One or more probe groups can contain the same detectable label whose intensity changes upon detection due to differences in probe concentration.

[0145] In some embodiments, the first and second fluids can each be in the form of droplets. Techniques known in the art for forming droplets can be used in the methods of the present invention. An exemplary method involves flowing a sample fluid containing nucleic acids across two opposing streams of flowing carrier fluid. The carrier fluid is immiscible with the sample fluid. The crossing of the sample fluid with the two opposing streams of flowing carrier fluid causes the sample fluid to split into individual sample droplets containing the first fluid. The carrier fluid can be any fluid that is immiscible with the sample fluid. An exemplary carrier fluid is oil. In certain embodiments, the carrier fluid comprises a surfactant, such as a fluorosurfactant. The same method can be applied to create individual droplets from a second fluid containing primer pairs (and in some embodiments, amplification reagents). Droplets containing the first fluid, droplets containing the second fluid, or both are formed and stored in a library for later integration.

[0146] In some embodiments, the nucleic acid contained in each of the merged / formed droplets is amplified, for example, by thermocycling the droplets under temperatures / conditions sufficient for a PCR reaction. The amplicons in the droplets can then be analyzed. In some embodiments, the method further includes digital PCR. In some embodiments, the nucleic acid-containing droplets can be combined with PCR reagents in a second fluid as described above to generate droplets containing Taq polymerase, A, C, G, and T deoxynucleotides, magnesium chloride, forward and reverse primers, a detectably labeled probe, and the target nucleic acid. In another embodiment, the first fluid can contain template DNA and a PCR master mix (defined below), and the second fluid can contain forward and reverse primers and a probe. The present disclosure is not limited by the composition of the first and second fluids for PCR or digital PCR. For example, in some embodiments, the template DNA is contained in the second fluid within the droplet. In some embodiments, multiplex primer pairs can be used in droplet-based digital PCR reactions.

[0147] However, the methods of the present disclosure are not particularly limited to any particular PCR detection method. It is noted that PCR methods other than ddPCR (e.g., qPCR) are useful for the method steps disclosed herein.

[0148] In other embodiments, the extracted nucleic acids are subjected to other processing. In some embodiments, the extracted nucleic acids are subjected to molecular profiling. For example, the commercially available Illumina RNA Access or custom NanoString nCounter assays can be used for molecular profiling of the extracted nucleic acids. In some embodiments, the extracted nucleic acids can be used to construct a nucleic acid library. In some embodiments, the extracted nucleic acids are subjected to sequencing, such as whole transcriptome RNA-seq or next-generation sequencing (NGS).

[0149] In a further aspect, the extracted nucleic acid is subjected to molecular target quantification and subtyping.

[0150] III. How to use The methods of the present disclosure can further be used to quantify, sequence, and / or determine cancer subtype using nucleic acid extracted from a sample.

[0151] In some embodiments, the abundance of two or more target nucleic acids or reference nucleic acids may be compared. In some embodiments, the target nucleic acids include one or more hormone targets, human epidermal growth factor receptor 2 (HER2) targets, and proliferation targets. In some embodiments, the hormone targets consist of nucleic acids or fragments thereof encoding estrogen receptor 1 (ESR1), progesterone receptor (PGR), B-cell lymphoma 2 (BCL2), signal peptide, CUB domain, and EGF-like domain-containing 2 (SCUBE2), or any combination thereof. In some embodiments, the HER2 target consists of nucleic acids or fragments thereof encoding human epidermal growth factor receptor 2 (HER2), growth factor receptor-bound protein 7 (GRB7), or any combination thereof. In some embodiments, the proliferation target consists of a nucleic acid or fragment thereof encoding the proliferation marker Ki-67 (MKI67), Aurora Kinase A (AURKA), Baculoviral IAP Repeat Containing 5 (BIRC5), Cyclin B1 (CCNB1), MYB Proto-Oncogene Like 2 (MYBL2), Thymidine Kinase 1 (TK1), or any combination thereof.

[0152] In some aspects, the disclosed methods comprise determining the levels of ESR1, PGR, BCL2, SCUBE2, HER2, GRB7, MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof. In some embodiments, the disclosed methods comprise determining the levels of ESR1, PGR, BCL2, SCUBE2, or any combination thereof. In some embodiments, the disclosed methods comprise determining the relative levels of HER2, GRB7, or any combination thereof. In some embodiments, the disclosed methods comprise determining the levels of MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof. In some embodiments, the disclosed methods comprise determining the levels of ESR1, PGR, BCL2, SCUBE2, HER2, GRB7, MKI67, AURKA, BIRC5, CCNB1, MYBL2, and TK1.

[0153] The level of the target nucleic acid can be determined using the methods described in Section II, such as ddPCR. However, any method known in the art can be used to determine the level of the target nucleic acid. By way of non-limiting example, the level of the target nucleic acid can be measured using RNA-seq, nanopore sequencing, nanostring, multiplex RT-PCR, singleplex RT-PCR, NASBA, fluorometry, or spectrophotometry.

[0154] In a further embodiment of the present disclosure, the method includes determining a z-score of the target nucleic acid. The z-score can be determined by the following formula:

number

number

[0155] In some embodiments, a target nucleic acid is considered to have increased levels if the z-score is greater than or equal to 1. In some embodiments, a target nucleic acid having increased levels comprises a z-score of about 1 to about 5. For example, the z-score is about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.

[0156] In certain embodiments, a target nucleic acid is considered to be reduced in levels if the z-score is <1. In some embodiments, reduced levels of a target nucleic acid comprise a z-score of about -0.1 to about -5. For example, z-scores can range from -0.1, -0.2, -0.3, -0.4, -0.5, -0.6, -0.7, -0.8, -0.9, -1, -1.1, -1.2, -1.3, -1.4, -1.5, -1.6, -1.7, -1.8, -1.9, -2, -2.1, -2.2, -2.3, -2.4, -2.5, -2.6, -2.7, -2.8, -2.9, -3, -3.1, -3.2, -3.3, -3.4, -3.5, -3.6, -3.7, -3.8, -3.9, -4, -4.1, -4.2, -4.3, -4.4, -4.5, -4.6, -4.7, -4.8, -4.9, or -5.

[0157] In further embodiments, the disclosed methods include determining the cancer subtype of a sample using the z-score of the target nucleic acid. In some embodiments, the cancer subtype determined using the disclosed methods consists of luminal A subtype (Lum A), luminal B subtype (Lum B), HER2 subtype, or triple negative subtype (TN).

[0158] In some embodiments, a sample is determined to be Lum A if the level of one or more of the hormone target nucleic acids or fragments thereof is determined to be elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding ESR1 is elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding PGR is elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding BCL2 is elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding SCUBE2 is elevated. In some embodiments, a sample is determined to be Lum A if the level of a nucleic acid or fragment thereof encoding ESR1, PGR, BCL2, SCUBE2 is elevated.

[0159] In some embodiments, a sample is determined to be Lum B if the level of one or more of a hormone target nucleic acid or fragment thereof and a proliferation target nucleic acid or fragment thereof is determined to be elevated. In some embodiments, a sample is determined to be Lum B if the level of a nucleic acid or fragment thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated, and if the level of a nucleic acid or fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated, or if any combination thereof is elevated. In some embodiments, a sample is determined to be Lum B if the level of a nucleic acid or fragment thereof encoding one or more of ESR1, PGR, BCL2, and SCUBE2 is elevated, and if the level of a nucleic acid or fragment thereof encoding one or more of MKI67, AURKA, BIRC5, CCNB1, MYBL2, and TK1 is elevated.

[0160] In some embodiments, if the level of one or more of HER target nucleic acids or fragments thereof is determined to be elevated, the sample is determined to be HER2. In some embodiments, if the level of nucleic acids or fragments thereof encoding HER2, GRB7, or a combination thereof is elevated, the sample is determined to be HER2. In some embodiments, if the level of nucleic acids or fragments thereof encoding HER2 is elevated, the sample is determined to be HER2. In some embodiments, if the level of nucleic acids or fragments thereof encoding GRB7 is elevated, the sample is determined to be HER2. In some embodiments, if the level of nucleic acids or fragments thereof encoding HER2 and GRB7 is elevated, the sample is determined to be HER2.

[0161] In some embodiments, a sample is determined to be TN if the level of one or more of the proliferation target nucleic acids or fragments thereof is determined to be elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding MKI67 is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding AURKA is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding BIRC5 is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding CCNB1 is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding MYBL2 is elevated. In some embodiments, a sample is determined to be TN if the level of a nucleic acid or fragment thereof encoding TK1 is elevated. In some embodiments, a sample is determined to be TN if the levels of nucleic acids encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, and TK1 or fragments thereof are elevated.

[0162] In some embodiments, the disclosed methods provide high sensitivity, accuracy, and / or reproducibility for cancer subtyping of old and / or deteriorated samples. In some embodiments, the sensitivity, accuracy, and / or reproducibility of cancer subtyping is at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or at least 100% higher than other methods for determining cancer subtype, such as immunohistochemistry (IHC), biopsy testing, IntClust algorithm determination, PAM50 assay, etc.

[0163] In further embodiments, the disclosed methods include verifying the quantification of the target nucleic acid and / or determining the cancer subtype. In some aspects, methods including distance geometry studies can be applied for verification. In some embodiments, the subtype determination can be confirmed with the assistance of a qualified pathologist.

[0164] The disclosed methods further include determining the genomic characteristics of the extracted nucleic acid. The disclosure further encompasses sequencing the extracted nucleic acid. In some embodiments, sequencing of the extracted nucleic acid is performed using one of the commercially available sequencing technologies, such as SBS (sequencing by synthesis) by Ulumina, chain termination method of DNA sequencing, or one of the commercially available next-generation sequencing technologies, such as SMRT (single-molecule real-time) sequencing by Pacific Biosciences, or Ion Torrent by ThermoFisher Scientific. TM sequencing, Roche Pyrosequencing (454), Applied Biosystems SOLiD (登録商標)For sequencing, suitable sequencing technology can be selected. In some embodiments, sequencing analysis can be performed by NGS and / or low-pass / ultra-low-pass WGS. Methods related to NGS and WGS are well known in the art.

[0165] In a further aspect, after nucleic acid sequencing, chromosome copy number alterations (CNAs) and / or mutations of target genes can be evaluated. In some embodiments, CNAs can be used to determine the tumor fraction of a sample. In other embodiments, CNAs can be used to determine the subtype of cancer. In some embodiments, CNAs are used to diagnose cancer in a subject.

[0166] In some embodiments, determining CNAs involves isolating DNA from a biological sample using the disclosed methods, sequencing the DNA using ULP-WGS, analyzing the DNA sequence using a statistical model (e.g., using ichorCNA), and characterizing the copy number alterations present in the sample, e.g., by generating a copy number alteration profile.

[0167] In some embodiments, determining mutations in the target gene includes creating a library using nucleic acids extracted using the methods described herein (e.g., using the QIAGEN QIAseq Human Breast Cancer Panel (DHS-001Z, 93 genes) library preparation kit), sequencing using NGS (e.g., Illumina HiSeq 3000, PE 150bp), and performing bioinformatic analysis (e.g., Qiagen, smCounter2-based bioinformatic analysis). In some embodiments, the methods of the present disclosure further include mutation analysis of the target gene. Mutation analysis can further include determining the mutation effect on protein function as a high effect, a medium effect, a low effect, and / or no mutation.

[0168] In some embodiments, the target genes consist of a panel of genes. In some embodiments, the genes consist of one or more genes disclosed in Tables 25-53. In some aspects, the genes involved include adenomatous polyposis coli (APC), ataxia telangiectasia mutated (ATM), ataxia telangiectasia and Rad3-related protein (ATR), BRCA1-associated RING domain 1 (BARD1), BLM RecQ-like helicase (BLM), breast cancer type 1 (BRAC1), breast cancer type 2 (BRAC2), CUB and Sushi Multiple Domains 1 (CSMD1), epidermal growth factor receptor (EGFR), Erb-B2 receptor tyrosine kinase 3 (ERBB3), fibroblast growth factor receptor 2 (FGFR2), GEN1 Holliday junction 5' flap endonuclease (GEN1), HECT and RLD domain-containing E3 ubiquitin protein ligase family member 1 (HERC1), lysine methyltransferase 2 C (KMT2 C), and Mitogen-Activated Protein kinase kinase 1 (MAP3K1), DNA mismatch repair protein Mlh1 (MLH1), MRE11 homolog, double-strand break repair nuclease (MRE11 A), mucin 16 (MUC16), nuclear receptor corepressor 1 (NCOR1), neurofibromatosis type 1 (NF1), palladin, cytoskeleton-associated protein (PALLD), phosphatidylinositol-4,5-bisphosphate 3-kinase, catalytic subunit alpha (PIK3CA), PMS1 homolog 2, mismatch repair system component (PMS2), Ret proto-oncogene (RET), septin 9 (SEPT9), spectrin repeat-containing nuclear envelope protein 1 (SYNE1), tumor protein P53 (TP53), or a combination thereof.

[0169] In some aspects, the target gene comprises APC. In some aspects, the target gene comprises ATM. In some aspects, the target gene comprises ATR. In some aspects, the target gene comprises BARD1. In some aspects, the target gene comprises BLM. In some aspects, the target gene comprises BRAC1. In some aspects, the target gene comprises BRAC2. In some aspects, the target gene comprises CSMD1. In some aspects, the target gene comprises EGFR. In some aspects, the target gene comprises ERBB3. In some aspects, the target gene comprises FGFR2, GEN1. In some aspects, the target gene comprises HERC1. In some aspects, the target gene comprises KMT2 C. In some aspects, the target gene comprises MAP3K1. In some aspects, the target gene comprises MLH1. In some aspects, the target gene comprises MRE11 A. In some aspects, the target gene comprises MUC16. In some aspects, the target gene comprises NCOR1. In some aspects, the target gene comprises NF1. In some aspects, the target gene comprises PALLD. In some aspects, the target gene comprises PIK3CA. In some aspects, the target gene comprises PMS2. In some aspects, the target gene comprises RET. In some aspects, the target gene comprises SEPT9. In some embodiments, the target gene comprises SYNE1. In some aspects, the target gene comprises TP53. In some embodiments, the target gene comprises APC, ATM, ATR, BARD1, BLM, BRAC1, BRAC2, CSMD1, EGFR, ERBB3, FGFR2, GEN1, HERC1, KMT2 C, MAP3K1, MLH1, MRE11 A, MUC16, NCOR1, NF1, PALLD, PIK3CA, PMS2, RET, SEPT9, SYNE1, TP53, or any combination thereof. In some embodiments, the target genes consist of APC, ATM, ATR, BARD1, BLM, BRAC1, BRAC2, CSMD1, EGFR, ERBB3, FGFR2, GEN1, HERC1, KMT2 C, MAP3K1, MLH1, MRE11 A, MUC16, NCOR1, NF1, PALLD, PIK3CA, PMS2, RET, SEPT9, SYNE1, and TP53.

[0170] In some embodiments, target gene mutation analysis can be used to determine the tumor fraction of a sample. In other embodiments, target gene mutation analysis can be used to determine the subtype of cancer. In some embodiments, target gene mutation analysis can be used to diagnose cancer in a subject.

[0171] The disclosed methods are not particularly limited to a particular cancer, and the disclosed methods can be adapted to determine the subtype of any cancer by selecting target nucleic acids associated with a particular cancer.

[0172] The disclosed methods can be used to diagnose, treat, or prevent disease in a subject. Identifying a cancer subtype can facilitate disease diagnosis, allow appropriate treatment, such as a therapeutic agent, or prevent the onset of the disease by administering a prophylactic therapeutic agent. In some embodiments, the disclosed methods can be used to guide cancer treatment and can include modifying the treatment based on analysis of the sample. For example, in some aspects, modifying the treatment can involve changing the amount of one or more therapeutic agents, changing the frequency of administration, or changing the duration over which one or more therapeutic agents are administered to the subject.

[0173] Additionally, the disclosed methods can be used to determine responsiveness to therapeutic agents. With reference to archived samples for which outcome data is known, knowledge gained from the disclosed methods can be used to assess responsiveness to therapeutic agents in cancer subtypes and guide treatment decisions. Additionally, the disclosed methods can be used to assess a subject's health status relative to cancer subtypes and guide treatment decisions based on knowledge gained from archived samples for which outcome data is known.

[0174] In further embodiments, the disclosed methods can be used to monitor side effects in subjects. Knowledge gained from the disclosed methods can be used to correlate side effects with cancer subtypes by referencing archived samples with known outcome data. In some embodiments, if no side effects are present, the disclosed methods can further include continuing treatment of the subject. In some embodiments, if side effects are present, the disclosed methods can further include modifying treatment steps. Methods for monitoring a subject's well-being can include both subjective and objective criteria (as discussed above). Such methods are known to those skilled in the art.

[0175] In some embodiments, various aspects of the disclosed methods can be automated using computer software analysis programs. Accordingly, the present disclosure further provides computer-implemented methods for detecting, comparing, and analyzing patterns of expression or levels of extracted nucleic acids for cancer diagnosis, cancer subtype determination, or treatment course determination in a subject. The analysis program can, for example, interface with a program that is part of an automated nucleic acid detection or quantification system, and data from the automated detection or quantification system can be directly fed to the analysis program. The computer-implemented program can, for example, be implemented to output the identity of nucleic acids in a sample and the degree of increase or decrease in nucleic acid abundance. In further embodiments, the computer-implemented program can be designed to output a cancer subtype based on the analysis of the nucleic acids. The interface between the analysis programs can be direct or indirect. In some embodiments, the program of the present disclosure can be designed to accept information regarding nucleic acid detection or quantification, perform data analysis, and output a cancer subtype assessment. In some aspects, the program of the present disclosure can further output a cancer diagnosis or treatment strategy.

[0176] IV. Kit The present disclosure also encompasses kits for carrying out methods according to one or more aspects of the invention. In some embodiments, the kits can include containers, organic solvents, proteinase K and / or lysis buffers, solutions and / or equipment for nucleic acid extraction, solutions and / or equipment for nucleic acid purification, solutions and / or equipment for nucleic acid amplification, manuals and / or instructions for carrying out the methods, or any combination thereof.

[0177] Common techniques The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, which are within the skill of the art. Such techniques are fully explained in the literature, e.g., Molecular Cloning: A Laboratory Manual, second edition (Sambrook, et al., 1989) Cold Spring Harbor Press; Oligonucleotide Synthesis (MJ Gait, ed. 1984); Methods in Molecular Biology, Humana Press; Cell Biology: A Laboratory Notebook (JE Cellis, ed., 1989) Academic Press; Animal Cell Culture (RI Freshney, ed. 1987); Introduction to Cell and Tissue Culture (JP Mather and PE Roberts, 1998) Plenum Press; Cell and Tissue Culture: Laboratory Procedures (A. Doyle, JB Griffiths, and DG Newell, eds. 1993-98) J. Wiley and Sons; Methods in Enzymology (Academic Press, Inc.); Handbook of Experimental Immunology (D.M. Weir and C.C. Blackwell, eds.): Gene Transfer Vectors for Mammalian Cells (JM Miller and MP Calos, eds., 1987); Current Protocols in Molecular Biology (FM Ausubel, et al. eds. 1987); PCR: The Polymerase Chain Reaction, (Mullis, et al., eds.1994); Current Protocols in Immunology (J. E. Coligan et al., eds., 1991); Short Protocols in Molecular Biology (Wiley and Sons, 1999); Immunobiology (C. A. Janeway and P. Travers, 1997); Antibodies (P. Finch, 1997); Antibodies: a practice approach (D. Catty., ed., IRL Press, 1988-1989); Monoclonal antibodies: a practical approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000); Using antibodies: a laboratory manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999); The Antibodies (M. Zanetti and J. D. Capra, eds. Harwood Academic Publishers, 1995); DNA Cloning: A practical Approach, Volumes I and II (D.N. Glover ed. 1985); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds.(1985); Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984); Animal Cell Culture (R.I. Freshney, ed. (1986); Immobilized Cells and Enzymes (lRL Press, (1986); and B. Perbal, A practical Guide To Molecular Cloning (1984); F.M. Ausubel et al. (eds.)。.

[0178] While several aspects have been described, those skilled in the art will recognize that various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the present disclosure. Moreover, in order to avoid unnecessarily obscuring the present disclosure, a number of well-known processes and elements have not been described. Accordingly, this specification should not be interpreted as limiting the scope of the present disclosure.

[0179] Those skilled in the art will appreciate that the presently disclosed embodiments are illustrative and do not teach by way of limitation. Accordingly, that which is contained in this specification or shown in the accompanying drawings should be construed as illustrative and not in a limiting sense. The following claims are intended to cover all general and specific features described herein, as well as all statements of the scope of methods and assemblies that may be said to fall therebetween as a matter of language.

[0180] It is intended that all matter contained in the above description and the following examples be interpreted as illustrative and not in a limiting sense, as various modifications can be made to the materials and methods described above without departing from the scope of the invention. [Example]

[0181] It should be appreciated by those of skill in the art that the techniques disclosed in the examples which follow represent techniques discovered by the inventors to function well in the practice of the invention. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the invention.

[0182] Through extensive testing and probe selection for challenging samples, this example provides a method for molecular subtyping of tumors from archival tissue. This example shows the results of hierarchical clustering analysis (HCA) on 31 specimens obtained from FFPE archival breast tissue blocks after processing them using the disclosed method. The tissue blocks were obtained from UAMS (20-year-old samples) and SWOG (30-year-old samples) along with two cell lines (MCF, MDA). The specimens included 11 TN (triple-negative), 8 luminal A (Lum A), 3 luminal B (Lum B), 6 normal lymph node (LN), 2 Her2-enriched (Her2), and 1 atypical ductal hyperplasia. The MCF sample was derived from the MCF-7 cell line (Lum B), and the MDA sample was derived from the MDA-MB-453 cell line (TN with weak Her2 expression). The probes used in this approach are listed in Table 1. [Table 1]

[0183] method. A. Protocol: Nucleic Acid Extraction from Archival Specimens Listed below are materials and equipment that can be used to practice the methods provided herein. 1. Specimen Tracking Worksheet 2. Tissue microarray (TMA) stylet-equipped needle punch 3. Razor blades, standard size 4. 2.0 mL screw-top microtubes (Sarstedt 72.694.006) 5. Low-bind Microtubes 6. 2 mL collection tube (Qiagen) 7. Light mineral oil NF (such as Geritrex brand from pharmacies) 8. Xylene (Sigma 247642) 9. Thermomixer (Eppendorf) 10. Microcentrifuge (Eppendorf) 11. Nanodrop Spectrophotometer 12. 100% EtOH 13. Nuclease-free water (Ambion AM9937) 14. ATL Buffer (Qiagen Cat #19076) 15. Proteinase K Solution* (Qiagen Cat #19133) 16. RNAseA (Qiagen Mat #1007885) 17. QIAQuick (purple) column (Qiagen #1018215) 18. PM Buffer (Qiagen Mat #1018139) 19. PE Buffer (Qiagen Mat #1015207) 20. Washing and Concentration Kit (Zymo Research) 21. RNeasy FFPE Kit* (Qiagen Cat #73504) 22. Agilent Bioanalyzer 2100 System or equivalent (e.g., Agilent Fragment Analyzer) 23. Agilent RNA 6000 Nano Kit (Cat #5067-1511) 24. LabNet Mini Incubator with Tube Rotisserie *Note: General guidelines for reagents and columns. Expiration dates are typically one year from the date of delivery to the lab, with the following exceptions: RNeasy FFPE Kit (9 months), Proteinase K Solution and RNAse A (2 years), and Light Mineral Oil NF (stamped on the bottle by the manufacturer). We recommend recording expiration dates on all component bottles / bags and checking the expiration date before each use.

[0184] The methods disclosed below can be used for pre-dissolving transfer of tissue into a tube. Part 1: Transferring pre-lysed tissue to a tube FFPE blocks 1. Remove one to three 0.6-1.0 mm needle core punches from the FFPE block and place them into labeled 2.0 mL screw-cap tubes. 2. Review the information on the Specimen Tracking Worksheet and verify that the tube label matches the Specimen Tracking Worksheet. Important: File the FFPE blocks in the appropriate location. 3. Go to section "Part 2" of the protocol. Unstained tissue section slides (without coverslips) 1. Apply 10 μL of light mineral oil to the tissue and wet it. Scrape the tissue off the slide with a razor blade and place it in a 2.0 mL screw-cap tube. This is an essential step. 2. IMPORTANT: Save the slides for future reference / use. Do not throw away the slides. 3. Proceed to step 2 of the "FFPE Block" protocol.

[0185] For stained coverslip slides containing cytology smears, the following method can be used. Stained slides with coverslips (including cytology smears) 1. Ensure that the slides are labeled with a pencil (any stickers or ink will be removed during this process). 2. In a fume hood, soak the slides in xylene overnight (or even better, over the weekend). Note: Use real xylene, not a xylene substitute. 3. Once loosened, carefully slide the coverslip off. Typically, the tissue will adhere to the glass slide, not the coverslip. 4. Place the uncovered slides in xylene for an additional 20 minutes to remove any remaining adhesive / mounting medium. Move the slide up and down 10 times. Note: Use genuine xylene, not a xylene substitute. 5. Immerse the slides in 100% EtOH for 10 minutes. 6. Allow the slides to air dry completely (approximately 15 minutes at room temperature). 7. To facilitate slide scraping, rehydrate the tissue by pipetting 10 μL of light mineral oil onto the slide. 8. Scrape the tissue from the slide into an appropriately labeled 2.0 mL microcentrifuge tube. 9. IMPORTANT: Save the slide for future reference / use. Do not throw away the scraped slide. 10. Proceed to step 2 of the "Part 1 FFPE Blocks" protocol.

[0186] The methods disclosed below can be used for partial lysis of tissue and RNA extraction. Part 2: Partial tissue lysis and RNA extraction RNA extraction using the Qiagen RNeasy FFPE kit 1. Add 0.8 mL of mineral oil to the sample tube. 2. Add 360 μL / 40 μL of Buffer PKD / Qiagen proteinase K mixture. 3. Incubate at 65°C in a heated rotisserie for 45 minutes (stat mode) to 16 hours (standard mode). 4. In stat mode, incubate at 80°C for 15 minutes. 5. Quickspin. 6. Add 25 μL of Qiagen proteinase K to the lower phase. 7. Incubate on a rotisserie at 65°C for 30 minutes. (Thaw DNase I on ice.) 8. Centrifuge at maximum speed for 15 minutes (for each sample setup / label: one 2 mL tube, one Qiagen column, five collection tubes for steps 21-27, one 1.5 mL tube for elution). 9. Transfer 250 μL of the aqueous phase (from the ~400 μL stock solution, leaving ~150 μL) to a new 2.0 mL labeled microcentrifuge tube, being careful not to disturb the pellet or aspirate the mineral oil. Save the remaining lysate tube for later DNA extraction. 10. Add 25 μL of DNase Booster Buffer and 10 μL of DNase I solution to the 250 μL portion and mix by inverting the tube. 11. Centrifuge briefly to remove any liquid remaining on the sides of the tube. 12. Incubate at room temperature for 15 minutes. Note: During this 15 minute incubation, continue incubating the lysate remaining from step 12 above for DNA extraction: see part 3 below. 13. Quickspin. 14. Add 500 μL of Buffer RBC and vortex to mix. 15. Quickspin. 16. Add 1200 μL of 100% EtOH, mix well by pipetting, and proceed immediately to the next step. 17. Transfer 700 µL of sample to a labeled RNeasy MinElute spin column and centrifuge at maximum speed for 15 seconds. 18. Set the flow-through in the collection tube aside and place the column in a new collection tube. 19. Repeat steps 20-21 until the entire sample has passed through the column. 20. Add 500 μL of Buffer RPE to the column and centrifuge at maximum speed for 15 seconds. 21. Set the flow-through in the collection tube aside and place the column in a new collection tube. 22. Add 500 μL of Buffer RPE to the column and centrifuge at maximum speed for 1 minute. 23. Set the flow-through in the collection tube aside and place the column in a new collection tube. 24. Open the spin column lid and centrifuge at maximum speed for 5 minutes. 25. Discard the collection tube and place the column in a new 1.5mL Low Bind microcentrifuge elution tube. 26. Add 30 μL of RNase-free water directly to the column membrane and incubate at room temperature for 1 minute. 27. Centrifuge at full speed for 1 minute to elute the RNA. 28. Quantify RNA using a Nanodrop spectrophotometer, 260 / 280, 260 / 230 (record ng / µL). 29. Assess RNA integrity using the Agilent RNA 6000 Nano Kit.

[0187] The methods disclosed below can be used to complete tissue lysis and DNA extraction. Part 3: Completion of tissue lysis and DNA extraction 1. Add 150 μL of tissue lysis solution to the remaining lysis solution not used for RNA extraction to make a total volume of 300 μL (tissue lysis solution = 120 μL ATL + 30 μL proteinase K [pro-K]). 2. Ensure the tube label is intact. 3. Cook in a rotisserie at 65°C for an additional 1-16 hours to ensure complete digestion of the tissue. 4. Add 25 μL of pro-K and repeat step 3 if any unlysed tissue remains. 5. Remove the tube from the rotisserie and give it a quick spin. 6. If performing an assay requiring RNA-free DNA: add 1.5 µL of RNase A, vortex for 3-5 seconds, quick spin, and incubate at RT for 5 minutes. 7. Add 490 μL of Binding Buffer PM and 10 μL of 3 M sodium acetate to a purple Qiaquick spin column (PQC), followed by 150 μL of aqueous lysate, pipetting up and down five times (thus leaving 150 μL of unused lysate for storage, direct bisulfite conversion, etc.). Note: Do not pipette out the organic phase. Ensure the tube label is intact. Store the remaining tube of lysate indefinitely. 8. Double Bind: Centrifuge the PQC at 2,000 rpm for 90 seconds and collect the flow-through. Note: Not all of the sample will pass through the filter at this point. Repeat step 7 if necessary. 9. The flow-through will contain unbound DNA. Reapply the flow-through to the same PQC and spin at 2,000 rpm for 90 seconds. Set the resulting flow-through aside and store until recovery of the eluted DNA is confirmed. 10. Add 700 μL of Buffer PE and centrifuge at 1,000 rpm for 15 seconds. Replace the collection tube. 11. Add 700 μL of 80% EtOH and centrifuge at maximum rpm for 60 seconds. 12. Replace the collection tube and centrifuge at maximum speed for 5 minutes. 13. Discard the collection tube. There will be ~1 μL of EtOH remaining on the side of the PQC column; remove it with a pipette. Next, remove the cap from the PQC and incubate it in a 65°C heat block for 5 minutes to evaporate the residual EtOH. During this incubation, label the low BIND DNA TUBES for DNA elution. 14. Preheat the AE in a heat block / chamber to 65°C. 15. Add 60 μL of AE elution buffer preheated to 65°C directly to the column filter and incubate in a 65°C heat block for 1 minute. 16. Centrifuge at maximum speed for 1 minute. 17. Double Elution: PQC contains uneluted DNA. Add the first elution mixture to the same PQC and centrifuge at full speed for 1 minute. 18. Check the sample label on the tube. 19. DNA eluted with Nanodrop: Record A260 / 280, A260 / 230, and ng / μL on the Sample Tracking Worksheet. 20. If the eluted DNA needs to be further concentrated and / or purified, proceed with the Zymo DNA Clean and Concentrate step. Note: Purification / concentration of FFPE DNA to remove impurities, including melanin, can be performed using the "Clean and Concentrate" kit. 21. Qubit TM Record the eluted DNA in ng / µL on the tracking worksheet. NOTE: Because DNA is stable in the lysis solution, lysis on a heated rotisserie can be extended over the weekend if necessary.

[0188] The following method can be used to clean (optionally) and concentrate the DNA. Optional DNA cleanup and concentration 1. Add 7 volumes of DNA Binding Buffer to the DNA sample. 2. Double bind: Load the mixture onto a Zymo-Spin column and centrifuge at full speed for 30 seconds. Reapply the flow-through and centrifuge at full speed for 30 seconds. 1. Place the column in a new collection tube and set the flow-through aside. 2. Add 200 μL of wash buffer and centrifuge at maximum speed for 30 seconds. 3. Discard the flow-through and repeat the wash step. 4. Discard the flow-through and centrifuge for an additional 30 seconds. 5. Place the column in a new collection tube. 6. Add 10 μL of dH2O (preheated to 65°C) and incubate in a thermomixer at 65°C for 1 minute. Note: Elution volume can be adjusted based on expected yield. 7. Centrifuge at maximum speed for 1 minute. 8. Double elution: Reapply the eluate to the column, incubate in a thermomixer at 65°C for 1 minute, and then centrifuge at maximum speed for 1 minute. 9. Re-quantify the DNA using Nanodrop.

[0189] Provided below is a ddPCR-based method for molecular target quantification in archived samples.

[0190] B. Protocol: ddPCR-Based Molecular Target Quantitation Method for Archival Samples The materials, samples, and controls used for ddPCR-based molecular target quantification in archival samples are listed below. 1. Samples and Controls 1) GeneCopoeia ORF cDNA clone (discontinued use, all controls transferred to cell lines (MCF7 & MB-231) i. Oncotype estrogen genes: ESR1 (A0322), PGR (A1694), BCL2 (H3307), SCUBE2 (T8649) ii. Proliferation genes: AURKA (Z0617), BIRC5 (A3492), CCNB1 (B0252), MYBL2 (B0073), TK1 (A8372) iii.Her2 gene: ERBB2 / HER2 (Z2866) iv. Reference genes: B2 M (I0035), CALM2 (Z0597), PUM1 (E0087)) 2) TissueScan, Breast Cancer cDNA Array I (OriGene, BCRT101) i. C11 (ER+PR+HER2+w) ii. C08 (triple negative) iii. F11 (ER+PG+HER2-) 3) Breast cancer cell lines i.MDA-MB-231 (ATCC, CRM-HTB-26 D), sample T25 ii. MCF-7 (ATCC, HTB-22), sample T24 4) 20-year-old UAMS breast cancer samples (13 samples) from which RNA was extracted using the custom protocol "Protocol: Nucleic Acid Extraction from Archival Specimens" disclosed in Section II. i. T1 - T13 5) S8897-extracted RNA (23 samples). RNA was extracted using the custom protocol "Protocol: Nucleic Acid Extraction from Archived Specimens" disclosed in Section II. i. Normal lymph nodes (13 samples): LN1-LN13 ii. Tumor (10 samples): T14 - T23

[0191] Below are listed the reagents and accessories required for the described method. 2. Reagents & Accessories cell culture 1) DMEM / F12 (Gibco, Cat #10565-108) 2) RPMI 1640 (Gibco, Cat #A10494-01) 3) Heat-inactivated fetal bovine serum, HI FBS (Gibco, Cat #10082-147) 4) DPBS, calcium-free, magnesium-free (Gibco, Cat. #14190-144) 5) Penicillin-Streptomycin Solution, 100x (Corning, Cat #30-002-CI) 6) Trypsin EDTA 1X, 0.05% Trypsin 0.53 mM EDTA (Corning, Cat #25-052-CI) Nucleic acid extraction and reverse transcription 7) Protocol: Nucleic acid extraction from archived specimens, custom protocol as disclosed in Section II. Reverse transcription 8) SuperScript VILO cDNA Synthesis Kit (Invitrogen, #11754-050) Preamplification by ddPCR 9) 20x TaqMan PreAmp Master Mix (ThermoFisher Scientific, Cat #4391128) 10) 2x SsoAdvanced PreAmp Supermix (Bio-Rad, Cat #172-5160) 11) Prime PCR Preamplifier Assay for Target Genes (BioRad) 12) 2x ddPCR Supermix for Probes, no dUTP (Bio-Rad, Cat #186-3023) 13) Low-concentration TE buffer or nuclease-free water (TEKNOVA, Cat #T0221) 14) TaqMan Gene Expression Assays (20X) for Target and Reference Genes [Table 2] [Table 3] 15) ddPCR droplet generation (DG) oil for probes (Bio-Rad, Cat #186-3005) 16) Probe ddPCR buffer control (Bio-Rad, Cat #186-3052) 17) DG8 Droplet Generator Cartridge (Bio-Rad, Cat #186-4008) 18) DG8 Gasket (Bio-Rad, Catalog No. 186-3009) 19) DG8 Droplet Generator Cartridge Holder (Bio-Rad, Cat #186-3051) 20) DNA LoBind Tube, 1.5ml (Eppendorf, Cat. #022431021) 21) 96-well plate (Eppendorf, Cat #951020303) 22) PCR Plate Heat Seal, Foil, Pierceable (Bio-Rad, Cat #181-4040) The equipment and software required for the described method are listed below.

[0192] 3. Equipment 1) CO2 incubator (ThermoFisher, Cat #51030-403) 2) Centrifuge 5804 R (Eppendorf, Cat # 022628146) 3) Microcentrifuge (Eppendorf, Cat #022620623) 4) QX200 Droplet Generator (Bio-Rad, Cat #186-4002) 5) QX200 droplet reader (Bio-Rad, Cat #186-4003) 6) T100 Thermal Cycler (Bio-Rad, Cat #186-1096) 7) PX1 PCR Plate Sealer (Bio-Rad, Cat #181-4000) 8) Nanodroplet 2000 Spectrophotometer (ThermoFisher Scientific, Cat #ND-2000) 9) Mini-centrifuge (Fisher Scientific, S67601B) Software: QuantaSoft Analysis Pro (version 1.0.596)

[0193] 4. Procedure The sample preparation procedure is shown below. 1. Sample preparation, controls, and FFPE samples 1) ORF cDNA clones were purchased from GeneCopoeia. 2) Breast cancer cDNA array was purchased from OriGene. 3) Breast cancer cell lines i. MDA-MB-231 and MCF-7 breast cancer cell lines were obtained from the American Type Culture Collection (ATCC) and maintained in DMEM / F12 and RPMI 1640 supplemented with 10% FBS and 1% penicillin / streptomycin, respectively. ii. RNA was extracted using the Quick DNA / RNA MiniPrep Plus kit (Zymo Research, D7003) according to the manufacturer's instructions. a. Harvesting: MDA-MB-231 and MCF-7 were harvested in DNase / RNase-free microcentrifuge tubes with PBS at a cell concentration of 0.5 x 106. b. Dissolution and purification. a) Pellet the cells by centrifugation at 200 xg for 3 minutes and suspend the cell pellet in DNA / RNA Lysis buffer. b) Transfer the lysed sample to a spin-away filter in a collection tube and centrifuge under standard conditions of 10,000 xg at room temperature for 30 seconds. Save the flow-through from the collection tube for RNA purification. c) Add an equal volume of 100% ethanol to the flow tube and mix thoroughly. d) Transfer the sample to a Zymo-Spin IIICG Column in a collection tube and centrifuge at 10,000 xg for 30 seconds at room temperature. Discard the flow-through. c. Washing and elution. a) Add 400 μl of DNA / RNA Prep Buffer to the column and centrifuge. Discard the flow-through. b) Add 700 μl of DNA / RNA Wash Buffer to the column and centrifuge. Discard the flow-through. c) Add 400 μl of DNA / RNA Wash Buffer to the column and centrifuge at 10,000 xg for 2 minutes to remove any remaining buffer. Carefully transfer the column to a new microcentrifuge tube. d) Add 50 μl of DNase / RNase-free water to the matrix in the center of the column, let it stand for 1 minute, then centrifuge to collect the RNA in a microtube. Store the RNA at -80°C until use. d. The purity and concentration of the extracted RNA were measured using nanodroplet 2000. iii. cDNA was synthesized using the SuperScript VILO cDNA Synthesis Kit (Invitrogen, #11754-050) according to the manufacturer's instructions. a. One microgram of RNA sample was subjected to reverse transcription as shown in Table 3, and the components were mixed well on ice. [Table 4] b. The samples were placed in a thermocycler and incubated using the following protocol: 25°C for 10 minutes, 42°C for 60 minutes, and 85°C for 5 minutes. c. The purity and concentration of the synthesized cDNA were measured using nanodroplet 2000. The A260 / A280 ratio was considered to be between 1.7 and 1.8 as "ideal DNA state," and between 1.8 and 2.0 as "pure DNA state." The cDNA was stored at -80°C until use. 4) UAMS breast cancer FFPE block (20 years ago) and SWOG S8897 breast cancer and normal LN slide (~30 years ago) i. Nucleic acid extraction was performed on the 20-year-old UAMS breast cancer FFPE block and the matched normal lymph node sample slide (~30 years old) using the custom protocol described in Section II, "Protocol: Nucleic Acid Extraction from Archival Specimens." Note: The 20-year-old UAMS breast cancer FFPE block was sectioned at 5-6 µm thickness and placed on a glass slide for nucleic acid extraction. The 30-year-old S8897 FFPE sample was pre-cut and mounted on a glass slide. ii. cDNA was synthesized using the SuperScript VILO cDNA Synthesis Kit (Invitrogen, #11754-050).

[0194] 2. Preparation of pre-amplification reaction cDNA products Below are described the preparation methods for preamplification reactions using SsoAdvanced (Method 1) and TaqMan (Method 2). Method 1: SsoAdvanced PreAmp Supermix (Bio-Rad) 1) Thaw SsoAdvanced PreAmp Supermix at room temperature. Mix thoroughly and centrifuge briefly to collect the solution at the bottom of the tube. Store on ice. 2) To create the preamplification assay pool, add 5 μl of PrimePCR PreAmp Assay for each target gene (final concentration: 0.01x) and nuclear-free water to a microcentrifuge tube to a total volume of 50 μl, mix thoroughly (Table 4), and centrifuge briefly. Prepare the preamplification reaction mix on ice according to the following instructions. 3) Adopt proper pipetting methods to ensure measurement precision and accuracy. [Table 5] 4) The reaction mixture was thoroughly mixed to make it homogeneous, and the sample was placed in a thermocycler and reacted as shown in Table 5. [Table 6] 5) After the run is complete, dilute the pre-amplified cDNA product 1:5 with TE buffer. The diluted pre-amplified cDNA product can be stored at -20°C for up to 12 months or at 4°C for up to 72 hours. 6) In an 8-strip PCR tube, prepare the reaction components shown in Table 6 for preamplification ddPCR using the above preamplification reaction mix and ddPCR using the original cDNA. [Table 7] 7) After capping the 8-strip PCR tubes, vortex the PCR tubes slightly and briefly centrifuge the PCR tubes to ensure the reaction contents are at the bottom of the wells. Method 2: TaqMan Preamp Supermix (Thermo Fisher Scientific) 1) Thaw the TaqMan PreAmp Master Mix at room temperature. Mix thoroughly and centrifuge briefly to collect the solution at the bottom of the tube. Store on ice. 2) To create a TaqMan assay pool, place equal volumes of 12 target genes and 3 standard reference volumes of TaqMan gene expression assays in a microcentrifuge tube and dilute each assay to a final concentration of 0.2x with 1x TE buffer. Store at 4°C for 30 days or at -20°C for 1 year. 3) Prepare the preamplification reaction mix on ice according to Table 7. Proper pipetting techniques must be used to ensure assay precision and accuracy. [Table 8] 4) The reaction mixture was mixed thoroughly to make it homogeneous, and the sample was placed in a thermocycler and reacted as shown in Table 8. [Table 9] After the run is complete, the pretreated cDNA product can be stored at -20°C for 7 days. 5) Prepare the preamplification ddPCR components using the preamplification reaction mix and original cDNA in an 8-strip PCR tube as shown in Table 9: [Table 10] 6) After capping the 8-strip PCR tubes, vortex the PCR tubes slightly and centrifuge them briefly to ensure the reaction mixture is at the bottom of the wells.

[0195] 3. ddPCR reaction droplet generation and PCR The procedure for droplet generation and ddPCR reaction is shown below. 7) After placing the DG8 droplet generator cartridge in the cartridge holder, transfer 20 μl of ddPCR reaction mix or preamplified ddPCR mix into the center well of the cartridge designated as the sample. To prevent droplet generation, do not pipette air bubbles into the well. 8) Dispense 70 μl of droplet generator oil into the well of the cartridge designated “oil (Bottom).” 9) Place the DG8 gasket on top of the holder / cartridge and insert the holder / cartridge into the QX200 droplet generator to generate millions of sample droplets. If droplets are successfully generated, the well will appear slightly opaque. 10) Transfer the 40 μl droplet from the cartridge holder to a 96-well PCR plate. 11) After sealing the plate with EasyPier Thermofoil for PCR plates, run ddPCR in a thermal cycler using the manufacturer's standard protocol as shown in Table 10. [Table 11]

[0196] 4. Droplet reading and analysis The method for reading and analyzing the droplets is as follows. 1) Fix the PCR plate containing the droplets onto the plate reader holder. 2) Enter the experimental information into the template and apply the following to each well: a. Sample name, Experiment - ABS (Absolute Quantitation), Supermix - ddPCR Supermix for Probes (without UTP) b. Target 1: Name, Channel 1 Unknown (FAM) c. Target 2: Name, Channel 2 Unknown (VIC) 3) After entering the experimental information, start the plate run. 4) Select the dye pair (FAM / VIC) to be used and the orientation of the wells to be read (columns or rows), and the QX200 will begin reading the droplet information. 5) After saving the ddPCR file, open QuantaSoft Analysis Pro to analyze the data. 6) Open the ddPCR file in QuantaSoft Analysis Pro software (BioRad, version 1.0) and set a threshold value for each well to separate positive and negative droplets. 7) Export the analyzed data to an Excel file.

[0197] 5. Molecular probe selection for gene targets A metaheuristic strategy similar to natural selection was employed for the selection of gene target probes. This approach was combined with extensive testing on challenging samples to optimize the selection of probes for each gene target. A minimax optimization strategy was utilized for generational survival of probes against gene targets. The selection procedure aimed to minimize the amplicon size of the selected probes and maximize the sensitivity and specificity performance of the probes in a set of reference materials. The reference materials consisted of cell lines and clinical samples with known molecular phenotypes.

[0198] C. Initial processing and normalization of ddPCR data Raw droplet digital polymerase chain reaction (ddPCR) data consist of multiple gene targets, each of which is composed of multiple probes. In essence, measurements from individual probes correspond to the mRNA gene expression levels detected in a sample. As a preprocessing step, for each sample, only the highest probe values ​​for a given gene target were selected for further processing. Next, each sample dataset was (self-)normalized individually using a method such as z-score:

number

number

[0199] D. Statistical Simulation of ddPCR Data The experimental data were resampled using a nonparametric bootstrap approach to derive more robust estimates by distance geometry plots of actual experimental samples (breast cancer subtypes, lymph nodes). A series of simulated (synthetic) datasets were created and used to approximate aspects of the sampling distribution, allowing some exploration of the variability expected when repeating ddPCR experiments. Bootstrapping is an efficient way to obtain estimates of the sampling distribution of experimental data without requiring distributional assumptions or building generative models to create new datasets. Bootstrapping was performed with replacement, meaning that the same data points may appear multiple times in the resampled dataset.

[0200] A multilevel approach was employed for the nonparametric bootstrap resampling procedure. Datasets from each breast cancer subtype (luminal A, luminal B, Her2-enriched, and basal) and lymph nodes were separated along with the probes corresponding to each molecular target. Each breast cancer subtype was resampled separately. Lymph nodes were also resampled separately as a group. After the resampling procedure, each simulated sample dataset (breast cancer subtype, lymph node) contained values ​​for all probes corresponding to all molecular targets. The simulated (synthetic) sample datasets were processed in the same way as the real sample datasets and subjected to further analysis. E. Distance Geometry and Visualization Approaches Applied to ddPCR Data

[0201] Distance geometry-type analyses of all data (real and synthetic) were performed using the R Statistical Computing and Graphics framework, version 4.1.2, and RStudio version 2022.02.0. Hierarchical clustering analysis (HCA) plots were performed, aggregated by sample. Additionally, HCA plots were also generated, aggregated by sample versus molecular target. In all cases, corresponding heatmaps were generated using the heatmap.2 function in the gplots v3.1.3 library, using default ("complete") clustering and the Euclidean distance measure. Principal component analysis (PCA) values ​​were calculated using the prcomp function in the stats library (part of the R core package), and visualizations were performed using the ggplot function in the ggplot2 v3.3.6 library. Finally, dot plots of molecular target (x-axis) versus z-score-transformed data (y-axis) for each ddPCR sample were visualized using the ggplot function in the ggplot2 library. F. RNA-seq sample preprocessing and bioinformatics RNA-seq sample preparation

[0202] SMARTer Stranded Total RNA-seq kit v2 - Pico (TaKaRa, cat # 634411) was used for RNA-seq sample preparation. An Illumina HiSeq 3000 with 75bp PE was used for all RNA NGS studies.

[0203] Bioinformatics Pipeline Processing RNA-seq samples were first demultiplexed, and FastQ files were created from BCL files using bcl2fastq2 (adapter trimming was additionally performed during conversion). FastQC was used to assess the quality of the FastQ files. STAR was used to align paired-end reads from each sample to the Ensembl Human reference genome build GRCm38 (using the STAR "2-pass" method). Quality control and assessment of the resulting BAM files were performed using QualiMap and RNA-SeQC. Picard was used to add read group information. Marking of duplicate reads and sorting of alignment files were also performed using Sambamba. BAM files from each sample were first processed using StringTie, using Ensembl gene annotations as a transcriptome guide (limiting output to only annotated genes). The StringTie option to output "Ballgown-ready" files was enabled.

[0204] G. DNA Sequencing Sample Preparation and Bioinformatics Genomics sample preparation and breast cancer panel processing The QIAGEN QIAseq Human Breast Cancer Panel (DHS-001Z, 93 genes) library preparation kit was used for targeted DNA-based assays, including tumor and normal (T / N) sequences. An Illumina HiSeq 3000 with a 150-bp PE was used for DNA NGS studies. The breast cancer panel, utilizing uniform molecular identifiers (UMIs), was performed at ~2000x coverage for tumor and 600x coverage for germline. The Qiagen web portal was used for smCounter2-based bioinformatics analysis. The aforementioned pipeline generates aligned reads in BAM format and detected variants in VCF format. Quality control and evaluation of the resulting BAM files were performed using QualiMap.

[0205] Whole genome sample preparation, alignment, and copy number variation analysis Whole-genome sequencing (WGS) libraries were constructed using the New England BioLabs (NEB) NEBNext Ultra II DNA library prep kit (NEB #E7645, E7103) and initially sequenced using either a low-pass (~7–15x) or ultra-low-pass (~0.3x) approach. The ultra-low-pass approach was utilized later in the study after the ichorCNA v0.2.0 analysis method for copy number analysis (CNA) was established at ~0.3x coverage for T / N. Regardless of sequencing depth, all samples were analyzed using ichorCNA.

[0206] After NGS, DNA samples were first demultiplexed, and FastQ files were created from BCL files using bcl2fastq2, with additional adapter trimming performed during conversion. FastQC was used to assess the quality of the FastQ files. FastQ paired-end files for each sample were aligned to the Ensembl Homo Sapiens reference genome (build GRCh37.75) using BWA v0.7.12. Quality control and assessment of BAM files were performed with QualiMap. BAM files were post-processed to mark duplicates and sort aligned reads via Sambamba. Copy number data were computationally estimated using the R library ichorCNA v0.2.0.

[0207] extraction analysis All study samples underwent gDNA and RNA extraction analysis using i) nanodroplets, ii) Qubit, iii) Fragment Analysis, and iv) functional assays assessing FFPE DNA quality through amplification metrics (Qiagen-QIAseq DNA QuantiMIZE or Roche-KAPPA NGS FFPE DNA QC).

[0208] Nanodrop's findings show that nucleic acids and proteins have absorbance maxima at 260 nm and 280 nm, and the absorbance ratio at these wavelengths has been used as a measure of purity. Generally, DNA purity is considered to be around 1.8, and RNA purity is around 2.0. The absorbance at 230 nm is considered to represent "other contaminants." While the purity ratio and spectral profile are important indicators of sample purity, the best indicator of DNA or RNA quality is functionality in downstream applications, such as amplification ability.

[0209] Regarding functional assay testing, for T1-T13 specimens (UAMS, 20 years ago, FFPE blocks), QuantiMIZE QC calls were high in 6 / 13 specimens and low in 7 / 13 specimens. For S8897 specimens (T14-T33), 8 specimens were high quality, 3 specimens were low quality, and 7 specimens were intermediate quality. Regarding fragmentation analysis, RNA (total RNA) showed extensive fragmentation, with the majority of fragments in the 10-40 nt size range. gDNA did not show extensive fragmentation, and in general, yields per case were found to be sufficient for NGS library construction for both tumor and LN.

[0210] Example 1: Molecular profiling of archival tissue specimens Figure 1A shows the S8897 specimen archive used in this study. The S8897 specimen archive comprises the largest number of patient specimens who did not receive treatment after surgery for small breast tumors. It is estimated that most patient specimens contain 4–6 glass slides of unstained formalin-fixed, paraffin-embedded (FFPE) material dating back more than 20–30 years. The criteria for the low-risk group are as follows: patients with a T less than 1 cm and no HR status or S-phase flow cytometry assessment (also known as the initial low-risk group). The specimens include glass slides of unstained FFPE material. Figure 1B shows a flow diagram outlining the archival tissue specimens and their molecular profiling history. This diagram outlines the overarching stages of this study, with success or failure, and identifies key figures, tables, and supplementary materials relevant to each study segment. All nucleic acid extractions from archival specimens ("A" in Figure 1B) were performed for both DNA and RNA according to the disclosed custom protocols (Tables 11–22, and Figures 2A–2D). Due to the difficulty of the fragmentation, several attempts were made to profile RNA (Figure 1B, B–E). The first attempt (Figure 1B, "B") was on specimens from the S8897 cohort (i.e., 30-year-old specimens sectioned on glass slides). The eight specimens with the highest RNA abundance were analyzed using two standard methods: i) Illumina RNA Access and a custom NanoString nCounter assay. After this external attempt failed, we performed molecular profiling on all of the specimens. [Table 12] [Table 13] *Tissue depleted. **Not a tumor [Table 14] *Tissue depleted. **LN tissue-free [Table 15] [Table 16] *Tissue depleted. **Not a tumor [Table 17] *Tissue depleted. **Not LN tissue [Table 18] [Table 19] *Tissue depleted. **Not a tumor [Table 20] *Tissue depleted. **Not LN tissue [Table 21] [Table 22] *Tissue depleted. **Not a tumor [Table 23] *Tissue depleted. **Not LN tissue

[0211] Bulk whole-transcriptome RNA-seq was performed successfully on the UAMS cohort (i.e., 20-year-old specimens from FFPE blocks) (Figure 1B, "C"). Details of these results are reported in Table 23A. This same RNA-seq approach (Figure 1B, "C") was also attempted on the available S8897 specimen but failed (Figure 1B, "D"). In this case, RNA library construction and sequencing were successful, but the insert size was found to be too small to map with the STAR aligner and Kallisto pseudoaligner. These mapping attempts also included technical recommendations from personnel contacts from the tool developers. [Table 24] [Table 25]

[0212] All RNA samples were then subjected to a custom-developed ddPCR-based RT-qPCR molecular target quantification and subtype calling approach (Figure 1B, "E"). This was successful for all samples, including 25 breast tumor samples (T1-T25) and 13 normal lymph node (LN) samples. Details of these results are reported in Table 23B. Excess LN tissue (final sample IDs, LN1-LN13) was used for ddPCR assay development. Data analysis was performed after normalization of the ddPCR data. [Table 26] [Table 27] [Table 28] [Table 29] [Table 30] [Table 31] * indicates assay not run. Abbreviations: Inv: invasive; NA: not applicable; N: negative; NL: normal; P: positive The following specimens (tumor blocks) are from the same patient: T1-4; T5-6; T10-13

[0213] Example 2: Visualization of breast tumor subtypes A custom assay visualization technique was developed for RNA tumor specimens and utilized for breast cancer tumor subtyping (Figures 3A–3D, 4A–4U, Table 24). Figures 3A–3D show four breast tumor subtype complexes. In each case, the x-axis represents molecular targets, and the y-axis represents z-score transformations of ddPCR counts. Hormonal targets are colored blue and are represented by ESR1, PGR, BCL2, and SCUB2. HER2 targets are colored orange and are represented by HER2 and GRB7. Proliferative targets are colored green and are represented by MKI67, AURKA, BIRC5, CCNB1, MYBL2, and TK1. Molecular targets are considered elevated if their z-score is ≥1. The luminal A (Lum A) subtype ("A" in Figure 1B) is defined by the presence of only one or more elevated hormonal targets. In this case, specimen T15 was classified as Lum A. An example of the Luminal B (Lum B) subtype ("B" in Figure 1B) is defined by the elevation of one or more hormone targets and proliferation targets. In this case, specimen T19 was classified as Lum B. The HER2 subtype ("C" in Figure 1B) is defined by the elevation of only one or more HER2 targets; specimen T23 was classified as HER2. An example of the triple-negative (TN) subtype ("D" in Figure 1B) is defined by the elevation of only one or more proliferation targets; specimen T21 was classified as TN. The assay visualization method was applied to tumor samples including UAS 20-year-old breast tumors (sectioned from FFPE blocks, specimen IDs T1-T13), S8897 breast tumors (20-30 years old, cut onto glass slides, specimen IDs T14-T23), and breast cancer cell lines (specimen IDs T24 (MCF-7 cell line) and T25 (MDA-MB-231 cell line)) as positive controls to determine breast cancer tumor subtypes as shown in Table 24. [Table 32]

[0214] Figure 3E further shows RNA-seq count data for 11 gene targets along with molecular phenotype call data from the PAM50, scmod2, and ddPCR assays. The RNA-seq, PAM50, and scmod2 assays were all performed during the same time period using adjacent tissue resections. The ddPCR assay was performed at a later date using tissue resections from different parts of the FFPE block or from a different block than the case. Due to tissue heterogeneity, assays using tissue from different areas of the tissue block may introduce technical and biological variability within the tumor tissue. The Sample Name column contains 13 breast cancer samples from the 20-year-old UAS cohort. The next four columns (ESR1, PGR, BCL2, and SCUBE2) in blue are gene symbols for estrogen-related molecular markers. The next two columns (ERBB2, GRB7 [orange]) are gene symbols for the Her2-enriched column, and the next five columns (MKI67, AURKA, BIRC5, CCNB1, MYBL2 [green]) are gene symbols for proliferation markers. The last three columns are analysis results from the PAM50, scmod2, and ddPCR assays. The integer values ​​listed under estrogen, Her2-enriched markers, and proliferation markers are from unnormalized RNA-seq count data; all samples were run on the same RNA-seq assay.

[0215] Example 3: Distance Geometry Analysis Distance geometry studies (Figures 5A–5F) were performed to validate unsupervised clustering by breast tumor subtype (i.e., Luminal A, Luminal B, HER2, TN) and, in some cases, tissue-based clustering to determine whether the ddPCR assay distinguished between LN tissue and breast tumors. These analyses also included targeting the molecular classes (hormonal, HER2, and proliferation) used to subtype breast tumors. Results of three distance geometry analyses were reported for T1–T25 specimens. Figure 5A shows specimen-based HCA of 25 breast cancer specimens, demonstrating good grouping of the luminal and triple-negative (TN) cohorts. Figure 5B shows the same 25 specimens after PCA, again demonstrating good separation of the luminal specimens (blue circles) and triple-negative specimens (red circles). Figure 5C shows the HCA of the same 25 breast cancer specimens on the X-axis and the 12 molecular markers used in the ddPCR assay (for molecular subtyping of breast cancer) on the Y-axis. This analysis also showed good discrimination between Luminal-based and triple-negative (TN) samples.

[0216] Figures 5D–5F show the results of processing and analysis by ddPCR assay for 25 breast tumor specimens (T1–T25) and 13 normal human lymph node specimens (LN1–LN13) from the S8897 cohort. Similar to the S8897 breast cancer specimens, these lymph node specimens were also unstained FFPE specimens mounted on glass microscope slides and stored without special processing or storage conditions for over 30 years. Analysis of 38 specimens (T1–T25, LN1–LN13) using distance geometric methods, such as HCA and PCA, was performed as follows. Figure 5D displays specimen-based HCA of these 38 specimens, demonstrating good grouping of the following cohorts: i) TN (triple-negative), ii) luminal, and iii) a cohort consisting of a portion of luminal A and B specimens and 12 / 13 lymph nodes. Figure 5E shows the same 38 samples after PCA, again showing good separation of luminal (blue circles) vs. TN (red) vs. lymph node (LN, yellow) samples, with fewer luminal A and B samples. Figure 5F shows the 38 HCA samples on the x-axis and the 12 molecular markers on the y-axis. This analysis again showed good discrimination between TN breast cancer samples (red circles) and luminal samples (blue), as well as lymph node samples with scattered luminal A and B (yellow).

[0217] Statistical simulations were performed to derive more robust estimates through additional distance geometry studies of actual experimental samples (Figures 6A-6F). Figures 6A-6C show the results of statistical simulations using distance geometry methods (HCA, PCA) with nonparametric bootstrap to create 500 synthetic samples from ddPCR assay data for each PAM50 subtype and analyze tumor samples (T1-T25). Figure 6A displays the sample-based HCA of actual breast cancer samples (T1-T25) and synthetic breast cancer samples, demonstrating good grouping and discrimination across all PAM50 cohorts (Lum A, Lum B, Her2, and TN). Figure 6B shows the results of PCA including 25 actual breast cancer samples and 500 synthetic samples for each PAM50 subtype, demonstrating good separation of PAM50 subtypes. Figure 6C shows the HCA of 25 real breast cancer samples with 500 synthetic samples per PAM50 subtype on the x-axis and the 12 molecular markers used in the ddPCR assay (for breast cancer molecular subtyping) on ​​the y-axis. This analysis also showed good discrimination of the four PAM50 subtypes (Lum A, Lum B, Her2, and TN).

[0218] Figures 6D-6F show the results of distance geometric analysis (HCA, PCA) of tumor and lymph node samples (T1-T25, LN1-LN13). Nonparametric bootstrap statistical simulations were performed to generate 500 synthetic samples and 500 synthetic normal lymph node samples from ddPCR assay data for each PAM50 subtype. Figure 6D displays specimen-based HCA of 38 real and synthetic samples, demonstrating excellent grouping of all PAM50 cohorts (Lum A, Lum B, Her2, TN) and lymph nodes. Figure 6E displays the results of PCA including 38 real and 500 synthetic samples for each PAM50 subtype and lymph node, demonstrating excellent separation of lymph nodes by PAM50 subtype. Finally, Figure 6F shows the 500 synthetic samples per PAM50 subtype and lymph node on the X-axis and the 12 molecular markers used in the ddPCR assay (for breast cancer molecular subtyping) on ​​the Y-axis for the 38 actual HCA samples. This analysis also showed good discrimination between the four PAM50 subtypes (Lum A, Lum B, Her2, and TN) and normal lymph nodes.

[0219] Example 4: DNA samples All DNA samples (Figure 1B, "F") underwent NGS analysis using a MyChip normal LN, if available, as a germline comparison. DNA analysis was performed on tumors only if LN tissue was unavailable. Tumor samples from T1–T13 (UAMS cohort) and T32 (S8897 cohort) did not have LNs. For T1–T13 specimens, a total of 13 tumor block specimens were available, but this tissue was obtained from six unique patients. Specifically, T1–T4 consisted of four tumor blocks from the same patient, T5 and T6 consisted of two tumor blocks from unique patients, and T10–T13 consisted of four tumor blocks from unique patients. The large amount of tissue provided sufficient material for assay development, including sensitivity, accuracy, reproducibility, cross-correlation with IHC, and comparison of RNA-seq and ddPCR RT-qPCR. Because high genetic diversity is established within a single solid tumor, the use of these specific specimen types and assay methods was also useful for assessing tumor heterogeneity.

[0220] For DNA NGS analysis, library preparation and NGS were successful, but subclonal heterogeneity resulted in variable findings, possibly related to fixation artifacts, potentially resulting in a high false positive rate (FPR). To mitigate these findings, we modified our approach by targeting a 93-gene panel utilizing unique molecular identifiers (UMIs) to reduce false positives, which was successful, as shown in Figure 7A, Figures 8-38, and Tables 25-53. Low-pass and, in some cases, ultra-low-pass copy number variation (CNV) WGS was also performed, with success, as shown in Figure 7A, Figures 8-38, and Tables 25-53 for T1-T33 samples.

[0221] The top row in Figure 7A is shared by the IntClust chromosome band and specimen, which are the final specimen IDs for this set of clinical and molecular results. Specimens T24 and T25 are not included because they are cell line control materials. The IntClust subtype row displays the molecular subtype of the breast tumor as determined by the IntClust classifier. The next row, the gene ddPCR subtype, displays the top 27 genes from a UMI-based gene panel of 91 TCGA breast cancer driver genes and color-codes them based on the mutation's impact on protein function: red indicates high impact, yellow indicates moderate impact, green indicates low impact, and black indicates no mutation. The ddPCR assay subtypes are displayed and color-coded as Luminal A (LumA, blue), Luminal B (LumB, green), triple-negative (TN, red), Her2 (pink), and if the specimen was primarily DCIS. Specimen T17 was not tumorous. [Table 33] [Table 34] [Table 35] Table 36 Table 37 Table 38 Table 39 Table 40 Table 41 Table 42 Table 43 Table 44 Table 45 Table 46 Table 47 Table 48 Table 49 Table 50 Table 51 Table 52 Table 53 Table 54 Table 55 Table 56 Table 57 Table 58 Table 59 Table 60 Table 61

[0222] Example 5: Clinical and genomic characteristics The clinical and genomic characteristics of all breast cancer tumor specimens are shown in Figure 7. The chromosome bands (Ch Bands) and specimen (T1–T33) identifiers used in the IntClust algorithm are listed, corresponding to the first column and entire row, respectively. Specimens T24 and T25 were RNA controls and were not included. The next row lists the IntClust subtypes, which directly map / correlate to PAM50 subtypes. Figure 39 explains the copy number mapping criteria to PAM50 subtypes. As background, the tissue specimens T1–T13 used in the IntClust subtype analysis (i.e., low-pass or ultra-low-pass WGS) were from different tumor blocks or significantly different block locations than those used in the ddPCR subtype assay. This is due to the consumption of FFPE blocks over time. Therefore, solid tissue / tumor heterogeneity is likely, which may help explain the differences in tumor subtypes reported by IntClust and ddPCR, particularly for T1 and T4 specimens. Next is the row heading "Gene ddPCR Subtype." The ddPCR molecular subtypes are listed as determined by a custom assay; as previously mentioned, for eight specimens (T26–T33), RNA was consumed before running this assay. The next side of this figure reports a matrix of findings for the top 27 genes from all specimens (T1–T33), with mutations color-coded according to their impact on protein function. Histopathology (HistoPath) was determined by a pathologist (UAMS) and is reported in the following row, color-coded. Tumor grade was also reported as determined by a pathologist at UAMS.

[0223] Example 6: Analysis of DNA samples Each DNA tumor specimen underwent four analyses; an example for specimen T6 is shown in Figure 8. Results for all specimens (T1-T33) are shown in Figures 9-38. Figure 8A details the results for copy number alterations in ichorCNAs, with the x-axis representing chromosomes and the y-axis representing the log2 ratio of copy numbers between the tumor specimen and matched normal LN (if present) or ichor pool normal (if absent), along with a color-coded copy number change. The chromosomal bands associated with IntClust and the associated copy number classification are shown in Figure 8B; specimen T6 was classified as Luminal A by IntClust 8 criteria (Figure 39).

[0224] Next, the top 27 gene targets from the gene panel are color-coded by the mutation's imprinting on protein function, as shown in Figure 8C. Mutation impact coding includes high impact on protein function (red), moderate (yellow), low (green), or no mutation (black). Finally, details of the mutations in the selected DNA panel are shown in Table 25: gene symbol, transcript, protein and mutation annotation, classification, impact, frequency, depth, and molecular tag (UMI) depth (MT depth). If available, COSMIC IDs are reported, further indicating the clinical relevance of the mutations.

[0225] Example 7: Further correlation of specimens with molecular and pathological results A correlation summary of the pathological and molecular results for the study samples is shown in Table 54. The column-based description in this table begins with the initial specimen ID because the samples are from two different cohorts, S8897 and UAMS, and include samples from breast tumors, normal LNs, and cell lines. As previously mentioned, specimens T1–T13 are tumor block specimens from six unique patients, and the initial specimen IDs are color-coded to indicate their origin / membership among the six patients. The final specimen IDs serve to standardize the specimen names for tumor specimens (T1–T33) and LN1–LN13, which are LN specimens used in ddPCR experiments and assay development. All LN specimens are from the S8897 cohort. For tumor specimens, T1–T13 were from the UAMS specimen cohort, T24 and T25 were from cancer cell lines used as ddPCR controls, and T14–T22 and T26–T33 were from the S8897 cohort. [Table 62] [Table 63] [Table 64] [Table 65] [Table 66] * indicates that the assay was not run. The following specimens (tumor blocks) are from the same patient: T1-4; T5-6; T10-13 Abbreviations: Inv, invasive; NA, not applicable; N, negative; NL, normal; P, positive

[0226] Immunohistochemistry (IHC) results are reported in Table 54 for ER, PR, and Her2. The results are listed as either positive (Pos) or negative (Neg) as interpreted by the pathologist based on the staining characteristics of specimens T1 through T13 (Figures 40 through 66 - H&E, IHC). Note: IHC for Ki-67 was attempted but was unsuccessful. An "*" in any column of Table 54 indicates that the assay was not performed. Further results from these same specimens (T1 through T13) are reported in Table 54 under i) PAM50 Subtype (RNAseq) and ii) scmod2 Subtype (RNAseq), specifying the tumor subtypes determined by these two methods. The ddPCR Subtype and IntClust Subtype columns list the breast tumor subtypes determined by these methods for specimens T1 through T33.

[0227] Comparing the concordance of molecular results for specimens T1-T13, IHC:PAM50 was 11 / 13 (~85%). IHC evaluation was incomplete due to unsuccessful Ki67 staining. IHC:scmod2 was 13 / 13 (100%), IHC:ddPCR was 11 / 13 (~85%), and IHC:IntClust was 13 / 13 (100%). PAM50:scmod2, 11 / 13 (~85%); ddPCR:PAM50, 10 / 13 (~77%); ddPCR:scmod2, 11 / 13 (~85%). Comparing the concordance of molecular results between ddPCR and IntClust for specimens T1-T23, 20 / 23 (~87%) concordance was observed. The "Comments" column provides additional notable information about the test specimen. The last five columns: i) Invasive, In Situ, ii) Tumor %, iii) Normal %, iv) Tumor Grade, v) Comments correspond to the pathologist's reading of each specimen.

[0228] overview In the disclosed example, we evaluated the feasibility of robust specimen recovery from challenging archives. From the S8897 specimen (a roughly 30-year-old FFPE specimen mounted on a glass slide and stored at room temperature without any special processing), tumor gDNA results showed a 95% (18 / 19) success rate. That is, DNA was extracted, and an NGS library was constructed, validated, and sequenced. The success rate for LN gDNA was also 95% (18 / 19). RNA results showed extensive fragmentation, but effective measurement suggested analytical validity in all cases. Due to the small insert size from the S8897 RNA specimen, a custom ddPCR RT-qPCR assay was developed for tumor subtype classification. The UAMS specimen, excised from a 20-year-old FFPE block and stored at room temperature, showed a 100% recovery rate (13 / 13) for both gDNA and RNA without any special processing, and NGS was also successful.

[0229] The disclosed examples further evaluated molecular profiling approaches for extracted DNA and RNA. Low-ion-path WGS (~15x) and, in some cases, ultra-low-ion-path (~0.3x) CNVs were achieved in 31 / 31 samples. LN tissue, when available and matched to the sample, was used as a normal control. Initially, a DNA panel using eight tumor-lymph node paired samples (44 TCGA genes) was successful, but demonstrated a significant degree of subclonal findings and a high false-positive rate (FPR), likely due to the sample age and fixation. A UMI-based repetitive DNA NGS panel was performed on 31 tumor samples with LN matched pairs (when available) for 93 TCGA breast cancer driver genes at high coverage (>2000x). The success rate for the NGS tumor panel library was 30 / 31 (~97%). This effectively resolved the high FPR, clarifying the results, many of which reported findings including COSMIC identifiers, confirming clinical relevance. Because eight S8897 RNA preps were consumed by the initial failed method, surrogate samples from UAMS showing extensive fragmentation were utilized. These samples were successfully subjected to RNA expression profiling for molecular tumor subtyping. The research team successfully analyzed intrinsic subtypes by comparing IHC with PAM50 in 11 / 13 tumors, scmod2 in 13 / 13 tumors, ddPCR in 11 / 13 tumors, and IntClust in 13 / 13 tumors. Furthermore, for 13 samples, RNA-seq count data was used to calculate OncoType gene groups classified into hormonal (ESR1, PGR, BCL2, SCUBE2), HER2 (GRB7, HER2), and proliferation (MKI67, AURKA, BIRC5, CCNB1, MYBL2) groups. PAM50, scmod2, ddPCR, and IHC were compared, but Ki67 could not be specifically analyzed. The findings of the RNA-seq data for the hormone, HER2, and proliferation groups showed high concordance with the IHC staining results for ER, PR, and HER2, respectively.

[0230] The RNA in the disclosed examples was more degraded with insert sizes less than 100 nt, and therefore required custom ddPCR assays for breast tumor subtyping. Using DNA-based IntClust subtyping and highly degraded RNA-based ddPCR assay subtyping, a strong match was obtained.

[0231] This disclosure further provides a potential role for molecular characterization of archived FFPE tissue specimens, generating a high level of evidence to facilitate patient care. For example, while the overall favorable prognosis of patients within the S8897 cohort is encouraging, the question arises as to why some patients developed distant metastases despite apparently similar clinicopathological conditions. This disclosure further demonstrates that the quality of nucleic acid extraction and characterization was satisfactory.

[0232] We provide here a thorough nucleic acid extraction with QA / QC metrics and NGS with DNA for copy number variation (CNV), mutation analysis, and high-resolution molecular profiling of RNA for breast tumor subtype classification, performed by RNA-seq for some samples and by ddPCR RT-qPCR for all samples. The correlation between tumor subtype classification by ddPCR and the IntClust assay was complementary and robust given the orthogonal approaches of using RNA and DNA in the breast tumor subtype prediction process.

[0233] Thus, the examples provided herein demonstrate that macromolecular profiling can be successfully achieved using FFPE specimens stored for 20-30 years under less-than-ideal conditions, such as seconds on glass slides. This analysis demonstrates that there are no technical impediments to proceeding with larger-scale analyses of the remaining specimens in the "good" cohort of S8897 or other tissue banks within the many clinical trial institutions that maintain tissue banks.

[0234] In summary, reliable gDNA recovery was achieved in over 90% of specimens. Low-depth whole-genome DNA sequencing was performed for copy number profiling, and a targeted DNA sequencing approach using unique molecular identifiers (UMIs) was used for mutation analysis. RNA recovery showed extensive fragmentation. Some specimens were sufficient for RNA-seq analysis. All specimens were sufficient for ddPCR-based RT-qPCR analysis, and these results were compared with a molecular subtyping method using ER, HER2, and proliferation-related markers. The expression of each of these three gene groups by molecular profiling correlated highly with the results of MycIHC in the same specimens.

[0235] equivalent product While several inventive aspects have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for obtaining the functions and / or results and / or one or more advantages described herein, and each such variation and / or modification is deemed to be within the scope of the inventive aspects described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the particular use or application for which the teachings of the present invention are / are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation equivalents to the specific inventive aspects described herein. Accordingly, the foregoing embodiments are presented by way of example only, and it will be understood that, within the scope of the appended claims and their equivalents, the inventive aspects may be practiced otherwise than as specifically described and claimed. The inventive aspects of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is within the inventive scope of the present disclosure, provided that such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0236] All definitions defined and used herein should be understood to take precedence over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0237] All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter for which each is cited, and in some cases may include the entire specification.

Claims

1. 1. A method for molecular subtyping a cancer sample obtained from a subject, the method comprising: a) Mineral oil was used untreated to facilitate solubilization of old and degraded FFPE sample material; b) digesting the sample with proteinase; c) incubating the sample digested in step a) with DNase; d) incubating the mixture of step b) in a guanidine salt-based buffer; e) enriching and isolating RNA; f) pre-proliferate; and g) Perform digital droplet PCR.

2. The method of claim 1, wherein the sample is a formalin-fixed, paraffin-embedded sample.

3. 10. The method of claim 1, wherein the sample is at least 5 years old.

4. 10. The method of claim 1, wherein the sample is digested with proteinase K.

5. 5. The method of claim 4, wherein the sample is digested at about 65°C to about 70°C for about 75 minutes.

6. 2. The method of claim 1, wherein step a) further comprises separating the sample into an aqueous phase by centrifugation and separating the aqueous phase from the residual lysate, wherein the aqueous phase is used in step b).

7. 10. The method of claim 1, wherein the RNA is concentrated and isolated using a spin column.

8. 7. The method of claim 6, wherein the residual lysate is used for DNA extraction.

9. 9. The method of claim 8, wherein the proteinase is incubated with the residual lysate at about 65°C to about 70°C for about 13 hours to about 18 hours, thereby producing a digested tissue lysate.

10. 10. The method of claim 9, wherein the digested tissue lysate is incubated with RNase.

11. The method of claim 10, further comprising concentrating and isolating the DNA using a spin column.

12. 12. The method of claim 11, further comprising the steps of pre-amplifying the isolated DNA and performing ddPCR.

13. 10. The method of any one of the preceding claims, wherein the target nucleic acid and optionally the reference nucleic acid are quantified.

14. 14. The method of claim 13, wherein the target nucleic acid or a fragment thereof encodes estrogen receptor 1 (ESR1), progesterone receptor (PGR), B-cell lymphoma 2 (BCL2), signal peptide, CUB domain and EGF-like domain containing 2 (SCUBE2), human epidermal growth factor receptor 2 (HER2), growth factor receptor-bound protein 7 (GRB7), marker of proliferation Ki-67 (MKI67), aurora kinase A (AURKA), baculoviral IAP repeat containing 5 (BIRC5), cyclin B1 (CCNB1), MYB proto-oncogene like 2 (MYBL2), thymidine kinase 1 (TK1), or a combination thereof.

15. The method of claim 1 , wherein the cancer sample is a breast cancer sample.

16. 2. The method of claim 1, wherein the subtype comprises luminal A subtype (Lum A), luminal B subtype (Lum B), HER2 subtype (HER2), or triple negative subtype (TN).

17. 17. The method of claim 1 or claim 16, wherein a sample is determined to be Lum A if the level of one or more of the target nucleic acids or fragments thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated in the sample.

18. 17. The method of claim 1 or claim 16, wherein a sample is determined to be Lum B if the level of a nucleic acid or a fragment thereof encoding ESR1, PGR, BCL2, SCUBE2, or any combination thereof is elevated, and if the level of a nucleic acid or a fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated in the sample.

19. 17. The method of claim 1 or claim 16, wherein the sample is determined to be HER2 if the level of a nucleic acid or fragment thereof encoding HER2, GRB7, or a combination thereof is elevated in the sample.

20. The method of claim 1 or claim 16, wherein the sample is determined to be TN if the level of a nucleic acid or a fragment thereof encoding MKI67, AURKA, BIRC5, CCNB1, MYBL2, TK1, or any combination thereof is elevated in the sample.

21. 15. The method of claim 13 or 14, wherein the amount of the target nucleic acid is compared to the outcome or treatment response of the subject.