Non-invasive differentiation of lung cancer histological subtypes by DELFI-derived cell-free DNA fragmentation patterns

By analyzing genome-wide cfDNA fragmentation patterns and clinical characteristics, the method non-invasively subtypes SCLC and NSCLC, facilitating personalized treatment and improving patient outcomes.

JP2026506641APending Publication Date: 2026-02-25DELFI DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025546413
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-21
Filing Date
2024-02-12
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

There is a need for non-invasive methods to identify and subtype small cell lung cancer (SCLC) and non-small cell lung cancer (NSCLC) to improve patient outcomes, as tumor biopsies are difficult and current methods are invasive.

Method used

Characterizing genome-wide patterns of cell-free DNA (cfDNA) fragmentation using low-coverage whole-genome sequencing in conjunction with clinical and demographic characteristics, analyzing transcription factor binding sites to determine cfDNA fragment coverage scores, and applying machine learning for subtyping.

Benefits of technology

Enables non-invasive subtyping of SCLC and NSCLC, allowing for personalized treatment strategies and improved patient outcomes by accurately distinguishing between subtypes like neuroendocrine-high and adenocarcinoma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026506641000001_ABST
    Figure 2026506641000001_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for improving diagnostic applications by using genome-wide patterns of plasma-derived fragmented cell-free DNA (cfDNA) derived by low-coverage whole genome sequencing. In particular, the present invention provides a new and effective method for subtyping non-small cell lung cancer in a subject as low-neuroendocrine non-small cell lung cancer or high-neuroendocrine non-small cell lung cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 445,284, filed February 13, 2023, U.S. Provisional Patent Application No. 63 / 470,101, filed May 31, 2023, and U.S. Provisional Patent Application No. 63 / 528,237, filed July 21, 2023. The disclosures of these prior applications are incorporated herein by reference in their entireties and are considered part of the disclosure of this application.

[0002] Technical Field The present invention generally relates to the non-invasive diagnosis and subtyping of small cell lung cancer (SCLC), more specifically to the analysis of genome-wide patterns of fragmented cell-free DNA (cfDNA) combined with the clinical and demographic characteristics of individual patients.The present invention also relates to cell-free DNA fragmentation profiling as a method for assessing tumor fraction and treatment monitoring in non-small cell lung cancer (NSCLC).Given the practical difficulty of performing tumor biopsy of SCLSC and NSCLC, there is a need for a new approach for feasible and rapid method of non-invasively subtyping SCLC and NSCLC. [Background technology]

[0003] Small cell lung cancer (SCLC) is an aggressive malignancy with a poor prognosis. Although SCLC is clinically managed as a single cancer type, emerging evidence supports that these subtypes (high neuroendocrine and low neuroendocrine) acquire diverse transcriptomes and epigenetic states. Furthermore, different SCLC subtypes respond differently to specific treatments, such as immunotherapy, in the case of the low neuroendocrine-high inflammatory SCLC subtype. Non-small cell lung cancer (NSCLC) is any type of epithelial lung cancer other than small cell lung cancer. The most common types of NSCLC are squamous cell carcinoma, large cell carcinoma, and adenocarcinoma, although several other less common types exist, and all types can occur in rare histological variants. As a class, NSCLC is typically less sensitive to chemotherapy and radiation therapy than small cell lung cancer. Patients with resectable disease may be cured with surgery or chemotherapy followed by surgery, as well as chemotherapy followed by surgery. Although local control can be achieved with radiation therapy in many patients with unresectable disease, cure remains a relatively minority of patients. Patients with unresectable locally advanced disease may achieve long-term survival with radiation therapy in combination with chemotherapy. Patients with progressive metastatic disease may achieve improved survival and palliation of symptoms with chemotherapy, targeted agents, and other supportive care. Novel, noninvasive methods for identifying and subtyping non-small cell and small cell lung cancer are needed to improve patient outcomes. Summary of the Invention

[0004] The present invention is based on the breakthrough discovery that the use of low-coverage whole-genome sequencing to characterize genome-wide patterns of cell-free DNA (cfDNA) fragmentation in plasma can improve cancer diagnosis when analyzed in combination with certain clinical and demographic characteristics of individual patients.

[0005] In one embodiment, the invention provides a method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of neuroendocrine-high small cell lung cancer.

[0006] In one embodiment, the invention provides a method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is between about 10× and 0.1×; mapping the sequenced fragments to the genome to obtain genomic sections of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic sections of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a neuroendocrine-high small cell lung cancer subtype.

[0007] In one embodiment, the invention provides a method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; and subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence that map to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a neuroendocrine-high small cell lung cancer subtype. In some embodiments, the methods use identified transcription factors ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

[0008] In some embodiments, the methods use machine learning to subtype small cell lung cancer in a subject.

[0009] In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

[0010] In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window.

[0011] In some embodiments, the method uses a shift of 1 in all coverage calculations to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1).

[0012] In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.

[0013] Some embodiments further include using a general additive model to smooth coverage and correct for GC bias.

[0014] In some embodiments, the genomic intervals are non-overlapping.

[0015] In some embodiments, the genomic intervals each contain thousands to millions of base pairs.

[0016] In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval.

[0017] In some embodiments, the cfDNA fragmentation profile comprises a median fragment size.

[0018] In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.

[0019] Some embodiments further comprise administering to a subject identified as having high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating said cancer type.

[0020] In some embodiments, the therapeutic agent is an immunotherapy.

[0021] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is between about 30× and 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence using the length and amount of the cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the subtype of adenocarcinoma or squamous cell carcinoma.

[0022] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is between about 10× and 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the adenocarcinoma or squamous cell carcinoma subtype.

[0023] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the adenocarcinoma or squamous cell carcinoma subtype.

[0024] In some embodiments, the identified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

[0025] In some embodiments, the method uses machine learning to subtype small cell lung cancer in a subject.

[0026] In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

[0027] In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window.

[0028] In some embodiments, all coverage calculations are shifted by 1 to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1).

[0029] In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.

[0030] Some embodiments further include using a general additive model to smooth coverage and correct for GC bias.

[0031] In some embodiments, the genomic intervals are non-overlapping.

[0032] In some embodiments, the genomic intervals each contain thousands to millions of base pairs.

[0033] In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval.

[0034] In some embodiments, the cfDNA fragmentation profile comprises a median fragment size.

[0035] In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.

[0036] Some embodiments further comprise administering to a subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating that type of cancer.

[0037] In some embodiments, the therapeutic agent is an immunotherapy.

[0038] Some embodiments further include quantifying the circulating tumor fraction, which reflects the performance of the major allele fraction (MAF) in relation to Response Evaluation Criteria in Solid Tumors (RECIST) assessment in both adenocarcinoma and squamous cell carcinoma. [Brief explanation of the drawings]

[0039] [Figure 1] 1 illustrates the DELFI scores of patients diagnosed with SCLC in the study of Example 1. [Figure 2] We illustrate that genome-wide fragmentation profiles were significantly consistent between pre-treatment, post-treatment, progression, and response time points, suggesting that genome-wide circulating tumor DNA fragment size is a powerful method for detecting changes during immunotherapy monitoring of SCLC cancer treatment. [Figure 3] Genome-wide fragmentation profiles (and corresponding DELFI scores at the bottom) showing partial differences between the two main subtypes of SCLC cases, i.e., high neuroendocrine and low neuroendocrine, are illustrated, suggesting that genome-wide circulating tumor DNA fragment size may be a powerful method to detect subtype classification during immunotherapy treatment monitoring of SCLC cancer treatment. [Figure 4] Clustering of the two subtypes of differential genes between high and low neuroendocrine SCLC cases using publicly available data (Lissa et al., 2022) and principal component analysis (DELFI 30X plasma WGS data) illustrates differential genes between these two cases. PCA (principal component analysis) analysis data revealed distinct clusters between pre-treatment samples of high and low neuroendocrine SCLC. [Figure 5] A-B illustrate that Eigencor analysis revealed significant correlations between principal components 1 and 2 and the neuroendocrine status of the samples. [Figure 6]Panels A-B illustrate the potential of the DELFI assay to subtype SCLC using TFBS and genome-wide cfDNA fragmentation analysis at ASCL1 binding sites in pretreatment NCI samples. Panel A illustrates the observation of distinct clusters of SCLC samples corresponding to different ASCL1 activation levels. Importantly, the two clusters perfectly corresponded to distinct SCLC subtypes. The primary factor driving the differences between samples was in fact the difference in fragment coverage identified at the center of the ASCL1 binding site. Panel B illustrates that other clinical and sample characteristics had little impact on the clustering of these samples. Overall, these data demonstrate how DELFI analysis can perfectly distinguish between SCLC subtypes using fragment coverage at TFBS. [Figure 7] An example patient is shown, supporting the hypothesis that the detected signal originates from tumor-infiltrating lymphocytes. Patient NCI-0422, diagnosed with SCLC of the inflammatory subtype, was treated with a combination of durvalumab and olaparib. NCI-0422 responded well to treatment but was later diagnosed with progressive disease. Variable genomic binding is detectable both pretreatment and at progression. Peaks are concentrated at positions where TFBSs and TSSs are mixed. [Figure 8] A-B illustrate ASCL1 binding sites between responders and non-responders. A shows 500 cell type-specific bins from PMD data to calculate leukocyte proportions. [Figure 9] Illustrates a machine learning model for predicting high- and low-neuroendocrine SCLC. [Figure 10] AB illustrate that the methods described herein can distinguish between neuroendocrine and non-neuroendocrine SCLCs by calculating read depth ratios at SCLC-specific genomic coordinates (A). B shows that the methods described herein appear to be able to better subtype SCLC cases. [Figure 11]The read depth ratios at SCLC-specific genomic coordinates were transformed into 2D PCA, illustrating the separation of samples within the neuroendocrine and non-neuroendocrine groups. Most samples clearly fall into distinct clusters. [Figure 12] This figure illustrates that PCA derived from read depth ratios at SCLC-specific genomic coordinates can separate pretreatment samples into specific SCLC subtypes (A, N, P, Y). The true SCLC subtypes were calculated using tissue RNA-seq differential gene expression analysis. [Figure 13] Illustrates how fragmentation profiles can suggest lung cancer status and treatment response patterns. [Figure 14] This figure shows that the model-derived DELFI-TFs exhibit a strong correlation with mutant allele frequencies across samples. [Figure 15] Figure 1 illustrates that DELFI-TF accurately quantifies circulating tumor fraction and reflects the performance of MAF in relation to RECIST assessment. [Figure 16] Illustrates altered fragmentation in tumor-derived cfDNA. [Figure 17] AC illustrate that DELFI-TF accurately detects circulating tumor fraction without being confounded by clonal hematopoiesis. [Figure 18] This demonstrates that cfDNA fragmentation patterns accurately distinguish NSCLC subtypes. [Figure 19] We present the development of the DELFI-TF model at the proof-of-concept stage. [Figure 20] 8 is an example of a computer 800 that can be used to implement the methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0040] The present invention is based on the breakthrough discovery that characterizing genome-wide patterns of cell-free DNA (cfDNA) fragmentation in plasma using low-coverage whole-genome sequencing improves cancer diagnosis when analyzed in conjunction with certain clinical and demographic characteristics of individual patients.

[0041] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of neuroendocrine-high small cell lung cancer.

[0042] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of neuroendocrine-high small cell lung cancer.

[0043]

[0003] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9x to 0.1x; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at identified transcription factor binding sites; determining the length and amount of the cfDNA fragments; analyzing the genomic intervals of mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of neuroendocrine-high small cell lung cancer.

[0044] In some embodiments, the methods use identified transcription factors ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

[0045] In some embodiments, the method uses machine learning to subtype small cell lung cancer in a subject.

[0046] In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

[0047] In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window.

[0048] In some embodiments, the method uses a shift of 1 in all coverage calculations to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1).

[0049] In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.

[0050] Some embodiments further include using a general additive model to smooth coverage and correct for GC bias.

[0051] In some embodiments, the genomic intervals are non-overlapping.

[0052] In some embodiments, the genomic intervals each contain thousands to millions of base pairs.

[0053] In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval.

[0054] In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.

[0055] Some embodiments further include administering to a subject identified as having high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating that type of cancer.

[0056] In some embodiments, the therapeutic agent is an immunotherapy.

[0057] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequencing fragments, wherein the genome coverage is between about 30× and 0.1×; and mapping the sequencing fragments to the genome to obtain genomic intervals of sequences mapped at identified transcription factor binding sites.

[0058] The method comprises: determining the length and amount of cfDNA fragments; analyzing the genomic interval of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding site using the length and amount of cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some embodiments, the method uses the identified transcription factors ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype the small cell lung cancer in the subject. In some embodiments, the subtyping is performed by calculating log2 (read depth ratio) over a 20bp window, starting from a position 2kb away from the identified transcription factor binding site. In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window. In some embodiments, the method uses a shift of 1 in all coverage calculations to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median for each window. Some embodiments further include using a general additive model to smooth the coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.Some embodiments further comprise administering to a subject identified as having a high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating said cancer type, hi some embodiments, the therapeutic agent is an immunotherapy.

[0059] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequencing fragments, wherein the genome coverage is between about 10× and 0.1×; and mapping the sequencing fragments to the genome to obtain genomic intervals of sequences mapped at identified transcription factor binding sites.

[0060] The method comprises: determining the length and amount of cfDNA fragments; analyzing the genomic interval of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding site using the length and amount of cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some embodiments, the method uses the identified transcription factors ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype the small cell lung cancer in the subject. In some embodiments, the subtyping is performed by calculating log2 (read depth ratio) over a 20bp window, starting from a position 2kb away from the identified transcription factor binding site. In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window. In some embodiments, the method uses a shift of 1 in all coverage calculations to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median for each window. Some embodiments further include using a general additive model to smooth the coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.Some embodiments further comprise administering to a subject identified as having a high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating said cancer type, hi some embodiments, the therapeutic agent is an immunotherapy.

[0061] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequencing fragments, wherein the genome coverage is between about 9× and 0.1×; and mapping the sequencing fragments to the genome to obtain genomic intervals of sequences mapped at identified transcription factor binding sites.

[0062] The method comprises: determining the length and amount of cfDNA fragments; analyzing the genomic interval of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding site using the length and amount of cfDNA fragments; and subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some embodiments, the method uses the identified transcription factors ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype the small cell lung cancer in the subject. In some embodiments, the subtyping is performed by calculating log2 (read depth ratio) over a 20bp window, starting from a position 2kb away from the identified transcription factor binding site. In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window. In some embodiments, the method uses a shift of 1 in all coverage calculations to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median for each window. Some embodiments further include using a general additive model to smooth the coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.Some embodiments further comprise administering to a subject identified as having a high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating said cancer type, hi some embodiments, the therapeutic agent is an immunotherapy.

[0063] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is between about 30× and 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence using the length and amount of the cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the subtype of adenocarcinoma or squamous cell carcinoma.

[0064] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; and subjecting the sequencing library to whole-genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the adenocarcinoma or squamous cell carcinoma subtype.

[0065] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence using the length and amount of the cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the subtype of adenocarcinoma or squamous cell carcinoma.

[0066] In some embodiments, the identified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

[0067] In some embodiments, the method uses machine learning to subtype small cell lung cancer in a subject.

[0068] In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

[0069] In some embodiments, the log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of a 2 kb window.

[0070] In some embodiments, all coverage calculations are shifted by 1 to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1).

[0071] In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.

[0072] Some embodiments further include using a general additive model to smooth coverage and correct for GC bias.

[0073] In some embodiments, the genomic intervals are non-overlapping.

[0074] In some embodiments, the genomic intervals each contain thousands to millions of base pairs.

[0075] In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval.

[0076] In some embodiments, the cfDNA fragmentation profile comprises a median fragment size.

[0077] In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution.

[0078] Some embodiments further comprise administering to a subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating said cancer type.

[0079] In some embodiments, the therapeutic agent is an immunotherapy.

[0080] Some embodiments further include quantifying circulating tumor fraction, which reflects the performance of major allele frequency (MAF) associated with assessment by Response Evaluation Criteria in Solid Tumors (RECIST) in both adenocarcinoma and squamous cell carcinoma.

[0081] In one embodiment, the present invention provides a method for non-invasively subtyping a subject's non-small cell lung cancer as adenocarcinoma or squamous cell carcinoma, comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole-genome sequencing to obtain sequenced fragments, with a genome coverage of about 30x to 0.1x; mapping the sequenced fragments to the genome to obtain genomic sections of sequences mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments; analyzing the genomic sections of the mapped sequences using the length and amount of the cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; and subtyping the subject's small cell lung cancer based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of adenocarcinoma or squamous cell carcinoma. In some embodiments, the identified transcription factors are ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype small cell lung cancer in a subject. In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site. In some embodiments, the log2(read depth ratio) is the total coverage relative to a coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of the 2 kb window. In some embodiments, all coverage calculations are shifted by 1 to avoid division by 0. In some embodiments, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor) + 1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.Some embodiments further include using a general additive model to smooth coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution. Some embodiments further include administering to a subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some embodiments, the therapeutic agent is an immunotherapy. Some embodiments further include quantifying circulating tumor fraction, which reflects the performance of major allele frequency (MAF) in relation to assessment by Response Evaluation Criteria in Solid Tumors (RECIST) in both adenocarcinoma and squamous cell carcinoma.

[0082] In one embodiment, the invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is between about 10× and 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of sequence mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments and analyzing the genomic intervals of the mapped sequence to establish a cfDNA fragment coverage score at the identified transcription factor binding sites using the length and amount of the cfDNA fragments; and subtyping the non-small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates the adenocarcinoma or squamous cell carcinoma subtype. In some embodiments, the identified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype small cell lung cancer in subjects. In some embodiments, subtype classification is performed by calculating log2(read depth ratio) over a 20bp window, starting from a position 2kb away from the identified transcription factor binding site. In some embodiments, log2(read depth ratio) is the total coverage relative to the coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500bp anchors at both ends of the 2kb window. In some embodiments, all coverage calculations are shifted by 1 to avoid division by 0. In some embodiments, this adjustment is expressed as f(x)=log2((window depth+1) / (median(anchor)+1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median for each window.Some embodiments further include using a general additive model to smooth coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution. Some embodiments further include administering to a subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some embodiments, the therapeutic agent is an immunotherapy. Some embodiments further include quantifying circulating tumor fraction, which reflects the performance of major allele frequency (MAF) in relation to assessment by Response Evaluation Criteria in Solid Tumors (RECIST) in both adenocarcinoma and squamous cell carcinoma.

[0083] In one embodiment, the present invention provides a method for non-invasively subtyping a subject's non-small cell lung cancer as adenocarcinoma or squamous cell carcinoma, comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole-genome sequencing to obtain sequenced fragments, with a genome coverage of about 9x to 0.1x; mapping the sequenced fragments to the genome to obtain genomic sections of sequences mapped to identified transcription factor binding sites; determining the length and amount of the cfDNA fragments; analyzing the genomic sections of the mapped sequences using the length and amount of the cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; and subtyping the subject's small cell lung cancer based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a subtype of adenocarcinoma or squamous cell carcinoma. In some embodiments, the identified transcription factors are ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some embodiments, the method uses machine learning to subtype small cell lung cancer in a subject. In some embodiments, subtyping is performed by calculating the log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site. In some embodiments, the log2(read depth ratio) is the total coverage relative to a coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of the 5' and 3' 500 bp anchors at both ends of the 2 kb window. In some embodiments, all coverage calculations are shifted by 1 to avoid division by 0. In some embodiments, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor) + 1). In some embodiments, to obtain a fragment coverage score, the log2(read depth ratio) is calculated for each transcription factor binding site, and then aggregated by calculating the median value for each window.Some embodiments further include using a general additive model to smooth coverage and correct for GC bias. In some embodiments, the genomic intervals are non-overlapping. In some embodiments, the genomic intervals each comprise thousands to millions of base pairs. In some embodiments, a cfDNA fragmentation profile is determined within each genomic interval. In some embodiments, the cfDNA fragmentation profile comprises a median fragment size. In some embodiments, the cfDNA fragmentation profile comprises a fragment size distribution. Some embodiments further include administering to a subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some embodiments, the therapeutic agent is an immunotherapy. Some embodiments further include quantifying circulating tumor fraction, which reflects the performance of major allele frequency (MAF) in relation to assessment by Response Evaluation Criteria in Solid Tumors (RECIST) in both adenocarcinoma and squamous cell carcinoma.

[0084] Before the present compositions and methods are described, it is to be understood that this invention is not limited to the particular methods and systems described, as such compositions, methods, and systems may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the invention will be limited only in the appended claims.

[0085] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, it will be apparent to persons skilled in the art upon reading this disclosure and so forth that reference to "the method" includes one or more methods and / or steps of the type described herein.

[0086] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are now described.

[0087] Described herein is a method for non-invasively subtyping non-small cell lung cancer (NSCLC) using cell-free DNA (cfDNA) fragmentomes. NSCLC often exists in distinct subtypes, such as adenocarcinoma and squamous cell carcinoma. This novel, non-invasive DELFI-based approach distinguishes between adenocarcinoma and squamous cell carcinoma subtypes of NSCLC by examining cell-free fragments that reflect the NSCLC-specific epigenetic state. No other cfDNA-based liquid biopsy method is known to be able to identify NSCLC subtypes without access to tissue clinical data.

[0088] A DELFI-based approach using cfDNA LC-WGS data can predict NSCLC subtypes. First, the DELFI machine learning classifier detects the presence of cancer in patients with NSCLC. By examining both genome-wide fragmentation profiles and corresponding DELFI-Tumor Fraction (TF) scores, we identify preliminary differences between the two major clinical subtypes of NSCLC cases: adenocarcinoma and squamous cell carcinoma. DELFI-TF accurately detects circulating tumor fraction without being confounded by clonal hematopoiesis. Furthermore, DELFI-TF accurately quantifies circulating tumor fraction, reflecting the performance of MAF in relation to RECIST assessment in both adenocarcinoma and squamous cell carcinoma. We used publicly available data to perform principal component analysis (PCA) and clustered the two subtypes using whole-genome fragmentation data. Figure 19 shows the development of the DELFI-TF model in the proof-of-concept phase. Second, we apply a list of differentially accessible transcription start sites (TSSs), thereby enabling us to distinguish between adenocarcinoma and squamous cell carcinoma. A list of differentially accessible transcription start sites can be obtained from public databases, such as, but not limited to, UCSC or Ensembl. The creation of such a list is similar to that performed for RNA-Seq experiments, except that instead of transcripts per million, read depth ratio is used as a proxy for expression. In some embodiments, the site with the largest and most significant logarithmic fold change is selected from 20% of the test cohort, and then applied to the remaining cohort at the same locus, followed by hierarchical clustering to determine whether the subtypes remain together. In another embodiment, the most highly expressed transcripts are selected from TCGA for a given cancer type, and those transcripts that are expressed at any level in AML (a blood cancer that is an imperfect surrogate for normal blood) are excluded from the list. Third, we investigated whether the most differential DELFI-TSSs exhibit short / long changes to further confirm that these cfDNA molecules are tumor-derived.Finally, we demonstrate the diagnostic accuracy of our approach among identified cluster samples using receiver operating characteristics (ROCs), which represent the sensitivity and specificity of the DELFI-fragmentome approach for identifying NSCLC subtypes. Given the practical difficulties in performing tumor biopsies for NSCLC, we believe this approach may be a feasible and rapid method for subtyping NSCLC in a noninvasive manner.

[0089] This document provides methods and materials for determining a cfDNA fragmentation profile in a mammal (e.g., within a sample obtained from a mammal). As used herein, the terms "fragmentation profile," "position-dependent differences in fragmentation patterns," and "differences in fragment size and coverage in a position-dependent manner across the genome" are synonymous and can be used interchangeably. In some cases, determining a cfDNA fragmentation profile in a mammal can be used to identify a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from the mammal) can be subjected to low-coverage whole-genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and evaluated to determine the cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal with cancer is more heterogeneous (e.g., in fragment length) than the cfDNA fragmentation profile of a healthy mammal (e.g., a mammal without cancer). Thus, this document also provides methods and materials for evaluating, monitoring, and / or treating a mammal (e.g., a human) with or suspected of having cancer. In some cases, this document provides methods and materials for identifying a mammal as having cancer. For example, a sample (e.g., a blood sample) collected from a mammal can be evaluated to determine the presence of cancer in the mammal, and optionally the tissue of origin of the cancer, based at least in part on the cfDNA fragmentation profile of the mammal. In some cases, this document provides methods and materials for monitoring a mammal as having cancer. For example, a sample (e.g., a blood sample) collected from a mammal can be evaluated to determine the presence of cancer in the mammal, based at least in part on the cfDNA fragmentation profile of the mammal. In some cases, this document provides methods and materials for identifying a mammal as having cancer and administering one or more cancer therapies to the mammal to treat the mammal.For example, a sample (e.g., a blood sample) taken from a mammal can be evaluated to determine whether the mammal has cancer based at least in part on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.

[0090] The cfDNA fragmentation profile may include one or more cfDNA fragmentation patterns. The cfDNA fragmentation pattern may include any suitable cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, but are not limited to, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and cfDNA fragment coverage. In some cases, the cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and cfDNA fragment coverage. In some cases, the cfDNA fragmentation profile may be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in a genome-wide window). In some cases, the cfDNA fragmentation profile may be a profile of a targeted region. The targeted region may be any suitable portion of the genome (e.g., a chromosomal region). As described herein, examples of chromosomal regions that can determine cfDNA fragmentation profiles include, but are not limited to, parts of chromosomes (e.g., parts of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or 14q) and chromosomal arms (e.g., parts of 8q, 13q, 11q, and / or 3p).In some cases, cfDNA fragmentation profiles can include two or more target region profiles.

[0091] In some cases, cfDNA fragmentation profiles can be used to identify variations (e.g., changes) in cfDNA fragment lengths. The changes can be genome-wide changes or changes in one or more target regions / locuses. The target region can be any region that contains one or more cancer-specific changes. Examples of cancer-specific changes and their chromosomal locations include, but are not limited to, those shown in Table 3 (Appendix C) and Table 6 (Appendix F). In some cases, the fragmentation profile of cfDNA can be used to identify (e.g., simultaneously identify) between about 10 and about 500 alterations (e.g., between about 25 and about 500, between about 50 and about 500, between about 100 and about 500, between about 200 and about 500, between about 300 and about 500, between about 10 and about 400, between about 10 and about 300, between about 10 and about 200, between about 10 and about 100, between about 10 and about 50, between about 20 and about 400, between about 30 and about 300, between about 40 and about 200, between about 50 and about 100, between about 20 and about 100, between about 25 and about 75, between about 50 and about 250, or between about 100 and about 200 alterations).

[0092] In some cases, cfDNA fragmentation profiles can be used to detect tumor-derived DNA. For example, cfDNA fragmentation profiles can be used to detect tumor-derived DNA by comparing the cfDNA fragmentation profile of a mammal with or suspected of having cancer with a reference cfDNA fragmentation profile (e.g., the cfDNA fragmentation profile of a healthy mammal and / or the nucleosomal DNA fragmentation profile of healthy cells collected from a mammal with or suspected of having cancer). In some cases, the reference cfDNA fragmentation profile is a profile previously generated from a healthy mammal. For example, the methods provided herein can be used to determine a reference cfDNA fragmentation profile in a healthy mammal, and the reference cfDNA fragmentation profile can be stored (e.g., in a computer or other electronic storage medium) for future comparison with a test cfDNA fragmentation profile in a mammal with or suspected of having cancer. In some cases, the reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined across the entire genome. In some cases, the reference cfDNA fragmentation profile (e.g., a stored cfDNA fragmentation profile) of a healthy mammal is determined across a subgenomic interval.

[0093] In some cases, the cfDNA fragmentation profile can be used to identify a mammal (e.g., a human) as having cancer (e.g., colon cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer).

[0094] The cfDNA fragmentation profile can include a cfDNA fragment size pattern. The cfDNA fragments can be of any appropriate size. For example, the cfDNA fragments can be about 50 base pairs (bp) to about 400 bp in length. As described herein, a mammal with cancer can have a cfDNA fragment size pattern that includes a median cfDNA fragment size that is shorter than the median cfDNA fragment size in healthy mammals. A healthy mammal (e.g., a mammal without cancer) can have a cfDNA fragment size with a median cfDNA fragment size of about 166.6 bp to about 167.2 bp (e.g., about 166.9 bp). In some cases, a mammal with cancer can have a cfDNA fragment size that is, on average, about 1.28 bp to about 2.49 bp (e.g., about 1.88 bp) shorter than the cfDNA fragment size in healthy mammals. For example, a mammal with cancer can have a cfDNA fragment size with a median cfDNA fragment size of about 164.11 bp to about 165.92 bp (e.g., about 165.02 bp).

[0095] The cfDNA fragmentation profile may include a cfDNA fragment size distribution. As described herein, a mammal with cancer may have a cfDNA size distribution that is more variable than that of a healthy mammal. In some cases, the size distribution may exist within a target region. A healthy mammal (e.g., a mammal without cancer) may have a cfDNA fragment size distribution of a target region in which the cfDNA fragment size distribution of the target region is about 1 or less. In some cases, a mammal with cancer may have a cfDNA fragment size distribution of a target region that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp or more longer, or any number of base pairs longer within these ranges) than that of a healthy mammal. In some cases, a mammal with cancer may have a cfDNA fragment size distribution in a target region that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 bp or more shorter, or any number of base pairs shorter within these ranges) than the cfDNA fragment size distribution in the target region in a healthy mammal. In some cases, a mammal with cancer may have a cfDNA fragment size distribution in a target region that is about 47 bp shorter to about 30 bp longer than the cfDNA fragment size distribution in the target region in a healthy mammal. In some cases, a mammal with cancer may have a cfDNA fragment size distribution in a target region that differs, on average, by 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20 bp or more in cfDNA fragment length. In some cases, a mammal with cancer may have a cfDNA fragment size distribution in a target region that differs, on average, by about 13 bp in cfDNA fragment length. In some cases, the size distribution may be a genome-wide size distribution.Healthy mammals (for example, mammals without cancer) may have genome-wide very similar distribution of short cfDNA fragments and long cfDNA fragments.In some cases, mammals with cancer may have genome-wide one or more changes (for example, increase and decrease) in cfDNA fragment size.The one or more changes may be in any suitable chromosomal region of genome.For example, changes may occur in a part of chromosome.Examples of chromosomal parts that may contain one or more changes in cfDNA fragment size include, but are not limited to, 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q parts.For example, changes may occur across chromosome arms (for example, the entire chromosome arms).

[0096] The cfDNA fragmentation profile may include the ratio of small cfDNA fragments to large cfDNA fragments, and the correlation of the fragment ratio with a reference fragment ratio. Herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the small cfDNA fragments may be about 100 bp to about 150 bp. Herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the large cfDNA fragments may be about 151 bp to 220 bp. As described herein, a mammal with cancer may have a lower fragment ratio correlation (e.g., a correlation of the cfDNA fragment ratio with a reference DNA fragment ratio, such as a DNA fragment ratio from one or more healthy mammals) than a healthy mammal (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower, or more). A healthy mammal (e.g., a mammal without cancer) may have a fragment ratio correlation (e.g., a correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals) of about 1 (e.g., about 0.96). In some cases, a mammal with cancer may have a fragment ratio correlation (e.g., a correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals) that is, on average, about 0.19 to about 0.30 (e.g., about 0.25) lower than the fragment ratio correlation in healthy mammals (e.g., a correlation of cfDNA fragment ratios to a reference DNA fragment ratio, such as DNA fragment ratios from one or more healthy mammals).

[0097] cfDNA fragmentation profile can include coverage of all fragments.Coverage of all fragments can include coverage windows (for example, non-overlapping windows).In some cases, coverage of all fragments can include windows of small fragments (for example, fragments with a length of about 100bp to about 150bp).In some cases, coverage of all fragments can include windows of large fragments (for example, fragments with a length of about 151bp to about 220bp).

[0098] The cfDNA fragmentation profile can be obtained using a suitable method. In some cases, cfDNA from a mammal (e.g., a mammal having or suspected of having cancer) can be processed into a sequencing library for whole genome sequencing (e.g., low-coverage whole genome sequencing), mapped to the genome, and analyzed to determine cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows covering the genome. The windows can be of any suitable size. For example, the windows can be thousands to millions of bases in length. As a non-limiting example, the windows can be about 5 megabases (Mb) in length. Any suitable number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. A cfDNA fragmentation profile can be determined within each window. In some cases, the cfDNA fragmentation profile can be obtained as described in Example 1. In some cases, the cfDNA fragmentation profile can be obtained as shown in Figure 1.

[0099] In some cases, methods and materials described herein can also comprise machine learning.For example, machine learning can be used to identify altered fragmentation profile (for example, using cfDNA fragment coverage, cfDNA fragment size, chromosome coverage, mtDNA).

[0100] In some cases, the methods and materials described herein may be the only method used to identify a mammal (e.g., a human) as having cancer (e.g., colon cancer, lung cancer, breast cancer, stomach cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). For example, determining a cfDNA fragmentation profile may be the only method used to identify a mammal as having cancer.

[0101] In some cases, the methods and materials described herein can be used in conjunction with one or more additional methods used to identify a mammal (e.g., a human) as having cancer (e.g., colon cancer, lung cancer, breast cancer, stomach cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). Examples of methods used to identify a mammal as having cancer include, but are not limited to, identifying one or more cancer-specific sequence alterations, identifying one or more chromosomal alterations (e.g., aneuploidies and rearrangements), and identifying other cfDNA alterations. For example, determining a cfDNA fragmentation profile, along with identifying one or more cancer-specific mutations in the mammal's genome, can be used to identify a mammal as having cancer. For example, determining a cfDNA fragmentation profile, along with identifying one or more aneuploidies in the mammal's genome, can be used to identify a mammal as having cancer.

[0102] In some embodiments, this document also provides methods and materials for evaluating, monitoring, and / or treating a mammal (e.g., a human) that has or is suspected of having cancer. In some cases, this document provides methods and materials for identifying that a mammal has cancer. For example, a sample (e.g., a blood sample) collected from a mammal can be evaluated to determine whether the mammal has cancer based at least in part on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for identifying the site (e.g., anatomical site or tissue of origin) of cancer in a mammal. For example, a sample (e.g., a blood sample) collected from a mammal can be evaluated to determine the tissue of origin of cancer in the mammal based at least in part on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for determining that a mammal has cancer and administering one or more cancer therapies to the mammal to treat the mammal. For example, a sample (e.g., a blood sample) collected from a mammal can be evaluated to determine whether the mammal has cancer, based at least in part on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments are administered to the mammal. In some cases, this document provides methods and materials for treating a mammal with cancer. For example, one or more cancer treatments can be administered to a mammal identified as having cancer (e.g., based at least in part on the cfDNA fragmentation profile of the mammal) to treat the mammal. In some cases, during or after the course of cancer treatment (e.g., any of the cancer treatments described herein), the mammal can be monitored (or selected for enhanced monitoring) and / or undergo further diagnostic testing.In some cases, monitoring may include evaluating a mammal having or suspected of having cancer, e.g., by evaluating a sample (e.g., a blood sample) taken from the mammal to determine a cfDNA fragmentation profile for the mammal, as described herein, and changes in the cfDNA fragmentation profile over time can be used to identify a response to treatment and / or to identify the mammal as having cancer (e.g., residual cancer).

[0103] Any suitable mammal can be evaluated, monitored, and / or treated as described herein.The mammal can be a mammal with cancer.The mammal can be a mammal suspected of having cancer.Examples of mammals that can be evaluated, monitored, and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice, and rats.For example, a human with or suspected of having cancer can be evaluated to determine a cfDNA fragmentation profile as described herein, and optionally can be treated with one or more cancer treatments as described herein.

[0104] Any suitable sample from a mammal can be evaluated (e.g., evaluated for DNA fragmentation patterns) as described herein. In some cases, the sample can include DNA (e.g., genomic DNA). In some cases, the sample can include cfDNA (e.g., circulating tumor DNA (ctDNA)). In some cases, the sample can be a liquid sample (e.g., liquid biopsy). Examples of samples that can include DNA and / or polypeptides include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic membrane, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph, cyst fluid, stool, ascites, Pap smear, breast milk, and exhaled breath condensate. For example, a plasma sample can be evaluated to determine a cfDNA fragmentation profile as described herein.

[0105] As described herein, the sample from mammal that is evaluated (for example, that is evaluated for DNA fragmentation pattern) can contain any suitable amount of cfDNA.In some cases, sample can contain limited amount of DNA.For example, cfDNA fragmentation profile can be obtained from the sample that contains less DNA than is usually required for other cfDNA analysis methods described in, for example, Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; and Newman et al., 2016 Nat Biotechnol 34:547).

[0106] In some cases, the sample may be treated (e.g., to isolate and / or purify DNA and / or polypeptides from the sample). For example, DNA isolation and / or purification may include cell lysis (e.g., with detergents and / or surfactants), protein removal (e.g., with proteases), and / or RNA removal (e.g., with RNases). As another example, polypeptide isolation and / or purification may include cell lysis (e.g., with detergents and / or surfactants), DNA removal (e.g., with DNases), and / or RNA removal (e.g., with RNases).

[0107] Additional methods are described in US Pat. Nos. 10,982,279 and 10,975,431, the disclosures of which are considered part of the disclosure of this application and are incorporated herein by reference in their entireties. Example of a Hardware Implementation

[0108] Figure 20 illustrates an example of a computer 800 that can be used to implement the methods described herein. For example, the computer 800 can include a machine learning system that trains a machine learning model for subtyping small cell lung cancer or non-small cell lung cancer, as described above, or a portion or combination thereof. The computer 800 can be any electronic device that executes software applications generated from compiled instructions, including, but not limited to, a personal computer, a server, a smartphone, a media player, an electronic tablet, a game console, an email device, etc. In some implementations, the computer 800 can include one or more processors 802, one or more input devices 804, one or more display devices 806, one or more network interfaces 808, and one or more computer-readable media 812. Each of these components can be coupled by a bus 810, or in some embodiments, these components can be distributed across multiple physical locations and coupled by a network.

[0109] The display device 806 may be of any known display technology, including, but not limited to, displays using liquid crystal display (LCD) or light-emitting diode (LED) technology. The processor(s) 802 may use any known processor technology, including, but not limited to, graphics processors and multi-core processors. The input device(s) 804 may be of any known input device technology, including, but not limited to, a keyboard (including a virtual keyboard), a mouse, a trackball, a camera, a touch-sensitive pad, or a display. The bus 810 may be of any known internal or external bus technology, including, but not limited to, ISA, EISA, PCI, PCI Express, USB, Serial ATA, or FireWire. The computer-readable medium 812 may be any non-transitory medium involved in providing instructions to the processor(s) 804 for execution, including, but not limited to, non-volatile storage media (e.g., optical disks, magnetic disks, flash drives, etc.) or volatile media (e.g., SDRAM, ROM, etc.).

[0110] The computer-readable medium 812 may include various instructions 814 for implementing an operating system (e.g., Mac OS, Windows, Linux), which may be multi-user, multi-processing, multi-tasking, multi-threading, real-time, etc. The operating system may perform basic tasks including, but not limited to, recognizing input from the input device 804, sending output to the display device 806, keeping track of files and directories on the computer-readable medium 812, controlling controllable peripheral devices (disk drives, printers, etc.) directly or through I / O controllers, and managing traffic on the bus 810. The network communication instructions 816 may establish and maintain network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, Ethernet, telephony, etc.).

[0111] Machine learning instructions 818 may include instructions that enable computer 800 to function as a machine learning system and / or train a machine learning model to generate DMS values ​​as described herein. Application(s) 820 may be applications that use or implement the processes described herein and / or other processes. The processes may also be implemented in operating system 814. For example, application 820 and / or the operating system may create tasks in the application as described herein.

[0112] The features described herein may be implemented in one or more computer programs executable on a programmable system including at least one programmable processor connected to receive data and instructions from, and transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a particular activity or bring about a particular result. Computer programs may be written in any type of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0113] Processors suitable for executing a program of instructions may include, by way of example, general-purpose and special-purpose microprocessors, the sole processor, or one of multiple processors or cores of any type of computer. Generally, a processor may receive instructions and data from a read-only memory, a random-access memory, or both. The basic elements of a computer may include a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer may include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data may include, by way of example, all forms of non-volatile memory, such as semiconductor memory devices, such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, application-specific integrated circuits (ASICs).

[0114] To provide for user interaction, these functions may be implemented on a computer with a display device, such as an LED or LCD monitor, to display information to the user, and a keyboard and pointing device, such as a mouse or trackball, to allow the user to provide input to the computer.

[0115] This functionality may be implemented in a computer system that includes back-end components such as a data server, a computer system that includes middleware components such as an application server or an Internet server, a computer system that includes front-end components such as a client computer having a graphical user interface or an Internet browser, or any combination thereof. The components of this system may be connected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include, for example, the telephone network, a LAN, a WAN, and the computers and networks forming the Internet.

[0116] A computer system may include clients and servers. Clients and servers may generally be remote from each other and typically interact through a network. The relationship of client and server may arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0117] One or more features or steps of the disclosed embodiments may be implemented using an application programming interface (API), which may define one or more parameters passed between a calling application and other software code (e.g., an operating system, library routine, function) that provides a service, provides data, or performs an operation or calculation.

[0118] An API may be implemented as one or more calls in program code that receive or send one or more parameters through a parameter list or other structure based on a calling convention defined in an API specification. A parameter may be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters may be implemented in any programming language. This programming language may define the vocabulary and calling conventions that programmers employ to access functions that support the API.

[0119] In some implementations, calls to the API may report to the application the capabilities of the device on which the application is running, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, etc.

[0120] While various embodiments have been described above, it should be understood that they have been presented by way of example, and not limitation. It will be apparent to one skilled in the relevant art(s) that various changes in form and detail can be made therein without departing from the spirit and scope. Indeed, after reading the above description, it will be apparent to one skilled in the relevant art(s) how to implement alternative embodiments. For example, other steps can be added to or deleted from the described flows, and other components can be added to or deleted from the described systems. Accordingly, other implementations are within the scope of the following claims.

[0121] Furthermore, it should be understood that any diagrams highlighting features and advantages are presented for illustrative purposes only, and that the disclosed techniques and systems are each sufficiently flexible and configurable that they may be utilized in ways other than those illustrated.

[0122] In this specification, claims, and drawings, the term "at least one" is often used, but terms such as "a," "an," "the," and "said" also mean "at least one" or "said at least one" in this specification, claims, and drawings.

[0123] Finally, it is Applicant's intention that only those claims which include the phrase "means for" or "step for" be construed under 35 U.S.C. 112(f). Any claim which does not expressly include the phrase "means for" or "step for" shall not be construed under 35 U.S.C. 112(f).

[0124] The methods and systems described herein are useful for subtyping non-small cell lung cancer or small cell lung cancer in a subject, and optionally treating the cancer subtype in the subject. Any suitable subject, such as a mammal, can be evaluated and / or treated as described herein. Examples of mammals that can be evaluated and / or treated as described herein include, but are not limited to, humans, primates such as monkeys, dogs, cats, horses, cows, pigs, sheep, mice, and rats. For example, a human with or suspected of having cancer can be evaluated using the methods described herein and optionally treated with one or more cancer treatments described herein. The methods disclosed herein can include administering to a subject identified as having a particular type of cancer a therapeutic agent appropriate for treating that type of cancer.

[0125] When treating a subject having or suspected of having cancer as described herein, the subject can be administered one or more cancer therapies. The cancer treatment can be any suitable cancer treatment. One or more cancer therapies described herein can be administered to the subject at any appropriate frequency (e.g., one or more times over a period of several days to several weeks). Examples of cancer therapies include, but are not limited to, surgical intervention, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormonal therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., T cells having chimeric antigen receptors and / or wild-type or modified T cell receptors), targeted therapy (e.g., kinase inhibitors, antibodies, bispecific antibodies) such as administration of kinase inhibitors (e.g., kinase inhibitors that target specific genetic abnormalities such as translocations or mutations), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or any combination of the above. In some embodiments, cancer treatment can reduce the severity of the cancer, reduce the symptoms of the cancer, and / or reduce the number of cancer cells present in the subject.

[0126] In some embodiments, the cancer therapeutic agent can be a chemotherapeutic agent. Non-limiting examples of chemotherapeutic agents include amsacrine, azacitidine, azathioprine, bevacizumab (or an antigen-binding fragment thereof), bleomycin, busulfan, carboplatin, capecitabine, chlorambucil, cisplatin, cyclophosphamide, cytarabine, dacarbazine, daunorubicin, docetaxel, doxifluridine, doxorubicin, epirubicin, erlotinib hydrochloride, etoposide, fludarabine, floxuridine, fludarabine, fluorouracil, gemcitabine, hydroxyurea, riboflavin ...

[0013] Examples of anti-cancer drug treatments include idarubicin, idarubicin, ifosfamide, irinotecan, lomustine, mechlorethamine, melphalan, mercaptopurine, methotrexate, mitomycin, mitoxantrone, oxaliplatin, paclitaxel, pemetrexed, procarbazine, all-trans retinoic acid, streptozocin, tafluposide, temozolomide, teniposide, thioguanine, topotecan, uramustine, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, and combinations thereof. Additional examples of anti-cancer drug treatments are known in the art, see, for example, the American Society of Clinical Oncology (ASCO), the European Society for Medical Oncology (ESMO), or the National Comprehensive Cancer Network (NCCN) treatment guidelines.

[0127] In various embodiments, DNA is present in a biological sample collected from a subject and used in the methods of the present invention. The biological sample can be virtually any type of biological sample containing DNA. The biological sample is typically a liquid, such as whole blood or a portion thereof containing circulating cfDNA. In embodiments, the sample contains DNA from a tumor or a liquid biopsy, including, but not limited to, amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), earwax, chyle, oozing fluid, endolymph, perilymph, feces, exhaled breath, gastric acid, gastric juice, lymph, mucus (including nasal discharge and sputum), pericardial fluid, ascites, pleural fluid, pus, ocular discharge, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomit, prostatic fluid, nipple aspirate, tears, sweat, buccal swab, cell lysate, gastrointestinal fluid, biopsy tissue, urine, or other biological fluids. In one embodiment, the sample comprises DNA from circulating tumor cells.

[0128] As disclosed above, the biological sample can be a blood sample. The blood sample can be collected using methods known in the art, such as fingerstick or venous blood collection. Preferably, the blood sample is about 0.1-20 ml, or about 1-15 ml, with a blood volume of about 10 ml. Smaller amounts of blood can be used, and free DNA circulating in the blood can be used. Microsampling, sampling by needle biopsy, catheter, stool, or generation of bodily fluids containing DNA are also potential sources of biological samples.

[0129] The methods and systems of the present disclosure utilize nucleic acid sequence information, and therefore may include any method or sequencing device for performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, and insertion tag sequencing. In some embodiments, the methods or systems of the present disclosure utilize systems such as those provided by Illumina, Inc. (including but not limited to HiSeq™ X10, HiSeq™ 1000, HiSeq™ 2000, HiSeq™ 2500, Genome Analyzers™, MiSeq™, NextSeq, and NovaSeq 6000 systems), Applied Biosystems Life Technologies (SOLiD™ System, ion Proton™ Sequencer, ion Proton™ Sequencer), or Genapsys or BGI MGI and other systems. Nucleic acid analysis can also be performed with systems provided by Oxford Nanopore Technologies (GridiON™, MinION™) or Pacific Biosciences (Pacbio™ RS II or Sequel I or II).

[0130] The present invention includes systems for performing the steps of the disclosed methods, which are described in part in terms of functional components and various processing steps. Such functional components and processing steps may be realized by any number of components, operations, and techniques configured to perform the specified functions and achieve various results. For example, the present invention may employ various biological samples, biomarkers, elements, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing standards, statistical analyses, regression analyses, and the like, which may perform various functions.

[0131] Accordingly, the present invention further provides a non-invasive system for subtyping small cell lung cancer or non-small cell lung cancer, in various embodiments, the system including (a) a sequencer configured to generate a low-coverage whole genome sequencing dataset for a sample, and (b) a computer system and / or processor having functionality for performing the methods of the present invention.

[0132] In some embodiments, this computer system further comprises one or more additional modules.For example, this system can comprise one or more suitable genetic component analysis, for example, extraction and / or isolation units that can be operated to select cfDNA fragments of specific sizes.

[0133] In some embodiments, the computer system further comprises a visual display device, which may be operable to display the curve fit line, the reference curve fit line, and / or a comparison of the two.

[0134] The method for non-invasive subtyping of small cell lung cancer or non-small cell lung cancer according to various aspects of the present invention can be implemented in any suitable manner, such as using a computer program running on a computer system. As discussed herein, exemplary systems according to various aspects of the present invention can be implemented in combination with a conventional computer system, such as a remotely accessible application server, network server, personal computer, or workstation, including a processor and random access memory. The computer system also preferably includes additional storage or information storage systems, such as a mass storage system, and a user interface, such as a conventional monitor, keyboard, and tracking device. However, the computer system can include any suitable computer system and related equipment and can be configured in any suitable manner. In one embodiment, the computer system comprises a stand-alone system. In another embodiment, the computer system is part of a network of computers, including a server and a database.

[0135] The software necessary to receive, process, and analyze information may be implemented in a single device or in multiple devices. This software may be accessible over a network so that information storage and processing is performed remotely from the user. Systems and their various elements according to various embodiments of the present invention provide functions and operations that facilitate detection and / or analysis, such as data collection, processing, analysis, reporting, and / or diagnosis. For example, in this embodiment, a computer system executes a computer program that can receive, store, retrieve, analyze, and report information related to the human genome or regions thereof. This computer program may include multiple modules that perform various functions or operations, such as a processing module that processes raw data to generate supplemental data and an analysis module that analyzes the raw data and supplemental data to generate quantitative assessments of disease state models and / or diagnostic information.

[0136] The processing performed by the system may include any suitable process for facilitating the analysis and / or subtyping of small cell lung cancer or non-small cell lung cancer. In one embodiment, the system is configured to establish a disease subtype model and / or calculate a disease subtype in the patient. Determining or identifying a disease subtype may include generating any useful information regarding the patient's condition related to the disease, such as making a diagnosis, providing information useful for diagnosis, assessing the stage or progression of the disease, identifying conditions that may indicate susceptibility to the disease, identifying whether further testing is recommended, predicting and / or evaluating the effectiveness of one or more treatment programs, or otherwise assessing the patient's disease state, likelihood of disease, or other health aspects.

[0137] The following examples are provided to further illustrate the advantages and features of the present invention and are not intended to limit the scope of the invention. While the examples are typical of those that might be used, other procedures, methods, or techniques known to those skilled in the art may alternatively be used. [Example]

[0138] Example 1 Analysis of small cell lung cancer subtypes by cell-free DNA fragmentome BACKGROUND: Small cell lung cancer (SCLC) is an aggressive form of lung cancer that is strongly associated with smoking and exposure to other environmental chemicals. Patients with SCLC suffer from a 5-year survival rate of less than 8% (Gay, C. M. et al. Patterns of transcription factor programs and immune pathway activation define four major subtypes of SCLC with distinct therapeutic vulnerabilities. Cancer Cell 39, 346-360.e7 (2021)). SCLC is a heterogeneous tumor type composed of tumor cells with both neuroendocrine and non-neuroendocrine features. SCLC primarily has two subtypes: high-grade neuroendocrine (NE-high) and low-grade neuroendocrine (NE-low) (Gay et al.). The NE-high subtype is characterized by activation of lineage-specific transcription factors such as ASCL1 and NEUROD1. The NE-low subtype is characterized by non-neuroendocrine factors such as POU2F3 (SCLC-P) or high abundance of inflammatory T cells (SCLC-I). Despite this molecular and clinical heterogeneity, SCLC is treated as a single disease and predicts poor outcomes. Recent results have shown that the inflammatory group (SCLC-I) has distinct biological characteristics and responds better to immunotherapy compared to other SCLC subtypes (Gay et al. and Lissa, D. et al. Heterogeneity of neuroendocrine transcriptional states in metastatic small cell lung cancer). (Nat Commun 13, 2023 (2022)). DELFI-based cell-free fragmentome accurately distinguishes between NE-high and NE-low SCLC subtypes in a noninvasive manner. The ultimate goal of DELFI SCLC subtyping is to guide the selection of the optimal treatment for each patient diagnosed with advanced SCLC.

[0139] Methods: Circulating cell-free DNA (cfDNA) was isolated from plasma samples of patients diagnosed with recurrent SCLC and treated with durvalumab and olaparib in a phase II trial (NCT02484404). Using immunohistochemistry and genomics data from pretreatment tissue biopsies, SCLC subtypes were classified into high NE (n=10) and low NE (n=5). To estimate tumor gene expression profiles in cfDNA, we investigated genome-wide signals of tissue-specific transcription factors differentially regulated in SCLC and applied a novel DELFI-based approach to provide information on molecular subtypes of SCLC. Specifically, DELFI calculates the fragment distribution for each TFBS within a 2kb window and scales it between 0 and 1 independently for each sample. Next, we performed principal component analysis in R to define clusters based on the TFBS sites of ASCL1 (Mathios, D. et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun 12, 5060 (2021)). Clinical information was examined using orthogonal analysis.

[0140] Results: DELFI's proprietary fragmentomics platform detected SCLC cases (N=47) with high sensitivity, with a median DELFI score of 1.0 (95% CI 0.99-1), consistent with existing data for SCLC samples within DELFI. Genome-wide cfDNA fragmentation analysis of ASCL1 binding sites (approximately 12,000 genomic coordinates) in SCLC patients revealed reduced coverage of transcription factor binding sites in SCLC patients compared with non-cancer and NSCLC samples. Differences in the center of ASCL1 binding sites were the primary driver of differentiation between pretreatment SCLC samples. Examination of the fragment distribution of ASCL1 binding sites was able to distinguish high-NE from low-NE samples. Patients classified as low-NE by the DELFI subtyping system (treated with I / O) had better treatment responses than patients classified as high-NE by the DELFI classification system. Differential pseudogene expression identified thousands of highly enriched transcription start sites (TSSs) in low-NE cases compared with high-NE cases. Low-NE samples showed a strong enrichment of the T-lymphocyte-specific factor KLF4, reflecting the neutrophil-to-lymphocyte ratio of this SCLC subtype. These data reflect an inflammatory phenotype and perhaps explain why these cancer subtypes are more responsive to I / O therapy.

[0141] Table 1. Patient characteristics and relevant clinical information TIFF2026506641000002.tif117135

[0142] Example 2 Use of the DELFI circulating cfDNA-based approach to subtype small cell lung cancer (SCLC) in a non-invasive manner using the cell-free DNA fragmentome.

[0143] SCLC is an aggressive malignancy with a poor prognosis. Although SCLC is clinically managed as a single cancer type, emerging evidence supports that these subtypes of SCLC (high neuroendocrine and low neuroendocrine) acquire diverse transcriptional and epigenetic states. Moreover, different SCLC subtypes respond to specific treatments, such as immunotherapy in the case of low neuroendocrine high inflammatory SCLC subtypes.

[0144] The novel, noninvasive DELFI-based approach described herein interrogates cell-free fragments that reflect the unique epigenetic status of SCLC and distinguishes between neuroendocrine-rich and neuroendocrine-poor SCLC subtypes. This novel method can identify patients who are more likely to respond to immunotherapies, such as immune checkpoint blockade.

[0145] Using WGS data from plasma-derived cfDNA, we investigated fragment coverage at specific genomic coordinates to identify specific subtypes of SCLC cases. Targeted analysis based on DELFI can distinguish between high- and low-neuroendocrine SCLC cases.

[0146] SCLC subtype can be predicted from a DELFI-based approach using cfDNA LC-WGS data.

[0147] First, the DELFI machine learning classifier detects the presence of cancer in patients with SCLC. We investigated both genome-wide fragmentation profiles and their corresponding DELFI scores to identify preliminary differences between two major clinical subtypes of SCLC cases: high neuroendocrine and low neuroendocrine. We then performed principal component analysis using publicly available data to cluster the two subtypes.

[0148] Second, using publicly available data, we confirmed that these clusters of SCLC samples exhibited reduced total fragment coverage in ASCL1 transcription factor binding sites, classifying these cases as high-neuroendocrine SCLC.

[0149] Third, we also investigated whether one of these clusters of SCLC samples displays reduced fragment coverage at genomic binding sites regulated by hematopoietic transcription factors (hyponeuroendocrine SCLC).

[0150] Fourth, we distinguished between high and low neuroendocrine cases based on fragment length distribution in T cell-specific partially methylated domains.

[0151] Finally, the receiver operating characteristics (ROC), which represent the sensitivity and specificity of the DELFI-fragmentome approach for identifying SCLC subtypes, were used to determine the diagnostic accuracy of our approach among the identified cluster samples.

[0152] Given the limited access to SCLC tumor biopsies, this approach may be a viable method for subtyping SCLC in a noninvasive manner, offering promise for both pharmaceutical companies in clinical trials evaluating novel immunotherapeutic agents and clinicians seeking to select patients diagnosed with SCLC who are most likely to respond to immunotherapy.

[0153] The DELFI scores of patients diagnosed with SCLC in this study are shown in Figure 1.

[0154] Figure 2 illustrates that genome-wide fragmentation profiles were significantly consistent among pre-treatment, post-treatment, progression, and response time points, suggesting that genome-wide circulating tumor DNA fragment size is a powerful method for detecting changes during immunotherapy monitoring of SCLC cancer treatment.

[0155] Figure 3 illustrates the genome-wide fragmentation profiles (and corresponding DELFI scores at the bottom) showing the partial differences between the two main subtypes of SCLC cases, i.e., high and low neuroendocrine subtypes, suggesting that genome-wide circulating tumor DNA fragment size may be a powerful method for detecting subtypes during immunotherapy treatment monitoring of SCLC cancer treatment.

[0156] Figure 4 illustrates the differential genes between these two cases, clustering the two subtypes of differential genes between high and low neuroendocrine SCLC cases using publicly available data (Lissa et al., 2022) and principal component analysis (DELFI 30X plasma WGS data). PCA analysis data revealed distinct clusters between pre-treatment samples of high and low neuroendocrine SCLC.

[0157] FIG. 5 illustrates that Eigencor analysis revealed significant correlations between principal components 1 and 2 and the neuroendocrine status of the samples.

[0158] Figures 6A-B illustrate the potential of the DELFI assay to subtype SCLC using TFBS fragments. We performed genome-wide cfDNA fragmentation analysis at ASCL1 binding sites within pre-treatment NCI samples. Figure 6A illustrates the observation of distinct clusters of SCLC samples corresponding to different ASCL1 activation levels. Importantly, the two clusters perfectly corresponded to distinct SCLC subtypes. The primary factor driving the differences between samples was in fact the difference in fragment coverage identified at the center of the ASCL1 binding site. Figure 6B illustrates that other clinical and sample characteristics had little impact on the clustering of these samples. Overall, these data demonstrate that DELFI analysis can perfectly distinguish SCLC subtypes using TFBS fragment coverage.

[0159] Figure 7 illustrates an example patient, supporting the hypothesis that the detected signal originates from tumor-infiltrating lymphocytes. Patient NCI-0422, diagnosed with SCLC of the inflammatory subtype, was treated with a combination of durvalumab and olaparib. NCI-0422 responded well to treatment but was later diagnosed with progressive disease. Variable genomic binding is detectable both pretreatment and at progression. Peaks are concentrated at positions where TFBSs and TSSs intersect.

[0160] Figures 8A-B illustrate ASCL1 binding sites between responders and non-responders. Figure 8A shows 500 cell type-specific bins from PMD data to calculate leukocyte percentages. Figure 8B shows lymphocyte tissue-specific bins.

[0161] Figure 9 illustrates the machine learning model for predicting high and low neuroendocrine SCLC.

[0162] Subtype classification was performed by calculating the log2(read depth ratio) over a 20-bp window, starting 2 kb from the ASCL1 TFBS site. The read depth ratio is the total coverage over a coverage correction factor (i.e., log2(read depth / correction factor)), calculated using the median coverage of the 5' and 3' 500-bp anchors at either end of the 2 kb window. A general additive model was then used to smooth the coverage and correct for GC bias. To validate the approach described herein, neuroendocrine status was derived from the mean NE50 score per subject / treatment status, with positive values ​​classified as NE+. Subtype status was derived from the aggregation of the mean values ​​for each of the four subtypes (ASCL1, NEUROD1, POUF23, YAP1), such that each subject / treatment status had one value for each subtype. The subtype with the highest score was then selected as true. Fragmentomics were calculated according to a previously described method (Cristiano et al.) and scores are based on the most recent models available.

[0163] Figures 10A-B illustrate that the methods described herein can distinguish between neuroendocrine and non-neuroendocrine by calculating read depth ratios at SCLC-specific genomic coordinates (Figure 10A). Figure 10B shows that the methods described herein appear to be better able to subtype SCLC cases.

[0164] Figure 11 illustrates the transformation of read depth ratios at SCLC-specific genomic coordinates into 2D PCA, separating samples within the neuroendocrine and non-neuroendocrine groups. Most samples clearly fall into distinct clusters.

[0165] Figure 12 illustrates that PCA derived from read depth ratios at SCLC-specific genomic coordinates can separate pretreatment samples into specific SCLC subtypes (A, N, P, Y). The true SCLC subtypes were calculated using tissue RNA-seq differential gene expression analysis.

[0166] The fragmentomics platform described herein detects small cell lung cancer (SCLC) with high sensitivity. The DELFI SCLC subtyping assay can subtype SCLC samples without the need for clinical knowledge. The DELFI SCLC subtyping assay identified four SCLC sample groups with different levels of ASCL1 activation. Samples with high ASCL1 activation were identified as belonging to the neuroendocrine subtype (SCLC-A and SCLC-N), while samples with low ASCL1 signal were identified as belonging to the non-neuroendocrine subtype (SCLC-P and SCLC-Y). Example 3

[0167] Cell-free DNA fragmentation profiling as a method for assessing tumor fraction and monitoring treatment in NSCLC

[0168] cfDNA Fragmentomics Background: Traditionally, monitoring has been performed using imaging, and many individuals do not have easy access to hospitals with appropriate imaging equipment and expertise, and such procedures require traveling significant distances, making continuous monitoring difficult. It does not require prior knowledge of where the tumor is located. Furthermore, it does not require prior knowledge of the somatic mutations the tumor harbors. The costs associated with liquid biopsies are inexpensive compared to imaging and should improve as the cost of sequencing continues to decrease.

[0169] cfDNA mutations have limited signal and can be confounded by clonal hematopoiesis. DELFI captures many features of the cfDNA landscape. cfDNA fragmentation patterns are determined by the underlying chromatin organization. cfDNA fragmentation profiles are highly consistent in healthy individuals and altered in patients with cancer.

[0170] Overview of the DELFI-TF model and its application to monitoring NSCLC patients undergoing treatment: cfDNA aliquots from plasma samples of CRC patients were analyzed for RAS mutation status using ddPCR and low-pass WGS sequencing. WGS data were aligned and fragment size distributions were obtained for 504 5 Mb bins across the genome. A Bayesian regression model was trained and cross-validated on RAS MT samples using fragmentomics features. Figure 20 shows the development of the DELFI-TF model at the proof-of-concept stage.

[0171] The DELFI-TF model trained on CRC was applied to a real-world NSCLC cohort. Table 2. Clinical characteristics of the NSCLC cohort TIFF2026506641000003.tif144128

[0172] Results: Figure 13 illustrates how fragmentation profiles suggest lung cancer status and treatment response patterns.

[0173] FIG. 14 illustrates that the model-derived DELFI-TFs exhibit a strong correlation with variant allele frequencies across samples.

[0174] FIG. 15 illustrates that DELFI-TF accurately quantifies circulating tumor fraction and reflects the performance of MAF in relation to RECIST assessment.

[0175] Figure 16 illustrates that fragmentation is altered in tumor-derived cfDNA.

[0176] Figures 17A-C illustrate that DELFI-TF accurately detects circulating tumor fraction without being confounded by clonal hematopoiesis.

[0177] Figure 18 depicts that cfDNA fragmentation patterns accurately distinguish NSCLC subtypes.

[0178] Conclusions: Genome-wide cfDNA fragment profiles are abnormal in patients with cancer.

[0179] The cfDNA fragmentation score (DELFI-TF) is highly correlated with known mutant allele frequencies. cfDNA fragmentation predicts RECIST status. cfDNA fragmentation is not confounded by clonal hematopoiesis. cfDNA fragmentation signatures can noninvasively distinguish histological subtypes of lung cancer.

Claims

1. 1. A non-invasive method for subtyping small cell lung cancer in a subject as low neuroendocrine small cell lung cancer or high neuroendocrine small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic sections of sequence that map to the identified transcription factor binding sites; analyzing the genomic interval of the mapped sequence to determine the length and amount of cfDNA fragments and to use the length and amount of cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a decrease in the aggregate cfDNA fragment coverage score at the identified transcription factor binding sites indicates a high neuroendocrine small cell lung cancer subtype; A method comprising:

2. 2. The method of claim 1, wherein the identified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

3. 10. The method of claim 1, wherein the method uses machine learning to subtype small cell lung cancer in the subject.

4. 2. The method of claim 1, wherein subtyping is performed by calculating log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

5. 5. The method of claim 4, wherein the log2(read depth ratio) is the total coverage relative to a coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of 5' and 3' 500 bp anchors at both ends of the 2 kb window.

6. The method of claim 5 , wherein all coverage calculations are shifted by one to avoid division by zero.

7. 6. The method of claim 5, wherein the log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median value for each window to obtain a fragment coverage score.

8. The method of claim 1 , further comprising using a general additive model to smooth coverage and correct for GC bias.

9. The method of claim 1 , wherein the genomic intervals are non-overlapping.

10. 2. The method of claim 1, wherein the genomic intervals each comprise thousands to millions of base pairs.

11. 2. The method of claim 1, wherein a cfDNA fragmentation profile is determined within each genomic interval.

12. 12. The method of claim 11, wherein the cfDNA fragmentation profile comprises a median fragment size.

13. 12. The method of claim 11, wherein the cfDNA fragmentation profile comprises a fragment size distribution.

14. 10. The method of claim 1, further comprising administering to the subject identified as having a high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for said treatment of said cancer subtype.

15. 15. The method of claim 14, wherein the therapeutic agent is an immunotherapy.

16. 1. A non-invasive method for subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, comprising: processing cfDNA fragments from a sample obtained from the subject to generate a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic sections of sequence that map to the identified transcription factor binding sites; analyzing the genomic interval of the mapped sequence to determine the length and amount of cfDNA fragments and to use the length and amount of cfDNA fragments to establish a cfDNA fragment coverage score at the identified transcription factor binding sites; subtyping the small cell lung cancer in the subject based on transcription factor activation, wherein a cumulative decrease in cfDNA fragment coverage score at the identified transcription factor binding sites indicates an adenocarcinoma or squamous cell carcinoma subtype; A method comprising:

17. 17. The method of claim 16, wherein the identified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.

18. 17. The method of claim 16, wherein the method uses machine learning to subtype small cell lung cancer in the subject.

19. 17. The method of claim 16, wherein subtyping is performed by calculating log2(read depth ratio) over a 20 bp window, starting from a position 2 kb away from the identified transcription factor binding site.

20. 20. The method of claim 19, wherein the log2(read depth ratio) is the total coverage relative to a coverage correction factor (i.e., log2(read depth / correction factor)) calculated using the median coverage of 5' and 3' 500 bp anchors at both ends of the 2 kb window.

21. 20. The method of claim 19, wherein all coverage calculations are shifted by one to avoid division by zero.

22. 20. The method of claim 19, wherein the log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median value for each window to obtain a fragment coverage score.

23. 17. The method of claim 16, further comprising using a general additive model to smooth coverage and correct for GC bias.

24. 17. The method of claim 16, wherein the genomic intervals are non-overlapping.

25. 17. The method of claim 16, wherein the genomic intervals each comprise thousands to millions of base pairs.

26. 17. The method of claim 16, wherein a cfDNA fragmentation profile is determined within each genomic interval.

27. 27. The method of claim 26, wherein the cfDNA fragmentation profile comprises a median fragment size.

28. 27. The method of claim 26, wherein the cfDNA fragmentation profile comprises a fragment size distribution.

29. 17. The method of claim 16, further comprising administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent appropriate for said treatment of said cancer type.

30. 30. The method of claim 29, wherein the therapeutic agent is an immunotherapy.

31. 17. The method of claim 16, further comprising quantifying circulating tumor fraction, which reflects the performance of major allele frequency (MAF) associated with assessment by Response Evaluation Criteria in Solid Tumors (RECIST) in both adenocarcinoma and squamous cell carcinoma.