DELFI-derived free DNA fragmentation pattern to differentiate histological subtypes of lung cancer in non-invasive manner
By analyzing the fragmentation pattern of free DNA in plasma through low-coverage whole-genome sequencing and combining it with individual characteristics, the problem of subtyping of small cell and non-small cell lung cancer was solved, and non-invasive and rapid subtyping and personalized treatment guidance were achieved.
Patent Information
- Application Number
- CN202480016300.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-21
- Filing Date
- 2024-02-12
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies make it difficult to effectively subtype small cell lung cancer and non-small cell lung cancer in a non-invasive manner, resulting in limited treatment options and effects.
The fragmentation pattern of free DNA in plasma was analyzed by low-coverage whole-genome sequencing. Combined with the clinical and demographic characteristics of individual patients, cfDNA fragment coverage scores of designated transcription factor binding sites were used to perform subtyping of small cell lung cancer and non-small cell lung cancer.
It achieves non-invasive and rapid lung cancer subtyping, guides personalized treatment selection, and improves treatment efficacy and patient outcomes.
Smart Images

Figure CN120752348A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 USC § 119(e) to U.S. Provisional Patent Application Serial No. 63 / 445,284, filed February 13, 2023, U.S. Provisional Patent Application Serial No. 63 / 470,101, filed May 31, 2023, and U.S. Provisional Patent Application Serial No. 63 / 528,237, filed July 21, 2023. The disclosures of the prior applications are considered part of the disclosure of the present application and are herein incorporated by reference in their entirety into the disclosure of the present application. Technical Field
[0002] The present invention generally relates to non-invasive diagnosis and subtyping of small cell lung cancer (SCLC), and more specifically to analysis of genome-wide patterns of fragmented cell-free DNA (cfDNA) in conjunction with clinical and demographic characteristics of individual patients. The present invention also relates to cell-free DNA fragmentation profiling as a method for tumor score assessment and treatment monitoring in non-small cell lung cancer (NSCLC). Given the logistical difficulties of performing SCLC and NSCLC tumor biopsies, new methods are needed to achieve feasible and rapid methods for subtyping SCLC and NSCLC in a non-invasive manner. Background Art
[0003] Small cell lung cancer (SCLC) is an aggressive malignancy with a poor prognosis. Although SCLC is clinically considered a single cancer type, emerging evidence supports that subtypes of SCLC (neuroendocrine-high versus neuroendocrine-low) acquire distinct transcriptional and epigenetic states. Furthermore, different SCLC subtypes respond to specific therapies, such as immunotherapy in the case of the neuroendocrine-low, highly inflammatory SCLC subtype. Non-small cell lung cancer (NSCLC) is any type of epithelial lung cancer other than SCLC. The most common types of NSCLC are squamous cell carcinoma, large cell carcinoma, and adenocarcinoma, but several other types occur less frequently, and all may present with unusual histologic variants. As a group, NSCLC is generally less sensitive to chemotherapy and radiotherapy than SCLC. Patients with resectable disease can be cured with surgery or surgery followed by chemotherapy, or chemotherapy followed by surgery. In many patients with unresectable disease, local control can be achieved with radiotherapy, but relatively few patients are cured. Patients with locally advanced, unresectable disease may achieve long-term survival with the combination of radiotherapy and chemotherapy. Patients with advanced metastatic disease can achieve improved survival and symptom relief with chemotherapy, targeted agents, and other supportive measures. Novel, noninvasive methods are needed to identify and subtype non-small cell lung cancer and small cell lung cancer to improve patient outcomes. Summary of the Invention
[0004] The present invention is based on the groundbreaking discovery that characterizing genome-wide patterns of fragmentation of cell-free DNA (cfDNA) in plasma using low-coverage whole-genome sequencing can improve cancer diagnosis when analyzed in conjunction with certain clinical and demographic characteristics of individual patients.
[0005] In one embodiment, the present invention provides a method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequences at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of a high neuroendocrine small cell lung cancer subtype.
[0006] In one embodiment, the present invention provides a method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequences at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a reduction in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype.
[0007] In one embodiment, the present invention provides a method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the genomic interval of the mapped sequence to determine the length and amount of cfDNA fragments, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at the specified transcription factor binding site; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some aspects, the method uses a specified transcription factor is ASCL1, NEUROD1, POUF23, YAP1 or any combination thereof.
[0008] In some aspects, the method uses machine learning to subtype small cell lung cancer in a subject.
[0009] In some aspects, subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
[0010] In some aspects, the log2(read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of the 2 kb window (i.e., log2(read depth / correction factor)).
[0011] In some aspects, the method uses an offset of 1 for all coverage calculations to avoid division by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1).
[0012] In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score.
[0013] Some aspects further include using a generalized additive model to smooth coverage and correct for GC bias.
[0014] In some aspects, the genomic intervals are non-overlapping.
[0015] In some aspects, the genomic intervals each comprise thousands to millions of base pairs.
[0016] In some aspects, a cfDNA fragmentation profile is determined within each genomic interval.
[0017] In some aspects, the cfDNA fragmentation profile comprises a median fragment size.
[0018] In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution.
[0019] Some aspects further comprise administering to a subject identified as having the high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating the cancer type.
[0020] In some aspects, the therapeutic agent is an immunotherapy.
[0021] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
[0022] In one embodiment, the present invention provides a method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequences at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a reduction in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates an adenocarcinoma or squamous cell carcinoma subtype.
[0023] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
[0024] In some aspects, the designated transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.
[0025] In some aspects, the method uses machine learning to subtype small cell lung cancer in the subject.
[0026] In some aspects, subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
[0027] In some aspects, the log2(read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of the 2 kb window (i.e., log2(read depth / correction factor)).
[0028] In some aspects, all coverage calculations are offset by 1 to avoid division by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1).
[0029] In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score.
[0030] Some aspects further include using a generalized additive model to smooth coverage and correct for GC bias.
[0031] In some aspects, the genomic intervals are non-overlapping.
[0032] In some aspects, the genomic intervals each comprise thousands to millions of base pairs.
[0033] In some aspects, a cfDNA fragmentation profile is determined within each genomic interval.
[0034] In some aspects, the cfDNA fragmentation profile comprises a median fragment size.
[0035] In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution.
[0036] Some aspects further comprise administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type.
[0037] In some aspects, the therapeutic agent is an immunotherapy.
[0038] Some aspects further include quantifying a circulating tumor fraction that reflects major allele fraction (MAF) expression associated with Response Evaluation Criteria in Solid Tumors (RECIST) assessment in both adenocarcinoma and squamous cell carcinoma. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 The DELFI scores of patients diagnosed with SCLC in the study of Example 1 are shown.
[0040] Figure 2 We demonstrated that genome-wide fragmentation profiles were remarkably consistent between pre-treatment, post-treatment, progression, and response time points, suggesting that genome-wide circulating tumor DNA fragment size is a powerful method for detecting changes during immunotherapy treatment in SCLC cancer.
[0041] Figure 3 Genome-wide fragmentation profiles (and corresponding DELFI scores below) are presented, showing some differences between the two major subtypes of SCLC cases (neuroendocrine-high and neuroendocrine-low); this suggests that genome-wide circulating tumor DNA fragment size may be a powerful method for detecting subtyping during immunotherapy treatment in monitoring SCLC cancer.
[0042] Figure 4 Differentially expressed genes between high and low neuroendocrine SCLC cases were shown using publicly available data (Lissa et al. 2022) and principal component analysis (based on DELFI 30X plasma WGS data) to cluster the two subtypes. PCA analysis of the data revealed distinguishable clusters between SCLC high and low NE pre-treatment samples.
[0043] Figures 5A to 5B Feature correlation analysis is shown, revealing significant correlations between principal components 1 and 2 and the NE status of the samples.
[0044] Figures 6A to 6B We demonstrated the potential of the DELFI assay using TFBSs to target SCLC subtypes, performing genome-wide cfDNA fragmentation analysis at ASCL1 binding sites in pretreatment NCI samples. Figure 5A We show that distinct clusters of SCLC samples were observed corresponding to different levels of ASCL1 activation. Importantly, these two clusters corresponded perfectly to distinct SCLC subtypes. The primary driver of differences between samples was actually differences in coverage of fragments identified at the center of the ASCL1 binding site. Figure 6B We show that other clinical or sample characteristics have negligible influence on the clustering of these samples. Overall, these data demonstrate how DELFI analysis can perfectly distinguish SCLC subtypes using fragment coverage at TFBSs.
[0045] Figure 7 This study presents a patient example supporting the hypothesis that the detected signal originates from tumor-infiltrating lymphocytes. Patient NCI-0422 was diagnosed with SCLC with an inflammatory subtype and received durvalumab combined with olaparib. NCI-0422 responded well to treatment but was later diagnosed with disease progression. Variable genomic binding was detected before treatment and at the time of progression. Peaks were concentrated at intermingled regions of TFBS and TSS.
[0046] Figures 8A to 8B ASCL1 binding sites between responders and non-responders are shown. Figure 8A Shown are 500 cell type-specific bins from the PMD data to determine the fraction of leukocytes.
[0047] Figure 9 A machine learning model for predicting high versus low neuroendocrine SCLC was demonstrated.
[0048] Figures 10A to 10B We demonstrate that the method described here can distinguish neuroendocrine from non-neuroendocrine SCLCs by calculating read depth ratios at SCLC-specific genomic coordinates. Figure 10A ). Figure 10B It was shown that the method described herein appears to better subtype SCLC cases.
[0049] Figure 11 This figure shows that the read depth ratios at SCLC-specific genomic coordinates were transformed into a two-dimensional PCA to classify the samples into neuroendocrine and non-neuroendocrine groups. Most samples were clearly divided into different clusters.
[0050] Figure 12The results show that PCA derived from read depth ratios at SCLC-specific genomic coordinates can classify pre-treatment samples into specific SCLC subtypes (A, N, P, Y). The true picture of SCLC subtypes was calculated using tissue RNA-seq gene expression differential analysis.
[0051] Figure 13 We demonstrate how fragmentation profiles can inform lung cancer status and treatment response patterns.
[0052] Figure 14 We show that the model-derived DELFI-TFs exhibit strong correlation with mutant allele frequencies across samples.
[0053] Figure 15 We demonstrated that DELFI-TF accurately quantified circulating tumor fraction, which reflected MAF expression in association with RECIST assessment.
[0054] Figure 16 demonstrated that tumor-derived cfDNA has altered fragmentation.
[0055] Figures 17A to 17C We demonstrated that DELFI-TF accurately detected the circulating tumor fraction without interference from clonal hematopoiesis.
[0056] Figure 18 Depicted that cfDNA fragmentation patterns accurately distinguish NSCLC subtypes.
[0057] Figure 19 Depicts proof-of-concept DELFI-TF model development.
[0058] Figure 20
[0066] Figure 8 is an example computer 800 that can be used to implement the methods described herein. DETAILED DESCRIPTION
[0059] The present invention is based on the groundbreaking discovery that characterizing genome-wide patterns of fragmentation of cell-free DNA (cfDNA) in plasma using low-coverage whole-genome sequencing improves cancer diagnosis when analyzed in conjunction with certain clinical and demographic characteristics of individual patients.
[0060] The present invention describes a non-invasive method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, the method comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is approximately 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of a high neuroendocrine small cell lung cancer subtype.
[0061] The present invention describes a non-invasive method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, the method comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of a high neuroendocrine small cell lung cancer subtype.
[0062] The present invention describes a non-invasive method for subtyping small cell lung cancer in a subject as low neuroendocrine or high neuroendocrine small cell lung cancer, the method comprising: processing cfDNA fragments of a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is approximately 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype.
[0063] In some aspects, the method uses a specified transcription factor that is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.
[0064] In some aspects, the method uses machine learning to subtype small cell lung cancer in the subject.
[0065] In some aspects, subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
[0066] In some aspects, the log2(read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of the 2 kb window (i.e., log2(read depth / correction factor)).
[0067] In some aspects, the method uses an offset of 1 for all coverage calculations to avoid division by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1).
[0068] In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score.
[0069] Some aspects further include using a generalized additive model to smooth coverage and correct for GC bias.
[0070] In some aspects, the genomic intervals are non-overlapping.
[0071] In some aspects, the genomic intervals each comprise thousands to millions of base pairs.
[0072] In some aspects, a cfDNA fragmentation profile is determined within each genomic interval.
[0073] In some aspects, the cfDNA fragmentation profile comprises a median fragment size. In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution.
[0074] Some aspects further comprise administering to a subject identified as having the high neuroendocrine small cell lung cancer subtype a therapeutic agent suitable for treating the cancer type.
[0075] In some aspects, the therapeutic agent is an immunotherapy.
[0076] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequences at specified transcription factor binding sites; Analyze the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at a specified transcription factor binding site; and based on transcription factor activation, subtype the small cell lung cancer in the subject; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some aspects, the method uses a specified transcription factor that is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some aspects, subtyping is performed by calculating the log2 (read depth ratio) of a 20bp window spanning 2kb from the specified transcription factor binding site. In some aspects, the log2 (read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some aspects, the method uses an offset of 1 for all coverage calculations to avoid division by zero. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor) + 1). In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each contain thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile includes a median fragment size. In some aspects, the cfDNA fragmentation profile includes a fragment size distribution. Some aspects further include administering a therapeutic agent suitable for treating the cancer type to a subject identified as having a high neuroendocrine small cell lung cancer subtype. In some aspects, the therapeutic agent is an immunotherapy.
[0077] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequences at specified transcription factor binding sites; Analyze the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at a specified transcription factor binding site; and based on transcription factor activation, subtype the small cell lung cancer in the subject; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some aspects, the method uses a specified transcription factor that is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some aspects, subtyping is performed by calculating the log2 (read depth ratio) of a 20bp window spanning 2kb from the specified transcription factor binding site. In some aspects, the log2 (read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some aspects, the method uses an offset of 1 for all coverage calculations to avoid division by zero. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor) + 1). In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each contain thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile includes a median fragment size. In some aspects, the cfDNA fragmentation profile includes a fragment size distribution. Some aspects further include administering a therapeutic agent suitable for treating the cancer type to a subject identified as having a high neuroendocrine small cell lung cancer subtype. In some aspects, the therapeutic agent is an immunotherapy.
[0078] Described herein is a non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-low or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; Analyze the genomic intervals of the mapped sequences to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at a specified transcription factor binding site; and based on transcription factor activation, subtype the small cell lung cancer in the subject; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates a high neuroendocrine small cell lung cancer subtype. In some aspects, the method uses a specified transcription factor that is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some aspects, subtyping is performed by calculating the log2 (read depth ratio) of a 20bp window spanning 2kb from the specified transcription factor binding site. In some aspects, the log2 (read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some aspects, the method uses an offset of 1 for all coverage calculations to avoid division by zero. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor) + 1). In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each contain thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile includes a median fragment size. In some aspects, the cfDNA fragmentation profile includes a fragment size distribution. Some aspects further include administering a therapeutic agent suitable for treating the cancer type to a subject identified as having a high neuroendocrine small cell lung cancer subtype. In some aspects, the therapeutic agent is an immunotherapy.
[0079] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
[0080] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
[0081] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genome coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain genomic intervals of mapped sequence at specified transcription factor binding sites; analyzing the genomic intervals of the mapped sequence to determine cfDNA fragment length and amount, thereby establishing a cfDNA fragment coverage score at the specified transcription factor binding site using the cfDNA fragment length and amount; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
[0082] In some aspects, the designated transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.
[0083] In some aspects, the method uses machine learning to subtype small cell lung cancer in the subject.
[0084] In some aspects, subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
[0085] In some aspects, the log2(read depth ratio) is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of the 2 kb window (i.e., log2(read depth / correction factor)).
[0086] In some aspects, all coverage calculations are offset by 1 to avoid division by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1).
[0087] In some aspects, this log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score.
[0088] Some aspects further include using a generalized additive model to smooth coverage and correct for GC bias.
[0089] In some aspects, the genomic intervals are non-overlapping.
[0090] In some aspects, the genomic intervals each comprise thousands to millions of base pairs.
[0091] In some aspects, a cfDNA fragmentation profile is determined within each genomic interval.
[0092] In some aspects, the cfDNA fragmentation profile comprises a median fragment size.
[0093] In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution.
[0094] Some aspects further comprise administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type.
[0095] In some aspects, the therapeutic agent is an immunotherapy.
[0096] Some aspects further include quantifying a circulating tumor fraction that reflects major allele frequency (MAF) expression associated with Response Evaluation Criteria in Solid Tumors (RECIST) assessment in both adenocarcinoma and squamous cell carcinoma.
[0097] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genomic coverage is about 30× to 0.1×; mapping the sequenced fragments to the genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the genomic interval of the mapped sequence to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at the specified transcription factor binding site; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates an adenocarcinoma or squamous cell carcinoma subtype. In some aspects, the specified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some respects, subtyping is performed by calculating the log2 (read depth ratio) across a 20bp window of 2kb starting from the specified transcription factor binding site. In some respects, the log2 (read depth ratio) is that total coverage is divided by a coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some respects, all coverage calculations are offset by 1 to avoid being divided by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1). In some aspects, the log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each comprise thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile comprises a median fragment size. In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution. Some aspects further include administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some aspects, the therapeutic agent is an immunotherapy. Some aspects further include quantifying a circulating tumor score that reflects the major allele frequency (MAF) performance associated with Response Evaluation Criteria in Solid Tumors (RECIST) evaluation in both adenocarcinoma and squamous cell carcinoma.
[0098] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genomic coverage is about 10× to 0.1×; mapping the sequenced fragments to the genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the genomic interval of the mapped sequence to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at the specified transcription factor binding site; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates an adenocarcinoma or squamous cell carcinoma subtype. In some aspects, the specified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some respects, subtyping is performed by calculating the log2 (read depth ratio) across a 20bp window of 2kb starting from the specified transcription factor binding site. In some respects, the log2 (read depth ratio) is that total coverage is divided by a coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some respects, all coverage calculations are offset by 1 to avoid being divided by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1). In some aspects, the log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each comprise thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile comprises a median fragment size. In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution. Some aspects further include administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some aspects, the therapeutic agent is an immunotherapy. Some aspects further include quantifying a circulating tumor score that reflects the major allele frequency (MAF) performance associated with Response Evaluation Criteria in Solid Tumors (RECIST) evaluation in both adenocarcinoma and squamous cell carcinoma.
[0099] In one embodiment, the present invention provides a method for non-invasively subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; subjecting the sequencing library to whole genome sequencing to obtain sequenced fragments, wherein the genomic coverage is about 9× to 0.1×; mapping the sequenced fragments to the genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the genomic interval of the mapped sequence to determine cfDNA fragment length and amount, thereby using the cfDNA fragment length and amount to establish a cfDNA fragment coverage score at the specified transcription factor binding site; and subtyping the small cell lung cancer in the subject based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site indicates an adenocarcinoma or squamous cell carcinoma subtype. In some aspects, the specified transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof. In some aspects, the method uses machine learning to subtype the small cell lung cancer in the subject. In some respects, subtyping is performed by calculating the log2 (read depth ratio) across a 20bp window of 2kb starting from the specified transcription factor binding site. In some respects, the log2 (read depth ratio) is that total coverage is divided by a coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500bp anchor points at the ends of the 2kb window (i.e., log2 (read depth / correction factor)). In some respects, all coverage calculations are offset by 1 to avoid being divided by 0. In some aspects, this adjustment is expressed as f(x) = log2((window depth + 1) / (median(anchor point) + 1). In some aspects, the log2(read depth ratio) is calculated for each transcription factor binding site and then aggregated by calculating the median for each window to obtain a fragment coverage score. Some aspects further include using a generalized additive model to smooth the coverage and correct for GC bias. In some aspects, the genomic intervals are non-overlapping. In some aspects, the genomic intervals each comprise thousands to millions of base pairs. In some aspects, a cfDNA fragmentation profile is determined within each genomic interval. In some aspects, the cfDNA fragmentation profile comprises a median fragment size. In some aspects, the cfDNA fragmentation profile comprises a fragment size distribution. Some aspects further include administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type. In some aspects, the therapeutic agent is an immunotherapy. Some aspects further include quantifying a circulating tumor score that reflects the major allele frequency (MAF) performance associated with Response Evaluation Criteria in Solid Tumors (RECIST) evaluation in both adenocarcinoma and squamous cell carcinoma.
[0100] Before describing the compositions and methods of the present invention, it should be understood that the present invention is not limited to the particular methods and systems described, as such methods and systems may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, as the scope of the present invention is limited only by the appended claims.
[0101] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "the method" includes one or more methods and / or steps of the type described herein that will become apparent to one skilled in the art upon reading this disclosure and so forth.
[0102] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, the preferred methods and materials are now described.
[0103] This article describes a method for noninvasively subtyping non-small cell lung cancer (NSCLC) using cell-free DNA (cfDNA) fragmentomes. NSCLC typically presents as distinct subtypes, such as adenocarcinoma and squamous cell carcinoma. This novel, noninvasive DELFI-based approach distinguishes adenocarcinoma from squamous cell carcinoma by analyzing cell-free fragments that reflect the unique epigenetic state of NSCLC. Other cfDNA-based liquid biopsies cannot distinguish NSCLC subtypes without tissue clinical data.
[0104] NSCLC subtypes can be predicted using a DELFI-based approach using cfDNA LC-WGS data. First, the DELFI machine learning classifier detected the presence or absence of cancer in NSCLC patients. Both the whole-genome fragmentation profile and the corresponding DELFI-tumor fraction (TF) score were studied to determine preliminary differences between the two major clinical subtypes of NSCLC cases: adenocarcinoma and squamous cell carcinoma. DELFI-TF accurately detected circulating tumor fraction without being confounded by clonal hematopoiesis. Furthermore, DELFI-TF accurately quantified circulating tumor fraction, which reflects MAF performance associated with RECIST assessment in both adenocarcinoma and squamous cell carcinoma. Publicly available data were used to perform principal component analysis to cluster the two subtypes using whole-genome fragmentation data. Figure 19The development of a proof-of-concept DELFI-TF model is described. Second, a list of differentially accessible transcription start sites (TSSs) was applied and was able to distinguish adenocarcinomas from squamous cell carcinomas. Lists of differentially accessible transcription start sites can be obtained from public databases such as, but not limited to, UCSC or Ensembl. The creation of such lists is similar to that done for RNA-Seq experiments, except that read depth ratios are used instead of transcripts per million as a surrogate for expression. In some aspects, the loci with the largest and most significant log-fold changes were selected from 20% of the test cohort and applied to the remaining cohort at the same loci, followed by hierarchical clustering to determine whether subtypes remained together. In another aspect, the most highly expressed transcripts from TCGA for a given cancer type were selected, and those transcripts that were also expressed at any level in AML (a blood cancer, serving as an imperfect surrogate for normal blood) were filtered out from the list. Third, the most differential DELFI-TSSs were investigated to see whether they exhibited short / long variations, further confirming that these cfDNA molecules were of tumor origin. Finally, receiver operating characteristics (ROCs) representing the sensitivity and specificity of the DELFI-fragment panel approach for identifying NSCLC subtypes were used to determine the diagnostic accuracy of our approach across the identified cluster samples. Given the logistical difficulties of performing NSCLC tumor biopsies, we believe this approach may be a feasible and rapid method for subtyping NSCLC in a non-invasive manner.
[0105] This document provides methods and materials for determining the cfDNA fragmentation profile in a mammal (e.g., in a sample obtained from the mammal). As used herein, the terms "fragmentation profile," "position-dependent differences in the fragmentation pattern," and "differences in fragment size and coverage across the genome in a position-dependent manner" are equivalent and can be used interchangeably. In some cases, determining the cfDNA fragmentation profile in a mammal can be used to identify the mammal as having cancer. For example, low-coverage whole-genome sequencing can be performed on cfDNA fragments obtained from a mammal (e.g., a sample obtained from the mammal), and the sequenced fragments can be mapped to the genome (e.g., in non-overlapping windows) and assessed to determine the cfDNA fragmentation profile. As described herein, the cfDNA fragmentation profile of a mammal with cancer is more heterogeneous (e.g., fragment lengths are more heterogeneous) than the cfDNA fragmentation profile of a healthy mammal (e.g., a mammal without cancer). Therefore, this document also provides methods and materials for assessing, monitoring, and / or treating a mammal (e.g., a human) having or suspected of having cancer. In some cases, this document provides methods and materials for identifying whether a mammal has cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine the presence or absence of cancer in the mammal and, optionally, the tissue of origin, based at least in part on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for monitoring whether a mammal has cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine the presence or absence of cancer in the mammal based at least in part on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for identifying whether a mammal has cancer and administering one or more cancer treatments to the mammal to treat the mammal. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine whether a mammal has cancer based at least in part on the mammal's cfDNA fragmentation profile, and one or more cancer treatments can be administered to the mammal.
[0106] The cfDNA fragmentation profile may include one or more cfDNA fragmentation patterns. The cfDNA fragmentation pattern may include any appropriate cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, but are not limited to, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and coverage of cfDNA fragments. In some cases, the cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) median fragment sizes, fragment size distributions, ratio of small cfDNA fragments to large cfDNA fragments, and coverage of cfDNA fragments. In some cases, the cfDNA fragmentation profile may be a whole-genome cfDNA profile (e.g., a whole-genome cfDNA profile in a window across the entire genome). In some cases, the cfDNA fragmentation profile may be a targeted region profile. The targeted region may be any appropriate portion of the genome (e.g., a chromosomal region). Examples of chromosomal regions for which cfDNA fragmentation profiles can be determined as described herein include, but are not limited to, a portion of a chromosome (e.g., a portion of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and / or 14q) and a chromosome arm (e.g., a chromosome arm of 8q, 13q, 11q, and / or 3p). In some cases, a cfDNA fragmentation profile can include two or more targeted region profiles.
[0107] In some cases, cfDNA fragmentation profiles can be used to identify changes (e.g., alterations) in cfDNA fragment length. Alterations can be genome-wide or in one or more targeted regions / locus. Target regions can be any region containing one or more cancer-specific alterations. Examples of cancer-specific alterations and their chromosomal locations include, but are not limited to, those shown in Table 3 (Appendix C) and those shown in Table 6 (Appendix F). In some cases, a cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) about 10 alterations to about 500 alterations (e.g., about 25 to about 500, about 50 to about 500, about 100 to about 500, about 200 to about 500, about 300 to about 500, about 10 to about 400, about 10 to about 300, about 10 to about 200, about 10 to about 100, about 10 to about 50, about 20 to about 400, about 30 to about 300, about 40 to about 200, about 50 to about 100, about 20 to about 100, about 25 to about 75, about 50 to 250, or about 100 to about 200 alterations).
[0108] In some cases, a cfDNA fragmentation profile can be used to detect tumor-derived DNA. For example, a cfDNA fragmentation profile can be used to detect tumor-derived DNA by comparing the cfDNA fragmentation profile of a mammal having or suspected of having cancer to a reference cfDNA fragmentation profile (e.g., a cfDNA fragmentation profile of a healthy mammal and / or a nucleosomal DNA fragmentation profile of healthy cells from a mammal having or suspected of having cancer). In some cases, the reference cfDNA fragmentation profile is a previously generated profile from a healthy mammal. For example, the methods provided herein can be used to determine a reference cfDNA fragmentation profile in a healthy mammal, and this reference cfDNA fragmentation profile can be stored (e.g., on a computer or other electronic storage medium) for future comparison with a test cfDNA fragmentation profile in a mammal having or suspected of having cancer. In some cases, the reference cfDNA fragmentation profile of a healthy mammal (e.g., a stored cfDNA fragmentation profile) is determined genome-wide. In some cases, the reference cfDNA fragmentation profile of a healthy mammal (e.g., a stored cfDNA fragmentation profile) is determined at a subgenomic interval scale.
[0109] In some cases, the cfDNA fragmentation profile can be used to identify whether a mammal (e.g., a human) has cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer).
[0110] A cfDNA fragmentation profile can include a cfDNA fragment size profile. cfDNA fragments can be of any suitable size. For example, cfDNA fragments can range in length from about 50 base pairs (bp) to about 400 bp. As described herein, a cfDNA fragment size profile for a mammal with cancer can contain a median cfDNA fragment size that is shorter than the median cfDNA fragment size for a healthy mammal. A healthy mammal (e.g., a mammal without cancer) can have a median cfDNA fragment size of about 166.6 bp to about 167.2 bp (e.g., about 166.9 bp). In some cases, a mammal with cancer can have cfDNA fragment sizes that are, on average, about 1.28 bp to about 2.49 bp (e.g., about 1.88 bp) shorter than the cfDNA fragment sizes for a healthy mammal. For example, a mammal with cancer can have a median cfDNA fragment size of about 164.11 bp to about 165.92 bp (e.g., about 165.02 bp).
[0111] A cfDNA fragmentation profile can include a cfDNA fragment size distribution. As described herein, a mammal with cancer can have a cfDNA size distribution that is more variable than the cfDNA fragment size distribution of a healthy mammal. In some cases, the size distribution can be within a targeted region. A healthy mammal (e.g., a mammal without cancer) can have a targeted region cfDNA fragment size distribution of about 1 or less than about 1. In some cases, a mammal with cancer can have a targeted region cfDNA fragment size distribution that is longer (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp, or any number of base pairs in between) than the targeted region cfDNA fragment size distribution of a healthy mammal. In some cases, a mammal with cancer can have a targeted region cfDNA fragment size distribution that is shorter (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50 or more bp, or any number of base pairs in between) than the targeted region cfDNA fragment size distribution of a healthy mammal. In some cases, a mammal with cancer may have a targeted region cfDNA fragment size distribution that is approximately 47 bp smaller and approximately 30 bp longer than the targeted region cfDNA fragment size distribution of a healthy mammal. In some cases, a mammal with cancer may have a targeted region cfDNA fragment size distribution in which the lengths of the cfDNA fragments differ by an average of 10, 11, 12, 13, 14, 15, 15, 17, 18, 19, 20, or more bp. For example, a mammal with cancer may have a targeted region cfDNA fragment size distribution in which the lengths of the cfDNA fragments differ by an average of approximately 13 bp. In some cases, the size distribution may be a genome-wide size distribution. A healthy mammal (e.g., a mammal without cancer) may have a very similar distribution of short and long cfDNA fragments across the genome. In some cases, a mammal with cancer may have one or more alterations (e.g., increases and decreases) in cfDNA fragment size across the genome. The one or more alterations may be in any suitable chromosomal region of the genome. For example, the alteration may be in a portion of a chromosome. Examples of chromosome portions that may contain one or more changes in cfDNA fragment size include, but are not limited to, portions of 2q, 4p, 5p, 6q, 7p, 8q, 9q, 10q, 11q, 12q, and 14q. For example, the change may be across an entire chromosome arm (e.g., the entire chromosome arm).
[0112] A cfDNA fragmentation profile can include a ratio of small cfDNA fragments to large cfDNA fragments and a correlation of the fragment ratios to a reference fragment ratio. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the small cfDNA fragments can be between about 100 bp and about 150 bp. As used herein, with respect to the ratio of small cfDNA fragments to large cfDNA fragments, the length of the large cfDNA fragments can be between about 151 bp and 220 bp. As described herein, a mammal with cancer can have a fragment ratio correlation (e.g., a correlation of the cfDNA fragment ratio to a reference DNA fragment ratio, such as a DNA fragment ratio from one or more healthy mammals) that is lower (e.g., 2-fold lower, 3-fold lower, 4-fold lower, 5-fold lower, 6-fold lower, 7-fold lower, 8-fold lower, 9-fold lower, 10-fold lower, or more) than a healthy mammal. A healthy mammal (e.g., a mammal without cancer) can have a fragment ratio correlation (e.g., a correlation of the cfDNA fragment ratio to a reference DNA fragment ratio, such as a DNA fragment ratio from one or more healthy mammals) of approximately 1 (e.g., approximately 0.96). In some cases, a mammal having cancer can have a fragmentation ratio correlation (e.g., correlation of cfDNA fragmentation ratio with a reference DNA fragmentation ratio, such as DNA fragmentation ratio from one or more healthy mammals) that is, on average, about 0.19 to about 0.30 (e.g., about 0.25) lower than the correlation of fragmentation ratios of healthy mammals (e.g., correlation of cfDNA fragmentation ratio with a reference DNA fragmentation ratio, such as DNA fragmentation ratio from one or more healthy mammals).
[0113] The cfDNA fragmentation profile may include coverage of all fragments. Coverage of all fragments may include windows of coverage (e.g., non-overlapping windows). In some cases, coverage of all fragments may include windows of small fragments (e.g., fragments with a length of about 100 bp to about 150 bp). In some cases, coverage of all fragments may include windows of large fragments (e.g., fragments with a length of about 151 bp to about 220 bp).
[0114] Any suitable method can be used to obtain a cfDNA fragmentation profile. In some cases, cfDNA from a mammal (e.g., a mammal having or suspected of having cancer) can be processed into a sequencing library that can be subjected to whole genome sequencing (e.g., low coverage whole genome sequencing), mapped to the genome, and analyzed to determine the cfDNA fragment lengths. The mapped sequences can be analyzed in non-overlapping windows covering the genome. The windows can be of any suitable size. For example, the length of the window can be thousands to millions of bases. As a non-limiting example, the window can be about 5 megabases (Mb) long. Any appropriate number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. The cfDNA fragmentation profile can be determined within each window. In some cases, a cfDNA fragmentation profile can be obtained as described in Example 1. In some cases, a cfDNA fragmentation profile can be obtained as described in Example 1. Figure 1 The cfDNA fragmentation profile was obtained as shown.
[0115] In some cases, the methods and materials described herein can also include machine learning. For example, machine learning can be used to identify altered fragmentation patterns (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, and mtDNA).
[0116] In some cases, the methods and materials described herein can be the sole method for identifying whether a mammal (e.g., a human) has cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). For example, determining the cfDNA fragmentation profile can be the sole method for identifying whether a mammal has cancer.
[0117] In some cases, the methods and materials described herein can be used in conjunction with one or more additional methods for identifying a mammal (e.g., a human) with cancer (e.g., colorectal cancer, lung cancer, breast cancer, gastric cancer, pancreatic cancer, bile duct cancer, and / or ovarian cancer). Examples of methods for identifying whether a mammal has cancer include, but are not limited to, identifying one or more cancer-specific sequence changes, identifying one or more chromosome changes (e.g., aneuploidy and rearrangement), and identifying other cfDNA changes. For example, determining a cfDNA fragmentation profile can be used in conjunction with identifying one or more cancer-specific mutations in a mammal's genome to identify whether a mammal has cancer. For example, determining a cfDNA fragmentation profile can be used in conjunction with identifying one or more aneuploidies in a mammal's genome to identify whether a mammal has cancer.
[0118] In some aspects, this document also provides methods and materials for assessing, monitoring, and / or treating mammals (e.g., humans) that have or are suspected of having cancer. In some cases, this document provides methods and materials for identifying whether a mammal has cancer. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine whether the mammal has cancer based, at least in part, on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for identifying the location (e.g., anatomical site or tissue of origin) of a cancer in a mammal. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine the tissue of origin of a cancer in a mammal based, at least in part, on the mammal's cfDNA fragmentation profile. In some cases, this document provides methods and materials for identifying whether a mammal has cancer and administering one or more cancer therapies to the mammal to treat the mammal. For example, a sample obtained from a mammal (e.g., a blood sample) can be assessed to determine whether the mammal has cancer based, at least in part, on the mammal's cfDNA fragmentation profile, and administering one or more cancer therapies to the mammal. In some cases, this document provides methods and materials for treating a mammal with cancer. For example, one or more cancer treatments (e.g., based at least in part on the mammal's cfDNA fragmentation profile) can be administered to a mammal identified as having cancer to treat the mammal. In some cases, during or after the course of cancer treatment (e.g., any cancer treatment described herein), the mammal can undergo monitoring (or be selected for enhanced monitoring) and / or further diagnostic testing. In some cases, monitoring can include assessing a mammal having or suspected of having cancer, for example, by assessing a sample (e.g., a blood sample) obtained from the mammal to determine the mammal's cfDNA fragmentation profile as described herein, and changes in the cfDNA fragmentation profile over time can be used to identify response to treatment and / or identify whether the mammal has cancer (e.g., residual cancer).
[0119] Any suitable mammal can be assessed, monitored, and / or treated as described herein. The mammal can be a mammal suffering from cancer. The mammal can be a mammal suspected of suffering from cancer. Examples of mammals that can be assessed, monitored, and / or treated as described herein include, but are not limited to, humans, primates (such as monkeys), dogs, cats, horses, cows, pigs, sheep, mice, and rats. For example, a human suffering from or suspected of suffering from cancer can be assessed to determine a cfDNA fragmentation profile as described herein, and optionally treated with one or more cancer treatment methods as described herein.
[0120] Any suitable sample from a mammal can be assessed as described herein (e.g., to assess DNA fragmentation patterns). In some cases, the sample can comprise DNA (e.g., genomic DNA). In some cases, the sample can comprise cfDNA (e.g., circulating tumor DNA (ctDNA)). In some cases, the sample can be a fluid sample (e.g., a liquid biopsy). Examples of samples that can contain DNA and / or polypeptides include, but are not limited to, blood (e.g., whole blood, serum, or plasma), amniotic fluid, tissue, urine, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, bile, lymph fluid, cyst fluid, stool, ascites, Pap smear, breast milk, and breath condensate. For example, a plasma sample can be assessed to determine a cfDNA fragmentation pattern as described herein.
[0121] A sample from a mammal to be assessed (e.g., assessed for DNA fragmentation pattern) as described herein can contain any suitable amount of cfDNA. In some cases, the sample can contain a limited amount of DNA. For example, a cfDNA fragmentation profile can be obtained from a sample containing less DNA than is typically required for other cfDNA analysis methods, such as those described in, for example, Phallen et al., 2017 Sci Transl Med 9; Cohen et al., 2018 Science 359:926; Newman et al., 2014 Nat Med 20:548; and Newman et al., 2016 Nat Biotechnol 34:547).
[0122] In some cases, the sample can be treated (e.g., to isolate and / or purify DNA and / or polypeptides from the sample). For example, DNA isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), protein removal (e.g., using proteases), and / or RNA removal (e.g., using RNases). As another example, polypeptide isolation and / or purification can include cell lysis (e.g., using detergents and / or surfactants), DNA removal (e.g., using DNases), and / or RNA removal (e.g., using RNases).
[0123] Additional methods are described in US Patent Nos. 10,982,279 and 10,975,431, the disclosures of which are considered part of the disclosure of the present application and are incorporated herein by reference in their entirety.
[0124] Example Hardware Implementation Figure 20An example computer 800 that can be used to implement the methods described herein is shown. For example, in some embodiments, the computer 800 may include a machine learning system that trains a machine learning model to subtype small cell lung cancer or non-small cell lung cancer, or a portion or combination thereof, as described above. The computer 800 can be any electronic device that runs a software application program obtained from compiled instructions, including but not limited to a personal computer, a server, a smart phone, a media player, an electronic tablet, a game console, an email device, etc. In some embodiments, the computer 800 may include one or more processors 802, one or more input devices 804, one or more display devices 806, one or more network interfaces 808, and one or more computer-readable media 812. Each of these components can be coupled by a bus 810, and in some embodiments, these components can be distributed between multiple physical locations and coupled by a network.
[0125] The display device 806 can be any known display technology, including but not limited to a display device using a liquid crystal display (LCD) or light emitting diode (LED) technology. The processor 802 can use any known processor technology, including but not limited to a graphics processor and a multi-core processor. The input device 804 can be any known input device technology, including but not limited to a keyboard (including a virtual keyboard), a mouse, a trackball, a camera, and a touch-sensitive pad or display. The bus 810 can be any known internal or external bus technology, including but not limited to ISA, EISA, PCI, PCI Express, USB, serial ATA, or FireWire. The computer-readable medium 812 can be any non-temporary medium that participates in providing instructions to the processor 804 for execution, including but not limited to non-volatile storage media (e.g., optical disks, magnetic disks, flash drives, etc.) or volatile media (e.g., SDRAM, ROM, etc.).
[0126] Computer-readable media 812 may include various instructions 814 for implementing an operating system (e.g., Mac OS®, Windows®, Linux). The operating system may be a multi-user, multi-processing, multi-tasking, multi-threaded, real-time operating system, etc. The operating system may perform basic tasks, including, but not limited to: recognizing input from input device 804; sending output to display device 806; keeping track of files and directories on computer-readable media 812; controlling peripheral devices (e.g., disk drives, printers, etc.) that may be controlled directly or through an I / O controller; and managing traffic on bus 810. Network communication instructions 816 may establish and maintain network connections (e.g., software for implementing communication protocols such as TCP / IP, HTTP, Ethernet, telephony, etc.).
[0127] Machine learning instructions 818 may include instructions that enable computer 800 to function as a machine learning system and / or train a machine learning model to generate DMS values as described herein. Application 820 may be an application that uses or implements the processes and / or other processes described herein. The processes may also be implemented in operating system 814. For example, application 820 and / or the operating system may create tasks in the applications described herein.
[0128] The described features can be implemented in one or more computer programs that are executable on a programmable system that includes at least one programmable processor coupled to receive data and instructions from and to send data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written in any programming language, including compiled or interpreted languages (e.g., Objective-C, Java), and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0129] For example, processors suitable for executing a program of instructions include general-purpose and special-purpose microprocessors, as well as the sole processor or one of multiple processors or cores of any type of computer. Typically, a processor may receive instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer may include a processor for executing instructions and one or more memories for storing instructions and data. Typically, a computer may also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data may include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into an ASIC (Application Specific Integrated Circuit).
[0130] To provide for interaction with a user, features may be implemented on a computer having a display device (such as an LED or LCD monitor) for displaying information to the user and a keyboard and pointing device (such as a mouse or trackball) through which the user can provide input to the computer.
[0131] Features can be implemented on a computer system that includes back-end components such as a data server, or middleware components such as an application server or an Internet server, or front-end components such as a client computer with a graphical user interface or an Internet browser, or any combination thereof. The components of the system can be connected by any form of digital data communication (such as a communication network) or medium for digital data communication. Examples of communication networks include, for example, telephone networks, LANs, WANs, and computers and the network that forms the Internet.
[0132] Computer systems can include clients and servers. Clients and servers can generally be remote from each other and can typically interact through a network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0133] One or more features or steps of the disclosed embodiments can be implemented using an application programming interface (API). The API can define one or more parameters that are passed between a calling application and other software code (e.g., an operating system, a library routine, a function) that provides a service, provides data, or performs an operation or calculation.
[0134] An API can be implemented as one or more calls in program code that send or receive one or more parameters via a parameter list or other structure based on a calling convention defined in an API specification document. A parameter can be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters can be implemented in any programming language. A programming language can define the vocabulary and calling conventions that programmers will use to access functions that support the API.
[0135] In some implementations, the API call may report to the application the capabilities of the device on which the application is running, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, and the like.
[0136] Although various embodiments have been described above, it should be understood that the embodiments are presented by way of example and not limitation. It will be apparent to those skilled in the relevant art that various changes in form and detail may be made thereto without departing from the spirit and scope. In fact, after reading the above description, it will be apparent to those skilled in the relevant art how to implement alternative embodiments. For example, additional steps may be provided, or steps may be eliminated from the described process, and additional components may be added to or removed from the described system. Therefore, other specific implementations are within the scope of the following claims.
[0137] In addition, it should be understood that any drawings highlighting features and advantages are presented for illustrative purposes only. The disclosed methods and systems are each sufficiently flexible and configurable such that they can be utilized in ways other than those shown.
[0138] Although the term "at least one" may often be used in the specification, claims, and drawings, the terms "a," "an," "the," "said," etc. also mean "at least one" or "the at least one" in the specification, claims, and drawings.
[0139] Finally, it is Applicant's intention that only claims that include the express language "means for" or "step for" will be construed under 35 USC 112(f). Claims that do not expressly include the phrase "means for" or "step for" will not be construed under 35 USC 112(f).
[0140] The presently described methods and systems can be used to subtype non-small cell lung cancer or small cell lung cancer in a subject, and optionally treat a subtype of cancer in the subject. Any suitable subject, such as a mammal, can be assessed and / or treated as described herein. Examples of some mammals that can be assessed and / or treated as described herein include, but are not limited to, humans, primates (such as monkeys), dogs, cats, horses, cattle, pigs, sheep, mice, and rats. For example, a person having or suspected of having cancer can be assessed using the methods described herein, and optionally can be treated with one or more cancer treatments as described herein. The methods disclosed herein can include administering a therapeutic agent suitable for treating the cancer type to a subject identified as having the cancer type.
[0141] When treating a subject suffering from or suspected of having a cancer as described herein, one or more cancer treatments can be administered to the subject. The cancer treatment can be any appropriate cancer treatment. One or more cancer treatments described herein can be administered to the subject at any appropriate frequency (e.g., once or multiple times over a period of days to weeks). Examples of cancer treatments include, but are not limited to, surgical intervention, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy (e.g., chimeric antigen receptors and / or T cells with wild-type or modified T cell receptors), targeted therapy, such as administration of kinase inhibitors (e.g., kinase inhibitors targeting specific genetic damage such as translocations or mutations) (e.g., kinase inhibitors, antibodies, bispecific antibodies), signal transduction inhibitors, bispecific antibodies or antibody fragments (e.g., BiTEs), monoclonal antibodies, immune checkpoint inhibitors, surgery (e.g., surgical resection), or any combination thereof. In some aspects, cancer treatment can reduce the severity of cancer, alleviate the symptoms of cancer, and / or reduce the number of cancer cells present in the subject.
[0142] In some aspects, the cancer treatment can be a chemotherapeutic agent. Non-limiting examples of chemotherapeutic agents include amsacrine, azacitidine, axathioprine, bevacizumab (or an antigen-binding fragment thereof), bleomycin, busulfan, carboplatin, capecitabine, chlorambucil, cisplatin, cyclophosphamide, cytarabine, dacarbazine, daunorubicin, docetaxel, doxifluridine, doxorubicin, epirubicin, erlotinib, dapoxetine ... hydrochloride), etoposide, fiudarabine, floxuridine, fludarabine, fluorouracil, gemcitabine, hydroxyurea, idarubicin, ifosfamide, irinotecan, lomustine, mechlorethamine, melphalan, mercaptopurine, methotrexate, mitomycin, mitoxantrone
[00145] The present invention relates to oxantrone, oxaliplatin, paclitaxel, pemetrexed, procarbazine, all-trans retinoic acid, streptozocin, tafluposide, temozolomide, teniposide, tioguanine, topotecan, uramustine, valrubicin, vinblastine, vincristine, vindesine, vinorelbine, and combinations thereof.Additional examples of anti-cancer therapies are known in the art; see, for example, therapy guidelines from the American Society of Clinical Oncology (ASCO), the European Society for Medical Oncology (ESMO), or the National Comprehensive Cancer Network (NCCN).
[0143] In various aspects, DNA is present in a biological sample taken from a subject and used in the methods of the present invention. Biological sample can actually be any type of biological sample including DNA. Biological sample is typically a fluid, such as whole blood or a portion thereof with circulating cfDNA. In an embodiment, the sample includes DNA from a tumor or liquid biopsy, the tumor or liquid biopsy such as but not limited to amniotic fluid, aqueous humor, vitreous humor, blood, whole blood, fractionated blood, plasma, serum, breast milk, cerebrospinal fluid (CSF), cerumen (earwax), chyle, corns, endolymph, perilymph, feces, breath, gastric acid, gastric juice, lymph, mucus (including nasal drainage and mucus), pericardial fluid, peritoneal fluid, pleural fluid, pus, nasal discharge, saliva, exhaled breath condensate, sebum, semen, sputum, sweat, synovial fluid, tears, vomitus, prostatic fluid, nipple aspirate, tears, sweat, cheek swab, cell lysate, gastrointestinal fluid, biopsy tissue and urine or other biological fluids. In one aspect, the sample comprises DNA from circulating tumor cells.
[0144] As disclosed above, the biological sample can be a blood sample. The blood sample can be obtained using methods known in the art, such as finger puncture or phlebotomy. Suitably, the blood sample is about 0.1 to 20 ml, or alternatively about 1 to 15 ml, wherein the volume of blood is about 10 ml. Smaller amounts can also be used, as well as circulating free DNA in the blood. Microsampling and sampling by needle biopsy, catheter, excretion or generation of body fluids containing DNA are also potential sources of biological samples.
[0145] The methods and systems of the present disclosure utilize nucleic acid sequence information and thus may include any method or sequencing apparatus for performing nucleic acid sequencing, including nucleic acid amplification, polymerase chain reaction (PCR), nanopore sequencing, 454 sequencing, and insertion tag sequencing. In some aspects, the methods or systems of the present disclosure utilize systems such as those provided by Illumina, Inc. (including but not limited to HiSeq TM X10, HiSeq TM 1000, HiSeq TM 2000, HiSeq TM 2500, Genome Analyzers TM , MiSeq TM, NextSeq, NovaSeq 6000 systems), systems provided by Applied Biosystems Life Technologies (SOLiD TM System, Ion PGM TM Sequencer, Ion Proton TM Sequencer) or systems provided by Genapsys or BGI MGI, among others. Nucleic acid analysis can also be performed using systems provided by Oxford Nanopore Technologies (GridiON TM 、MiniON TM ) or Pacific Biosciences (Pacbio TM RS II or Sequel I or II).
[0146] The present invention includes systems for performing the steps of the disclosed methods and is described in part in terms of functional components and various processing steps. Such functional components and processing steps can be implemented by any number of components, operations, and techniques configured to perform the specified functions and achieve various results. For example, the present invention can employ various biological samples, biomarkers, components, materials, computers, data sources, storage systems and media, information collection techniques and processes, data processing standards, statistical analysis, regression analysis, etc. that can perform various functions.
[0147] Therefore, the present invention further provides a non-invasive system for subtyping small cell lung cancer or non-small cell lung cancer. In various aspects, the system includes: (a) a sequencer configured to generate a low-coverage whole-genome sequencing dataset of a sample; and (b) a computer system and / or processor having the function of executing the method of the present invention.
[0148] In some aspects, the computer system further comprises one or more additional modules. For example, the system may include one or more extraction units and / or separation units operable to select an appropriate genetic component for analysis (e.g., cfDNA fragments of a specific size).
[0149] In some aspects, the computer system further comprises a visual display device. The visual display device can be operable to display the curve fit line, the reference curve fit line, and / or a comparison of the two.
[0150] The method for non-invasive subtyping of small cell lung cancer or non-small cell lung cancer according to various aspects of the present invention can be implemented in any suitable manner, for example, using a computer program running on a computer system. As discussed herein, the exemplary systems according to various aspects of the present invention can be implemented in conjunction with a computer system, which is, for example, a conventional computer system comprising a processor and random access memory, such as a remotely accessible application server, network server, personal computer or workstation. The computer system also suitably includes another memory device or information storage system, such as a mass storage system and a user interface, such as a conventional monitor, keyboard and tracking device. However, the computer system may include any suitable computer system and associated equipment, and may be configured in any suitable manner. In one embodiment, the computer system includes a stand-alone system. In another embodiment, the computer system is part of a computer network including a server and a database.
[0151] Can be implemented in single device or in multiple devices, implement the software that needs to be used for receiving, processing and analyzing information.This software can be accessed through a network so that the storage and processing of information occur remotely with respect to the user.System according to various aspects of the present invention and its various elements provide the function and operation that are convenient to detect and / or analyze, such as data collection, processing, analysis, reporting and / or diagnosis.For example, in this aspect, computer system execution computer program, described computer program can receive, store, search, analyze and report the information relevant to human genome or its region.Computer program can comprise multiple modules that perform various functions or operations, such as the processing module that is used to process raw data and generate supplementary data and the analysis module that is used to analyze raw data and supplementary data to generate the quantitative assessment of disease state model and / or diagnostic information.
[0152] The procedures performed by the system may include any suitable process that facilitates analysis and / or subtyping of small cell lung cancer or non-small cell lung cancer. In one embodiment, the system is configured to model disease subtypes and / or determine a patient's disease subtype. Determining or identifying a disease subtype may include generating any useful information about the patient's condition relative to the disease, such as performing a diagnosis, providing information that aids in diagnosis, assessing the stage or progression of the disease, identifying conditions that may indicate susceptibility to the disease, identifying whether further testing would be recommended, predicting and / or assessing the efficacy of one or more treatment programs, or otherwise assessing the patient's disease status, likelihood of disease, or other health aspects.
[0153] The following examples are provided to further illustrate the advantages and features of the present invention, but are not intended to limit the scope of the present invention. Although this example is typical in that it can be used, other procedures, methods or techniques known to those skilled in the art can also be used alternatively.
[0154] Example 1 Dissecting small cell lung cancer subtypes using cell-free DNA fragmentomes Background: Small cell lung cancer (SCLC) is an aggressive form of lung cancer that is closely associated with smoking and exposure to other environmental chemicals. The 5-year survival rate of patients with SCLC is less than 8% (Gay, CM et al. Transcription factor programs and immune pathway activation patterns define four major SCLC subtypes with distinct therapeutic vulnerabilities. Cancer Cell 39, 346-360.e7(2021). SCLC is a heterogeneous tumor type composed of tumor cells with neuroendocrine and non-neuroendocrine characteristics. There are two main subtypes of SCLC: high-grade neuroendocrine (High-NE) and low-grade neuroendocrine (Low-NE). et al. The high NE subtype is characterized by lineage-specific transcription factors ( ASCL1 and NEUROD1 ) activation. Low NE is characterized by non-neuroendocrine factors such as POU2F3 Despite this molecular and clinical heterogeneity, SCLC is still treated as a single entity with predictably poor outcomes. Recent results suggest that the inflammatory group (SCLC-I) has unique biological properties and responds better to immunotherapy than other SCLC subtypes (Gay et al., 2011). et al. and Lissa, D. et al. Heterogeneity of neuroendocrine transcriptional states in metastatic small cell lung cancer and its patient-derived models. Nat Commun 13, 2023 (2022). The DELFI-based free fragment panel accurately distinguishes high-NE and low-NE SCLC subtypes in a noninvasive manner. The ultimate goal of DELFI SCLC subtyping is to provide optimal treatment options for every patient diagnosed with advanced SCLC.
[0155] Methods: In a phase II trial (NCT02484404), circulating cell-free DNA (cfDNA) was isolated from plasma samples of patients diagnosed with recurrent SCLC who received durvalumab plus olaparib. SCLC subtypes were stratified into high-negative-sense (NE) (n=10) and low-negative-sense (n=5) using pretreatment tissue biopsy immunohistochemistry and genomic data. To infer tumor gene expression profiles from cfDNA, we investigated genome-wide signals of differentially regulated tissue-specific transcription factors in SCLC and applied a novel DELFI-based approach to inform SCLC molecular subtypes. Specifically, DELFI computes the fragment distribution of each TFBS within a 2kb window and independently scales it between 0 and 1 for each sample. Principal component analysis was then performed in R, and clusters were defined based on ASCL1 TFBS loci (Mathios, D. et al. Detection and characterization of lung cancer using cell-free DNA fragmentation panels. Nat Commun 12, 5060 (2021)). Clinical information was examined in an orthogonal analysis.
[0156] Results: DELFI's proprietary fragmentomics platform detected SCLC cases (N=47) with high sensitivity, with a median DELFI score of 1.0 (95% CI 0.99-1), consistent with previous data from DELFI's in-house SCLC samples. Genome-wide cfDNA fragmentation analysis of ASCL1 binding sites (approximately 12,000 genomic coordinates) in SCLC patients revealed reduced coverage near transcription factor binding sites in SCLC patients compared with non-cancer and NSCLC samples. Differences in the center of ASCL1 binding sites were also the primary driver for distinguishing pre-treatment SCLC samples. Studying the fragment distribution of ASCL1 binding sites enabled differentiation between high- and low-NE samples. Patients classified as low-NE by the DELFI subtyping classifier (who received I / O therapy) had a better treatment response than those classified as high-NE by DELFI. Differential pseudogene expression identified thousands of highly enriched transcription start sites (TSSs) in low-NE cases compared with high-NE cases. Low NE samples showed high enrichment of the T lymphocyte-specific factor KLF4, reflecting the neutrophil-to-lymphocyte ratio of this SCLC subtype. These data reflect an inflammatory phenotype, which may explain why these cancer subtypes are more likely to respond to I / O therapy.
[0157] Table 1. Patient Characteristics and Related Clinical Information Example 2 Using a DELFI circulating cfDNA-based approach, we non-invasively characterize small cell populations using cell-free DNA fragmentation panels. Lung cancer (SCLC) is subtyped.
[0158] SCLC is an aggressive malignancy with a poor prognosis. Although clinically considered a single cancer type, emerging evidence supports that subtypes of SCLC (neuroendocrine-high versus neuroendocrine-low) acquire distinct transcriptional and epigenetic states. Furthermore, different SCLC subtypes respond to specific therapies, such as immunotherapy in the case of the neuroendocrine-low, highly inflammatory SCLC subtype.
[0159] The novel, noninvasive DELFI-based approach described here distinguishes between neuroendocrine-high and neuroendocrine-low SCLC subtypes by studying cell-free fragments that reflect the unique epigenetic state of SCLC. This novel approach can identify patients who are likely to respond to immunotherapy, such as immune checkpoint blockade.
[0160] Using WGS data from plasma-derived cfDNA, we investigated fragment coverage at specific genomic coordinates to reveal specific subtypes of SCLC cases. Targeted DELFI-based analysis was able to distinguish between neuroendocrine-high and neuroendocrine-low SCLC cases.
[0161] SCLC subtypes can be predicted using cfDNA LC-WGS data according to a DELFI-based approach.
[0162] First, the DELFI machine learning classifier detected whether SCLC patients had cancer. Both the genome-wide fragmentation profiles and the corresponding DELFI scores were examined to identify preliminary differences between the two major clinical subtypes of SCLC cases: high neuroendocrine and low neuroendocrine. Subsequently, principal component analysis was performed on publicly available data to cluster the two subtypes.
[0163] Second, using publicly available data, we determined that a cluster of these SCLC samples exhibited reduced total fragment coverage at ASCL1 transcription factor binding sites, thereby classifying these cases as neuroendocrine-high SCLC.
[0164] Third, we investigated whether one of these clusters of SCLC samples exhibited reduced fragment coverage at genomic binding sites regulated by hematopoietic transcription factors (neuroendocrine-low SCLC).
[0165] Fourth, high neuroendocrine cases were distinguished from low neuroendocrine cases based on the fragment length distribution of T cell-specific partially methylated domains.
[0166] Finally, receiver operating characteristics (ROCs) representing the sensitivity and specificity of the DELFI-fragment panel method for identifying SCLC subtypes were used to evaluate the diagnostic accuracy of our method across the identified cluster samples.
[0167] Given the limited access to SCLC tumor biopsies, this assay is thought to be a viable way to subtype SCLC in a noninvasive manner. It is a promising test both for pharmaceutical companies evaluating new immunotherapy drugs during clinical trials and for clinicians seeking to select which patients diagnosed with SCLC have the best chance of responding to immunotherapy.
[0168] The DELFI scores of patients diagnosed with SCLC in this study were as follows Figure 1 shown.
[0169] Figure 2 We demonstrated that genome-wide fragmentation profiles were remarkably consistent between pre-treatment, post-treatment, progression, and response time points, suggesting that genome-wide circulating tumor DNA fragment size is a powerful method for detecting changes during immunotherapy treatment in SCLC cancer.
[0170] Figure 3 Genome-wide fragmentation profiles (and corresponding DELFI scores below) are presented, showing some differences between the two major subtypes of SCLC cases (neuroendocrine-high and neuroendocrine-low); this suggests that genome-wide circulating tumor DNA fragment size may be a powerful method for detecting subtyping during immunotherapy treatment in monitoring SCLC cancer.
[0171] Figure 4 Differentially expressed genes between high and low neuroendocrine SCLC cases were shown using publicly available data (Lissa et al. 2022) and principal component analysis (based on DELFI 30X plasma WGS data) to cluster the two subtypes. PCA analysis of the data revealed distinguishable clusters between SCLC high and low NE pre-treatment samples.
[0172] Figure 5 shows the feature correlation analysis that revealed significant correlations between principal components 1 and 2 and the NE status of the samples.
[0173] Figures 6A to 6B We demonstrated the potential of the DELFI assay using TFBSs to target SCLC subtypes, performing genome-wide cfDNA fragmentation analysis at ASCL1 binding sites in pretreatment NCI samples. Figure 5A We show that distinct clusters of SCLC samples were observed corresponding to different levels of ASCL1 activation. Importantly, these two clusters corresponded perfectly to distinct SCLC subtypes. The primary driver of differences between samples was actually differences in coverage of fragments identified at the center of the ASCL1 binding site. Figure 6BWe show that other clinical or sample characteristics have negligible influence on the clustering of these samples. Overall, these data demonstrate how DELFI analysis can perfectly distinguish SCLC subtypes using fragment coverage at TFBSs.
[0174] Figure 7 This study presents a patient example supporting the hypothesis that the detected signal originates from tumor-infiltrating lymphocytes. Patient NCI-0422 was diagnosed with SCLC with an inflammatory subtype and received durvalumab combined with olaparib. NCI-0422 responded well to treatment but was later diagnosed with disease progression. Variable genomic binding was detected before treatment and at the time of progression. Peaks were concentrated at intermingled regions of TFBS and TSS.
[0175] Figures 8A to 8B ASCL1 binding sites between responders and non-responders are shown. Figure 8A Shown are 500 cell type-specific bins from the PMD data to determine the fraction of leukocytes. Figure 8B Lymphocyte tissue-specific compartments are shown.
[0176] Figure 9 A machine learning model for predicting high versus low neuroendocrine SCLC was demonstrated.
[0177] Subtyping was performed by calculating the log2 (read depth ratio) across a 20-bp window spanning 2 kb from the ASCL1 TFBS site. The read depth ratio is the total coverage divided by the coverage correction factor, which is calculated using the median coverage of the 500-bp anchors 5' and 3' to the ends of the 2 kb window (i.e., log2 (read depth / correction factor)). A generalized additive model was then used to smooth the coverage and correct for GC bias. To validate the method described here, neuroendocrine status was derived from the mean NE50 score for each subject / treatment status, with positive values classified as NE+. Subtype status was derived by summing the mean values for each of the four subtypes (ASCL1, NEUROD1, POUF23, and YAP1), resulting in a single value for each subtype for each subject / treatment status. The highest-scoring subtype was then selected as the true subtype. Fragmentomics was calculated as previously described (Cristiano et al.), and scores were applied to the latest available model.
[0178] Figures 10A to 10B We demonstrate that the method described here can distinguish neuroendocrine from non-neuroendocrine SCLCs by calculating read depth ratios at SCLC-specific genomic coordinates. Figure 10A ). Figure 10B It was shown that the method described herein appears to better subtype SCLC cases.
[0179] Figure 11 This figure shows that the read depth ratios at SCLC-specific genomic coordinates were transformed into a two-dimensional PCA to classify the samples into neuroendocrine and non-neuroendocrine groups. Most samples were clearly divided into different clusters.
[0180] Figure 12 The results show that PCA derived from read depth ratios at SCLC-specific genomic coordinates can classify pre-treatment samples into specific SCLC subtypes (A, N, P, Y). The true picture of SCLC subtypes was calculated using tissue RNA-seq gene expression differential analysis.
[0181] The fragmentomics platform described in this article detects small cell lung cancer (SCLC) with high sensitivity. The DELFI SCLC Subtyping Assay enables subtyping of SCLC samples without clinical knowledge. The DELFI SCLC Subtyping Assay identified four groups of SCLC samples with varying levels of ASCL1 activation. Samples with high ASCL1 activation were confirmed to belong to the neuroendocrine subtype (SCLC-A and SCLC-N), while samples with low ASCL1 signal were confirmed to belong to the non-neuroendocrine subtype (SCLC-P and SCLC-Y).
[0182] Example 3 Cell-free DNA fragmentation profiling as a method for tumor score assessment and treatment monitoring in NSCLC cfDNA fragmentomics Background: Traditionally, monitoring has been done through imaging, but many people do not have easy access to hospitals with appropriate imaging equipment and expertise and need to travel long distances for such procedures, making continuous monitoring difficult. in advance The location of the tumor is known. Additionally, there is no need to know in advance what somatic mutations the tumor harbors. Compared to imaging, the costs associated with liquid biopsies are low, and as sequencing costs continue to decline, these costs should become even lower.
[0183] Mutations in cfDNA have limited signal and may be confounded by clonal hematopoiesis. DELFI captures many features across the entire cfDNA landscape. The cfDNA fragmentation pattern is determined by the underlying chromatin organization. The cfDNA fragmentation profile is highly consistent in healthy individuals but altered in cancer patients.
[0184] Overview of the DELFI-TF model and its application to monitoring NSCLC patients during therapy: cfDNA aliquots from plasma samples from CRC patients were analyzed for RAS mutation status using ddPCR and low-pass WGS sequencing. The WGS data were aligned, and fragment size distributions were obtained for 504 5-Mb bins across the genome. A Bayesian regression model was trained using the fragmentomic signature and cross-validated for RAS MT samples. Figure 20Depicts proof-of-concept DELFI-TF model development.
[0185] The CRC-trained DELFI-TF model was applied to a real-world NSCLC cohort.
[0186] Table 2 - Clinical characteristics of the NSCLC cohort result: Figure 13 We demonstrate how fragmentation profiles can inform lung cancer status and treatment response patterns.
[0187] Figure 14 We show that the model-derived DELFI-TFs exhibit strong correlation with mutant allele frequencies across samples.
[0188] Figure 15 We demonstrated that DELFI-TF accurately quantified circulating tumor fraction, which reflected MAF expression in association with RECIST assessment.
[0189] Figure 16 demonstrated that tumor-derived cfDNA has altered fragmentation.
[0190] Figures 17A to 17C We demonstrated that DELFI-TF accurately detected the circulating tumor fraction without interference from clonal hematopoiesis.
[0191] Figure 18 Depicted that cfDNA fragmentation patterns accurately distinguish NSCLC subtypes.
[0192] Conclusion: The genome-wide cfDNA fragmentation profiles of cancer patients are abnormal The cfDNA fragmentation score (DELFI-TF) is highly correlated with known mutant allele frequencies. cfDNA fragmentation predicts RECIST status. cfDNA fragmentation is not confounded by clonal hematopoiesis. cfDNA fragmentation signatures can noninvasively differentiate lung cancer histological subtypes.
Claims
1. A non-invasive method for subtyping small cell lung cancer in a subject as neuroendocrine-poor or neuroendocrine-high small cell lung cancer, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; performing whole genome sequencing on the sequencing library to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to a genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the mapped sequenced genomic intervals to determine cfDNA fragment lengths and amounts, thereby establishing a cfDNA fragment coverage score at a specified transcription factor binding site using the cfDNA fragment lengths and amounts; as well as The small cell lung cancer in the subject is subtyped based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of a high neuroendocrine small cell lung cancer subtype. 2 . The method of claim 1 , wherein the designated transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.
3. The method of claim 1, wherein the method uses machine learning to subtype small cell lung cancer in the subject.
4. The method of claim 1, wherein subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
5. The method of claim 4, wherein the log2(read depth ratio) is the total coverage divided by a coverage correction factor, which is calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of a 2 kb window (i.e., log2(read depth / correction factor)). The method of claim 5 , wherein all coverage calculations are offset by 1 to avoid division by zero.
7. The method of claim 5, wherein the log2(read depth ratio) is calculated for each transcription factor binding site and then summarized by calculating the median of each window to obtain a fragment coverage score.
8. The method of claim 1, further comprising using a generalized additive model to smooth coverage and correct for GC bias.
9. The method of claim 1, wherein the genomic intervals are non-overlapping.
10. The method of claim 1, wherein the genomic intervals each comprise thousands to millions of base pairs.
11. The method of claim 1 , wherein a cfDNA fragmentation profile is determined within each genomic bin.
12. The method of claim 11, wherein the cfDNA fragmentation profile comprises a median fragment size.
13. The method of claim 11, wherein the cfDNA fragmentation profile comprises a fragment size distribution.
14. The method of claim 1, further comprising administering to the subject identified as having the neuroendocrine-high small cell lung cancer subtype a therapeutic agent suitable for treating the cancer subtype.
15. The method of claim 14, wherein the therapeutic agent is an immunotherapy.
16. A non-invasive method for subtyping non-small cell lung cancer in a subject as adenocarcinoma or squamous cell carcinoma, the method comprising: processing cfDNA fragments from a sample obtained from the subject and generating a sequencing library; performing whole genome sequencing on the sequencing library to obtain sequenced fragments, wherein the genome coverage is about 30× to 0.1×; mapping the sequenced fragments to a genome to obtain a genomic interval of mapped sequence at a specified transcription factor binding site; analyzing the mapped sequenced genomic intervals to determine cfDNA fragment lengths and amounts, thereby establishing a cfDNA fragment coverage score at a specified transcription factor binding site using the cfDNA fragment lengths and amounts; as well as The small cell lung cancer in the subject is subtyped based on transcription factor activation; wherein a decrease in the total cfDNA fragment coverage score at the specified transcription factor binding site is indicative of an adenocarcinoma or squamous cell carcinoma subtype.
17. The method of claim 16, wherein the designated transcription factor is ASCL1, NEUROD1, POUF23, YAP1, or any combination thereof.
18. The method of claim 16, wherein the method uses machine learning to subtype small cell lung cancer in the subject.
19. The method of claim 16, wherein subtyping is performed by calculating the log2 (read depth ratio) across a 20 bp window 2 kb from the specified transcription factor binding site.
20. The method of claim 19, wherein the log2(read depth ratio) is the total coverage divided by a coverage correction factor calculated using the median coverage of the 5' and 3' 500 bp anchors at the ends of a 2 kb window (i.e., log2(read depth / correction factor)).
21. The method of claim 19, wherein all coverage calculations are offset by 1 to avoid division by zero.
22. The method of claim 19, wherein the log2(read depth ratio) is calculated for each transcription factor binding site and then summarized by calculating the median value for each window to obtain a fragment coverage score.
23. The method of claim 16, further comprising using a generalized additive model to smooth coverage and correct for GC bias.
24. The method of claim 16, wherein the genomic intervals are non-overlapping.
25. The method of claim 16, wherein the genomic intervals each comprise thousands to millions of base pairs.
26. The method of claim 16, wherein a cfDNA fragmentation profile is determined within each genomic bin.
27. The method of claim 26, wherein the cfDNA fragmentation profile comprises a median fragment size.
28. The method of claim 26, wherein the cfDNA fragmentation profile comprises a fragment size distribution.
29. The method of claim 16, further comprising administering to the subject identified as having an adenocarcinoma or squamous cell carcinoma subtype a therapeutic agent suitable for treating the cancer type.
30. The method of claim 29, wherein the therapeutic agent is an immunotherapy.
31. The method of claim 16, further comprising quantifying a circulating tumor score that reflects major allele frequency (MAF) performance associated with Response Evaluation Criteria in Solid Tumors (RECIST) assessment in both adenocarcinoma and squamous cell carcinoma.
Citation Information
Patent Citations
Cell-free DNA for assessing and / or treating cancer
US10975431B2
Cell-free DNA for assessing and / or treating cancer
US10982279B2