Methods and systems for tumor informed circulating tumor fraction estimation
The method addresses the challenge of accurately determining circulating tumor fraction in liquid biopsies by combining tissue-based and blood-based profiling, enhancing the precision of cancer monitoring and treatment decisions.
Patent Information
- Application Number
- JP2025062086
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-04-03
- Publication Date
- 2025-10-17
AI Technical Summary
Conventional liquid biopsy assays face challenges in accurately determining the circulating tumor fraction due to cancer heterogeneity and chromosomal variations, leading to inaccurate measurement of cancer features and limitations in comprehensive genomic analysis.
A method and system for estimating circulating tumor fraction using a combination of tissue-based comprehensive genomic profiling and non-custom blood-based profiling, involving targeted panel sequencing of both solid tumor and liquid biopsy samples to determine variant allele frequencies and tumor fraction.
Provides accurate and sensitive estimates of circulating tumor fraction, enabling better monitoring of treatment response, disease recurrence, and progression, facilitating more precise treatment decisions and reducing the invasiveness of cancer monitoring.
Smart Images

Figure 2025158958000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 574,758, entitled "Methods and Systems for Tumor Informed Circulating Tumor Fraction Estimation," filed April 4, 2024, which is incorporated herein by reference.
[0002] The present disclosure relates generally to the use of tumor-informed liquid biopsy data to estimate a subject's circulating tumor fraction to provide clinical support for personalized cancer treatment. [Background technology]
[0003] Precision oncology is the practice of tailoring cancer therapy to an individual's unique genomic, epigenetic, and / or transcriptomic profile. Personalized cancer therapy builds on traditional treatment regimens used to treat cancer based solely on the overall classification of the cancer, for example, treating all breast cancer patients with one treatment and all lung cancer patients with a second treatment. This field arose from the frequent observation that different patients diagnosed with the same type of cancer, such as breast cancer, responded very differently to common treatment regimens. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that improve predictions of how individual cancers will respond to specific treatment modalities.
[0004] There is growing evidence that cancer patients who receive genetically guided therapy have better outcomes. For example, studies have shown that targeted therapy results in significant improvements in progression-free cancer survival. See, e.g., Radovich M. et al., Oncotarget, 7(35):56491-500 (2016). Similarly, a report from the IMPACT trial, a large (n=1307) retrospective analysis of consecutive, prospectively molecularly profiled patients with advanced cancer who participated in a large personalized medicine trial, showed that patients receiving tumor-biologically matched targeted therapy had a 16.2% response rate, compared with a 5.2% response rate for patients receiving unmatched therapy. Tsimberidou et al., ASCO 2018, Abstract LBA2553 (2018).
[0005] Indeed, therapies targeted to specific genomic alterations are already standard of care in some tumor types, as suggested by the National Comprehensive Cancer Network (NCCN) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer. In practice, the implementation of these targeted therapies requires determining the diagnostic marker status in each eligible cancer patient. While this can be achieved for some known mutations associated with treatment recommendations in the NCCN guidelines using individual assays or small next-generation sequencing (NGS) panels, the increasing number of actionable genomic alterations and the increasing complexity of diagnostic classifiers necessitates a more comprehensive assessment of each patient's cancer genome, epigenome, and / or transcriptome.
[0006] For example, evidence suggests that the use of combination therapies, in which each component is matched to actionable genomic alterations, holds the greatest potential for treating individual cancers. To date, retrospective studies of cancer patients treated with one or more therapeutic regimens have revealed that patients receiving therapies matched to a higher proportion of genomic alterations experienced a higher frequency of stable disease (e.g., longer time to recurrence), longer time to treatment failure, and better overall survival. Wheeler et al., 2016, Cancer Res., 76:3690-701. Therefore, comprehensive assessment of each cancer patient's genome, epigenome, and / or transcriptome should maximize the benefits offered by precision oncology by facilitating more fine-tuned combination therapies, off-label use of novel drugs, and / or tissue-independent immunotherapies. See, e.g., Schwaederle et al., 2015, J Clin Oncol., 33(32):3817-25; Schwaederle et al., 2016, JAMA Oncol., 2(11):1452-59; and Wheler et al., 2016, Cancer Res., 76(13):3690-701. Furthermore, the use of comprehensive next-generation sequencing analysis of cancer genomes facilitates better access and larger patient pools for clinical trial enrollment. Coyne et al., 2017, Curr. Probl. Cancer, 41(3):182-93; and Markman, Oncology, 31(3):158, 168.
[0007] To address the need for more comprehensive characterization of an individual's cancer genome, the use of large-scale NGS genomic analysis is increasing. See, e.g., Fernandes et al., Clinics, 72(10):588-94. Recent studies have shown that 30-40% of patients undergoing large-scale NGS genomic analysis subsequently receive clinical care based on the assay results, which is limited by, at least, the identification of actionable genomic alterations, the availability of drugs to treat the identified actionable genomic alterations, and the subject's clinical condition. See, e.g., Ross et al., 2015, JAMA Oncol., 1(1):40-49; Ross et al., 2015, Arch. Pathol. Lab Med., 139:642-49; Hirshfield KM et al., Oncologist, 2016, 21(11):1315-25; and Groisberg et al., 2017, Oncotarget, 8:39254-67.
[0008] However, these large-scale NGS genomic analyses are traditionally performed on solid tumor samples. For example, each of the studies referenced in the above paragraph performed NGS analysis of FFPE tumor blocks from patients. Solid tissue biopsies represent a well-known and proven methodology that provides a high degree of accuracy and therefore remain the gold standard for diagnosis and identification of predictive biomarkers. Nevertheless, the use of solid tissue materials for large-scale NGS genomic analyses of cancer has significant limitations. For example, tumor biopsies are subject to sampling bias caused by spatial and / or temporal genetic heterogeneity, e.g., between two regions of a single tumor and / or between different cancer tissues (e.g., between a primary tumor site and a metastatic tumor site, or between two different primary tumor sites). Such inter- or intratumor heterogeneity can lead to overlooking subclonal mutations or newly emerging mutations when using localized tissue biopsies, and sampling bias can worsen over time as subclonal populations further evolve and / or shift in dominance.
[0009] In addition, obtaining solid tissue biopsies often requires invasive surgical procedures, for example, when the primary tumor site is located in an internal organ. These procedures are expensive, time-consuming, and can involve significant risks to the patient, for example, when the patient's health is poor and they may not be able to tolerate invasive medical procedures, and / or when the tumor is located in a particularly sensitive or inoperable location, such as the brain or heart. Furthermore, the amount of tissue that can be procured depends on multiple factors, including tumor location, tumor size, patient vulnerability, and the risk of biopsy-related comorbidities, such as bleeding and infection. For example, a recent study reported that tissue samples in the majority of patients with advanced non-small cell lung cancer are limited to small biopsies, and in up to 31% of patients, no samples can be obtained at all. (Ilie and Hofman, Transl. Lung Cancer Res., 5(4):420-23 (2016)). Even when tissue biopsies are obtained, the sample may be too insufficient for comprehensive testing.
[0010] Furthermore, methods of tissue collection, preservation (e.g., formalin fixation), and / or archiving of tissue biopsies can result in sample degradation and variable-quality DNA. This, in turn, leads to inaccuracies in downstream assays and analyses, including next-generation sequencing (NGS) for biomarker identification. Ilie and Hofman, Transl Lung Cancer Res., 5(4):420-23 (2016).
[0011] Additionally, the invasiveness of the biopsy procedure, the time and expense associated with obtaining the sample, and the impaired state of cancer patients undergoing therapy make repeated testing of cancer tissue impractical, if not impossible. As a result, solid tissue biopsy analysis is not suitable for many monitoring schemes that benefit cancer patients, such as disease progression analysis, treatment efficacy assessment, disease recurrence monitoring, and other techniques that require data from several time points.
[0012] Cell-free DNA (cfDNA) has been identified in various body fluids, such as serum, plasma, and urine. Chan et al., 2003, Ann. Clin. Biochem., 40(Pt 2):122-30. This cfDNA is derived from all types of necrotic or apoptotic cells, including germline cells, hematopoietic cells, and pathological (e.g., cancer) cells. Advantageously, genomic alterations in cancer tissues can be identified from cfDNA isolated from cancer patients. See, for example, Stroun et al., 1989, Oncology, 46(5):318-22; Goessl et al., 2000, Cancer Res., 60(21):5941-45; and Frenel et al., 2015, Clin. Cancer Res. 21(20):4586-96. Thus, one approach to overcoming the problems presented by the use of solid tissue biopsies described above is to analyze cell-free nucleic acids (e.g., cfDNA) and / or nucleic acids in circulating tumor cells present in biological fluids, for example, via liquid biopsy.
[0013] Specifically, liquid biopsies offer several advantages over traditional solid tissue biopsy analysis. For example, because bodily fluids can be collected in a minimally or non-invasive manner, sample collection is simpler, faster, safer, and less expensive than solid tumor biopsies. Such methods require only small sample volumes (e.g., 10 mL or less of whole blood per biopsy), reducing the discomfort and risk of complications experienced by patients during traditional tissue biopsies. In fact, liquid biopsy samples can be collected with limited or no assistance from a medical professional and can be performed in almost any location. Furthermore, liquid biopsy samples can be collected from any patient, regardless of the location of their cancer, their overall health, and any previous biopsy collections. This enables the analysis of cancer genomes in patients for whom solid tumor samples cannot be easily and / or safely obtained. In addition, because cell-free DNA in bodily fluids originates from many different types of tissue in a patient, the genomic alterations present in the pool of cell-free DNA represent various distinct clonal subpopulations of the target cancer tissue, facilitating a more comprehensive analysis of the target cancer genome than is possible from one or more sections of a single solid tumor sample.
[0014] Liquid biopsies also enable serial genetic testing prior to cancer detection, during early stages of cancer progression, throughout the course of treatment, and during remission, for example, to monitor for disease recurrence. The ability to perform serial testing via non-invasive liquid biopsies throughout the course of disease may prove beneficial to many patients, for example, by monitoring patient response to therapy, the emergence of new actionable genomic alterations, and / or drug resistance changes. This type of information allows medical professionals to more quickly adjust and update treatment regimens, for example, facilitating more timely intervention in the event of disease progression. See, e.g., Ilie and Hofman, 2016, Transl. Lung Cancer Res., 5(4):420-23.
[0015] While liquid biopsies are a promising tool for improving outcomes using precision oncology, the use of cell-free DNA for assessing a subject's cancer genome presents significant challenges. For example, one challenge associated with liquid biopsies is accurately determining the tumor fraction in a sample. This difficulty stems at least in part from cancer heterogeneity and the increased frequency of large-scale chromosomal duplications and deletions found in cancer. As a result, the frequency of genomic alterations from cancer tissue varies from locus to locus based on at least (i) their prevalence in different subclonal populations of a subject's cancer and (ii) their location within the genome relative to large-scale chromosomal copy number variations. The difficulty of accurately determining the tumor fraction of a liquid biopsy sample impacts the accurate measurement of various cancer features that have been shown to have diagnostic value for the analysis of solid tumor biopsies. These include allele ratios, copy number variations, global mutation burden, and the frequency of aberrant methylation patterns, all of which correlate with the percentage of DNA fragments originating from cancer tissue as opposed to healthy tissue.
[0016] The information disclosed in this Background section is merely for the purpose of enhancing understanding of the general background of the present invention and should not be construed as an acknowledgment or any form of suggestion that this information forms prior art already known to those skilled in the art. Summary of the Invention
[0017] Quantitative estimation of circulating tumor fraction in liquid biopsy samples is a promising area of clinical development for monitoring therapeutic molecular genetic response and correlates with patient outcomes. However, in light of the above background, there is a need in the art for improved methods and systems for supporting clinical decisions using liquid biopsy assays in precision oncology. In particular, there is a need in the art for improved methods and systems for determining accurate circulating tumor fraction estimates (ctFE) in liquid biopsy assays. The present disclosure solves this need and other needs in the art by providing a method and system for estimating circulating tumor fraction in liquid biopsy samples using a combination of tissue-based comprehensive genomic profiling (CGP) and non-custom blood-based profiling.
[0018] For example, in one aspect, the present disclosure provides methods, systems programmed to perform such methods, and computer-readable media storing instructions for carrying out such methods, for estimating circulating tumor fraction in a subject.
[0019] The method includes obtaining a first plurality of nucleic acid sequences from a solid tumor sample of a subject, the first plurality of nucleic acid sequences including a nucleic acid sequence corresponding to each respective locus in a plurality of loci in genomic DNA.
[0020] In some such embodiments, the first plurality of nucleic acid sequences is determined from a second panel enrichment sequencing reaction using a second plurality of probes that includes, for each respective locus in the plurality of loci, a corresponding probe in the second plurality of probes that hybridizes to the respective locus.
[0021] In some such embodiments, the plurality of loci are sequenced in the second panel enrichment sequencing reaction to an average sequencing depth of at least 50x, 75x, 100x, 125x, 500x, or 1000x.
[0022] In some such embodiments, the second plurality of probes enriches loci from at least 50 genes.
[0023] The method also includes obtaining, for each respective locus in the plurality of loci, a second plurality of nucleic acid sequences including a corresponding nucleic acid sequence for each cell-free DNA fragment in the plurality of cell-free DNA fragments obtained from the liquid biopsy sample from the first panel enrichment sequencing assay, using a first plurality of probes including a corresponding probe that hybridizes to the respective locus.
[0024] In some such embodiments, the first plurality of probes and the second plurality of probes are different.
[0025] In some such embodiments, the plurality of loci are sequenced in the first panel enrichment sequencing reaction to an average sequencing depth of at least 300x, 400x, 500x, 700x, or 1000x.
[0026] In some such embodiments, the first plurality of probes enriches loci from at least 50 genes.
[0027] In some such embodiments, the identities of the first plurality of probes are non-custom to the subject.
[0028] In some such embodiments, the solid tumor sample is collected prior to collecting the liquid biopsy sample.
[0029] In some such embodiments, the solid tumor sample and the liquid biopsy sample are collected within six months of each other.
[0030] In some such embodiments, the liquid biopsy sample is blood.
[0031] In some such embodiments, the liquid biopsy sample comprises the subject's blood, whole blood, peripheral blood, plasma, serum, or lymph.
[0032] The method also includes identifying one or more somatic mutations in the first plurality of nucleic acid sequences, each respective somatic mutation in the one or more somatic mutations being at corresponding one or more nucleotide positions in corresponding loci in the plurality of one or more loci.
[0033] In some such embodiments, the identifying includes identifying a plurality of candidate somatic mutations by comparing each nucleic acid sequence in the first plurality of nucleic acid sequences to nucleic acid sequences in a third plurality of nucleic acid sequences obtained from a sequencing reaction of genomic DNA from non-cancerous tissue of the subject.
[0034] In some such embodiments, the identifying further comprises excluding one or more respective candidate somatic mutations in the plurality of candidate somatic mutations determined to have an acentric variant allele fraction in the first plurality of sequences.
[0035] In some such embodiments, the eliminating comprises fitting a VAF for each respective candidate somatic mutation in the plurality of candidate somatic mutations in the first plurality of sequences to a distribution, and eliminating candidate somatic mutations having a corresponding VAF that is outside the spread for the distribution.
[0036] In some such embodiments, the distribution is a normal distribution, a beta distribution, a beta prime distribution, a lognormal distribution, or a gamma distribution.
[0037] In some such embodiments, the distribution is a normal distribution.
[0038] In some such embodiments, the dispersion is a multiple of the standard deviation of the distribution.
[0039] In some such embodiments, the identifying further comprises eliminating one or more respective candidate somatic mutations in the plurality of candidate somatic mutations having a nucleotide position that does not correspond to any probe in the first plurality of probes.
[0040] The method also includes, for each respective somatic mutation of the one or more somatic mutations, determining a corresponding variant allele frequency (VAF) in the liquid biopsy sample, where the corresponding VAF is determined from (i) the frequency of each somatic mutation in the second plurality of nucleic acid sequences and (ii) the frequency of the wild-type allele at the one or more nucleotide positions corresponding to each somatic mutation in the second plurality of nucleic acid sequences, thereby determining a set of VAFs for the one or more somatic mutations in the liquid biopsy sample.
[0041] In some embodiments, an estimate of circulating tumor fraction for a subject is determined based on a set of VAFs for one or more somatic mutations.
[0042] In some such embodiments, the estimate of circulating tumor fraction is representative of a set of VAFs for one or more somatic mutations.
[0043] In some such embodiments, the representative value is the median value.
[0044] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief explanation of the drawings]
[0045] [Figure 1A] 1A, 1B, 1C, and 1D collectively illustrate a block diagram of an exemplary computing device for estimating circulating tumor fraction of a liquid biopsy sample from targeted panel sequencing data, according to some embodiments of the present disclosure. [Figure 1B]1A, 1B, 1C, and 1D collectively illustrate a block diagram of an exemplary computing device for estimating circulating tumor fraction of a liquid biopsy sample from targeted panel sequencing data, according to some embodiments of the present disclosure. [Figure 1C] 1A, 1B, 1C, and 1D collectively illustrate a block diagram of an exemplary computing device for estimating circulating tumor fraction of a liquid biopsy sample from targeted panel sequencing data, according to some embodiments of the present disclosure. [Figure 1D] 1A, 1B, 1C, and 1D collectively illustrate a block diagram of an exemplary computing device for estimating circulating tumor fraction of a liquid biopsy sample from targeted panel sequencing data, according to some embodiments of the present disclosure. [Figure 2A] 1 illustrates an exemplary workflow for generating a clinical report based on information generated from the analysis of one or more patient samples, according to some embodiments of the present disclosure. [Figure 2B] 1 illustrates an example of a distributed diagnostic environment for collecting and evaluating patient data for precision oncology purposes, according to some embodiments of the present disclosure. [Figure 3] 1 provides an exemplary flowchart of processes and features for liquid biopsy sample collection and analysis for use in precision oncology, according to some embodiments of the present disclosure. [Figure 4A]Figures 4A, 4B, 4C, 4D, and 4E collectively illustrate an example bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in the bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of the bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 4B] Figures 4A, 4B, 4C, 4D, and 4E collectively illustrate an example bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in the bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of the bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 4C]Figures 4A, 4B, 4C, 4D, and 4E collectively illustrate an example bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in the bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of the bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 4D] Figures 4A, 4B, 4C, 4D, and 4E collectively illustrate an example bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in the bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of the bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 4E]Figures 4A, 4B, 4C, 4D, and 4E collectively illustrate an example bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in the bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of the bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 5A] Figures 5A, 5B, 5C, and 5D collectively provide a flowchart of processes and features for estimating circulating tumor fraction based on tumor information of a liquid biopsy sample from targeted panel sequencing data according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. [Figure 5B] Figures 5A, 5B, 5C, and 5D collectively provide a flowchart of processes and features for estimating circulating tumor fraction based on tumor information of a liquid biopsy sample from targeted panel sequencing data according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. [Figure 5C] Figures 5A, 5B, 5C, and 5D collectively provide a flowchart of processes and features for estimating circulating tumor fraction based on tumor information of a liquid biopsy sample from targeted panel sequencing data according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. [Figure 5D]Figures 5A, 5B, 5C, and 5D collectively provide a flowchart of processes and features for estimating circulating tumor fraction based on tumor information of a liquid biopsy sample from targeted panel sequencing data according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. [Figure 6A] Figures 6A and 6B show the distribution of Kolmogorov-Smirnov p-values for beta, beta prime (inverse beta), gamma, lognormal (heavy-tailed), and normal distributions fitted to untransformed (6A) and log2-transformed (6B) VAFs for candidate mutations identified in solid tumor samples, as described in Example 3. [Figure 6B] Figures 6A and 6B show the distribution of Kolmogorov-Smirnov p-values for beta, beta prime (inverse beta), gamma, lognormal (heavy-tailed), and normal distributions fitted to untransformed (6A) and log2-transformed (6B) VAFs for candidate mutations identified in solid tumor samples, as described in Example 3. [Figure 7A] 7A and 7B show examples of normal distributions fitted to solid tumor VAFs of solid tumor samples (7A) and distributions of VAFs for somatic mutations identified in a liquid biopsy assay (7B), as described in Example 3. [Figure 7B] 7A and 7B show examples of normal distributions fitted to solid tumor VAFs of solid tumor samples (7A) and distributions of VAFs for somatic mutations identified in a liquid biopsy assay (7B), as described in Example 3. [Figure 8A] We show that tumor-informed ctDNA tumor fractions (TFs) correlate with tumor-naive ctDNA TFs after removing specimens without tumor-informed mutations when using a 105-gene liquid biopsy panel enrichment sequencing assay. [Figure 8B] The number of mutations in tumor samples based on tumor information is shown. [Figure 9A]We show that tumor-informed ctDNA TFs correlate with tumor-naive ctDNA TFs after removing specimens without tumor-informed somatic mutations when using a 523-gene liquid biopsy panel enrichment sequencing assay. [Figure 9B] The number of variants in the tumor specimens based on tumor information is shown. [Figure 10] We show the accuracy of tumor-informed cfDNA TF estimates in a cohort of tumor samples with companion LPG-WGS, using ichorCNA, mean VAF, and tumor-naive ctDNA TFs as controls. [Figure 11A] Figures 11A and 11B show that tumor-informed ctDNA TFs correlate with tumor-naive ctDNA TFs when using 105-gene (11A) and 523-gene (11B) liquid biopsy panel enrichment sequencing assays, requiring more than four tumor-informed somatic mutations. [Figure 11B] Figures 11A and 11B show that tumor-informed ctDNA TFs correlate with tumor-naive ctDNA TFs when using 105-gene (11A) and 523-gene (11B) liquid biopsy panel enrichment sequencing assays, requiring more than four tumor-informed somatic mutations. [Figure 12A] Performance metrics for tumor-informed ctDNA fraction estimation are shown. LOB(95) and LOB(99) were 0%. LOD hit rates were 100% at the lowest titer evaluated for each assay. [Figure 12B] Figures 12B and 12C show that a 100x bootstrap LOB calculated from presumably healthy subjects yields variant count distributions similar to those observed in the 105-gene (12A) and 523-gene (12B) liquid biopsy panel enrichment sequencing assays. [Figure 12C]Figures 12B and 12C show that a 100x bootstrap LOB calculated from presumably healthy subjects yields variant count distributions similar to those observed in the 105-gene (12A) and 523-gene (12B) liquid biopsy panel enrichment sequencing assays. [Figure 13A] Figures 13A and 13B show that the LODs calculated from titered Seraseq ctDNA reference material have low titer-to-titer variability and a strong linear relationship when using the 105-gene (13A) and 523-gene (13B) liquid biopsy panel enrichment sequencing assays. [Figure 13B] Figures 13A and 13B show that the LODs calculated from titered Seraseq ctDNA reference material have low titer-to-titer variability and a strong linear relationship when using the 105-gene (13A) and 523-gene (13B) liquid biopsy panel enrichment sequencing assays. [Figure 14A] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14B] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14C] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14D]Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14E] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14F] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14G] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14H] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14I] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14J]Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14K] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14L] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 14M] Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14J, 14K, 14L, and 14M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 15A] 15A, 15B, and 15C collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 15B] 15A, 15B, and 15C collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 15C] 15A, 15B, and 15C collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0046] Like reference numbers refer to corresponding parts throughout the several views of the drawings. Detailed Description
[0047] Introduction.
[0048] As mentioned above, conventional liquid biopsy assays do not provide accurate determination of circulating tumor fraction estimates (ctFE). For example, low-pass whole genome sequencing can be used to estimate tumor fraction, but somatic variant sequences are poorly identified from low-pass whole genome sequencing data, especially from samples with low tumor fraction. Therefore, conventional liquid biopsy assays typically use targeted panel sequencing to achieve the higher sequence coverage required to identify somatic variants present at low levels in samples. However, targeted panel sequencing data may not cover a large enough portion of the genome to accurately estimate tumor fraction. Rather, tumor fraction estimates obtained using variant allele fractions (VAFs) from targeted panel sequencing data are noisy due to variant tissue source and capture bias.
[0049] Collectively, these factors result in highly variable concentrations of ctDNA from patient to patient, potentially locus to locus, confounding accurate measurement of disease indicators and actionable genomic alterations. Furthermore, the quantity and quality of cfDNA obtained from liquid biopsy samples is highly dependent on the specific methodology used to collect, store, sequence, and normalize the sequencing data. Accurate ctFE offers several benefits for liquid biopsy applications, including classifying variants as somatic or germline, detecting clinically relevant copy number changes, and / or using ctFE as a biomarker.
[0050] For example, because up to 30% of breast cancer patients and up to 55% of lung cancer patients, as well as a significant portion of patients in other cancer cohorts, relapse after initial treatment, the ability to detect metastasis and disease recurrence earlier in these patients could significantly improve patient outcomes. Colleoni et al., 2016, “Annual Hazard Rates of Recurrence for breast Cancer During 24 years of Follow-Up: Results From the International Breast Cancer Study Group Trials I to V,” J Clin Oncol, (34), pg. 927; Yates et al., 2017, “Genomic Evolution of Breast Cancer Metastasis and Relapse,” Cancer Cell, (32), pg. 169; Uramoto et al., 2014, “Recurrence after surgery in patients with NSCLC,” Transl Lung Cancer Res, (3), pg. 242; Taunk et al., 2017, “Immunotherapy and radiation therapy for operable early stage and locally advanced non-small cell lung cancer,” Transl Lung Cancer Res, (6), pg. 178. Indeed, recent retrospective and prospective studies have shown that ctDNA following completion of treatment or surgery can serve as a biomarker of disease recurrence in many cancer types, including breast, lung, melanoma, bladder, and colon cancer.Coombes et al., 2019, “Personalized Detection of Circulating Tumor DNA Antedates breast Cancer Metastatic Recurrence,” Clin Cancer Res, (25), pg. 4255; Tie et al., 2019, “Circulating Tumor DNA Analyses as Markers of Recurrence Risk and Benefit of Adjuvant Therapy for Stage III Colon Cancer,” JAMA Oncol, print; McEvoy et al., 2019, “Monitoring melanoma recurrence with circulating tumor DNA: a proof of concept from three case studies,” Oncotarget, (10), pg. 113; Christensen et al. 2019, “Early Detection of Metastatic Relapse and Monitoring of Therapeutic Efficacy by Ultra-Deep Sequencing of Plasma Cell-Free DNA in Patients With Urothelial Bladder Carcinoma,” J Clin Oncol, (37), See pg. 1547; Isaksson et al., 2019, “Pre-operative plasma cell-free circulating tumor DNA and serum protein tumor markers as predictors of lung adenocarcinoma recurrence,” Acta Oncol, (58), pg. 1079. Higher ctFE is associated with radiological disease progression and increased number of metastases.
[0051] Furthermore, ctFE correlates with important clinical outcomes and provides a minimally invasive method for monitoring patients for response to treatment, disease recurrence, and disease progression. However, traditional methods used to determine ctFE in liquid biopsy samples rely on low-pass whole-genome sequencing, which often cannot be used for variant detection (see, e.g., Adalsteinsson et al., “Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors,” (2017) Nature Communications Nov 6;8(1):1324, doi:10.1038 / s41467-017-00965-y; and ichorCNA, the Broad Institute, available online at github.com / broadinstitute / ichorCNA). Other traditional approaches use variant allele fraction (VAF) to estimate tumor fraction, but such approaches are confounded by variant tissue source and capture bias, resulting in high levels of noise. In addition, conventional methods for determining tumor purity estimates in solid tumor biopsy samples rely solely on on-target probe regions, which often cannot be used in conjunction with targeted gene panels containing a small number of genes.
[0052] Advantageously, the present disclosure provides a sensitive, tumor-specific, non-customized approach to infer ctDNA TFs, in which all patient samples are analyzed with the same panel / assay, as opposed to custom methods in which each patient has a customized set of probes, typically based on the patient's previous sequencing results.
[0053] Advantageously, as shown by the data presented in the Examples herein, linearity improves with increasing panel size and number of variants. These results suggest that tumor-informed ctDNA TFs can be utilized to improve the sensitivity of existing methods for estimating tumor fraction and aid in treatment decisions using comprehensive NGS genomic profiling of tissues and fluids.
[0054] Improved methods for obtaining accurate circulating tumor fraction estimates offer several benefits over liquid biopsies. Advantageously, a more robust ctFE improves the classification accuracy of detected variants as somatic or germline variants (e.g., any variant detected with a ctFE or lower can be classified as a somatic variant with high confidence). In addition, accurate ctFE can significantly improve the sensitivity of detecting clinically relevant copy number variations, including integer copy number calls. Furthermore, in some embodiments, ctFE is used as a biomarker for tumor burden, metastasis, disease progression, or treatment resistance. For example, ctFE has been shown to correlate with tumor volume and change in response to treatment.
[0055] As a result, the methods and systems disclosed herein provide a sensitive, cost-effective, and minimally invasive way to monitor patients for response to treatment, disease burden, relapse, progression, and / or new resistance mutations, which can lead to better patient care. Serial ctFE monitoring, when used as part of a course of care, can predict objective measures of progression in at-risk individuals. Due to the cost and convenience of sampling, the methods and systems disclosed herein can be applied at shorter time intervals than radiological methods, allowing for more timely intervention in the event of disease progression.
[0056] Additionally, the methods and systems disclosed herein provide benefits to clinicians by generating more accurate variant calls and / or informative ctFE biomarkers that can aid in predicting clinical outcomes and / or selecting appropriate treatment regimens in patients.
[0057] Identifying actionable genomic alterations in a patient's cancer genome is a challenging and computationally intensive task. Determining various predictive indices useful for precision oncology, such as variant allele ratios, copy number variation, tumor mutation burden, and microsatellite instability status, requires the analysis of hundreds of millions to billions of sequenced nucleic acid bases. A typical bioinformatics pipeline established for this purpose involves at least five stages of analysis: assessing the quality of raw next-generation sequencing data, generating collapsed nucleic acid fragment sequences and aligning them to a reference genome, detecting structural variants in the aligned sequence data, annotating identified variants, and visualizing the data. Each of these steps is computationally intensive.
[0058] For example, the time and space computational complexity of the overall simple global sequence alignment and local pairwise sequence alignment algorithms is quadratic (e.g., a quadratic problem) and increases rapidly with the size (n and m) of the nucleic acid sequences being compared. Specifically, the time and space complexity of these sequence alignment algorithms can be estimated as O(mn). Where O is the upper limit of the asymmetry growth rate of the algorithm, n is the number of bases in the first nucleic acid sequence, and m is the number of bases in the second nucleic acid sequence. See Baichoo and Ouzounis, BioSystems, 156-157:72-85 (2017), the contents of which are incorporated herein by reference in their entirety for all purposes. Considering that the human genome contains more than 3 billion bases, these alignment algorithms are computationally intensive, especially when used to analyze next-generation sequencing (NGS) data, which can generate more than 3 billion sequence reads per reaction.
[0059] This is especially true when performed in liquid biopsy assays, because liquid biopsy samples contain a complex mixture of short DNA fragments originating from many different germline (e.g., healthy) and diseased tissues (e.g., cancerous). Therefore, the cellular origin of sequence reads is unknown, and to provide relevant information about a subject's cancer, sequence signals originating from cancerous cells, which may comprise multiple subclonal populations, must be computationally deconvolved from signals originating from germline and hematopoietic lineages. Thus, in addition to the computationally demanding process required to align sequence reads to the human genome, there is the computational problem of determining whether a particular aberrant signal, e.g., one or more sequence reads corresponding to a genomic alteration, is (i) an artifact and (ii) originates from a cancerous source in the subject. This becomes even more challenging during the early stages of cancer, when treatment is likely most effective, when only a small amount of ctDNA is diluted by germline and hematopoietic DNA.
[0060] Advantageously, the present disclosure provides various systems and methods for improving computational identification of actionable genomic alterations from liquid biopsy samples of cancer patients. Specifically, the present disclosure improves the accuracy of circulating tumor fraction estimated from targeted panel sequencing. Furthermore, because the methods described herein do not need to process data from two different sequencing reactions, the present disclosure reduces the computational budget for accurately estimating circulating tumor fraction and identifying actionable variants. As mentioned above, the disclosed methods and systems are necessarily implemented by computers due to their complexity and stringent computational requirements, thus solving problems in computational technology.
[0061] The methods and systems described herein also improve precision oncology methods for assigning and / or administering treatments by improving the accuracy of estimating circulating tumor fraction. Accurate ctFE can be reported as a biomarker and / or used in downstream analyses to identify therapeutically actionable variants that can be included in clinical reports for patient and / or clinician review. Additionally, ctFE and any therapeutically actionable variants identified using ctFE can be matched to appropriate therapies and / or clinical trials, allowing for more accurate treatment allocation. Improved accuracy of biomarker detection increases the likelihood of efficacy and reduces the risk of patients receiving unnecessary or potentially harmful regimens due to misdiagnosis.
[0062] Definition.
[0063] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, a human (e.g., a male human, a female human, a fetus, a pregnant woman, a child, or the like), a non-human mammal, or a non-human animal. Any human or non-human animal can serve as a subject, including, but not limited to, a mammal, a reptile, a bird, an amphibian, a fish, a ungulate, a ruminant, a bovine (e.g., a cow), an equine (e.g., a horse), a caprine and ovine (e.g., a sheep, a goat), a porcine (e.g., a pig), a camelid (e.g., a camel, a llama, an alpaca), a monkey, an ape (e.g., a gorilla, a chimpanzee), a ursidae (e.g., a bear), a fowl, a dog, a cat, a mouse, a rat, a fish, a dolphin, a whale, and a shark. In some embodiments, the subject is a male or female (e.g., a man, a woman, or a child) of any age.
[0064] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" describe a sample derived from non-diseased tissue. In some embodiments, such a sample is derived from a subject that does not have a particular condition (e.g., cancer). In other embodiments, such a sample is an internal control, e.g., from a subject that may or may not have a particular disease (e.g., cancer), but is derived from the subject's healthy tissue. For example, if a liquid or solid tumor sample is obtained from a subject with cancer, an internal control sample can be obtained from the subject's healthy tissue, e.g., a white blood cell sample from a subject that does not have a blood cancer, or a solid germline tissue sample from the subject. Thus, a reference sample can be obtained from the subject or from a database, e.g., from a second subject that does not have a particular disease (e.g., cancer).
[0065] As used herein, the terms "cancer," "cancer tissue," or "tumor" refer to an abnormal mass of tissue, including both a solid mass (e.g., in the case of solid tumors) or a fluid mass (e.g., in the case of blood cancers), in which the growth of the mass exceeds and is out of step with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are well differentiated, have characteristically slower growth than malignant tumors, and may remain localized to the site of origin. In addition, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors may be poorly differentiated (dysplastic) and have characteristically rapid growth accompanied by progressive infiltration, invasion, and destruction of surrounding tissue. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is out of step with that of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained from or derived from a tumor in a subject, as described herein.
[0066] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, bile duct cancer, adrenal cancer, neural carcinoma, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, non-small cell lung cancer, thymus cancer, tumors, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, serous ovarian cancer, HR+ breast cancer, serous uterine cancer, endometrial cancer, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.
[0067] As used herein, the term "cancer state" or "cancer condition" refers to characteristics of a cancer patient's condition, such as diagnostic status, cancer type, cancer location, cancer primary origin, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject, such as age, sex, weight, race, personal habits (e.g., smoking, alcohol use, diet), other relevant medical conditions (e.g., high blood pressure, dry skin, other diseases), current medications, allergies, relevant medical history, current side effects of cancer treatments and other medications, etc., are used to further describe the subject's cancer state or condition.
[0068] As used herein, the term "liquid biopsy" sample refers to a liquid sample obtained from a subject that contains cell-free DNA. Examples of liquid biopsy samples include, but are not limited to, a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal material, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. In some embodiments, the liquid biopsy sample is a cell-free sample, e.g., a cell-free blood sample. In some embodiments, the liquid biopsy sample is obtained from a subject with cancer. In some embodiments, the liquid biopsy sample is collected from a subject with an unknown cancerous condition, e.g., for use in determining the subject's cancerous condition. Similarly, in some embodiments, the liquid biopsy is collected from a subject with a non-cancerous disorder, e.g., cardiovascular disease. In some embodiments, the liquid biopsy is collected from a subject with an unknown non-cancerous disorder, e.g., for use in determining the subject's non-cancerous disorder condition.
[0069] As used herein, the terms "cell-free DNA" and "cfDNA" refer interchangeably to DNA fragments circulating within a subject's body (e.g., bloodstream) and derived from one or more healthy cells and / or one or more cancer cells. These DNA molecules are found outside of cells in bodily fluids, such as a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal matter, saliva, sweat, perspiration, tears, pleural fluid, pericardial fluid, or peritoneal fluid, and are considered to be fragments of genomic DNA excreted from healthy cells and / or cancer cells, for example, during apoptosis and dissolution of the cell envelope.
[0070] As used herein, the term "locus" refers to a location (e.g., site) within a genome, for example, on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position on a particular chromosome within a genome. In some embodiments, a locus refers to a group of nucleotide positions within a genome. In some cases, a locus is defined by a mutation (e.g., a substitution, insertion, deletion, inversion, or translocation) of consecutive nucleotides within a cancer genome. In some cases, a locus is defined by a gene, a subgenic structure (e.g., a regulatory element, exon, intron, or a combination thereof), or a defined span of a chromosome. Because normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus located on an autosome, for example, one copy on the maternal autosome and one copy on the paternal autosome.
[0071] As used herein, the term "allele" refers to a specific sequence of one or more nucleotides at a chromosomal locus. In haploid organisms, a subject has one allele at every chromosomal locus. In diploid organisms, a subject has two alleles at every chromosomal locus.
[0072] As used herein, the term "base pair" or "bp" refers to a unit consisting of two nucleic acid bases joined together by hydrogen bonds. Generally, the size of an organism's genome is measured in base pairs because DNA is typically double-stranded. However, some viruses have single-stranded DNA or RNA genomes.
[0073] As used herein, the terms "genomic alteration," "mutation," and "variant" refer to detectable changes in the genetic material of one or more cells. Genomic alterations, mutations, or variants can refer to various types of changes in a cell's genetic material, including changes in the primary genomic sequence at single or multiple nucleotide positions, e.g., single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels (e.g., nucleotide insertions or deletions), DNA rearrangements (e.g., inversions or translocations of portions of a chromosome or multiple chromosomes), copy number variations of loci (e.g., exons, genes, or large spans of chromosomes) (CNVs), partial or complete changes in the ploidy of a cell, and changes in the epigenetic information of a genome, such as altered DNA methylation patterns. In some embodiments, a mutation is a change in a cell's genetic information relative to one or more "normal" alleles found in a particular reference genome or population of a species of interest. For example, mutations can be found in both a subject's germline cells (e.g., non-cancerous "normal" cells) and in a subject's abnormal cells (e.g., pre-cancerous or cancerous cells). Thus, mutations in a subject's germline (e.g., found in substantially all "normal cells" in a subject) are identified relative to the reference genome for the subject's species. However, many loci in a species' reference genome are associated with several variant alleles that are significantly represented in the subject's population and are not associated with a pathological condition, e.g., such that they are not considered "mutations." In contrast, in some embodiments, mutations in a subject's cancer cells can be identified relative to either the subject's reference genome or the subject's own germline genome. In certain cases, identifying both types of variants can be beneficial. For example, in some cases, mutations present in both the subject's cancer genome and the subject's germline are beneficial for precision oncology when the mutations are so-called "driver mutations" that contribute to the initiation and / or development of cancer. However, in other cases, mutations present in both the subject's cancer genome and the subject's germline are not beneficial for precision oncology when the mutations are so-called "passenger mutations" that do not contribute to the initiation and / or development of cancer.Similarly, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is beneficial for precision oncology, e.g., if the mutation is a driver mutation and / or the mutation facilitates a therapeutic approach, e.g., by distinguishing cancer cells from normal cells in a therapeutically actionable manner. However, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is not beneficial for precision oncology, e.g., if the mutation is a passenger mutation and / or the mutation fails to distinguish cancer cells from germline cells in a therapeutically actionable manner.
[0074] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the dominant allele represented at that chromosomal locus within a population of a species (e.g., a "wild-type" sequence) or an allele predefined within a reference genome of the species.
[0075] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is not the dominant allele represented at that locus within a population of a species (e.g., not a "wild-type" sequence) or is not a predefined allele within a reference sequence construct for the species (e.g., a reference genome or set of reference genomes). In some cases, sequence isoforms found within a population of a species that do not affect the change in the protein encoded by the genome or that result in amino acid substitutions that do not substantially affect the function of the encoded protein are not variant alleles.
[0076] As used herein, the terms "variant allele fraction," "VAF," "allele fraction," or "AF" refer to the number of times a variant or mutant allele is observed (e.g., the number of reads supporting the candidate variant allele) divided by the total number of times the position was sequenced (e.g., the total number of reads covering the candidate locus).
[0077] As used herein, the term "germline variant" refers to genetic variants inherited from maternal and paternal DNA. Germline variants can be determined through an adapted tumor-normal calling pipeline.
[0078] As used herein, the term "somatic variant" refers to a variant that arises as a result of dysregulation, e.g., mutation, of a cellular process associated with a neoplastic cell. Somatic variants can be detected by subtraction from a matched normal sample.
[0079] As used herein, the term "single nucleotide variant" or "SNV" refers to the substitution of one nucleotide with a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y may be represented as "X>Y." For example, a cytosine to thymine SNV may be represented as "C>T."
[0080] As used herein, the terms "insertions and deletions" or "indels" refer to variants that result from the gain or loss of DNA base pairs within the analyzed region.
[0081] As used herein, the term "copy number variation" or "CNV" refers to a process for detecting large structural changes in the genome associated with tumor aneuploidy and other dysregulation of repair systems. These processes are used to detect large insertions or deletions of entire genomic regions. CNV is defined as a structural insertion or deletion whose size is greater than a certain base pair ("bp"), such as 500 bp.
[0082] As used herein, the term "gene fusion" refers to the product of large-scale chromosomal abnormalities that result in the production of chimeric proteins. These expressed products may be non-functional or may be highly overactive or underactive. This may lead to deleterious effects in cancer, such as a hyperproliferative or anti-apoptotic phenotype.
[0083] As used herein, the term "loss of heterozygosity" refers to the loss of one copy of a segment (e.g., including part or all of one or more genes) of the genome of a diploid subject (e.g., a human) in a tissue of a subject, e.g., cancer tissue, or the loss of one copy of a sequence encoding a functional gene product in the genome of a diploid subject. As used herein, when referring to a metric representing loss of heterozygosity across a subject's genome, the loss of heterozygosity is caused by the loss of one copy of various segments in the subject's genome. Genome-wide loss of heterozygosity can be estimated without sequencing the entire genome of a subject, and methods for such estimation based on gene panel targeted sequencing methodologies have been described in the art. Thus, in some embodiments, a metric representing loss of heterozygosity across the genome of a subject's tissue is expressed as a single value, e.g., a percentage or proportion of the genome. In some cases, a tumor may be composed of various subclonal populations, each of which may have different degrees of loss of heterozygosity across their respective genomes. Thus, in some embodiments, genome-wide loss of heterozygosity of a cancer tissue refers to the average loss of heterozygosity across a heterogeneous tumor population. As used herein, when referring to the metric of loss of heterozygosity for a particular gene, for example, a DNA repair protein (e.g., BRCA1 or BRCA2), such as a protein involved in the homologous DNA recombination pathway, loss of heterozygosity refers to the complete or partial loss of one copy of the gene encoding the protein in the genome of a tissue, and / or a mutation in one copy of the gene that prevents translation of the full-length gene product, such as a frameshift or truncation mutation (generating a premature stop codon in the gene) in the gene of interest. In some cases, tumors are composed of various subclonal populations, each of which may have a different mutation status in the gene of interest. Thus, in some embodiments, loss of heterozygosity for a particular gene of interest is represented by the average loss of heterozygosity for the gene across all sequenced subclonal populations of the cancer tissue.In other embodiments, loss of heterozygosity for a particular gene of interest is represented by the number of unique occurrences of loss of heterozygosity in the gene of interest across all sequenced subclonal populations of the cancer tissue (e.g., the number of unique frameshift and / or truncating mutations in the gene identified in the sequencing data).
[0084] As used herein, the term "microsatellite" refers to a short, repeated sequence of DNA. The minimal nucleotide repeating unit of a microsatellite is referred to as a "repeated unit" or "repeat unit." In some embodiments, the stability of a microsatellite locus is assessed by comparing a metric of the distribution of the number of repeat units at the microsatellite locus to a reference number or distribution.
[0085] As used herein, the term "microsatellite instability" or "MSI" refers to a genetic hypermutability state associated with various cancers resulting from impaired DNA mismatch repair (MMR) in a subject. Among other phenotypes, MSI causes changes in the size of microsatellite loci, e.g., changes in the number of repeat units at a microsatellite locus, during DNA replication. Thus, the size of microsatellite repeats is altered in MSI cancers compared to the size of corresponding microsatellite repeats in the germline of the cancer subject. The term "microsatellite instability-high" or "MSI-H" refers to a cancer (e.g., tumor) state with a significant MMR defect that results in microsatellite loci having lengths significantly different from those of corresponding microsatellite loci in normal cells of the same individual. The term "microsatellite stability" or "MSS" refers to a cancer (e.g., tumor) state without a significant MMR defect, such that there is no significant difference between the lengths of microsatellite loci in cancer cells and those of corresponding microsatellite loci in normal (e.g., non-cancerous) cells of the same individual. The term "microsatellite indeterminate" or "MSE" refers to a cancer (e.g., tumor) state with an intermediate microsatellite length phenotype that cannot be clearly classified as MSI-H or MSS based on the statistical cutoffs used to define these two categories.
[0086] As used herein, the term "gene product" refers to a specific genomic locus, e.g., an RNA (e.g., mRNA or miRNA) or protein molecule transcribed or translated from a specific gene. A genomic locus can be identified using the gene name, chromosomal location, or any other genetic mapping metric.
[0087] As used herein, the terms "expression level," "abundance level," or simply "abundance" refer to the amount of a gene product (an RNA species, e.g., mRNA or miRNA, or a protein molecule) transcribed or translated by a cell, or the average amount of a gene product transcribed or translated across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g., a particular gene. However, in some embodiments, the expression level may refer to the amount of a particular isoform of an mRNA or protein corresponding to a particular gene that gives rise to multiple mRNA or protein isoforms. A genomic locus can be identified using gene name, chromosomal location, or any other genetic mapping metric.
[0088] As used herein, the term "ratio" refers to any comparison of a first metric X, or a first mathematical transform thereof X' (e.g., a measurement of the number of units of a genomic sequence in a first one or more biological samples, or a first mathematical transform thereof), to another metric Y, or a second mathematical transform thereof Y' (e.g., the number of units of each genomic sequence in a second one or more biological samples, or a second mathematical transform thereof), such as X / Y, Y / X, log N (X / Y), log N (Y / X), X' / Y, Y / X', log N (X' / Y) or log N (Y / X'), X / Y', Y' / X, log N (X / Y'), log N (Y' / X), X' / Y', Y' / X', log N (X' / Y') or log N (Y' / X'), where N is any real number greater than 1, and exemplary mathematical transformations of X and Y include, but are not limited to, raising X or Y to the Zth power, multiplying X or Y by a constant Q, where Z and Q are any real numbers, and / or taking the M-based logarithm of X and / or Y, where M is a real number greater than 1. In one non-limiting example, X is multiplied by squaring X (X2 ) before the ratio calculation, and Y is calculated by raising Y to the power of 3.2 (Y 3.2 ) and the ratio of X and Y is calculated as log2(X' / Y').
[0089] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound, e.g., a gene product (RNA species, e.g., mRNA or miRNA, or a protein molecule) or a nucleic acid fragment with a particular property (e.g., aligning to a particular locus or encompassing a particular allele), measured in a sample to a second amount of the compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a species of a compound to the total amount of the compound in the same sample. For example, it is the ratio of the amount of mRNA transcripts encoding a particular gene in the sample (e.g., aligning to a particular region of the exome) to the total amount of mRNA transcripts in the sample. In other embodiments, relative abundance refers to the ratio of the amount of a compound or species of a compound in a first sample to the amount of the compound of that species of compound in a second sample. For example, it is the ratio of the normalized amount of mRNA transcripts encoding a particular gene in the first sample to the normalized amount of mRNA transcripts encoding the particular gene in the second sample and / or reference sample.
[0090] As used herein, the terms "sequencing," "determining a sequence," and the like refer to any biochemical process that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.
[0091] As used herein, the term "gene sequence" refers to a record of the series of nucleotides present in a subject's RNA or DNA as determined by sequencing nucleic acid from the subject.
[0092] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any nucleic acid sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (a "single-end read") or from both ends of a nucleic acid fragment (e.g., a paired-end read, a double-end read). The length of a sequence read is often related to a particular sequencing technology. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., an average, median, or mean length of about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads are an average, median, or mean length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, Iore® sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina® parallel sequencing can provide sequence reads that are less variable, e.g., the majority of sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a series of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, for example, hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.
[0093] As used herein, the term "read segment" refers to any form of nucleotide sequence read, including raw sequence reads obtained directly from nucleic acid sequencing technology or sequences derived therefrom, e.g., aligned sequence reads, collapsed sequence reads, or stitched sequence reads.
[0094] As used herein, the term "number of reads" refers to the total number of nucleic acid reads generated, which may or may not equal the number of nucleic acid molecules generated during a nucleic acid sequencing reaction.
[0095] As used herein, the terms "read depth," "sequencing depth," or "depth" may refer to the total number of unique nucleic acid fragments encompassing a particular locus or region of a subject's genome that are sequenced in a particular sequencing reaction. Sequencing depth can be expressed as "Y-fold," e.g., 50-fold, 100-fold, etc., where "Y" refers to the number of unique nucleic acid fragments encompassing a particular locus that are sequenced in a sequencing reaction. In such cases, Y is necessarily an integer, since it represents the actual sequencing depth of a particular locus. Alternatively, read depth, sequencing depth, or depth may refer to a representative value (e.g., an average or mode) of the number of unique nucleic acid fragments encompassing one of multiple loci or regions of a subject's genome that are sequenced in a particular sequencing reaction. For example, in some embodiments, sequencing depth refers to the average depth of all loci across a chromosome arm, a targeted sequencing panel, an exome, or the entire genome. In such cases, Y may be expressed as a fraction or decimal to refer to the average coverage across multiple loci. When an average depth is listed, the actual depth of any particular locus may differ from the overall listed depth. A metric can be determined that provides a range of sequencing depths that encompasses a defined percentage of the total number of loci. For example, a range of sequencing depths that encompasses 90%, 95%, or 99% of the loci. As will be understood by those skilled in the art, various sequencing technologies provide various sequencing depths. For example, low-pass whole genome sequencing can refer to a technology that provides a sequencing depth of less than 5x, less than 4x, less than 3x, or less than 2x, for example, about 0.5x to about 3x.
[0096] As used herein, the term "sequencing breadth" refers to a particular reference exome (e.g., a human reference exome), a particular reference genome (e.g., a human reference genome), or what fraction of a portion of an exome or genome has been analyzed. Sequencing breadth can be expressed as a fraction, decimal, or percentage and is generally calculated as (number of loci analyzed / total number of loci in the reference exome or reference genome). The denominator of the fraction can be the repeat-masked genome, and thus 100% can correspond to the entire reference genome minus the masked portion. A repeat-masked exome or genome can refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads align with unmasked portions of the exome or genome). In some embodiments, any portion of the exome or genome can be masked, and thus sequencing breadth can be assessed for any desired portion of the reference exome or genome. In some embodiments, "wide-ranging sequencing" refers to sequencing / analysis of at least 0.1% of the exome or genome.
[0097] As used herein, the terms "sequence ratio" and "coverage ratio" refer interchangeably to any measurement of the number of units of a genome sequence in a first one or more biological samples (e.g., a test sample and / or a tumor sample) compared to the number of units of the respective genome sequence in a second one or more biological samples (e.g., a reference sample and / or a control sample). In some embodiments, the sequence ratio is a copy ratio, a log2-transformed copy ratio (e.g., a log2 copy ratio), a coverage ratio, a base ratio, an allele ratio (e.g., a variant allele ratio), and / or a tumor ploidy. In some embodiments, the sequence ratio is a log N is the conversion copy ratio, where N is any real number greater than 1.
[0098] As used herein, the term "sequencing probe" refers to a molecule that binds to a nucleic acid with an affinity based on the predicted nucleotide sequence of the RNA or DNA present at that locus.
[0099] As used herein, the term "target panel" or "target gene panel" refers to a combination of probes for sequencing (e.g., by next-generation sequencing) nucleic acids present in a biological sample from a subject (e.g., a tumor sample, a liquid biopsy sample, a germline tissue sample, a white blood cell sample, or a tumor or tissue organoid sample) selected to map to one or more loci of interest on one or more chromosomes. Exemplary loci / gene sets that can be analyzed using a targeted panel and are useful for, e.g., precision oncology via solid or liquid biopsy assays, are listed in Table 1. Another exemplary loci / gene sets that can be analyzed using a targeted panel and are useful for, e.g., precision oncology via solid or liquid biopsy assays, are listed in Table 2. In some embodiments, in addition to loci useful for precision oncology, the targeted panel includes one or more probes for sequencing one or more of loci associated with different disease states, loci used for internal control purposes, or loci from pathogenic organisms (e.g., oncogenic pathogens).
[0100] As used herein, the term "reference exome" refers to any sequenced or otherwise characterized exome of any tissue from any organism or pathogen, whether partial or complete, that can be used to reference an identified sequence from a subject. Typically, the reference exome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference exomes for human subjects, as well as many other organisms, are provided in the online genome browser hosted by the National Center for Biotechnology Information ("NCBI"). "Exome" refers to the complete transcriptional profile of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference exome is often an assembled or partially assembled exome sequence from an individual or multiple individuals. In some embodiments, a reference exome is an assembled or partially assembled exome sequence from one or more human individuals. A reference exome can be considered a representative example of a species' set of expressed genes. In some embodiments, a reference exome includes sequences assigned to chromosomes.
[0101] As used herein, the term "reference genome" refers to any sequenced or otherwise characterized genome of any organism or pathogen, whether partial or complete, that can be used to reference identified sequences from a subject. Typically, a reference genome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference genomes used for human subjects, as well as many other organisms, are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz ("UCSC"). "Genome" refers to the complete genetic information of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of a species' set of genes. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38). For haploid genomes, only one nucleotide may be present at each locus. For diploid genomes, heterozygous loci may be identified, and each heterozygous locus may have two alleles, with either allele being able to match for alignment to the locus.
[0102] As used herein, the term "bioinformatics pipeline" refers to a series of processing steps used to determine characteristics of a subject's genome or exome based on sequencing data of the subject's genome or exome. A bioinformatics pipeline can be used to determine characteristics of a subject's germline genome or exome and / or a subject's cancer genome or exome. In some embodiments, the pipeline extracts information related to genomic alterations in a subject's cancer genome, which is useful for guiding clinical decisions for precision oncology from sequencing results of a subject-derived biological sample, such as a tumor sample, a liquid biopsy sample, or a reference normal sample. Certain processing steps in bioinformatics can be "connected," meaning that the results of a first respective processing step are useful and / or essential for the execution of a second downstream processing step. For example, in some embodiments, a bioinformatics pipeline includes a first respective processing step for identifying genomic alterations unique to a subject's cancer genome, and a second respective processing step for determining a metric useful for precision oncology, such as tumor mutation burden, using the amount and / or identity of the identified genomic alterations. In some embodiments, the bioinformatics pipeline includes a reporting stage that generates a report of relevant and / or actionable information identified by upstream stages of the pipeline, which may or may not further include recommendations to support clinical therapy decisions.
[0103] As used herein, the term "limit of detection" or "LOD" refers to the minimum amount of a feature that can be identified with a certain level of confidence. Thus, the level of detection can be used to describe the amount of a substance that must be present for a particular assay to reliably detect it. The level of detection can also be used to describe the level of support required for an algorithm to reliably identify genomic alterations based on sequencing data. For example, the minimum number of unique sequence reads required to support the identification of sequence variants such as SNVs.
[0104] As used herein, the terms "BAM file" or "binary file containing an alignment map" refer to a file that stores sequencing data aligned to a reference sequence (e.g., a reference genome or exome). In some embodiments, a BAM file is a compressed binary version of a SAM (sequence alignment map) file that contains, for each of a plurality of unique sequence reads, an identifier for the sequence read, information about the nucleotide sequence, information about the alignment of the sequence to the reference sequence, and optionally metrics about the quality of the sequence read and / or the quality of the sequence alignment. While a BAM file generally relates to a file having a particular format, for brevity, unless otherwise specified, it is used herein to simply refer to a file of any format that contains information about sequence alignments.
[0105] As used herein, the term "representative value" refers to a central or representative value of a distribution of values. Non-limiting examples of representative values include the arithmetic mean, weighted mean, median, central hinge, trimean, geometric mean, geometric median, Winsorized mean, median, and mode of a distribution of values.
[0106] As used herein, the term "dispersion" refers to the degree to which data points in a dataset fitted to a distribution vary or spread out from the mean (or another representative value) of the fitted distribution. Variance, standard deviation, mean absolute deviation (MAD), interquartile range (IQR), range, coefficient of variation (CV), median absolute deviation (MAD), skewness (which measures asymmetry, but is also related to the spread of data in a distribution), kurtosis (which measures "tail spread," but also describes aspects of how spread out the data is in a distribution), and Gini coefficient (which measures the inequality or variation of a distribution) are non-limiting examples of dispersion.
[0107] As used herein, the term "positive predictive value" or "PPV" refers to the likelihood that a variant will be correctly called, given that the variant was called by the assay. PPV can be expressed as (number of true positives) / (number of false positives + number of true positives).
[0108] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, e.g., a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining copy number variations of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to one of skill in the art can be used to detect any of the nucleic acid characteristics described herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of a mutation in a nucleic acid at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). Assays or methods can have a particular sensitivity and / or specificity, and their relative utility as diagnostic tools can be measured using ROC-AUC statistics.
[0109] As used herein, the term "classification" may refer to any number or other letter associated with a particular characteristic of a sample. For example, in some embodiments, the term "classification" may refer to the type of cancer in a subject, the stage of cancer in a subject, the prognosis of cancer in a subject, tumor burden, the presence of tumor metastasis in a subject, and the like. Classifications may be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" may refer to a predetermined number used in a calculation. For example, a cutoff size may refer to a size above which fragments are excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms may be used in either of these contexts.
[0110] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.
[0111] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers indicative of cancer.
[0112] As used herein, "actionable genomic alteration" or "actionable variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor proportion) that is known or believed to be associated with a therapeutic course of action that is more likely to result in a positive outcome in cancer patients with an actionable variant than in similarly located cancer patients without an actionable variant. For example, administration of an EGFR inhibitor (e.g., afatinib, erlotinib, gefitinib) is more effective in treating non-small cell lung cancer in patients with an EGFR mutation in exons 19 / 21 than in patients without an EGFR mutation in exons 19 / 21. Thus, an EGFR mutation in exons 19 / 21 is an actionable variant. In some cases, an actionable variant is associated with improved treatment outcomes only in one specific cancer type or a group of specific cancer types. In other cases, actionable variants are associated with improved treatment outcomes in virtually all cancer types.
[0113] As used herein, "variant of uncertain significance" or "VUS" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) whose impact on disease development / progression is unknown.
[0114] As used herein, "benign variant" or "likely benign variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed not to contribute to disease development / progression.
[0115] As used herein, "pathogenic variant" or "likely pathogenic variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed to contribute to disease development / progression.
[0116] As used herein, an "effective amount" or "therapeutically effective amount" is an amount sufficient to affect beneficial or desired clinical results during treatment. An effective amount can be administered to a subject in one or more doses. In terms of treatment, an effective amount is an amount sufficient to palliate, improve, stabilize, reverse, or slow the progression of a disease, or otherwise reduce the pathological consequences of a disease. An effective amount is generally determined by a physician on a case-by-case basis and is within the skill of one of ordinary skill in the art. Several factors are typically considered when determining an appropriate dosage to achieve an effective amount. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.
[0117] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used in the description of the present invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be as inclusive as the term "comprising."
[0118] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "response to determining" or "response to detecting," depending on the context. Similarly, the phrase "when determined" or "when [a described condition or event] is detected" may be interpreted to mean "upon determining," or "response to determining," or "upon detection [of [a described condition or event]," or "response to detecting [a described condition or event]," depending on the context.
[0119] Also, while terms such as first, second, etc. may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both subjects, but are not the same subject. Furthermore, the terms "subject," "user," and "patient" are used interchangeably herein.
[0120] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure, including example systems, methods, techniques, instruction sequences, and computing machine program products embodying illustrative embodiments. However, the following illustrative discussion is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. Features described herein are not limited by the illustrated ordering of acts or events, as some acts may occur in different orders and / or contemporaneously with other acts or events.
[0121] The embodiments provided herein are chosen and described to best explain the principles and their practical applications, thereby enabling those skilled in the art to best utilize the various embodiments with various modifications suited to the particular uses contemplated. In some instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. In other instances, it will be apparent to one skilled in the art that the present disclosure may be practiced without one or more of the specific details.
[0122] It will be understood that in the development of any such actual implementation, numerous implementation-specific decisions will be made to achieve the designer's specific goals, such as adherence to use case and business-related constraints, and that these specific goals will vary from implementation to implementation and from designer to designer. It will further be understood that such a design effort may be complex and time-consuming, but would nevertheless be a routine undertaking for one of ordinary skill in the art having the benefit of this disclosure.
[0123] 1 illustrates an exemplary system embodiment.
[0124] Having provided an overview of some aspects of the present disclosure and some definitions used herein, details of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay will now be described in conjunction with FIGS. 1A-1D. Collectively, FIGS. 1A-1D illustrate configurations of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay, according to some embodiments of the present disclosure. Advantageously, the exemplary system shown in FIGS. 1A-1D improves upon conventional methods for providing clinical support for personalized cancer therapy by determining tumor-informed circulating tumor fraction estimates using a non-custom liquid biopsy assay.
[0125] 1A is a block diagram illustrating a system according to some embodiments. Device 100 in some embodiments includes one or more processing units (CPUs) 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including, for example, a display 108 and / or input 110 (e.g., a mouse, touchpad, keyboard, etc.), non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. One or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes a CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU 102. Persistent memory 112 and the non-volatile memory devices within non-persistent memory 112 comprise non-transitory computer-readable storage media. In some embodiments, non-persistent memory 111, or alternatively, non-transitory computer-readable storage media, sometimes in conjunction with persistent memory 112, store the following programs, modules, and data structures, or a subset thereof: an operating system 116, which includes procedures for handling various basic system services and for performing hardware-dependent tasks; a network communication module (or instructions) 118 for connecting the system 100 with other devices and / or communication networks 105; a study patient data store 120 for storing one or more sets of features from a patient (e.g., subject); a bioinformatics module 140 for processing the sequencing data and extracting features from the sequencing data, e.g., from a liquid biopsy sequencing assay; a feature analysis module 160 for assessing patient features, such as genomic alterations, complex genomic features, and clinical features; and A reporting module 180 for generating and transmitting reports that provide clinical support for personalized cancer therapy.
[0126] While FIGS. 1A-1D depict "system 100," the diagrams are intended as a functional illustration of various features that may be present in a computer system, rather than as a structural schematic of the embodiments described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art. Furthermore, while FIG. 1 depicts certain data and modules in non-persistent memory 111, some or all of these data and modules may be in persistent memory 112. For example, in various embodiments, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for implementing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various embodiments.
[0127] In some implementations, non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of system 100 that is addressable by system 100 such that system 100 can retrieve all or a portion of such data when needed.
[0128] 1A, system 100 is depicted as a single computer containing all of the functionality for providing clinical support for personalized cancer therapy. However, while a single machine is illustrated, the term "system" should also be construed to include any collection of machines that individually or jointly execute a set (or sets) of instructions for implementing any one or more of the methodologies discussed herein.
[0129] For example, in some embodiments, system 100 includes one or more computers. In some embodiments, functionality for providing clinical support for personalized cancer therapy is spread across any number of networked computers and / or resident on each of several networked computers and / or hosted on one or more virtual machines at remote locations accessible over communications network 105. For example, different portions of the various modules and data stores illustrated in Figures 1A-1D may be stored and / or executed on various instances of processing devices and / or processing servers / databases (e.g., processing devices 224, 234, 244, and 254, processing server 262, and database 264) within distributed diagnostic environment 210 illustrated in Figure 2B.
[0130] The system may operate in the capacity of a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment. The system may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, server, network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine.
[0131] In another embodiment, the system comprises a virtual machine including modules for executing instructions for performing any one or more of the methodologies disclosed herein. In computing, a virtual machine (VM) is an emulation of a computer system that is based on a computer architecture and provides the functionality of a physical computer. Some such embodiments may involve dedicated hardware, software, or a combination of hardware and software.
[0132] Those skilled in the art will appreciate that any of a wide variety of different computer topologies may be used in an application, and that all such topologies are within the scope of the present disclosure.
[0133] Examination patient data storage unit (120)
[0134] Referring to FIG. 1B , in some embodiments, the system (e.g., system 100) includes a patient data store 120 that stores data for patients 121-1 through 121-M (e.g., cancer patients or patients being tested for cancer), including one or more of sequencing data 122, feature data 125, and clinical assessments 139. These data are used and / or generated by various processes stored in bioinformatics module 140 and feature analysis module 160 of system 100 to generate reports that ultimately provide clinical support for personalized cancer therapy for the patients. While the feature range of patient data 121 across all patients may be informationally dense, an individual patient's feature set may be sparse across the collective feature range of all features across all patients. That is, data stored for one patient may include a different set of features than data stored for another patient. Furthermore, although illustrated as a single data structure in FIG. 1B , different sets of patient data may be stored in different databases or modules spread across one or more system memories.
[0135] In some embodiments, sequencing data 122 from one or more sequencing reactions 122-i, including multiple sequence reads 123-i-1 through 123-iK, is stored in a test patient data store 120. The data store may contain different sets of sequencing data from a single subject corresponding to different samples from the patient, e.g., tumor samples, liquid biopsy samples, tumor organoids derived from the patient's tumor, and / or normal samples, and / or samples acquired at different times, e.g., while monitoring the progression, regression, remission, and / or recurrence of cancer in the subject. The sequence reads may be in any suitable file format, e.g., BCL, FASTA, FASTQ, etc. In some embodiments, the sequencing data 122 is accessed by a sequencing data processing module 141, which performs various preprocessing, genome alignment, and demultiplexing operations, as described in detail below with reference to a bioinformatics module 140. In some embodiments, sequence data aligned to a reference construct, eg, BAM file 124, is stored in the test patient data store 120.
[0136] In some embodiments, the test patient data store 120 includes feature data 125 that is useful, for example, for identifying clinical support for personalized cancer therapy. In some embodiments, the feature data 125 includes patient personal characteristics 126, such as patient name, date of birth, gender, ethnicity, physical address, smoking status, alcohol consumption characteristics, anthropomorphic data, etc.
[0137] In some embodiments, the feature data 125 includes patient medical history data 127, such as cancer diagnosis information (e.g., date of initial diagnosis, date of metastasis diagnosis, cancer staging, tumor characterization, tissue of origin, previous treatments and outcomes, adverse effects of treatments, history of treatment groups, history of clinical trials, previous and current medications, history of surgery, etc.), previous or current symptoms, previous or current treatments, previous treatment outcomes, previous disease diagnoses, diabetic status, depression diagnosis, other physical or mental illness diagnoses, and family history. In some embodiments, the feature data 125 includes clinical features 128, such as pathology data 128-1, medical imaging data 128-2, and tissue culture and / or tissue organoid culture data 128-3.
[0138] In some embodiments, additional clinical features, such as previous laboratory test results, are stored in the laboratory patient data store 120. Historical data 127 and clinical features may be collected from the patient from a variety of sources, including directly from the patient's electronic medical record (EMR) or electronic health record (EHR), or curated from other sources, such as fields from various laboratory records (e.g., gene sequencing reports).
[0139] In some embodiments, the feature data 125 includes genomic features 131 of the patient. Non-limiting examples of genomic features include allele status 132 (e.g., allele identity at one or more loci, support for wild-type or variant alleles at one or more loci, support for SNV / MNV at one or more loci, support for indels at one or more loci, and / or support for genetic rearrangements at one or more loci), allele fraction 133 (e.g., the ratio of variant to reference alleles (or vice versa)), methylation status 132 (e.g., the distribution of methylation patterns at one or more loci and / or support for abnormal methylation patterns at one or more loci), genome copy number 135 (e.g., copy number values at one or more loci and / or support for abnormal (increased or decreased) copy number at one or more loci), tumor mutational burden 136 (e.g., a measure of the number of mutations in a subject's cancer genome), and microsatellite instability status 137 (e.g., a measure of repeat unit length at one or more microsatellite loci and / or a classification of MSI status for a patient's cancer). In some embodiments, one or more of genomic features 131 are determined by a nucleic acid bioinformatics pipeline, e.g., as described in more detail below with reference to Figures 4A-4E. In particular, in some embodiments, feature data 125 includes circulating tumor fraction estimates 131-i determined using improved methods for determining circulating tumor fraction estimates, described in more detail below with reference to Figures 1C, 1D, and 4E. In some embodiments, one or more of genomic features 131 are obtained from an external testing source that is not connected to the bioinformatics pipeline, e.g., as described below.
[0140] 1B , in some embodiments, feature data 125 further includes data 138 from other "-omics" research fields. Non-limiting examples of "-omics" research fields that may yield feature data useful for providing clinical support for personalized cancer therapy include transcriptomics, epigenomics, proteomics, metabolomics, metabonomics, microbiomics, lipidomics, glycomics, cellomics, and organoidomics.
[0141] In some embodiments, still other features may include, but are not limited to, those listed above, including features derived from machine learning approaches based at least in part on the evaluation of any relevant molecular or clinical features, considered alone or in combination. For example, in some embodiments, one or more latent features learned from the evaluation of a cancer patient training dataset improve the diagnostic and prognostic capabilities of various analysis algorithms in feature analysis module 160.
[0142] Those skilled in the art will know of other types of features that are useful for providing clinical support for personalized cancer therapy. The above list of features is merely representative and should not be construed as limiting.
[0143] In some embodiments, the test patient data store 120 includes clinical assessment data 139 for the patient, e.g., based on the feature data 125 collected for the subject. In some embodiments, the clinical assessment data 139 includes a catalog of actionable variants and features 139-1 (e.g., genomic alterations and composite metrics based on genomic features known or thought to be targetable by one or more specific cancer therapies), matched therapies 139-2 (e.g., therapies known or thought to be particularly beneficial for treating subjects with actionable variants), and / or clinical reports 139-3 generated for the subject, e.g., based on the identified actionable variants and features 139-1 and / or matched therapies 139-2.
[0144] In some embodiments, clinical assessment data 139 is generated by analysis of feature data 125 using various algorithms in feature analysis module 160, as described in further detail below. In some embodiments, clinical assessment data 139 is generated, modified, and / or verified by evaluation of feature data 125 by a clinician, e.g., an oncologist. For example, in some embodiments, a clinician (e.g., in clinical environment 220) uses feature analysis module 160 or directly accesses laboratory patient data store 120 to evaluate feature data 125 and recommend a personalized cancer treatment for the patient. Similarly, in some embodiments, a clinician (e.g., in clinical environment 220) reviews recommendations determined using feature analysis module 160 and, for example, approves, rejects, or modifies the recommendations before they are sent to a medical professional treating the cancer patient.
[0145] Bioinformatics module (140).
[0146] Referring again to FIG. 1A, the system (e.g., system 100) includes a bioinformatics module 140, which includes a feature extraction module 145 and optional auxiliary data processing constructs, such as a sequence data processing module 141 and / or one or more reference sequence constructs 158 (e.g., a reference genome, exome, or target panel construct that includes reference sequences for multiple loci targeted by the sequencing panel).
[0147] In some embodiments, bioinformatics module 140 includes a sequence data processing module 141 that includes instructions for processing sequence reads, e.g., raw sequence reads 123 from one or more sequencing reactions 122, prior to analysis by various feature extraction algorithms, as described in detail below. In some embodiments, sequence data processing module 141 includes one or more preprocessing algorithms 142 that prepare the data for analysis. In some embodiments, preprocessing algorithm 142 includes instructions for converting the file format of sequence reads from the output of a sequencer (e.g., a BCL file format) into a file format compatible with downstream analysis of the sequences (e.g., a FASTQ or FASTA file format). In some embodiments, the preprocessing algorithm 142 includes instructions for assessing the quality of sequence reads (e.g., by examining quality metrics such as Phred score, base calling error probability, quality (Q) score, and the like) and / or removing sequence reads that do not meet a threshold quality (e.g., an estimated base calling accuracy of at least 80%, at least 90%, at least 95%, at least 99%, at least 99.5%, at least 99.9%, or more). In some embodiments, the preprocessing algorithm 142 includes instructions for filtering sequence reads for one or more properties, for example, removing sequences that fail to meet a lower or upper size threshold or removing duplicate sequence reads.
[0148] In some embodiments, the sequence data processing module 141 includes one or more alignment algorithms 143 for aligning preprocessed sequence reads 123 to a reference sequence construct 158, such as a reference genome, exome, or target panel construct. Many algorithms for aligning sequencing data to a reference construct, such as BWA, Blat, SHRiMP, LastZ, and MAQ, are known in the art. One example of a sequence read alignment package is the Burrows-Wheeler Alignment Tool (BWA), which uses the Burrows-Wheeler Transform (BWT) to align short sequence reads to a large reference construct, allowing for mismatches and gaps. The sequence read alignment package imports raw or preprocessed sequence reads 122, for example, in BCL, FASTA, or FASTQ file format, and outputs aligned sequence reads 124, for example, in SAM or BAM file format.
[0149] In some embodiments, the sequence data processing module 141 includes one or more demultiplexing algorithms 144 to split sequence read or sequence alignment files generated from pooled nucleic acid sequencing reactions into separate sequence read or sequence alignment files, each corresponding to a different source of nucleic acid in the nucleic acid sequencing pool. For example, due to the cost of sequencing reactions, it is common practice to pool nucleic acids from multiple samples into a single sequencing reaction. Nucleic acids from each sample are tagged with sample-specific and / or molecule-specific sequence tags (e.g., UMIs) that are sequenced along with the molecule. In some embodiments, the demultiplexing algorithm 144 sorts these sequence tags within the sequence read or sequence alignment files and demultiplexes the sequencing data into separate files for each sample included in the sequencing reaction.
[0150] The bioinformatics module 140 includes a feature extraction module 145 that includes instructions for identifying diagnostic features, e.g., genomic features 131, from sequencing data 122 of one or more biological samples from a subject, e.g., a solid tumor sample, a liquid biopsy sample, or a normal tissue (e.g., control) sample. For example, in some embodiments, the feature extraction algorithm compares the identity of one or more nucleotides at a locus from the sequencing data 122 to the identity of the nucleotide at that locus in a reference sequence construct (e.g., a reference genome, exome, or targeted panel construct) to determine whether the subject has a variant at that locus. In some embodiments, the feature extraction algorithm evaluates data other than the raw sequence to identify genomic alterations in the subject, e.g., allele ratios, relative copy numbers, repeat unit distribution, etc.
[0151] For example, in some embodiments, feature extraction module 145 includes one or more variant identification modules that include instructions for various variant calling processes. In some embodiments, variants in a subject's germline are identified using, for example, germline variant identification module 146. In some embodiments, variants in a cancer genome, e.g., somatic variants, are identified using, for example, somatic variant identification module 150. Although separate germline and somatic variant identification modules are illustrated in FIG. 1A, in some embodiments, they are integrated into a single module. In some embodiments, the variant identification module includes instructions for identifying one or more of nucleotide variants (e.g., single nucleotide variants (SNVs) and multiple nucleotide variants (MNVs)) using one or more SNV / MNV calling algorithms (e.g., algorithms 147 and / or 151), indels (e.g., nucleotide insertions or deletions) using one or more indel calling algorithms (e.g., algorithms 148 and / or 152), and genomic rearrangements (e.g., nucleotide sequence inversions, translocations, and fusions) using one or more genomic rearrangement calling algorithms (e.g., algorithms 149 and / or 153).
[0152] The SNV / MNV algorithm 147 can identify single-nucleotide substitutions occurring at specific positions in the genome. For example, at a specific base position or locus in the human genome, a C nucleotide may occur in most individuals, while in a minority of individuals, the position is occupied by an A. This means that a SNP exists at this specific position, and the two possible nucleotide variations, C or A, are said to be alleles for this position. SNPs underlie differences in human susceptibility to a wide range of diseases (e.g., sickle cell anemia, beta-thalassemia, and cystic fibrosis resulting from SNPs). Disease severity and how the body responds to treatment are also manifestations of genetic variation. For example, a single-nucleotide mutation in the APOE (apolipoprotein E) gene is associated with a lower risk of Alzheimer's disease. A single-nucleotide variant (SNV) is a variation in a single nucleotide without any frequency restriction and can occur somatically. Somatic single-nucleotide mutations (e.g., caused by cancer) may also be referred to as single-nucleotide changes. MNP (multiple nucleotide polymorphism) modules can identify substitutions of consecutive nucleotides at specific positions in the genome.
[0153] The indel calling algorithm 148 can identify insertions or deletions of bases in the genome of organisms classified among minor genetic variations. Indels typically measure 1 to 10,000 base pairs in length, while microindels are defined as indels resulting in a net change of 1 to 50 nucleotides. Indels can be contrasted with SNPs or point mutations. While indels insert and / or delete nucleotides from a sequence, point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels, which are insertions and / or deletions, can be used as genetic markers in natural populations, particularly in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except near highly repetitive regions, including homopolymers and microsatellites.
[0154] Genome rearrangement algorithms149 can identify hybrid genes formed from two previously separated genes. This can occur as a result of translocations, interstitial deletions, or chromosomal inversions. Gene fusions can play an important role in tumorigenesis. Fusion genes can contribute to tumorigenesis because they can produce abnormal proteins that are much more active than non-fusion genes. Often, fusion genes are cancer-causing oncogenes, including BCR-ABL, TEL-AML1 (ALL with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG, an interstitial deletion on chromosome 21 that frequently occurs in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by disrupting androgen receptor (AR) signaling and inhibiting AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematologic cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a fusion gene specific and characteristic of high-grade serous ovarian cancer. Oncogenic fusion genes can lead to gene products with new or distinct functions from the two fusion partners. Alternatively, proto-oncogenes can be fused to strong promoters, thereby setting up oncogenic function through upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where oncogenes are juxtaposed to immunoglobulin gene promoters. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events. Because chromosomal translocations play such an important role in neoplasia, a dedicated database of chromosome aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer.
[0155] In some embodiments, feature extraction module 145 includes instructions for identifying one or more complex genomic alterations (e.g., features that incorporate more than changes in the primary sequence of the genome) in a subject's cancer genome. For example, in some embodiments, feature extraction module 145 includes modules for identifying one or more of copy number variation (e.g., copy number variation analysis module 153), microsatellite instability status (e.g., microsatellite instability analysis module 154), tumor mutation burden (e.g., tumor mutation burden analysis module 155), tumor ploidy (e.g., tumor ploidy analysis module 156), and homologous recombination pathway deficiency (e.g., homologous recombination pathway analysis module 157).
[0156] For example, referring to FIG. 1D , in some embodiments, feature extraction module 145 includes tumor fraction estimation module 145-tf. In some embodiments, tumor fraction estimation module 145-tf includes sequence ratio data structure 145-tf-r that includes a plurality of sequence ratios (e.g., coverage ratios) obtained from sequencing a subject's test liquid biopsy sample. In some embodiments, sequence ratio data structure 145-tf-r includes sequence ratios that are used as inputs to determine a tumor fraction estimate for the test liquid biopsy sample. In some embodiments, tumor fraction estimation module 145-tf also includes tumor purity algorithm construct 145-tf-a, which, for example, performs maximum likelihood estimation (e.g., an expectation-maximization algorithm) to calculate an estimate of circulating tumor fraction. The tumor purity algorithm construct 145-tf-a includes an optional input data filtration construct 145-tf-k (e.g., for filtering out one or more inputs migrated from the sequence ratio data structure based on a minimum probe threshold or a location on a sex chromosome), and a plurality of model parameters 145-tf-d (e.g., 145-tf-d-1, 145-tf-d-2, ...) used to execute the algorithm. In some embodiments, the model parameters include: a predicted sequence ratio for a set of copy states at a given tumor purity; a distance (e.g., error) from a test sequence ratio to the most recent predicted sequence ratio at a given tumor purity; a minimum distance (e.g., minimum error) from a test sequence ratio to the most recent predicted sequence ratio at a given tumor purity (e.g., an assigned test copy state selected from the minimum distance predicted copy states); and / or a tumor purity score (e.g., a weighted sum of errors).
[0157] 1C , tumor fraction estimation module 145-tf is used to obtain one or more circulating tumor fraction estimates 131-i, which are included as feature data 125 in test patient data store 120. For example, in some embodiments, multiple circulating tumor fraction estimates are obtained from test liquid biopsy samples of subjects 131-i-cf (e.g., 131-i-cf-1, 131-i-cf-2, ..., 131-i-cf-N). In some embodiments, multiple circulating tumor fraction estimates are obtained from a single patient at different collection times.
[0158] A feature analysis module (160).
[0159] 1A , the system (e.g., system 100) includes a feature analysis module 160, which includes one or more genomic variation interpretation algorithms 161, one or more optional clinical data analysis algorithms 165, an optional treatment curation algorithm 165, and an optional recommendation validation module 167. In some embodiments, feature analysis module 160 uses one or more analysis algorithms (e.g., algorithms 162, 163, 164, and 165) to evaluate feature data 125 to identify actionable variants and characteristics 139-1 and corresponding matched treatments 139-2 and / or clinical trials. The identified actionable variants and characteristics 139-1 and corresponding matched treatments 139-2, which are optionally stored in test patient data store 120, are then curated by feature analysis module 160 to generate a clinical report 139-3, which is optionally verified by a user, e.g., a clinician, before being sent to a medical professional, e.g., an oncologist, treating the patient.
[0160] In some embodiments, the genomic variation interpretation algorithm 161 includes instructions for, e.g., evaluating the impact of one or more genomic features 131 of a subject identified by the feature extraction module 145 on the characteristics of the patient's cancer and / or whether one or more targeted cancer therapies may improve the patient's clinical outcome. For example, in some embodiments, the one or more genomic variant analysis algorithms 163 evaluate various genomic features 131 by querying a database, e.g., a look-up table (“LUT”), of actionable genomic alterations, targeted therapies associated with the actionable genomic alterations, and any other conditions that should be met before administering the targeted therapy to a subject with an actionable genomic alteration. For example, evidence suggests that depatuxizumab mafodotin (an anti-EGFR mAb conjugated to monomethyl auristatin F) has improved efficacy for treating recurrent glioblastoma with EGFR focal amplification. van den Bent et al., 2017, Cancer Chemother Pharmacol., 80(6):1209-17. Thus, the LUT of actionable genomic alterations has an entry for a focal amplification of the EGFR gene, indicating that depatuxizumab mafodotin is a targeted therapy for glioblastomas with focal gene amplification (e.g., recurrent glioblastomas). In some cases, the LUT may also include contraindications to the associated targeted therapy, such as adverse drug interactions or personal characteristics that contraindicate administration of a particular targeted therapy.
[0161] In some embodiments, the genomic alteration interpretation algorithm 161 determines whether a particular genomic feature 131 should be reported to a medical professional treating a cancer patient. In some embodiments, a genomic feature 131 (e.g., genomic alterations and composite features) is reported if there is clinical evidence that the feature significantly impacts cancer biology, affects cancer prognosis, and / or impacts pharmacogenomics, for example, by indicating or contraindicating a particular treatment approach. For example, the genomic variation interpretation algorithm 161 may classify a particular CNV feature 135 as "reportable," meaning, e.g., that the CNV has been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "non-reportable," meaning, e.g., that the CNV has not been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "no evidence," meaning, e.g., that there is no evidence to support the CNV being "reportable" or "non-reportable"; or as "conflicting evidence," meaning, e.g., that there is evidence to support both the CNV being "reportable" and the CNV being "non-reportable."
[0162] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more pathogenic variant analysis algorithms 162 that evaluate various genomic features to identify the presence of an oncogenic pathogen associated with the patient's cancer and / or targeted therapies associated with oncogenic pathogen infection in the cancer. For example, RNA expression patterns in some cancers are associated with the presence of oncogenic pathogens that are instrumental in inducing the cancer. See, e.g., U.S. Patent No. 11,043,304, the contents of which are incorporated by reference herein in their entirety for all purposes. In some cases, recommended treatments for cancer differ when the cancer is associated with oncogenic pathogen infection than when it is not. Thus, in some embodiments, for example, if feature data 125 includes RNA abundance data for the patient's cancer, one or more pathogenic variant analysis algorithms 162 evaluate the RNA abundance data for the patient's cancer to determine whether a signature indicative of the presence of an oncogenic pathogen in the cancer is present in the data. Similarly, in some embodiments, bioinformatics module 140 includes an algorithm that searches for the presence of pathogenic nucleic acid sequences in sequencing data 122. See, for example, U.S. Patent Application No. 17 / 800,492, filed August 17, 2022, entitled "Systems and Methods for Detecting Viral DNA from Sequencing," the contents of which are incorporated by reference herein in their entirety for all purposes. Accordingly, in some embodiments, one or more pathogenic variant analysis algorithms 162 evaluate whether the presence of an oncogenic pathogen in a subject is associated with an actionable treatment for the infection. In some embodiments, system 100 queries a database, e.g., a look-up table ("LUT"), of actionable oncogenic pathogen infections, targeted therapies associated with actionable infections, and any other conditions that must be met before administering the targeted therapy to a subject infected with an oncogenic pathogen. In some cases, the LUT may also include contraindications to the associated targeted therapy, such as adverse drug interactions or personal characteristics that contraindicate administration of a particular targeted therapy.
[0163] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more multi-feature analysis algorithms 164 that evaluate multiple features to classify cancers with respect to the efficacy of one or more targeted therapies. For example, in some embodiments, feature analysis module 160 includes one or more classifiers trained on feature data, one or more clinical therapies, and their associated clinical outcomes for multiple training subjects to classify cancers based on predicted clinical outcomes following one or more therapies.
[0164] In some embodiments, the classifier is implemented as an artificial intelligence engine and may include a gradient boosting model, a random forest model, a neural network (NN), a regression model, a naive Bayes model, and / or a machine learning algorithm (MLA). The MLA or NN may be trained from a training dataset that includes one or more features 125, including personal characteristics 126, medical history 127, clinical features 128, genomic features 131, and / or other "mix" features 138. MLA includes supervised algorithms (i.e., algorithms where the features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, nearest neighbor clustering, unsupervised algorithms (i.e., algorithms where the features / classifications in the dataset are not annotated) using Apriori, average clustering, principal component analysis, random forests, adaptive boosting, and semi-supervised algorithms (i.e., algorithms where an incomplete number of features / classifications in the dataset are annotated) using generative approaches (e.g., mixtures of Gaussian distributions, mixtures of multinomial distributions, hidden Markov models), sparse separation, graph-based approaches (e.g., min-cut, harmonic functions, manifold normalization), heuristic approaches, or support vector machines.
[0165] NNs include conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long-short-term memory networks, or other neural models where the training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample.
[0166] Although MLAs and neural networks identify distinct approaches to machine learning, these terms may be used interchangeably herein. Thus, unless explicitly stated otherwise, a reference to an MLA may include a corresponding NN, or vice versa. Training may include providing an optimized dataset, labeling these features as they occur in patient records, and training the MLA to predict or classify based on new inputs. Artificial NNs are efficient computing models that have demonstrated their strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators, i.e., they can represent a wide variety of functions given appropriate parameters.
[0167] In some embodiments, system 100 includes a classifier training module that includes instructions for training one or more untrained or partially trained classifiers based on feature data from a training dataset. In some embodiments, system 100 also includes a database of training data for use in training the one or more classifiers. In other embodiments, the classifier training module accesses a remote storage device that hosts the training data. In some embodiments, the training data includes a set of training features, including, but not limited to, the various types of feature data 125 illustrated in FIG. 1B. In some embodiments, the classifier training module uses patient data 121, for example, when test patient data store 120 also stores records of treatments administered to patients and patient outcomes after treatment.
[0168] In some embodiments, feature analysis module 160 includes one or more clinical data analysis algorithms 165 that evaluate the clinical features 128 of the cancer to identify targeted therapies that may benefit the subject. For example, in some embodiments, where feature data 125 includes pathology data 128-1, for example, one or more clinical data analysis algorithms 165 evaluate the data to determine whether an actionable therapy is indicated, for example, based on the histopathology of a tumor biopsy from the subject, which may indicate a particular cancer type and / or stage of the cancer. In some embodiments, system 100 queries a database, e.g., a look-up table (“LUT”), of actionable clinical features (e.g., pathology features), targeted therapies associated with the actionable features, and any other conditions that must be met before administering the targeted therapy to a subject associated with the actionable clinical feature 128 (e.g., pathology feature 128-1). In some embodiments, system 100 directly evaluates clinical features 128 (e.g., pathology feature 128-1) to determine whether a patient's cancer is susceptible to a particular therapeutic agent. Further details about exemplary methods, systems, and algorithms for classifying cancer and identifying targeted therapies based on clinical data, such as pathology data 128-1, imaging data 138-2, and / or tissue culture / organoid data 128-3, are discussed in, for example, U.S. Patent Nos. 10,957,041, 10,957,445, 11,244,763, 11,848,107, and 11,145,416, the contents of which are incorporated herein by reference in their entirety and for all purposes.
[0169] In some embodiments, feature analysis module 160 includes a clinical trial module that evaluates the test patient data 121 to determine whether the patient is eligible for inclusion in clinical trials for cancer therapies, e.g., clinical trials that are currently recruiting patients, clinical trials that have not yet begun recruiting patients, and / or ongoing clinical trials that may recruit additional patients in the future. In some embodiments, the clinical trial module evaluates the test patient data 121 to determine whether clinical trial results, e.g., results of ongoing clinical trials and / or results of completed clinical trials, are relevant to the patient. For example, in some embodiments, system 100 queries a database of clinical trials, e.g., active and / or completed clinical trials, e.g., a look-up table (“LUT”), and compares the patient data 121 to the clinical trial inclusion criteria stored in the database to identify clinical trials that closely and / or exactly match the patient's data 121. In some embodiments, records of matching clinical trials, e.g., clinical trials for which the patient may be eligible and / or that may inform individualized treatment decisions for the patient, are stored in clinical evaluation database 139.
[0170] In some embodiments, the feature analysis module 160 includes a therapy curation algorithm 166 that assembles the actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials identified for the patient, as described above. In some embodiments, the therapy curation algorithm 166 evaluates certain criteria related to which actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials should be reported and / or whether certain matched therapies, considered alone or in combination, may be contraindicated for the patient, for example, based on the patient's personal characteristics 126 and / or known drug-drug interactions. In some embodiments, the therapy curation algorithm then generates one or more clinical reports 139-3 for the patient. In some embodiments, the therapy curation algorithm generates a first clinical report 139-3-1 that is reported to the medical professional treating the patient, and a second clinical report 139-3-2 that is not communicated to the medical professional but can be used to improve various algorithms within the system.
[0171] In some embodiments, the feature analysis module 160 includes a recommendation validation module 167 that includes an interface that allows a clinician to review, modify, and approve the clinical report 139-3 before the report is sent to a medical professional, e.g., an oncologist, treating the patient.
[0172] In some embodiments, each of one or more of the feature collection, sequencing module, bioinformatics module (e.g., including a variation module, structural variant calling, and data processing module), classification module, and outcome module is communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In some alternative embodiments, each of the feature collection, variation module, structural variant calling, and feature store is communicatively coupled to each other for independent communication without sharing a data bus.
[0173] Further details regarding modules and feature collection systems and exemplary embodiments are discussed in U.S. Patent No. 11,830,587, which is incorporated herein by reference in its entirety.
[0174] Exemplary Methods
[0175] For example, details of a system 100 for providing clinical support aimed at personalized cancer therapy with improved circulating tumor fraction estimates are disclosed, and details regarding the processes and features of the system according to various embodiments of the present disclosure are disclosed below. Specifically, exemplary processes are described below with reference to FIGS. 2A, 3, 4A-4E, 5A-5F, and 6A-6G. In some embodiments, such processes and features of the system are performed by modules 118, 120, 140, 160, and / or 170, as shown in FIG. 1A. With reference to these methods, the systems described herein (e.g., system 100) include instructions for determining an improved and accurate circulating tumor fraction estimate compared to conventional methods for obtaining circulating tumor fraction estimates.
[0176] Figure 2B: Distributed diagnostic and clinical environment.
[0177] In some aspects, the methods described herein for providing clinical support for personalized cancer therapy are implemented across a distributed diagnostic / clinical environment, for example, as illustrated in Figure 2B. However, in some embodiments, the improved methods described herein for providing clinical support for personalized cancer therapy (e.g., by determining accurate circulating tumor fraction estimates) are implemented at a single location, e.g., in a single computing system or environment, although ancillary procedures that support the methods described herein and / or procedures that further utilize the results of the methods described herein may be implemented across the distributed diagnostic / clinical environment.
[0178] FIG. 2B illustrates an example of a distributed diagnostic / clinical environment 210. In some embodiments, the distributed diagnostic / clinical environment is connected via a communications network 105. In some embodiments, one or more biological samples, such as one or more liquid biopsy samples, solid tumor biopsies, normal tissue samples, and / or control samples, are collected from a subject in a clinical environment 220, such as a doctor's office, hospital, or medical clinic, or in a home health care setting (not depicted). Advantageously, solid tumor samples should be collected within a clinical setting, while liquid biopsy samples can be obtained in a minimally invasive manner and are more easily collected outside of a traditional clinical setting. In some embodiments, the one or more biological samples, or portions thereof, are processed within the clinical environment 220 where collection occurred using a processing device 224, such as a nucleic acid sequencer to obtain sequencing data, a microscope to obtain pathology data, a mass spectrometer to obtain proteomic data, etc. In some embodiments, one or more biological samples or portions thereof are sent to one or more external environments, e.g., a sequencing laboratory 230, a pathology laboratory 240, and / or a molecular biology laboratory 250, each of which includes a processing device 234, 244, and 254, respectively, for generating subject biological data 121. Each environment includes a communication device 222, 232, 242, and 252, respectively, for communicating the subject biological data 121 to a processing server 262 and / or database 264, which may be located in yet another environment, e.g., a processing / storage center 260. Thus, in some embodiments, different portions of the systems and methods described herein are performed by different processing devices located in different physical environments.
[0179] Thus, in some embodiments, the method for providing clinical support for personalized cancer therapy, e.g., with improved circulating tumor fraction estimates, is implemented across one or more environments, as illustrated in FIG. 2B . For example, in some such embodiments, a liquid biopsy sample is collected in a clinical setting 220 or a home health care setting. The sample, or a portion thereof, is sent to a sequencing laboratory 230, where raw sequence reads 123 of nucleic acids in the sample are generated by a sequencer 234. The raw sequencing data 123 is communicated, for example, from a communication device 232 to a database 264 in a processing / storage center 260, where a processing server 262 extracts features from the sequence reads by performing one or more processes in a bioinformatics module 140, thereby generating a genomic feature 131 for the sample. The processing server 262 may then analyze the identified features by performing one or more processes in a feature analysis module 160, thereby generating a clinical assessment 139, including a clinical report 139-3. The clinician may access the clinical report 139-3 via the recommendation validation module 167, for example, in the processing / storage center 260 or through the communication network 105. After final approval, the clinical report 139-3 is transmitted to a medical professional, e.g., an oncologist, in the clinical environment 220, who uses the report to support clinical decision-making for the patient's personalized cancer treatment.
[0180] Figure 2A: Exemplary workflow for precision oncology
[0181] 2A is a flowchart of an exemplary workflow 200 for collecting and analyzing data to generate a clinical report 139 for supporting clinical decision-making in precision oncology. Advantageously, the methods described herein improve this process by improving various stages within feature extraction 206, including, for example, determining circulating tumor fraction estimates.
[0182] Briefly, the workflow begins with patient intake and sample collection 201, in which one or more liquid biopsy samples, one or more tumor biopsies, and one or more normal and / or control tissue samples are collected from a patient (e.g., in a clinical setting 220 or a home health care setting, as illustrated in FIG. 2B ). In some embodiments, personal data 126 corresponding to the patient and a record of the one or more biological samples obtained (e.g., patient identifier, patient clinical data, sample type, sample identifier, cancer status, etc.) are entered into a data analysis platform, e.g., laboratory patient data store 120. Thus, in some embodiments, the methods disclosed herein include obtaining one or more biological samples from one or more subjects, e.g., cancer patients. In some embodiments, the subject is a human, e.g., a human cancer patient.
[0183] Sequence reads are then generated from the sequencing library or pool of sequencing libraries (312). Sequencing data can be obtained by any methodology known in the art, such as sequencing-by-synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing-by-ligation (SOLiD sequencing), NANOPORE® sequencing (Oxford Nanopore Technologies), or next-generation sequencing (NGS) technology, such as paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators. In some embodiments, sequencing is performed using next-generation sequencing technology, such as short-read technology. In other embodiments, long-read sequencing or another sequencing method known in the art is used.
[0184] Homologous recombination status (HRD):
[0185] In some embodiments, for example, analysis of aligned sequence reads in SAM or BAM format includes analysis of whether the cancer is homologous recombination deficient (HRD status 137-3) using the homologous recombination pathway analysis module 157.
[0186] Homologous recombination (HR) is a normal, highly conserved DNA repair process that allows the exchange of genetic information between identical or closely related DNA molecules. It is most widely used by cells to accurately repair harmful breaks (e.g., lesions) that occur on both strands of DNA. DNA damage can arise from exogenous (external) sources, such as ultraviolet light, radiation, or chemical damage, or from endogenous (internal) sources, such as errors in DNA replication or other cellular processes that generate DNA damage. Double-strand breaks are a type of DNA damage. The use of poly(ADP-ribose) polymerase (PARP) inhibitors in patients with HRD impairs both pathways of DNA repair, leading to cell death (apoptosis). The efficacy of PARP inhibitors is improved not only in ovarian cancers that exhibit germline or somatic BRCA mutations, but also in cancers where HRD is caused by other underlying etiologies.
[0187] In some embodiments, HRD status can be determined by inputting features correlated with HRD status into a classifier trained to distinguish between cancers with homologous recombination pathway defects and cancers without homologous recombination pathway defects. For example, in some embodiments, the features include one or more of: (i) the heterozygous state of a first plurality of DNA damage repair genes in the genome of the subject's cancer tissue; (ii) a measure of loss of heterozygosity across the genome of the subject's cancer tissue; (iii) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's cancer tissue; and (iv) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's non-cancerous tissue. In some embodiments, all four of the features described above are used as features in the HRD classifier. For more detailed information about HRD classifiers using these and other features, see U.S. Patent Application Publication No. 2020 / 0255909, the contents of which are incorporated herein by reference in their entirety and for all purposes.
[0188] Parallel Testing
[0189] Unless otherwise specified, as used herein, the term "parallel" when referring to an assay refers to a period of 0 to 90 days. In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 90 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 60 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 30 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 21 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 14 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 7 days (e.g., the biological samples are collected within that period).In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for hematological cancers, and a non-cancerous sample) are conducted within a 0-3 day period (e.g., the biological samples are collected within that period).
[0190] In some embodiments, liquid biopsy assays are used in parallel with solid tumor assays to provide more comprehensive information about patient variants. For example, blood samples and solid tumor samples may be sent to a laboratory for evaluation. Solid tumor samples may be analyzed using a bioinformatics pipeline to generate solid tumor results. For example, solid tumor assays are described in U.S. Patent No. 11 / 705,226, the contents of which are incorporated herein by reference in their entirety for all purposes. Solid tumor cancer types may include, for example, non-small cell lung cancer, colorectal cancer, or breast cancer. Alterations identified in tumor / matched normal results may include, for example, EGFR+ for non-small cell lung cancer, HER2+ for breast cancer, or KRAS G12C for some cancers.
[0191] In some embodiments, a blood sample may be divided into a first portion and a second portion. The blood sample and the first portion of the solid tumor sample may be analyzed using a bioinformatics pipeline to generate a tumor / matched normal result. The second portion of the blood sample may be analyzed using a bioinformatics pipeline to generate a liquid biopsy result. For example, a blood sample may be analyzed using improvements in at least somatic variant identification, e.g., as described herein in the section entitled "Variant Identification." For example, a blood sample may be analyzed using improvements in focused copy number identification, e.g., as described herein in the section entitled "Copy Number Variation." For example, a blood sample may be analyzed using improvements in circulating tumor fraction determination, e.g., using the methods disclosed herein.
[0192] In some embodiments, a therapy is identified for further consideration following the results of the tumor or tumor / matched normal along with the results of the liquid biopsy. In one example, when the results overall suggest that the patient has HER2+ breast cancer, the cancer drug neratinib is identified along with the test results for further consideration by the prescribing physician.
[0193] In some embodiments, solid tumor or tumor / matched normal assays are ordered in parallel, their results delivered in parallel, and analyzed in parallel.
[0194] Systems and methods for improving circulating tumor fraction estimates
[0195] An overview of a method for providing clinical support for personalized cancer therapy is described above with reference to Figures 2-4. Below, for example, a system and method for improving circulating tumor fraction estimates within the context of the above-described methods and systems is described with reference to Figures 5A-5D.
[0196] Many of the embodiments described below in conjunction with Figures 5A-5D relate to analysis carried out using genomic data from solid tumor samples and the sequencing data of cfDNA obtained from liquid biopsy samples of subjects such as cancer patients. Generally, these embodiments are independent and therefore do not depend on any specific DNA sequencing method. However, in some embodiments, the methods described below include generating sequencing data.
[0197] As described herein, in some embodiments, the methods described herein (e.g., method 500 shown in Figures 5A-5D) include one or more data collection steps in addition to data analysis and downstream steps. As described herein, for example, with reference to Figures 2 and 3, in some embodiments, the methods include collection of a solid tumor biopsy from a subject, a liquid biopsy sample, and optionally one or more matched biological samples (e.g., matched cancerous samples from the subject). Similarly, as described herein, for example, with reference to Figures 2 and 3, in some embodiments, the methods include extraction of genomic DNA (gDNA) from the solid tumor biopsy and, optionally, a matched non-cancerous sample. Similarly, as described herein, for example, with reference to Figures 2 and 3, in some embodiments, the methods include isolation of cfDNA fragments from the liquid biopsy sample (cfDNA). Similarly, as described herein, e.g., with reference to Figures 2 and 3, in some embodiments, the method includes nucleic acid sequencing of gDNA extracted from a solid tumor biopsy, a liquid biopsy sample, and / or cfDNA, optionally from one or more matched biological samples from the subject (e.g., matched cancerous and / or matched non-cancerous samples derived from the subject).
[0198] However, in other embodiments, the methods described herein begin with obtaining nucleic acid sequencing results, such as raw or collapsed sequence reads of gDNA from a solid tumor biopsy, a liquid biopsy sample (cfDNA), and optionally cfDNA from one or more matched biological samples from the subject (e.g., matched cancerous and / or matched non-cancerous samples from the subject), from which genomic features (e.g., variant allele identification and / or variant allele proportion) necessary to estimate the circulating tumor fraction can be determined. For example, in some embodiments, sequencing data 122 of patient 121 is accessed and / or downloaded by system 100 via network 105.
[0199] In some embodiments, the method further comprises isolating a plurality of cell-free nucleic acids from the liquid biopsy sample of the subject prior to sequencing. In some embodiments, the sequencing is multiplexed sequencing. In some embodiments, the sequencing is short-read sequencing or long-read sequencing.
[0200] In some embodiments, pathology data 128-1 collected during a clinical evaluation includes, for example, visual features identified by a pathologist's examination of a specimen (e.g., a solid tumor biopsy) on a stained H&E or IHC slide. In some embodiments, the sample is a solid tissue biopsy sample. In some embodiments, the tissue biopsy sample is formalin-fixed tissue (FFT), e.g., formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the tissue biopsy sample is an FFPE or FFT block. In some embodiments, the tissue biopsy sample is a fresh-frozen tissue biopsy. The tissue biopsy sample can be prepared in thin sections (e.g., by cutting and / or mounting on slides) to facilitate pathology review (e.g., by staining with immunohistochemical stains for IHC review and / or with hematoxylin and eosin stains for H&E pathology review). For example, analysis of slides for H&E or IHC staining may reveal characteristics such as tumor infiltration, programmed death-ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.
[0201] Similarly, in some embodiments, the methods described herein begin with obtaining genomic features necessary for filtering clonal hematopoietic variants from sequencing a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). For example, in some embodiments, (i) one or more fragment length indicators, (ii) one or more features determined from the variant allele fraction of a candidate somatic variant and the ctFE of the liquid biopsy sample, or the variant allele fraction of a candidate somatic variant and the ctFE of the liquid biopsy sample, and (iii) one or more indicators of clonal hematopoietic prevalence at a first nucleotide position are accessed and / or downloaded by system 100 via network 105.
[0202] Similarly, in some embodiments, the methods described herein begin with obtaining genomic features necessary to estimate circulating tumor fraction (e.g., variant allele identification and / or variant allele fraction) for sequencing data for gDNA from a solid tumor biopsy, a liquid biopsy sample (cfDNA), and, optionally, cfDNA from one or more matched biological samples from the subject (e.g., matched cancerous and / or matched non-cancerous samples from the subject). For example, in some embodiments, variant allele count and / or variant allele fraction for sequencing data 122 of patient 121 is accessed and / or downloaded by system 100 via network 105.
[0203] 5A-5D collectively provide a flowchart of processes and features for determining an estimate of a subject's circulating tumor fraction, according to some embodiments of the present disclosure.
[0204] The present disclosure provides a method 500 for estimating circulating tumor fraction for a subject from panel enriched sequencing data for a plurality of sequences.
[0205] Block 502. Referring to block 502, in some embodiments, the method includes obtaining a first plurality of nucleic acid sequences from a solid tumor sample from the subject, the first plurality of nucleic acid sequences comprising a nucleic acid sequence corresponding to each respective locus in a plurality of loci in genomic DNA.
[0206] In some embodiments, the solid tumor sample is a specific type of cancer. Non-limiting examples of cancer types include lung cancer, breast cancer, ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, cholangiocarcinoma, adrenal cancer, neural carcinoma, neuroblastoma, basal cell carcinoma, brain cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, gastric cancer, bone cancer, thymoma, These include prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, serous ovarian cancer, HR+ breast cancer, serous uterine cancer, endometrial cancer, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.
[0207] In some embodiments, the cancer is a specific stage of a specific type of cancer. In some such embodiments, the stage of a specific type of cancer is the stage of the cancer at which the subject was diagnosed before treatment. Cancers are typically staged to determine the extent of its spread and guide treatment decisions. The stage of a cancer refers to the extent to which the cancer has grown and spread from its original location. Each cancer has its own criteria for determining the stage, but generally relies on determining the size of the primary tumor (T) and whether it has invaded nearby tissues, evaluating lymph node metastasis (N) to find signs of whether the cancer has spread to nearby lymph nodes, and evaluating distant metastasis (M) to indicate whether the cancer has spread to distant organs or tissues. Metastasis refers to the spread of cancer from the primary site to other parts of the body. In some embodiments, the staging system used is the cancer TNM system, which combines T, N, and M information to assign a stage. In some embodiments, stages are indicated using Roman numerals (I, II, III, IV) and may have subcategories (e.g., stage IIA, stage IIB) to provide more precise information. In a brief overview of these stages, in stage 0, the cancer is in situ, meaning it is confined to the layer of cells where it began and has not invaded nearby tissues; in stage I, the cancer is localized and small in size; in stage II, the cancer may be larger and / or have spread to nearby lymph nodes but is still relatively localized; in stage III, the cancer has typically spread further into nearby tissues and may have metastasized to more lymph nodes; and in stage IV, the cancer has spread to distant organs or tissues and exhibits metastasis. This is often the most advanced stage. The specific criteria for each stage can vary depending on the type of cancer.For details on TNM staging of breast cancer, see, for example, Part et al., 2011, "Clinical relevance of TNM staging system according to breast cancer subtypes," Annals of Oncology 22(7), pp. 1554-1560, which is incorporated herein by reference. In addition, some cancers have their own staging systems tailored to their characteristics. See the internet URL cancer.gov / about-cancer / diagnosis-staging / staging.
[0208] In some embodiments, the first plurality of nucleic acid sequences is obtained from a first sequencing reaction performed at a read depth of at least 1x. In some embodiments, the first sequencing reaction is a panel enrichment sequencing reaction performed at a read depth of at least 2x, at least 3x, at least 4x, at least 5x, at least 10x, at least 25x, at least 50x, at least 100x, at least 250x, 400x, 500x, 600x, or more. In some embodiments, the first sequencing reaction is a panel-based or whole genome sequencing reaction performed at a read depth of 1000x or less, 500x or less, 100x or less, 50x or less, or less. In some embodiments, the first sequencing reaction is performed at a read depth of 1x to 500x, 1x to 100x, or 1x to 50x. In some embodiments, the first sequencing reaction is performed at a read depth of 2.5x to 500x, 2.5x to 100x, or 2.5x to 50x. In some embodiments, the first sequencing reaction is performed at a read depth of 5x to 500x, 5x to 100x, or 5x to 50x. In some embodiments, the first sequencing reaction is performed at a read depth of 10x to 500x, 10x to 100x, or 10x to 50x.
[0209] In some embodiments, the first plurality of sequence reads are sequence reads from a panel enrichment sequencing reaction, comprising a first subset of sequence reads corresponding to the cfDNA fragments targeted by one or more probes in the target enrichment panel.In some embodiments, each cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to each probe sequence in the plurality of probe sequences, and the probe sequences are used to enrich the cell-free DNA fragments in the liquid biopsy sample in the panel enrichment sequencing reaction.In some embodiments, the plurality of probe sequences are mapped to 150 or less genes in the human genome.
[0210] In some embodiments, the first plurality of nucleic acid sequences is obtained from a plurality of sequence reads obtained by whole genome or whole exome sequencing. In some such embodiments, whole exome capture is performed using an automated system using a liquid handling robot. Because many loci are being sequenced, whole genome sequencing, and to some extent whole exome sequencing, is typically performed at a lower sequencing depth than smaller targeted panel sequencing reactions. For example, in some embodiments, whole genome or whole exome sequencing is performed to an average sequencing depth of at least 3x, at least 5x, at least 10x, at least 15x, at least 20x, or more. In some embodiments, low-pass whole genome sequencing (LPWGS) technology is used for whole genome or whole exome sequencing. LPWGS is typically performed to an average sequencing depth of about 0.25x to about 5x, more typically to an average sequencing depth of about 0.5x to about 3x.
[0211] Due to differences in sequencing methodology, data obtained from targeted panel sequencing is better suited for certain analyses than data obtained from whole genome / whole exome sequencing, and vice versa. For example, due to the higher sequencing depth achieved by targeted panel sequencing, the resulting sequence data is better suited for identifying variant alleles present in a sample at low allele rates, for example, less than 20%. In contrast, data generated from whole genome / whole exome sequencing is better suited for estimating whole genome metrics, such as tumor mutation burden, because the whole genome is better represented by the sequencing data. Thus, in some embodiments, nucleic acid samples, such as cfDNA, gDNA, or mRNA samples, are evaluated using both targeted panel sequencing and whole genome / whole exome sequencing (e.g., LPG Seq).
[0212] In some embodiments, raw sequence reads resulting from a sequencing reaction are output from the sequencer in a native file format, e.g., a BCL file. In some embodiments, the native file is passed directly to a bioinformatics pipeline (e.g., variant analysis 206), the components of which are described in detail below. In other embodiments, preprocessing is performed before passing the sequences to the bioinformatics platform. For example, in some embodiments, the format of the sequence read file is converted from the native file format (e.g., BCL) to a file format compatible with one or more algorithms used in the bioinformatics pipeline (e.g., FASTQ or FASTA). In some embodiments, the raw sequence reads are filtered to remove sequences that do not meet one or more quality thresholds. In some embodiments, raw sequence reads generated from the same unique nucleic acid molecule in a sequencing read are collapsed into a single sequence read representing that molecule, e.g., using the UMI described above. In some embodiments, one or more of these preprocessing activities are performed within the bioinformatics pipeline itself.
[0213] In one example, a sequencer can generate a BCL file. The BCL file can contain raw image data for multiple patient samples to be sequenced. The BCL image data is an image of the flow cell for each cycle during sequencing. Cycles can be implemented by illuminating the patient sample with specific wavelengths of electromagnetic radiation to generate multiple images that can be processed into base calls via a BCL-to-FASTQ processing algorithm that identifies which base pairs are present in each cycle. The resulting FASTQ file contains the entirety of the reads for each patient sample, paired with a quality metric ranging from 0 to 64, for example, where 64 is the best quality and 0 is the worst quality.
[0214] The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants or copy number changes are present in the sample. Each FASTQ file contains reads, which may be paired-end or single reads and may be short or long reads. Each read represents the sequence of one detected nucleotide in a nucleic acid molecule or copy of a nucleic acid molecule isolated from a patient sample, as detected by a sequencer. Each read in a FASTQ file is also associated with a quality assessment. The quality assessment may reflect the likelihood of an error occurring during the sequencing procedure that affected the associated read. In some embodiments, the paired-end sequencing results for each isolated nucleic acid sample are contained in a pair of FASTQ files, separated for efficiency. Thus, in some embodiments, the forward (read 1) and reverse (read 2) sequences for each isolated nucleic acid sample are stored separately, but in the same order and under the same identifier.
[0215] In various embodiments, the bioinformatics pipeline can filter the FASTQ data from the corresponding sequence data file for each respective biological sample. Such filtering can include correcting or masking sequencer errors, and removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, over-represented sequences, biases caused by library preparation, amplification, or capture, and other errors.
[0216] Although workflow 200 illustrates obtaining a biological sample, extracting nucleic acids from the biological sample, and sequencing the isolated nucleic acids, in some embodiments, the sequencing data used in the improved systems and methods described herein (e.g., including improved methods for determining accurate circulating tumor fraction estimates) is obtained by receiving previously generated sequence reads in electronic format.
[0217] In some embodiments, sequencing of the plurality of nucleic acids in the subject's biopsy sample is performed at a central laboratory or sequencing facility. In some such embodiments, the method includes accessing one or more sequencing datasets and / or one or more auxiliary files in electronic format via a cloud-based interface. For example, the datasets can be obtained by running a bioinformatics pipeline using a tumor BAM file, a normal BAM file, a human reference genome file, a target region BED file, a list of mappable regions of the genome, and / or a blacklist of recurrent problem regions of the genome.
[0218] In some embodiments, obtaining the dataset includes accessing the dataset in electronic format via a cloud-based interface. For example, the dataset can include one or more outputs from a bioinformatics pipeline (e.g., CNVkit output ".cns" and / or ".cnr").
[0219] Additional methods and embodiments for sequencing nucleic acid, including aligning and pre-processing sequence reads, are described in more detail herein. Additional methods and embodiments for implementing the method of the present disclosure in distributed diagnostic and clinical environments are described in detail above (see exemplary method: Figure 2B: distributed diagnostic and clinical environments).As will be clear to those skilled in the art, other embodiments and / or any combination, substitution, addition or deletion thereof are possible.
[0220] 2B, nucleic acid sequencing of one or more samples collected from a subject is performed during wet-lab processing 204, for example, in some embodiments in a sequencing laboratory 230. An exemplary workflow for nucleic acid sequencing is illustrated in FIG. 3. In some embodiments, one or more biological samples obtained in the sequencing laboratory 230 are registered (302) to track the samples and data throughout the sequencing process.
[0221] Next, nucleic acids, e.g., RNA and / or DNA, are extracted from one or more biological samples (304). Methods for isolating nucleic acids from biological samples are known in the art and depend on the type of nucleic acid being isolated (e.g., cfDNA, DNA, and / or RNA) and the type of sample from which the nucleic acid is isolated (e.g., liquid biopsy samples, leukocyte buffy coat preparations, formalin-fixed paraffin-embedded (FFPE) solid tissue samples, and fresh-frozen solid tissue samples). The selection of any particular nucleic acid isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the sample type, sample condition, type of nucleic acid to be sequenced, and sequencing technology used.
[0222] Many techniques for DNA isolation, e.g., genomic DNA isolation, from tissue samples are known in the art, such as organic extraction, silica adsorption, and anion exchange chromatography. Similarly, many techniques for RNA isolation, e.g., mRNA isolation, from tissue samples are known in the art. For example, acid guanidinium thiocyanate-phenol-chloroform extraction (see, e.g., Chomczynski and Sacchi, 2006, Nat Protoc, 1(2):581-85, incorporated herein by reference) and silica bead / glass fiber adsorption (see, e.g., Poeckh et al., 2008, Anal Biochem., 373(2):253-62, incorporated herein by reference). The selection of any particular DNA or RNA isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the tissue type, tissue condition (e.g., fresh, frozen, formalin-fixed, paraffin-embedded (FFPE)), and the type of nucleic acid analysis to be performed.
[0223] Referring to block 508 below, in some embodiments where the biological sample is a liquid biopsy sample, e.g., a blood or plasma sample, cfDNA is isolated from the blood sample using a commercially available reagent containing proteinase K to produce a liquid solution of cfDNA.
[0224] In some embodiments, the isolated DNA molecules are mechanically sheared to an average length using an ultrasonicator (e.g., a Covaris ultrasonicator). In some embodiments, the isolated nucleic acid molecules are analyzed to determine their fragment size, for example, through gel electrophoresis and / or the use of a device such as a LabChip GX Touch. Those skilled in the art will understand the appropriate range of fragment sizes based on the sequencing technology being employed, as different sequencing technologies have different fragment size requirements for robust sequencing. In some embodiments, quality control tests are performed on the extracted nucleic acids (e.g., DNA and / or RNA) to, for example, assess nucleic acid concentration and / or fragment size. Sizing DNA fragments, such as determining whether DNA fragments require additional shearing before sequencing, provides valuable information for use in downstream processing.
[0225] Wet-lab processing 204 then includes preparing a nucleic acid library from the isolated nucleic acids (e.g., cfDNA, DNA, and / or RNA). For example, in some embodiments, a DNA library (e.g., a gDNA and / or cfDNA library) is prepared from the isolated DNA from one or more biological samples. In some embodiments, the DNA library is prepared using a commercially available library preparation kit, such as a KAPA Hyper Prep Kit, a New England Biolabs (NEB) kit, or a similar kit.
[0226] In some embodiments, adaptors (e.g., UDI adaptors such as Roche SeqCap double-ended adaptors, or UMI adaptors such as full-length or short Y adaptors) are ligated onto nucleic acid molecules during library preparation. In some embodiments, the adaptors comprise unique molecular identifiers (UMIs), which are short nucleic acid sequences (e.g., 3-10 base pairs) that are added to the ends of DNA fragments during adaptor ligation. In some embodiments, the UMIs are degenerate base pairs that serve as unique tags that can be used to identify sequence reads derived from specific DNA fragments. In some embodiments, patient-specific indicators are also added to nucleic acid molecules, for example, when multiplex sequencing is used to sequence DNA from multiple samples (e.g., from the same or different subjects) in a single sequencing reaction. In some embodiments, patient-specific indicators are short nucleic acid sequences (e.g., 3-20 nucleotides) that are added to the ends of DNA fragments during library construction that serve as unique tags that can be used to identify sequence reads derived from specific patient samples. Examples of identifier sequences are described in Kivioja et al., 2011, Nat. Methods 9(1):72-74, and Islam et al., 2014, Nat. Methods 11(2):163-66, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0227] In some embodiments, the adapters include a PCR primer landing site designed for efficient binding of a PCR or second-strand synthesis primer used during the sequencing reaction. In some embodiments, the adapters include an anchor binding site to facilitate binding of the DNA molecule to an anchor oligonucleotide molecule on a sequencer flow cell, serving as a seed for the sequencing process by providing a starting point for the sequencing reaction. During PCR amplification after adapter ligation, the UMI, patient index, and binding site are replicated along with the attached DNA fragment. This provides a method for identifying sequence reads derived from the same original fragment in downstream analysis.
[0228] In some embodiments, the DNA library is amplified and purified using commercially available reagents (e.g., Axygen MAG PCR cleanup beads). In some such embodiments, the concentration and / or amount of DNA molecules is then quantified using a fluorescent dye and fluorescence microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In some embodiments, library amplification is performed on a device (e.g., Illumina C-Bot2), and the resulting flow cells containing the amplified target capture DNA library are sequenced to a specific on-target depth selected by the user on a next-generation sequencer (e.g., Illumina HiSeq 4000 or Illumina NovaSeq 6000). In some embodiments, DNA library preparation is performed using an automated system that uses a liquid handling robot (e.g., SciClone NGSx).
[0229] In some embodiments, where the feature data 125 includes the methylation state 132 of one or more genomic locations, nucleic acids (e.g., cfDNA) isolated from a biological sample are processed to convert unmethylated cytosines to uracil, for example, before generating a sequencing library. Thus, when the nucleic acids are sequenced, all cytosines called in the sequencing reaction are necessarily methylated, since unmethylated cytosines are converted to uracil and would therefore be called thymidine rather than cytosine in the sequencing reaction. Commercially available kits are available for the bisulfite-mediated conversion of methylated cytosines to uracil. Commercially available kits are also available for the enzymatic conversion of methylated cytosines to uracil.
[0230] In some embodiments, wet-lab processing 204 includes pooling (308) DNA molecules from multiple libraries corresponding to different samples from the same and / or different patients to form a sequencing pool of DNA libraries. When the pool of DNA libraries is sequenced, the resulting sequence reads correspond to nucleic acids isolated from the multiple samples. The sequence reads can be separated into different sequence read data structures (e.g., files) corresponding to the various samples represented by the sequencing reads based on unique identifiers present in the appended nucleic acid fragments. In this manner, a single sequencing reaction can generate sequence reads from multiple samples. Advantageously, this allows for the processing of more samples per sequencing reaction.
[0231] In some embodiments, wet-lab processing 204 includes enriching 310 a sequencing library or pool of sequencing libraries for target nucleic acids, e.g., nucleic acids encompassing loci that are useful for precision oncology and / or used as internal controls for sequencing or bioinformatics processes. In some embodiments, enrichment is achieved by hybridizing target nucleic acids in the sequencing library to probes that hybridize to the target sequences and then isolating the captured nucleic acids from off-target nucleic acids that are not bound by the capture probes. In some embodiments, some off-target nucleic acids will remain in the final sequencing pool.
[0232] In some embodiments, the first plurality of sequence reads obtained from the sequencing described above comprises at least 10,000 sequence reads, at least 50,000 sequence reads, at least 100,000 sequence reads, at least 500,000 sequence reads, at least 1 million sequence reads, at least 5 million sequence reads, at least 10 million sequence reads, or more. In some embodiments, the first plurality of sequence reads comprises 1 billion or less sequence reads, 500 million or less sequence reads, 100 million or less sequence reads, 50 million or less sequence reads, 10 million or less sequence reads, 5 million or less sequence reads, 1 million or less sequence reads, or less. In some embodiments, the first plurality of sequence reads is between 10,000 and 1 billion sequence reads, between 10,000 and 500 million sequence reads, between 10,000 and 100 million sequence reads, between 10,000 and 50 million sequence reads, between 10,000 and 10 million sequence reads, between 10,000 and 5 million sequence reads, or between 10,000 and 1 million sequence reads. In some embodiments, the first plurality of sequence reads is between 100,000 and 1 billion sequence reads, between 100,000 and 500 million sequence reads, between 100,000 and 100 million sequence reads, between 100,000 and 50 million sequence reads, between 100,000 and 10 million sequence reads, between 100,000 and 5 million sequence reads, or between 100,000 and 1 million sequence reads. In some embodiments, the first plurality of sequence reads is between 500,000 and 1 billion sequence reads, between 500,000 and 500 million sequence reads, between 500,000 and 100 million sequence reads, between 500,000 and 50 million sequence reads, between 500,000 and 10 million sequence reads, between 500,000 and 5 million sequence reads, or between 500,000 and 1 million sequence reads.In some embodiments, the first plurality of sequence reads is between 1 million and 1 billion sequence reads, between 1 million and 500 million sequence reads, between 1 million and 100 million sequence reads, between 1 million and 50 million sequence reads, between 1 million and 10 million sequence reads, or between 1 million and 5 million sequence reads.
[0233] In some embodiments, the genomic DNA from the subject's solid tumor sample comprises a first plurality of nucleic acid fragments. In some embodiments, the first plurality of nucleic acid fragments comprises at least 1,000 DNA fragments, at least 5,000 DNA fragments, at least 10,000 DNA fragments, at least 50,000 DNA fragments, at least 100,000 DNA fragments, at least 500,000 DNA fragments, at least 1 million DNA fragments, at least 5 million DNA fragments, or more. In some embodiments, the first plurality of nucleic acid fragments comprises no more than 100 million DNA fragments, no more than 50 million DNA fragments, no more than 10 million DNA fragments, no more than 5 million DNA fragments, no more than 1 million DNA fragments, no more than 500,000 DNA fragments, or no more than 100,000 DNA fragments. In some embodiments, the first plurality of DNA fragments is between 1,000 DNA fragments and 500 million DNA fragments, between 1,000 DNA fragments and 100 million DNA fragments, between 1,000 DNA fragments and 50 million DNA fragments, between 1,000 DNA fragments and 10 million DNA fragments, between 1,000 DNA fragments and 5 million DNA fragments, between 1,000 DNA fragments and one million DNA fragments, between 1,000 DNA fragments and 500,000 DNA fragments, between 1,000 DNA fragments and 250,000 DNA fragments, or between 1,000 DNA fragments and 100,000 DNA fragments. In some embodiments, the first plurality of DNA fragments is between 5,000 DNA fragments and 500 million DNA fragments, between 5,000 DNA fragments and 100 million DNA fragments, between 5,000 DNA fragments and 50 million DNA fragments, between 5,000 DNA fragments and 10 million DNA fragments, between 5,000 DNA fragments and 5 million DNA fragments, between 5,000 DNA fragments and one million DNA fragments, between 5,000 DNA fragments and 500,000 DNA fragments, between 5,000 DNA fragments and 250,000 DNA fragments, or between 5,000 DNA fragments and 100,000 DNA fragments.In some embodiments, the first plurality of DNA fragments is between 10,000 DNA fragments and 500 million DNA fragments, between 10,000 DNA fragments and 100 million DNA fragments, between 10,000 DNA fragments and 50 million DNA fragments, between 10,000 DNA fragments and 10 million DNA fragments, between 10,000 DNA fragments and 5 million DNA fragments, between 10,000 DNA fragments and one million DNA fragments, between 10,000 DNA fragments and 500,000 DNA fragments, between 10,000 DNA fragments and 250,000 DNA fragments, or between 10,000 DNA fragments and 100,000 DNA fragments. In some embodiments, the first plurality of DNA fragments is between 25,000 DNA fragments and 500 million DNA fragments, between 25,000 DNA fragments and 100 million DNA fragments, between 25,000 DNA fragments and 50 million DNA fragments, between 25,000 DNA fragments and 10 million DNA fragments, between 25,000 DNA fragments and 5 million DNA fragments, between 25,000 DNA fragments and one million DNA fragments, between 25,000 DNA fragments and 500,000 DNA fragments, between 25,000 DNA fragments and 250,000 DNA fragments, or between 25,000 DNA fragments and 100,000 DNA fragments.
[0234] In some embodiments, the sequencing data is processed (e.g., using sequence data processing module 141) to prepare data for genomic feature identification 385. For example, in some embodiments described above, the sequencing data is in the native file format provided by the sequencer. Thus, in some embodiments, the system (e.g., system 100) applies preprocessing algorithms 142 to convert the file format into one recognized by one or more upstream processing algorithms (318). For example, BCL file output from a sequencer can be converted to FASTQ file format using bcl2fastq or bcl2fastq2 conversion software (Illumina®). The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants, copy number changes, etc., are present in the sample.
[0235] In some embodiments, other preprocessing functions are performed, such as filtering sequence reads 122 based on desired quality, e.g., size and / or quality of base calls. In some embodiments, quality control checks are performed to ensure that data is sufficient for variant calling. For example, entire reads, individual nucleotides, or multiple nucleotides likely to have errors can be discarded based on a quality assessment associated with the reads in the FASTQ file, the known error rate of the sequencer, and / or a comparison between each nucleotide in the read and one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools, such as Skewer. See Jiang et al., 2014, BMC Bioinformatics 15(182):1-12. FASTQ files can be analyzed for quality control and rapid assessment of reads by sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC, or another similar software program. For paired-end reads, the reads can be merged.
[0236] For efficiency, in some embodiments, the paired-end sequencing results for each isolate are contained in a separate pair of FASTQ files. The forward (read 1) and reverse (read 2) sequences for each tumor and normal isolate are stored separately but in the same order and under the same identifier. See, for example, Figure 4C. In various embodiments, the bioinformatics pipeline may filter the FASTQ data from each isolate. Such filtering may include correcting or masking sequencer errors, as well as removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, overrepresented sequences, biases caused by library preparation, amplification, or capture, and other errors. See, for example, Figure 4D.
[0237] Similarly, in some embodiments, sequencing (312) is performed on a pool of nucleic acid sequencing libraries prepared from different biological samples, e.g., from the same or different patients. Accordingly, in some embodiments, the system demultiplexes (320) the data (e.g., using a demultiplexing algorithm 144) to separate sequence reads into separate files for each sequencing library included in the sequencing pool, e.g., based on UMIs or patient identifier sequences added to the nucleic acid fragments during sequencing library preparation, as described above. In some embodiments, the demultiplexing algorithm is part of the same software package as one or more preprocessing algorithms 142. For example, commercially available conversion software includes instructions for both converting the native file format output from the sequencer and demultiplexing the sequence reads 122 output from the reaction.
[0238] The sequence reads are then aligned (322) to a reference sequence construct 158, such as a reference genome, reference exome, or other reference construct prepared for a particular sequencing reaction, using, for example, an alignment algorithm 143. For example, in some embodiments, individual sequence reads 123 in electronic form (e.g., in a FASTQ file) are aligned to a reference sequence construct for the species of interest (e.g., a reference human genome) by identifying the sequence in the region of the reference sequence construct that best matches the sequence of nucleotides in the sequence read. In some embodiments, the sequence reads are aligned to the reference exome or reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information may also include the sequence read length, which may be determined from the start and end positions. The region in the reference genome may relate to a gene or a segment of a gene. Any of a variety of alignment tools can be used for this task.
[0239] For example, local sequence alignment algorithms compare different lengths of subsequences in query sequence (e.g., sequence reads) with subsequences in target sequence (e.g., reference constructs) to create the best alignment for each part of query sequence.In contrast, global sequence alignment algorithms align the entire sequence, for example, end to end.Examples of local sequence alignment algorithms include Smith-Waterman algorithm.
[0240] In some embodiments, the read mapping process begins by constructing an index of either the reference genome or the read, which is then used to retrieve a set of positions in the reference sequence to which the read is more likely to align. Once this subset of possible mapping positions is identified, alignment is performed within these candidate regions using slower, more sensitive algorithms. See, for example, Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14:184, and Flicek and Birney, 2009, "Sense from sequence reads: methods for alignment and assembly," Nat Methods 6 (Suppl. 11), S6-S12, each of which is incorporated herein by reference. In some embodiments, the mapping tool methodology utilizes a hash table or the Burrows-Wheeler transform (BWT). See, for example, Li and Homer, 2010, “A survey of sequence alignment algorithms for next-generation sequencing,” Brief Bioinformatics 11, pp. 473-483, which is incorporated herein by reference.
[0241] Other software programs designed to align reads include, for example, Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and / or programs using the Smith-Waterman algorithm. Candidate reference genomes include, for example, HG19, GRCh38, hg38, GRCh37, and / or other reference genomes developed by the Genome Reference Consortium. In some embodiments, the alignment generates a SAM file, which stores the start and end positions of each read, along with its coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome.
[0242] For example, in some embodiments, each read in the FASTQ file is aligned to the location in the human genome that best matches the sequence of nucleotides in the read. Many software programs designed to align reads exist, including Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and programs using the Smith-Waterman algorithm. Alignment can be directed using a reference genome (e.g., HG19, GRCh38, HG38, GRCh37, or other reference genomes developed by the Genome Reference Consortium) by comparing the nucleotide sequence in each read to portions of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that most likely corresponds to the sequence in the read. In some embodiments, one or more SAM files are generated for the alignment, which store the start and end positions of each read according to their coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. The SAM file can be converted to a BAM file. In some embodiments, the BAM file is sorted, and duplicate reads are marked for deletion, resulting in a de-duplicated BAM file.
[0243] In some embodiments, adapter-trimmed FASTQ files are aligned to the 19th edition of the Human Reference Genome Build (HG19). After alignment, reads are grouped by alignment position and UMI family and collapsed into a consensus sequence. Bases with insufficient quality or significant discrepancies between family members (e.g., when it is uncertain whether a base is adenine, cytosine, guanine, etc.) can be replaced with N to represent a wildcard nucleotide type. PHRED scores are then scaled based on the initial base call estimates combined across all family members. After single-strand consensus generation, a double-stranded consensus sequence is generated by comparing the forward and reverse PCR products with the mirrored UMI sequence. In various embodiments, a consensus can be generated across read pairs; otherwise, a single-stranded consensus call is used. After consensus calling, filtering is performed to remove low-quality consensus fragments. The consensus fragments are then realigned to the human reference genome using BWA. A BAM output file is generated after the realignment, which is then sorted and indexed by alignment position.
[0244] In some embodiments, the sequencing data is normalized to account for, for example, pulldown, amplification, and / or sequencing bias (e.g., mappability, GC bias, etc.).
[0245] In some embodiments, the SAM file generated after alignment is converted to a BAM file 124. Thus, after preprocessing the sequencing data generated for the pooled sequencing reactions, a BAM file is generated for each sequencing library present in the master sequencing pool. In some embodiments, the BAM file is sorted and duplicate reads are marked for removal, resulting in a de-duplicated BAM file. For example, a tool such as SamBAMBA marks and filters duplicate alignments in the sorted BAM file.
[0246] The alignment file (e.g., BAM file 124) prepared as described above is then passed to a feature extraction module 145, where the sequences are analyzed to identify genomic alterations (e.g., SNV / MNV, indels, genomic rearrangements, copy number variations, etc.) and / or determine various characteristics of the patient's cancer (e.g., MSI status, TMB, tumor ploidy, HRD status, tumor fraction, tumor purity, methylation patterns, etc.) (324). Many software packages for identifying genomic alterations are known in the art. Generally, these software packages identify variants in the sorted SAM or BAM file 124 relative to one or more reference sequence constructs 158. The software package then outputs a file, e.g., a raw VCF (variant call format), that lists the called variants (e.g., genomic features 131) and identifies positions relative to the reference sequence construct (e.g., where the sequence of the sample nucleic acid differs from the corresponding sequence in the reference construct). In some embodiments, the system 100 digests the contents of the native output file and inputs feature data 125 into the test patient data store 120. In other embodiments, the native output file serves as a record of these genomic features 131 in the test patient data store 120.
[0247] In general, the systems described herein can employ any combination of available variant calling software packages and internally developed variant identification algorithms. In some embodiments, the output of a particular algorithm in a variant calling software is further evaluated, for example, to improve variant identification. Thus, in some embodiments, system 100 employs an available variant calling software package to implement some or all of the functionality of one or more of the algorithms shown in feature extraction module 145.
[0248] In various embodiments, the detected genetic variants and genetic features are analyzed as a form of quality control. For example, a pattern of detected genetic variants or features may indicate problems related to the sample, the sequencing procedure, and / or the bioinformatics pipeline (e.g., sample contamination, mislabeling of the sample, changes in reagents, changes in the sequencing procedure and / or the bioinformatics pipeline, etc.).
[0249] FIG. 4E illustrates an exemplary workflow for genomic feature identification (324). This particular workflow is merely an example of one possible set and arrangement of algorithms for feature extraction from sequencing data 124. Generally, any combination of modules and algorithms, for example, those illustrated in FIG. 1A , of feature extraction module 145, can be used in a bioinformatics pipeline. For example, in some embodiments, an architecture useful in the methods and systems described herein includes at least one of the modules or variant calling algorithms shown in feature extraction module 145. In some embodiments, an architecture includes at least two, three, four, five, six, seven, eight, nine, ten, or more of the modules or variant calling algorithms shown in feature extraction module 145. Furthermore, in some embodiments, feature extraction modules and / or algorithms not illustrated in FIG. 1A are utilized in the methods and systems described herein.
[0250] In some embodiments, obtaining, enrolling, storing, preparing, processing, and / or analyzing the biopsy sample of block 502 from the subject includes any of the methods and / or embodiments described above in this disclosure. In some embodiments, the sequencing reaction includes any of the methods and / or embodiments described above in this disclosure.
[0251] In some embodiments, all or substantially all of the aligned sequence reads are evaluated to identify candidate sequence variants (e.g., candidate somatic and / or germline sequence variants). In other embodiments, a subset of the aligned sequence reads is evaluated to identify candidate sequence variants, such as those discussed below in block 504. For example, in one embodiment, a targeted panel sequencing reaction is used to generate sequencing data 122, and only sequence reads corresponding to the target panel (on-target reads) are evaluated to identify candidate sequence variants. In some embodiments, a targeted panel sequencing reaction is used to generate sequencing data 122, and a subset of sequence reads corresponding to a subset of the target panel are evaluated to identify candidate sequence variants. In some embodiments, regardless of whether the sequencing reaction is a targeted panel sequencing reaction, a whole exome sequencing reaction, or a whole genome sequencing reaction, a subset of sequence reads corresponding to a subset of genes is evaluated to identify candidate sequence variants. In some embodiments, a subset of sequence reads corresponding to a defined set of regions within the genome, such as one or more genes, one or more introns, one or more exons, or subregions of one or more introns and / or exons, associated with the etiology of cancer, are evaluated to identify candidate sequence variants.
[0252] Alternatively, in some embodiments, regardless of which subset of aligned sequence reads is evaluated to identify candidate sequence variants, only a subset of the candidate sequence variants are further validated. For example, in some embodiments, only candidate sequence variants corresponding to a target panel (on-target reads) are validated. Similarly, in some embodiments, only candidate sequence variants corresponding to a subset of the target panel are validated. Similarly, in some embodiments, regardless of whether the sequencing reaction is a targeted panel sequencing reaction, a whole exome sequencing reaction, or a whole genome sequencing reaction, only candidate sequence variants corresponding to a subset of genes are validated. Similarly, in some embodiments, only candidate variants corresponding to a set of defined regions within the genome, such as one or more genes, one or more introns, one or more exons, or subregions of one or more introns and / or exons associated with the etiology of cancer, are validated.
[0253] In some embodiments, enrichment is performed before pooling multiple nucleic acid sequencing libraries, however, in other embodiments, enrichment is performed after pooling the nucleic acid sequencing libraries, which has the advantage of reducing the number of enrichment assays that need to be performed.
[0254] In some embodiments, enrichment is performed before generating a nucleic acid sequencing library. This has the advantage that fewer reagents are required to perform both enrichment (because there are fewer target sequences at this point before library amplification) and library production (because there are fewer nucleic acid molecules to tag and amplify after enrichment). However, this increases the possibility of pull-down bias and / or the possibility that slight variations in the enrichment protocol will result in less consistent results.
[0255] In some embodiments, nucleic acid libraries are pooled (two or more DNA libraries can be mixed to create a pool) and treated with a reagent to reduce off-target capture. The pool can be dried and resuspended in a centrifugal concentrator. The DNA library or pool can be hybridized to a probe set (e.g., a probe set specific to a panel including at least 100, 600, 1,000, 10,000, etc. loci of the 19,000 known human genes) and amplified using commercially available reagents. For example, in some embodiments, the pool is incubated in an incubator, PCR machine, water bath, or other temperature-regulating device to allow the probes to hybridize. The pool can then be mixed with streptavidin-coated beads or another means to capture hybridized DNA probe molecules, such as DNA molecules representing exons of the human genome and / or genes selected for the gene panel.
[0256] Pools may be amplified and purified two or more times using commercially available reagents. Pools or DNA libraries may be analyzed to determine the concentration or quantity of DNA molecules, for example, by using fluorescent dyes and a fluorescence microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In one example, DNA library preparation and / or capture is performed using an automated system that uses a liquid handling robot.
[0257] In some embodiments, for example, when whole genome sequencing is used, nucleic acid sequencing libraries are not subjected to target enrichment before sequencing, so as to obtain sequencing data for substantially all competent nucleic acids in the sequencing library.Similarly, in some embodiments, for example, when whole genome sequencing is used, nucleic acid sequencing libraries are not mixed because of the associated processing power limitations in obtaining meaningful sequencing depth across the whole genome.However, in other embodiments, for example, when low-pass whole genome sequencing (LPWGS) is used, nucleic acid sequencing libraries can still be pooled because a very low average sequencing coverage is achieved across each genome, for example, about 0.5x to about 5x.
[0258] In some embodiments, a plurality of nucleic acid probes (e.g., a probe set) is used to enrich for one or more target sequences in a nucleic acid sample (e.g., an isolated nucleic acid sample or a nucleic acid sequencing library), e.g., where one or more target sequences are beneficial for precision oncology. For example, in some embodiments, one or more of the target sequences encompass loci associated with actionable alleles; i.e., variations in the target sequences are relevant to targeted therapeutic approaches. In some embodiments, one or more of the target sequences and / or one or more characteristics of the target sequences are used in a classifier trained to distinguish between two or more cancer states.
[0259] Block 504. Referring to block 504, in some embodiments, the first plurality of nucleic acid sequences is determined from a second panel enrichment sequencing reaction using a second plurality of probes, the second plurality of probes including, for each respective locus in the plurality of loci, a corresponding probe in the second plurality of probes that hybridizes to the respective locus. In other embodiments, the first plurality of nucleic acid sequences is obtained by whole genome sequencing of genomic DNA from a solid tumor sample.
[0260] Advantageously, enriching target sequences prior to sequencing nucleic acids significantly reduces the cost and time associated with sequencing, facilitates multiplex sequencing by allowing multiple samples to be mixed together for a single sequencing reaction, and significantly reduces the computational burden of aligning the resulting sequence reads as a result of significantly reducing the total amount of nucleic acid analyzed from each sample.
[0261] Block 506. Referring to block 506, in some embodiments, the plurality of loci are sequenced in the second panel enrichment sequencing reaction at an average sequence depth of at least 50x, 75x, 100x, 125x, 500x, or 1000x. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of at least 100x. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 100x, at least 500x, at least 1000x, at least 5000x, at least 10,000x, at least 50,000x, or more. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of 100,000x or less, 50,000x or less, 10,000x or less, 5000x or less, or less. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 100x to 50,000x, 100x to 10,000x, 100x to 5000x, 100x to 1000x, or 100x to 500x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 500x to 50,000x, 500x to 10,000x, 500x to 5000x, or 500x to 1000x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 1000x to 50,000x, 1000x to 10,000x, or 1000x to 5000x. In some embodiments, the sequencing depth threshold is a minimum depth selected by a user or practitioner.
[0262] Block 508. Referring to block 508, in some embodiments, the second plurality of probes enriches loci from at least 50 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches at least 25 genes, at least 50 genes, at least 100 genes, at least 250 genes, at least 500 genes, at least 1000 genes, at least 2500 genes, at least 5000 genes, or more. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 40,000 genes or less, 20,000 genes or less, 10,000 genes or less, 5000 genes or less, 2500 genes or less, 1000 genes or less, or fewer genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 genes to 10,000 genes, 25 genes to 5000 genes, 25 genes to 2500 genes, 25 genes to 1000 genes, 25 genes to 500 genes, or 25 genes to 250 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 10,000 genes, 50 genes to 5000 genes, 50 genes to 2500 genes, 50 genes to 1000 genes, 50 genes to 500 genes, or 50 genes to 250 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches between 100 genes and 10,000 genes, between 100 genes and 5000 genes, between 100 genes and 2500 genes, between 100 genes and 1000 genes, between 100 genes and 500 genes, or between 100 genes and 250 genes.
[0263] In some embodiments, the multiple probe sequences used in the first panel enrichment sequencing reaction to enrich cell-free DNA fragments in liquid biopsy samples are collectively mapped to at least 25 different genes in human reference genome.In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches at least 25 human genes, at least 50 human genes, at least 100 human genes, at least 250 human genes, at least 500 human genes, at least 1000 human genes, at least 2500 human genes, at least 5000 human genes or more.In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 40,000 or less human genes, 20,000 or less human genes, 10,000 or less human genes, 5000 or less human genes, 2500 or less human genes, 1000 or less human genes or less. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 10,000 human genes, 25 to 5000 human genes, 25 to 2500 human genes, 25 to 1000 human genes, 25 to 500 human genes, or 25 to 250 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 to 10,000 human genes, 50 to 5000 human genes, 50 to 2500 human genes, 50 to 1000 human genes, 50 to 500 human genes, or 50 to 250 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches between 100 and 10,000 human genes, between 100 and 5000 human genes, between 100 and 2500 human genes, between 100 and 1000 human genes, between 100 and 500 human genes, or between 100 and 250 human genes.
[0264] Thus, in some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of at least 100x. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of at least 100x, at least 500x, at least 1000x, at least 5000x, at least 10,000x, at least 50,000x, or more. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of 100,000x or less, 50,000x or less, 10,000x or less, 5000x or less, or less. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of 100x to 50,000x, 100x to 10,000x, 100x to 5000x, 100x to 1000x, or 100x to 500x. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of 500x to 50,000x, 500x to 10,000x, 500x to 5000x, or 500x to 1000x. In some embodiments, the second panel enrichment sequencing reaction is performed at a read depth of 1000x to 50,000x, 1000x to 10,000x, or 1000x to 5000x.
[0265] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches (contains probes for) 50 to 150 genes. In some embodiments, the multiple probe sequences used to enrich cell-free DNA fragments in the second panel enrichment sequencing reaction collectively map to 25 to 150 different genes in the human reference genome. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches 50 to 150 genes, 100 to 200 genes, 150 to 300 genes, or 250 to 500 genes. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches 50 to 1000 genes, 60 to 800 genes, 70 to 700 genes, 80 to 600 genes, or 90 to 500 genes. In some embodiments, each of the enriched genes in the sequencing panel is a human gene.
[0266] In some embodiments, the second plurality of probe sequences used to enrich nuclear fragments in sample in the second panel enrichment sequencing reaction are collectively mapped to at least 25 different genes in human reference genome.In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches at least 25 human genes, at least 50 human genes, at least 100 human genes, at least 250 human genes, at least 500 human genes, at least 1000 human genes, at least 2500 human genes, at least 5000 human genes or more.In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches 1,000 or less human genes, 500 or less human genes, 250 or less human genes, 200 or less human genes, 175 or less human genes, 100 or less human genes or less. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 1,000 human genes, 25 to 500 human genes, 10 to 250 human genes, 10 to 200 human genes, 5 to 150 human genes, or 5 to 100 human genes. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 400 human genes, 30 to 500 human genes, 50 to 300 human genes, 5 to 95 human genes, 15 to 130 human genes, or 15 to 165 human genes. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 human genes to 600 human genes, 40 human genes to 80 human genes, 35 human genes to 95 human genes, 45 human genes to 80 human genes, 20 human genes to 80 human genes, or 20 human genes to 120 human genes.
[0267] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 100 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the sequencing panel enriches only for genes in Table 1, while in other embodiments, the sequencing panel enriches for some genes that are in Table 1 and some genes that are not in Table 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for between 25 different genes and 150 different genes listed in Table 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 150 genes, 100 genes to 200 genes, or 150 genes to 300 genes listed in Table 1.
[0268] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 100 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the sequencing panel enriches only for genes in Table 2, while in other embodiments, the sequencing panel enriches for some genes that are in Table 2 and some genes that are not in Table 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for between 25 different genes and 150 different genes listed in Table 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 150 genes, 100 genes to 200 genes, or 150 genes to 300 genes listed in Table 2.
[0269] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set includes probes targeting one or more of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 5 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 10 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 25 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 50 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 75 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 100 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting all of the genes listed in Table 1.
[0270] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set includes probes targeting one or more of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 5 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 10 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 25 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 50 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 75 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 100 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting all of the genes listed in Table 2.
[0271] Table 1. Example of a 105-gene panel. [Table 1]
[0272] Table 2. Example of a 523-gene panel. [Table 2-1] [Table 2-2] [Table 2-3]
[0273] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in List 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in List 1. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all genes listed in List 1, for 10 to all genes listed in List 1, or for 20 to all genes listed in List 1.
[0274] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set consists of or includes probes targeting one or more of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 5 of the genes listed in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 10 of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 25 of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 50 of the genes listed in List 1. In some such embodiments, the probe set consists of or includes probes targeting all of the genes in List 1.
[0275] リスト1:AKT1 (14q32.33), ALK (2p23.2-23.1), APC (5q22.2), AR (Xq12), ARAF (Xp11.3), ARID1A (1p36.11), ATM (11q22.3), BRAF (7q34), BRCA1 (17q21.31), BRCA2 (13q13.1), CCND1 (11q13.3), CCND2 (12p13.32), CCNE1 (19q12), CDH1 (16q22.1), CDK4 (12q14.1), CDK6 (7q21.2), CDKN2A (9p21.3), CTNNB1 (3p22.1), DDR2 (1q23.3), EGFR (7p11.2), ERBB2 (17q12), ESR1 (6q25.1-25.2), EZH2 (7q36.1), FBXW7 (4q31.3), FGFR1 (8p11.23), FGFR2 (10q26.13), FGFR3 (4p16.3), GATA3 (10p14), GNA11 (19p13.3), GNAQ (9q21.2), GNAS (20q13.32), HNF1A (12q24.31), HRAS (11p15.5), IDH1 (2q34), IDH2 (15q26.1), JAK2 (9p24.1), JAK3 (19p13.11), KIT (4q12), KRAS (12p12.1), MAP2K1 (15q22.31), MAP2K2 (19p13.3), MAPK1 (22q11.22), MAPK3 (16p11.2), MET (7q31.2), MLH1 (3p22.2), MPL (1p34.2), MTOR (1p36.22), MYC (8q24.21), NF1 (17q11.2), NFE2L2 (2q31.2), NOTCH1 (9q34.3), NPM1 (5q35.1), NRAS (1p13.2), NTRK1 (1q23.1), NTRK3 (15q25.3), PDGFRA (4q12), PIK3CA (3q26.32), PTEN (10q23.31), PTPN11 (12q24.13), RAF1 (3p25.2), RB1 (13q14.2), RET (10q11.21), RHEB (7q36.1), RHOA (3p21.31), RIT1 (1q22), ROS1 (6q22.1), SMAD4 (18q21.2), SMO (7q32.1), STK11 (19p13.3), TERT (5p15.33), TP53 (17p13.1), TSC1 (9q34.13), and VHL (3p25.3). .
[0276] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in List 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in List 2. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all genes listed in List 2, for 10 to all genes listed in List 2, or for 20 to all genes listed in List 2.
[0277] In some embodiments, the probe set includes probes targeting no more than 50 genes, no more than 100 genes, no more than 150 genes, or no more than 200 genes. In some such embodiments, the probe set consists of or includes probes targeting one or more of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 5 of the genes listed in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 10 of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 25 of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 50 of the genes listed in List 2. In some such embodiments, the probe set consists of or includes probes targeting all of the genes in List 2.
[0278] Package 2:ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1 (FAM123B), APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, C11orf30 (EMSY), C17orf39 (GID4), CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274 (PD-L1), CD70, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, EZH2, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R,IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A, KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1 (MEK1), MAP2K2 (MEK2), MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYC, MYCL (MYCL1), MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NSD3 (WHSC1L1), NT5C2, NTRK1, NTRK2, NTRK3, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1 (PD-1), PDCD1LG2 (PD-L2), PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1,PRKAR1A,PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC,STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, ncRNA, TGFBR2, TIPARP, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WT1, XPO1, XRCC2, ZNF217, and ZNF703. ,
[0279] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, 50, or all of the genes listed in FIG. 14 (any combination of FIG. 14A, FIG. 14B, FIG. 14C, FIG. 14D, FIG. 14E, FIG. 14F, FIG. 14G, FIG. 14H, FIG. 14I, FIG. 14j, FIG. 14K, FIG. 14L, and FIG. 14M, collectively "FIG. 14"). In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in FIG. 14. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all of the genes listed in FIG. 14, for 10 to all of the genes listed in FIG. 14, or for 20 to all of the genes listed in FIG. 14. Although Figure 14 and List 2 provide the same genes, in preferred embodiments, Figure 14 indicates the types of variants that are in such genes.
[0280] In some embodiments, a probe set includes probes targeting no more than 50 genes, no more than 100 genes, no more than 150 genes, or no more than 200 genes. In some such embodiments, a probe set consists of or includes probes targeting one or more of the genes in Figure 14. In some such embodiments, a probe set consists of or includes probes targeting at least five of the genes listed in Figure 14. In some such embodiments, a probe set consists of or includes probes targeting at least 10 of the genes in Figure 14. In some such embodiments, a probe set consists of or includes probes targeting at least 25 of the genes in Figure 14. In some such embodiments, a probe set consists of or includes probes targeting at least 50 of the genes listed in Figure 14. In some such embodiments, a probe set consists of or includes probes targeting all of the genes in Figure 14.
[0281] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in any of Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14j, 14K, 14L, and 14M.
[0282] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14j, 14K, 14L, and 14M.
[0283] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14j, 14K, 14L, and 14M.
[0284] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 150 different genes listed in any of Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14j, 14K, 14L, and 14M. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 to 150 genes, 100 to 200 genes, or 150 to 300 genes listed in any of Figures 14A, 14B, 14C, 14D, 14E, 14F, 14G, 14H, 14I, 14j, 14K, 14L, and 14M.
[0285] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, 50, or all of the genes listed in FIG. 15 (any combination of FIGS. 15A, 15B, and 15C, collectively "FIG. 15"). In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in FIG. 15. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all of the genes listed in FIG. 14, 10 to all of the genes listed in FIG. 15, or 20 to all of the genes listed in FIG. 15.
[0286] In some such embodiments, the probe set consists of or includes probes targeting one or more of the genes in Figure 15. In some such embodiments, the probe set consists of or includes probes targeting at least five of the genes listed in Figure 15. In some such embodiments, the probe set consists of or includes probes targeting at least 10 of the genes in Figure 15. In some such embodiments, the probe set consists of or includes probes targeting at least 25 of the genes in Figure 15. In some such embodiments, the probe set consists of or includes probes targeting at least 50 of the genes listed in Figure 15. In some such embodiments, the probe set consists of or includes probes targeting all of the genes in Figure 15.
[0287] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in any of Figure 145, Figure 15B, and Figure 15C.
[0288] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 15A, 15B, and 15C.
[0289] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 15A, 15B, and 15C.
[0290] In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 150 different genes listed in any of Figures 15A, 15B, and 15C. In some embodiments, the second panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 to 150 genes, 100 to 200 genes, or 150 to 300 genes listed in any of Figures 15A, 15B, and 15C.
[0291] Generally, probes for enrichment of nucleic acids (e.g., cfDNA obtained from a liquid biopsy sample) comprise DNA, RNA, or modified nucleic acid structures having a base sequence complementary to a locus of interest. For example, a probe designed to hybridize to a locus in a cfDNA molecule can comprise a sequence complementary to either strand, since cfDNA molecules are double-stranded. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 consecutive bases of the locus of interest. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 20, 25, 30, 40, 50, 75, 100, 150, 200, or more consecutive bases of the locus of interest.
[0292] Target panel provides several benefits for nucleic acid sequencing.For example, in some embodiments, for example, the algorithm for distinguishing between a first cancer state and a second cancer state can be trained with a smaller and more informative data set (for example, fewer genes), which leads to more computationally efficient training of the classifier for distinguishing between a first cancer state and a second cancer state.This improvement in computational efficiency due to the reduction in the size of the gene set for distinguishing can be advantageously used to accelerate classifier training, or can be used to improve the performance of such classifier (for example, through more extensive training of classifier).
[0293] In some embodiments, the gene panel is a whole exome panel that analyzes the exome of a biological sample. In some embodiments, the gene panel is a whole genome panel that analyzes the genome of a subject.
[0294] In some embodiments, the probe comprises an additional nucleic acid sequence that does not share any homology with the locus of interest. For example, in some embodiments, the probe also comprises an identifier sequence, e.g., a nucleic acid sequence comprising a unique molecular identifier (UMI), that is specific to a particular sample or subject. Examples of identifier sequences are described in Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, which are incorporated herein by reference. Similarly, in some embodiments, the probe also comprises a primer nucleic acid sequence useful for amplifying the nucleic acid molecule of interest, e.g., using PCR. In some embodiments, the probe also comprises a capture sequence designed to hybridize to an anti-capture sequence to retrieve the nucleic acid molecule of interest from the sample.
[0295] Similarly, in some embodiments, each probe comprises a non-nucleic acid affinity moiety covalently bound to a nucleic acid molecule complementary to a target locus for retrieving the target nucleic acid molecule. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, such as a dipstick or magnetic bead, for retrieving the target nucleic acid. In some embodiments, the methods described herein include amplifying the nucleic acid bound to the probe set before further analysis, such as sequencing. Methods for amplifying nucleic acids, for example, by PCR, are well known in the art.
[0296] Next-generation sequencing generates millions of short reads (e.g., sequence reads) for each biological sample. Thus, in some embodiments, the multiple sequence reads obtained by next-generation sequencing of cfDNA molecules are DNA sequence reads. In some embodiments, the sequence reads have an average length of at least 50 nucleotides. In other embodiments, the sequence reads have an average length of at least 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, or more nucleotides.
[0297] In some embodiments, sequencing is performed after enriching nucleic acids (e.g., cfDNA, gDNA, and / or RNA) containing multiple predetermined target sequences, e.g., human genes and / or non-coding sequences associated with cancer. Advantageously, sequencing a nucleic acid sample enriched for target nucleic acids, rather than all nucleic acids isolated from a biological sample, significantly reduces the average time and cost of the sequencing reaction. Thus, in some preferred embodiments, the methods described herein include obtaining multiple sequence reads of nucleic acids hybridized to a probe set for hybrid capture enrichment (e.g., for one or more genes listed in Table 1, or one or more genes listed in Table 2, one or more genes listed in List 1, one or more genes listed in List 2, or one or more genes listed in Figure 14).
[0298] In some embodiments, the second panel targeted sequencing is performed to an average on-target depth of at least 50x, at least 100x, at least 125x, at least 150x, at least 500x, at least 750x, at least 1000x, at least 2500x, at least 10,000x, or more. In some embodiments, the samples are further evaluated for uniformity above a sequencing depth threshold (e.g., 95% of all target base pairs at 300x sequencing depth). In some embodiments, the sequencing depth threshold is a minimum depth selected by a user or practitioner.
[0299] Block 509. Referring to block 509, in some embodiments, a first plurality of probes, including a corresponding probe that hybridizes to each respective locus in the plurality of loci, is used to obtain a second plurality of nucleic acid sequences, including a corresponding nucleic acid sequence for each cell-free DNA fragment in the plurality of cell-free DNA fragments obtained from the liquid biopsy sample from the first panel enrichment sequencing assay.
[0300] In some embodiments, the second plurality of nucleic acids comprises at least 10,000 sequences, at least 50,000 sequences, at least 100,000 sequences, at least 500,000 sequences, at least 1 million sequences, at least 5 million sequences, at least 10 million sequences, or more. In some embodiments, the second plurality of sequences comprises no more than 1 billion sequences, no more than 500 million sequences, no more than 100 million sequences, no more than 50 million sequences, no more than 10 million sequences, no more than 5 million sequences, no more than 1 million sequences, or fewer. In some embodiments, the second plurality of sequences is between 10,000 sequences and 1 billion sequences, between 10,000 sequences and 500 million sequences, between 10,000 sequences and 100 million sequences, between 10,000 sequences and 50 million sequences, between 10,000 sequences and 10 million sequences, between 10,000 sequences and 5 million sequences, or between 10,000 sequences and 1 million sequences. In some embodiments, the second plurality of sequences is between 100,000 sequences and 1 billion sequences, between 100,000 sequences and 500 million sequences, between 100,000 sequences and 100 million sequences, between 100,000 sequences and 50 million sequences, between 100,000 sequences and 10 million sequences, between 100,000 sequences and 5 million sequences, or between 100,000 sequences and 1 million sequences. In some embodiments, the second plurality of sequences is 500,000 to 1 billion sequences, 500,000 to 500 million sequences, 500,000 to 100 million sequences, 500,000 to 50 million sequences, 500,000 to 10 million sequences, 500,000 to 5 million sequences, or 500,000 to 1 million sequences. In some embodiments, the second plurality of sequences is 1 million to 1 billion sequences, 1 million to 500 million sequences, 1 million to 100 million sequences, 1 million to 50 million sequences, 1 million to 10 million sequences, or 1 million to 5 million sequences.
[0301] In some embodiments, the first plurality of probes used to enrich cell-free DNA fragments in a liquid biopsy sample in the first panel enrichment sequencing reaction are collectively mapped to at least 25 different genes in a human reference genome. In some embodiments, the first plurality of probes are collectively mapped to at least 50, at least 100, at least 250, at least 500, or at least 1000 different genes in a human reference genome. In some embodiments, the first plurality of probes are collectively mapped to at least 10 of the genes listed in Table 1. In some embodiments, the first plurality of probes are collectively mapped to at least 20, 25, 30, 40, 50, 60, 75, 100, or all 105 of the genes listed in Table 1, Table 2, List 1, List 2, and / or Figure 14.
[0302] For example, in some embodiments, the target enrichment panel (first plurality of probes) of block 509 includes any of the embodiments of the second plurality of probes described in block 504 for the second panel enrichment sequencing reaction.
[0303] In some embodiments, the target enrichment panel of block 509 includes probes targeting one or more genetic loci, e.g., exon or intron loci. In some embodiments, the target enrichment panel of block 509 includes probes targeting one or more non-protein-coding loci, e.g., regulatory loci, miRNA loci, and other non-coding loci, e.g., found to be associated with cancer. In some embodiments, the plurality of loci targeted in block 509 includes at least 25, 50, 100, 150, 200, 250, 300, 350, 400, 500, 750, 1000, 2500, 5000, or more human genomic loci.
[0304] In some embodiments, the target enrichment panel of block 509 includes probes targeting one or more of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least five of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least 10 of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least 25 of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least 50 of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least 75 of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting at least 100 of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15. In some embodiments, the target enrichment panel of block 509 includes probes targeting all of the genes listed in Table 1, Table 2, List 1, List 2, FIG. 14, and / or FIG. 15.
[0305] In some embodiments, obtaining, enrolling, storing, preparing, processing, and / or analyzing the liquid biopsy sample from the subject comprises any of the methods and / or embodiments described above in this disclosure. In some embodiments, the sequencing reaction comprises any of the methods and / or embodiments described above in this disclosure.
[0306] Block 510. Referring to block 510, in some embodiments, the first plurality of probes and the second plurality of probes are different. That is, in some embodiments, a different probe set is used to enrich genomic DNA sequences from a solid tumor sample than the probe set used to enrich cell-free DNA from a liquid biopsy sample.
[0307] Indeed, in some embodiments, the only requirement regarding the relationship between probe sets used to enrich DNA from solid tumor samples is that there be some overlap in the genomic regions drawn, such that variable allele frequencies for one or more somatic mutations identified from the solid tumor sample can be used to estimate the circulating tumor fraction of the liquid biopsy sample.
[0308] In some embodiments, the only requirement regarding the relationship between the probe sets used to enrich DNA from a solid tumor sample is that there is some overlap in the genes enriched by the probe set used in block 509 and the probe set used in block 504. In some embodiments, the first plurality of nucleic acid sequences is obtained using whole genome sequencing, and therefore there is overlap between the first plurality of nucleic acid sequences in block 502 and the second plurality of nucleic acid sequences in block 509.
[0309] In some embodiments, the first plurality of nucleic acid sequences and the second plurality of nucleic acid sequences of block 509 each collectively map to 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more of the same gene.
[0310] In some embodiments, the first plurality of nucleic acid sequences and the second plurality of nucleic acid sequences of block 509 each collectively map to between 5 and 500, between 10 and 1000, between 20 and 500, between 30 and 2000, or between 40 and 400 of the same genes.
[0311] Block 512. Referring to block 512, in some embodiments, the plurality of loci are sequenced in the first panel enrichment sequencing reaction at an average sequencing depth of at least 500x. In some embodiments, the first panel enrichment sequencing reaction is performed at a read depth of at least 1,000x. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 100x, at least 500x, at least 1000x, at least 5000x, at least 10,000x, at least 50,000x, or more. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of 100,000x or less, 50,000x or less, 10,000x or less, 5000x or less, or less. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 100x to 50,000x, 100x to 10,000x, 100x to 5000x, 100x to 1000x, or 100x to 500x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 500x to 50,000x, 500x to 10,000x, 500x to 5000x, or 500x to 1000x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 1000x to 50,000x, 1000x to 10,000x, or 1000x to 5000x.
[0312] Block 514. Referring to block 514, in some embodiments, the first plurality of probes enriches loci from at least 50 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches at least 25 genes, at least 50 genes, at least 100 genes, at least 250 genes, at least 500 genes, at least 1000 genes, at least 2500 genes, at least 5000 genes, or more. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 40,000 genes or less, 20,000 genes or less, 10,000 genes or less, 5000 genes or less, 2500 genes or less, 1000 genes or less, or fewer genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 genes to 10,000 genes, 25 genes to 5000 genes, 25 genes to 2500 genes, 25 genes to 1000 genes, 25 genes to 500 genes, or 25 genes to 250 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 10,000 genes, 50 genes to 5000 genes, 50 genes to 2500 genes, 50 genes to 1000 genes, 50 genes to 500 genes, or 50 genes to 250 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches between 100 genes and 10,000 genes, between 100 genes and 5000 genes, between 100 genes and 2500 genes, between 100 genes and 1000 genes, between 100 genes and 500 genes, or between 100 genes and 250 genes.
[0313] In some embodiments, the multiple probe sequences used in the second panel enrichment sequencing reaction to enrich genomic DNA sequences from solid tumor samples are collectively mapped to at least 25 different genes in human reference genome.In some embodiments, the panel enrichment sequencing reaction uses the sequencing panel that enriches at least 25 human genes, at least 50 human genes, at least 100 human genes, at least 250 human genes, at least 500 human genes, at least 1000 human genes, at least 2500 human genes, at least 5000 human genes or more.In some embodiments, the panel enrichment sequencing reaction uses the sequencing panel that enriches 40,000 or less human genes, 20,000 or less human genes, 10,000 or less human genes, 5000 or less human genes, 2500 or less human genes, 1000 or less human genes or less. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 10,000 human genes, 25 to 5000 human genes, 25 to 2500 human genes, 25 to 1000 human genes, 25 to 500 human genes, or 25 to 250 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 to 10,000 human genes, 50 to 5000 human genes, 50 to 2500 human genes, 50 to 1000 human genes, 50 to 500 human genes, or 50 to 250 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches between 100 and 10,000 human genes, between 100 and 5000 human genes, between 100 and 2500 human genes, between 100 and 1000 human genes, between 100 and 500 human genes, or between 100 and 250 human genes.
[0314] Block 516. Referring to block 516, in some embodiments, the identity of the first plurality of probes is non-customized to the subject. Advantageously, the methods and systems described herein facilitate tumor-informed estimation of tumor fraction without requiring custom sequencing of liquid biopsy samples. In this manner, sequencing data from any number of commercially available panel enrichment sequencing assays for evaluating solid tumor biopsies can be used to inform the evaluation of any number of commercially available panel enrichment sequencing assays on liquid biopsy samples.
[0315] Block 518. Referring to block 518, in some embodiments, the solid tumor sample is collected prior to collection of the liquid biopsy sample. In some embodiments, the solid tumor sample is collected at least one day before collection of the liquid biopsy sample. In some embodiments, the solid tumor sample is collected at least one week, at least one month, at least two months, at least three months, at least four months, at least five months, at least six months, at least nine months, at least twelve months, at least eighteen months, at least two years, or more before collection of the liquid biopsy sample. In some embodiments, the solid tumor sample is collected five years or less before collection of the liquid biopsy sample. In some embodiments, the solid tumor sample is collected four years or less, three years or less, two years or less, 18 months or less, 12 months or less, nine months or less, six months or less, five months or less, four months or less, three months or less, two months or less, one month or less, three weeks or less, two weeks or less, one week or less, or less than before collection of the liquid biopsy sample. In some embodiments, the solid tumor sample is collected 1 day to 5 years before the liquid biopsy sample is collected. In some embodiments, the solid tumor sample is collected 1 day to 2 years, 1 day to 1 year, 1 day to 6 months, 1 day to 3 months, 1 day to 1 month, or 1 day to 1 week before the liquid biopsy sample is collected. In some embodiments, the solid tumor biopsy and liquid biopsy sample are collected on the same day. In some embodiments, the liquid biopsy sample is collected before the solid tumor biopsy is collected.
[0316] Block 520. Referring to block 520, in some embodiments, the solid tumor sample and the liquid biopsy sample are taken within 6 months of each other. In some embodiments, the solid tumor sample and the liquid biopsy sample are taken within 1 day, 1 week, 2 weeks, 1 month, 3 months, 6 months, 9 months, 12 months, 18 months, 2 years, 3 years, 4 years, or 5 years of each other.
[0317] Blocks 522-524. Referring to block 522, in some embodiments, the liquid biopsy sample is blood. Referring to block 524, in some embodiments, the liquid biopsy sample comprises blood, whole blood, peripheral blood, plasma, serum, or lymph from the subject. In some embodiments, one or more of the biological samples obtained from the patient are biological liquid samples, also referred to as liquid biopsy samples. In some embodiments, one or more of the biological samples obtained from the patient are selected from blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of a testicle), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple effluent, aspirates from different parts of the body (e.g., thyroid, breast), and the like. In some embodiments, the liquid biopsy sample comprises blood and / or saliva. In some embodiments, the liquid biopsy sample is peripheral blood. In some embodiments, one or more blood samples are collected from the subject in a commercially available blood collection container. In some embodiments, a saliva sample is collected from a patient in a commercially available saliva collection container.
[0318] In some embodiments, the volume of the first liquid biopsy sample is less than 30 mL. In some embodiments, the volume of the liquid biopsy sample is between 1 mL and 50 mL, between 2 mL and 40 mL, between 3 mL and 35 mL, or between 5 mL and 31 mL. For example, in some embodiments, the liquid biopsy sample has a volume of about 1 mL, about 2 mL, about 3 mL, about 4 mL, about 5 mL, about 6 mL, about 7 mL, about 8 mL, about 9 mL, about 10 mL, about 11 mL, about 12 mL, about 13 mL, about 14 mL, about 15 mL, about 16 mL, about 17 mL, about 18 mL, about 19 mL, about 20 mL, or more.
[0319] As described herein, cfDNA is a particularly useful source of biological data for various embodiments of the methods and systems described herein because it can be easily obtained from various bodily fluids. Advantageously, the use of bodily fluids facilitates continuous monitoring due to the ease of collection, as these fluids can be collected by non-invasive or minimally invasive methodologies. This is in contrast to methods that rely on solid tissue samples, such as biopsies, which often require invasive surgical procedures. Furthermore, because bodily fluids such as blood circulate throughout the body, cfDNA populations represent samples of many different tissue types from many different locations.
[0320] In some embodiments, a liquid biopsy sample is separated into two different samples, for example, in some embodiments, a blood sample is separated into a plasma sample containing cfDNA and a buffy coat preparation containing white blood cells.
[0321] In some embodiments, multiple liquid biopsy samples are obtained from each subject at intervals over a period of time (e.g., using serial testing). For example, in some such embodiments, the time between obtaining liquid biopsy samples from each subject is at least 1 day, at least 2 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 6 months, or at least 1 year.
[0322] Liquid biopsy samples contain cell-free nucleic acids, including cell-free DNA (cfDNA).As described above, the cfDNA isolated from cancer patients includes DNA derived from cancer cells, also referred to as circulating tumor DNA (ctDNA), cfDNA derived from germline (e.g., healthy or non-cancerous) cells, and cfDNA derived from hematopoietic cells (e.g., leukocytes).The relative proportions of cancerous and non-cancerous cfDNA present in liquid biopsy samples vary depending on the characteristics of the patient's cancer (e.g., type, stage, lineage, genomic profile, etc.).
[0323] In some embodiments, cell-free DNA is isolated from a liquid biological sample using commercially available reagents, including digestion with proteinase K. In some embodiments, the selective binding properties of a silica membrane are used to extract cell-free DNA from a first liquid biological sample using a circulating nucleic acid kit. In some such embodiments, the liquid biological sample is dissolved in an optimized buffer and adjusted to binding conditions. The liquid biological sample is then loaded directly onto a spin column. In this step, cell-free DNA binds to the silica membrane, and a washing step removes contaminants. Finally, pure cell-free DNA is eluted in a small volume of low-salt buffer for downstream applications. See Hai et al., 2022, "Whole-genome circulating tumor DNA methylation landscape reveals sensitive biomarkers of breast cancer," MedComm (2020) Sep 3(3): e134, which is incorporated herein by reference.
[0324] In some embodiments, adaptors, such as unique dual index (UDI) adaptors, are ligated onto cell-free DNA fragments. In some embodiments, adaptors having unique molecular indexes (UMIs), which are short nucleic acid sequences (e.g., 4-10 base pairs), are ligated onto cell-free DNA fragments. In some embodiments, the UDI adaptors comprise a UMI. In some embodiments, the UMI is a degenerate base pair that serves as a unique tag that can be used to identify sequence reads derived from a specific DNA fragment. In some embodiments, for example, when multiplex sequencing is used to sequence cell-free DNA fragments from multiple samples (e.g., from the same or different subjects) in a single sequencing reaction, a patient-specific index is also added to the nucleic acid molecule. In some embodiments, the sample-specific index is a short nucleic acid sequence (e.g., 3-20 nucleotides) added to the end of the cell-free DNA fragment during library construction that serves as a unique tag that can be used to identify sequence reads derived from a specific patient sample.
[0325] In some embodiments, the adapters include a PCR primer landing site designed for efficient binding of a PCR or second-strand synthesis primer used during the sequencing reaction. In some embodiments, the adapters include an anchor binding site to facilitate binding of cell-free DNA fragments to anchor oligonucleotide molecules on a sequencer flow cell, serving as a seed for the sequencing process by providing a starting point for the sequencing reaction. During PCR amplification after adapter ligation, the UMI, patient index, and binding site are replicated along with the attached cell-free DNA fragment. This provides a method for identifying sequence reads derived from the same original cell-free DNA fragment in downstream analysis.
[0326] In some embodiments, the sequence reads in the first plurality of sequence reads are trimmed to remove sequencing adaptors, amplification primers, and low-quality bases at the ends of the reads. This can be done, for example, using trim_galore (version 0.4.2) or cutadept. See internet URLs bioinformatics.babraham.ac.uk / projects / trim_galore and Martin, 2011, "Cutadapt removes adaptor sequences from high-throughput sequencing reads," EMBnet.journal, [Sl] 17(1), pp. 10-12, respectively.
[0327] In some embodiments, the cell-free DNA fragments are amplified and purified using commercially available reagents. In some such embodiments, the concentration and / or amount of the cell-free DNA fragments are quantified using a fluorescent dye and a fluorescent microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In some embodiments, library amplification is performed on a device (e.g., Illumina C-Bot2), and the resulting flow cells containing the amplified cell-free DNA fragments are sequenced.
[0328] In some embodiments, sequencing is performed on a next-generation sequencer (e.g., Illumina HiSeq 4000, Illumina NovaSeq 6000, Oxford Nanopore, Biomodal) to a specific on-target depth selected by the user. In some embodiments, sequencing is performed using sequencing-by-synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), or sequencing-by-ligation (SOLiD sequencing).
[0329] Referring to FIG. 2 , in some embodiments, biological samples collected from a patient are optionally sent to various analytical environments (e.g., sequencing lab 230, pathology lab 240, and / or molecular biology lab 250) for processing (e.g., data collection) and / or analysis (e.g., feature extraction). Wet-lab processing 204 may include sample cataloging (e.g., enrollment), clinical characterization of one or more samples (e.g., pathology review), and nucleic acid sequence analysis (e.g., extraction, library preparation, capture+hybridization, pooling, and sequencing). In some embodiments, the workflow includes clinical analysis of one or more biological samples collected from the subject, for example, in pathology lab 240 and / or molecular and cell biology lab 250, to generate clinical features such as pathology features 128-3, image data 128-3, and / or tissue culture / organoid data 128-3.
[0330] In some embodiments, pathology data 128-1 collected during a clinical evaluation includes, for example, visual features identified by a pathologist's examination of a specimen (e.g., a solid tumor biopsy) on a stained H&E or IHC slide. In some embodiments, the sample is a solid tissue biopsy sample. In some embodiments, the tissue biopsy sample is formalin-fixed tissue (FFT), e.g., formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the tissue biopsy sample is an FFPE or FFT block. In some embodiments, the tissue biopsy sample is a fresh-frozen tissue biopsy. The tissue biopsy sample can be prepared in thin sections (e.g., by cutting and / or mounting on slides) to facilitate pathology review (e.g., by staining with immunohistochemical stains for IHC review and / or with hematoxylin and eosin stains for H&E pathology review). For example, analysis of slides for H&E or IHC staining may reveal characteristics such as tumor infiltration, programmed death-ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.
[0331] In some embodiments, liquid samples (e.g., blood) collected from patients (e.g., in EDTA-containing collection tubes) are prepared (e.g., by smearing) on slides for pathology review. In some embodiments, grossly dissected FFPE tissue sections, which may be mounted on histopathology slides from solid tissue samples (e.g., tumor or normal tissue), are analyzed by a pathologist. In some embodiments, tumor samples are evaluated to determine, for example, the tumor purity of the sample, the tumor cellularity rate as a ratio of tumor to normal nuclei, etc. For each section, background tissue can be excluded or removed so that the section meets a tumor purity threshold, e.g., at least 20% of the nuclei in the section are tumor nuclei, or at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the nuclei in the section are tumor nuclei.
[0332] In some embodiments, pathology data 128-1 is extracted using a computational approach to digital pathology in addition to or instead of visual inspection, providing, for example, morphological features extracted from digital images of stained tissue samples. In some embodiments, pathology data 128-1 includes features determined using machine learning algorithms to evaluate pathology data collected as described above.
[0333] Further details regarding methods, systems, and algorithms for using pathology data to classify cancer and identify targeted treatments are discussed, for example, in U.S. Patent Nos. 10,957,041, 11,244,763, 11,848,107, and 11,145,416, the contents of each of which are incorporated herein by reference in their entirety and for all purposes.
[0334] In some embodiments, the image data 128-2 collected during clinical evaluation includes features identified by review of in vitro and / or in vivo imaging results (e.g., of a tumor site), e.g., tumor size, differences in tumor size over time (e.g., during treatment or other changes), etc. In some embodiments, the image data 128-2 includes features determined using machine learning algorithms to evaluate the image data collected as described above.
[0335] Further details regarding methods, systems, and algorithms for using medical images to classify cancers and identify targeted treatments are discussed, for example, in U.S. Patent Nos. 10,957,041, 11,244,763, 11,848,107, and 11,145,416, the contents of each of which are incorporated herein by reference in their entirety and for all purposes.
[0336] In some embodiments, the tissue culture / organoid data 128-3 collected during clinical evaluation includes features identified by evaluation of cultured tissue from a subject. For example, in some embodiments, tissue samples (e.g., tumor tissue, normal tissue, or both) obtained from a patient are cultured (e.g., in liquid culture, solid culture, and / or organoid culture) and various features, such as cell morphology, growth characteristics, genomic alterations, and / or drug sensitivity, are evaluated. In some embodiments, the tissue culture / organoid data 128-3 includes features determined using machine learning algorithms to evaluate the tissue culture / organoid data collected as described above. Examples of culturing tissue organoids (e.g., individual tumor organoids) and their feature extraction are described in PCT Publication WO 2021 / 081253 and U.S. Patent No. 11,629,385, the contents of each of which are incorporated herein by reference in their entirety and for all purposes.
[0337] In some embodiments, the method further includes obtaining a liquid biopsy sample from a sample repository or sample database (e.g., BioIVT, TSC Biosample Repository, BioLINCC, etc.). In some embodiments, the liquid biopsy sample is obtained from the subject at least 1 hour, at least 2 hours, at least 12 hours, at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year prior to processing and / or sequencing the liquid biopsy sample. In some such embodiments, the liquid biopsy sample is fresh, frozen, dried, and / or fixed. In some embodiments, the liquid biopsy sample is processed and / or sequenced at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year prior to obtaining the first dataset. For example, in some embodiments, sequencing data for the liquid biopsy sample is obtained from a data repository (e.g., GenBank, NCBI Assembly, DNA DataBank of Japan, European Nucleotide Archive, European Variation Archive, etc.).
[0338] Block 525. Referring to block 525 of Figure 5C, method 500 also includes identifying one or more somatic mutations in the first plurality of nucleic acid sequences, each respective somatic variant in the one or more somatic variants at corresponding one or more nucleotide positions at corresponding loci in the plurality of one or more loci.
[0339] Example 1 provides an illustration of block 525.
[0340] 2A , in some embodiments, nucleic acid sequencing data 122 generated from one or more patient samples (e.g., a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from a subject) is evaluated (e.g., via variant analysis 206) using bioinformatics module 140 of system 100, e.g., in a bioinformatics pipeline, to identify genomic alterations in the patient's cancer genome. An exemplary overview of a bioinformatics pipeline is described herein with respect to FIGS. 4A-4E. Advantageously, in some embodiments, the present disclosure improves bioinformatics pipelines such as pipeline 206 by improving circulating tumor fraction estimates. In some embodiments where sequencing data 122 was obtained using panel-based sequencing, analysis is limited to the plurality of loci encompassed by the panel-based sequencing.
[0341] 4A shows an exemplary bioinformatics pipeline 206 for providing clinical support for precision oncology (e.g., used for feature extraction in the workflows illustrated in FIGS. 2A and 3). As shown in FIG. 4A, sequencing data 122 (e.g., sequence reads 314, a first plurality of nucleic acid sequences including a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from a subject) obtained from wet-lab processing 204 is input into the pipeline.
[0342] In various embodiments, the bioinformatics pipeline includes a circulating tumor DNA (ctDNA) pipeline for analyzing liquid biopsy samples. In some embodiments, the pipeline detects SNVs, INDELs, copy number amplifications / deletions, and genomic rearrangements (e.g., fusions). In some embodiments, the pipeline employs unique molecular identifier (UMI)-based consensus base calling as a method of error suppression and Bayesian trinucleotide context-based position-level error suppression. In various embodiments, the pipeline is capable of detecting variants with a variant allele fraction of 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.4%, or 0.5%.
[0343] As shown in Example 1, in some embodiments according to block 525, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more somatic mutations are identified using a first plurality of nucleic acid sequences that includes a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from the subject. In some embodiments, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 or more somatic mutations are identified using a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from the subject.
[0344] In some embodiments according to block 525, fewer than 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 somatic mutations are identified using a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from the subject.
[0345] In some embodiments according to block 525, between 30, between 3 and 40, between 3 and 50, between 3 and 60, between 3 and 70, between 3 and 80, between 3 and 90, between 3 and 100, between 3 and 110, between 3 and 120, between 3 and 130, between 3 and 140, between 3 and 150, between 3 and 160, between 3 and 170, between 3 and 180, between 3 and 190, between 3 and 200, between 3 and 210, between 3 and 220, between 3 and 230, between 3 and 240, between 3 and 250, between 3 and 260, between 3 and 270, between 3 and 280, between 3 and 290, or between 3 and 300 somatic mutations are identified using a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus of a plurality of loci in genomic DNA derived from a solid tumor sample from the subject.
[0346] In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency of at least 0.1 percent for each respective somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency of at least 1 percent for each respective somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency of at least 2 percent for each respective somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency of at least 5 percent for each respective somatic mutation.
[0347] In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 0.1 percent but less than 60 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 0.1 percent but less than 50 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 0.1 percent but less than 40 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 0.1 percent but less than 30 percent for each somatic mutation.
[0348] In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 5 percent but less than 60 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 5 percent but less than 50 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 5 percent but less than 40 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 525 has a variant allele frequency greater than 5 percent but less than 30 percent for each somatic mutation.
[0349] Block 526. Referring to block 526, in some embodiments, the identifying of block 525 includes identifying a plurality of candidate somatic mutations by comparing each nucleic acid sequence in the first plurality of nucleic acid sequences to nucleic acid sequences in a third plurality of nucleic acid sequences obtained from a sequencing reaction of genomic DNA from non-cancerous tissue of the subject. That is, in some embodiments, the process for identifying somatic mutations includes filtering out mutations present in the germline of the subject, even if they are pathogenic in nature or likely to be pathogenic.
[0350] In some such embodiments, a dedicated normal sample is collected from the patient for simultaneous processing with the liquid biopsy sample. Generally, the normal sample is of non-cancerous tissue and can be collected using any of the tissue collection methods described herein. In some embodiments, oral cells collected from the inside of the patient's cheek are used as the normal sample. Oral cells can be collected by placing an absorbent material, such as a cotton swab, into the subject's mouth and rubbing it against the cheek for, for example, at least 15 seconds or at least 30 seconds. The swab is then removed from the patient's mouth and inserted into a tube so that the tip of the tube is immersed in a liquid that serves to extract the oral cells from the absorbent material. An example of an oral cell recovery and collection device is provided in U.S. Patent No. 9,138,205, the contents of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, oral swab DNA is used as a source of normal DNA in circulating hematologic malignancies.
[0351] Block 528. Referring to block 528, in some embodiments, the identifying of blocks 525 and / or 526 further includes filtering out one or more respective somatic mutations determined to have an off-center variant allele fraction in the first plurality of sequences among the somatic mutations identified using the first plurality of sequence reads.
[0352] For example, in some embodiments, somatic variants with very high or very low VAFs compared to the VAFs for the majority of identified somatic variants are excluded from analysis because they likely do not represent the true circulating tumor fraction. For example, somatic variants with VAFs significantly lower than the VAFs for the majority of somatic variants may result from small subclonal cell populations within the tumor. Figure 7A illustrates this. As further described in block 530 below and in Example 1, the identified somatic variants are fitted to a distribution (a normal distribution in the case of Figure 7A) based on their respective VAF values in the first plurality of sequence reads. Variants whose VAF values fall more than two standard deviations from the mean of this normal distribution are then excluded from the somatic variants analyzed using the following liquid biopsy data to calculate the circulating tumor fraction.
[0353] Block 530. Referring to block 530, in some embodiments, the filtering includes fitting a VAF for each respective somatic mutation obtained using the first plurality of sequences (after removing germline mutations) to a distribution, and filtering out candidate somatic mutations with a corresponding VAF outside the spread of the distribution.
[0354] In some embodiments, the quality of this distribution fit is assessed using a distribution quality methodology such as a Kolmogorov-Smirnov goodness-of-fit test, a chi-square test, an Anderson-Darling test, a Shapiro-Wilk test, a Lilliefors test, a Cramer-von Mises test, a Jarque-Bera test, a Kuiper V test, or a Watson U test. For disclosure of the Anderson-Darling, Kolmogorov-Smirnov, Cramer-von Mises, Kuiper V, and Watson U goodness-of-fit tests, see Jantschi and Bolboaca, 2017, "Performance of Shannon's Entropy Statistic in Assessment of Distribution Data," and references cited therein, which are incorporated herein by reference. See also Huber-Carol et al., (ed.), March 8, 2002, "Goodness-of-Fit Test and Model Validity (Statistics for Industry and Technology), Birkhauser, ISBN-10 0817642099, which is incorporated herein by reference.
[0355] In some embodiments, in order for a distribution to be used to remove outliers from the mutations considered in the analysis of somatic mutations from liquid biopsies in block 539, the quality of the fit must have a p-value greater than a threshold amount. In some embodiments, the threshold amount is 0.05, as in Example 1. In other words, if the distribution quality methodology assesses the goodness of fit to a distribution to be greater than p=0.05, the distribution can be used to identify outliers to the distribution (e.g., those greater than two standard deviations from the distribution's mean) and remove them from the somatic mutations used in block 539. If the distribution quality methodology assesses the goodness of fit to a distribution to be less than p=0.05, the distribution is not used to identify outliers to the distribution on the basis that the VAFs of somatic mutations derived from solid tumor samples from the subject do not map well to the distribution. In some embodiments, the threshold amount is 0.10, meaning that the distribution quality methodology must recognize that the VAFs of mutations derived from solid tumor samples from the subject fit the distribution with a p-value of at least 0.10. In some embodiments, the threshold is 0.15, 0.20, 0.25, 0.50 or greater.
[0356] Blocks 532-534. Referring to block 532, in some embodiments, the VAF values of somatic mutations from genomic DNA derived from a solid tumor sample from the subject are fitted to a distribution such as a normal distribution, a beta distribution, a beta prime distribution, a lognormal distribution, or a gamma distribution. Referring to block 534, in some embodiments, the VAF values of somatic mutations from genomic DNA derived from a solid tumor sample from the subject are fitted to a normal distribution.
[0357] In some embodiments, the type of distribution is predetermined.For example, in Example 1, it is determined that normal distribution is best suited to the majority of subjects analyzed in that example.Therefore, in some embodiments, normal distribution is always used.In some embodiments, several different types of distribution are evaluated, and the distribution with the best fit is selected.In some embodiments, any of the distribution types disclosed herein is used.
[0358] In some embodiments, the VAF values of somatic mutations are fitted to a distribution using maximum likelihood estimation, where the parameters of the distribution maximize the likelihood of the observed data (the VAF values of somatic mutations) under an assumed distribution (e.g., a normal distribution).
[0359] In some embodiments, the VAF values of somatic mutations are fitted to a distribution using a method of moments algorithm, in which sample moments (e.g., mean, variance) are mapped to theoretical moments of a selected distribution.
[0360] In some embodiments, the VAF values of somatic mutations are fitted to a distribution using Bayesian estimation. Bayesian estimation provides a probabilistic framework for estimating the parameters of a distribution by treating the VAF values of somatic mutations as random variables. Instead of finding a single point estimate for a parameter (as in MLE), Bayesian methods take into account observed data and prior knowledge (prior distribution) to generate a distribution of possible parameter values (posterior distribution).
[0361] In some embodiments, the VAF values of somatic mutations are fitted to a distribution using a least squares method, in which the sum of the squared differences between the observed VAF values and those predicted by a selected distribution is minimized. See, e.g., Levie, 2000, "Curve Fitting with Least Squares. Critical Reviews in Analytical Chemistry," 30(1), 59-74, which is incorporated herein by reference.
[0362] In some embodiments, the VAF values of somatic mutations are fitted to a distribution using a quantile matching algorithm (e.g., by matching quantiles of the observed VAF values with quantiles of a theoretical distribution), where the sum of squared differences between the observed VAF values and those predicted by a theoretical distribution (e.g., a normal distribution) is minimized. See Sgouropoulos et al., 2015, "Matching a Distribution by Matching Quantiles Estimation," Journal of the American Statistical Association, 110(510), pp. 742-759, which is incorporated herein by reference.
[0363] In some embodiments, the VAF values of somatic mutations are fitted to the distribution using the expectation-maximization (EM) algorithm. The EM algorithm is an iterative method for finding maximum likelihood estimates in the presence of latent (unobserved) variables. The EM algorithm alternates between two steps: an expectation (E-step), in which expected values of latent variables are estimated given current parameter estimates, and a maximization (M-step), in which a likelihood function for the parameters based on the estimated latent variables is maximized. See Koch, 2013, "Robust estimation by expectation maximization algorithm," Journal of Geodesy 87, pp. 107-116, which is incorporated herein by reference.
[0364] In some embodiments, the VAF values of somatic mutations are fitted to the distribution using a non-parametric method such as kernel density estimation. Non-parametric methods do not assume a specific form for the distribution of the VAF values of somatic mutations. One example is kernel density estimation (KDE), which is used to estimate the probability density function (PDF) of the VAF values of somatic mutations directly from the observed VAF values without assuming an underlying parametric distribution. See Chen et al., 2017, "A tutorial on kernel density estimation and recent advances," Biostatistics & Epidemiology, 1(1), 161-187, which is incorporated herein by reference.
[0365] Block 536. Referring to block 536, in some embodiments, the dispersion is a multiple of the standard deviation for the probability distribution. In some embodiments, the threshold for excluding somatic variants is a VAF that exceeds a threshold number of standard deviations from a representative value (e.g., the mean or median) of the set of VAFs for each respective candidate somatic mutation in the first plurality of sequences (of the solid tumor sequencing reaction). In some embodiments, the threshold number of standard deviations is at least 2 standard deviations. In some embodiments, the threshold number of standard deviations is at least 0.5, at least 0.75, at least 1, at least 1.25, at least 1.5, at least 1.75, at least 2, at least 2.5, at least 3, at least 4, or more standard deviations from a representative value (e.g., the mean or median) of the set of VAFs for each respective candidate somatic mutation in the first plurality of sequences. In some embodiments, the threshold number of standard deviations is 4 standard deviations or less from a representative value (e.g., mean or median) for the set of VAFs for each respective candidate somatic mutation of the plurality of candidate somatic mutations in the first plurality of sequences. In some embodiments, the threshold number of standard deviations is 5 or less, 4 or less, 3 or less, 2.5 or less, 2 or less, 1.75 or less, 1.5 or less, 1.25 or less, 1 or less, or less than a representative value (e.g., mean or median) for the set of VAFs for each respective candidate somatic mutation of the plurality of candidate somatic mutations in the first plurality of sequences.
[0366] In some embodiments, dispersion is a multiple of the standard deviation around a representative value of the distribution (eg, the mean, median, average, etc.).
[0367] In some embodiments, the dispersion is a multiple of the mean absolute deviation (MAD) around the center of the distribution.
[0368] In some embodiments, the dispersion is the interquartile range (IQR), range, coefficient of variation (CV) range, skewness range, kurtosis range, or Gini index (e.g., centered at the median value of the distribution) within the distribution.
[0369] Block 538. Referring to block 538, in some embodiments, the identifying of block 524 further includes excluding one or more respective candidate somatic mutations in the plurality of candidate somatic mutations having a nucleotide position that does not correspond to any probe in the first plurality of probes. That is, in embodiments in which different panel enrichment sequencing reactions are used to sequence the solid tumor biopsy genomic DNA and the liquid biopsy cfDNA, somatic variants identified from the solid tumor biopsy that are not within a region targeted by the enrichment panel for the liquid biopsy sequencing reaction are excluded from the analysis because the locus is not enriched in the liquid biopsy sequencing reaction.
[0370] Block 539. Referring to block 539, in some embodiments, a determination of a corresponding variant allele frequency (VAF) in the liquid biopsy sample is performed for each respective somatic mutation in the one or more somatic mutations, wherein the corresponding VAF is determined from the frequency of each somatic variant in the second plurality of nucleic acid sequences at the corresponding one or more nucleotide positions for each somatic variant in the second plurality of nucleic acid sequences, thereby determining a set of VAFs for the one or more somatic mutations in the liquid biopsy sample. In some embodiments, the one or more somatic mutations is at least one somatic mutation.
[0371] In some embodiments, a set of variant allele frequencies (VAFs) is formed, including a respective VAF for each respective somatic mutation in one or more somatic mutations (identified according to the previous blocks of FIG. 5 above), with VAFs having measurable values in a liquid biopsy sample determined from the frequency of each respective somatic mutation in the second plurality of nucleic acid sequences. In some embodiments, a solid assay is used to identify the set of somatic mutations according to block 525, and block 539 serves to determine the variant allele frequency of each of these somatic mutations in a corresponding liquid biopsy using the second plurality of nucleic acid sequences. Thus, in some embodiments, there is no minimum VAF for each respective somatic mutation identified according to block 539 for the set of VAFs and identified using the second plurality of sequence reads. In some alternative embodiments, each respective somatic mutation identified according to block 539 for the set of VAFs and identified using the second plurality of sequence reads has a variant allele frequency of at least 0.1 percent for each somatic mutation. In some embodiments, each respective somatic mutation identified according to block 539 for the set of VAFs and identified using the second plurality of sequence reads has a variant allele frequency of at least 5 percent for the respective somatic mutation.
[0372] In some embodiments, the one or more somatic mutations are at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, or more somatic mutations. In some embodiments, the one or more somatic mutations are or were 5,000 or fewer somatic variants, 2,500 or fewer somatic variants, 1,000 or fewer somatic variants, 500 or fewer somatic variants, or 250 or fewer somatic variants. In some embodiments, the one or more somatic variants are between 1 and 5,000 somatic variants. In some embodiments, the one or more somatic variants are between 1 and 2500, 1 and 1000, 1 and 500, 1 and 250, 2 and 5000, 2 and 2500, 2 and 1000, 2 and 500, 2 and 250, 5 and 5000, 5 and 2500, 5 and 1000, 5 and 500, or 5 and 250 somatic mutations.
[0373] Block 540. Referring to block 540, in some embodiments, an estimate of circulating tumor fraction for the subject is determined based on a set of VAFs. Advantageously, this set is tumor informative in that each of the somatic mutations contributing to the set was also found in the tumor sample. Advantageously, this set is also filtered by the distribution of VAFs of somatic mutations in the liquid biopsy to exclude outliers. In some embodiments, after such filtering, the one or more somatic mutations contributing to the set of VAFs are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more somatic mutations. In some embodiments, after such filtering, the one or more somatic mutations contributing to the set of VAFs are multiple somatic mutations.
[0374] Blocks 542-544. Referring to block 542, in some embodiments, the estimate of circulating tumor fraction is a representative value of a set of VAFs.
[0375] For example, in some embodiments, the ctFE is the arithmetic mean, weighted mean, midpoint, central hinge, adjusted mean, geometric mean, geometric median, winsorized mean, median, or mode of the VAFs for multiple somatic mutations. Referring to block 544, in some embodiments, the representative value is the median.
[0376] In some embodiments, the one or more somatic mutations is a single somatic mutation, and the ctFE is the VAF for the single somatic mutation.
[0377] Average VAF Method. In some embodiments, the estimate of circulating tumor fraction is determined using the average VAF method. In some embodiments, the average VAF method is calculated as follows:
number
number
[0378] VAF-based CFT estimation. In some embodiments, an estimate of circulating tumor fraction is determined using a VAF-based method. The VAF-based method for estimating circulating tumor fraction (CTF) utilizes the VAF (proportion of sequencing reads with mutations in a given genomic region) of each somatic mutation that contributes to a set of VAFs. The method assumes that the VAF of a somatic mutation observed in a cfDNA liquid biopsy sample is a function of both the tumor DNA fraction (CTF) in the tumor and the mutant allele frequency. According to VAF-based CFT estimation, the formula relating the VAF of a mutation in cfDNA to the CTF is as follows:
number
number
number
[0379] If there are multiple candidate somatic mutations contributing to the set of VAFs, an estimate of the CTF for each somatic mutation contributing to the set of VAFs can be calculated separately, and then all estimates can be averaged to obtain a final estimate of the CTF across all somatic mutations.
[0380] In some embodiments, the obtained estimates of circulating tumor fraction are used for further downstream analysis and biomarker detection (e.g., calculating variant allele fraction, variant calling, and / or identifying other indicators). In some embodiments, the obtained estimates of circulating tumor fraction are used as indicators for disease detection, diagnosis, and / or treatment. In some embodiments, the obtained estimates of circulating tumor fraction are included in clinical reports available to patients or clinicians. In some embodiments, the obtained estimates of circulating tumor fraction are used to select appropriate therapies and / or clinical trials for evaluation of treatment response.
[0381] Thus, in some embodiments, the method also includes generating a report about the subject (e.g., for use by a physician) that includes the circulating tumor fraction for the subject. In some embodiments, the report further includes a matched therapy (e.g., treatment and / or clinical trial) for the subject based on the reported circulating tumor fraction for the subject.
[0382] In some embodiments, the methods described herein include generating a clinical report 139-3 (e.g., a patient report), providing clinical support for personalized cancer therapy, and / or using information curated from the sequencing of the liquid biopsy sample described above. In some embodiments, the report is provided to the patient, physician, medical professional, or researcher in a digital copy (e.g., a JSON object, a PDF file, or an image on a website or portal), a hard copy (e.g., printed on paper or other tangible medium). For example, the report object, such as a JSON object, can be used for further processing and / or display. For example, information from the report object can be used to generate a clinical test report for return to the prescribing physician. In some embodiments, the report is presented as text, as audio (e.g., prerecorded or streaming), as an image, or in another format, and / or any combination thereof.
[0383] In some embodiments, the report includes information related to specific features of the patient's cancer, such as detected genetic variants, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities. In some embodiments, other features of the patient's sample and / or clinical record are also included in the report. For example, in some embodiments, the clinical report includes information regarding one or more of clinical variants, such as copy number variants (e.g., in potentially actionable genes CCNE1, CD274 (PD-L1), EGFR, ERBB2 (HER2), MET, MYC, BRCA1, and / or BRCA2), fusions, translocations, and / or rearrangements (e.g., in potentially actionable genes ALK, ROS1, RET, NTRK1, FGFR2, FGFR3, NTRK2, and / or NTRK3), pathogenic single nucleotide polymorphisms, insertion-deletions (e.g., somatic / tumor and / or germline / normal), therapeutic biomarkers, microsatellite instability status, and / or tumor mutational burden.
[0384] In some embodiments, the results are used to design patient biological cell line tests, such as tumor organoid experiments.For example, organoids can be genetically enginee...
Claims
1. 1. A method for determining an estimate of circulating tumor fraction for a subject, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, A) obtaining a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus in a plurality of loci of genomic DNA derived from a solid tumor sample from the subject; B) obtaining a second plurality of nucleic acid sequences for each cell-free DNA fragment in a plurality of cell-free DNA fragments obtained from the liquid biopsy sample from the first panel enrichment sequencing assay using a first plurality of probes, the second plurality of probes including a corresponding probe that hybridizes to each locus in the plurality of loci; C) identifying one or more somatic mutations in the first plurality of nucleic acid sequences, wherein each respective somatic mutation in the one or more somatic mutations is at one or more corresponding nucleotide positions at a corresponding locus in the plurality of one or more loci; D) forming a set of VAFs, the set including a VAF for each respective somatic mutation in the one or more somatic mutations, determined from the frequency of each somatic mutation in the second plurality of nucleic acid sequences; E) determining an estimate of the circulating tumor fraction for the subject based on the set of VAFs.
2. 2. The method of claim 1, wherein the first plurality of nucleic acid sequences is determined from a second panel enrichment sequencing reaction using a second plurality of probes that includes, for each respective locus in the plurality of loci, a corresponding probe in the second plurality of probes that hybridizes to the respective locus.
3. 3. The method of claim 2, wherein the plurality of loci are sequenced in the second panel enrichment sequencing reaction to an average sequencing depth of at least 100-fold.
4. 4. The method of claim 2 or claim 3, wherein the second plurality of probes enriches loci from at least 50 genes.
5. 14. The method of claim 2 or claim 3, wherein the second plurality of probes enriches loci from at least 50 genes in Table 1, Table 2, List 1, List 2, FIG. 14, or FIG.
15.
6. The method of any one of claims 1 to 5, wherein the first plurality of probes and the second plurality of probes are different.
7. 7. The method of claim 1, wherein the plurality of loci are sequenced in the first panel enrichment sequencing reaction to an average sequencing depth of at least 500-fold.
8. The method of any one of claims 1 to 8, wherein the first plurality of probes enriches loci from at least 50 genes.
9. The method of any one of claims 1 to 9, wherein the identities of the first plurality of probes are non-custom to the subject.
10. The method of any one of claims 1 to 10, wherein the solid tumor sample is taken prior to taking the liquid biopsy sample.
11. 11. The method of any one of claims 1 to 10, wherein the solid tumor sample and the liquid biopsy sample are taken within 6 months of each other.
12. The method of any one of claims 1 to 12, wherein the liquid biopsy sample is blood.
13. The method of any one of claims 1 to 12, wherein the liquid biopsy sample comprises blood, whole blood, peripheral blood, plasma, serum, or lymph of the subject.
14. 14. The method of any one of claims 1-13, wherein said identifying C) comprises identifying a plurality of candidate somatic mutations by comparing each nucleic acid sequence in the first plurality of nucleic acid sequences to nucleic acid sequences in a third plurality of nucleic acid sequences obtained from a sequencing reaction of genomic DNA from non-cancerous tissue of the subject.
15. 15. The method of claim 14, wherein said identifying step C) further comprises eliminating one or more respective candidate somatic mutations in said plurality of candidate somatic mutations determined to have an acentric variant allele fraction in said first plurality of sequences.
16. 16. The method of claim 15, wherein said eliminating comprises fitting a VAF for each respective candidate somatic mutation in said plurality of candidate somatic mutations in said first plurality of sequences to a distribution; and eliminating candidate somatic mutations having a corresponding VAF that is outside a spread for said distribution.
17. 17. The method of claim 16, wherein the distribution is a normal distribution, a beta distribution, a beta prime distribution, a lognormal distribution, or a gamma distribution.
18. 17. The method of claim 16, wherein the fitting comprises applying maximum likelihood estimation, method of moments, Bayesian inference, least squares, quantile matching, or expectation maximization to the VAF of each respective candidate somatic mutation in the plurality of candidate somatic mutations.
19. 18. The method of claim 16 or claim 17, wherein the distribution is a normal distribution.
20. 20. The method of any one of claims 16 to 19, wherein the dispersion is a multiple of the standard deviation about a representative value of the distribution.
21. 20. The method of any one of claims 16 to 19, wherein the dispersion is a multiple of the mean absolute deviation (MAD) about a representative value of the distribution.
22. 20. The method of any one of claims 16 to 19, wherein the dispersion is the interquartile range (IQR), range, coefficient of variation (CV) range, skewness range, kurtosis range, or Gini index within the distribution.
23. 16. The method of claim 15, wherein said eliminating comprises: determining a distribution of the VAF for each respective candidate somatic mutation in the plurality of candidate somatic mutations in the first plurality of sequences using a non-parametric approach; and eliminating candidate somatic mutations having a corresponding VAF that is outside a spread for the distribution.
24. 24. The method of claim 23, wherein the non-parametric technique is kernel density estimation.
25. 25. The method of any one of claims 14 to 24, wherein said identifying (C) further comprises eliminating one or more respective candidate somatic mutations in said plurality of candidate somatic mutations having a nucleotide position that does not correspond to any probe in said first plurality of probes.
26. 26. The method of any one of claims 1 to 25, wherein the estimate of the circulating tumor fraction is representative of the set of VAFs.
27. 27. The method of claim 26, wherein the representative value is a median value.
28. 26. The method of any one of claims 1 to 25, wherein the estimate of the circulating tumor fraction is determined from the set of VAFs using the mean VAF method.
29. 26. The method of any one of claims 1 to 25, wherein the estimate of the circulating tumor fraction is determined from the set of VAFs using a VAF-based CFT estimation.
30. The method comprises: The method of any one of claims 1 to 29, further comprising F) reporting said estimate of said circulating tumor fraction for said subject.
31. 31. The method of claim 30, wherein said reporting F) further comprises reporting a matched treatment recommendation for the subject in response to determining that the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
32. 32. The method of any one of claims 1-31, further comprising administering a cancer therapeutic agent to the subject when the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
33. 33. The method of any one of claims 1 to 32, further comprising altering a cancer therapeutic agent therapy regimen administered to the subject when the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
34. 31. The method of claim 30, wherein said reporting F) further comprises reporting a matched clinical trial recommendation for the subject in response to determining that the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
35. 35. The method of any one of claims 1-34, further comprising enrolling the subject in a clinical trial when the estimate of the circulating tumor fraction for the subject meets a clinical trial threshold.
36. The method of any one of claims 1 to 35, wherein the one or more somatic mutations comprise two or more somatic mutations.
37. The method of any one of claims 1 to 35, wherein the one or more somatic mutations comprise five or more somatic mutations.
38. The method of any one of claims 1 to 35, wherein the one or more somatic mutations comprise 10 or more somatic mutations.
39. The method of any one of claims 1 to 35, wherein the one or more somatic mutations comprise 20 or more somatic mutations.
40. 40. The method of any one of claims 1 to 39, wherein said estimate of circulating tumor fraction for a subject has a detection limit of less than 0.1%.
41. 1. A computer system comprising: one or more processors, and A computer system comprising: a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by the one or more processors, cause the processors to perform the method of any one of claims 1 to 40.
42. A non-transitory computer readable storage medium storing program code instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 40.
43. 1. A method for determining an estimate of circulating tumor fraction for a subject, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, A) obtaining a second plurality of nucleic acid sequences for each cell-free DNA fragment in a plurality of cell-free DNA fragments obtained from a liquid biopsy sample from a first panel enrichment sequencing assay using a first plurality of probes, the second plurality of probes including a corresponding probe that hybridizes to each respective locus in the plurality of loci; B) identifying a plurality of somatic mutations in the first plurality of nucleic acid sequences, wherein each respective somatic mutation in the plurality of somatic mutations is at corresponding one or more nucleotide positions at corresponding loci of the plurality of loci; C) removing from the plurality of somatic mutations each respective somatic mutation flagged as an outlier in the distribution of variant allele fractions determined for somatic mutations using a solid tumor sample from the subject; D) removing each respective somatic mutation not identified as having a variant allele proportion using the solid tumor sample from the plurality of somatic mutations; E) after removing C) and removing D), forming a set of VAFs including the respective VAFs for each respective somatic mutation in the plurality of somatic mutations; F) determining an estimate of the circulating tumor fraction for the subject based on the set of VAFs.
44. obtaining a first plurality of nucleic acid sequences comprising a corresponding nucleic acid sequence for each respective locus in a plurality of loci of genomic DNA from a solid tumor sample from the subject; identifying one or more somatic mutations in the first plurality of nucleic acid sequences, wherein each respective somatic mutation in the one or more somatic mutations is located at one or more nucleotide positions at corresponding loci in the plurality of one or more loci and is identified as having a variant allele proportion based on the first plurality of nucleic acid sequences.
44. The method of claim 43.
45. 45. The method of Claim 44, wherein the first plurality of nucleic acid sequences is determined from a second panel enrichment sequencing reaction using a second plurality of probes that includes, for each respective locus in the plurality of loci, a corresponding probe in the second plurality of probes that hybridizes to the respective locus.
46. 46. The method of claim 45, wherein the plurality of loci are sequenced in the second panel enrichment sequencing reaction to an average sequencing depth of at least 100-fold.
47. 47. The method of claim 45 or claim 46, wherein the second plurality of probes enriches loci from at least 50 genes.
48. 47. The method of claim 45 or claim 46, wherein the second plurality of probes enriches loci from at least 50 genes in Table 1, Table 2, List 1, List 2, Figure 14, or Figure 15.
49. 49. The method of any one of claims 45 to 48, wherein the first plurality of probes and the second plurality of probes are different.
50. 50. The method of any one of claims 43 to 49, wherein the plurality of loci are sequenced in the first panel enrichment sequencing reaction to an average sequencing depth of at least 500-fold.
51. 51. The method of any one of claims 43-50, wherein the first plurality of probes enriches loci from at least 50 genes.
52. 52. The method of any one of claims 43 to 51, wherein the identities of the first plurality of probes are non-custom to the subject.
53. 53. The method of any one of claims 43 to 52, wherein the solid tumor sample is taken prior to taking the liquid biopsy sample.
54. 54. The method of any one of claims 43 to 53, wherein the solid tumor sample and the liquid biopsy sample are taken within six months of each other.
55. 55. The method of any one of claims 43 to 54, wherein the liquid biopsy sample is blood.
56. 55. The method of any one of claims 43 to 54, wherein the liquid biopsy sample comprises blood, whole blood, peripheral blood, plasma, serum, or lymph of the subject.
57. 57. The method of any one of claims 43 to 56, wherein the distribution is a normal distribution, a beta distribution, a beta prime distribution, a lognormal distribution, or a gamma distribution.
58. 57. The method of any one of claims 43 to 56, wherein the distribution is a normal distribution.
59. 59. The method of any one of claims 43 to 58, wherein somatic mutations are flagged as outliers of the distribution based on their spread.
60. 60. The method of claim 59, wherein the dispersion is a multiple of the standard deviation about a representative value of the distribution.
61. 60. The method of claim 59, wherein the dispersion is a multiple of the mean absolute deviation (MAD) about a representative value of the distribution.
62. 60. The method of claim 59, wherein the dispersion is an interquartile range (IQR), range, coefficient of variation (CV) range, skewness range, kurtosis range, or Gini index within the distribution.
63. 63. The method of any one of claims 43 to 62, wherein the method further comprises filtering out one or more respective somatic mutations in the plurality of somatic mutations having a nucleotide position that does not correspond to any probe in the first plurality of probes.
64. 64. The method of any one of claims 43 to 63, wherein the estimate of the circulating tumor fraction is representative of the set of VAFs.
65. 65. The method of claim 64, wherein the representative value is a median value.
66. 64. The method of any one of claims 43 to 63, wherein the estimate of the circulating tumor fraction is determined from the set of VAFs using a mean VAF method.
67. 64. The method of any one of claims 43 to 63, wherein the estimate of the circulating tumor fraction is determined from the set of VAFs using a VAF-based CFT estimation.
68. The method comprises: G) reporting said estimate of said circulating tumor fraction for said subject.
69. 69. The method of claim 68, wherein said reporting G) further comprises reporting a matched treatment recommendation to the subject in response to determining that the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
70. 70. The method of any one of claims 43-69, further comprising administering a cancer therapeutic agent to the subject when the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
71. 71. The method of any one of claims 43-70, further comprising altering a cancer therapeutic agent therapy regimen administered to the subject when the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
72. 69. The method of claim 68, wherein said reporting g) further comprises reporting a matched clinical trial recommendation for the subject in response to determining that the estimate of the circulating tumor fraction for the subject meets a treatment threshold.
73. 73. The method of any one of claims 43-72, further comprising enrolling the subject in a clinical trial when the estimate of the circulating tumor fraction for the subject meets a clinical trial threshold.
74. 74. The method of any one of claims 43-73, wherein the plurality of somatic mutations comprises five or more somatic mutations prior to said removing C) and removing D).
75. 74. The method of any one of claims 43-73, wherein the plurality of somatic mutations comprises 10 or more somatic mutations prior to said eliminating C) and eliminating D).
76. 74. The method of any one of claims 43-73, wherein the one or more somatic mutations comprise 20 or more somatic mutations prior to said removing C) and removing D).
77. 77. The method of any one of claims 43 to 76, wherein said estimate of circulating tumor fraction for a subject has a detection limit of less than 0.1%.
78. 1. A computer system comprising: one or more processors, and 78. A computer system comprising: a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by said one or more processors, cause said processors to perform the method of any one of claims 43 to 77.
79. A non-transitory computer readable storage medium storing program code instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 43 to 77.