Methods and systems for determining blood tumor mutational burden in liquid biopsy assay

The method and system for determining bTMB in liquid biopsies address the inaccuracies of conventional assays by using panel enrichment sequencing and bioinformatics to estimate circulating tumor fraction, enhancing precision oncology treatment decisions.

JP2025124606APending Publication Date: 2025-08-26テンパスエーアイインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025019889
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-30
Filing Date
2025-02-10
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Conventional liquid biopsy assays struggle to accurately determine tumor mutational burden (TMB) due to challenges such as dilution of circulating tumor DNA and the use of small target panels, which limits the identification of actionable genomic alterations, especially in precision oncology.

Method used

A method and system for determining blood tumor mutation burden (bTMB) using panel enrichment sequencing reactions with a circulating tumor fraction threshold and filtering criteria to identify relevant mutations, incorporating bioinformatics pipelines for somatic mutation detection and circulating tumor fraction estimation.

Benefits of technology

Provides accurate bTMB determination from small targeted panel sequencing-based liquid biopsies, enabling improved clinical decision-making in precision oncology, including treatment recommendations and clinical trial enrollment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124606000001_ABST
    Figure 2025124606000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for determining a blood tumor mutational burden (bTMB) for a test subject.SOLUTION: Systems and methods for determining a blood tumor mutational burden (bTMB) for a test subject are provided, in which a plurality of nucleic acid sequences is obtained from a panel-enriched sequencing reaction. The plurality of nucleic acid sequences comprises a corresponding sequence for each cell-free DNA fragment in a plurality of cell-free DNA fragments obtained from a liquid biopsy sample of the test subject. Each respective cell-free DNA fragment in the plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences, the probe sequences being used in the panel-enriched sequencing reaction to enrich cell-free DNA fragments in the liquid biopsy sample. Using the panel-enriched sequencing reaction, it is determined that a circulating tumor fraction (ctFE) exceeds a threshold ctFE value. In response to this determination, the bTMB of the test subject is calculated from the panel-enriched sequencing reaction and reported.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 552,644, entitled "Methods and Systems for Determining Blood Tumor Mutational Burden in a Liquid Biopsy Assay," filed February 12, 2024, which is incorporated herein by reference.

[0002] The present disclosure relates generally to the use of cell-free DNA sequencing data to provide clinical support aimed at personalized cancer treatment. [Background technology]

[0003] Precision oncology is the practice of tailoring cancer therapy to an individual's unique genomic, epigenetic, and / or transcriptomic profile. Personalized cancer therapy builds on traditional treatment regimens used to treat cancer based solely on the overall classification of the cancer, for example, treating all breast cancer patients with one treatment and all lung cancer patients with a second treatment. This field arose from the frequent observation that different patients diagnosed with the same type of cancer, such as breast cancer, responded very differently to common treatment regimens. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that improve predictions of how individual cancers will respond to specific treatment modalities.

[0004] There is growing evidence that cancer patients who receive genetically guided therapy have better outcomes. For example, studies have shown that targeted therapy results in significant improvements in progression-free cancer survival. See, e.g., Radovich et al., Oncotarget, 7(35):56491-500 (2016). Similarly, a report from the IMPACT trial, a large (n=1307) retrospective analysis of consecutive, prospectively molecularly profiled patients with advanced cancer who participated in a large personalized medicine trial, showed that patients receiving tumor-biologically matched targeted therapy had a 16.2% response rate, compared with a 5.2% response rate for patients receiving unmatched therapy. Tsimberidou et al., ASCO 2018, Abstract LBA2553 (2018).

[0005] Indeed, therapies targeted to specific genomic alterations are already standard of care in some tumor types, as suggested by the National Comprehensive Cancer Network (NCCN) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer. In practice, the implementation of these targeted therapies requires determining the diagnostic marker status in each eligible cancer patient. While this can be achieved for some known mutations associated with treatment recommendations in the NCCN guidelines using individual assays or small next-generation sequencing (NGS) panels, the increasing number of actionable genomic alterations and the increasing complexity of diagnostic classifiers necessitates a more comprehensive assessment of each patient's cancer genome, epigenome, and / or transcriptome.

[0006] For example, evidence suggests that the use of combination therapies, in which each component is matched to actionable genomic alterations, holds the greatest potential for treating individual cancers. To date, retrospective studies of cancer patients treated with one or more therapeutic regimens have revealed that patients receiving therapies matched to a higher proportion of genomic alterations experienced a higher frequency of stable disease (e.g., longer time to recurrence), longer time to treatment failure, and better overall survival. (Wheeler et al., Cancer Res., 76:3690-701 (2016)) Therefore, comprehensive assessment of each cancer patient's genome, epigenome, and / or transcriptome should maximize the benefits offered by precision oncology by facilitating more fine-tuned combination therapies, off-label use of novel drugs, and / or tissue-independent immunotherapies. See, e.g., Schwaederle et al., J Clin Oncol., 33(32):3817-25 (2015); Schwaederle et al., JAMA Oncol., 2(11):1452-59 (2016); and Wheler et al., Cancer Res., 76(13):3690-701 (2016). Furthermore, the use of comprehensive next-generation sequencing analysis of cancer genomes facilitates better access and larger patient pools for clinical trial enrollment. Coyne et al., Curr. Probl. Cancer, 41(3):182-93 (2017); and Markman, Oncology, 31(3):158,168.

[0007] To address the need for more comprehensive characterization of an individual's cancer genome, the use of large-scale NGS genomic analysis is increasing. See, for example, Fernandes et al., Clinics, 72(10):588-94. Recent studies have shown that 30-40% of patients undergoing large-scale NGS genomic analysis subsequently receive clinical care based on the assay results, which is limited by, at least, the identification of actionable genomic alterations, the availability of drugs to treat those alterations, and the subject's clinical condition. Ross et al., JAMA Oncol., 1(1):40-49 (2015); Ross et al., Arch. Pathol. Lab Med., 139:642-49 (2015); Hirshfield et al., Oncologist, 21(11):1315-25 (2016); and Groisberg et al., Oncotarget, 8:39254-67 (2017).

[0008] However, these large-scale NGS genomic analyses are traditionally performed on solid tumor samples. For example, each of the studies referenced in the above paragraph performed NGS analysis of FFPE tumor blocks from patients. Solid tissue biopsies represent a well-known and proven methodology that provides a high degree of accuracy and therefore remain the gold standard for diagnosis and identification of predictive biomarkers. Nevertheless, the use of solid tissue materials for large-scale NGS genomic analyses of cancer has significant limitations. For example, tumor biopsies are subject to sampling bias caused by spatial and / or temporal genetic heterogeneity, e.g., between two regions of a single tumor and / or between different cancer tissues (between a primary tumor site and a metastatic tumor site, or between two different primary tumor sites). Such inter- or intratumor heterogeneity can lead to overlooking subclonal mutations or newly emerging mutations when using localized tissue biopsies, and sampling bias can worsen over time as subclonal populations further evolve and / or shift in dominance.

[0009] In addition, obtaining solid tissue biopsies often requires invasive surgical procedures, for example, when the primary tumor site is located in an internal organ. These procedures are expensive, time-consuming, and can involve significant risks to the patient, for example, when the patient's health is poor and they cannot tolerate invasive medical procedures, and / or the tumor is located in a particularly sensitive or inoperable location, such as the brain or heart. Furthermore, the amount of tissue that can be procured depends on multiple factors, including tumor location, tumor size, patient vulnerability, and the risk of biopsy-related comorbidities, such as bleeding and infection. For example, a recent study reported that tissue samples in the majority of patients with advanced non-small cell lung cancer are limited to small biopsies, and in up to 31% of patients, no samples can be obtained at all. (Ilie and Hofman, Transl. Lung Cancer Res., 5(4):420-23 (2016)). Even when tissue biopsies are obtained, the sample may be too insufficient for comprehensive testing.

[0010] Furthermore, methods of tissue collection, preservation (e.g., formalin fixation), and / or archiving of tissue biopsies can result in sample degradation and variable-quality DNA. This, in turn, leads to inaccuracies in downstream assays and analyses, including next-generation sequencing (NGS) for biomarker identification. Ilie and Hofman, Transl Lung Cancer Res., 5(4):420-23 (2016).

[0011] Additionally, the invasiveness of the biopsy procedure, the time and expense associated with obtaining the sample, and the impaired state of cancer patients undergoing therapy make repeated testing of cancer tissue impractical, if not impossible. As a result, solid tissue biopsy analysis is not suitable for many monitoring schemes that benefit cancer patients, such as disease progression analysis, treatment efficacy assessment, disease recurrence monitoring, and other techniques that require data from several time points.

[0012] Cell-free DNA (cfDNA) has been identified in various body fluids, such as serum, plasma, and urine. Chan et al., 2003, Ann. Clin. Biochem., 40(Pt 2):122-30. This cfDNA is derived from all types of necrotic or apoptotic cells, including germline cells, hematopoietic cells, and pathological (e.g., cancer) cells. Advantageously, genomic alterations in cancer tissues can be identified from cfDNA isolated from cancer patients. See, for example, Stroun et al., 1989, Oncology, 46(5):318-22; Goessl et al., 2000, Cancer Res., 60(21):5941-45; and Frenel et al., 2015, Clin. Cancer Res. 21(20):4586-96. Thus, one approach to overcoming the problems presented by the use of solid tissue biopsies described above is to analyze cell-free nucleic acids (e.g., cfDNA) and / or nucleic acids in circulating tumor cells present in biological fluids, for example, via liquid biopsy.

[0013] Specifically, liquid biopsies offer several advantages over traditional solid tissue biopsy analysis. For example, because bodily fluids can be collected in a minimally or non-invasive manner, sample collection is simpler, faster, safer, and less expensive than solid tumor biopsies. Such methods require only small sample volumes (e.g., 10 mL or less of whole blood per biopsy), reducing the discomfort and risk of complications experienced by patients during traditional tissue biopsies. In fact, liquid biopsy samples can be collected with limited or no assistance from a medical professional and can be performed in almost any location. Furthermore, liquid biopsy samples can be collected from any patient, regardless of the location of their cancer, their overall health, and any previous biopsy collections. This enables the analysis of cancer genomes in patients for whom solid tumor samples cannot be easily and / or safely obtained. In addition, because cell-free DNA in bodily fluids originates from many different types of tissue in a patient, the genomic alterations present in the pool of cell-free DNA represent various distinct clonal subpopulations of the target cancer tissue, facilitating a more comprehensive analysis of the target cancer genome than is possible from one or more sections of a single solid tumor sample.

[0014] Liquid biopsies also enable serial genetic testing prior to cancer detection, during early stages of cancer progression, throughout the course of treatment, and during remission, for example, to monitor for disease recurrence. The ability to perform serial testing via non-invasive liquid biopsies throughout the course of disease may prove beneficial to many patients, for example, by monitoring patient response to therapy, the emergence of new actionable genomic alterations, and / or drug resistance changes. This type of information allows medical professionals to more quickly adjust and update treatment regimens, for example, facilitating more timely intervention in the event of disease progression. See, e.g., Ilie and Hofman, 2016, Transl. Lung Cancer Res. 5(4):420-23.

[0015] While liquid biopsies are a promising tool for improving outcomes using precision oncology, there are significant challenges inherent in using cell-free DNA to assess a subject's cancer genome. For example, there is a highly variable signal-to-noise ratio from one liquid biopsy sample to the next. This occurs because cfDNA originates from a variety of different cells, both healthy and diseased, in a subject. Depending on the stage and type of cancer in any particular subject, the proportion of cfDNA fragments derived from cancer cells (the "tumor fraction" or "ctDNA fraction" of a sample / subject) can range from nearly 0% to well over 50%. Other factors, including tumor type and mutational profile, can also affect the amount of DNA released from cancer tissue. For example, cfDNA clearance through the liver and kidneys is affected by a variety of factors, including renal dysfunction or other tissue-damaging factors (e.g., chemotherapy, surgery, and / or radiation therapy).

[0016] The information disclosed in this Background section is merely for the purpose of enhancing understanding of the general background of the present invention and should not be construed as an acknowledgment or any form of suggestion that this information forms prior art already known to those skilled in the art. Summary of the Invention

[0017] Given the above background, there is a need in the art for improved methods and systems for supporting clinical decisions in precision oncology using liquid biopsy assays. In particular, there is a need in the art for improved methods and systems for determining tumor mutational burden (TMB) based on liquid biopsy assays. Tumor mutational burden (TMB) is defined as the total number of somatic mutations per defined region of the tumor genome and is a pan-tumor biomarker for immune checkpoint inhibitor (ICI) response in patients with advanced cancer. While not intending to be limited to any particular theory, the potential clinical benefit of this biomarker is based on the hypothesis that highly mutated tumors produce high-quality neoantigens that increase T-cell reactivity and therefore may also improve response to immune checkpoint inhibitor treatment. For example, Aggrawal et al., 2023, "Assessment of Tumor Mutational Burden and Outcomes in Patients With Diverse Advanced Cancers Treated With Immunotherapy," JAMA Network Open 6(5):e2311181, found that subjects with high TMB of non-small cell lung cancer (NSCLC), bladder cancer, melanoma, and colorectal cancer each had a higher one-year survival rate than subjects with the same cancer types who had low TMB. Accordingly, one aspect of the present disclosure provides a method for determining liquid biopsy tumor mutation burden (lTMB) for a subject.

[0018] In some such embodiments, a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors obtains a plurality of nucleic acid sequences from a panel enrichment sequencing reaction, the plurality of nucleic acid sequences including sequences corresponding to each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from a test subject. Each cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences, and the probe sequences are used to enrich the cell-free DNA fragments in the liquid biopsy sample in the panel enrichment sequencing reaction. In some embodiments, the plurality of probe sequences map to 150 or fewer genes in the human genome.

[0019] In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of at least 1,000X.

[0020] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches between 50 genes and 150 genes.

[0021] In some embodiments, the multiple probe sequences used to enrich cell-free DNA fragments in a liquid biopsy sample in a panel enrichment sequencing reaction collectively map to 25 to 150 different genes in the human reference genome.

[0022] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in Table 1.

[0023] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in List 1.

[0024] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in List 2.

[0025] In some embodiments, the liquid biopsy sample is a blood sample, hi some embodiments, the liquid biopsy sample is blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal material, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject.

[0026] In some embodiments, the liquid biopsy sample is a cell-free sample, such as a cell-free blood sample.

[0027] Using the panel enrichment sequencing reactions, the circulating tumor fraction (ctFE) is determined to be above a threshold ctFE value. In some such embodiments, the threshold ctFE value is 0.01.

[0028] In response to determining that the ctFE is above the threshold, the lTMB for the subject is calculated from the panel enrichment sequencing reactions.

[0029] In some embodiments, the calculation comprises determining the number of genetic variants present in the plurality of nucleic acid sequences.

[0030] In some embodiments, the number of genetic variants present in the plurality of nucleic acid sequences is the number of unique genetic variants present in the plurality of nucleic acid sequences that meet one or more qualification criteria in the set of qualification criteria.

[0031] In some such embodiments, an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant be a missense variant, a combination of a missense variant and a splice region variant, a frameshift variant, a stop-loss variant, a splice acceptor variant, an in-frame insertion variant, an in-frame deletion variant, a combination of a frameshift variant and a splice region variant, a disruptive in-frame insertion variant, or a disruptive in-frame deletion variant.

[0032] In some embodiments, an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant has a variant allele frequency greater than 0.005 (0.5%) in the liquid biopsy sample.

[0033] In some embodiments, an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant have a variant allele frequency less than 1.0 (100%) in the liquid biopsy sample.

[0034] In some embodiments, a qualifying criterion in the set of qualifying criteria is the requirement that each genetic variant have a variant allele frequency (VAF) in a liquid biopsy sample that is one of: (i) greater than 0.01 (1%) and less than 0.4 (40%); (ii) greater than 0.6 (60%) and less than 0.9 (90%); or (iii) greater than 0.4 (40%) and less than 0.60 (60%), provided that |VAF-ctFE| / ctFE<1; or (iv) greater than 0.9 (90%), provided that |VAF-ctFE| / ctFE<1.

[0035] In some embodiments, an eligibility criterion in the set of eligibility criteria is selection by a medical professional of a genetic variant that is present in the plurality of nucleic acid sequences.

[0036] In some such embodiments, the calculation further comprises normalizing the number of genetic variants present in the plurality of nucleic acid sequences by the coverage of the plurality of probe sequences. In some such embodiments, the coverage is between 0.1 megabases and 0.4 megabases. In some such embodiments, the coverage is between 0.15 megabases and 0.3 megabases.

[0037] In some embodiments, lTMB is reported for a subject.

[0038] In some embodiments, the reporting further includes reporting a matched treatment recommendation for the subject in response to determining that the lTMB meets the treatment threshold.

[0039] In some embodiments, the reporting includes comparing the subject's lTMB to a severity threshold and reporting a qualitative status of either high lTMB (lTBM-H) or low lTMB (lTMB-L) based on the comparison.

[0040] In some embodiments, a subject's lTMB is reported only if the lTMB meets a reporting threshold.

[0041] In some embodiments, the immunotherapeutic agent is administered to a subject only if the subject's lTMB meets the therapeutic threshold.

[0042] In some embodiments, the lTMB reports matched clinical trial recommendations to the subject in response to determining that the clinical trial threshold is met.

[0043] In some embodiments, the method further comprises enrolling the subject only if the subject's lTMB meets the clinical trial threshold.

[0044] In some embodiments, the method further includes using the lTMB to identify a concordant lTMB based on a defined correlation between (i) detection of somatic mutations in cell-free DNA from liquid biopsy samples from the cohort of training subjects and (ii) detection of somatic mutations in genomic DNA from solid tumor biopsy samples from the cohort of training subjects. Further, reporting further includes reporting the concordant lTMB.

[0045] In some embodiments, the subject's electronic medical record is updated to include the subject's lTMB.

[0046] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief explanation of the drawings]

[0047] [Figure 1A] 1A, 1B, 1C, and 1D collectively illustrate a block diagram of an exemplary computing device for determining a liquid biopsy tumor mutation burden (lTMB) of a subject, according to some embodiments of the present disclosure. [Figure 1B] Same as above. [Figure 1C] Same as above. [Figure 1D] Same as above. [Figure 2A] FIG. 2A illustrates an exemplary workflow for generating a clinical report based on information generated from the analysis of one or more patient samples, according to some embodiments of the present disclosure. [Figure 2B] FIG. 2B illustrates an example of a distributed diagnostic environment for collecting and evaluating patient data for precision oncology purposes, according to some embodiments of the present disclosure. [Figure 3]3 provides an exemplary flowchart of processes and features related to the collection and analysis of liquid biopsy samples for use in precision oncology, according to some embodiments of the present disclosure. In the flowchart, dashed boxes represent optional elements. [Figure 4A] Figures 4A, 4B, 4C, 4D, 4E, 4F1, 4F2, 4F3, 4G1, 4G2, and 4G3 collectively illustrate an example of a bioinformatics pipeline for precision oncology. Figure 4A provides a general flowchart of processes and features in a bioinformatics pipeline according to some embodiments of the present disclosure. Figure 4B provides an overview of a bioinformatics pipeline run using either a liquid biopsy sample alone or a liquid biopsy sample and a matched normal sample. Figure 4C illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier according to some embodiments of the present disclosure. Figure 4D illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. Figure 4E shows the process for obtaining tumor and normal BAM alignment files. Figure 4F1 provides a flowchart of a method for validating copy number variation according to some embodiments of the present disclosure, where dashed boxes represent optional method portions. Figure 4F2 provides a flowchart of a method for validating somatic sequence variants in a subject with cancer disease according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. Figures 4G1, 4G2, and 4G3 illustrate methods of variant detection according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. Figure 4F3 provides an overview of a method for estimating circulating tumor fraction in a liquid biopsy sample based on targeted panel sequencing data according to some embodiments of the present disclosure, with dashed boxes representing optional portions of the method. [Figure 4B] Same as above. [Figure 4C] Same as above. [Figure 4D] Same as above. [Figure 4E] Same as above. [Figure 4F1] Same as above. [Figure 4F2] Same as above. [Figure 4F3] Same as above. [Figure 4G1] Same as above. [Figure 4G2] Same as above. [Figure 4G3] Same as above. [Figure 5A] Figures 5A, 5B, 5C, 5D, 5E, 5F, 5G and 5H collectively show the results of a comparison of circulating tumor fraction estimates (ctFE) to variant allele fractions (VAF) using the Off-Target Tumor Estimation Routine (OTTER) method, according to various embodiments of the present disclosure. [Figure 5B] Same as above. [Figure 5C] Same as above. [Figure 5D] Same as above. [Figure 5E] Same as above. [Figure 5F] Same as above. [Figure 5G] Same as above. [Figure 5H] Same as above. [Figure 6A] 6A and 6B collectively show the results of assessing ctFE and mutation status according to cancer type, according to various embodiments of the present disclosure. [Figure 6B] Same as above. [Figure 7A] 7A, 7B, and 7C collectively show results assessing the association between ctFE and progressive disease status, according to various embodiments of the present disclosure. [Figure 7B] Same as above. [Figure 7C] Same as above. [Figure 8A] 8A, 8B, and 8C collectively show results comparing recent clinical response results with ctFE, according to various embodiments of the present disclosure. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 9A]FIGS. 9A, 9B, 9C, 9D, and 9E collectively provide a flowchart of processes and features for determining tumor mutation burden from a liquid biopsy sample, according to some embodiments of the present disclosure, where dashed boxes represent optional method portions. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 9E] Same as above. [Figure 10A] Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10J, 10K, 10L and 10M collectively show examples of nucleic acids targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 10D] Same as above. [Figure 10E] Same as above. [Figure 10F] Same as above. [Figure 10G] Same as above. [Figure 10H] Same as above. [Figure 10I] Same as above. [Figure 10J] Same as above. [Figure 10K] Same as above. [Figure 10L] Same as above. [Figure 10M] Ibid. Like reference numerals refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION OF THE INVENTION

[0048] Introduction

[0049] Tumor mutation burden (TMB) is a measure of the total number of somatic mutations in cancer and has been used as a biomarker to identify cancer patients who are likely to respond well to immune checkpoint blockade (ICB) therapy. However, conventional liquid biopsy assays, also referred to herein as blood TMB (bTMB) or liquid biopsy TMB (lTMB), have not provided accurate TMB determination. Many reasons have been suggested for the particular difficulty in determining bTMB, including the dilution of circulating tumor DNA (ctDNA) and the use of small target panels for sequencing cell-free DNA (cfDNA) enriched in genomic regions that are mutation hotspots. Therefore, there is a need in the art for improved methods for determining bTMB from liquid biopsy assays, especially those that use small target panels (e.g., targeting less than 250 genes and / or targeting less than 1 Mb) to enrich genomic regions of interest.

[0050] Advantageously, the present invention discloses methods and systems that provide accurate bTMB determination from small targeted panel sequencing-based liquid biopsy assays. In some embodiments, the improved performance of the methods and systems described herein is based, at least in part, on the incorporation of a circulating tumor fraction threshold from which bTMB is calculated. In some embodiments, the improved performance of the methods and systems described herein is based, at least in part, on the incorporation of filtering criteria to identify which mutations detected in the liquid biopsy assay should be used for bTMB determination.

[0051] In one aspect, the present disclosure provides a method for carrying out such a method, and corresponding system and non-transitory computer-readable medium (CRM), for determining the blood tumor mutation burden (bTMB) of a subject.In some embodiments, the method, system and CRM described herein are incorporated into the framework of liquid biopsy assay.For example, Figure 2A shows an exemplary scheme of the combined wet lab and bioinformatics process for liquid biopsy analysis.

[0052] In some embodiments, the methods, systems, and CRMs described herein acquire data from one or more different parts of a bioinformatics pipeline for liquid biopsy assays. In some embodiments, the methods and systems described herein use somatic mutations identified in a separate workstream of a liquid biopsy assay. For example, Figures 4F2 and 4G illustrate methods 400-2 and 450, respectively, which use a dynamic variant count threshold to identify somatic sequence variants from cfDNA sequencing data. In some embodiments, the somatic variants so identified are used to determine bTMB using the methods, systems, and CRMs described herein. For details of methods for identifying somatic mutations from cfDNA samples, see, for example, U.S. Patent No. 11,475,981, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0053] Similarly, Figure 4F3 shows a method 400-3 for estimating circulating tumor fraction of a liquid biopsy sample by fitting a simulated tumor fraction to copy number status generated from cfDNA sequencing data. In some embodiments, such circulating tumor fraction estimates (ctFE) are used in the methods, systems, and CRMs described herein, for example, in determining the threshold for determining bTMB. For details of methods for estimating circulating tumor fraction of a liquid biopsy sample, see, for example, U.S. Patent No. 11,211,147, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0054] In one aspect, the present disclosure provides a method, as well as systems and CRMs for performing all or part of such a method, for obtaining a plurality of nucleic acid sequences from a panel enrichment sequencing reaction, the plurality of nucleic acid sequences including corresponding sequences for each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from a subject, such as sequencing data 122-1 described herein with respect to sequencing 312 performed during the wet-lab portion 204 of the exemplary liquid biopsy scheme shown in FIG.

[0055] Each cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in the plurality of probe sequences, and the probe sequences are used to enrich cell-free DNA fragments in liquid biopsy samples in a panel enrichment sequencing reaction. Examples of gene sets targeted by such probes are described, for example, in Table 1 provided herein and with reference to Figure 6 in PCT Patent Application Publication WO2023 / 164713. This document is incorporated herein by reference in its entirety for all purposes. In some embodiments, the plurality of probe sequences are mapped to 150 or fewer genes in the human genome.

[0056] The method also includes using a panel enrichment sequencing reaction to determine that the circulating tumor fraction (ctFE) is above a threshold ctFE value. In some embodiments, a circulating tumor fraction estimate prepared according to a method described herein, e.g., see method 400-3 shown in FIG. 4F3, is used in thresholding the bTMB determination. Similarly, in some embodiments, a ctFE prepared according to the method described in U.S. Provisional Patent Application No. 18 / 930,786, filed October 31, 2023, entitled "ESTIMATION OF CIRCULATING TUMOR FRACTION USING OFF-TARGET READS OF TARGETED-PANEL SEQUENCING," is used in thresholding the bTMB determination. The document is incorporated by reference in its entirety and for all purposes.

[0057] The method also includes calculating the lTMB of the subject from the panel enrichment sequencing reaction in response to determining that the ctFE is above a threshold. In some embodiments, the calculation includes determining the number of genetic variants present in the plurality of nucleic acid sequences. In some embodiments, the genetic variants are verified somatic variants determined according to the verification method described herein, for example, with respect to methods 400-2 or 450 illustrated in Figures 4F2 and 4G, respectively. In some embodiments, the genetic variants are somatic variants verified using other methods, for example, as described in U.S. Patent No. 11,475,981. In some embodiments, the calculation further includes normalizing the number of genetic variants present in the plurality of nucleic acid sequences by the coverage of the plurality of probe sequences, for example, the panel described in Table 1 herein or the panel described in PCT Patent Application Publication WO2023 / 164713. In some embodiments, the method also includes reporting the bTMB of the subject.

[0058] As described herein, in some embodiments, the methods described herein include one or more data collection steps in addition to data analysis and downstream steps. For example, as described below, such as with reference to Figures 2 and 3, in some embodiments, the methods include collection of a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). Similarly, as described below, such as with reference to Figures 2 and 3, in some embodiments, the methods include extraction of DNA from a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). Similarly, as described below, such as with reference to Figures 2 and 3, in some embodiments, the methods include nucleic acid sequencing of DNA from a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject).

[0059] However, in other embodiments, the methods described herein begin with nucleic acid sequencing results, such as raw or collapsed sequence reads of DNA from a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or non-cancerous samples from the subject), from which statistics (e.g., bin-level sequence ratios, segment-level sequence ratios, and segment-level variance measures) necessary for focal CNV validation can be determined. For example, in some embodiments, sequence data 122 of patient 121 is accessed and / or downloaded by system 100 via network 105.

[0060] Identifying actionable genomic alterations in a patient's cancer genome is a challenging and computationally intensive task. For example, determining various predictive indices useful for precision oncology, such as variant allele ratios, copy number variation, tumor mutation burden, and microsatellite instability status, requires the analysis of hundreds of millions to hundreds of billions of sequenced nucleic acid bases. A typical bioinformatics pipeline established for this purpose involves at least five stages of analysis: assessing the quality of raw next-generation sequencing data, generating corrected nucleic acid fragment sequences and aligning these sequences to a reference genome, detecting structural variants in the aligned sequence data, annotating identified variants, and visualizing the data. See Wadapurkar and Vyas, Informatics in Medicine Unlocked, 11:75-82 (2018), the contents of which are incorporated herein by reference in their entirety and for all purposes. Each of these steps is computationally intensive.

[0061] For example, the time and space computational complexity of the overall algorithms of simple global sequence alignment and local pairwise sequence alignment is quadratic (e.g., a quadratic problem) and increases rapidly with the size (n and m) of the nucleic acid sequences being compared. Specifically, the time and space complexity of these sequence alignment algorithms can be estimated as O(mn). Where O is the upper limit of the asymmetry growth rate of the algorithm, n is the number of bases in the first nucleic acid sequence, and m is the number of bases in the second nucleic acid sequence. See Baichoo and Ouzounis, BioSystems, 156-157:72-85 (2017). The contents of this document are incorporated herein by reference in their entirety for all purposes. Considering that the human genome contains more than 3 billion bases, these alignment algorithms are computationally intensive, especially when used to analyze next-generation sequencing (NGS) data, which can generate more than 3 billion sequence reads per reaction.

[0062] This is especially true when performed in liquid biopsy assays, because liquid biopsy samples contain a complex mixture of short DNA fragments originating from many different germline (e.g., healthy) and diseased tissues (e.g., cancerous tissue). Therefore, when the cellular origin of a sequence read is unknown, sequence signals originating from cancerous cells, which may comprise multiple subclonal populations, must be computationally deconvolved from signals originating from germline and hematopoietic lineages to provide relevant information about the subject's cancer. Thus, in addition to the computationally demanding process required to align sequence reads to the human genome, there is the computational problem of determining whether a particular aberrant signal, e.g., one or more sequence reads corresponding to a genomic alteration, is (i) an artifact and (ii) originates from a cancerous source in the subject. This becomes even more challenging during the early stages of cancer, when treatment is likely most effective, when only small amounts of ctDNA are diluted by germline and hematopoietic DNA.

[0063] In addition to the computational complexity involved in aligning sequencing data to a human reference genome, the method involves dividing multiple aligned sequence reads into "bins" (e.g., regions of a predetermined span of base pairs corresponding to the reference genome), determining the copy ratio of each bin by calculating the differential read depth between the experimental sample and the reference sample, and grouping a subset of adjacent bins that share a copy ratio into segments. By grouping bins into segments, each chromosome is divided into regions of equal copy number, minimizing noise in the data. These methods essentially implement change-point or edge detection algorithms, which are either time-limited or computationally intensive. For example, in some embodiments, segmentation is performed using circular binary segmentation, which calculates statistics for each genomic position. The statistic comprises the likelihood ratio of the null hypothesis (no change in copy ratio at each position) to the alternative (one change in copy ratio at each position); if the statistic is greater than a predetermined distribution threshold, the null hypothesis is rejected. In particular, in circular binary segmentation, chromosomes are assumed to be circular; therefore, the calculation is performed recursively for each position (e.g., each bin) along the circumference to identify all change points across the entire chromosome length. Furthermore, for each position (e.g., bin) under investigation, a permutation approach is used to generate a reference distribution, where the copy ratios of multiple bins are randomized (typically 10,000 times). In some embodiments utilizing bins approximately 100-150 bases long, spanning the billions of bases of the human reference genome, the number of permutations required to implement this recursive method makes it computationally intensive. See, e.g., Olshen et al., Biostatistics 5, 4, 557-572 (2004), doi:10.1093 / biostatistics / kxh008, which is incorporated herein by reference in its entirety.

[0064] Definition.

[0065] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, a human (e.g., a male human, a female human, a fetus, a pregnant woman, a child, or the like), a non-human mammal, or a non-human animal. Any human or non-human animal can serve as a subject, including, but not limited to, a mammal, a reptile, a bird, an amphibian, a fish, a ungulate, a ruminant, a bovine (e.g., a cow), an equine (e.g., a horse), a caprine and ovine (e.g., a sheep, a goat), a porcine (e.g., a pig), a camelid (e.g., a camel, a llama, an alpaca), a monkey, an ape (e.g., a gorilla, a chimpanzee), a ursidae (e.g., a bear), a fowl, a dog, a cat, a mouse, a rat, a fish, a dolphin, a whale, and a shark. In some embodiments, the subject is a male or female (e.g., a man, a woman, or a child) of any age.

[0066] As used herein, "control," "control sample," "reference," "reference sample," "normal," and "normal sample" describe a sample derived from non-diseased tissue. In some embodiments, such a sample is derived from a subject that does not have a particular condition (e.g., cancer). In other embodiments, such a sample is an internal control, e.g., from a subject that may or may not have a particular disease (e.g., cancer), but is derived from the subject's healthy tissue. For example, if a liquid or solid tumor sample is obtained from a subject with cancer, an internal control sample can be obtained from the subject's healthy tissue, e.g., a white blood cell sample from a subject that does not have a blood cancer, or a solid germline tissue sample from the subject. Thus, a reference sample can be obtained from the subject or from a database, e.g., from a second subject that does not have a particular disease (e.g., cancer).

[0067] As used herein, the terms "cancer," "cancer tissue," or "tumor" refer to an abnormal mass of tissue, including both a solid mass (e.g., in the case of solid tumors) or a fluid mass (e.g., in the case of blood cancers), in which the growth of the mass exceeds and is out of step with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are well differentiated, have characteristically slower growth than malignant tumors, and may remain localized to the site of origin. In addition, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors may be poorly differentiated (dysplastic) and have characteristically rapid growth accompanied by progressive infiltration, invasion, and destruction of surrounding tissue. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is out of step with that of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained from or derived from a tumor in a subject, as described herein.

[0068] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, bile duct cancer, adrenal cancer, neural carcinoma, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, non-small cell lung cancer, thymus cancer, tumors, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, serous ovarian cancer, HR+ breast cancer, serous uterine cancer, endometrial cancer, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.

[0069] As used herein, the term "cancer state" or "cancer condition" refers to characteristics of a cancer patient's condition, such as diagnostic status, cancer type, cancer location, cancer primary origin, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject, such as age, sex, weight, race, personal habits (e.g., smoking, alcohol use, diet), other relevant medical conditions (e.g., high blood pressure, dry skin, other diseases), current medications, allergies, relevant medical history, current side effects of cancer treatments and other medications, etc., are used to further describe the subject's cancer state or condition.

[0070] As used herein, the term "liquid biopsy" sample refers to a liquid sample obtained from a subject that contains cell-free DNA. Examples of liquid biopsy samples include, but are not limited to, a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal material, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. In some embodiments, the liquid biopsy sample is a cell-free sample, e.g., a cell-free blood sample. In some embodiments, the liquid biopsy sample is obtained from a subject with cancer. In some embodiments, the liquid biopsy sample is collected from a subject with an unknown cancerous condition, e.g., for use in determining the subject's cancerous condition. Similarly, in some embodiments, the liquid biopsy is collected from a subject with a non-cancerous disorder, e.g., cardiovascular disease. In some embodiments, the liquid biopsy is collected from a subject with an unknown non-cancerous disorder, e.g., for use in determining the subject's non-cancerous disorder condition.

[0071] As used herein, the terms "cell-free DNA" and "cfDNA" refer interchangeably to DNA fragments circulating within a subject's body (e.g., bloodstream) and derived from one or more healthy cells and / or one or more cancer cells. These DNA molecules are found outside of cells in bodily fluids, such as a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal matter, saliva, sweat, perspiration, tears, pleural fluid, pericardial fluid, or peritoneal fluid, and are considered to be fragments of genomic DNA excreted from healthy cells and / or cancer cells, for example, during apoptosis and dissolution of the cell envelope.

[0072] As used herein, the term "locus" refers to a location (e.g., site) within a genome, for example, on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position on a particular chromosome within a genome. In some embodiments, a locus refers to a group of nucleotide positions within a genome. In some cases, a locus is defined by a mutation (e.g., a substitution, insertion, deletion, inversion, or translocation) of consecutive nucleotides within a cancer genome. In some cases, a locus is defined by a gene, a subgenic structure (e.g., a regulatory element, exon, intron, or a combination thereof), or a defined span of a chromosome. Because normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus located on an autosome, for example, one copy on the maternal autosome and one copy on the paternal autosome.

[0073] As used herein, the term "allele" refers to a specific sequence of one or more nucleotides at a chromosomal locus. In haploid organisms, a subject has one allele at every chromosomal locus. In diploid organisms, a subject has two alleles at every chromosomal locus.

[0074] As used herein, the term "base pair" or "bp" refers to a unit consisting of two nucleic acid bases joined together by hydrogen bonds. Generally, the size of an organism's genome is measured in base pairs because DNA is typically double-stranded. However, some viruses have single-stranded DNA or RNA genomes.

[0075] As used herein, the terms "genomic alteration," "mutation," and "variant" refer to detectable changes in the genetic material of one or more cells. Genomic alterations, mutations, or variants can refer to various types of changes in a cell's genetic material, including changes in the primary genomic sequence at single or multiple nucleotide positions, e.g., single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels (e.g., nucleotide insertions or deletions), DNA rearrangements (e.g., inversions or translocations of portions of a chromosome or multiple chromosomes), copy number variations of loci (e.g., exons, genes, or large spans of chromosomes) (CNVs), partial or complete changes in the ploidy of a cell, and changes in the epigenetic information of a genome, such as altered DNA methylation patterns. In some embodiments, a mutation is a change in a cell's genetic information relative to one or more "normal" alleles found in a particular reference genome or population of a species of interest. For example, mutations can be found in both a subject's germline cells (e.g., non-cancerous "normal" cells) and in a subject's abnormal cells (e.g., pre-cancerous or cancerous cells). Thus, mutations in a subject's germline (e.g., found in substantially all "normal" cells in a subject) are identified relative to the reference genome for the subject's species. However, many loci in a species' reference genome are associated with several variant alleles that are significantly represented in the subject's population and are not associated with a pathological condition, e.g., such that they are not considered "mutations." In contrast, in some embodiments, mutations in a subject's cancer cells can be identified relative to either the subject's reference genome or the subject's own germline genome. In certain cases, identifying both types of variants can be beneficial. For example, in some cases, mutations present in both the subject's cancer genome and the subject's germline can be beneficial for precision oncology when the mutations are so-called "driver mutations" that contribute to the initiation and / or development of cancer. However, in other cases, mutations present in both the subject's cancer genome and the subject's germline are not beneficial for precision oncology when, for example, the mutations are so-called "passenger mutations" that do not contribute to the initiation and / or development of cancer.Similarly, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is beneficial for precision oncology, e.g., if the mutation is a driver mutation and / or the mutation facilitates a therapeutic approach, e.g., by distinguishing cancer cells from normal cells in a therapeutically actionable manner. However, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is not beneficial for precision oncology, e.g., if the mutation is a passenger mutation and / or the mutation fails to distinguish cancer cells from germline cells in a therapeutically actionable manner.

[0076] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the dominant allele represented at that chromosomal locus within a population of a species (e.g., a "wild-type" sequence) or an allele predefined within a reference genome of the species.

[0077] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is not the dominant allele represented at that locus within a population of a species (e.g., not a "wild-type" sequence) or is not a predefined allele within a reference sequence construct for the species (e.g., a reference genome or set of reference genomes). In some cases, sequence isoforms found within a population of a species that do not affect the change in the protein encoded by the genome or that result in amino acid substitutions that do not substantially affect the function of the encoded protein are not variant alleles.

[0078] As used herein, the terms "variant allele fraction," "VAF," "allele fraction," or "AF" refer to the number of times a variant or mutant allele is observed (e.g., the number of reads supporting the candidate variant allele) divided by the total number of times the position was sequenced (e.g., the total number of reads covering the candidate locus).

[0079] As used herein, the term "germline variant" refers to genetic variants inherited from maternal and paternal DNA. Germline variants can be determined through an adapted tumor-normal calling pipeline.

[0080] As used herein, the term "somatic variant" refers to a variant that arises as a result of dysregulation, e.g., mutation, of a cellular process associated with a neoplastic cell. Somatic variants can be detected by subtraction from a matched normal sample.

[0081] As used herein, the term "single nucleotide variant" or "SNV" refers to the substitution of one nucleotide with a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y may be represented as "X>Y." For example, a cytosine to thymine SNV may be represented as "C>T."

[0082] As used herein, the terms "insertions and deletions" or "indels" refer to variants that result from the gain or loss of DNA base pairs within the analyzed region.

[0083] As used herein, the term "copy number variation" or "CNV" refers to a process for detecting large structural changes in the genome associated with tumor aneuploidy and other dysregulation of repair systems. These processes are used to detect large insertions or deletions of entire genomic regions. CNV is defined as a structural insertion or deletion whose size is greater than a certain base pair ("bp"), such as 500 bp.

[0084] As used herein, the term "gene fusion" refers to the product of large-scale chromosomal abnormalities that result in the production of chimeric proteins. These expressed products may be non-functional or may be highly overactive or underactive. This may lead to deleterious effects in cancer, such as a hyperproliferative or anti-apoptotic phenotype.

[0085] As used herein, the term "loss of heterozygosity" refers to the loss of one copy of a segment (e.g., including part or all of one or more genes) of the genome of a diploid subject (e.g., a human) in a tissue of a subject, e.g., cancer tissue, or the loss of one copy of a sequence encoding a functional gene product in the genome of a diploid subject. As used herein, when referring to a metric representing loss of heterozygosity across a subject's genome, the loss of heterozygosity is caused by the loss of one copy of various segments in the subject's genome. Genome-wide loss of heterozygosity can be estimated without sequencing the entire genome of a subject, and methods for such estimation based on gene panel targeted sequencing methodologies have been described in the art. Thus, in some embodiments, a metric representing loss of heterozygosity across the genome of a subject's tissue is expressed as a single value, e.g., a percentage or proportion of the genome. In some cases, a tumor may be composed of various subclonal populations, each of which may have different degrees of loss of heterozygosity across their respective genomes. Thus, in some embodiments, genome-wide loss of heterozygosity of a cancer tissue refers to the average loss of heterozygosity across a heterogeneous tumor population. As used herein, when referring to the metric of loss of heterozygosity for a particular gene, for example, a DNA repair protein (e.g., BRCA1 or BRCA2), such as a protein involved in the homologous DNA recombination pathway, loss of heterozygosity refers to the complete or partial loss of one copy of the gene encoding the protein in the genome of a tissue, and / or a mutation in one copy of the gene that prevents translation of the full-length gene product, such as a frameshift or truncation mutation (generating a premature stop codon in the gene) in the gene of interest. In some cases, tumors are composed of various subclonal populations, each of which may have a different mutation status in the gene of interest. Thus, in some embodiments, loss of heterozygosity for a particular gene of interest is represented by the average loss of heterozygosity for the gene across all sequenced subclonal populations of the cancer tissue.In other embodiments, loss of heterozygosity for a particular gene of interest is represented by the number of unique occurrences of loss of heterozygosity in the gene of interest across all sequenced subclonal populations of the cancer tissue (e.g., the number of unique frameshift and / or truncating mutations in the gene identified in the sequencing data).

[0086] As used herein, the term "microsatellite" refers to a short, repeated sequence of DNA. The minimal nucleotide repeating unit of a microsatellite is referred to as a "repeated unit" or "repeat unit." In some embodiments, the stability of a microsatellite locus is assessed by comparing a metric of the distribution of the number of repeat units at the microsatellite locus to a reference number or distribution.

[0087] As used herein, the term "microsatellite instability" or "MSI" refers to a genetic hypermutability state associated with various cancers resulting from impaired DNA mismatch repair (MMR) in a subject. Among other phenotypes, MSI causes changes in the size of microsatellite loci, e.g., changes in the number of repeat units at a microsatellite locus, during DNA replication. Thus, the size of microsatellite repeats is altered in MSI cancers compared to the size of corresponding microsatellite repeats in the germline of the cancer subject. The term "microsatellite instability-high" or "MSI-H" refers to a cancer (e.g., tumor) state with a significant MMR defect that results in microsatellite loci having lengths significantly different from those of corresponding microsatellite loci in normal cells of the same individual. The term "microsatellite stability" or "MSS" refers to a cancer (e.g., tumor) state without a significant MMR defect, such that there is no significant difference between the lengths of microsatellite loci in cancer cells and those of corresponding microsatellite loci in normal (e.g., non-cancerous) cells of the same individual. The term "microsatellite indeterminate" or "MSE" refers to a cancer (e.g., tumor) state with an intermediate microsatellite length phenotype that cannot be clearly classified as MSI-H or MSS based on the statistical cutoffs used to define these two categories.

[0088] As used herein, the term "gene product" refers to a specific genomic locus, e.g., an RNA (e.g., mRNA or miRNA) or protein molecule transcribed or translated from a specific gene. A genomic locus can be identified using the gene name, chromosomal location, or any other genetic mapping metric.

[0089] As used herein, the terms "expression level," "abundance level," or simply "abundance" refer to the amount of a gene product (an RNA species, e.g., mRNA or miRNA, or a protein molecule) transcribed or translated by a cell, or the average amount of a gene product transcribed or translated across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g., a particular gene. However, in some embodiments, the expression level may refer to the amount of a particular isoform of an mRNA or protein corresponding to a particular gene that gives rise to multiple mRNA or protein isoforms. A genomic locus can be identified using gene name, chromosomal location, or any other genetic mapping metric.

[0090] As used herein, the term "ratio" refers to any comparison of a first metric X, or a first mathematical transform thereof X' (e.g., a measurement of the number of units of a genomic sequence in a first one or more biological samples, or a first mathematical transform thereof), to another metric Y, or a second mathematical transform thereof Y' (e.g., the number of units of each genomic sequence in a second one or more biological samples, or a second mathematical transform thereof), such as X / Y, Y / X, log N (X / Y), log N (Y / X), X' / Y, Y / X', log N (X' / Y) or log N (Y / X'), X / Y', Y' / X, log N (X / Y'), log N (Y' / X), X' / Y', Y' / X', log N (X' / Y') or log N (Y' / X'), where N is any real number greater than 1, and exemplary mathematical transformations of X and Y include, but are not limited to, raising X or Y to the Zth power, multiplying X or Y by a constant Q, where Z and Q are any real numbers, and / or taking the M-based logarithm of X and / or Y, where M is a real number greater than 1. In one non-limiting example, X is multiplied by squaring X (X2 ) before the ratio calculation, and Y is calculated by raising Y to the power of 3.2 (Y 3.2 ) and the ratio of X and Y is calculated as log2(X' / Y').

[0091] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound, e.g., a gene product (RNA species, e.g., mRNA or miRNA, or a protein molecule) or a nucleic acid fragment with a particular property (e.g., aligning to a particular locus or encompassing a particular allele), measured in a sample to a second amount of the compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a species of a compound to the total amount of the compound in the same sample. For example, it is the ratio of the amount of mRNA transcripts encoding a particular gene in the sample (e.g., aligning to a particular region of the exome) to the total amount of mRNA transcripts in the sample. In other embodiments, relative abundance refers to the ratio of the amount of a compound or species of a compound in a first sample to the amount of the compound of that species of compound in a second sample. For example, it is the ratio of the normalized amount of mRNA transcripts encoding a particular gene in the first sample to the normalized amount of mRNA transcripts encoding the particular gene in the second sample and / or reference sample.

[0092] As used herein, the terms "sequencing," "determining a sequence," and the like refer to any biochemical process that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.

[0093] As used herein, the term "gene sequence" refers to a record of the series of nucleotides present in a subject's RNA or DNA as determined by sequencing nucleic acid from the subject.

[0094] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any nucleic acid sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (a "single-end read") or from both ends of a nucleic acid fragment (e.g., a paired-end read, a double-end read). The length of a sequence read is often related to a particular sequencing technology. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., an average, median, or mean length of about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads are an average, median, or mean length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, Iore® sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina® parallel sequencing can provide sequence reads that are less variable, e.g., the majority of sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a series of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, for example, hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0095] As used herein, the term "read segment" refers to any form of nucleotide sequence, including raw sequence reads obtained directly from nucleic acid sequencing technology or sequences derived therefrom, e.g., aligned sequence reads, corrected sequence reads, or spliced ​​sequence reads.

[0096] As used herein, the term "number of reads" refers to the total number of nucleic acid reads generated, which may or may not equal the number of nucleic acid molecules generated during a nucleic acid sequencing reaction.

[0097] As used herein, the terms "read depth," "sequencing depth," or "depth" may refer to the total number of unique nucleic acid fragments encompassing a particular locus or region of a subject's genome that are sequenced in a particular sequencing reaction. Sequencing depth can be expressed as "Y-fold," e.g., 50-fold, 100-fold, etc., where "Y" refers to the number of unique nucleic acid fragments encompassing a particular locus that are sequenced in a sequencing reaction. In such cases, Y is necessarily an integer, since it represents the actual sequencing depth of the particular locus. Alternatively, read depth, sequencing depth, or depth may refer to a measure of central tendency (e.g., the mean or mode) of the number of unique nucleic acid fragments encompassing one of multiple loci or regions of a subject's genome that are sequenced in a particular sequencing reaction. For example, in some embodiments, sequencing depth refers to the average depth of all loci across a chromosome, targeted sequencing panel, exome, or entire genome. In such cases, Y may be expressed as a fraction or decimal to refer to the average coverage across multiple loci. When an average depth is listed, the actual depth of any particular locus may differ from the overall listed depth. A metric can be determined that provides a range of sequencing depths that encompasses a defined percentage of the total number of loci. For example, a range of sequencing depths that encompasses 90%, 95%, or 99% of the loci. As will be understood by those skilled in the art, various sequencing technologies provide various sequencing depths. For example, low-pass whole genome sequencing can refer to a technology that provides a sequencing depth of less than 5x, less than 4x, less than 3x, or less than 2x, for example, about 0.5x to about 3x.

[0098] As used herein, the term "sequencing breadth" refers to a particular reference exome (e.g., a human reference exome), a particular reference genome (e.g., a human reference genome), or what proportion of a portion of an exome or genome has been analyzed. Sequencing breadth can be expressed as a fraction, decimal, or percentage and is generally calculated as (number of loci analyzed / total number of loci in the reference exome or reference genome). The denominator of the fraction can be the repeat-masked genome, and thus 100% can correspond to the entire reference genome minus the masked portion. A repeat-masked exome or genome can refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads align with unmasked portions of the exome or genome). In some embodiments, any portion of the exome or genome can be masked, and thus sequencing breadth can be assessed for any desired portion of the reference exome or genome. In some embodiments, "wide-area sequencing" refers to sequencing / analysis of at least 0.1% of the exome or genome.

[0099] As used herein, the terms "sequence ratio" and "coverage ratio" refer interchangeably to any measurement of the number of units of a genome sequence in a first one or more biological samples (e.g., a test sample and / or a tumor sample) compared to the number of units of the respective genome sequence in a second one or more biological samples (e.g., a reference sample and / or a control sample). In some embodiments, the sequence ratio is a copy ratio, a log2-transformed copy ratio (e.g., a log2 copy ratio), a coverage ratio, a base ratio, an allele ratio (e.g., a variant allele ratio), and / or a tumor ploidy. In some embodiments, the sequence ratio is a log N is the conversion copy ratio, where N is any real number greater than 1.

[0100] As used herein, the term "sequencing probe" refers to a molecule that binds to a nucleic acid with an affinity based on the predicted nucleotide sequence of the RNA or DNA present at that locus.

[0101] As used herein, the term "target panel" or "target gene panel" refers to a combination of probes for sequencing (e.g., by next-generation sequencing) nucleic acids present in a biological sample from a subject (e.g., a tumor sample, a liquid biopsy sample, a germline tissue sample, a white blood cell sample, or a tumor or tissue organoid sample) selected to map to one or more loci of interest on one or more chromosomes. An exemplary set of genes that can be analyzed using a targeted panel and are useful for precision oncology, e.g., via solid or liquid biopsy assays, is listed in Table 1. Another exemplary set of genes that can be analyzed using a targeted panel and are useful for precision oncology, e.g., via solid or liquid biopsy assays, is listed in Table 2. In some embodiments, in addition to loci useful for precision oncology, the targeted panel includes one or more probes for sequencing one or more of loci associated with different disease states, loci used for internal control purposes, or loci from pathogenic organisms (e.g., oncogenic pathogens).

[0102] As used herein, the term "reference exome" refers to any sequenced or otherwise characterized exome of any tissue from any organism or pathogen, whether partial or complete, that can be used to reference an identified sequence from a subject. Typically, the reference exome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference exomes for human subjects, as well as many other organisms, are provided in the online genome browser hosted by the National Center for Biotechnology Information ("NCBI"). "Exome" refers to the complete transcriptional profile of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference exome is often an assembled or partially assembled exome sequence from an individual or multiple individuals. In some embodiments, a reference exome is an assembled or partially assembled exome sequence from one or more human individuals. A reference exome can be considered a representative example of a species' set of expressed genes. In some embodiments, a reference exome includes sequences assigned to chromosomes.

[0103] As used herein, the term "reference genome" refers to any sequenced or otherwise characterized genome of any organism or pathogen, whether partial or complete, that can be used to reference identified sequences from a subject. Typically, a reference genome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference genomes used for human subjects, as well as many other organisms, are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz ("UCSC"). "Genome" refers to the complete genetic information of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of a species' set of genes. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38). For haploid genomes, only one nucleotide may be present at each locus. For diploid genomes, heterozygous loci may be identified, and each heterozygous locus may have two alleles, with either allele being able to match for alignment to the locus.

[0104] As used herein, the term "bioinformatics pipeline" refers to a series of processing steps used to determine characteristics of a subject's genome or exome based on sequencing data of the subject's genome or exome. A bioinformatics pipeline can be used to determine characteristics of a subject's germline genome or exome and / or a subject's cancer genome or exome. In some embodiments, the pipeline extracts information related to genomic alterations in a subject's cancer genome, which is useful for guiding clinical decisions for precision oncology from sequencing results of a subject-derived biological sample, such as a tumor sample, a liquid biopsy sample, or a reference normal sample. Certain processing steps in bioinformatics can be "connected," meaning that the results of a first respective processing step are useful and / or essential for the execution of a second downstream processing step. For example, in some embodiments, a bioinformatics pipeline includes a first respective processing step for identifying genomic alterations unique to a subject's cancer genome, and a second respective processing step for determining a metric useful for precision oncology, such as tumor mutation burden, using the amount and / or identity of the identified genomic alterations. In some embodiments, the bioinformatics pipeline includes a reporting stage that generates a report of relevant and / or actionable information identified by upstream stages of the pipeline, which may or may not further include recommendations to support clinical therapy decisions.

[0105] As used herein, the term "limit of detection" or "LOD" refers to the minimum amount of a feature that can be identified with a certain level of confidence. Thus, the level of detection can be used to describe the amount of a substance that must be present for a particular assay to reliably detect it. The level of detection can also be used to describe the level of support required for an algorithm to reliably identify genomic alterations based on sequencing data. For example, the minimum number of unique sequence reads required to support the identification of sequence variants such as SNVs.

[0106] As used herein, the terms "BAM file" or "binary file containing an alignment map" refer to a file that stores sequencing data aligned to a reference sequence (e.g., a reference genome or exome). In some embodiments, a BAM file is a compressed binary version of a SAM (sequence alignment map) file that contains, for each of a plurality of unique sequence reads, an identifier for the sequence read, information about the nucleotide sequence, information about the alignment of the sequence to the reference sequence, and optionally metrics about the quality of the sequence read and / or the quality of the sequence alignment. While a BAM file generally relates to a file having a particular format, for brevity, unless otherwise specified, it is used herein to simply refer to a file of any format that contains information about sequence alignments.

[0107] As used herein, the term "measure of central tendency" refers to the central or representative value of a distribution of values. Non-limiting examples of measures of central tendency include the arithmetic mean, weighted mean, median, central area, central hinge, trimean, geometric mean, geometric median, Winsorized mean, median, and mode of a distribution of values.

[0108] As used herein, the term "positive predictive value" or "PPV" refers to the likelihood that a variant will be correctly called, given that the variant was called by the assay. PPV can be expressed as (number of true positives) / (number of false positives + number of true positives).

[0109] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, e.g., a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining copy number variations of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to one of skill in the art can be used to detect any of the nucleic acid characteristics described herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of a mutation in a nucleic acid at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). Assays or methods can have a particular sensitivity and / or specificity, and their relative utility as diagnostic tools can be measured using ROC-AUC statistics.

[0110] As used herein, the term "classification" may refer to any number or other letter associated with a particular characteristic of a sample. For example, in some embodiments, the term "classification" may refer to the type of cancer in a subject, the stage of cancer in a subject, the prognosis of cancer in a subject, tumor burden, the presence of tumor metastasis in a subject, and the like. Classifications may be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" may refer to a predetermined number used in a calculation. For example, a cutoff size may refer to a size above which fragments are excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms may be used in either of these contexts.

[0111] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.

[0112] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.

[0113] As used herein, "actionable genomic alteration" or "actionable variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor proportion) that is known or believed to be associated with a therapeutic course of action that is more likely to result in a positive outcome in cancer patients with an actionable variant than in similarly located cancer patients without an actionable variant. For example, administration of an EGFR inhibitor (e.g., afatinib, erlotinib, gefitinib) is more effective in treating non-small cell lung cancer in patients with an EGFR mutation in exons 19 / 21 than in patients without an EGFR mutation in exons 19 / 21. Thus, an EGFR mutation in exons 19 / 21 is an actionable variant. In some cases, an actionable variant is associated with improved treatment outcomes only in one specific cancer type or a group of specific cancer types. In other cases, actionable variants are associated with improved treatment outcomes in virtually all cancer types.

[0114] As used herein, "variant of uncertain significance" or "VUS" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) whose impact on disease development / progression is unknown.

[0115] As used herein, "benign variant" or "likely benign variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed not to contribute to disease development / progression.

[0116] As used herein, "pathogenic variant" or "likely pathogenic variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed to contribute to disease development / progression.

[0117] As used herein, an "effective amount" or "therapeutically effective amount" is an amount sufficient to affect beneficial or desired clinical results during treatment. An effective amount can be administered to a subject in one or more doses. In terms of treatment, an effective amount is an amount sufficient to palliate, improve, stabilize, reverse, or slow the progression of a disease, or otherwise reduce the pathological consequences of a disease. An effective amount is generally determined by a physician on a case-by-case basis and is within the skill of one of ordinary skill in the art. Several factors are typically considered when determining an appropriate dosage to achieve an effective amount. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.

[0118] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. As used herein, the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or the claims, such terms are intended to be as inclusive as the term "comprising."

[0119] As used herein, the term "if" may be interpreted to mean "when" or "upon," "in response to detecting," or "in response to determining," depending on the context. Similarly, the phrase "when determined" or "when [a described condition or event] is detected" may be interpreted to mean "upon determining," or "in response to determining," or "upon detection [of [a described condition or event]," or "in response to detecting [a described condition or event]," depending on the context.

[0120] Additionally, while terms such as "first," "second," and the like may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both subjects, but are not the same subject. Furthermore, the terms "subject," "user," and "patient" are used interchangeably herein.

[0121] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure, including example systems, methods, techniques, instruction sequences, and computing machine program products embodying illustrative embodiments. However, the following illustrative discussion is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. Features described herein are not limited by the illustrated ordering of acts or events, as some acts may occur in different orders and / or contemporaneously with other acts or events.

[0122] The embodiments provided herein are chosen and described to best explain the principles and their practical applications, thereby enabling those skilled in the art to best utilize the various embodiments with various modifications suited to the particular uses contemplated. In some instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. In other instances, it will be apparent to one skilled in the art that the present disclosure may be practiced without one or more of the specific details.

[0123] It will be understood that in the development of any such actual implementation, numerous implementation-specific decisions will be made to achieve the designer's specific goals, such as adherence to use case and business-related constraints, and that these specific goals will vary from implementation to implementation and from designer to designer. It will further be understood that such a design effort may be complex and time-consuming, but would nevertheless be a routine undertaking for one of ordinary skill in the art having the benefit of this disclosure.

[0124] 1 illustrates an exemplary system embodiment.

[0125] Having provided an overview of some aspects of the present disclosure and some definitions used herein, details of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay will now be described in conjunction with Figures 1A, 1B, 1C, and 1D, which collectively illustrate the topology of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay, according to some embodiments of the present disclosure. Advantageously, the exemplary system shown in Figures 1A, 1B, 1C, and 1D improves upon conventional methods by determining accurate liquid biopsy tumor mutation burden (lTMB), providing clinical support for personalized cancer therapy.

[0126] 1A is a block diagram illustrating a system according to some embodiments. Device 100 in some embodiments includes one or more processing units (CPUs) 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including, for example, a display 108 and / or input 110 (e.g., a mouse, touchpad, keyboard, etc.), non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. One or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes a CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU 102. Persistent memory 112 and the non-volatile memory devices within non-persistent memory 112 comprise non-transitory computer-readable storage media. In some embodiments, non-persistent memory 111, or alternatively, non-transitory computer-readable storage media, sometimes in conjunction with persistent memory 112, store the following programs, modules, and data structures, or a subset thereof: an operating system 116, including procedures for handling various basic system services and for performing hardware-dependent tasks; a network communication module (or instructions) 118 for connecting the system 100 with other devices and / or communication networks 105; a test patient data store 120 for storing one or more sets of features from a patient (e.g., subject); a bioinformatics module 140 for processing the sequencing data and extracting features from the sequencing data, e.g., from a liquid biopsy sequencing assay; a feature analysis module 160 for assessing patient features, such as genomic alterations, complex genomic features, and clinical features; and A reporting module 180 for generating and transmitting reports that provide clinical support for personalized cancer therapy.

[0127] While FIGS. 1A, 1B, 1C, and 1D depict "system 100," the figures are intended as a functional description of various features that may be present in a computer system, rather than as a structural schematic of the embodiments described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art. Furthermore, while FIG. 1 depicts certain data and modules in non-persistent memory 111, some or all of these data and modules may be in persistent memory 112. For example, in various embodiments, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for implementing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various embodiments.

[0128] In some implementations, non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of system 100 that is addressable by system 100 such that system 100 can retrieve all or a portion of such data when needed.

[0129] 1A, system 100 is depicted as a single computer containing all of the functionality for providing clinical support for personalized cancer therapy. However, while a single machine is illustrated, the term "system" should also be construed to include any collection of machines that individually or jointly execute a set (or sets) of instructions for implementing any one or more of the methodologies discussed herein.

[0130] For example, in some embodiments, system 100 includes one or more computers. In some embodiments, functionality for providing clinical support for personalized cancer therapy is spread across any number of networked computers and / or resident on each of several networked computers and / or hosted on one or more virtual machines at remote locations accessible over communications network 105. For example, different portions of the various modules and data stores illustrated in Figures 1A, 1B, 1C1, 1D1, 1C2, 1D2, 1E2, 1F2, 1C3, and 1D3 may be stored and / or executed on various instances of processing devices and / or processing servers / databases (e.g., processing devices 224, 234, 244, and 254, processing server 262, and database 264) within distributed diagnostic environment 210 illustrated in Figure 2B.

[0131] The system may operate in the capacity of a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment. The system may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, server, network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine.

[0132] In another embodiment, the system comprises a virtual machine including modules for executing instructions for performing any one or more of the methodologies disclosed herein. In computing, a virtual machine (VM) is an emulation of a computer system that is based on a computer architecture and provides the functionality of a physical computer. Some such embodiments may involve dedicated hardware, software, or a combination of hardware and software.

[0133] Those skilled in the art will appreciate that any of a wide variety of different computer topologies may be used in an application, and that all such topologies are within the scope of the present disclosure.

[0134] Subject patient data storage unit (120)

[0135] Referring to FIG. 1B , in some embodiments, the system (e.g., system 100) includes a patient data store 120 that stores data for patients 121-1 through 121-M (e.g., cancer patients or patients being tested for cancer), including one or more of sequencing data 122, feature data 125, and clinical assessments 139. These data are used and / or generated by various processes stored in bioinformatics module 140 and feature analysis module 160 of system 100 to generate reports that ultimately provide clinical support for personalized cancer therapy for the patients. While the feature range of patient data 121 across all patients may be informationally dense, an individual patient's feature set may be sparse across the collective feature range of all features across all patients. That is, data stored for one patient may include a different set of features than data stored for another patient. Furthermore, although illustrated as a single data structure in FIG. 1B , different sets of patient data may be stored in different databases or modules spread across one or more system memories.

[0136] In some embodiments, sequencing data 122 from one or more sequencing reactions 122-i, including multiple sequence reads 123-i-1 through 123-iK, is stored in a test patient data store 120. The data store may contain different sets of sequencing data from a single subject corresponding to different samples from the patient, e.g., tumor samples, liquid biopsy samples, tumor organoids derived from the patient's tumor, and / or normal samples, and / or samples acquired at different times, e.g., while monitoring the progression, regression, remission, and / or recurrence of cancer in the subject. The sequence reads may be in any suitable file format, e.g., BCL, FASTA, FASTQ, etc. In some embodiments, the sequencing data 122 is accessed by a sequencing data processing module 141, which performs various preprocessing, genome alignment, and demultiplexing operations, as described in detail below with reference to a bioinformatics module 140. In some embodiments, sequence data aligned to a reference construct, eg, BAM file 124, is stored in the test patient data store 120.

[0137] In some embodiments, the test patient data store 120 includes feature data 125 that is useful, for example, for identifying clinical support for personalized cancer therapy. In some embodiments, the feature data 125 includes patient personal characteristics 126, such as patient name, date of birth, gender, ethnicity, physical address, smoking status, alcohol consumption characteristics, anthropomorphic data, etc.

[0138] In some embodiments, the feature data 125 includes patient medical history data 127, such as cancer diagnosis information (e.g., date of initial diagnosis, date of metastasis diagnosis, cancer staging, tumor characterization, tissue of origin, previous treatments and outcomes, adverse effects of treatments, history of treatment groups, history of clinical trials, previous and current medications, history of surgery, etc.), previous or current symptoms, previous or current treatments, previous treatment outcomes, previous disease diagnoses, diabetic status, depression diagnosis, other physical or mental illness diagnoses, and family history. In some embodiments, the feature data 125 includes clinical features 128, such as pathology data 128-1, medical imaging data 128-2, and tissue culture and / or tissue organoid culture data 128-3.

[0139] In some embodiments, additional clinical features, such as previous laboratory test results, are stored in the laboratory patient data store 120. Historical data 127 and clinical features may be collected from the patient from a variety of sources, including directly from the patient's electronic medical record (EMR) or electronic health record (EHR), or curated from other sources, such as fields from various laboratory records (e.g., gene sequencing reports).

[0140] In some embodiments, the feature data 125 includes patient genomic features 131. Non-limiting examples of genomic features include allele status 132 (e.g., allele identity at one or more loci, support for wild type or variant alleles at one or more loci, support for SNV / MNS at one or more loci, support for indels at one or more loci, and / or support for gene rearrangements at one or more loci), allele proportion 133 (e.g., ratio of variant to reference allele (or vice versa)), methylation status 134 (e.g., distribution of methylation patterns at one or more loci and / or support for aberrant methylation patterns at one or more loci), genomic copy number 135 (e.g., copy number status at one or more loci and / or support for aberrant methylation patterns at one or more loci), and / or allele proportion 136 (e.g., ratio of variant to reference allele (or vice versa)). genomic features 131 include, for example, aberrant (increased or decreased) copy number support at a locus above 131, tumor mutation burden 136 (e.g., a measure of the number of mutations in a subject's cancer genome), and microsatellite instability status 137 (e.g., a measure of repeat unit length at one or more microsatellite loci and / or a classification of MSI status for the patient's cancer). In some embodiments, one or more of the genomic features 131 are determined, for example, by a nucleic acid bioinformatics pipeline, as described in more detail below with reference to FIG. 4 (e.g., FIGS. 4A-E, 4F1, 4F2, and 4F3). In particular, in some embodiments, feature data 125 includes, for example, genomic copy number 135 (e.g., aberrant (increased or decreased) copy number support at a locus above 131), tumor mutation burden 136 (e.g., a measure of the number of mutations in a subject's cancer genome), and microsatellite instability status 137 (e.g., a measure of repeat unit length at one or more microsatellite loci and / or a classification of MSI status for the patient's cancer). 121-1 (135-1) variant allele fraction 133 and / or circulating tumor fraction estimates 131-i, which are determined using improved methods for analyzing copy number variation (CNV), validating somatic sequence variants, and / or determining circulating tumor fraction estimates using copy number variation analysis module 153, as described in more detail below with reference to Figures 1 and 4 (e.g., Figures 1C1, 1D1, 4F1; Figures 1C2, 1D2, and 4F2; and / or Figures 1C3, 1D3, and 4F3). In some embodiments, one or more of genomic features 131 are obtained from external testing sources not connected to the bioinformatics pipeline, for example, as described below.

[0141] For example, with reference to FIG. 1C1 , one or more genomic features 131, according to some embodiments of the present disclosure, include a genome copy number 135, including a liquid biopsy genome copy number 135-cf and an optional tumor biopsy genome copy number 135-t. In some embodiments, the liquid biopsy genome copy number 135-cf is determined by a nucleic acid bioinformatics pipeline (e.g., as described in more detail below with reference to FIGS. 4A-E and 4F1 ) using sequence reads 123 obtained from sequencing cell-free nucleic acids derived from the liquid biopsy sample. In some embodiments, the liquid biopsy genome copy number includes multiple copy number annotations (e.g., 135-cf-1, 135-cf-2, ... 135-cf-N), where each copy number annotation corresponds to a genomic target (e.g., a gene or genomic region). In some embodiments, the copy number annotations include a qualitative status and / or a quantitative copy number. In some further embodiments, the optional tumor biopsy genomic copy number 135-t is determined by a nucleic acid bioinformatics pipeline using multiple sequence reads 123 obtained from sequencing nucleic acids from a biopsy of a tumor (e.g., tumor tissue). In some embodiments, the optional tumor biopsy genomic copy number includes multiple optional copy number annotations (e.g., 135-1-t-1, 135-1-t-2, ... 135-1-tO), where each copy number annotation corresponds to a genomic target (e.g., a gene or genomic region).

[0142] 1B, in some embodiments, feature data 125 further includes data 138 from other "-omics" research fields. Non-limiting examples of "-omics" research fields that may yield feature data useful for providing clinical support for personalized cancer therapy include transcriptomics, epigenomics, proteomics, metabolomics, metabonomics, microbiomics, lipidomics, glycomics, cellomics, and organomics.

[0143] In some embodiments, still other features may include, but are not limited to, those listed above, including features derived from machine learning approaches based at least in part on the evaluation of any relevant molecular or clinical features, considered alone or in combination. For example, in some embodiments, one or more latent features learned from the evaluation of a cancer patient training dataset improve the diagnostic and prognostic capabilities of various analysis algorithms in feature analysis module 160.

[0144] Those skilled in the art will know of other types of features that are useful for providing clinical support for personalized cancer therapy. The above list of features is merely representative and should not be construed as limiting.

[0145] In some embodiments, the test patient data store 120 includes clinical assessment data 139 for the patient, e.g., based on the feature data 125 collected for the subject. In some embodiments, the clinical assessment data 139 includes a catalog of actionable variants and features 139-1 (e.g., genomic alterations and composite metrics based on genomic features known or thought to be targetable by one or more specific cancer therapies), matched therapies 139-2 (e.g., therapies known or thought to be particularly beneficial for treating subjects with actionable variants), and / or clinical reports 139-3 generated for the subject, e.g., based on the identified actionable variants and features 139-1 and / or matched therapies 139-2.

[0146] In some embodiments, clinical assessment data 139 is generated by analysis of feature data 125 using various algorithms in feature analysis module 160, as described in further detail below. In some embodiments, clinical assessment data 139 is generated, modified, and / or verified by evaluation of feature data 125 by a clinician, e.g., an oncologist. For example, in some embodiments, a clinician (e.g., in clinical environment 220) uses feature analysis module 160 or directly accesses laboratory patient data store 120 to evaluate feature data 125 and recommend a personalized cancer treatment for the patient. Similarly, in some embodiments, a clinician (e.g., in clinical environment 220) reviews recommendations determined using feature analysis module 160 and, for example, approves, rejects, or modifies the recommendations before they are sent to a medical professional treating the cancer patient.

[0147] Bioinformatics module (140).

[0148] Referring again to FIG. 1A, the system (e.g., system 100) includes a bioinformatics module 140, which includes a feature extraction module 145 and optional auxiliary data processing constructs, such as a sequence data processing module 141 and / or one or more reference sequence constructs 158 (e.g., a reference genome, exome, or target panel construct that includes reference sequences for multiple loci targeted by the sequencing panel).

[0149] In some embodiments, bioinformatics module 140 includes a sequence data processing module 141 that includes instructions for processing sequence reads, e.g., raw sequence reads 123 from one or more sequencing reactions 122-i, prior to analysis by various feature extraction algorithms, as described in detail below. In some embodiments, sequence data processing module 141 includes one or more preprocessing algorithms 142 that prepare the data for analysis. In some embodiments, preprocessing algorithm 142 includes instructions for converting the file format of sequence reads from the output of a sequencer (e.g., a BCL file format) into a file format compatible with downstream analysis of the sequences (e.g., a FASTQ or FASTA file format). In some embodiments, the preprocessing algorithm 142 includes instructions for assessing the quality of sequence reads (e.g., by examining quality metrics such as Phred score, base calling error probability, quality (Q) score, and the like) and / or removing sequence reads that do not meet a threshold quality (e.g., an estimated base calling accuracy of at least 80%, at least 90%, at least 95%, at least 99%, at least 99.5%, at least 99.9%, or more). In some embodiments, the preprocessing algorithm 142 includes instructions for filtering sequence reads for one or more properties, for example, removing sequences that fail to meet a lower or upper size threshold or removing duplicate sequence reads.

[0150] In some embodiments, the sequence data processing module 141 includes one or more alignment algorithms 143 for aligning the preprocessed sequence reads 123 to a reference sequence construct 158, such as a reference genome, exome, or target panel construct. Many algorithms for aligning sequencing data to a reference construct are known in the art, such as BWA, Blat, SHRiMP, LastZ, and MAQ. One example of a sequence read alignment is the Burrows-Wheeler Alignment Tool (BWA), which uses the Burrows-Wheeler Transform (BWT) to align short sequence reads to a large reference construct, allowing for mismatches and gaps. Li and Durbin, Bioinformatics, 25(14):1754-60 (2009), the contents of which are incorporated herein by reference in their entirety for all purposes. The sequence read alignment package imports raw or pre-processed sequence reads 122, for example, in BCL, FASTA, or FASTQ file format, and outputs aligned sequence reads 124, for example, in SAM or BAM file format.

[0151] In some embodiments, the sequence data processing module 141 includes one or more demultiplexing algorithms 144 to split sequence read or sequence alignment files generated from pooled nucleic acid sequencing reactions into separate sequence read or sequence alignment files, each corresponding to a different source of nucleic acid in the nucleic acid sequencing pool. For example, due to the cost of sequencing reactions, it is common practice to pool nucleic acids from multiple samples into a single sequencing reaction. Nucleic acids from each sample are tagged with sample-specific and / or molecule-specific sequence tags (e.g., UMIs) that are sequenced along with the molecule. In some embodiments, the demultiplexing algorithm 144 sorts these sequence tags within the sequence read or sequence alignment files and demultiplexes the sequencing data into separate files for each sample included in the sequencing reaction.

[0152] The bioinformatics module 140 includes a feature extraction module 145 that includes instructions for identifying diagnostic features, e.g., genomic features 131, from sequencing data 122 of one or more biological samples from a subject, e.g., a solid tumor sample, a liquid biopsy sample, or a normal tissue (e.g., control) sample. For example, in some embodiments, the feature extraction algorithm compares the identity of one or more nucleotides at a locus from the sequencing data 122 to the identity of the nucleotide at that locus in a reference sequence construct (e.g., a reference genome, exome, or targeted panel construct) to determine whether the subject has a variant at that locus. In some embodiments, the feature extraction algorithm evaluates data other than the raw sequence to identify genomic alterations in the subject, e.g., allele ratios, relative copy numbers, repeat unit distribution, etc.

[0153] For example, in some embodiments, feature extraction module 145 includes one or more variant identification modules that include instructions for various variant calling processes. In some embodiments, variants in a subject's germline are identified using, for example, germline variant identification module 146. In some embodiments, variants in a cancer genome, e.g., somatic variants, are identified using, for example, somatic variant identification module 150. Although separate germline and somatic variant identification modules are illustrated in FIG. 1A, in some embodiments, they are integrated into a single module. In some embodiments, the variant identification module includes instructions for identifying nucleotide variants (e.g., single nucleotide variants (SNVs) and multiple nucleotide variants (MNVs)) using one or more SNV / MNV calling algorithms (e.g., algorithms 147 and / or 151), indels (e.g., nucleotide insertions or deletions) using one or more indel calling algorithms (e.g., algorithms 148 and / or 152), and genomic rearrangements (e.g., nucleotide sequence inversions, translocations, and fusions) using one or more genomic rearrangement calling algorithms (e.g., algorithms 149 and / or 153).

[0154] The SNV / MNV algorithm 147 can identify single-nucleotide substitutions occurring at specific positions in the genome. For example, at a specific base position or locus in the human genome, a C nucleotide may occur in most individuals, while in a minority of individuals, the position is occupied by an A. This means that a SNP exists at this specific position, and the two possible nucleotide variations, C or A, are said to be alleles for this position. SNPs underlie differences in human susceptibility to a wide range of diseases (e.g., sickle cell anemia, beta-thalassemia, and cystic fibrosis resulting from SNPs). Disease severity and how the body responds to treatment are also manifestations of genetic variation. For example, a single-nucleotide mutation in the APOE (apolipoprotein E) gene is associated with a lower risk of Alzheimer's disease. A single-nucleotide variant (SNV) is a variation in a single nucleotide without any frequency restriction and can occur somatically. Somatic single-nucleotide mutations (e.g., caused by cancer) may also be referred to as single-nucleotide changes. MNP (multiple nucleotide polymorphism) modules can identify substitutions of consecutive nucleotides at specific positions in the genome.

[0155] The indel calling algorithm 148 can identify insertions or deletions of bases in the genome of organisms classified among minor genetic variations. Indels typically measure 1 to 10,000 base pairs in length, while microindels are defined as indels resulting in a net change of 1 to 50 nucleotides. Indels can be contrasted with SNPs or point mutations. While indels insert and / or delete nucleotides from a sequence, point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels, which are insertions and / or deletions, can be used as genetic markers in natural populations, particularly in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except near highly repetitive regions, including homopolymers and microsatellites.

[0156] Genome rearrangement algorithms149 can identify hybrid genes formed from two previously separated genes. This can occur as a result of translocations, interstitial deletions, or chromosomal inversions. Gene fusions can play an important role in tumorigenesis. Fusion genes can contribute to tumorigenesis because they can produce abnormal proteins that are much more active than non-fusion genes. Often, fusion genes are cancer-causing oncogenes, including BCR-ABL, TEL-AML1 (ALL with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG, an interstitial deletion on chromosome 21 that frequently occurs in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by disrupting androgen receptor (AR) signaling and inhibiting AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematologic cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a fusion gene specific and characteristic of high-grade serous ovarian cancer. Oncogenic fusion genes can lead to gene products with new or distinct functions from the two fusion partners. Alternatively, proto-oncogenes can be fused to strong promoters, thereby setting up oncogenic function through upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where oncogenes are juxtaposed to immunoglobulin gene promoters. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events. Because chromosomal translocations play such an important role in neoplasia, a dedicated database of chromosome aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer.

[0157] In some embodiments, feature extraction module 145 includes instructions for identifying one or more complex genomic alterations (e.g., features that incorporate more than changes in the primary sequence of the genome) in a subject's cancer genome. For example, in some embodiments, feature extraction module 145 includes copy number variation (e.g., copy number variation analysis module 153), microsatellite instability status (e.g., microsatellite instability analysis module 154), tumor mutation burden (e.g., tumor mutation burden analysis module 155), tumor ploidy (e.g., tumor ploidy analysis module 156), and homologous recombination pathway deficiency (e.g., homologous recombination pathway analysis module 157).

[0158] For example, referring to FIG. 1D , in some embodiments, feature extraction module 145 includes tumor fraction estimation module 145-tf. In some embodiments, tumor fraction estimation module 145-tf includes sequence ratio data structure 145-tf-r that includes a plurality of sequence ratios (e.g., coverage ratios) obtained from sequencing a subject's test liquid biopsy sample. In some embodiments, sequence ratio data structure 145-tf-r includes sequence ratios that are used as inputs to determine a tumor fraction estimate for the test liquid biopsy sample. In some embodiments, tumor fraction estimation module 145-tf includes tumor purity algorithm construct 145-tf-a, which, for example, performs maximum likelihood estimation (e.g., an expectation-maximization algorithm) to calculate an estimate of circulating tumor fraction. The tumor purity algorithm construct 145-tf-a includes an optional input data filtration construct 145-tf-k (e.g., for filtering out one or more inputs migrated from the sequence ratio data structure based on a minimum probe threshold or a location on a sex chromosome), and a plurality of model parameters 145-tf-d (e.g., 145-tf-d-1, 145-tf-d-2, ...) used to execute the algorithm. In some embodiments, the model parameters include: a predicted sequence ratio for a set of copy statuses at a given tumor purity; a distance (e.g., error) from a test sequence ratio to the most recent predicted sequence ratio at a given tumor purity; a minimum distance (e.g., minimum error) from a test sequence ratio to the most recent predicted sequence ratio at a given tumor purity (an assigned test copy status selected from the minimum distance predicted copy statuses); and / or a tumor purity score (e.g., a sum of weighted errors).

[0159] 1C , tumor fraction estimation module 145-tf is used to obtain one or more circulating tumor fraction estimates 131-i, which are included as feature data 125 in subject patient data store 120. For example, in some embodiments, multiple circulating tumor fraction estimates are obtained from subject liquid biopsy samples 131-i-cf (e.g., 131-i-cf-1, 131-i-cf-2, ..., 131-i-cf-N). In some embodiments, multiple circulating tumor fraction estimates are obtained from a single patient at different collection times.

[0160] A feature analysis module (160).

[0161] 1A , the system (e.g., system 100) includes a feature analysis module 160, which includes one or more genomic variation interpretation algorithms 161, one or more optional clinical data analysis algorithms 165, an optional treatment curation algorithm 165, and an optional recommendation validation module 167. In some embodiments, feature analysis module 160 uses one or more analysis algorithms (e.g., algorithms 162, 163, 164, and 165) to evaluate feature data 125 to identify actionable variants and characteristics 139-1 and corresponding matched treatments 139-2 and / or clinical trials. The identified actionable variants and characteristics 139-1 and corresponding matched treatments 139-2, which are optionally stored in test patient data store 120, are then curated by feature analysis module 160 to generate a clinical report 139-3, which is optionally verified by a user, e.g., a clinician, before being sent to a medical professional, e.g., an oncologist, treating the patient.

[0162] In some embodiments, the genomic variation interpretation algorithm 161 includes instructions for, for example, evaluating the impact of one or more genomic features 131 of a subject identified by the feature extraction module 145 on the characteristics of the patient's cancer and / or whether one or more targeted cancer therapies may improve the patient's clinical outcome. For example, in some embodiments, the one or more genomic variant analysis algorithms 163 evaluate various genomic features 131 by querying a database, e.g., a look-up table (“LUT”), of actionable genomic alterations, targeted therapies associated with the actionable genomic alterations, and any other conditions that should be met before administering the targeted therapy to a subject with an actionable genomic alteration. For example, evidence suggests that depatuxizumab mafodotin (an anti-EGFR mAb conjugated to monomethyl auristatin F) has improved efficacy for treating recurrent glioblastoma with EGFR focal amplification. van den Bent et al., 2017, Cancer Chemother Pharmacol., 80(6):1209-17. Thus, the LUT of actionable genomic alterations has an entry for a focal amplification of the EGFR gene, indicating that depatuxizumab mafodotin is a targeted therapy for glioblastoma with focal gene amplification (e.g., recurrent glioblastoma). In some cases, the LUT may also include contraindications to the associated targeted therapy, such as drug interactions or personal characteristics that contraindicate administration of a particular targeted therapy.

[0163] In some embodiments, the genomic alteration interpretation algorithm 161 determines whether a particular genomic feature 131 should be reported to a medical professional treating a cancer patient. In some embodiments, a genomic feature 131 (e.g., genomic alterations and composite features) is reported if there is clinical evidence that the feature significantly impacts cancer biology, affects cancer prognosis, and / or impacts pharmacogenomics, for example, by indicating or contraindicating a particular treatment approach. For example, the genomic variation interpretation algorithm 161 may classify a particular CNV feature 135 as "reportable," meaning, e.g., that the CNV has been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "non-reportable," meaning, e.g., that the CNV has not been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "no evidence," meaning, e.g., that there is no evidence to support the CNV being "reportable" or "non-reportable"; or as "conflicting evidence," meaning, e.g., that there is evidence to support both the CNV being "reportable" and the CNV being "non-reportable."

[0164] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more pathogenic variant analysis algorithms 162 that evaluate various genomic features to identify the presence of an oncogenic pathogen associated with the patient's cancer and / or targeted therapies associated with oncogenic pathogen infection in the cancer. For example, RNA expression patterns in some cancers are associated with the presence of oncogenic pathogens that are instrumental in inducing the cancer. See, e.g., U.S. Patent No. 11,043,304, the contents of which are incorporated by reference herein in their entirety for all purposes. In some cases, recommended treatments for cancer differ when the cancer is associated with oncogenic pathogen infection than when it is not. Thus, in some embodiments, for example, if feature data 125 includes RNA abundance data for the patient's cancer, one or more pathogenic variant analysis algorithms 162 evaluate the RNA abundance data for the patient's cancer to determine whether a signature indicative of the presence of an oncogenic pathogen in the cancer is present in the data. Similarly, in some embodiments, bioinformatics module 140 includes an algorithm that searches for the presence of pathogenic nucleic acid sequences in sequencing data 122. See, e.g., U.S. patent application Ser. No. 17 / 800,492, filed August 17, 2022, entitled "SYSTEMS AND METHODS FOR DETECTING VIRAL DNA FROM SEQUENCING," the contents of which are incorporated by reference herein in their entirety and for all purposes. Thus, in some embodiments, one or more pathogenic variant analysis algorithms 162 evaluate whether the presence of an oncogenic pathogen in a subject is associated with an actionable treatment for the infection. In some embodiments, system 100 queries a database, e.g., a look-up table (LUT), of actionable oncogenic pathogen infections, targeted therapies associated with actionable infections, and any other conditions that must be met before administering the targeted therapy to a subject infected with an oncogenic pathogen. In some cases, the LUT may also include contraindications to the associated targeted therapy, for example, drug interactions or personal characteristics that contraindicate the administration of a particular targeted therapy.

[0165] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more multi-feature analysis algorithms 164 that evaluate multiple features to classify cancers with respect to the efficacy of one or more targeted therapies. For example, in some embodiments, feature analysis module 160 includes one or more classifiers trained on feature data, one or more clinical therapies, and their associated clinical outcomes for multiple training subjects to classify cancers based on predicted clinical outcomes following one or more therapies.

[0166] In some embodiments, the classifier is implemented as an artificial intelligence engine and may include a gradient boosting model, a random forest model, a neural network (NN), a regression model, a naive Bayes model, and / or a machine learning algorithm (MLA). The MLA or NN may be trained from a training dataset that includes one or more features 125, including personal characteristics 126, medical history 127, clinical features 128, genomic features 131, and / or other "mix" features 138. MLA includes supervised algorithms (i.e., algorithms where the features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, nearest neighbor clustering, unsupervised algorithms (i.e., algorithms where the features / classifications in the dataset are not annotated) using Apriori, average clustering, principal component analysis, random forests, adaptive boosting, and semi-supervised algorithms (i.e., algorithms where an incomplete number of features / classifications in the dataset are annotated) using generative approaches (e.g., mixtures of Gaussian distributions, mixtures of multinomial distributions, hidden Markov models), sparse separation, graph-based approaches (e.g., min-cut, harmonic functions, manifold normalization), heuristic approaches, or support vector machines.

[0167] NNs include conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long-short-term memory networks, or other neural models where the training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample.

[0168] MLAs and neural networks identify distinct approaches to machine learning, and these terms may be used interchangeably herein. Thus, unless explicitly stated otherwise, a reference to an MLA may include a corresponding NN, or vice versa. Training may include providing an optimized dataset, labeling these features as they occur in patient records, and training the MLA to predict or classify based on new inputs. Artificial NNs are efficient computing models that have demonstrated their strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators, i.e., they can represent a wide variety of functions given appropriate parameters.

[0169] In some embodiments, system 100 includes a classifier training module that includes instructions for training one or more untrained or partially trained classifiers based on feature data from a training dataset. In some embodiments, system 100 also includes a database of training data for use in training the one or more classifiers. In other embodiments, the classifier training module accesses a remote storage device that hosts the training data. In some embodiments, the training data includes a set of training features, including, but not limited to, the various types of feature data 125 illustrated in FIG. 1B. In some embodiments, the classifier training module uses patient data 121, for example, when test patient data store 120 also stores records of treatments administered to patients and patient outcomes after treatment.

[0170] In some embodiments, feature analysis module 160 includes one or more clinical data analysis algorithms 165 that evaluate the clinical features 128 of the cancer to identify targeted therapies that may benefit the subject. For example, in some embodiments, where feature data 125 includes pathology data 128-1, for example, one or more clinical data analysis algorithms 165 evaluate the data to determine whether an actionable therapy is indicated, for example, based on the histopathology of a tumor biopsy from the subject, which may indicate a particular cancer type and / or stage of the cancer. In some embodiments, system 100 queries a database, e.g., a look-up table (“LUT”), of actionable clinical features (e.g., pathology features), targeted therapies associated with the actionable features, and any other conditions that must be met before administering the targeted therapy to a subject associated with the actionable clinical feature 128 (e.g., pathology feature 128-1). In some embodiments, system 100 directly evaluates clinical features 128 (e.g., pathology feature 128-1) to determine whether a patient's cancer is susceptible to a particular therapeutic agent. Further details about exemplary methods, systems, and algorithms for classifying cancer and identifying targeted therapies based on clinical data, such as pathology data 128-1, imaging data 138-2, and / or tissue culture / organoid data 128-3, are discussed in, for example, U.S. Patent Nos. 10,957,041, 10,957,445, 11,244,763, 11,848,107, and 11,145,416, the contents of which are incorporated herein by reference in their entirety and for all purposes.

[0171] In some embodiments, feature analysis module 160 includes a clinical trial module that evaluates the test patient data 121 to determine whether the patient is eligible for inclusion in clinical trials for cancer therapies, e.g., clinical trials that are currently recruiting patients, clinical trials that have not yet begun recruiting patients, and / or ongoing clinical trials that may recruit additional patients in the future. In some embodiments, the clinical trial module evaluates the test patient data 121 to determine whether clinical trial results, e.g., results of ongoing clinical trials and / or results of completed clinical trials, are relevant to the patient. For example, in some embodiments, system 100 queries a database of clinical trials, e.g., active and / or completed clinical trials, e.g., a look-up table (“LUT”), and compares the patient data 121 to the clinical trial inclusion criteria stored in the database to identify clinical trials that closely and / or exactly match the patient's data 121. In some embodiments, records of matching clinical trials, e.g., clinical trials for which the patient may be eligible and / or that may inform individualized treatment decisions for the patient, are stored in clinical evaluation database 139.

[0172] In some embodiments, the feature analysis module 160 includes a therapy curation algorithm 166 that assembles the actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials identified for the patient, as described above. In some embodiments, the therapy curation algorithm 166 evaluates certain criteria related to which actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials should be reported and / or whether certain matched therapies, considered alone or in combination, may be contraindicated for the patient, for example, based on the patient's personal characteristics 126 and / or known drug-drug interactions. In some embodiments, the therapy curation algorithm then generates one or more clinical reports 139-3 for the patient. In some embodiments, the therapy curation algorithm generates a first clinical report 139-3-1 that is reported to the medical professional treating the patient, and a second clinical report 139-3-2 that is not communicated to the medical professional but can be used to improve various algorithms within the system.

[0173] In some embodiments, the feature analysis module 160 includes a recommendation validation module 167 that includes an interface that allows a clinician to review, modify, and approve the clinical report 139-3 before the report is sent to a medical professional, e.g., an oncologist, treating the patient.

[0174] In some embodiments, each of one or more of the feature collection, sequencing module, bioinformatics module (including, for example, a variation module, structural variant calling and data processing module), classification module, and outcome module is communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In some alternative embodiments, each of the feature collection, variation module, structural variant calling and feature store is communicatively coupled to each other for independent communication without sharing a data bus.

[0175] Further details regarding modules and feature collection systems and exemplary embodiments are discussed in U.S. Patent No. 11,830,587, which is incorporated herein by reference in its entirety.

[0176] Exemplary Methods

[0177] For example, details of a system 100 for providing clinical support for personalized cancer therapy with improved liquid biopsy tumor mutation burden determination are disclosed, and details regarding the system's processes and features according to various embodiments of the present disclosure are disclosed below. Specifically, exemplary processes are described below with reference to Figures 3, 4, and 9 (e.g., Figures 3, 4A-E, and 9A-9E). In some embodiments, such processes and features of the system are performed by modules 118, 120, 140, 160, and / or 170, as shown in Figure 1A. With reference to these methods, the systems described herein (e.g., system 100) include instructions for determining liquid biopsy tumor mutation burden that are improved over conventional methods for determining liquid biopsy tumor mutation burden.

[0178] Figure 2B: Distributed diagnostic and clinical environment.

[0179] In some aspects, the methods described herein for providing clinical support for personalized cancer therapy are implemented across a distributed diagnostic / clinical environment, for example, as illustrated in Figure 2B. However, in some embodiments, the improved methods described herein for supporting clinical decisions in precision oncology using liquid biopsy assays (e.g., by determining liquid biopsy tumor mutation burden) are implemented at a single location, e.g., in a single computing system or environment, although ancillary procedures that support the methods described herein and / or procedures that further utilize the results of the methods described herein may be implemented across a distributed diagnostic / clinical environment.

[0180] FIG. 2B illustrates an example of a distributed diagnostic / clinical environment 210. In some embodiments, the distributed diagnostic / clinical environment is connected via a communications network 105. In some embodiments, one or more biological samples, such as one or more liquid biopsy samples, solid tumor biopsies, normal tissue samples, and / or control samples, are collected from a subject in a clinical environment 220, such as a doctor's office, hospital, or medical clinic, or in a home health care setting (not depicted). Advantageously, solid tumor samples should be collected within a clinical setting, while liquid biopsy samples can be obtained in a minimally invasive manner and are more easily collected outside of a traditional clinical setting. In some embodiments, the one or more biological samples, or portions thereof, are processed within the clinical environment 220 where collection occurred using a processing device 224, such as a nucleic acid sequencer to obtain sequencing data, a microscope to obtain pathology data, a mass spectrometer to obtain proteomic data, etc. In some embodiments, one or more biological samples or portions thereof are sent to one or more external environments, e.g., a sequencing laboratory 230, a pathology laboratory 240, and / or a molecular biology laboratory 250, each of which includes a processing device 234, 244, and 254, respectively, for generating subject biological data 121. Each environment includes a communication device 222, 232, 242, and 252, respectively, for communicating the subject biological data 121 to a processing server 262 and / or database 264, which may be located in yet another environment, e.g., a processing / storage center 260. Thus, in some embodiments, different portions of the systems and methods described herein are performed by different processing devices located in different physical environments.

[0181] Thus, in some embodiments, a method for providing clinical support for personalized cancer therapy, for example, with improved determination of liquid biopsy tumor mutation burden, is implemented across one or more environments, as illustrated in FIG. 2B . For example, in some such embodiments, a liquid biopsy sample is collected in a clinical setting 220 or a home health care setting. The sample, or a portion thereof, is sent to a sequencing laboratory 230, where raw sequence reads 123 of nucleic acids in the sample are generated by a sequencer 234. The raw sequencing data 123 is communicated, for example, from a communication device 232 to a database 264 in a processing / storage center 260, where a processing server 262 extracts features from the sequence reads by performing one or more processes in a bioinformatics module 140, thereby generating a genomic signature 131 for the sample. The processing server 262 can then analyze the identified features by performing one or more processes in a feature analysis module 160, thereby generating a clinical assessment 139, including a clinical report 139-3. The clinician may access the clinical report 139-3 via the recommendation validation module 167, for example, in the processing / storage center 260 or through the communication network 105. After final approval, the clinical report 139-3 is transmitted to a medical professional, e.g., an oncologist, in the clinical environment 220, who uses the report to support clinical decision-making for the patient's personalized cancer treatment.

[0182] Figure 2A: Exemplary workflow for precision oncology

[0183] 2A is a flowchart of an exemplary workflow 200 for collecting and analyzing data to generate a clinical report 139 for supporting clinical decision-making in precision oncology. Advantageously, the methods described herein improve this process by improving various stages within feature extraction 206, including, for example, determining liquid biopsy tumor mutation burden.

[0184] Briefly, the workflow begins with patient intake and sample collection 201, in which one or more liquid biopsy samples, one or more tumor biopsies, and one or more normal and / or control tissue samples are collected from a patient (e.g., in a clinical setting 220 or a home health care setting, as illustrated in FIG. 2B ). In some embodiments, personal data 126 corresponding to the patient and a record of the one or more biological samples obtained (e.g., patient identifier, patient clinical data, sample type, sample identifier, cancer status, etc.) are entered into a data analysis platform, e.g., laboratory patient data store 120. Thus, in some embodiments, the methods disclosed herein include obtaining one or more biological samples from one or more subjects, e.g., cancer patients. In some embodiments, the subject is a human, e.g., a human cancer patient.

[0185] Sequence reads are then generated from the sequencing library or pool of sequencing libraries (312). Sequencing data can be obtained by any methodology known in the art, such as sequencing-by-synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing-by-ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or next-generation sequencing (NGS) technology, such as paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators. In some embodiments, sequencing is performed using next-generation sequencing technology, such as short-read technology. In other embodiments, long-read sequencing or another sequencing method known in the art is used.

[0186] Referring again to FIG. 2A, nucleic acid sequencing data 122 generated from one or more patient samples is then evaluated in a bioinformatics pipeline (e.g., via variant analysis 206), e.g., using bioinformatics module 140 of system 100, to identify genomic alterations and other metrics in the patient's cancer genome. An exemplary overview of a bioinformatics pipeline is described below with respect to FIG. 4 (e.g., FIGS. 4A-E, 4F1-3, and / or 4G1-3). Advantageously, in some embodiments, the present disclosure improves bioinformatics pipelines such as pipeline 206 by improving methods and systems for copy number variation validation, somatic sequence variant validation, and / or circulating tumor fraction estimate determination.

[0187] 4A shows an exemplary bioinformatics pipeline 206 for providing clinical support for precision oncology (e.g., used for feature extraction in the workflows illustrated in FIGS. 2A and 3). As shown in FIG. 4A, sequencing data 122 (e.g., sequence reads 314) obtained from wet-lab processing 204 are input into the pipeline.

[0188] 4A shows an exemplary bioinformatics pipeline 206 for providing clinical support for precision oncology (e.g., used for feature extraction in the workflows illustrated in FIGS. 2A and 3). As shown in FIG. 4A, sequencing data 122 (e.g., sequence reads 314) obtained from wet-lab processing 204 are input into the pipeline.

[0189] In various embodiments, the bioinformatics pipeline includes a circulating tumor DNA (ctDNA) pipeline for analyzing liquid biopsy samples. The pipeline can detect SNVs, INDELs, copy number amplifications / deletions, and genomic rearrangements (e.g., fusions). The pipeline can employ a unique molecular landmark (UMI)-based consensus-based error suppression method and Bayesian trinucleotide context-based position-level error suppression. In various embodiments, this can detect variants with a variant allele fraction of 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.4%, or 0.5%.

[0190] Homologous recombination status (HRD):

[0191] In some embodiments, for example, analysis of aligned sequence reads in SAM or BAM format includes analysis of whether the cancer is homologous recombination deficient (HRD status 137-3) using the homologous recombination pathway analysis module 157.

[0192] Homologous recombination (HR) is a normal, highly conserved DNA repair process that allows the exchange of genetic information between identical or closely related DNA molecules. It is most widely used by cells to accurately repair harmful breaks (e.g., lesions) that occur on both strands of DNA. DNA damage can arise from exogenous (external) sources, such as ultraviolet light, radiation, or chemical damage, or from endogenous (internal) sources, such as errors in DNA replication or other cellular processes that generate DNA damage. Double-strand breaks are a type of DNA damage. The use of poly(ADP-ribose) polymerase (PARP) inhibitors in patients with HRD impairs both pathways of DNA repair, leading to cell death (apoptosis). The efficacy of PARP inhibitors is improved not only in ovarian cancers that exhibit germline or somatic BRCA mutations, but also in cancers where HRD is caused by other underlying etiologies.

[0193] In some embodiments, HRD status can be determined by inputting features correlated with HRD status into a classifier trained to distinguish between cancers with homologous recombination pathway defects and cancers without homologous recombination pathway defects. For example, in some embodiments, the features include one or more of: (i) the heterozygous state of a first plurality of DNA damage repair genes in the genome of the subject's cancer tissue; (ii) a measure of loss of heterozygosity across the genome of the subject's cancer tissue; (iii) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's cancer tissue; and (iv) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's non-cancerous tissue. In some embodiments, all four of the features described above are used as features in the HRD classifier. For more detailed information about HRD classifiers using these and other features, see U.S. Patent No. 10,975,445, the contents of which are incorporated herein by reference in their entirety and for all purposes.

[0194] Parallel Testing

[0195] Unless otherwise specified, as used herein, the term "parallel" when referring to an assay refers to a period of 0 to 90 days. In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 90 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 60 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 30 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 21 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 14 days (e.g., the biological samples are collected within that period). In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for blood-related cancers, and a non-cancerous sample) are conducted within a period of 0 to 7 days (e.g., the biological samples are collected within that period).In some embodiments, parallel studies using different biological samples from the same subject (e.g., two or more of a liquid biopsy sample, a cancerous tissue, e.g., a solid tumor sample, or a blood sample for hematological cancers, and a non-cancerous sample) are conducted within a 0-3 day period (e.g., the biological samples are collected within that period).

[0196] In some embodiments, liquid biopsy assays can be used in parallel with solid tumor assays to generate more comprehensive information about patient variants. For example, blood samples and solid tumor samples can be sent to a laboratory for evaluation. Solid tumor samples can be analyzed using a bioinformatics pipeline to generate solid tumor results. For example, solid tumor assays are described in U.S. Pat. No. 11,705,226, the contents of which are incorporated herein by reference in their entirety for all purposes. Solid tumor cancer types can include, for example, non-small cell lung cancer, colorectal cancer, or breast cancer. Alterations identified in tumor / matched normal results can include, for example, EGFR+ for non-small cell lung cancer, HER2+ for breast cancer, or KRAS G12C for some cancers.

[0197] In some embodiments, a blood sample may be divided into a first portion and a second portion. The blood sample and the first portion of the solid tumor sample may be analyzed using a bioinformatics pipeline to generate a tumor / matched normal result. The second portion of the blood sample may be analyzed using a bioinformatics pipeline to generate a liquid biopsy result. For example, a blood sample may be analyzed using improvements in at least somatic variant identification, e.g., as described herein in the section "Variant Identification." For example, a blood sample may be analyzed using improvements in focused copy number identification, e.g., as described herein in the section entitled "Copy Number Variation." For example, a blood sample may be analyzed using improvements in circulating tumor fraction determination, e.g., as described herein in the sections entitled "Systems and Methods for Improved Circulating Tumor Fraction Estimates" and / or "Systems and Methods for Improved Somatic Sequence Variant Validation."

[0198] A treatment may be identified for further consideration following the results of the liquid biopsy, along with the tumor or tumor / matched normal results. For example, if the overall results suggest that the patient has HER2+ breast cancer, neratinib may be identified along with the test results for further consideration by the prescribing physician.

[0199] Solid tumor or tumor / matched normal assays may be ordered in parallel, and their results may be delivered and analyzed in parallel.

[0200] Methods for improved determination of tumor mutation burden - Patents.com

[0201] An overview of a method for providing clinical support for personalized cancer treatment is described above with reference to Figures 2-4. Below, for example, within the context of the above-described methods and systems, a system and method for improving the determination of tumor mutation burden in a subject is described with reference to Figure 9.

[0202] Many of the embodiments described below in conjunction with Figure 9 relate to the analysis carried out using the sequencing data of the cfDNA obtained from the liquid biopsy sample of a subject, such as a cancer patient.Generally, these embodiments are independent and therefore do not depend on any specific DNA sequencing method.However, in some embodiments, the method described below comprises generating sequencing data.

[0203] As described herein, in some embodiments, the methods described herein (e.g., method 900 shown in FIG. 9 ) include one or more data collection steps in addition to data analysis and downstream steps. As described herein, e.g., with reference to FIGS. 2 and 3 , in some embodiments, the methods include collection of a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). Similarly, as described herein, e.g., with reference to FIGS. 2 and 3 , in some embodiments, the methods include extraction of DNA from a liquid biopsy sample (cfDNA) from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). Similarly, as described herein, e.g., with reference to FIGS. 2 and 3 , in some embodiments, the methods include nucleic acid sequencing of DNA from a liquid biopsy (cfDNA) sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). Advantageously, the method and system described herein can accurately classify the lineage of variants as either somatic or hematopoietic based on the sequencing data of only cfDNA fragments.Therefore, in some embodiments, the matched cancerous sample and / or matched non-cancerous sample from subject are not used in the method described herein.

[0204] However, in other embodiments, the methods described herein begin with obtaining nucleic acid sequencing results, such as raw or corrected sequence reads of DNA from a liquid biopsy sample (cfDNA) from a subject, and optionally one or more matched biological samples from the subject (e.g., matched cancerous and / or matched non-cancerous samples from the subject), from which genomic features necessary for detecting clonal hematopoietic and / or solid tumor variants can be determined. For example, in some embodiments, sequence data 122 of patient 121 is accessed and / or downloaded by system 100 via network 105.

[0205] In some embodiments, the method further comprises isolating a plurality of cell-free nucleic acids from the liquid biopsy sample of the subject prior to sequencing. In some embodiments, the sequencing is multiplexed sequencing. In some embodiments, the sequencing is short-read sequencing or long-read sequencing.

[0206] Similarly, in some embodiments, the methods described herein begin with obtaining genomic features necessary for filtering clonal hematopoietic variants from sequencing a liquid biopsy sample from a subject, and optionally one or more matched biological samples (e.g., matched cancerous and / or matched non-cancerous samples from the subject). For example, in some embodiments, (i) one or more fragment length indicators, (ii) one or more features determined from the variant allele fraction of a candidate somatic variant and the ctFE of the liquid biopsy sample, or the variant allele fraction of a candidate somatic variant and the ctFE of the liquid biopsy sample, and (iii) one or more indicators of clonal hematopoietic prevalence at a first nucleotide position are accessed and / or downloaded by system 100 via network 105.

[0207] 9A-9E collectively provide a flowchart of processes and features for determining the liquid biopsy tumor mutation burden (lTMB) of a subject's cell-free DNA from a liquid biopsy assay, according to some embodiments of the present disclosure (block 902).

[0208] Block 902. In some embodiments, referring to block 902, the method includes obtaining a nucleic acid sequence corresponding to each cell-free DNA (cfDNA) fragment in the plurality of DNA fragments (e.g., cfDNA fragments) from a plurality of sequence reads of a sequencing reaction of the plurality of DNA fragments from one or more biological samples derived from a subject.

[0209] In some embodiments, the plurality of sequence reads are sequence reads from a panel enrichment sequencing reaction, comprising a first subset of sequence reads corresponding to the cfDNA fragments targeted by one or more probes in the target enrichment panel.In some embodiments, each cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to each probe sequence in the plurality of probe sequences, and the probe sequences are used to enrich the cell-free DNA fragments in the liquid biopsy sample in the panel enrichment sequencing reaction.In some embodiments, the plurality of probe sequences are mapped to 150 or less genes in the human genome.

[0210] 2B, nucleic acid sequencing of one or more samples collected from a subject is performed during wet-lab processing 204, for example, in a sequencing laboratory 230. An exemplary workflow for nucleic acid sequencing is illustrated in FIG. 3. In some embodiments, one or more biological samples obtained in the sequencing laboratory 230 are registered (302) to track the samples and data throughout the sequencing process.

[0211] Next, nucleic acids, e.g., RNA and / or DNA, are extracted from one or more biological samples (304). Methods for isolating nucleic acids from biological samples are known in the art and depend on the type of nucleic acid being isolated (e.g., cfDNA, DNA, and / or RNA) and the type of sample from which the nucleic acid is isolated (e.g., liquid biopsy samples, leukocyte buffy coat preparations, formalin-fixed paraffin-embedded (FFPE) solid tissue samples, and fresh-frozen solid tissue samples). The selection of any particular nucleic acid isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the sample type, sample condition, type of nucleic acid to be sequenced, and sequencing technology used.

[0212] Many techniques for DNA isolation, e.g., genomic DNA isolation, from tissue samples are known in the art, such as organic extraction, silica adsorption, and anion exchange chromatography. Similarly, many techniques for RNA isolation, e.g., mRNA isolation, from tissue samples are known in the art. For example, acid guanidine thiocyanate-phenol-chloroform extraction (see, e.g., Chomczynski and Sacchi, 2006, Nat Protoc, 1(2):581-85, incorporated herein by reference) and silica bead / glass fiber adsorption (see, e.g., Poeckh et al., 2008, Anal Biochem., 373(2):253-62, incorporated herein by reference). The selection of any particular DNA or RNA isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the tissue type, tissue condition (e.g., fresh, frozen, formalin-fixed, paraffin-embedded (FFPE)), and the type of nucleic acid analysis to be performed.

[0213] In some embodiments where the biological sample is a liquid biopsy sample, e.g., a blood or plasma sample, cfDNA is isolated from the blood sample using commercially available reagents containing proteinase K to produce a liquid solution of cfDNA.

[0214] In some embodiments, the isolated DNA molecules are mechanically sheared to an average length using an ultrasonicator (e.g., a Covaris ultrasonicator). In some embodiments, the isolated nucleic acid molecules are analyzed to determine their fragment size, for example, through gel electrophoresis and / or the use of a device such as a LabChip GX Touch. Those skilled in the art will understand the appropriate range of fragment sizes based on the sequencing technology being employed, as different sequencing technologies have different fragment size requirements for robust sequencing. In some embodiments, quality control tests are performed on the extracted nucleic acids (e.g., DNA and / or RNA) to, for example, assess nucleic acid concentration and / or fragment size. Sizing DNA fragments, such as determining whether DNA fragments require additional shearing before sequencing, provides valuable information for use in downstream processing.

[0215] Wet-lab processing 204 then includes preparing a nucleic acid library from the isolated nucleic acids (e.g., cfDNA, DNA, and / or RNA). For example, in some embodiments, a DNA library (e.g., a gDNA and / or cfDNA library) is prepared from the isolated DNA from one or more biological samples. In some embodiments, the DNA library is prepared using a commercially available library preparation kit, such as a KAPA Hyper Prep Kit, a New England Biolabs (NEB) kit, or a similar kit.

[0216] In some embodiments, adaptors (e.g., UDI adaptors such as Roche SeqCap double-ended adaptors, or UMI adaptors such as full-length or short Y adaptors) are ligated onto nucleic acid molecules during library preparation. In some embodiments, the adaptors comprise unique molecular identifiers (UMIs), which are short nucleic acid sequences (e.g., 3-10 base pairs) that are added to the ends of DNA fragments during adaptor ligation. In some embodiments, the UMIs are degenerate base pairs that serve as unique tags that can be used to identify sequence reads derived from specific DNA fragments. In some embodiments, patient-specific indicators are added to nucleic acid molecules, for example, when multiplex sequencing is used to sequence DNA from multiple samples (e.g., from the same or different subjects) in a single sequencing reaction. In some embodiments, patient-specific indicators are short nucleic acid sequences (e.g., 3-20 nucleotides) that are added to the ends of DNA fragments during library construction that serve as unique tags that can be used to identify sequence reads derived from specific patient samples. Examples of identifier sequences are described, for example, in Kivioja et al., 2011, Nat. Methods 9(1):72-74, and Islam et al., 2014, Nat. Methods 11(2):163-66, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0217] In some embodiments, the adapters include a PCR primer landing site designed for efficient binding of a PCR or second-strand synthesis primer used during the sequencing reaction. In some embodiments, the adapters include an anchor binding site to facilitate binding of the DNA molecule to an anchor oligonucleotide molecule on a sequencer flow cell, serving as a seed for the sequencing process by providing a starting point for the sequencing reaction. During PCR amplification after adapter ligation, the UMI, patient index, and binding site are replicated along with the attached DNA fragment. This provides a method for identifying sequence reads derived from the same original fragment in downstream analysis.

[0218] In some embodiments, the DNA library is amplified and purified using commercially available reagents (e.g., Axygen MAG PCR cleanup beads). In some such embodiments, the concentration and / or amount of DNA molecules is then quantified using a fluorescent dye and fluorescence microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In some embodiments, library amplification is performed on a device (e.g., Illumina C-Bot2), and the resulting flow cells containing the amplified target capture DNA library are sequenced to a specific on-target depth selected by the user on a next-generation sequencer (e.g., Illumina HiSeq 4000 or Illumina NovaSeq 6000). In some embodiments, DNA library preparation is performed using an automated system that uses a liquid handling robot (e.g., SciClone NGSx).

[0219] In some embodiments where the feature data 125 includes the methylation state 132 of one or more genomic locations, nucleic acids (e.g., cfDNA) isolated from a biological sample are processed to convert unmethylated cytosines to uracil, for example, before generating a sequencing library. Thus, when the nucleic acids are sequenced, all cytosines called in the sequencing reaction are necessarily methylated, since unmethylated cytosines are converted to uracil and would therefore be called thymidine rather than cytosine in the sequencing reaction. Commercially available kits, such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, and EZ DNA Methylation™-Lightning kits (available from Zymo Research Corp., Irvine, CA), are available for bisulfite-mediated conversion of methylated cytosines to uracil. Commercially available kits, such as the APOBEC-Seq kit (available from NEBiolabs, Ipswich, Mass.), are also available for the enzymatic conversion of methylated cytosine to uracil.

[0220] In some embodiments, wet-lab processing 204 includes pooling (308) DNA molecules from multiple libraries corresponding to different samples from the same and / or different patients to form a sequencing pool of DNA libraries. When the pool of DNA libraries is sequenced, the resulting sequence reads correspond to nucleic acids isolated from the multiple samples. The sequence reads can be separated into different sequence read files corresponding to the various samples represented by the sequencing reads based on unique identifiers present in the added nucleic acid fragments. In this manner, a single sequencing reaction can generate sequence reads from multiple samples. Advantageously, this allows for the processing of more samples per sequencing reaction.

[0221] In some embodiments, wet-lab processing 204 includes enriching 310 a sequencing library or pool of sequencing libraries for target nucleic acids, e.g., nucleic acids encompassing loci that are useful for precision oncology and / or used as internal controls for sequencing or bioinformatics processes. In some embodiments, enrichment is achieved by hybridizing target nucleic acids in the sequencing library to probes that hybridize to the target sequences and then isolating the captured nucleic acids from off-target nucleic acids that are not bound by the capture probes. Of course, some off-target nucleic acids will remain in the final sequencing pool.

[0222] In some embodiments, the plurality of sequence reads obtained from the above-described sequencing comprises at least 10,000 sequence reads, at least 50,000 sequence reads, at least 100,000 sequence reads, at least 500,000 sequence reads, at least 1 million sequence reads, at least 5 million sequence reads, at least 10 million sequence reads, or more. In some embodiments, the plurality of sequence reads comprises 1 billion or less sequence reads, 500 million or less sequence reads, 100 million or less sequence reads, 50 million or less sequence reads, 10 million or less sequence reads, 5 million or less sequence reads, 1 million or less sequence reads, or less. In some embodiments, the plurality of sequence reads is between 10,000 and 1 billion sequence reads, between 10,000 and 500 million sequence reads, between 10,000 and 100 million sequence reads, between 10,000 and 50 million sequence reads, between 10,000 and 10 million sequence reads, between 10,000 and 5 million sequence reads, or between 10,000 and 1 million sequence reads. In some embodiments, the plurality of sequence reads is between 100,000 and 1 billion sequence reads, between 100,000 and 500 million sequence reads, between 100,000 and 100 million sequence reads, between 100,000 and 50 million sequence reads, between 100,000 and 10 million sequence reads, between 100,000 and 5 million sequence reads, or between 100,000 and 1 million sequence reads. In some embodiments, the plurality of sequence reads is between 500,000 and 1 billion sequence reads, between 500,000 and 500 million sequence reads, between 500,000 and 100 million sequence reads, between 500,000 and 50 million sequence reads, between 500,000 and 10 million sequence reads, between 500,000 and 5 million sequence reads, or between 500,000 and 1 million sequence reads.In some embodiments, the plurality of sequence reads is between 1 million and 1 billion sequence reads, between 1 million and 500 million sequence reads, between 1 million and 100 million sequence reads, between 1 million and 50 million sequence reads, between 1 million and 10 million sequence reads, or between 1 million and 5 million sequence reads.

[0223] In some embodiments, the plurality of DNA (e.g., cfDNA) fragments comprises at least 1,000 DNA (e.g., cfDNA) fragments, at least 5,000 DNA (e.g., cfDNA) fragments, at least 10,000 DNA (e.g., cfDNA) fragments, at least 50,000 DNA (e.g., cfDNA) fragments, at least 100,000 DNA (e.g., cfDNA) fragments, at least 500,000 DNA (e.g., cfDNA) fragments, at least 1 million DNA (e.g., cfDNA) fragments, at least 5 million DNA (e.g., cfDNA) fragments, or more. In some embodiments, the plurality of DNA (e.g., cfDNA) fragments comprises 100 million or fewer DNA (e.g., cfDNA) fragments, 50 million or fewer DNA (e.g., cfDNA) fragments, 10 million or fewer DNA (e.g., cfDNA) fragments, 5 million or fewer DNA (e.g., cfDNA) fragments, 1 million or fewer DNA (e.g., cfDNA) fragments, 500,000 or fewer DNA (e.g., cfDNA) fragments, 100,000 or fewer DNA (e.g., cfDNA) fragments, or fewer. In some embodiments, the plurality of DNA (e.g., cfDNA) fragments is between 1000 DNA (e.g., cfDNA) fragments and 500 million DNA (e.g., cfDNA) fragments, between 1000 DNA (e.g., cfDNA) fragments and 100 million DNA (e.g., cfDNA) fragments, between 1000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments, between 1000 DNA (e.g., cfDNA) fragments and 10 million DNA (e.g., cfDNA) fragments, , cfDNA) fragments to 5 million DNA (e.g., cfDNA) fragments, 1000 DNA (e.g., cfDNA) fragments to one million DNA (e.g., cfDNA) fragments, 1000 DNA (e.g., cfDNA) fragments to 500,000 DNA (e.g., cfDNA) fragments, 1000 DNA (e.g., cfDNA) fragments to 250,000 DNA (e.g., cfDNA) fragments, or 1000 DNA (e.g., cfDNA) fragments to 100,000 DNA (e.g., cfDNA).In some embodiments, the plurality of DNA (e.g., cfDNA) fragments is between 5,000 DNA (e.g., cfDNA) fragments and 500 million DNA (e.g., cfDNA) fragments, between 5,000 DNA (e.g., cfDNA) fragments and 100 million DNA (e.g., cfDNA) fragments, between 5,000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments, between 5,000 DNA (e.g., cfDNA) fragments and 10 million DNA (e.g., cfDNA) fragments, or between 5,000 DNA (e.g., cfDNA) fragments and 10 million DNA (e.g., cfDNA) fragments. , cfDNA) fragments to 5 million DNA (e.g., cfDNA) fragments, 5,000 DNA (e.g., cfDNA) fragments to one million DNA (e.g., cfDNA) fragments, 5,000 DNA (e.g., cfDNA) fragments to 500,000 DNA (e.g., cfDNA) fragments, 5,000 DNA (e.g., cfDNA) fragments to 250,000 DNA (e.g., cfDNA) fragments, or 5,000 DNA (e.g., cfDNA) fragments to 100,000 DNA (e.g., cfDNA). In some embodiments, the plurality of DNA (e.g., cfDNA) fragments is between 10,000 DNA (e.g., cfDNA) fragments and 500 million DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and 100 million DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and 10 million DNA (e.g., cfDNA) fragments, or between 10,000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments. For example, between 10,000 DNA (e.g., cfDNA) fragments and 5 million DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and one million DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and 500,000 DNA (e.g., cfDNA) fragments, between 10,000 DNA (e.g., cfDNA) fragments and 250,000 DNA (e.g., cfDNA) fragments, or between 10,000 DNA (e.g., cfDNA) fragments and 100,000 DNA (e.g., cfDNA).In some embodiments, the plurality of DNA (e.g., cfDNA) fragments is between 25,000 DNA (e.g., cfDNA) fragments and 500 million DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 100 million DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 10 million DNA (e.g., cfDNA) fragments, or between 25,000 DNA (e.g., cfDNA) fragments and 50 million DNA (e.g., cfDNA) fragments. For example, between 5 million DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 1 million DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 500,000 DNA (e.g., cfDNA) fragments, between 25,000 DNA (e.g., cfDNA) fragments and 250,000 DNA (e.g., cfDNA) fragments, or between 25,000 DNA (e.g., cfDNA) fragments and 100,000 DNA (e.g., cfDNA).

[0224] In some embodiments, obtaining, enrolling, storing, preparing, processing, and / or analyzing the biopsy sample from the subject comprises any of the methods and / or embodiments described above in this disclosure. In some embodiments, the sequencing reaction comprises any of the methods and / or embodiments described above in this disclosure.

[0225] In some embodiments, all or substantially all of the aligned sequence reads are evaluated to identify candidate sequence variants (e.g., candidate somatic and / or germline sequence variants). In other embodiments, a subset of the aligned sequence reads is evaluated to identify candidate sequence variants. For example, in one embodiment, a targeted panel sequencing reaction is used to generate sequencing data 122, and only sequence reads corresponding to the target panel (on-target reads) are evaluated to identify candidate sequence variants. In some embodiments, a targeted panel sequencing reaction is used to generate sequencing data 122, and a subset of sequence reads corresponding to a subset of the target panel are evaluated to identify candidate sequence variants. In some embodiments, regardless of whether the sequencing reaction is a targeted panel sequencing reaction, a whole exome sequencing reaction, or a whole genome sequencing reaction, a subset of sequence reads corresponding to a subset of genes are evaluated to identify candidate sequence variants. In some embodiments, a subset of sequence reads corresponding to a defined set of regions within the genome, such as one or more genes, one or more introns, one or more exons, or subregions of one or more introns and / or exons, associated with the etiology of cancer, are evaluated to identify candidate sequence variants.

[0226] Alternatively, in some embodiments, regardless of which subset of aligned sequence reads is evaluated to identify candidate sequence variants, only a subset of the candidate sequence variants are further validated. For example, in some embodiments, only candidate sequence variants corresponding to a target panel (on-target reads) are validated. Similarly, in some embodiments, only candidate sequence variants corresponding to a subset of the target panel are validated. Similarly, in some embodiments, regardless of whether the sequencing reaction is a targeted panel sequencing reaction, a whole exome sequencing reaction, or a whole genome sequencing reaction, only candidate sequence variants corresponding to a subset of genes are validated. Similarly, in some embodiments, only candidate variants corresponding to a set of defined regions within the genome, such as one or more genes, one or more introns, one or more exons, or subregions of one or more introns and / or exons associated with the etiology of cancer, are validated.

[0227] In some embodiments, enrichment is performed before pooling multiple nucleic acid sequencing libraries, however, in other embodiments, enrichment is performed after pooling the nucleic acid sequencing libraries, which has the advantage of reducing the number of enrichment assays that need to be performed.

[0228] In some embodiments, enrichment is performed before generating a nucleic acid sequencing library. This has the advantage that fewer reagents are required to perform both enrichment (because there are fewer target sequences at this point before library amplification) and library production (because there are fewer nucleic acid molecules to tag and amplify after enrichment). However, this increases the possibility of pull-down bias and / or the possibility that small variations in the enrichment protocol will result in inconsistent results.

[0229] In some embodiments, nucleic acid libraries are pooled (two or more DNA libraries can be mixed to create a pool) and treated with a reagent to reduce off-target capture, e.g., human COT-1 and / or IDT xGen Universal Blocker. The pool can be dried in a centrifugal concentrator and resuspended. The DNA library or pool can be hybridized to a probe set (e.g., a probe set specific to a panel including at least 100, 600, 1,000, 10,000, etc. loci of the 19,000 known human genes) and amplified using commercially available reagents (e.g., KAPA HiFi HotStart ReadyMix). For example, in some embodiments, the pool is incubated in an incubator, PCR machine, water bath, or other temperature-regulating device to allow the probes to hybridize. The pool can then be mixed with streptavidin-coated beads or another means to capture hybridized DNA probe molecules, such as DNA molecules representing exons of the human genome and / or genes selected for the gene panel.

[0230] Pools can be amplified and purified more than once using commercially available reagents, such as the KAPA HiFi Library Amplification Kit and Axygen MAG PCR cleanup beads, respectively. Pools or DNA libraries can be analyzed to determine the concentration or quantity of DNA molecules, for example, by using fluorescent dyes (e.g., PicoGreen pool quantification) and a fluorescent microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In one example, DNA library preparation and / or capture is performed using an automated system that uses a liquid handling robot (e.g., SciClone NGSx).

[0231] In some embodiments, for example, when whole genome sequencing is used, nucleic acid sequencing libraries are not subjected to target enrichment before sequencing, so as to obtain sequencing data for substantially all competent nucleic acids in the sequencing library.Similarly, in some embodiments, for example, when whole genome sequencing is used, nucleic acid sequencing libraries are not mixed because of the associated processing power limitations in obtaining meaningful sequencing depth across the whole genome.However, in other embodiments, for example, when low-pass whole genome sequencing (LPWGS) is used, nucleic acid sequencing libraries can be pooled because very low average sequencing coverage is achieved across each genome, for example, about 0.5x to about 5x.

[0232] In some embodiments, multiple nucleic acid probes (e.g., probe sets) are used to enrich for one or more target sequences in a nucleic acid sample (e.g., an isolated nucleic acid sample or a nucleic acid sequencing library), e.g., where one or more target sequences are beneficial for precision oncology. For example, in some embodiments, one or more of the target sequences encompass a locus associated with an actionable allele. That is, variations in the target sequence are relevant to targeted therapeutic approaches. In some embodiments, one or more of the target sequences and / or one or more characteristics of the target sequences are used in a classifier trained to distinguish between two or more cancer states.

[0233] Block 904. Referring to block 904, in some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 1,000x. In some embodiments, the panel targeted sequencing is performed to an average on-target depth of at least 500x, at least 750x, at least 1000x, at least 2500x, at least 500x, at least 10,000x, or more. In some embodiments, the samples are further evaluated for uniformity above a sequencing depth threshold (e.g., 95% of all target base pairs at 300x sequencing depth). In some embodiments, the sequencing depth threshold is a minimum depth selected by a user or practitioner.

[0234] Advantageously, enriching target sequences before sequencing nucleic acids significantly reduces the cost and time associated with sequencing, facilitates multiplex sequencing by allowing multiple samples to be mixed together for a single sequencing reaction, and significantly reduces the computational burden of aligning the resulting sequence reads as a result of significantly reducing the total amount of nucleic acid analyzed from each sample. Thus, in some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 1,000 times. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 100 times, at least 500 times, at least 1,000 times, at least 5,000 times, at least 10,000 times, at least 50,000 times, or more. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of 100,000 times or less, 50,000 times or less, 10,000 times or less, 5,000 times or less, or less. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 100x to 50,000x, 100x to 10,000x, 100x to 5000x, 100x to 1000x, or 100x to 500x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 500x to 50,000x, 500x to 10,000x, 500x to 5000x, or 500x to 1000x. In some embodiments, the panel enrichment sequencing reactions are performed at a read depth of 1000x to 50,000x, 1000x to 10,000x, or 1000x to 5000x.

[0235] In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of at least 1x. In some embodiments, the panel enrichment sequencing reaction is performed at a read depth of at least 2x, at least 3x, at least 4x, at least 5x, at least 10x, at least 25x, at least 50x, at least 100x, at least 250x, or more. In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of 1000x or less, 500x or less, 100x or less, 50x or less, or less. In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of 1x to 500x, 1x to 100x, or 1x to 50x. In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of 2.5x to 500x, 2.5x to 100x, or 2.5x to 50x. In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of 5x to 500x, 5x to 100x, or 5x to 50x. In some embodiments, the total cfDNA fragment sequencing reaction is performed at a read depth of 10x to 500x, 10x to 100x, or 10x to 50x.

[0236] Blocks 906 and 908. Referring to block 906, in some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches (includes probes for) 50 to 150 genes. Referring to block 908, in some embodiments, the multiple probe sequences used to enrich cell-free DNA fragments in the liquid biopsy sample in the panel enrichment sequencing reaction collectively map to 25 to 150 different genes in the human reference genome. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 50 to 150 genes, 100 to 200 genes, 150 to 300 genes, or 250 to 500 genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 50 to 1000 genes, 60 to 800 genes, 70 to 700 genes, 80 to 600 genes, or 90 to 500 genes. In some embodiments, each of the enriched genes in the sequencing panel is a human gene.

[0237] In some embodiments, the multiple probe sequences used in panel enrichment sequencing reaction to enrich cell-free DNA fragments in liquid biopsy samples are collectively mapped to at least 25 different genes in human reference genome.In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches at least 25 human genes, at least 50 human genes, at least 100 human genes, at least 250 human genes, at least 500 human genes, at least 1000 human genes, at least 2500 human genes, at least 5000 human genes or more.In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches 1,000 or less human genes, 500 or less human genes, 250 or less human genes, 200 or less human genes, 175 or less human genes, 100 or less human genes or less. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 1,000 human genes, 25 to 500 human genes, 10 to 250 human genes, 10 to 200 human genes, 5 to 150 human genes, or 5 to 100 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 400 human genes, 30 to 500 human genes, 50 to 300 human genes, 5 to 95 human genes, 15 to 130 human genes, or 15 to 165 human genes. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 human genes to 600 human genes, 40 human genes to 80 human genes, 35 human genes to 95 human genes, 45 human genes to 80 human genes, 20 human genes to 80 human genes, or 20 human genes to 120 human genes.

[0238] Blocks 910-912. Referring to block 910, in some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 100 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 1. In some embodiments, the sequencing panel enriches only for genes in Table 1, while in other embodiments, the sequencing panel enriches for some genes that are in Table 1 and some genes that are not in Table 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for between 25 different genes and 150 different genes listed in Table 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 150 genes, 100 genes to 200 genes, or 150 genes to 300 genes listed in Table 1.

[0239] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 100 genes, at least 10, 20, 30, 40, or 50 genes listed in Table 2. In some embodiments, the sequencing panel enriches only for genes in Table 2, while in other embodiments, the sequencing panel enriches for some genes that are in Table 2 and some genes that are not in Table 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for between 25 different genes and 150 different genes listed in Table 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 genes to 150 genes, 100 genes to 200 genes, or 150 genes to 300 genes listed in Table 2.

[0240] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set includes probes targeting one or more of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 5 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 10 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 25 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 50 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 75 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting at least 100 of the genes listed in Table 1. In some such embodiments, the probe set includes probes targeting all of the genes listed in Table 1.

[0241] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set includes probes targeting one or more of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 5 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 10 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 25 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 50 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 75 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting at least 100 of the genes listed in Table 2. In some such embodiments, the probe set includes probes targeting all of the genes listed in Table 2.

[0242] [Table 1]

[0243] [Table 2-1] [Table 2-2]

[0244] Block 912. Referring to block 912, in some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in List 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in List 1. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all genes listed in List 1, for 10 to all genes listed in List 1, or for 20 to all genes listed in List 1.

[0245] In some embodiments, the probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, the probe set consists of or includes probes targeting one or more of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 5 of the genes listed in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 10 of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 25 of the genes in List 1. In some such embodiments, the probe set consists of or includes probes targeting at least 50 of the genes listed in List 1. In some such embodiments, the probe set consists of or includes probes targeting all of the genes in List 1.

[0246] リスト1:AKT1(14q32.33)、ALK(2p23.2-23.1)、APC(5q22.2)、AR(Xq12)、ARAF(Xp11.3)、ARID1A(1p36.11)、ATM(11q22.3)、BRAF(7q34)、BRCA1(17q21.31)、BRCA2(13q13.1)、CCND1(11q13.3)、CCND2(12p13.32)、CCNE1(19q12)、CDH1(16q22.1)、CDK4(12q14.1)、CDK6(7q21.2)、CDKN2A(9p21.3)、CTNNB1(3p22.1)、DDR2(1q23.3)、EGFR(7p11.2)、ERBB2(17q12)、ESR1(6q25.1-25.2)、EZH2(7q36.1)、FBXW7(4q31.3)、FGFR1(8p11.23)、FGFR2(10q26.13)、FGFR3(4p16.3)、GATA3(10p14)、GNA11(19p13.3)、GNAQ(9q21.2)、GNAS(20q13.32)、HNF1A(12q24.31)、HRAS(11p15.5)、IDH1(2q34)、IDH2(15q26.1)、JAK2(9p24.1)、JAK3(19p13.11)、KIT(4q12)、KRAS(12p12.1)、MAP2K1(15q22.31)、MAP2K2(19p13.3)、MAPK1(22q11.22)、MAPK3(16p11.2)、MET(7q31.2)、MLH1(3p22.2)、MPL(1p34.2)、MTOR(1p36.22)、MYC(8q24.21)、NF1(17q11.2)、NFE2L2(2q31.2)、NOTCH1(9q34.3)、NPM1(5q35.1)、NRAS(1p13.2)、NTRK1(1q23.1)、NTRK3(15q25.3)、PDGFRA(4q12)、PIK3CA(3q26.32)、PTEN(10q23.31)、PTPN11(12q24.13)、RAF1(3p25.2)、RB1(13q14.2)、RET(10q11.21)、RHEB(7q36.1)、RHOA(3p21.31)、RIT1(1q22)、ROS1(6q22.1)、SMAD4(18q21.2)、SMO(7q32.1)、STK11(19p13.3)、TERT(5p15.33)、TP53(17p13.1), TSC1 (9q34.13), and VHL (3p25.3).

[0247] Block 914. Referring to block 914, in some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in List 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in List 2. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all genes listed in List 2, for 10 to all genes listed in List 2, or for 20 to all genes listed in List 2.

[0248] In some embodiments, the probe set includes probes targeting no more than 50 genes, no more than 100 genes, no more than 150 genes, or no more than 200 genes. In some such embodiments, the probe set consists of or includes probes targeting one or more of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 5 of the genes listed in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 10 of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 25 of the genes in List 2. In some such embodiments, the probe set consists of or includes probes targeting at least 50 of the genes listed in List 2. In some such embodiments, the probe set consists of or includes probes targeting all of the genes in List 2.

[0249] <h2 style=";text-align:left;direction:ltr">Lineage 2: ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1(FAM123B), APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AX L、BAP1、BARD1、BCL2、BCL2L1、BCL2L2、BCL6、BCOR、BCORL1、BRAF、BRCA1、BR CA2、BRD4、BRIP1、BTG1、BTG2、BTK、C11orf30(EMSY)、C17orf39(GID4)、CAL R、CARD11、CASP8、CBFB、CBL、CCND1、CCND2、CCND3、CCNE1、CD22、CD274(PD- L1)、CD70、CD79A、CD79B、CDC73、CDH1、CDK12、CDK4、CDK6、CDK8、CDKN1A、CD KN1B、CDKN2A、CDKN2B、CDKN2C、CEBPA、CHEK1、CHEK2、CIC、CREBBP、CRKL、CS F1R、CSF3R、CTCF、CTNNA1、CTNNB1、CUL3、CUL4A、CXCR4、CYP17A1、DAXX、DDR1 、DDR2、DIS3、DNMT3A、DOT1L、EED、EGFR、EP300、EPHA3、EPHB1、EPHB4、ERBB2 、ERBB3、ERBB4、ERCC4、ERG、ERRFI1、ESR1、EZH2、FAM46C、FANCA、FANCC、FAN CG、FANCL、FAS、FBXW7、FGF10、FGF12、FGF14、FGF19、FGF23、FGF3、FGF4、FGF 6、FGFR1、FGFR2、FGFR3、FGFR4、FH、FLCN、FLT1、FLT3、FOXL2、FUBP1、GABRA6 GATA3, GATA4, GATA6, GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A, KMT2D(MLL2), KRAS, LTK, LYN, MAF, MAP2K1(MEK1),<h2 style=";text-align:left;direction:ltr">MAP2K2(MEK2)、MAP2K4、MAP3K1、MAP3K13、MAPK1、MCL1、MDM2、MDM 4、MED12、MEF2B、MEN1、MERTK、MET、MITF、MKNK1、MLH1、MPL、MRE11 A、MSH2、MSH3、MSH6、MST1R、MTAP、MTOR、MUTYH、MYC、MYCL(MYCL1) 、MYCN、MYD88、NBN、NF1、NF2、NFE2L2、NFKBIA、NKX2-1、NOTCH1、NOT CH2、NOTCH3、NPM1、NRAS、NSD3(WHSC1L1)、NT5C2、NTRK1、NTRK2、NTRK3、P2RY8、PALB2、PARK2、PARP1、PARP2、PARP3、PAX5、PBRM1、PDC D1(PD-1)、PDCD1LG2(PD-L2)、PDGFRA、PDGFRB、PDK1、PIK3C2B、PIK3C2G、PIK3CA、PIK3CB、PIK3R1、PIM1、PMS2、POLD1、POLE、PPARG、P PP2R1A、PPP2R2A、PRDM1、PRKAR1A、PRKCI、PTCH1、PTEN、PTPN11、P TPRO、QKI、RAC1、RAD21、RAD51、RAD51B、RAD51C、RAD51D、RAD52、RA D54L、RAF1、RARA、RB1、RBM10、REL、RET、RICTOR、RNF43、ROS1、RPT OR、SDHA、SDHB、SDHC、SDHD、SETD2、SF3B1、SGK1、SMAD2、SMAD4、SMA RCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, ncRNA, TGFBR2, TIPARP, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WT1, XPO1, XRCC2, ZNF217, ZNF703.<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">

[0250] <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr">ブロック914。<h2 style=";text-align:left;direction:ltr"> Referring to block 914, in some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in FIG. 10 (any combination of FIGS. 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M). In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 50 different genes listed in FIG. 10. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 5 to all genes listed in FIG. 10, for 10 to all genes listed in FIG. 10, or for 20 to all genes listed in FIG. 10. While FIG. 10 and List 2 provide the same genes, in preferred embodiments, FIG. 10 indicates the types of variants that are such genes.

[0251] In some embodiments, a probe set includes probes targeting 50 or fewer genes, 100 or fewer genes, 150 or fewer genes, or 200 or fewer genes. In some such embodiments, a probe set consists of or includes probes targeting one or more of the genes in FIG. 10. In some such embodiments, a probe set consists of or includes probes targeting at least five of the genes listed in FIG. 10. In some such embodiments, a probe set consists of or includes probes targeting at least 10 of the genes in FIG. 10. In some such embodiments, a probe set consists of or includes probes targeting at least 25 of the genes in FIG. 10. In some such embodiments, a probe set consists of or includes probes targeting at least 50 of the genes listed in FIG. 10. In some such embodiments, a probe set consists of or includes probes targeting all of the genes in FIG. 10.

[0252] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10, 20, 30, 40, or 50 genes listed in any of Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M.

[0253] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M.

[0254] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 genes listed in any of Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M.

[0255] In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 25 to 150 different genes listed in any of Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M. In some embodiments, the panel enrichment sequencing reaction uses a sequencing panel that enriches for 50 to 150 genes, 100 to 200 genes, or 150 to 300 genes listed in any of Figures 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10H, 10I, 10j, 10K, 10L, and 10M.

[0256] Generally, probes for enrichment of nucleic acids (e.g., cfDNA obtained from a liquid biopsy sample) comprise DNA, RNA, or modified nucleic acid structures having a base sequence complementary to a locus of interest. For example, a probe designed to hybridize to a locus in a cfDNA molecule can comprise a sequence complementary to either strand, since cfDNA molecules are double-stranded. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 consecutive bases of the locus of interest. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 20, 25, 30, 40, 50, 75, 100, 150, 200, or more consecutive bases of the locus of interest.

[0257] Target panel provides several benefits for nucleic acid sequencing.For example, in some embodiments, for example, the algorithm for distinguishing between a first cancer state and a second cancer state can be trained with a smaller and more informative data set (for example, fewer genes), which leads to more computationally efficient training of the classifier for distinguishing between a first cancer state and a second cancer state.This improvement in computational efficiency due to the reduction in the size of the gene set for distinguishing can be advantageously used to accelerate classifier training, or can be used to improve the performance of such classifier (for example, through more extensive training of classifier).

[0258] In some embodiments, the gene panel is a whole-exome panel that analyzes the exome of a biological sample. In some embodiments, the gene panel is a whole-genome panel that analyzes the genome of a specimen. In some preferred embodiments, the gene panel is optimized for use with liquid biopsy samples (e.g., to provide clinical decision support for solid tumors). See, for example, Table 1 above.

[0259] In some embodiments, the probe comprises an additional nucleic acid sequence that does not share any homology with the locus of interest. For example, in some embodiments, the probe also comprises an identifier sequence, e.g., a nucleic acid sequence comprising a unique molecular identifier (UMI), that is specific to a particular sample or subject. For examples of identifier sequences, see, e.g., Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, which are incorporated herein by reference. Similarly, in some embodiments, the probe also comprises a primer nucleic acid sequence useful for amplifying the nucleic acid molecule of interest, e.g., using PCR. In some embodiments, the probe also comprises a capture sequence designed to hybridize to an anti-capture sequence to retrieve the nucleic acid molecule of interest from the sample.

[0260] Similarly, in some embodiments, each probe comprises a non-nucleic acid affinity moiety covalently bound to a nucleic acid molecule complementary to a target locus for retrieving the target nucleic acid molecule. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, such as a dipstick or magnetic bead, for retrieving the target nucleic acid. In some embodiments, the methods described herein include amplifying the nucleic acid bound to the probe set before further analysis, such as sequencing. Methods for amplifying nucleic acids, for example, by PCR, are well known in the art.

[0261] Next-generation sequencing generates millions of short reads (e.g., sequence reads) for each biological sample. Thus, in some embodiments, the multiple sequence reads obtained by next-generation sequencing of cfDNA molecules are DNA sequence reads. In some embodiments, the sequence reads have an average length of at least 50 nucleotides. In other embodiments, the sequence reads have an average length of at least 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, or more nucleotides.

[0262] In some embodiments, sequencing is performed after enriching nucleic acids (e.g., cfDNA, gDNA, and / or RNA) containing multiple predetermined target sequences, e.g., human genes and / or non-coding sequences associated with cancer. Advantageously, sequencing a nucleic acid sample enriched for target nucleic acids, rather than all nucleic acids isolated from a biological sample, significantly reduces the average time and cost of the sequencing reaction. Thus, in some preferred embodiments, the methods described herein include obtaining multiple sequence reads of nucleic acids hybridized to a probe set for hybrid capture enrichment (e.g., for one or more genes listed in Table 1, or one or more genes listed in Table 2, one or more genes listed in List 1, one or more genes listed in List 2, or one or more genes listed in Figure 10).

[0263] In some embodiments, panel targeted sequencing is performed to an average on-target depth of at least 500x, at least 750x, at least 1000x, at least 2500x, at least 500x, at least 10,000x, or more. In some embodiments, samples are further evaluated for uniformity above a sequencing depth threshold (e.g., 95% of all target base pairs at 300x sequencing depth). In some embodiments, the sequencing depth threshold is a minimum depth selected by a user or practitioner.

[0264] In some embodiments, sequence reads are obtained by whole genome or whole exome sequencing methodologies. In some such embodiments, whole exome capture is performed using an automated system that uses a liquid-handling robot (e.g., SciClone NGSx). Because many loci are being sequenced, whole genome sequencing, and to some extent whole exome sequencing, is typically performed at a lower sequencing depth than smaller, targeted panel sequencing reactions. For example, in some embodiments, whole genome or whole exome sequencing is performed to an average sequencing depth of at least 3x, at least 5x, at least 10x, at least 15x, at least 20x, or more. In some embodiments, low-pass whole genome sequencing (LPWGS) technology is used for whole genome or whole exome sequencing. LPWGS is typically performed to an average sequencing depth of about 0.25x to about 5x, more typically to an average sequencing depth of about 0.5x to about 3x.

[0265] Due to differences in sequencing methodology, data obtained from targeted panel sequencing may be more suitable for certain analyses than data obtained from whole genome / whole exome sequencing, and vice versa. For example, due to the higher sequencing depth achieved by targeted panel sequencing, the resulting sequence data is more suitable for identifying variant alleles present at low allele rates in samples, for example, less than 20%. In contrast, data generated from whole genome / whole exome sequencing is more suitable for estimating whole genome metrics, such as tumor mutation burden, because the entire genome is well represented by the sequencing data. Therefore, in some embodiments, a nucleic acid sample, such as cfDNA, gDNA, or mRNA sample, is evaluated using both targeted panel sequencing and whole genome / whole exome sequencing (e.g., LPG Seq).

[0266] In some embodiments, raw sequence reads resulting from a sequencing reaction are output from the sequencer in a native file format, e.g., a BCL file. In some embodiments, the native file is passed directly to a bioinformatics pipeline (e.g., variant analysis 206), the components of which are described in detail below. In other embodiments, preprocessing is performed before passing the sequences to the bioinformatics platform. For example, in some embodiments, the format of the sequence read file is converted from the native file format (e.g., BCL) to a file format compatible with one or more algorithms used in the bioinformatics pipeline (e.g., FASTQ or FASTA). In some embodiments, the raw sequence reads are filtered to remove sequences that do not meet one or more quality thresholds. In some embodiments, raw sequence reads generated from the same unique nucleic acid molecule in a sequencing read are collated into a single sequence read representing that molecule, e.g., using the UMI described above. In some embodiments, one or more of these preprocessing activities are performed within the bioinformatics pipeline itself.

[0267] In one example, a sequencer can generate a BCL file. The BCL file can contain raw image data for multiple patient samples to be sequenced. The BCL image data is an image of the flow cell for each cycle during sequencing. Cycles can be implemented by illuminating the patient sample with specific wavelengths of electromagnetic radiation to generate multiple images that can be processed into base calls via a BCL-to-FASTQ processing algorithm that identifies which base pairs are present in each cycle. The resulting FASTQ file contains the entirety of the reads for each patient sample, paired with a quality metric ranging from 0 to 64, for example, where 64 is the best quality and 0 is the worst quality. In embodiments where both liquid biopsy samples and normal tissue samples are sequenced, the sequence reads in the corresponding FASTQ files can be matched so that a liquid biopsy-normal analysis can be performed.

[0268] The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants or copy number variations are present in a sample. Each FASTQ file contains reads, which may be paired-end or single reads and may be short or long reads. Each read represents the sequence of one detected nucleotide in a nucleic acid molecule or copy of a nucleic acid molecule isolated from a patient sample, as detected by a sequencer. Each read in a FASTQ file is also associated with a quality assessment. The quality assessment may reflect the likelihood of an error occurring during the sequencing procedure that affected the associated read. In some embodiments, the paired-end sequencing results for each isolated nucleic acid sample are contained in a pair of FASTQ files, separated for efficiency. Thus, in some embodiments, the forward (read 1) and reverse (read 2) sequences for each isolated nucleic acid sample are stored separately, but in the same order and under the same identifier.

[0269] In various embodiments, the bioinformatics pipeline can filter the FASTQ data from the corresponding sequence data file for each respective biological sample. Such filtering can include correcting or masking sequencer errors, and removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, over-represented sequences, biases caused by library preparation, amplification, or capture, and other errors.

[0270] Although workflow 200 illustrates obtaining a biological sample, extracting nucleic acids from the biological sample, and sequencing the isolated nucleic acids, in some embodiments, the sequencing data used in the improved systems and methods described herein (e.g., including improved methods for determining accurate circulating tumor fraction estimates) is obtained by receiving previously generated sequence reads in electronic format.

[0271] In some embodiments, sequencing of the plurality of cell-free nucleic acids in the subject's liquid biopsy sample is performed at a central laboratory or sequencing facility. In some such embodiments, the method includes accessing one or more sequencing datasets and / or one or more auxiliary files in electronic format via a cloud-based interface. For example, the datasets can be obtained by running a bioinformatics pipeline using a tumor BAM file, a normal BAM file, a human reference genome file, a target region BED file, a list of mappable regions of the genome, and / or a blacklist of recurrent problem regions of the genome.

[0272] In some embodiments, obtaining the dataset includes accessing the dataset in electronic format via a cloud-based interface. For example, the dataset can include one or more outputs from a bioinformatics pipeline (e.g., CNVkit output ".cns" and / or ".cnr").

[0273] Additional methods and embodiments for sequencing nucleic acids, including alignment and pre-processing sequence reads, are described in more detail above (see exemplary method: Figure 2A: Example of a workflow for precision oncology). Additional methods and embodiments for implementing the methods of the present disclosure in distributed diagnostic and clinical environments are described in more detail above (see exemplary method: Figure 2B: Distributed diagnostic and clinical environments). As will be apparent to those skilled in the art, other embodiments and / or any combination, substitution, addition, or deletion thereof are possible.

[0274] In some embodiments, the subject is a patient with cancer. In some such embodiments, the cancer is a solid tumor cancer. In some embodiments, the cancer is selected from the group consisting of ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, cholangiocarcinoma, adrenal cancer, neural carcinoma, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, tumor of unknown primary, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, gastric cancer, bone cancer, non-small cell lung cancer, thymoma, low grade glioma, thyroid cancer ... In some embodiments, the subject is a patient undergoing a clinical trial, the cancer being studied is selected from the group consisting of glioma, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer, head and neck squamous cell carcinoma, meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, HER2-negative breast cancer, solid tumor, serous ovarian cancer, HR+ breast cancer, serous uterine cancer, endometrial cancer, endometrial cancer, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, or papillary renal cell carcinoma.

[0275] In some embodiments, the sequencing data is processed (e.g., using sequence data processing module 141) to prepare data for genomic feature identification 385. For example, in some embodiments described above, the sequencing data is in the native file format provided by the sequencer. Thus, in some embodiments, the system (e.g., system 100) applies preprocessing algorithms 142 to convert the file format into one recognized by one or more upstream processing algorithms (318). For example, BCL file output from a sequencer can be converted to FASTQ file format using bcl2fastq or bcl2fastq2 conversion software (Illumina®). The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants, copy number changes, etc., are present in the sample.

[0276] In some embodiments, other preprocessing functions are performed, such as filtering sequence reads 122 based on desired quality, e.g., size and / or quality of base calls. In some embodiments, quality control checks are performed to ensure that data is sufficient for variant calling. For example, entire reads, individual nucleotides, or multiple nucleotides likely to have errors can be discarded based on a quality assessment associated with the read in the FASTQ file, the known error rate of the sequencer, and / or a comparison between each nucleotide in the read and one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools, such as Skewer. See Jiang et al., 2014, BMC Bioinformatics 15(182):1-12. FASTQ files can be analyzed for quality control and rapid assessment of reads by sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC, or another similar software program. For paired-end reads, the reads can be merged.

[0277] In some embodiments, when both a liquid biopsy sample and a normal tissue sample from a patient are sequenced, two FASTQ output files are generated: one for the liquid biopsy sample and one for the normal tissue sample. A "matched" (e.g., panel-specific) workflow is implemented to jointly analyze the matched liquid biopsy-normal FASTQ files. When a matched normal sample is not available from the patient, the FASTQ file from the liquid biopsy sample is analyzed in "tumor-only" mode. See, for example, Figure 4B. When two or more patient samples, such as a liquid biopsy sample and a normal tissue sample, are processed simultaneously on the same sequencer flow cell, differences in the sequence of the adapters used for each patient sample barcode the nucleic acids extracted from both samples, associating each read with the correct patient sample and facilitating assignment to the correct FASTQ file.

[0278] For efficiency, in some embodiments, the paired-end sequencing results for each isolate are contained in a separate pair of FASTQ files. The forward (read 1) and reverse (read 2) sequences for each tumor and normal isolate are stored separately but in the same order and under the same identifier. See, for example, Figure 4C. In various embodiments, the bioinformatics pipeline may filter the FASTQ data from each isolate. Such filtering may include correcting or masking sequencer errors, as well as removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, overrepresented sequences, biases caused by library preparation, amplification, or capture, and other errors. See, for example, Figure 4D.

[0279] Similarly, in some embodiments, sequencing (312) is performed on a pool of nucleic acid sequencing libraries prepared from different biological samples, e.g., from the same or different patients. Accordingly, in some embodiments, the system demultiplexes (320) the data (e.g., using a demultiplexing algorithm 144) to separate sequence reads into separate files for each sequencing library included in the sequencing pool, e.g., based on UMI or patient identifier sequences added to the nucleic acid fragments during sequencing library preparation, as described above. In some embodiments, the demultiplexing algorithm is part of the same software package as one or more preprocessing algorithms 142. For example, bcl2fastq or bcl2fastq2 conversion software (Illumina®) includes instructions for both converting the native file format output from the sequencer and demultiplexing the sequence reads 122 output from the reaction.

[0280] The sequence reads are then aligned (322) to a reference sequence construct 158, such as a reference genome, reference exome, or other reference construct prepared for a particular targeted panel sequencing reaction, using, for example, an alignment algorithm 143. For example, in some embodiments, individual sequence reads 123 in electronic form (e.g., in a FASTQ file) are aligned to a reference sequence construct for the species of interest (e.g., a reference human genome) by identifying the sequence in the region of the reference sequence construct that best matches the sequence of nucleotides in the sequence read. In some embodiments, the sequence reads are aligned to the reference exome or reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information may also include the sequence read length, which may be determined from the start and end positions. The region in the reference genome may relate to a gene or a segment of a gene. Any of a variety of alignment tools can be used for this task.

[0281] For example, local sequence alignment algorithms compare different lengths of subsequences in query sequence (e.g., sequence reads) with subsequences in target sequence (e.g., reference constructs) to create the best alignment for each part of query sequence.In contrast, global sequence alignment algorithms align the entire sequence, for example, end to end.Examples of local sequence alignment algorithms include the Smith-Waterman algorithm.

[0282] In some embodiments, the read mapping process begins by constructing an index of either the reference genome or the read, which is then used to retrieve a set of positions in the reference sequence to which the read is more likely to align. Once this subset of possible mapping positions is identified, alignment is performed within these candidate regions using slower, more sensitive algorithms. See, for example, Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14:184, and Flicek and Birney, 2009, "Sense from sequence reads: methods for alignment and assembly," Nat Methods 6 (Suppl. 11), S6-S12, each of which is incorporated herein by reference. In some embodiments, the mapping tool methodology utilizes a hash table or the Burrows-Wheeler transform (BWT). See, for example, Li and Homer, 2010, “A survey of sequence alignment algorithms for next-generation sequencing,” Brief Bioinformatics 11, pp. 473-483, which is incorporated herein by reference.

[0283] Other software programs designed to align reads include, for example, Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and / or programs using the Smith-Waterman algorithm. Candidate reference genomes include, for example, HG19, GRCh38, hg38, GRCh37, and / or other reference genomes developed by the Genome Reference Consortium. In some embodiments, the alignment generates a SAM file, which stores the start and end positions of each read, along with its coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome.

[0284] For example, in some embodiments, each read in the FASTQ file is aligned to the location in the human genome that best matches the sequence of nucleotides in the read. Many software programs designed to align reads exist, including Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and programs using the Smith-Waterman algorithm. Alignment can be directed using a reference genome (e.g., HG19, GRCh38, HG38, GRCh37, or other reference genomes developed by the Genome Reference Consortium) by comparing the nucleotide sequence in each read to portions of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that most likely corresponds to the sequence in the read. In some embodiments, one or more SAM files are generated for the alignment, which store the start and end positions of each read according to their coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. The SAM file can be converted to a BAM file. In some embodiments, the BAM file is sorted, and duplicate reads are marked for deletion, resulting in a de-duplicated BAM file.

[0285] In some embodiments, adapter-trimmed FASTQ files are aligned to the 19th edition of the Human Reference Genome Build (HG19). After alignment, reads are grouped by alignment position and UMI family and corrected to a consensus sequence. Bases with insufficient quality or significant discrepancies between family members (e.g., when it is uncertain whether the base is adenine, cytosine, guanine, etc.) can be replaced by N to represent the wild-type nucleotide type. PHRED scores are then scaled based on the initial base call estimates combined across all family members. After single-strand consensus generation, a double-stranded consensus sequence is generated by comparing the forward and reverse PCR products with the mirrored UMI sequence. In various embodiments, a consensus can be generated across read pairs; otherwise, a single-stranded consensus call is used. After consensus calling, filtering is performed to remove low-quality consensus fragments. The consensus fragments are then realigned to the human reference genome using BWA. A BAM output file is generated after the realignment, which is then sorted and indexed by alignment position.

[0286] In some embodiments, when both a liquid biopsy sample and a normal tissue sample are analyzed, this process generates a liquid biopsy BAM file (e.g., liquid BAM 124-1-i-cf) and a normal BAM file (e.g., germline BAM 124-1-ig), as illustrated in Figure 4A. In various embodiments, the BAM files can be analyzed to detect genetic variants and other genetic features, including single nucleotide variants (SNVs), copy number variants (CNVs), gene rearrangements, etc.

[0287] In some embodiments, the sequencing data is normalized to account for, for example, pulldown, amplification, and / or sequencing bias (e.g., mappability, GC bias, etc.).

[0288] In some embodiments, the SAM file generated after alignment is converted to a BAM file 124. Thus, after preprocessing the sequencing data generated for the pooled sequencing reactions, a BAM file is generated for each sequencing library present in the master sequencing pool. For example, as illustrated in FIG. 4A, separate BAM files are generated for each of three samples obtained from subject 1 at time i (e.g., tumor BAM 124-1-it corresponding to the alignment of sequence reads of nucleic acids isolated from solid tumor samples from subject 1, liquid BAM 124-1-i-cf corresponding to the alignment of sequence reads of nucleic acids isolated from liquid biopsy samples from subject 1, and germline BAM 124-1-ig corresponding to the alignment of sequence reads of nucleic acids isolated from normal tissue samples from subject 1), and one or more samples obtained from one or more additional subjects at time j (e.g., tumor BAM 124-2-jt corresponding to the alignment of sequence reads of nucleic acids isolated from solid tumor samples from subject 2). In some embodiments, BAM files are sorted and duplicate reads are marked for removal, resulting in a de-duplicated BAM file, for example, a tool such as SamBAMBA marks and filters duplicate alignments in the sorted BAM file.

[0289] Many of the embodiments described below in conjunction with Figure 4 relate to analyses performed using sequencing data derived from cancer patient cfDNA, e.g., obtained from the patient's liquid biopsy sample. Generally, these embodiments are independent and therefore do not rely on any particular sequencing data generation method, e.g., sample preparation, sequencing, and / or data preprocessing methodology. However, in some embodiments, the methods described herein include one or more features 204 of generating sequencing data, as illustrated in Figures 2A and 3.

[0290] The alignment file (e.g., BAM file 124) prepared as described above is then passed to a feature extraction module 145, where the sequences are analyzed to identify genomic alterations (e.g., SNV / MNV, indels, genomic rearrangements, copy number variations, etc.) and / or determine various characteristics of the patient's cancer (e.g., MSI status, TMB, tumor ploidy, HRD status, tumor fraction, tumor purity, methylation patterns, etc.) (324). Many software packages for identifying genomic alterations are known in the art. Generally, these software packages identify variants in the sorted SAM or BAM file 124 relative to one or more reference sequence constructs 158. The software package then outputs a file, e.g., a raw VCF (variant call format), that lists the called variants (e.g., genomic features 131) and identifies positions relative to the reference sequence construct (e.g., where the sequence of the sample nucleic acid differs from the corresponding sequence in the reference construct). In some embodiments, the system 100 digests the contents of the native output file and inputs feature data 125 into the test patient data store 120. In other embodiments, the native output file serves as a record of these genomic features 131 in the test patient data store 120.

[0291] In general, the systems described herein can employ any combination of available variant calling software packages and internally developed variant identification algorithms. In some embodiments, the output of a particular algorithm in a variant calling software is further evaluated, for example, to improve variant identification. Thus, in some embodiments, system 100 employs an available variant calling software package to implement some or all of the functionality of one or more of the algorithms shown in feature extraction module 145.

[0292] In some embodiments, as illustrated in FIG. 1A, separate algorithms (or the same algorithm implemented with different parameters) are applied to identify variants specific to a patient's cancer genome and variants present in the subject's germline. In other embodiments, variants are identified randomly and later classified as either germline or somatic, e.g., based on sequencing data, population data, or a combination thereof. In some embodiments, variants are classified as germline variants and / or non-actionable variants when they are represented in a population above a threshold level determined using a population database, e.g., ExAC or gnomAD. For example, in some embodiments, variants represented in at least 1% of alleles in a population are annotated as germline and / or non-actionable. In other embodiments, variants represented in at least 2%, at least 3%, at least 4%, at least 5%, at least 7.5%, at least 10%, or more of the alleles in a population are annotated as germline and / or non-actionable. In some embodiments, sequencing data from a matched sample from a patient, e.g., a normal tissue sample, is used to annotate variants identified in a cancer sample from a subject, i.e., variants present in both the cancer sample and the normal sample represent variants that were present in the germline before the patient developed cancer and can be annotated as germline variants.

[0293] In various embodiments, the detected genetic variants and genetic features are analyzed as a form of quality control. For example, a pattern of detected genetic variants or features may indicate problems related to the sample, the sequencing procedure, and / or the bioinformatics pipeline (e.g., sample contamination, mislabeling of the sample, changes in reagents, changes in the sequencing procedure and / or the bioinformatics pipeline, etc.).

[0294] FIG. 4E illustrates an exemplary workflow for genomic feature identification (324). This particular workflow is merely an example of one possible set and arrangement of algorithms for feature extraction from sequencing data 124. Generally, any combination of modules and algorithms, for example, those illustrated in FIG. 1A , of feature extraction module 145, can be used in a bioinformatics pipeline, particularly a bioinformatics pipeline for analyzing liquid biopsy samples. For example, in some embodiments, an architecture useful in the methods and systems described herein includes at least one of the modules or variant calling algorithms shown in feature extraction module 145. In some embodiments, an architecture includes at least two, three, four, five, six, seven, eight, nine, ten, or more of the modules or variant calling algorithms shown in feature extraction module 145. Furthermore, in some embodiments, feature extraction modules and / or algorithms not illustrated in FIG. 1A are utilized in the methods and systems described herein.

[0295] In some embodiments, the methods and systems described herein use somatic mutations identified in a separate workstream of a liquid biopsy assay. For example, Figures 4F2 and 4G (4G1-4G3) illustrate methods 400-2 and 450, respectively, which use a dynamic variant count threshold to identify somatic sequence variants from cfDNA sequencing data. In some embodiments, the somatic variants so identified are used to determine ITMB using the methods, systems, and CRMs described herein. For details of methods for identifying somatic mutations from cfDNA samples, see, for example, U.S. Patent No. 11,475,981, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0296] In one aspect, the present disclosure provides a method, as well as systems and CRMs for performing all or part of such a method, for obtaining a plurality of nucleic acid sequences from a panel enrichment sequencing reaction, the plurality of nucleic acid sequences including corresponding sequences for each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from a subject, such as sequencing data 122-1 described herein with respect to sequencing 312 performed during the wet-lab portion 204 of the exemplary liquid biopsy scheme shown in FIG.

[0297] In some embodiments, each cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to each probe sequence in the plurality of probe sequences, and the probe sequences are used to enrich cell-free DNA fragments in a liquid biopsy sample in a panel enrichment sequencing reaction. Examples of gene sets targeted by such probes are described, for example, in Table 1, Table 2, List 1, List 2, and Figure 10 provided herein, and Figure 6 in PCT Patent Application Publication WO2023 / 164713. This document is incorporated herein by reference in its entirety for all purposes. In some embodiments, the plurality of probe sequences maps to 150 or fewer genes in the human genome. In some embodiments, the plurality of probe sequences maps to 150 or fewer genes.

[0298] Block 916. Referring to block 916, in some embodiments, the liquid biopsy sample is a blood sample. For example, in some embodiments, the liquid biopsy sample comprises blood, whole blood, peripheral blood, plasma, serum, or lymph of the subject. In some alternative embodiments, the liquid biopsy sample is any of the embodiments described above (see definition: liquid biopsy and / or example method: Figure 2A: Example Precision Oncology Workflow).

[0299] In some embodiments, the liquid biopsy sample is blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal material, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid from a subject. In some embodiments, the liquid biopsy sample is a cell-free sample, such as a cell-free blood sample. In some embodiments, the liquid biopsy sample is obtained from a subject with cancer. In some embodiments, the liquid biopsy sample is collected from a subject whose cancer status is unknown. In some embodiments, the liquid biopsy is collected from a subject with a non-cancerous disorder, e.g., cardiovascular disease. In some embodiments, the liquid biopsy is collected from a subject whose status is unknown for a non-cancerous disorder.

[0300] In some embodiments, one or more of the biological samples obtained from the patient are biological liquid samples, also referred to as liquid biopsy samples. In some embodiments, one or more of the biological samples obtained from the patient are selected from blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testes), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple effusion, aspirates from different parts of the body (e.g., thyroid, breast), and the like. In some embodiments, the liquid biopsy sample comprises blood and / or saliva. In some embodiments, the liquid biopsy sample is peripheral blood. In some embodiments, the blood sample is collected from the patient in a commercially available blood collection container, e.g., using a PAXGENE® Blood DNA Tube. In some embodiments, the saliva sample is collected from the patient in a commercially available saliva collection container, e.g., using an ORAGENE® DNA Saliva Kit.

[0301] In some embodiments, the liquid biopsy sample has a volume of about 1 mL to about 50 mL. For example, in some embodiments, the liquid biopsy sample has a volume of about 1 mL, about 2 mL, about 3 mL, about 4 mL, about 5 mL, about 6 mL, about 7 mL, about 8 mL, about 9 mL, about 10 mL, about 11 mL, about 12 mL, about 13 mL, about 14 mL, about 15 mL, about 16 mL, about 17 mL, about 18 mL, about 19 mL, about 20 mL, or more.

[0302] Liquid biopsy samples contain cell-free nucleic acids, including cell-free DNA (cfDNA).As described above, cfDNA isolated from cancer patients includes DNA derived from cancer cells, also referred to as circulating tumor DNA (ctDNA), cfDNA derived from germline (e.g., healthy or non-cancerous) cells, and cfDNA derived from hematopoietic cells (e.g., leukocytes).The relative proportion of cancerous and non-cancerous cfDNA present in liquid biopsy samples varies depending on the characteristics of the patient's cancer (e.g., type, stage, lineage, genomic profile, etc.).As used herein, the "tumor burden" of a subject refers to the percentage of cfDNA derived from cancer cells.

[0303] As described herein, cfDNA is a particularly useful source of biological data for various embodiments of the methods and systems described herein because it can be easily obtained from various bodily fluids. Advantageously, the use of bodily fluids facilitates continuous monitoring due to the ease of collection, as these fluids can be collected by non-invasive or minimally invasive methodologies. This is in contrast to methods that rely on solid tissue samples, such as biopsies, which often require invasive surgical procedures. Furthermore, because bodily fluids such as blood circulate throughout the body, cfDNA populations represent samples of many different tissue types from many different locations.

[0304] In some embodiments, a liquid biopsy sample is separated into two different samples, for example, in some embodiments, a blood sample is separated into a plasma sample containing cfDNA and a buffy coat preparation containing white blood cells.

[0305] In some embodiments, multiple liquid biopsy samples are obtained from each subject at intervals over a period of time (e.g., using serial testing). For example, in some such embodiments, the time between obtaining liquid biopsy samples from each subject is at least 1 day, at least 2 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 6 months, or at least 1 year.

[0306] In some alternative embodiments, the one or more biological samples collected from the patient are solid tissue samples, such as solid tumor samples or solid normal tissue samples. Methods for obtaining solid tissue samples, e.g., of cancer and / or normal tissues, are known in the art and depend on the type of tissue being sampled. For example, bone marrow biopsy and isolation of circulating tumor cells can be used to obtain samples of blood cancers; endoscopic biopsy can be used to obtain samples of gastrointestinal, bladder, and lung cancers; needle biopsy (e.g., fine needle aspiration, core needle aspiration, vacuum-assisted biopsy, image-guided biopsy) can be used to obtain samples of subcutaneous tumors; skin biopsy (e.g., shave biopsy, punch biopsy, incisional biopsy, and excision biopsy) can be used to obtain samples of skin cancers; and surgical biopsy can be used to obtain samples of cancers affecting the patient's internal organs. In some embodiments, the solid tissue sample is formalin-fixed tissue (FFPE). In some embodiments, the solid tissue sample is grossly dissected formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the solid tissue sample is a fresh frozen tissue sample.

[0307] In some embodiments, a dedicated normal sample is collected from the patient for simultaneous processing with the liquid biopsy sample. Generally, the normal sample is of non-cancerous tissue and can be collected using any of the tissue collection methods described above. In some embodiments, oral cells collected from the inside of the patient's cheek are used as the normal sample. Oral cells can be collected by placing an absorbent material, such as a cotton swab, into the subject's mouth and rubbing it against the cheek for, for example, at least 15 seconds or at least 30 seconds. The swab is then removed from the patient's mouth and inserted into a tube so that the tip of the tube is immersed in a liquid that serves to extract the oral cells from the absorbent material. An example of an oral cell recovery and collection device is provided in U.S. Patent No. 9,138,205, the contents of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, oral swab DNA is used as a source of normal DNA in circulating hematologic malignancies.

[0308] Referring to FIG. 2 , in some embodiments, biological samples collected from a patient are optionally sent to various analytical environments (e.g., sequencing lab 230, pathology lab 240, and / or molecular biology lab 250) for processing (e.g., data collection) and / or analysis (e.g., feature extraction). Wet-lab processing 204 may include sample cataloging (e.g., enrollment), clinical characterization of one or more samples (e.g., pathology review), and nucleic acid sequence analysis (e.g., extraction, library preparation, capture+hybridization, pooling, and sequencing). In some embodiments, the workflow includes clinical analysis of one or more biological samples collected from the subject, for example, in pathology lab 240 and / or molecular and cell biology lab 250, to generate clinical features such as pathology features 128-3, image data 128-3, and / or tissue culture / organoid data 128-3.

[0309] In some embodiments, pathology data 128-1 collected during a clinical evaluation includes, for example, visual features identified by a pathologist's examination of a specimen (e.g., a solid tumor biopsy) on a stained H&E or IHC slide. In some embodiments, the sample is a solid tissue biopsy sample. In some embodiments, the tissue biopsy sample is formalin-fixed tissue (FFT), e.g., formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the tissue biopsy sample is an FFPE or FFT block. In some embodiments, the tissue biopsy sample is a fresh-frozen tissue biopsy. The tissue biopsy sample can be prepared in thin sections (e.g., by cutting and / or mounting on slides) to facilitate pathology review (e.g., by staining with immunohistochemical stains for IHC review and / or with hematoxylin and eosin stains for H&E pathology review). For example, analysis of slides for H&E or IHC staining may reveal characteristics such as tumor infiltration, programmed death-ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.

[0310] In some embodiments, liquid samples (e.g., blood) collected from patients (e.g., in EDTA-containing collection tubes) are prepared (e.g., by smearing) on ​​slides for pathology review. In some embodiments, grossly dissected FFPE tissue sections, which may be mounted on histopathology slides from solid tissue samples (e.g., tumor or normal tissue), are analyzed by a pathologist. In some embodiments, tumor samples are evaluated to determine, for example, the tumor purity of the sample, the tumor cellularity rate as a ratio of tumor to normal nuclei, etc. For each section, background tissue can be excluded or removed so that the section meets a tumor purity threshold, e.g., at least 20% of the nuclei in the section are tumor nuclei, or at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the nuclei in the section are tumor nuclei.

[0311] In some embodiments, pathology data 128-1 is extracted using a computational approach to digital pathology in addition to or instead of visual inspection, providing, for example, morphological features extracted from digital images of stained tissue samples. In some embodiments, pathology data 128-1 includes features determined using machine learning algorithms to evaluate pathology data collected as described above.

[0312] Further details regarding methods, systems, and algorithms for using pathology data to classify cancer and identify targeted treatments are discussed, for example, in U.S. Patent Nos. 10,957,041, 11,244,763, 11,848,107, and 11,145,416, the contents of which are incorporated herein by reference in their entirety and for all purposes.

[0313] In some embodiments, the image data 128-2 collected during clinical evaluation includes features identified by review of in vitro and / or in vivo imaging results (e.g., of a tumor site), e.g., tumor size, differences in tumor size over time (e.g., during treatment or other changes), etc. In some embodiments, the image data 128-2 includes features determined using machine learning algorithms to evaluate the image data collected as described above.

[0314] Further details regarding methods, systems, and algorithms for using medical images to classify cancers and identify targeted treatments are discussed, for example, in U.S. Patent Nos. 10,957,041, 11,244,763, 11,848,107, and 11,145,416, the contents of which are incorporated herein by reference in their entirety and for all purposes.

[0315] In some embodiments, the tissue culture / organoid data 128-3 collected during clinical evaluation includes features identified by evaluation of cultured tissue from a subject. For example, in some embodiments, tissue samples (e.g., tumor tissue, normal tissue, or both) obtained from a patient are cultured (e.g., in liquid culture, solid culture, and / or organoid culture) and various features, such as cell morphology, growth characteristics, genomic alterations, and / or drug sensitivity, are evaluated. In some embodiments, the tissue culture / organoid data 128-3 includes features determined using machine learning algorithms to evaluate the tissue culture / organoid data collected as described above. Examples of culturing tissue organoids (e.g., individual tumor organoids) and their feature extraction are described in PCT Publication WO 2021 / 081253 and U.S. Patent No. 11,629,385, the contents of each of which are incorporated herein by reference in their entirety and for all purposes.

[0316] In some embodiments, the method further includes obtaining a liquid biopsy sample from a sample repository or sample database (e.g., BioIVT, TSC Biosample Repository, BioLINCC, etc.). In some embodiments, the liquid biopsy sample is obtained from the subject at least 1 hour, at least 2 hours, at least 12 hours, at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year prior to processing and / or sequencing the liquid biopsy sample. In some such embodiments, the liquid biopsy sample is fresh, frozen, dried, and / or fixed. In some embodiments, the liquid biopsy sample is processed and / or sequenced at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year prior to obtaining the first dataset. For example, in some embodiments, sequencing data for the liquid biopsy sample is obtained from a data repository (e.g., GenBank, NCBI Assembly, DNA DataBank of Japan, European Nucleotide Archive, European Variation Archive).

[0317] Block 918. According to block 918, using the panel enrichment sequencing reaction, a circulating tumor fraction (ctFE) is determined to meet (e.g., exceed) a threshold ctFE value. ctFE refers to the proportion of tumor-derived material in the bloodstream of a cancer subject. This material may include circulating tumor cells (CTCs), cell-free DNA (cfDNA), or other components shed from the tumor into the subject's bodily fluids, such as the bloodstream. ctFE provides information about tumor burden, disease progression, treatment response, and the extent of minimal residual disease. Monitoring changes in CTF over time can help assess the effectiveness of cancer treatment, detect early signs of metastasis or recurrence, and guide treatment decisions.

[0318] 4F3 shows a method 400-3 for estimating circulating tumor fraction of a liquid biopsy sample by fitting a simulated tumor fraction to the copy number state generated from cfDNA sequencing data. In some embodiments, such circulating tumor fraction estimates (ctFE) are used in the methods, systems, and CRMs described herein. For details of methods for estimating circulating tumor fraction of a liquid biopsy sample, see, for example, U.S. Patent No. 11,211,147, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0319] In some embodiments, any of the methods used to calculate ctFE disclosed in US Pat. No. 11,211,147 are used.

[0320] In some embodiments, ctFE is determined according to any of the methods described in U.S. patent application Ser. No. 18 / 930,786, filed Oct. 29, 2024, entitled "ESTIMATION OF CIRCULATING TUMOR FRACTION USING OFF-TARGET READS OF TARGETED-PANEL SEQUENCING," the disclosure of which is incorporated herein by reference in its entirety and for all purposes. This ctFE value is used to threshold (gate) the bTMB determination (e.g., a bTMB determination is made only if the ctFE value meets a threshold).

[0321] In some embodiments, ctFE is determined using targeted sequencing of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or all of the genes in Table 1.

[0322] In some embodiments, ctFE is determined using targeted sequencing of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or all of the genes in Table 2.

[0323] In some embodiments, ctFE is determined using targeted sequencing of at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or all genes in Table 1, and the tumor targets in these genes are any combination of single nucleotide variants (SNVs), insertions / deletions (indels), copy number variants (CNVs), and gene rearrangements.

[0324] In some embodiments, the ctFE value is determined according to any of the methods described in Finkle, 2021, "Validation of a liquid biopsy assay with molecular and clinical profiling of circulating tumor DNA," npj Precision Oncology 5(63), the disclosure of which is incorporated herein by reference in its entirety and for all purposes. This ctFE value is used to threshold (gate) the bTMB determination (e.g., only make a bTMB determination if the ctFE value meets a threshold).

[0325] In some embodiments, the method also includes using a panel enrichment sequencing reaction to determine whether ctFE exceeds a threshold ctFE value. In some embodiments, the circulating tumor fraction estimate prepared according to the methods described herein, see, for example, method 400-3 shown in Figure 4F3, is used in thresholding or gating bTMB determination. In other words, bTMB determination is performed only if the calculated ctFE value for the subject exceeds a threshold value.

[0326] Block 920. Referring to block 920, in some such embodiments, the threshold ctFE value is 0.01. In some embodiments, the threshold ctFE value is 0.0025. In some embodiments, the threshold ctFE value is between 0.001 and 0.015. In some embodiments, the threshold ctFE value is between 0.005 and 0.015. In some embodiments, the threshold ctFE value is between 0.015 and 0.025. In some embodiments, the threshold ctFE value is between 0.025 and 0.035. In some embodiments, the threshold ctFE value is between 0.035 and 0.045. In some embodiments, the threshold ctFE value is between 0.045 and 0.055.

[0327] Block 922. Referring to block 922, in response to determining that the ctFE is above a threshold, the liquid biopsy tumor mutation burden (lTMB) of the subject is calculated from the panel enrichment sequencing reaction. As used herein, lTMB refers to the number of mutations in the cell-free DNA of the liquid biopsy sample that originate from the subject's tumor cells. In some embodiments, lTMB is typically calculated by determining the total number of mutations detected in the panel enrichment sequencing reaction per megabase (Mb) of DNA. In some embodiments, lTMB is reported as mutations per megabase (mut / Mb).

[0328] In some embodiments, all non-silent somatic coding variations, such as missense variants, indel variants, and stop-loss variants, with coverage greater than x100 and allele fraction greater than 5% are included in the count of non-synonymous variations.

[0329] Block 924. Referring to block 924, in some embodiments, the calculation includes counting a plurality of genetic variants present in a plurality of nucleic acid sequences. In such counting, since the panel enrichment sequencing reactions are enriched for genes known to be mutated in cancer cells, any somatic mutations detected in the panel enrichment sequencing reactions are presumed to arise from the tumor and therefore contribute to the count. In some embodiments, a check is made to ensure that the genetic variants present in the plurality of nucleic acid sequences are not germline, and any germline mutations are removed from the count. In some embodiments, no check is made to determine whether such genetic variants are germline. Rather, metrics related to variant allele frequency or other criteria, discussed below beginning with block 2026, are used to remove genetic variants from the count.

[0330] Block 926. Referring to block 926, in some embodiments, the number of genetic variants present in the plurality of nucleic acid sequences is the number of unique genetic variants present in the plurality of nucleic acid sequences that meet one or more eligibility criteria in the set of eligibility criteria, i.e., only genetic variants that meet one or more eligibility criteria in the set of eligibility criteria contribute to one of the plurality of genetic variants present in the plurality of nucleic acid sequences used in the lTMB calculation.

[0331] Block 928. Referring to block 928, in some embodiments, an eligibility criterion in the set of eligibility criteria is a requirement that each genetic variant be a missense variant, a combination of a missense variant and a splice region variant, a frameshift variant, a stop-loss variant, a splice acceptor variant, an in-frame insertion variant, an in-frame deletion variant, a combination of a frameshift variant and a splice region variant, a disruptive in-frame insertion variant, or a disruptive in-frame deletion variant.

[0332] In some embodiments, the set of eligibility criteria consists of any one, two, three, four, five, six, seven, eight, nine, or all ten of the eligibility criteria in the group consisting of: (i) the requirement that each genetic variant be a missense variant, (ii) a combination of a missense variant and a splice region variant, (iii) a frameshift variant, (iv) a stop-loss variant, (v) a splice acceptor variant, (vi) an in-frame insertion variant, (vii) an in-frame deletion variant, (viii) a combination of a frameshift variant and a splice region variant, (ix) a disruptive in-frame insertion variant, and (x) a disruptive in-frame deletion variant. In such embodiments, genetic variants that meet any of the eligibility criteria in the set of eligibility criteria are removed from the count of genetic variants.

[0333] In some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant be a missense variant.A missense variant in a nucleic acid sequence is a type of genetic mutation that causes a single base change in the nucleic acid sequence of a gene, causing the gene to replace one amino acid in the corresponding protein with another amino acid.This mutation occurs when the single base change changes the codon (a sequence of three nucleotides) of the gene, so that a different amino acid is encoded.

[0334] In some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant be a combination of a missense variant and a splice region variant. That is, in the panel enrichment sequencing reaction, the original nucleic acid molecule represented by the sequence read contains both a missense variant and a splice region variant. A splice region variant is a type of genetic mutation that occurs in the non-coding region of a gene, particularly in a sequence involved in the process of RNA splicing. RNA splicing is a process of gene expression in which introns (non-coding regions) are removed from pre-mRNA and exons (coding regions) are joined to form a mature mRNA transcript. Splice region variants affect this splicing process, potentially resulting in changes in the mRNA transcript and affecting protein expression. Splice region variants occur within intron sequences adjacent to exon-intron boundaries. These regions contain conserved sequences, such as splice donor sites (5' splice sites), splice acceptor sites (3' splice sites), and branchpoint sequences. These are essential for the splicing machinery to correctly recognize and splice out introns. Therefore, a combination of missense and splice site variants refers to the situation where genetic mutations affecting both the coding and non-coding regions (splice sites) of a gene are present within the same allele in tumor cells.

[0335] In some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant is a frameshift variant. A frameshift variant is a type of genetic mutation resulting from the insertion or deletion of a number of nucleotides in a DNA sequence that is not divisible by 3. This insertion or deletion changes the reading frame of the gene during translation (the mechanism by which the sequence of nucleotides is read in triplets known as codons), causing a disruption in the normal protein coding sequence.

[0336] In some embodiments, an eligibility criterion in the set of eligibility criteria is a requirement that each genetic variant be a stop-loss variant. Stop-loss variants, also known as non-stop mutations or read-through mutations, are a type of genetic mutation that affects a termination codon (STOP codon) in the coding region of a gene. A STOP codon is a signal in the genetic code that instructs ribosomes to stop translating the mRNA chain and release the newly synthesized protein. However, in the case of a stop-loss variant, a mutation occurs that changes the STOP codon into a codon that codes for an amino acid, thereby increasing the length of the protein encoded by the gene.

[0337] In some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant be a splice acceptor variant. A splice acceptor variant, also known as a splice site variant or splice donor mutation, is a type of genetic mutation that affects the splice acceptor site at the boundary between exons and introns in the DNA sequence of a gene. A splice acceptor site is a specific sequence of nucleotides located at the exon-intron boundary in the DNA sequence of a gene. The splice acceptor site functions in the process of RNA splicing, removing introns and joining exons to generate mature mRNA transcripts. The splice acceptor site marks the 3' end of an exon and signals the spliceosome to recognize and remove introns during splicing. A splice acceptor variant occurs when a mutation disrupts or changes the consensus sequence of a splice acceptor site. This disruption can prevent splice site recognition by the spliceosome, potentially resulting in errors in the splicing process. The presence of splice acceptor variants can result in aberrant splicing patterns, such as exon skipping, intron retention, or activation of cryptic splice sites. These aberrant splicing events can generate mRNA transcripts with altered exon composition, potentially resulting in the synthesis of abnormal or truncated proteins.

[0338] In some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant be an in-frame insertion variant. An in-frame insertion variant is a type of genetic mutation that inserts nucleotides into a DNA sequence in such a way that the reading frame of the gene remains intact. In other words, the number of inserted nucleotides is divisible by 3. This means that the insertion does not disrupt the triplet codon structure of the gene. Such insertion of an amino acid coding sequence can change the structure and function of a protein to various degrees, depending on the specific position of the insertion and the inserted sequence.

[0339] In some embodiments, the qualifying criterion in the qualifying criterion set is that each genetic variant must be an in-frame deletion variant.In-frame deletion variants are a type of genetic mutation that deletes some nucleotides from the DNA sequence, while the reading frame of the gene remains intact.In other words, the number of deleted nucleotides is divisible by 3.In other words, the triplet structure of the gene is maintained.The deletion of these amino acid coding sequences can cause the structure and function of the protein encoded by the nucleic acid to change to various degrees, depending on the specific position of the deletion and the deletion sequence.

[0340] In some embodiments, the qualifying criterion in the qualifying criterion set is that the mutation is a combination of frameshift variants and splice region variants.The combination of frameshift variants and splice region variants at the same gene locus may cause complex genetic changes that affect both the coding and non-coding regions of genes.Frameshift variants may cause the production of truncated proteins with altered or lost function, while splice region variants may further exacerbate these effects by causing abnormal splicing and the production of abnormal mRNA transcripts.

[0341] In some embodiments, a qualifying criterion in the set of qualifying criteria is the requirement that the mutation be a disruptive in-frame insertion variant, which is a type of genetic mutation that inserts nucleotides into a DNA sequence in a manner that disrupts the normal reading frame of the gene encoded by the DNA sequence while maintaining that reading frame to some extent.

[0342] In some embodiments, an eligibility criterion in the set of eligibility criteria is the requirement that the mutation be a disruptive in-frame deletion variant, a type of genetic mutation that deletes nucleotides from a DNA sequence in a manner that maintains the reading frame of the gene but significantly alters the protein sequence encoded by the gene, potentially affecting its structure and function.

[0343] Blocks 930-932. Referring to block 930, in some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant has a variant allele frequency greater than 0.005 (0.5%) in a liquid biopsy sample. Referring to block 2032, in some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant has a variant allele frequency less than 1.0 (100%) in a liquid biopsy sample. The variant allele frequency (VAF) of each genetic variant is calculated by dividing the number of unique DNA fragments represented by the plurality of nucleic acid sequences containing each genetic variant by the total number of unique DNA fragments represented by the plurality of nucleic acid sequences that map to the locus of each genetic variant in the human genome. Thus, for example, consider a case where the plurality of nucleic acid sequences includes sequence reads of 13 unique DNA fragments containing a genetic variant. In this case, the genetic variant is for one of the genes enriched by the panel enrichment sequencing reaction. Further, in this example, the plurality of nucleic acid sequences includes sequence reads of 1000 unique DNA fragments that map to the locus of a genetic variant, where the variant allele frequency of this genetic variant is 13 / 1000 or 0.013.

[0344] Block 934. Referring to block 934, in some embodiments, a qualifying criterion in the set of qualifying criteria is a requirement that each genetic variant have a variant allele frequency (VAF) in a liquid biopsy sample that is one of: (i) greater than 0.01 (1%) and less than 0.4 (40%); (ii) greater than 0.6 (60%) and less than 0.9 (90%); (iii) greater than 0.4 (40%) and less than 0.60 (60%), provided that |VAF-ctFE| / ctFE<1; or (iv) greater than 0.9 (90%), provided that |VAF-ctFE| / ctFE<1.

[0345] Here, VAF is understood to represent the percentage of DNA molecules in a sample that carry a particular genetic variant. VAF ranges from 0 (no variant) to 1 (100% of DNA carries the variant).

[0346] This filter seeks to retain genetic variants with a specific VAF range. The filter has four inclusion criteria. Meeting any one of the criteria causes the variant to contribute to the variant count. Variants that do not meet any of these criteria do not contribute to the variant count. Therefore, these criteria are retention criteria.

[0347] Criterion (i): Variants with a VAF between 1% and 40% are retained. Without intending to be limited by any particular theory, this range may represent variants occurring at relatively low frequencies, potentially suggesting subclonal tumor populations or rare mutations.

[0348] Criterion (ii): Variants with a VAF of 60% to 90% are retained. Without intending to be limited by any particular theory, variants in this range may be frequent enough to represent clonal events, but not frequent enough to suggest near-homozygosity. These variants may represent mutations present in a significant proportion of tumor cells.

[0349] Criterion (iii): Variants with a VAF between 40% and 60% are retained if:

number

number

[0350] Criterion (iv): Variants with a VAF greater than 90% are retained if:

number

[0351] In some embodiments, a qualifying criterion in the set of qualifying criteria is the requirement that each genetic variant have a variant allele frequency (VAF) in a liquid biopsy sample that is one of: (i) greater than 0.0025 and less than 0.4 (40%); (ii) greater than 0.6 (60%) and less than 0.9 (90%); (iii) greater than 0.4 (40%) and less than 0.60 (60%), provided that |VAF-ctFE| / ctFE<1; or (iv) greater than 0.9 (90%), provided that |VAF-ctFE| / ctFE<1.

[0352] In some embodiments, an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant has a variant allele frequency (VAF) in a liquid biopsy sample that is one of: (i) greater than 0.0025 and less than A; (ii) greater than B and less than C; (iii) greater than A and less than B, provided that |VAF-ctFE| / ctFE<1; or (iv) greater than C, provided that |VAF-ctFE| / ctFE<1. In some such embodiments, A is a value between 30% and 50%, B is a value between 55% and 70%, and C is a value between 80% and 95%.

[0353] In some embodiments, the eligibility criteria in the set of eligibility criteria are such that each genetic variant has a variant allele frequency (VAF) in the liquid biopsy sample that is one of the following: (i) greater than D and less than A, (ii) greater than B and less than C, (iii) greater than A and less than B, provided that |VAF - ctFE| / ctFE < 1, or (iv) greater than C, provided that |VAF - ctFE| / ctFE < 1. In some such embodiments, A is a value between 30% and 50%, B is a value between 55% and 70%, C is a value between 80% and 95%, and D is a value between 0.0002 and 0.02.

[0354] In some embodiments, the eligibility criteria in the set of eligibility criteria are such that each genetic variant has a variant allele frequency (VAF) in the liquid biopsy sample that is one of the following: (i) greater than D and less than A, (ii) greater than B and less than C, (iii) greater than A and less than B, provided that |VAF - ctFE| / ctFE < P, or (iv) greater than C, provided that |VAF - ctFE| / ctFE < Q. In some such embodiments, A is a value between 30% and 50%, B is a value between 55% and 70%, C is a value between 80% and 95%, D is a value between 0.0002 and 0.02, P is a value between 0.8 and 1.2, and Q is the same as or different from P and is a value between 0.8 and Q.

[0355] Criterion (i): Variants with VAF between D and A are retained, where D is selected from the range of 0.0002 to 0.02 and A is selected from the range of 30% to 50%.

[0356] Criterion (ii): Variants with VAF between B and C are retained, where B is selected from the range of 55% to 70% and C is selected from the range of 80% to 95%.

[0357] Criterion (iii): Variants with VAF between A and B are retained if:

Number

[0358] Criterion (iv): Variants with a VAF greater than C are retained if:

number

[0359] Block 936. Referring to block 936, in some embodiments, a qualifying criterion in the set of qualifying criteria is a selection by a medical professional of genetic variants present in the plurality of nucleic acid sequences. In other words, the medical professional specifically flags the genetic variants as those that should contribute to the lTMB.

[0360] Block 2038. Referring to block 938, in some such embodiments, the calculation further comprises normalizing the number of genetic variants present in the plurality of nucleic acid sequences by the coverage of the plurality of probe sequences. In some such embodiments, the count of the genetic variants is normalized by dividing the count of the genetic variants by the total size of the panel associated with the panel enrichment sequencing reaction. For example, if the panel enrichment sequencing reaction includes probes that collectively map to 3 megabases of the human genome, the lTMB is calculated by dividing the count of the genetic variants (that meet the applicable eligibility criteria in the set of eligibility criteria) by 3 Mb.

[0361] Blocks 940-942. Referring to block 940, in some embodiments, the coverage of the panel enrichment sequencing reaction is between 0.1 megabases and 0.4 megabases. Referring to block 2042, in some embodiments, the coverage of the panel enrichment sequencing reaction is between 0.15 megabases and 0.3 megabases. In some embodiments, the coverage of the panel enrichment sequencing reaction is between 0.1 megabases and 15 megabases. In some embodiments, the coverage of the panel enrichment sequencing reaction is between 0.2 megabases and 30 megabases. In some embodiments, the coverage of the panel enrichment sequencing reaction is at least 0.2 megabases, 0.3 megabases, 0.4 megabases, 0.5 megabases, 0.6 megabases, or 0.7 megabases.

[0362] Block 944. Referring to block 944, in some embodiments, the lTMB is reported for the subject. Examples of suitable reports are described above in conjunction with FIG. 1, e.g., clinical assessment 139-1, clinical reports 139-1-3, and reporting module 180.

[0363] In some embodiments, the method further includes generating a report (e.g., for use by a physician) that includes the lTMB of each subject's biological sample. In some such embodiments, the generated report further includes a matched therapy (e.g., treatment and / or clinical trial) based on the lTMB status of the sample.

[0364] In some embodiments, the methods further include disease screening and / or monitoring over multiple time points. For example, in some embodiments, the methods are used to monitor disease progression and / or recurrence after treatment, to assess the effectiveness of treatment, and / or to perform comparative studies using liquid biopsy samples and matched solid tissue samples.

[0365] When a subject is determined to have a high tumor mutational burden (TMB), it indicates that the tumor has a large number of genetic mutations. This has clinical implications with respect to treatment with immune checkpoint inhibitors (ICIs), e.g., anti-PD-1 / PD-L1 therapy or anti-CTLA-4 therapy. High TMB often correlates with increased tumor neoantigens, making the cancer more likely to respond to immunotherapy. Thus, in some embodiments, the methods of the present disclosure determine that a cancer subject has a high TMB and that the subject is treated with an immune checkpoint inhibitor (ICI), e.g., anti-PD-1 / PD-L1 therapy or anti-CTLA-4 therapy. In particular, in some embodiments, the methods of the present disclosure determine that a cancer subject has a high TMB and that the subject is treated with pembrolizumab (Keytruda). In some embodiments, the methods of the present disclosure determine that a cancer subject has a high TMB and that the subject is treated with an immune checkpoint inhibitor in addition to an applicable cancer therapeutic for the underlying cancer.

[0366] High TMB is a biomarker for the potential efficacy of ICIs. Drugs such as pembrolizumab (Keytruda) have tumor-agonistic activity and are FDA-approved for tumors with high TMB (more than 10 mutations per megabase), a broad range of cancer types.

[0367] In some embodiments, if the methods of the present disclosure determine that a cancer subject has high TMB, further testing, such as PD-L1 expression and immune infiltration, is performed to develop a comprehensive cancer treatment strategy.

[0368] In some embodiments, when a subject is treated with an immune checkpoint inhibitor, the subject is monitored more closely for response and possible immune-related adverse events (irAEs), which can affect organs such as the liver, skin, lungs, and endocrine glands.

[0369] In some embodiments, the cancer is lung cancer, metastatic melanoma, colorectal cancer, head and neck cancer, bladder cancer, endometrial cancer, or cancer of unknown primary, each of which has been shown to be responsive to ICIs if the cancer is MSI-high (MSI-H).

[0370] In some embodiments, the disclosed systems and methods determine that a subject has high TMB colon cancer. High TMB is often seen in mismatch repair deficient (dMMR) or microsatellite instability-high (MSI-H) colorectal cancer. This cancer responds well to pembrolizumab and nivolumab. Thus, in some embodiments, if the disclosed systems and methods determine that a subject has high TMB colon cancer, the subject is treated with pembrolizumab and / or nivolumab.

[0371] In some embodiments, the methods described herein include generating a clinical report 139-3 (e.g., a patient report), providing clinical support for personalized cancer therapy, and / or using information curated from the sequencing of the liquid biopsy sample described above. In some embodiments, the report is provided to the patient, physician, medical professional, or researcher in a digital copy (e.g., a JSON object, a PDF file, or an image on a website or portal), a hard copy (e.g., printed on paper or other tangible medium). For example, the report object, such as a JSON object, can be used for further processing and / or display. For example, information from the report object can be used to generate a clinical test report for return to the prescribing physician. In some embodiments, the report is presented as text, as audio (e.g., prerecorded or streaming), as an image, or in another format, and / or any combination thereof.

[0372] The report includes information related to specific features of the patient's cancer, such as ITMB, detected genetic variants, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities. In some embodiments, other features of the patient's sample and / or clinical record are also included in the report. For example, in some embodiments, the clinical report includes information regarding one or more of clinical variants, such as copy number variants (e.g., in potentially actionable genes CCNE1, CD274 (PD-L1), EGFR, ERBB2 (HER2), MET, MYC, BRCA1, and / or BRCA2), fusions, translocations, and / or rearrangements (e.g., in potentially actionable genes ALK, ROS1, RET, NTRK1, FGFR2, FGFR3, NTRK2, and / or NTRK3), pathogenic single nucleotide polymorphisms, insertion-deletions (e.g., somatic / tumor and / or germline / normal), therapeutic biomarkers, microsatellite instability status, and / or tumor mutational burden.

[0373] In some embodiments, the results are used to design patient biological cell line tests, such as tumor organoid experiments.For example, organoids can be genetically engineered to have the same characteristics as the specimen, and can be observed after exposure to treatment to determine whether the treatment can reduce the growth rate of organoids, and therefore likely to reduce the growth rate of the patient's cancer associated with the specimen.Similarly, in some embodiments, the results are used to control experiments on tumor organoids derived directly from patients.An example of such experiment is described in U.S. Provisional Patent Application No. 62 / 944,292, filed on December 5, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0374] As shown in Figure 2A, in some embodiments, the clinical report is checked for final verification, review, and approval by a medical professional (e.g., a pathologist), after which the clinical report is sent for action (e.g., utilizing precision oncology).

[0375] Long-term reporting. In various embodiments, the report may include and / or compare the results of multiple liquid biopsy tests and / or solid tumor tests (e.g., multiple tests related to the same patient). The results of multiple liquid biopsy tests and / or solid tumor tests may be displayed on the portal in various configurations that can be selected and / or customized by the observer. The tests may be performed at different time points, and the samples on which the tests are performed may be taken at different time points.

[0376] Download the results. Clinical and / or molecular data associated with a patient (e.g., information that would be included in a report) may be aggregated and available via a portal. Any portion of the report data may be available for download (e.g., as a CSV file) by the physician and / or patient. In various embodiments, the data may include data related to gene variants, RNA expression levels, immunotherapy markers (including MSI and TMB), RNA fusions, etc. In one embodiment, if a physician or medical facility orders multiple tests (which may all be associated with the same patient, or the tests may be associated with multiple patients), results associated with two or more tests may be aggregated and downloaded in a single file.

[0377] Block 946. Referring to block 946, in some embodiments, reporting further includes reporting a matched treatment recommendation for the subject in response to determining that the lTMB meets the treatment threshold. Such treatment thresholds will necessarily be application dependent. For example, the treatment threshold may vary depending on the type of cancer the test is being tested to monitor, the stage of the cancer the test is being tested to monitor, and / or the identity of the genes included in the sequencing panel of the panel enrichment sequencing reaction, specifying several non-limiting variables that may affect the value of the treatment threshold.

[0378] Block 948. Referring to block 948, in some embodiments, reporting includes comparing the subject's lTMB to a severity threshold and reporting a qualitative status of either high lTMB (lTBM-H) or low lTMB (lTMB-L) based on the comparison.

[0379] Block 950. Referring to block 950, in some embodiments, a subject's lTMB is reported only if the lTMB meets a reporting threshold. Such a reporting threshold is necessarily application-dependent. For example, the reporting threshold may vary depending on the type of cancer the test is being performed to monitor, the stage of the cancer the test is being performed to monitor, and / or the identity of the genes included in the sequencing panel of the panel enrichment sequencing reaction, specifying several non-limiting variables that may affect the value of the reporting threshold.

[0380] Block 952. Referring to block 952, in some embodiments, an immunotherapeutic agent is administered to a subject if the subject's lTMB meets a treatment threshold. Such a treatment threshold will necessarily be application dependent. For example, the treatment threshold may vary depending on the type of cancer the test is being tested to monitor, the stage of the cancer the test is being tested to monitor, and / or the identity of the genes included in the sequencing panel of the panel enrichment sequencing reaction, specifying several non-limiting variables that may affect the value of the treatment threshold.

[0381] Block 954. Referring to block 954, in some embodiments, in response to the lTMB determining that the clinical trial threshold (associated with the clinical trial) is met, a clinical trial recommendation matched to the subject is reported. In other words, the subject is recommended for the clinical trial. Such a clinical trial threshold is necessarily application-dependent. For example, the clinical trial threshold may vary depending on the type of cancer targeted for the clinical trial, the stage of the cancer targeted for the clinical trial, and / or the identity of the genes included in the sequencing panel of the panel enrichment sequencing reaction, specifying several non-limiting variables that may affect the value of the clinical trial threshold.

[0382] In some embodiments, the clinical report 139-3 further includes information regarding clinical trials for which the patient is eligible, treatments specific to the patient's cancer, and / or potential adverse treatment effects associated with particular characteristics of the patient's cancer, such as the patient's genetic mutations, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities, or other features of the patient's sample and / or clinical record. For example, in some embodiments, the clinical report includes patient information and analysis metrics, including cancer type and / or diagnosis, variant allele proportion, patient demographic and / or histological characteristics, matched treatments (e.g., FDA-approved and / or investigational), matched clinical trials, variants of unknown significance (VUS), low-coverage genes, panel information, specimen information, details about reported variants, patient's clinical history, status and / or availability of previous test results, and / or bioinformatics pipeline version.

[0383] In some embodiments, the results contained in the report, and / or any additional results (e.g., results from a bioinformatics pipeline), are used to query a database of clinical data to determine, for example, whether there is a trend in other patients with the same or similar characteristics that indicates that a particular treatment will or will not be effective (e.g., slow or stop the progression of cancer), or that there is an adverse effect of such treatment.

[0384] Block 956. Referring to block 956, in some embodiments, a subject is enrolled (in a respective clinical trial) only if the subject's lTMB meets a clinical trial threshold (associated with the respective clinical trial). Such clinical trial thresholds are necessarily application-dependent. For example, the clinical trial thresholds may vary depending on the type of cancer targeted for the clinical trial, the stage of the cancer targeted for the clinical trial, and / or the identity of the genes included in the sequencing panel of the panel enrichment sequencing reaction, specifying several non-limiting variables that may affect the value of the clinical trial threshold.

[0385] Block 958. Referring to block 958, in some embodiments, the method further includes using the lTMB to identify a matching lTMB based on a predetermined correlation (e.g., a correlation of at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, or at least 0.8) between (i) detection of somatic mutations in cell-free DNA from liquid biopsy samples from the cohort of training subjects and (ii) detection of somatic mutations in genomic DNA from solid tumor biopsy samples from the cohort of training subjects. In other words, the correlation is established between the tumor mutation burden detected by sequencing the liquid biopsy samples and the tumor mutation burden detected from sequencing the solid tumor samples based on matched samples from the training subject cohort (i.e., a liquid biopsy sample and a solid tumor sample are sequenced for each member of the training cohort). For example, in some embodiments, the correlation is an average correlation across multiple training subjects or a correlation modeled using the individual correlations of each training subject. Furthermore, reporting further includes reporting the matching lTMB.

[0386] Block 960. Referring to block 960, in some embodiments, the subject's electronic medical record is updated to include the subject's lTMB.

[0387] Subjects and biological samples.

[0388] In some embodiments, the liquid biopsy sample corresponds to a matched tumor sample (e.g., a solid tumor sample obtained from the subject). For example, in some embodiments, the method further includes obtaining a second dataset determined from sequencing a plurality of cell-free nucleic acids in the matched tumor sample from the subject. In some embodiments, the matched tumor sample is obtained from the subject in parallel with the liquid biopsy sample. In some embodiments, the matched tumor sample is obtained from the subject at a different time than the acquisition of the liquid biopsy sample. In some embodiments, the matched tumor sample is any of the embodiments described above (see example method; Figure 2A: Example Precision Oncology Workflow). In some embodiments, the method further includes obtaining the matched tumor sample from a sample repository or sample database (e.g., BioIVT, TSC Biosample Repository, BioLINCC, etc.). In some embodiments, the matched tumor sample is obtained from the subject at least 1 hour, at least 2 hours, at least 12 hours, at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year prior to the acquisition of the liquid biopsy sample. In some embodiments, the matched tumor sample is fresh, frozen, dried, and / or fixed.In some embodiments, the matched tumor sample is processed and / or sequenced at least 1 day, at least 2 days, at least 1 week, at least 1 month, or at least 1 year before obtaining the second data set.For example, in some embodiments, the sequencing data of the nucleic acids in the matched tumor sample is obtained from a data repository (e.g., GenBank, NCBI Assembly, DNA DataBank of Japan, European Nucleotide Archive, European Variation Archive).

[0389] In some embodiments, the one or more reference samples are non-cancerous samples. In some embodiments, the one or more reference samples are matched normal samples (e.g., normal samples obtained from the subject). In some embodiments, the matched normal samples are obtained from the subject in parallel with the liquid biopsy sample. In some embodiments, the matched normal samples are obtained from the subject at a different time than the liquid biopsy sample. In some embodiments, the matched normal samples are any of the embodiments described above (see example method; Figure 2A: Example Precision Oncology Workflow).

[0390] In some alternative embodiments, the one or more reference samples comprise a pool of normal (e.g., non-cancerous) samples obtained from multiple control subjects (e.g., healthy subjects). In some such embodiments, the method further comprises obtaining the one or more reference samples from a sample repository or sample database (e.g., BioIVT, TSC Biosample Repository, BioLINCC, etc.). In some embodiments, the one or more reference samples comprise a liquid biopsy sample containing multiple cell-free nucleic acids and / or a solid tissue sample containing multiple nucleic acids. In some embodiments, the one or more reference samples are processed and / or sequenced at least one day, at least two days, at least one week, at least one month, or at least one year prior to obtaining the first dataset. For example, in some such embodiments, sequencing data for the one or more reference samples is obtained from a data repository (e.g., GenBank, NCBI Assembly, DNA DataBank of Japan, European Nucleotide Archive, European Variation Archive).

[0391] In some embodiments, the cell-free nucleic acids (e.g., in the first liquid biopsy sample of the subject and one or more reference samples) comprise circulating tumor DNA (ctDNA). In some embodiments, the method further comprises isolating a plurality of cell-free nucleic acids from the liquid biopsy sample of the subject prior to sequencing. In some embodiments, the sequencing is multiplexed sequencing. In some embodiments, the sequencing is short-read sequencing or long-read sequencing.

[0392] In some embodiments, sequencing is panel enrichment sequencing reaction.In some such embodiments, sequencing reaction is carried out at a read depth of 100 times or more, 250 times or more, 500 times or more, 1000 times or more, 2500 times or more, 5000 times or more, 10,000 times or more, 20,000 times or more, or 30,000 times or more.In some embodiments, sequencing panel comprises 1 or more, 10 or more, 20 or more, 50 or more, 100 or more, 150 or more, 200 or more, 300 or more, 500 or more, or 1000 or more genes.In some embodiments, sequencing panel comprises one or more genes listed in Table 1. In some embodiments, the sequencing panel includes at least 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or all of the genes listed in Table 1. In some embodiments, the sequencing panel includes one or more genes selected from the group consisting of MET, EGFR, ERBB2, CD274, CCNE1, MYC, BRCA1, and BRCA2. In some embodiments, the sequencing panel includes at least 2, 3, 4, 5, 6, 7, or all eight of MET, EGFR, ERBB2, CD274, CCNE1, MYC, BRCA1, and BRCA2. In some embodiments, the sequencing reaction is a whole exome sequencing reaction.

[0393] In some embodiments, obtaining the first dataset further comprises aligning a plurality of sequence reads obtained from sequencing a plurality of cell-free nucleic acids in a first liquid biopsy sample of the subject to a human reference genome.

[0394] In some embodiments, on average, each bin in the plurality of bins has 2 or more, 3 or more, 5 or more, 10 or more, 15 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1000 or more, 10000 or more, or 100,000 or more sequence reads in the plurality of sequence reads, and the sequence reads are mapped to the portion of the reference genome corresponding to each bin, and each sequence read uniquely represents various molecules in the cell-free nucleic acids in the liquid biopsy sample. For example, in some embodiments, the plurality of cell-free nucleic acids in the liquid biopsy sample are sequenced using a sequencing method that utilizes a unique molecular identifier (UMI) for each cell-free nucleic acid in the liquid biopsy sample, and each sequence read in the plurality of sequence reads has a unique UMI. In such embodiments, sequence reads with the same UMI are packed (corrected) into one sequence read that carries the UMI.

[0395] Long-term testing

[0396] In some embodiments, one or more liquid biopsy assays described herein may be used to analyze patient samples obtained over the course of a patient's treatment. For example, blood samples may be obtained periodically and / or based on indications of response to treatment, disease recurrence, and / or disease progression. In some embodiments, one or more liquid biopsy assays may be performed on samples obtained from a patient monthly, every two months, every three months, every four months, every five months, every six to twelve months, etc. In some embodiments, the longitudinal use of liquid biopsy assays may be used to track clonal evolution to identify resistance mutations. In some embodiments, the longitudinal use of liquid biopsy assays may be used to track the evolution of mutations, such as EGFR mutations or APC mutations.

[0397] In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect new mechanisms of therapeutic resistance. In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect alterations in the AR gene. In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect alterations in the WNT pathway in mCRPC associated with resistance to enzalutimide and abiraterone. In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect ER mutations, such as those associated with resistance to endocrine therapy in breast cancer. In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect EGFR mutations involved in anti-EGFR therapy resistance (e.g., T790M) in NSCLC. In some embodiments, the longitudinal use of liquid biopsy assays may be used to detect KRAS, NRAS, MET, ERBB2, FLT3, or EGFR mutations associated with primary or acquired resistance to EGFR inhibitors in colorectal cancer. In some embodiments, longitudinal use of liquid biopsy assays may be used to assess genetic alterations from tumor cells exfoliated from primary tumors and metastatic sites.

[0398] In some embodiments, one or more blood samples may be collected from a patient in a home environment, for example, by a mobile phlebotomist.

[0399] For example, a first blood sample, a second blood sample, and a third blood sample may be taken from a patient during the course of treatment.

[0400] The present disclosure also provides a computer system comprising one or more processors and a non-transitory computer-readable medium containing computer-executable instructions that, when executed by the one or more processors, cause the processors to perform any of the methods and embodiments disclosed herein.

[0401] The present disclosure also provides a non-transitory computer-readable storage medium storing program code instructions that, when executed by a processor, cause the processor to perform any of the methods and embodiments disclosed herein.

[0402] Variant feature analysis

[0403] In some embodiments, predicted functional effects and / or clinical interpretations for one or more identified variants are curated using information from a variant database, hi some embodiments, a weighted heuristic model is used to characterize each variant.

[0404] In some embodiments, identified clinical variants are labeled as "potentially actionable," "biologically relevant," "variants of uncertain significance (VUS)," or "benign." Potentially actionable modifications are protein-modifying variants with associated therapies based on evidence from the medical literature. Biologically relevant modifications are protein-modifying variants that may be functionally significant or are recognized in the medical literature but are not associated with specific therapies. Variants of uncertain significance (VUS) are protein-modifying variants that have an unclear effect on function and / or for which there is insufficient evidence to determine their pathogenicity. In some embodiments, benign variants are not reported. In some embodiments, variants are identified by aligning the patient's DNA sequence to the human genome reference sequence version hg19 (GRCh37). In some embodiments, potentially actionable and biologically relevant somatic variants are provided in the clinical summary during report generation.

[0405] For example, in some embodiments, variant classification and reporting are performed, and detected variants are investigated according to criteria from known evolutionary models, functional data, clinical data, literature, and other research efforts, including tumor organoid experiments. In some embodiments, variants are prioritized and classified based on known gene-disease associations, hotspot regions within genes, internal and external somatic databases, primary literature, and other characteristics of somatic inducers. Variants can be added to patient (or sample, e.g., organoid sample) reports based on recommendations from the AMP / ASCO / CAP guidelines. Additional guidelines may be followed. Briefly, pathogenic variants with therapeutic significance, diagnostic significance, or prognostic significance may be prioritized in reporting. Non-functional pathogenic variants may be included as biologically relevant, followed by variants of uncertain significance. Translocations may be reported based on known gene fusion characteristics, associated breakpoints, and biological relevance. Evidence will be curated from public and private databases or studies and presented as 1) consensus guidelines, 2) clinical studies, or 3) case studies, with links to supporting literature. Germline alterations may be reported as incidental findings in a subset of genes, with patient consent. Genes recommended by the American College of Medical Genetics and Genomics (ACMG) and additional genes associated with cancer predisposition or drug resistance may be included.

[0406] It should be understood that the examples given above are illustrative and do not limit the use of the systems and methods described herein in combination with digital and laboratory medical platforms.

[0407] The results of the bioinformatics pipeline may be provided for report generation 208. Report generation may include variant scientific analysis, including interpretation of variants (including somatic and germline variants, if applicable) for pathogenicity and biological significance. Variant scientific analysis may also estimate microsatellite instability (MSI) or tumor mutational burden. Targeted therapies may be identified based on the gene, variant, and cancer type for further consideration and review by the ordering physician. In some aspects, clinical trials for which the patient may be eligible may be identified based on the mutation, cancer type, and / or medical history. Subsequent validation may occur, after which the report may be finalized for signature and delivery. In some embodiments, the first or second report may include additional data p...

Claims

1. 1. A method for determining liquid biopsy tumor mutation burden (lTMB) for a subject, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, A) obtaining from a panel enrichment sequencing reaction a plurality of nucleic acid sequences, the plurality of nucleic acid sequences comprising sequences corresponding to each cell-free DNA fragment in a first plurality of cell-free DNA fragments obtained from a liquid biopsy sample from the subject, wherein each respective cell-free DNA fragment in the first plurality of cell-free DNA fragments corresponds to a respective probe sequence in a plurality of probe sequences, the probe sequences being used in the panel enrichment sequencing reaction to enrich the cell-free DNA fragments in the liquid biopsy sample; and B) determining, using the panel enrichment sequencing reaction, that the circulating tumor fraction (ctFE) is above a threshold ctFE value; C) calculating the lTMB for the subject from the panel enrichment sequencing reactions in response to determining that the ctFE is above a threshold; and D) reporting the lTMB of the subject.

2. 10. The method of claim 1, wherein the panel enrichment sequencing reaction is performed at a read depth of at least 500x.

3. 3. The method of claim 1 or 2, wherein the panel enrichment sequencing reaction uses a sequencing panel that enriches between 50 and 150 genes.

4. 3. The method of claim 1 or 2, wherein the plurality of probe sequences used to enrich cell-free DNA fragments in the liquid biopsy sample in the panel enrichment sequencing reaction collectively map to between 25 different genes and 150 different genes in the human reference genome.

5. 5. The method of any one of claims 1 to 4, wherein the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in Table 1.

6. 6. The method of any one of claims 1 to 5, wherein the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in List 1.

7. 6. The method of any one of claims 1 to 5, wherein the panel enrichment sequencing reaction uses a sequencing panel that enriches for at least 10 genes listed in List 2.

8. 6. The method of any one of claims 1-5, wherein said calculating C) comprises determining the number of genetic variants present in the plurality of nucleic acid sequences.

9. 9. The method of claim 8, wherein said calculating C) further comprises normalizing the number of said plurality of genetic variants present in said plurality of nucleic acid sequences by the coverage of said plurality of probe sequences.

10. 10. The method of claim 9, wherein the coverage is between 0.1 megabases and 0.4 megabases.

11. 10. The method of claim 9, wherein the coverage is between 0.15 megabases and 0.3 megabases.

12. 12. The method of any one of claims 8 to 11, wherein the number of genetic variants present in the plurality of nucleic acid sequences is the number of unique genetic variants present in the plurality of nucleic acid sequences that meet one or more eligibility criteria in a set of eligibility criteria.

13. 13. The method of claim 12, wherein an eligibility criterion in the set of eligibility criteria is a requirement that each genetic variant be a missense variant, a combination of a missense variant and a splice region variant, a frameshift variant, a stop-loss variant, a splice acceptor variant, an in-frame insertion variant, an in-frame deletion variant, a combination of a frameshift variant and a splice region variant, a disruptive in-frame insertion variant, or a disruptive in-frame deletion variant.

14. 14. The method of claim 12 or 13, wherein an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant has a variant allele frequency greater than 0.005 (0.5%) in the liquid biopsy sample.

15. 15. The method of any one of claims 12 to 14, wherein an eligibility criterion in the set of eligibility criteria is the requirement that each genetic variant has a variant allele frequency below 1.0 (100%) in the liquid biopsy sample.

16. An eligibility criterion in the set of eligibility criteria is that each genetic variant has a variant allele frequency (VAF) in the liquid biopsy sample that is one of the following: (i) greater than 0.01 (1%) and less than 0.4 (40%); (ii) greater than 0.6 (60%) and less than 0.9 (90%); (iii) greater than 0.4 (40%) and less than 0.60 (60%), provided that: [Equation 1] or (iv) higher than 0.9 (90%), but [Equation 2] 14. The method of claim 12 or 13, wherein the necessary condition is that

17. The method of any one of claims 1 to 16, wherein the threshold ctFE value is 0.0025.

18. The method of any one of claims 1 to 17, wherein the liquid biopsy sample is a blood sample.

19. The method comprises: using the lTMB to identify a concordant lTMB based on a defined correlation between (i) detection of somatic mutations in cell-free DNA from liquid biopsy samples from each training subject in a cohort of training subjects, and (ii) detection of somatic mutations in genomic DNA from solid tumor biopsy samples from the cohort of training subjects; and The method of any one of claims 1 to 18, wherein said reporting D) further comprises reporting said matching lTMB.

20. 20. The method of any one of claims 1-19, wherein the reporting D) further comprises reporting a matched treatment recommendation for the subject in response to determining that the lTMB meets a treatment threshold.

21. The method of any one of claims 1 to 20, further comprising administering an immunotherapeutic agent to the subject when the lTMB of the subject meets a therapeutic threshold.

22. 22. The method of any one of claims 1-21, wherein the reporting D) further comprises reporting a clinical trial recommendation matched to the subject in response to determining that the lTMB meets a clinical trial threshold.

23. 23. The method of any one of claims 1-22, further comprising enrolling the subject when the lTMB of the subject meets a clinical trial threshold.

24. 24. The method of any one of claims 1 to 23, wherein the reporting D) comprises comparing the lTMB of the subject with a severity threshold, and reporting a qualitative status of either high lTMB (lTMB-H) or low lTMB (lTMB-L) based on the comparison.

25. 25. The method of any one of claims 1 to 24, wherein the lTMB of the subject is reported only if the lTMB meets a reporting threshold.

26. 26. The method of any one of claims 1-25, further comprising updating the subject's electronic medical record to include the lTMB of the subject.

27. 17. The method of any one of claims 12 to 16, wherein an eligibility criterion in said set of eligibility criteria is a selection by a medical professional of genetic variants present in said plurality of nucleic acid sequences.

28. 28. The method of any one of claims 1 to 27, wherein the plurality of probe sequences maps to no more than 150 genes in the human genome.

29. 1. A computer system comprising: one or more processors, and A computer system comprising: a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by the one or more processors, cause the processors to perform the method of any one of claims 1 to 28.

30. A non-transitory computer readable storage medium storing program code instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 28.