Probe sets for liquid biopsy assays

JP2025507673A5Pending Publication Date: 2026-03-19テンパスエーアイインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing liquid bioassay methods face signal-to-noise ratio fluctuations and low-concentration DNA detection difficulties in detecting and verifying cancer-specific genome changes, especially in the early stages of cancer, resulting in false-negative and incomplete cancer genome analysis.

Method used

Multi-group probe sets were used to enrich and hybridize cell free DNA at different dilution rates, adjust the detection limits of different gene loci, and reduce the need for multiple nucleic acid enrichment and sequencing tests.

Benefits of technology

The enrichment of multiple gene loci across different detection limits is achieved, the sensitivity and specificity of liquid biocheck is improved, experimental steps and costs are reduced, and the detection ability of early cancers is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000244_0000
    Figure 00000244_0000
  • Figure 00000244_0001
    Figure 00000244_0001
  • Figure 00000244_0002
    Figure 00000244_0002
Patent Text Reader

Abstract

Compositions for concentrating target nucleic acids and methods of using the compositions are provided. The compositions include a probe set and a plurality of nucleic acids. The probe set includes a first set of probes including a first plurality of probe species, each probe species targeting a respective genomic region in the first plurality of genomic regions and present in the composition at a first average molar concentration. The probe set further includes a second set of probes including a second plurality of probe species, each probe species targeting a respective genomic region in the second plurality of genomic regions and present in the composition at a second average molar concentration that is 5 to 8 times higher than the first average concentration. The plurality of nucleic acids includes cell-free nucleic acids from a biological sample of a subject or nucleic acids prepared therefrom.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 314,267, filed February 25, 2022, and U.S. Provisional Patent Application No. 63 / 387,262, filed December 13, 2022, the contents of which are incorporated by reference herein in their entirety for all purposes.

[0002] The present disclosure relates generally to improved probe sets and their use to enrich cell-free DNA data to provide clinical support for personalized treatment of disorders such as cancer. [Background technology]

[0003] Precision oncology is the practice of tailoring cancer therapy to an individual's unique genomic, epigenetic, and / or transcriptomic profile. Personalized cancer therapy builds on traditional treatment regimens used to treat cancer based solely on the overall classification of the cancer, for example, treating all breast cancer patients with one treatment and all lung cancer patients with a second treatment. This field arose from the frequent observation that different patients diagnosed with the same type of cancer, such as breast cancer, responded very differently to common treatment regimens. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that improve predictions of how individual cancers will respond to specific treatment modalities.

[0004] There is growing evidence that cancer patients who receive genetically guided therapy have better outcomes. For example, studies have shown that targeted therapy results in significant improvements in progression-free cancer survival. See, e.g., Radovich M. et al., Oncotarget, 7(35):56491-500 (2016). Similarly, a report from the IMPACT trial, a large (n=1307) retrospective analysis of consecutive, prospectively molecularly profiled patients with advanced cancer who participated in a large personalized medicine trial, showed that patients receiving targeted therapy matched to their tumor biology had a 16.2% response rate, compared to a 5.2% response rate for patients receiving unmatched therapy. Tsimberidou AM et al., ASCO 2018, Abstract LBA2553 (2018).

[0005] Indeed, therapies targeted to specific genomic alterations are already standard of care in some tumor types, as suggested by, for example, the National Comprehensive Cancer Network (NCCN) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer. In practice, the implementation of these targeted therapies requires determining the diagnostic marker status in each eligible cancer patient. While this can be achieved for some known mutations associated with treatment recommendations in the NCCN guidelines using individual assays or small next-generation sequencing (NGS) panels, the increasing number of actionable genomic alterations and the increasing complexity of diagnostic classifiers necessitates a more comprehensive assessment of each patient's cancer genome, epigenome, and / or transcriptome.

[0006] For example, evidence suggests that the use of combination therapies, in which each component is matched to actionable genomic alterations, holds the greatest potential for treating individual cancers. To date, retrospective studies of cancer patients treated with one or more therapeutic regimens have revealed that patients receiving therapies matched to a higher proportion of genomic alterations experienced a higher frequency of stable disease (e.g., longer time to recurrence), longer time to treatment failure, and better overall survival. Wheeler JJ et al., Cancer Res., 76:3690-701 (2016). Therefore, comprehensive assessment of each cancer patient's genome, epigenome, and / or transcriptome should maximize the benefits offered by precision oncology by facilitating more fine-tuned combination therapies, off-label use of novel drugs, and / or tissue-independent immunotherapies. See, for example, Schwaederle M. et al., J Clin Oncol., 33(32):3817-25(2015), Schwaederle M. et al., JAMA Oncol., 2(11):1452-59(2016), and Wheeler JJ et al., Cancer Res., 76(13):3690-701(2016). Furthermore, the use of comprehensive next-generation sequencing analysis of cancer genomes facilitates better access and larger patient pools for clinical trial enrollment. Coyne GO et al., Curr. Probl. Cancer, 41(3):182-93(2017), and Markman M., Oncology, 31(3):158,168.

[0007] To address the need for more comprehensive characterization of an individual's cancer genome, the use of large-scale NGS genomic analysis is increasing. See, e.g., Fernandes GS et al., Clinics, 72(10):588-94. Recent studies have shown that 30-40% of patients on whom large-scale NGS genomic analysis is performed subsequently receive clinical care based on the assay results, which is limited, at least, by the identification of actionable genomic alterations, the availability of drugs to treat the identified actionable genomic alterations, and the subject's clinical condition. Ross JS et al., JAMA Oncol.,1(1):40-49(2015), Ross JS et al.,Arch.Pathol.Lab Med.,139:642-49(2015), Hirshfield KM et al.,Oncologist,21(11):1315-25(2016), and Groisberg R.et al. al., Oncotarget, 8:39254-67 (2017).

[0008] However, these large-scale NGS genomic analyses are traditionally performed on solid tumor samples. For example, each of the studies referenced in the above paragraph performed NGS analysis of FFPE tumor blocks from patients. Solid tissue biopsies represent a well-known and proven methodology that provides a high degree of accuracy and therefore remain the gold standard for diagnosis and identification of predictive biomarkers. Nevertheless, the use of solid tissue materials for large-scale NGS genomic analyses of cancer has significant limitations. For example, tumor biopsies are subject to sampling bias caused by spatial and / or temporal genetic heterogeneity, e.g., between two regions of a single tumor and / or between different cancer tissues (between a primary tumor site and a metastatic tumor site, or between two different primary tumor sites). Such inter- or intratumor heterogeneity can lead to overlooking subclonal or emerging mutations when using localized tissue biopsies, and sampling bias can worsen over time as subclonal populations further evolve and / or shift in dominance.

[0009] In addition, obtaining solid tissue biopsies often requires invasive surgical procedures, for example, when the primary tumor site is located in an internal organ. These procedures are expensive, time-consuming, and can involve significant risks to the patient, for example, when the patient's health is poor and they cannot tolerate invasive medical procedures, and / or the tumor is located in a particularly sensitive or inoperable location, such as the brain or heart. Furthermore, the amount of tissue that can be procured depends on multiple factors, including tumor location, tumor size, patient vulnerability, and the risk of biopsy-related comorbidities, such as bleeding and infection. For example, a recent study reported that tissue samples in the majority of patients with advanced non-small cell lung cancer are limited to small biopsies, and in up to 31% of patients, no samples can be obtained at all. (Ilie and Hofman, Transl. Lung Cancer Res., 5(4):420-23 (2016)). Even when tissue biopsies are obtained, the sample may be too insufficient for comprehensive testing.

[0010] Furthermore, methods of tissue collection, preservation (e.g., formalin fixation), and / or archiving of tissue biopsies can result in sample degradation and variable-quality DNA. This, in turn, leads to inaccuracies in downstream assays and analyses, including next-generation sequencing (NGS) for biomarker identification. Ilie and Hofman, Transl Lung Cancer Res., 5(4):420-23 (2016).

[0011] Additionally, the invasiveness of the biopsy procedure, the time and expense associated with obtaining the sample, and the impaired state of cancer patients undergoing therapy make repeated testing of cancer tissue impractical, if not impossible. As a result, solid tissue biopsy analysis is not suitable for many monitoring schemes that benefit cancer patients, such as disease progression analysis, treatment efficacy assessment, disease recurrence monitoring, and other techniques that require data from several time points.

[0012] Cell-free DNA (cfDNA) has been identified in various body fluids, such as serum, plasma, and urine. Chan et al., Ann. Clin. Biochem., 40(Pt 2):122-30(2003). This cfDNA originates from all types of necrotic or apoptotic cells, including germline cells, hematopoietic cells, and pathological (e.g., cancer) cells. Advantageously, genomic alterations in cancer tissues can be identified from cfDNA isolated from cancer patients. See, for example, Stroun et al., Oncology, 46(5):318-22(1989), Goessl et al., Cancer Res., 60(21):5941-45(2000), and Frenel et al., Clin. Cancer Res. 21(20):4586-96(2015). Thus, one approach to overcoming the problems presented by the use of solid tissue biopsies described above is to analyze cell-free nucleic acids (e.g., cfDNA) and / or nucleic acids in circulating tumor cells present in biological fluids, for example, via liquid biopsy.

[0013] Specifically, liquid biopsies offer several advantages over traditional solid tissue biopsy analysis. For example, because bodily fluids can be collected in a minimally or non-invasive manner, sample collection is simpler, faster, safer, and less expensive than solid tumor biopsies. Such methods require only small sample volumes (e.g., 10 mL or less of whole blood per biopsy), reducing the discomfort and risk of complications experienced by patients during traditional tissue biopsies. In fact, liquid biopsy samples can be collected with limited or no assistance from medical professionals and can be performed almost anywhere. Furthermore, liquid biopsy samples can be collected from any patient, regardless of the location of their cancer, their overall health, and any previous biopsy collections. This enables the analysis of cancer genomes in patients for whom solid tumor samples cannot be easily and / or safely obtained. In addition, because cell-free DNA in bodily fluids originates from many different types of tissue in a patient, the genomic alterations present in the pool of cell-free DNA represent various distinct clonal subpopulations of the target cancer tissue, facilitating a more comprehensive analysis of the target cancer genome than is possible from one or more sections of a single solid tumor sample.

[0014] Liquid biopsies also enable serial genetic testing prior to cancer detection, during early stages of cancer progression, throughout the course of treatment, and during remission, for example, to monitor for disease recurrence. The ability to perform serial testing via noninvasive liquid biopsies throughout the course of disease may prove beneficial to many patients, for example, by monitoring patient response to therapy, the emergence of new actionable genomic alterations, and / or drug resistance changes. This type of information allows healthcare professionals to more quickly adjust and update treatment regimens, for example, facilitating more timely intervention in the event of disease progression. See, e.g., Ilie and Hofman, Transl. Lung Cancer Res., 5(4):420-23 (2016).

[0015] While liquid biopsies are a promising tool for improving outcomes using precision oncology, there are significant challenges inherent in using cell-free DNA for assessment of a subject's cancer genome. For example, there is a highly variable signal-to-noise ratio from one liquid biopsy sample to the next. This occurs because cfDNA originates from a variety of different cells, both healthy and diseased, in a subject. Depending on the stage and type of cancer in any particular subject, the fraction of cfDNA fragments derived from cancer cells (the "tumor fraction" or "ctDNA fraction" of the sample / subject) can range from nearly 0% to well over 50%. Other factors, including tumor type and mutation profile, can also affect the amount of DNA released from cancer tissue. For example, cfDNA clearance through the liver and kidneys is affected by a variety of factors, including renal dysfunction or other tissue-damaging factors (e.g., chemotherapy, surgery, and / or radiation therapy).

[0016] This, in turn, leads to problems in detecting and / or validating cancer-specific genomic alterations in liquid samples. This is especially true during the early stages of disease, when cancer therapy has a much higher success rate, because the tumor fraction in patients is at its lowest at this time. Thus, early-stage cancer patients may have ctDNA fractions below the limit of detection (LOD) of one or more beneficial genomic alterations, limiting clinical utility due to the risk of false negatives and / or providing an incomplete picture of the patient's cancer genome. Furthermore, because cancers and even individual tumors can be clonally diverse, actionable genomic alterations occurring in only a subset of clonal populations are diluted below the overall tumor fraction of the sample, hindering attempts to tailor combination therapies to various actionable mutations in the patient's cancer genome. As a result, most studies using liquid biopsy samples to date have focused on late-stage patients for assay validation and research.

[0017] Another challenge associated with liquid biopsies is accurately determining the tumor fraction in a sample. This difficulty stems at least from the heterogeneity of cancer and the increased frequency of large-scale chromosomal duplications and deletions found in cancer. As a result, the frequency of genomic alterations from cancer tissue varies from locus to locus based on at least (i) their prevalence in different subclonal populations of the cancer of interest and (ii) their location within the genome relative to large-scale chromosomal copy number variations. The difficulty of accurately determining the tumor fraction of a liquid biopsy sample impacts the accurate measurement of various cancer characteristics that have been shown to have diagnostic value for the analysis of solid tumor biopsies. These include allele ratios, copy number variations, global mutation burden, and the frequency of aberrant methylation patterns, all of which correlate with the percentage of DNA fragments originating from cancer tissue as opposed to healthy tissue.

[0018] Collectively, these factors result in highly variable concentrations of ctDNA from patient to patient, and potentially from locus to locus, confounding accurate measurement of disease indicators and actionable genomic alterations. Furthermore, the quantity and quality of cfDNA obtained from liquid biopsy samples is highly dependent on the specific methodology for collecting the sample, storing the sample, sequencing the sample, and normalizing the sequencing data.

[0019] While validation studies of existing liquid biopsy assays have demonstrated high sensitivity and specificity, few studies have corroborated results using orthogonal methods or across specific testing platforms, such as different NGS technologies and / or targeted panel sequencing versus whole genome / exome sequencing. Reports of liquid biopsy-based studies are limited by comparison with non-comprehensive tissue testing algorithms, including Sanger sequencing, small NGS hotspot panels, polymerase chain reaction (PCR), and fluorescent in situ hybridization (FISH), which may not include all NCCN guideline genes within their reportable range and are therefore inferior to more comprehensive liquid biopsy assays.

[0020] As the field of precision oncology continues to grow, many different nucleic acid sequencing-based assays have been developed to inform the diagnosis, prognosis, and treatment of a range of cancer conditions. For example, assays have been developed for the early identification of cancer, determining cancer type and / or cancer stage, predicting the primary tissue source of tumors of unknown origin, predicting treatment response to specific treatment regimens, etc. Each of these assays has its own sequencing requirements, for example, requiring sequencing information from a specific set of genomic loci at a specific minimum sequencing depth, and the performance of each assay is validated using nucleic acids prepared from a specific sample type according to a specific protocol and sequenced according to a specific sequencing technology. As clinical standard of care continues to become more specialized for each cancer type and each individual cancer, more of these nucleic acid-based sequencing assays will be relied upon to inform the treatment of a single patient. However, because each of these assays relies on different sequencing information collected using different preparation and sequencing methodologies, this requires an increasing amount of biological sample from the patient, which is particularly problematic when solid tumor samples are required for one or more of the assays. This also requires the performance of an increasing number of assays per patient, which increases healthcare costs and slows clinical decision making.

[0021] The information disclosed in this Background section is merely for the purpose of enhancing understanding of the general background of the present invention and should not be construed as an acknowledgment or any form of suggestion that this information forms prior art already known to those skilled in the art. Summary of the Invention

[0022] In view of the above background, there is a need in the art for improved compositions, methods, and systems for supporting clinical decisions in precision oncology using liquid biopsy assays. In particular, there is a need for improved liquid biopsy assays that can support an increasing number of molecular analyses and their different requirements in a cost-effective manner. For example, liquid biopsy assays that use probe sets that enrich for a wide range of genomic loci at different detection limits. Advantageously, the present disclosure solves this and other needs in the art by providing probe sets and cell-free DNA hybridization reactions that facilitate the enrichment of many genomic loci across different detection limits, eliminating the need to perform multiple nucleic acid enrichment and sequencing assays.

[0023] For example, in some embodiments, the improved methods and compositions described herein are based on the discovery that different genomic variant detection limits can be achieved in liquid biopsy assays by using probe sets with different stoichiometric ratios. However, as shown herein, there is no linear relationship between the concentration of enrichment probes in liquid biopsy assays and locus sequencing coverage / detection limit. For example, as reported in Example 11, the median detection limit for SNVs and MNVs in panel enrichment sequencing reactions of cell-free DNA for a first set of genes captured in a hybridization assay ("non-enriched" genes) using 0.7 fmol of each probe was 0.42%, while the median detection limit for SNVs and MNVs in the same panel enrichment sequencing reactions of cell-free DNA for a second set of genes captured in a hybridization assay ("enriched" genes) using 4.55 fmol of each probe was 0.25%. That is, even though the probes for the enriched genes were present at a 6.5-fold higher concentration in the hybridization reaction than the probes for the non-enriched genes, the detection limit for SNVs and MNVs in the enriched genes was less than 2-fold.

[0024] Furthermore, it was discovered that as the concentration of probes targeting the second set of genes was increased in the hybridization reaction, the sequencing coverage for the first set of genes decreased. For example, as shown in Figure 7 and reported in Example 3, increasing the molar concentration of enrichment probes targeting the second set of genes (the "enriched" genes) in the hybridization reaction resulted in both an increase in sequence coverage for the second set of genes and a decrease in sequence coverage for the first set of genes.

[0025] Thus, in some embodiments, the advantages provided by the methods and compositions described herein are realized, at least in part, through the discovery of stoichiometric ratios that allow variant detection in cell-free DNA to be tuned to different detection limits for different genomic loci in a single liquid biopsy reaction. Thus, in some embodiments of the methods and compositions described herein, a first set of enrichment probes is used at a first concentration (e.g., to facilitate detection of genomic variants represented in cell-free DNA at a first detection limit), and a second set of enrichment probes is used at a second concentration that is 5-8 times higher than the first concentration (e.g., to facilitate detection of genomic variants represented in cell-free DNA at a second, lower detection limit).

[0026] For example, in one aspect, the present disclosure provides a composition for enriching a target nucleic acid, the composition comprising a probe set and a plurality of nucleic acids. The probe set comprises a first set of polynucleotide probes (e.g., non-enhanced probes) that collectively target a first plurality of genomic regions at an average coverage of 0.75x to 1.25x, and the first set of polynucleotide probes comprises a first plurality of polynucleotide probe species. Each respective polynucleotide probe species in the first plurality of polynucleotide probe species targets a respective genomic region in the first plurality of genomic regions, and the polynucleotide probe species in the first plurality of polynucleotide probe species are present in the composition at a first average molar concentration.

[0027] The probe set includes a second set of polynucleotide probes (e.g., enrichment probes) that collectively target a second plurality of genomic regions at an average coverage of 0.75x to 1.25x, the second set of polynucleotide probes including a second plurality of polynucleotide probe species, each of which targets a respective genomic region in the second plurality of genomic regions, the polynucleotide probe species in the second plurality of polynucleotide probe species being present in the composition at a second average molar concentration, the second average molar concentration being 5-8 times higher than the first average concentration.

[0028] The plurality of nucleic acids includes cell-free nucleic acids from, or nucleic acids prepared from, a first biological sample of a first subject.

[0029] In another aspect, the present disclosure provides a method for concentrating a target nucleic acid. The method includes contacting a plurality of nucleic acids, including the target nucleic acid, with a probe set under hybridizing conditions, the probe set comprising a first set of polynucleotide probes (e.g., non-enhanced probes) collectively targeting a first plurality of genomic regions at an average coverage of 0.75x to 1.25x, the first set of polynucleotide probes comprising a first plurality of polynucleotide probe species. Each respective polynucleotide probe species in the first plurality of polynucleotide probe species targets a respective genomic region in the first plurality of genomic regions, and the polynucleotide probe species in the first plurality of polynucleotide probe species are present in the composition at a first average molar concentration. The probe set comprises a second set of polynucleotide probes (e.g., enhanced probes) collectively targeting a second plurality of genomic regions at an average coverage of 0.75x to 1.25x, the second set of polynucleotide probes comprising a second plurality of polynucleotide probe species. Each respective polynucleotide probe species in the second plurality of polynucleotide probe species targets a respective genomic region in the second plurality of genomic regions, and the polynucleotide probe species in the second plurality of polynucleotide probe species are present in the composition at a second average molar concentration, the second average molar concentration being 5 to 8 times higher than the first average concentration. The plurality of nucleic acids includes cell-free nucleic acids derived from or prepared from a first biological sample of a first subject.

[0030] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief explanation of the drawings]

[0031] [Figure 1A]1A-1D collectively illustrate block diagrams of an exemplary computing device for providing clinical support for personalized cancer therapy based on sequencing of cell-free DNA, in accordance with some embodiments of the present disclosure. [Figure 1B] 1A-1D collectively illustrate block diagrams of an exemplary computing device for providing clinical support for personalized cancer therapy based on sequencing of cell-free DNA, in accordance with some embodiments of the present disclosure. [Figure 2A] 1 illustrates an exemplary workflow for generating a clinical report based on information generated from the analysis of one or more patient samples, according to some embodiments of the present disclosure. [Figure 2B] 1 illustrates an example of a distributed diagnostic environment for collecting and evaluating patient data for the purposes of precision oncology, according to some embodiments of the present disclosure. [Figure 3] 1 provides an exemplary flowchart of processes and features for liquid biopsy sample collection and analysis for use in precision oncology, according to some embodiments of the present disclosure. [Figure 4A] 1 collectively illustrates exemplary steps of a bioinformatics pipeline for precision oncology, according to various embodiments of the present disclosure. 2 provides a summary flowchart of processes and features in a bioinformatics pipeline, according to some embodiments of the present disclosure. [Figure 4B] An overview of the bioinformatics pipeline performed using either liquid biopsy samples alone or liquid biopsy samples and matched normal samples is provided. [Figure 4C] FIG. 10 illustrates that paired-end reads from tumor and normal isolates are zipped and stored separately under the same sequence identifier, according to some embodiments of the present disclosure. [Figure 4D] 1 illustrates quality correction of FASTQ files according to some embodiments of the present disclosure. [Figure 4E] 1 illustrates a process for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 5A]1A-1D collectively illustrate flow charts of processes and features for generating cell-free DNA sequencing data using improved probe sets, according to some embodiments of the present disclosure. [Figure 5B] 1A-1D collectively illustrate flow charts of processes and features for generating cell-free DNA sequencing data using improved probe sets, according to some embodiments of the present disclosure. [Figure 6A] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6B] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6C] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6D] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6E] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6F] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6G] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6H] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6I]1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6J] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6K] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6L] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 6M] 1A-1D collectively illustrate exemplary nucleic acids targeted for enrichment and variant detection using one or more probes, according to some embodiments of the present disclosure. [Figure 7] 1 illustrates sequencing coverage obtained using a probe set comprising a first set of polynucleotide probes and a second set of polynucleotide probes at various concentration ratios, according to an embodiment of the present disclosure. [Figure 8] 1 illustrates the sensitivity of a probe set comprising a first set of polynucleotide probes and a second set of polynucleotide probes when performing singleplex and multiplex library hybridization enrichment according to embodiments of the present disclosure. [Figure 9] 1 illustrates a schematic diagram of relative unique read coverage and total read coverage obtained using a first set of polynucleotide probes and a second set of polynucleotide probes with various amounts of nucleic acid input, according to some embodiments of the present disclosure. [Figure 10A] 1A-1D collectively illustrate exemplary microsatellite regions in the human genome useful for determining the MSI status of a sample, according to some embodiments of the present disclosure. [Figure 10B]1A-1D collectively illustrate exemplary microsatellite regions in the human genome useful for determining the MSI status of a sample, according to some embodiments of the present disclosure. [Figure 11] Figure 1 illustrates post-duplicate removal coverage for each DNA input and probe ratio condition, divided by target type, according to an embodiment of the present disclosure. The Y-axis indicates coverage, and each dot represents one target region for one sample in the condition. The X-axis indicates conditions with 30 ng DNA input and probe ratio on the left and 10 ng DNA input and probe ratio on the right. Reinforced and non-reinforced targets were plotted separately to illustrate the difference in coverage for each probe ratio. Because there was no increase in coverage for the reinforced target with increasing probe ratio, the coverage of the reinforced target (left boxplot for each pair of boxplots) appeared to reach an upper limit. Non-reinforced coverage (right boxplot for each pair of boxplots) appeared to decrease at a probe ratio of 1:5:1 relative to a probe ratio of 1:1:1. [Figure 12] Figure 11 illustrates pre-duplicate removal coverage for each DNA input and probe ratio condition, divided by target type, according to an embodiment of the present disclosure. The Y-axis is coverage, and each dot represents one target region for one sample in the condition. The X-axis indicates conditions with 30 ng DNA input and probe ratio on the left and 10 ng DNA input and probe ratio on the right. Enhanced and non-enhanced targets were plotted separately to illustrate the difference in coverage for each probe ratio. Enhanced targets (left boxplot for each pair of boxplots) did not reach an upper limit, indicating that the probe ratio caused differences in coverage in some cases. However, as shown in Figure 11, there was an upper limit to the number of unique fragments identified in the samples. [Figure 13] 1 illustrates enriched pre-duplicate-removed read counts by enriched PCR duplication rate, according to an embodiment of the present disclosure. The enriched pre-duplicate-removed read counts (y-axis, 100 million) are plotted against the enriched PCR duplication rate (x-axis), with dots corresponding to one sample and color-coded based on the condition (DNA input and probe ratio). [Figure 14] 1 illustrates the number of unenhanced pre-duplicate-removed reads by the enhanced PCR duplication rate, according to an embodiment of the present disclosure. The number of unenhanced pre-duplicate-removed reads (y-axis, 100 million) is plotted against the unenhanced PCR duplication rate (x-axis), with dots corresponding to one sample and color-coded based on the condition (DNA input and probe ratio). [Figure 15] 1 illustrates the total number of reads (y-axis, 100 million) for each sample according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). The total number of reads was greater than the expected total number of reads expected for the combined reinforced and non-reinforced probes. [Figure 16] 1 illustrates the number of unique reads (y-axis, 100 million) for each sample according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). The number of unique reads was greater than the expected number of unique reads expected for the combined enhanced and non-enhanced probes, and there was no significant change in the unique reads identified for each probe ratio. [Figure 17] 1 illustrates PCR duplication rates (y-axis) by UMI, according to embodiments of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). PCR duplication rates were lower for the 30 ng DNA input, as expected. [Figure 18] 1 illustrates the on-target rate (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). The on-target rate appeared to decrease at higher probe ratios for the enhanced probes, indicating that the enhanced probes, in some cases, caused more off-target reads. [Figure 19]1 illustrates the on-target rate (y-axis) before duplicate removal according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe (legend) and DNA input (x-axis, 10 ng or 30 ng). The on-target rate before duplicate removal increased compared to the on-target rate after duplicate removal and decreased at probe ratios favoring the enhanced probe, consistent with the on-target rate after duplicate removal shown in FIG. 18. [Figure 20] Exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient is likely to be resistant to an immune cancer therapy. [Figure 21A] Collectively provided are exemplary microsatellite genomic regions that are useful for determining the microsatellite stability status of a patient, targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 21B] Collectively provided are exemplary microsatellite genomic regions that are useful for determining the microsatellite stability status of a patient, targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 21C] Collectively provided are exemplary microsatellite genomic regions that are useful for determining the microsatellite stability status of a patient, targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 21D] Collectively provided are exemplary microsatellite genomic regions that are useful for determining the microsatellite stability status of a patient, targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 21E]Collectively provided are exemplary microsatellite genomic regions that are useful for determining the microsatellite stability status of a patient, targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure. [Figure 22A] Figures 22A and 22B collectively illustrate pre- and post-duplication removal coverage for each DNA input and probe ratio condition, divided by enriched and non-enriched targets, according to embodiments of the present disclosure. Pre-duplication coverage is shown in Figure 22A. Post-duplication coverage is shown in Figure 22B. The Y-axis is coverage, and each dot represents one target region for one sample in the condition. The X-axis indicates conditions with 30 ng DNA input and probe ratio on the left and 10 ng DNA input and probe ratio on the right. Enriched (left boxplot for each pair of boxplots) and non-enriched (right boxplot for each pair of boxplots) coverage appeared to increase and decrease, respectively, in correlation with the ratio of enriched probes to other probes. [Figure 22B] Figures 22A and 22B collectively illustrate pre- and post-duplication removal coverage for each DNA input and probe ratio condition, divided by enriched and non-enriched targets, according to embodiments of the present disclosure. Pre-duplication coverage is shown in Figure 22A. Post-duplication coverage is shown in Figure 22B. The Y-axis is coverage, and each dot represents one target region for one sample in the condition. The X-axis indicates conditions with 30 ng DNA input and probe ratio on the left and 10 ng DNA input and probe ratio on the right. Enriched (left boxplot for each pair of boxplots) and non-enriched (right boxplot for each pair of boxplots) coverage appeared to increase and decrease, respectively, in correlation with the ratio of enriched probes to other probes. [Figure 23A]Figures 23A and 23B collectively illustrate enriched and unenriched pre-duplicate-removed read counts (y-axis, tens of millions and hundreds of millions, respectively) by enriched PCR duplication rate (x-axis, fractions), according to embodiments of the present disclosure. The enriched pre-duplicate-removed read counts are shown in Figure 23A. The unenriched pre-duplicate-removed read counts are shown in Figure 23B. Dots represent single samples, with colors based on condition (DNA input and probe ratio). The enriched pre-duplicate-removed read counts and PCR duplication rate generally appeared to increase with molar ratio. The unenriched pre-duplicate-removed read counts and PCR duplication rate appeared to be mutually exclusive, with PCR duplication rate primarily dependent on DNA input (lower rates for higher DNA inputs), while the pre-duplicate-removed read count decreased with increasing molar ratio, as expected in some cases. [Figure 23B] Figures 23A and 23B collectively illustrate enriched and unenriched pre-duplicate-removed read counts (y-axis, tens of millions and hundreds of millions, respectively) by enriched PCR duplication rate (x-axis, fractions), according to embodiments of the present disclosure. The enriched pre-duplicate-removed read counts are shown in Figure 23A. The unenriched pre-duplicate-removed read counts are shown in Figure 23B. Dots represent single samples, with colors based on condition (DNA input and probe ratio). The enriched pre-duplicate-removed read counts and PCR duplication rate generally appeared to increase with molar ratio. The unenriched pre-duplicate-removed read counts and PCR duplication rate appeared to be mutually exclusive, with PCR duplication rate primarily dependent on DNA input (lower rates for higher DNA inputs), while the pre-duplicate-removed read count decreased with increasing molar ratio, as expected in some cases. [Figure 24] 1 illustrates the total number of reads (y-axis, 100 million) according to an embodiment of the present disclosure. Dots represent single samples, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). There was a clear difference in the total number of reads between the two DNA input weights. [Figure 25] 1 illustrates the number of unique reads for each sample (y-axis, 10 million) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). [Figure 26]1 illustrates PCR duplication rates (y-axis) by UMI, according to embodiments of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). PCR duplication rates were lower for the 30 ng DNA input, as expected in some cases. [Figure 27] 1 illustrates the on-target rate (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 10 ng or 30 ng). The on-target rate appeared to decrease at higher probe ratios for the enhanced probes, indicating that the enhanced probes, in some cases, caused more off-target reads. [Figure 28] 27 illustrates the on-target rate (y-axis) before duplicate removal according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe (legend) and DNA input (x-axis, 10 ng or 30 ng). The on-target rate before duplicate removal was similar to the on-target rate after duplicate removal and decreased at probe ratios favoring the enhanced probe, consistent with the on-target rate after duplicate removal illustrated in FIG. 27. [Figure 29A] Collectively provided are exemplary viral genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant viral infection. [Figure 29B] Collectively provided are exemplary viral genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant viral infection. [Figure 30]1 illustrates the enriched-to-unenriched pre-overlap removal coverage ratio as a function of enrichment probe molar ratio, according to an embodiment of the present disclosure. Considering the equation for the line of best fit, the optimal ratio to achieve an enriched-to-unenriched coverage of 4:1 was 1:5.5:1. The R value of 0.993 indicated high confidence in the line of best fit. [Figure 31A] To illustrate the difference in coverage for each probe ratio according to an embodiment of the present disclosure, the pre-redundancy and post-redundancy coverage for each sequencing depth and probe ratio condition are collectively illustrated, divided by the enriched target (the left boxplot in each pair of boxplots) and the non-enriched target (the right boxplot in each pair of boxplots). Pre-redundancy coverage is shown in Figure 31A. Post-redundancy coverage is shown in Figure 31B. The Y axis is coverage, and each dot represents one target region for one sample in the condition. The X axis indicates conditions with 1x depth and probe ratio on the left and 2.5x depth + probe ratio on the right. [Figure 31B] To illustrate the difference in coverage for each probe ratio according to an embodiment of the present disclosure, the pre-redundancy and post-redundancy coverage for each sequencing depth and probe ratio condition are collectively illustrated, divided by the enriched target (the left boxplot in each pair of boxplots) and the non-enriched target (the right boxplot in each pair of boxplots). Pre-redundancy coverage is shown in Figure 31A. Post-redundancy coverage is shown in Figure 31B. The Y axis is coverage, and each dot represents one target region for one sample in the condition. The X axis indicates conditions with 1x depth and probe ratio on the left and 2.5x depth + probe ratio on the right. [Figure 32A] Figure 1 collectively illustrates the number of enriched and unenriched pre-duplicate-removed reads (y-axis, 100 million) by enriched PCR duplication rate (x-axis, fraction), according to an embodiment of the present disclosure. Dots represent single samples, with colors based on condition (DNA input and probe ratio). Higher sequencing depth resulted in higher PCR duplication rates, as expected in some cases, but this was more pronounced for enriched (10% higher) versus unenriched (5% higher). [Figure 32B] Figure 1 collectively illustrates the number of enriched and unenriched pre-duplicate-removed reads (y-axis, 100 million) by enriched PCR duplication rate (x-axis, fraction), according to an embodiment of the present disclosure. Dots represent single samples, with colors based on condition (DNA input and probe ratio). Higher sequencing depth resulted in higher PCR duplication rates, as expected in some cases, but this was more pronounced for enriched (10% higher) versus unenriched (5% higher). [Figure 33] 1 illustrates the total number of reads (y-axis, 100 million) according to an embodiment of the present disclosure. Dots represent single samples, and box plots are shown for each probe ratio (legend) and sequencing depth (x-axis, 1x or 2.5x). The total number of reads was higher than expected for the combined reinforced and non-reinforced probes. The total number of reads appeared to correspond to the sequencing depth, as expected in some cases. [Figure 34] 1 illustrates the number of unique reads for each sample (y-axis, 100 million) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and sequencing depth (x-axis, 1x or 2.5x). [Figure 35] 1 illustrates PCR duplication rate (y-axis) by UMI according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and sequencing depth (x-axis, 1x or 2.5x). [Figure 36] 1 illustrates the on-target rate (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and sequencing depth (x-axis, 1x or 2.5x). The on-target rate was lower at higher sequencing depths, and in some cases showed a higher number of reads (e.g., mapped away from the target region). [Figure 37] 1 illustrates the on-target rate (y-axis) before duplicate removal according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe (legend) and sequencing depth (x-axis, 1x or 2.5x). [Figure 38] 1 illustrates the median enriched-to-unenriched coverage ratio before duplicate removal as a function of enrichment probe molar ratio, according to an embodiment of the present disclosure. Considering the equation of the line of best fit, the optimal ratio to achieve 4:1 enriched-to-unenriched coverage was 1:5.5:1. The R2 values ​​of 0.974 and 0.979 for 1x and 2x sequencing depths, respectively, indicated high confidence in the line of best fit. [Figure 39A] To illustrate the difference in coverage according to an embodiment of the present disclosure, the pre- and post-duplication removal coverage for each DNA input is collectively illustrated, divided by enriched targets (left box plot), BRCA1 / 2 targets (middle box plot), and non-enriched targets (right box plot). The Y axis is coverage, and each dot represents one target region for one sample in a condition. The X axis indicates the condition. [Figure 39B] To illustrate the difference in coverage according to an embodiment of the present disclosure, the pre- and post-duplication removal coverage for each DNA input is collectively illustrated, divided by enriched targets (left box plot), BRCA1 / 2 targets (middle box plot), and non-enriched targets (right box plot). The Y axis is coverage, and each dot represents one target region for one sample in a condition. The X axis indicates the condition. [Figure 40A] Figures 40A and 40B collectively illustrate enrichment, BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) by enrichment PCR duplication rate (x-axis, fraction), according to embodiments of the present disclosure. Enriched targets are shown in Figure 40A. BRCA1 / 2 targets are shown in Figure 40B. Non-enriched targets are shown in Figure 40C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). Higher DNA input resulted in lower PCR duplication rates, as expected in some cases, with no change in pre-duplicate removal read counts. [Figure 40B]Figures 40A and 40B collectively illustrate enrichment, BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) by enrichment PCR duplication rate (x-axis, fraction), according to embodiments of the present disclosure. Enriched targets are shown in Figure 40A. BRCA1 / 2 targets are shown in Figure 40B. Non-enriched targets are shown in Figure 40C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). Higher DNA input resulted in lower PCR duplication rates, as expected in some cases, with no change in pre-duplicate removal read counts. [Figure 40C] Figures 40A and 40B collectively illustrate enrichment, BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) by enrichment PCR duplication rate (x-axis, fraction), according to embodiments of the present disclosure. Enriched targets are shown in Figure 40A. BRCA1 / 2 targets are shown in Figure 40B. Non-enriched targets are shown in Figure 40C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). Higher DNA input resulted in lower PCR duplication rates, as expected in some cases, with no change in pre-duplicate removal read counts. [Figure 41] 1 illustrates total reads (y-axis) according to an embodiment of the present disclosure. Dots represent single samples, and box plots are shown for each probe ratio (legend) and DNA input (30 ng or 45 ng). [Figure 42] 1 illustrates the number of unique reads for each sample (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 43] 1 illustrates PCR duplication rate (y-axis) by UMI, according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 44] 1 illustrates the on-target rate (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 45]1 illustrates the on-target rate (y-axis) before duplicate removal according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 46A] To illustrate the difference in coverage for each DNA input according to an embodiment of the present disclosure, the pre- and post-removal coverage for each DNA input is collectively illustrated, divided by enriched targets (left box plot), BRCA1 / 2 targets (middle box plot), and non-enriched targets (right box plot). The Y axis is coverage, and each dot represents one target region for one sample in a condition. The X axis indicates the condition. [Figure 46B] To illustrate the difference in coverage for each DNA input according to an embodiment of the present disclosure, the pre- and post-removal coverage for each DNA input is collectively illustrated, divided by enriched targets (left box plot), BRCA1 / 2 targets (middle box plot), and non-enriched targets (right box plot). The Y axis is coverage, and each dot represents one target region for one sample in a condition. The X axis indicates the condition. [Figure 47A] Figures 47A and 47B collectively illustrate enrichment by enriched PCR duplication rate (x-axis, fraction), BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) according to embodiments of the present disclosure. Enriched targets are shown in Figure 47A. BRCA1 / 2 targets are shown in Figure 47B. Non-enriched targets are shown in Figure 47C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). [Figure 47B] Figures 47A and 47B collectively illustrate enrichment by enriched PCR duplication rate (x-axis, fraction), BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) according to embodiments of the present disclosure. Enriched targets are shown in Figure 47A. BRCA1 / 2 targets are shown in Figure 47B. Non-enriched targets are shown in Figure 47C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). [Figure 47C]Figures 47A and 47B collectively illustrate enrichment by enriched PCR duplication rate (x-axis, fraction), BRCA1 / 2, and non-enriched pre-duplicate removal read counts (y-axis) according to embodiments of the present disclosure. Enriched targets are shown in Figure 47A. BRCA1 / 2 targets are shown in Figure 47B. Non-enriched targets are shown in Figure 47C. Dots represent single samples, with colors based on condition (DNA input and probe ratio). [Figure 48] 1 illustrates total reads (y-axis) according to an embodiment of the present disclosure. Dots represent single samples, and box plots are shown for each probe ratio (legend) and DNA input (30 ng or 45 ng). [Figure 49] 1 illustrates the number of unique reads for each sample (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 50] 1 illustrates PCR duplication rate (y-axis) by UMI, according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 51] 1 illustrates the on-target ratio (y-axis) according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe ratio (legend) and DNA input (x-axis, 1x or 2.5x). [Figure 52] 1 illustrates the on-target rate (y-axis) before duplicate removal according to an embodiment of the present disclosure. Individual samples are shown as dots, and box plots are shown for each probe (legend) and DNA input (x-axis, 30 ng or 45 ng). [Figure 53] 1 illustrates the change in allele frequency of insertion-deletion sites greater than 10 bp using a control single library and multiplexed libraries for hybridization reactions according to embodiments of the present disclosure. [Figure 54]Exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient is likely to be resistant to androgen receptor therapy, e.g., androgen receptor antagonists, or other therapies that target, modulate, or interact with the androgen receptor. [Figure 55] Exemplary genomic regions in the BRCA1 and BRCA2 genes that are targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure and that are useful for determining whether a patient has a clinically relevant homologous recombination repair deficiency mutation are provided. [Figure 56A] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56B] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56C] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56D] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56E]Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56F] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56G] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56H] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56I] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56J] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56K] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56L]Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56M] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56N] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56O] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56P] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56Q] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56R] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56S]Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56T] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56U] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56V] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56W] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 56X] Collectively provided are exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure that are useful for determining whether a patient has a clinically relevant copy number variation. [Figure 57A] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57B]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57C] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57D] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57E] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57F] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57G] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57H] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57I]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57J] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57K] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57L] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57M] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57N] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57O] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57P]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57Q] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57R] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57S] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57T] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57U] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57V] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57W]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57X] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 57Y] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58A] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58B] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58C] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58D] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58E]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58F] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58G] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58H] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58I] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58J] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58K] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58L]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58M] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58N] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58O] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58P] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58Q] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58R] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58S]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58T] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58U] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58V] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58W] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58X] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58YA] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58YB]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58Z] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AA] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AB] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AC] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Fig. 58AD] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AE] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AF]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AG] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AH] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AI] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AJ] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AK] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AL] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AM]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AN] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AO] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AP] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AQ] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AR] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AS] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AT]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AU] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AV] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AW] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AX] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AY] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58AZ] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BA]Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BB] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BC] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BD] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BE] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant. [Figure 58BF] Collectively, exemplary genomic regions targeted for enrichment and variant detection using one or more probes according to some embodiments of the present disclosure are provided that are useful for determining whether a patient has a clinically relevant variant.

[0032] Like reference numbers refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION OF THE INVENTION

[0033] Introduction One aspect of precision oncology and / or next-generation sequencing assay design is the selection and concentration of probes used to identify specific regions of the genome. In some cases, biological samples, such as liquid biopsy assays, contain nucleic acids from multiple different genomic regions, and two or more of the multiple different genomic regions may have different detection limits. For example, as described in the background section above, the fraction of cfDNA fragments derived from cancer cells in a liquid biopsy sample (the "tumor fraction" or "ctDNA fraction" of the sample / subject) can range from nearly 0% to well over 50%. In some cases, clonal heterogeneity can result in dilution of one or more clonal populations far below the overall tumor fraction of the sample. Furthermore, the frequency of genomic alterations from cancer tissue can vary from locus to locus based on at least (i) their prevalence in different subclonal populations of the subject's cancer and (ii) their location in the genome relative to large-scale chromosomal copy number variations. Thus, the amount of DNA released from cancer tissue for different genomic targets can vary greatly in a single given sample, hindering accurate measurement of disease indicators and actionable genomic alterations.

[0034] In particular, enrichment of regions with lower LODs using a standard probe panel for target genes can be difficult because genomic regions with low detection limits may be underrepresented in the resulting sequencing data. Therefore, the compositions, methods, and systems of the present disclosure provide improved probe sets that address the different LODs between genomic regions with a baseline of detection (e.g., non-enhanced genes) and genomic regions with a lower LOD than the baseline (e.g., enriched genes). In particular, the present disclosure describes improved probe sets obtained by adjusting the ratio of probe molar concentrations between enriched and non-enhanced genomic regions.

[0035] Advantageously, as described in Examples 2, 3, and 13 with reference to Figures 7 and 8, the disclosed compositions and methods enable hybridization and capture of non-enhanced genes with a detection limit of 0.01 (1%) allele fraction or less, and exemplary enriched genes with a detection limit of 0.0025 (0.25%) allele fraction or less. Using an improved ratio of enriched and non-enhanced probes, in some embodiments, sequencing coverage of enriched genes is increased without losing sensitivity to either enriched or non-enhanced probes when performing hybridization enrichment (e.g., DNA input) of multiplexed libraries. Thus, the compositions and methods described herein adjust the sensitivity and specificity of probes for target nucleic acid enrichment in a locus-specific manner to achieve higher accuracy of true variant calling in liquid biopsy assays. The improved performance of the disclosed compositions and methods, including enriched and non-enhanced probes, is further illustrated by the schematic diagram in Figure 9, in which the coverage of unique reads per total reads sequenced is increased when using enriched probes compared to non-enhanced probes. Even greater effectiveness can be observed as the sample nucleic acid input (e.g., mass of DNA input) increases relative to the amount of probe, with coverage of unique reads per total reads sequenced increasing further for enhanced probes with minimal loss of sensitivity relative to non-enhanced probes.

[0036] The methods and systems described herein also improve precision oncology methods for allocating and / or administering treatment through improved accuracy of variation detection. Identification of therapeutically actionable variants, which can be included in clinical reports for patient and / or clinician review and / or matched with appropriate therapies and / or clinical trials for treatment and / or monitoring, allows for more accurate allocation of treatment. Furthermore, elimination of false-positive variant detection reduces the risk of patients receiving unnecessary or potentially harmful regimens due to misdiagnosis.

[0037] definition As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, a human (e.g., a male human, a female human, a fetus, a pregnant woman, a child, or the like), a non-human mammal, or a non-human animal. Any human or non-human animal can serve as a subject, including, but not limited to, a mammal, a reptile, a bird, an amphibian, a fish, a ungulate, a ruminant, a bovine (e.g., a cow), an equine (e.g., a horse), a caprine and ovine (e.g., a sheep, a goat), a porcine (e.g., a pig), a camelid (e.g., a camel, a llama, an alpaca), a monkey, an ape (e.g., a gorilla, a chimpanzee), a ursidae (e.g., a bear), a fowl, a dog, a cat, a mouse, a rat, a fish, a dolphin, a whale, and a shark. In some embodiments, the subject is a male or female (e.g., a man, a woman, or a child) of any age.

[0038] As used herein, "control," "control sample," "reference," "reference sample," "normal," and "normal sample" describe a sample derived from non-diseased tissue. In some embodiments, such a sample is derived from a subject that does not have a particular condition (e.g., cancer). In other embodiments, such a sample is an internal control, e.g., from a subject that may or may not have a particular disease (e.g., cancer), but is derived from the subject's healthy tissue. For example, if a liquid or solid tumor sample is obtained from a subject with cancer, an internal control sample can be obtained from the subject's healthy tissue, e.g., a white blood cell sample from a subject that does not have a blood cancer, or a solid germline tissue sample from the subject. Thus, a reference sample can be obtained from the subject or from a database, e.g., from a second subject that does not have a particular disease (e.g., cancer).

[0039] As used herein, the terms "cancer," "cancer tissue," or "tumor" refer to an abnormal mass of tissue, including both a solid mass (e.g., in the case of solid tumors) or a fluid mass (e.g., in the case of blood cancers), in which the growth of the mass exceeds and is out of step with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are well differentiated, have characteristically slower growth than malignant tumors, and may remain localized to the site of origin. In addition, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors may be poorly differentiated (dysplastic) and have characteristically rapid growth accompanied by progressive infiltration, invasion, and destruction of surrounding tissue. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is out of step with that of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained from or derived from a tumor in a subject, as described herein.

[0040] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, bile duct cancer, adrenal cancer, neural carcinoma, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, non-small cell lung cancer, thymus cancer, tumors, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, serous ovarian cancer, HR+ breast cancer, serous uterine cancer, endometrial cancer, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.

[0041] As used herein, the term "cancer state" or "cancer condition" refers to characteristics of a cancer patient's condition, such as diagnostic status, cancer type, cancer location, cancer primary origin, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject, such as age, sex, weight, race, personal habits (e.g., smoking, alcohol use, diet), other relevant medical conditions (e.g., high blood pressure, dry skin, other diseases), current medications, allergies, relevant medical history, current side effects of cancer treatments and other medications, etc., are used to further describe the subject's cancer state or condition.

[0042] As used herein, the term "liquid biopsy" sample refers to a liquid sample obtained from a subject that contains cell-free DNA. Examples of liquid biopsy samples include, but are not limited to, a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal material, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. In some embodiments, the liquid biopsy sample is a cell-free sample, e.g., a cell-free blood sample. In some embodiments, the liquid biopsy sample is obtained from a subject with cancer. In some embodiments, the liquid biopsy sample is collected from a subject with an unknown cancerous condition, e.g., for use in determining the subject's cancerous condition. Similarly, in some embodiments, the liquid biopsy is collected from a subject with a non-cancerous disorder, e.g., cardiovascular disease. In some embodiments, the liquid biopsy is collected from a subject with an unknown non-cancerous disorder, e.g., for use in determining the subject's non-cancerous disorder condition.

[0043] As used herein, the terms "cell-free DNA" and "cfDNA" refer interchangeably to DNA fragments circulating in a subject's body (e.g., bloodstream) and derived from one or more healthy cells and / or one or more cancer cells. These DNA molecules are found outside of cells in bodily fluids such as a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal matter, saliva, sweat, perspiration, tears, pleural fluid, pericardial fluid, or peritoneal fluid, and are considered to be fragments of genomic DNA excreted from healthy cells and / or cancer cells, for example, during apoptosis and dissolution of the cell envelope.

[0044] As used herein, the term "locus" refers to a location (e.g., site) within a genome, for example, on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position on a particular chromosome within a genome. In some embodiments, a locus refers to a group of nucleotide positions within a genome. In some cases, a locus is defined by a mutation (e.g., a substitution, insertion, deletion, inversion, or translocation) of consecutive nucleotides within a cancer genome. In some cases, a locus is defined by a gene, a subgenic structure (e.g., a regulatory element, exon, intron, or a combination thereof), or a defined span of a chromosome. Because normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus located on an autosome, for example, one copy on the maternal autosome and one copy on the paternal autosome.

[0045] As used herein, the term "allele" refers to a specific sequence of one or more nucleotides at a chromosomal locus. In a haploid organism, a subject has one allele at every chromosomal locus. In a diploid organism, a subject has two alleles at every chromosomal locus.

[0046] As used herein, the term "base pair" or "bp" refers to a unit consisting of two nucleic acid bases joined together by hydrogen bonds. Generally, the size of an organism's genome is measured in base pairs because DNA is typically double-stranded. However, some viruses have single-stranded DNA or RNA genomes.

[0047] As used herein, the terms "genomic alteration," "mutation," and "variant" refer to detectable changes in the genetic material of one or more cells. Genomic alterations, mutations, or variants can refer to various types of changes in a cell's genetic material, including changes in the primary genomic sequence at single or multiple nucleotide positions, such as single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels (e.g., nucleotide insertions or deletions), DNA rearrangements (e.g., inversions or translocations of portions of a chromosome or multiple chromosomes), copy number variations of loci (e.g., exons, genes, or large spans of chromosomes) (CNVs), partial or complete changes in the ploidy of a cell, and changes in the epigenetic information of a genome, such as altered DNA methylation patterns. In some embodiments, a mutation is a change in a cell's genetic information relative to one or more "normal" alleles found in a particular reference genome or population of a species of interest. For example, mutations can be found in both a subject's germline cells (e.g., non-cancerous "normal" cells) and in a subject's abnormal cells (e.g., pre-cancerous or cancerous cells). Thus, mutations in a subject's germline (e.g., found in substantially all "normal" cells in a subject) are identified relative to the reference genome for the subject's species. However, many loci in a species' reference genome are associated with several variant alleles that are significantly represented in the subject's population and are not associated with a pathological condition, e.g., such that they are not considered "mutations." In contrast, in some embodiments, mutations in a subject's cancer cells can be identified relative to either the subject's reference genome or the subject's own germline genome. In certain cases, identifying both types of variants can be beneficial. For example, in some cases, mutations present in both the subject's cancer genome and the subject's germline can be beneficial for precision oncology when the mutations are so-called "driver mutations" that contribute to the initiation and / or development of cancer. However, in other cases, mutations present in both the subject's cancer genome and the subject's germline are not beneficial for precision oncology when, for example, the mutations are so-called "passenger mutations" that do not contribute to the initiation and / or development of cancer.Similarly, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is beneficial for precision oncology, e.g., if the mutation is a driver mutation and / or the mutation facilitates a therapeutic approach, e.g., by distinguishing cancer cells from normal cells in a therapeutically actionable manner. However, in some cases, a mutation present in a subject's cancer genome but not in the subject's germline is not beneficial for precision oncology, e.g., if the mutation is a passenger mutation and / or the mutation does not distinguish cancer cells from germline cells in a therapeutically actionable manner.

[0048] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the dominant allele represented at that chromosomal locus within a population of a species (e.g., a "wild-type" sequence) or a predefined allele within a reference genome of the species.

[0049] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is not the dominant allele represented at that locus within a population of a species (e.g., is not a "wild-type" sequence) or is not a predefined allele within a reference sequence construct for the species (e.g., a reference genome or set of reference genomes). In some cases, sequence isoforms found within a population of a species that result in amino acid substitutions that do not affect the alteration of a protein encoded by the genome or that do not substantially affect the function of the encoded protein are not variant alleles.

[0050] As used herein, the terms "variant allele fraction," "VAF," "allele fraction," or "AF" refer to the number of times a variant or mutant allele is observed (e.g., the number of reads supporting the candidate variant allele) divided by the total number of times the position was sequenced (e.g., the total number of reads covering the candidate locus).

[0051] As used herein, the terms "variant fragment count" and "variant allele fragment count" refer interchangeably to a quantification, e.g., raw count or normalized count, of the number of sequences representing unique cell-free DNA fragments encompassing a variant allele in a sequencing reaction. That is, the variant fragment count represents the count of sequence reads representing unique molecules in a liquid biopsy sample after overlapping sequence reads in the raw sequencing data have been collapsed, e.g., through the use of uniform molecular identifiers (UMIs) and bagging, as described herein.

[0052] As used herein, the term "germline variant" refers to a genetic variant inherited from maternal and paternal DNA. Germline variants can be determined through a matched tumor-normal calling pipeline.

[0053] As used herein, the term "somatic variant" refers to a variant that arises as a result of dysregulation, e.g., mutation, of a cellular process associated with a neoplastic cell. Somatic variants can be detected via subtraction from a matched normal sample.

[0054] As used herein, the term "single nucleotide variant" or "SNV" refers to the substitution of one nucleotide with a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y can be represented as "X>Y." For example, a cytosine to thymine SNV can be represented as "C>T."

[0055] As used herein, the terms "insertions and deletions" or "indels" refer to variants that result from the gain or loss of DNA base pairs within the analyzed region.

[0056] As used herein, the term "copy number variation" or "CNV" refers to a process for detecting large structural changes in the genome associated with tumor aneuploidy and other dysregulation of repair systems. These processes are used to detect large insertions or deletions of entire genomic regions. CNV is defined as a structural insertion or deletion whose size is greater than a certain base pair ("bp"), such as 500 bp.

[0057] As used herein, the term "gene fusion" refers to the product of large-scale chromosomal abnormalities that result in the production of chimeric proteins. These expressed products may be non-functional or may be highly overactive or underactive. This may lead to deleterious effects in cancer, such as a hyperproliferative or anti-apoptotic phenotype.

[0058] As used herein, the term "loss of heterozygosity" refers to the loss of one copy of a segment (e.g., including part or all of one or more genes) of the genome of a diploid subject (e.g., a human) in a tissue of a subject, e.g., cancer tissue, or the loss of one copy of a sequence encoding a functional gene product in the genome of a diploid subject. As used herein, when referring to a metric representing loss of heterozygosity across a subject's genome, the loss of heterozygosity is caused by the loss of one copy of various segments in the subject's genome. Genome-wide loss of heterozygosity can be estimated without sequencing the subject's entire genome, and methods for such estimation based on gene panel targeted sequencing methodologies have been described in the art. Thus, in some embodiments, a metric representing loss of heterozygosity across the genome of a subject's tissue is expressed as a single value, e.g., a percentage or fraction of the genome. In some cases, a tumor may be composed of various subclonal populations, each of which may have different degrees of loss of heterozygosity across their respective genomes. Thus, in some embodiments, genome-wide loss of heterozygosity of a cancer tissue refers to the average loss of heterozygosity across a heterogeneous tumor population. As used herein, when referring to the metric of loss of heterozygosity in a particular gene, for example, a DNA repair protein (e.g., BRCA1 or BRCA2), such as a protein involved in the homologous DNA recombination pathway, loss of heterozygosity refers to the complete or partial loss of one copy of the protein-encoding gene in the genome of a tissue, and / or a mutation in one copy of the gene that prevents translation of the full-length gene product, such as a frameshift or truncation mutation (generating a premature stop codon in the gene) in the gene of interest. In some cases, tumors are composed of various subclonal populations, each of which may have a different mutation status in the gene of interest. Thus, in some embodiments, loss of heterozygosity for a particular gene of interest is represented by the average value of loss of heterozygosity for the gene across all sequenced subclonal populations of the cancer tissue.In other embodiments, loss of heterozygosity for a particular gene of interest is represented by the number of unique occurrences of loss of heterozygosity in the gene of interest across all sequenced subclonal populations of cancer tissue (e.g., the number of unique frameshift and / or truncating mutations in the gene identified in the sequencing data).

[0059] As used herein, the term "microsatellite" refers to a short, repeated sequence of DNA. The minimal nucleotide repeating unit of a microsatellite is referred to as a "repeated unit" or "repeat unit." In some embodiments, the stability of a microsatellite locus is assessed by comparing a metric of the distribution of the number of repeat units at the microsatellite locus to a reference number or distribution.

[0060] As used herein, the term "microsatellite instability" or "MSI" refers to a genetic hypermutability state associated with various cancers resulting from impaired DNA mismatch repair (MMR) in a subject. Among other phenotypes, MSI causes changes in the size of microsatellite loci, e.g., changes in the number of repeat units at a microsatellite locus, during DNA replication. Thus, the size of the microsatellite repeat is altered in MSI cancers compared to the size of the corresponding microsatellite repeat in the germline of the cancer subject. The term "microsatellite instability-high" or "MSI-H" refers to a cancer (e.g., tumor) state with a significant MMR defect that results in microsatellite loci having lengths significantly different from those of corresponding microsatellite loci in normal cells of the same individual. The term "microsatellite stability" or "MSS" refers to a cancer (e.g., tumor) state without a significant MMR defect, such that there is no significant difference between the lengths of microsatellite loci in cancer cells and those of corresponding microsatellite loci in normal (e.g., non-cancer) cells of the same individual. The term "microsatellite indeterminate" or "MSE" refers to a cancer (e.g., tumor) state with an intermediate microsatellite length phenotype that cannot be clearly classified as MSI-H or MSS based on the statistical cutoffs used to define these two categories.

[0061] As used herein, the term "gene product" refers to a specific genomic locus, e.g., an RNA (e.g., mRNA or miRNA) or protein molecule transcribed or translated from a specific gene. A genomic locus can be identified using the gene name, chromosomal location, or any other genetic mapping metric.

[0062] As used herein, the term "ratio" refers to any comparison of a first metric X, or a first mathematical transform thereof X' (e.g., a measurement of the number of units of a genomic sequence in a first one or more biological samples, or a first mathematical transform thereof), to another metric Y, or a second mathematical transform thereof Y' (e.g., the number of units of each genomic sequence in a second one or more biological samples, or a second mathematical transform thereof), such as X / Y, Y / X, log N (X / Y), log N (Y / X), X' / Y, Y / X', log N (X' / Y) or log N (Y / X'), X / Y', Y' / X, log N (X / Y'), log N (Y' / X), X' / Y', Y' / X', log N (X' / Y') or log N (Y' / X'), where N is any real number greater than 1, and exemplary mathematical transformations of X and Y include, but are not limited to, raising X or Y to the Zth power, multiplying X or Y by a constant Q, where Z and Q are any real numbers, and / or taking the M-based logarithm of X and / or Y, where M is a real number greater than 1. In one non-limiting example, X is multiplied by squaring X (X 2 ) before the ratio calculation, and Y is calculated by raising Y to the power of 3.2 (Y 3.2 ) and the ratio of X and Y is calculated as log2(X' / Y').

[0063] As used herein, the terms "expression level," "abundance level," or simply "abundance" refer to the amount of a gene product (an RNA species, e.g., mRNA or miRNA, or a protein molecule) transcribed or translated by a cell, or the average amount of a gene product transcribed or translated across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g., a particular gene. However, in some embodiments, the expression level may refer to the amount of a particular isoform of an mRNA or protein corresponding to a particular gene that gives rise to multiple mRNA or protein isoforms. A genomic locus can be identified using gene name, chromosomal location, or any other genetic mapping metric.

[0064] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound, e.g., a gene product (e.g., an RNA species, e.g., mRNA or miRNA, or a protein molecule) or a nucleic acid fragment with a particular property (e.g., aligning to a particular locus or encompassing a particular allele), measured in a sample to a second amount of the compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a species of a compound to the total amount of the compound in the same sample. For example, it is the ratio of the amount of an mRNA transcript encoding a particular gene in the sample (e.g., aligning to a particular region of the exome) to the total amount of mRNA transcripts in the sample. In other embodiments, relative abundance refers to the ratio of the amount of a compound or species of a compound in a first sample to the amount of the compound of that species of compound in a second sample. For example, it is the ratio of the normalized amount of an mRNA transcript encoding a particular gene in a first sample to the normalized amount of an mRNA transcript encoding a particular gene in a second sample and / or a reference sample.

[0065] As used herein, the terms "sequencing," "determining a sequence," and the like refer to any biochemical process that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.

[0066] As used herein, the term "gene sequence" refers to a record of the series of nucleotides present in a subject's RNA or DNA as determined by sequencing nucleic acid from the subject.

[0067] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any nucleic acid sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (a "single-end read") or from both ends of a nucleic acid fragment (e.g., a paired-end read, a double-end read). The length of a sequence read is often related to a particular sequencing technology. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads are about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or greater in length. Nanopore® sequencing, for example, can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina® parallel sequencing, for example, can provide sequence reads that are less variable, e.g., the majority of sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a series of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, for example, hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0068] As used herein, the term "read segment" refers to a nucleotide sequence in any form, including raw sequence reads obtained directly from nucleic acid sequencing technology or sequences derived therefrom, e.g., aligned sequence reads, folded sequence reads, or spliced ​​sequence reads.

[0069] As used herein, the term "number of reads" refers to the total number of nucleic acid reads generated, which may or may not equal the number of nucleic acid molecules generated during a nucleic acid sequencing reaction.

[0070] As used herein, the terms "read depth," "sequencing depth," or "depth" may refer to the total number of unique nucleic acid fragments encompassing a particular locus or region of a subject's genome that are sequenced in a particular sequencing reaction. Sequencing depth can be expressed as "Y-fold," e.g., 50-fold, 100-fold, etc., where "Y" refers to the number of unique nucleic acid fragments encompassing a particular locus that are sequenced in a sequencing reaction. In such cases, Y is necessarily an integer, since it represents the actual sequencing depth of a particular locus. Alternatively, read depth, sequencing depth, or depth may refer to a measure of central tendency (e.g., the mean or mode) of the number of unique nucleic acid fragments encompassing one of multiple loci or regions of a subject's genome that are sequenced in a particular sequencing reaction. For example, in some embodiments, sequencing depth refers to the average depth of all loci across a chromosome, targeted sequencing panel, exome, or entire genome. In such cases, Y may be expressed as a fraction or decimal, since it refers to the average coverage across multiple loci. When an average depth is listed, the actual depth of any particular locus may differ from the overall listed depth. A metric can be determined that provides a range of sequencing depth within which a defined percentage of the total number of loci is included. For example, a range of sequencing depth within which 90%, 95%, or 99% of the loci are included. As will be understood by those skilled in the art, different sequencing technologies provide different sequencing depths. For example, low-pass whole genome sequencing can refer to a technology that provides a sequencing depth of less than 5x, less than 4x, less than 3x, or less than 2x, for example, about 0.5x to about 3x.

[0071] As used herein, the term "sequencing breadth" refers to what fraction of a particular reference exome (e.g., a human reference exome), a particular reference genome (e.g., a human reference genome), or a portion of an exome or genome has been analyzed. Sequencing breadth can be expressed as a fraction, decimal, or percentage and is generally calculated as (number of loci analyzed / total number of loci in the reference exome or reference genome). The denominator of the fraction can be the repeat-masked genome, and thus 100% can correspond to the entire reference genome minus the masked portion. A repeat-masked exome or genome can refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads align with unmasked portions of the exome or genome). In some embodiments, any portion of the exome or genome can be masked, and thus sequencing breadth can be assessed for any desired portion of the reference exome or genome. In some embodiments, "extensive sequencing" refers to sequencing / analysis of at least 0.1% of the exome or genome.

[0072] As used herein, the term "sequencing probe" refers to a molecule that binds to nucleic acids with an affinity based on the expected nucleotide sequence of the RNA or DNA present at that locus.

[0073] As used herein, the term "targeted panel" or "targeted gene panel" refers to a combination of probes for sequencing (e.g., by next-generation sequencing) nucleic acids present in a biological sample from a subject (e.g., a tumor sample, a liquid biopsy sample, a germline tissue sample, a white blood cell sample, or a tumor or tissue organoid sample) selected to map to one or more loci of interest on one or more chromosomes. Exemplary gene sets that can be analyzed using a targeted panel and are useful for precision oncology, e.g., via solid or liquid biopsy assays, are described in Lists 1-6. In some embodiments, in addition to loci beneficial for precision oncology, the targeted panel includes one or more probes for sequencing one or more of the following loci associated with different disease states, loci used for internal control purposes, or loci from pathogenic organisms (e.g., carcinogenic pathogens).

[0074] As used herein, the term "reference exome" refers to any sequenced or otherwise characterized exome of any tissue from any organism or pathogen, whether partial or complete, that can be used to reference an identified sequence from a subject. Typically, the reference exome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference exomes for human subjects, as well as many other organisms, are provided in the online genome browser hosted by the National Center for Biotechnology Information ("NCBI"). "Exome" refers to the complete transcriptional profile of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference exome is often an assembled or partially assembled exome sequence from an individual or multiple individuals. In some embodiments, a reference exome is an assembled or partially assembled exome sequence from one or more human individuals. A reference exome can be considered a representative example of a species' set of expressed genes. In some embodiments, a reference exome includes sequences assigned to chromosomes.

[0075] As used herein, the term "reference genome" refers to any sequenced or otherwise characterized genome of any organism or pathogen, whether partial or complete, that can be used to reference identified sequences from a subject. Typically, a reference genome is derived from a subject of the same species as the subject whose sequence is being evaluated. Exemplary reference genomes used for human subjects, as well as many other organisms, are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz ("UCSC"). "Genome" refers to the complete genetic information of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of a species' set of genes. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38). For haploid genomes, only one nucleotide may be present at each locus. For diploid genomes, heterozygous loci may be identified, and each heterozygous locus may have two alleles, with either allele being able to match for alignment to the locus.

[0076] As used herein, the term "bioinformatics pipeline" refers to a series of processing steps used to determine characteristics of a subject's genome or exome based on sequencing data of the subject's genome or exome. A bioinformatics pipeline can be used to determine characteristics of a subject's germline genome or exome and / or a subject's cancer genome or exome. In some embodiments, the pipeline extracts information related to genomic alterations in a subject's cancer genome, which is useful for guiding clinical decisions for precision oncology from sequencing results of a subject-derived biological sample, such as a tumor sample, a liquid biopsy sample, or a reference normal sample. Certain processing steps in bioinformatics can be "connected," meaning that the results of a first respective processing step are informative and / or essential for the execution of a second downstream processing step. For example, in some embodiments, a bioinformatics pipeline includes a first respective processing step for identifying genomic alterations unique to a subject's cancer genome, and a second respective processing step for using the amount and / or identity of the identified genomic alterations to determine a metric, such as tumor mutation burden, that is informative for precision oncology. In some embodiments, the bioinformatics pipeline includes a reporting stage that generates a report of relevant and / or actionable information identified by upstream stages of the pipeline, which may or may not further include recommendations to support clinical therapy decisions.

[0077] As used herein, the term "limit of detection" or "LOD" refers to the minimum amount of a feature that can be identified with a certain level of confidence. Thus, the level of detection can be used to describe the amount of a substance that must be present for a particular assay to reliably detect the substance. The level of detection can also be used to describe the level of support required for an algorithm to reliably identify genomic alterations based on sequencing data. For example, the minimum number of unique sequence reads required to support the identification of sequence variants such as SNVs.

[0078] As used herein, the terms "BAM file" or "binary file containing an alignment map" refer to a file that stores sequencing data aligned to a reference sequence (e.g., a reference genome or exome). In some embodiments, a BAM file is a compressed binary version of a SAM (sequence alignment map) file that contains, for each of a plurality of unique sequence reads, an identifier for the sequence read, information about the nucleotide sequence, information about the alignment of the sequence to the reference sequence, and optionally metrics about the quality of the sequence read and / or the quality of the sequence alignment. While a BAM file generally relates to a file having a particular format, for brevity, unless otherwise specified, it is used herein to simply refer to a file of any format that contains information about sequence alignments.

[0079] As used herein, the term "measure of central tendency" refers to the central or representative value of a distribution of values. Non-limiting examples of measures of central tendency include the arithmetic mean, weighted mean, median, central area, central hinge, trimean, geometric mean, geometric median, Winsorized mean, median, and mode of a distribution of values.

[0080] As used herein, the term "positive predictive value" or "PPV" refers to the likelihood that a variant will be correctly called, given that the variant was called by the assay. PPV can be expressed as (number of true positives) / (number of false positives + number of true positives).

[0081] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, e.g., a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining copy number variations of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the nucleic acid characteristics described herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of a mutation in a nucleic acid at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). Assays or methods can have a particular sensitivity and / or specificity, and their relative utility as diagnostic tools can be measured using ROC-AUC statistics.

[0082] As used herein, the term "classification" may refer to any number or other letter associated with a particular characteristic of a sample. For example, in some embodiments, the term "classification" may refer to the type of cancer in a subject, the stage of cancer in a subject, the prognosis of cancer in a subject, tumor burden, the presence of tumor metastasis in a subject, and the like. Classifications may be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" may refer to a predetermined number used in a calculation. For example, a cutoff size may refer to a size above which fragments are excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms may be used in either of these contexts.

[0083] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.

[0084] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity can characterize the ability of a method to correctly identify one or more markers that are indicative of cancer.

[0085] As used herein, "actionable genomic alteration" or "actionable variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed to be associated with a therapeutic course of action that is more likely to result in a positive outcome in cancer patients with an actionable variant than in similarly located cancer patients without an actionable variant. For example, administration of an EGFR inhibitor (e.g., afatinib, erlotinib, gefitinib) is more effective in treating non-small cell lung cancer in patients with an EGFR mutation in exons 19 / 21 than in patients without an EGFR mutation in exons 19 / 21. Thus, an EGFR mutation in exons 19 / 21 is an actionable variant. In some cases, the actionable variant is associated with an improved therapeutic outcome in only one specific cancer type or a group of specific cancer types, while in other cases, the actionable variant is associated with an improved therapeutic outcome in substantially all cancer types.

[0086] As used herein, "variant of uncertain significance" or "VUS" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) whose impact on disease development / progression is unknown.

[0087] As used herein, "benign variant" or "likely benign variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed not to contribute to disease development / progression.

[0088] As used herein, "pathogenic variant" or "likely pathogenic variant" refers to a genomic alteration (e.g., SNV, MNV, indel, rearrangement, copy number variation, or ploidy variation) or the value of another cancer metric derived from nucleic acid sequencing data (e.g., tumor mutational burden, MSI status, or tumor fraction) that is known or believed to contribute to disease development / progression.

[0089] As used herein, an "effective amount" or "therapeutically effective amount" is an amount sufficient to affect beneficial or desired clinical results during treatment. An effective amount can be administered to a subject in one or more doses. In terms of treatment, an effective amount is an amount sufficient to palliate, improve, stabilize, reverse, or slow the progression of a disease, or otherwise reduce the pathological consequences of a disease. An effective amount is generally determined by a physician on a case-by-case basis and is within the skill of one of ordinary skill in the art. Several factors are typically considered when determining an appropriate dosage to achieve an effective amount. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.

[0090] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. As used herein, the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or the claims, such terms are intended to be as inclusive as the term "comprising."

[0091] As used herein, the term "if" may be interpreted to mean "when" or "upon," "in response to detecting," or "in response to determining," depending on the context. Similarly, the phrase "when determined" or "when [a described condition or event] is detected" may be interpreted to mean "upon determining," or "in response to determining," or "upon detection [of [a described condition or event]," or "in response to detecting [a described condition or event]," depending on the context.

[0092] Additionally, while terms such as "first," "second," and the like may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both subjects, but are not the same subject. Furthermore, the terms "subject," "user," and "patient" are used interchangeably herein.

[0093] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure, including example systems, methods, techniques, instruction sequences, and computing machine program products embodying illustrative embodiments. However, the following illustrative discussion is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. Features described herein are not limited by the illustrated ordering of acts or events, as some acts may occur in different orders and / or contemporaneously with other acts or events.

[0094] The embodiments provided herein are chosen and described to best explain the principles and their practical applications, thereby enabling those skilled in the art to best utilize the various embodiments with various modifications suited to the particular uses contemplated. In some instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. In other instances, it will be apparent to one skilled in the art that the present disclosure may be practiced without one or more of the specific details.

[0095] It will be understood that in the development of any such actual implementation, numerous implementation-specific decisions will be made to achieve the designer's specific goals, such as adherence to use case and business-related constraints, and that these specific goals will vary from implementation to implementation and from designer to designer. It will further be understood that such a design effort may be complex and time-consuming, but would nevertheless be a routine undertaking for one of ordinary skill in the art having the benefit of this disclosure.

[0096] Exemplary System Embodiments Having provided an overview of some aspects of the present disclosure and some definitions used herein, details of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay will now be described in conjunction with Figures 1A-B. Figures 1A-B collectively illustrate the topology of an exemplary system for providing clinical support for personalized cancer therapy using a liquid biopsy assay, according to some embodiments of the present disclosure. Advantageously, the exemplary system illustrated in Figures 1A-B improves upon conventional methods for providing clinical support for personalized cancer therapy by validating somatic sequence variants in a test subject with a cancerous condition.

[0097] 1A is a block diagram illustrating a system according to some embodiments. Device 100 in some embodiments includes one or more processing units (CPUs) 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including, for example, a display 108 and / or input 110 (e.g., a mouse, touchpad, keyboard, etc.), non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. One or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communications between system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes a CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU 102. Persistent memory 112 and the non-volatile memory devices within non-persistent memory 112 comprise non-transitory computer-readable storage media. In some implementations, non-persistent memory 111, or alternatively, non-transitory computer-readable storage media, sometimes in conjunction with persistent memory 112, store the following programs, modules, and data structures, or a subset thereof: an operating system 116, including procedures for handling various basic system services and for performing hardware-dependent tasks; a network communication module (or instructions) 118 for connecting the system 100 with other devices and / or communication networks 105; a test patient data store 120 for storing one or more sets of features from a patient (e.g., subject); a bioinformatics module 140 for processing the sequencing data and extracting features from the sequencing data, e.g., from a liquid biopsy sequencing assay; a feature analysis module 160 for assessing patient features, such as genomic alterations, complex genomic features, and clinical features; and A reporting module 180 for generating and transmitting reports that provide clinical support for personalized cancer therapy.

[0098] While FIGS. 1A-B depict "system 100," the diagrams are intended as a functional illustration of various features that may be present in a computer system, rather than as a structural schematic of the embodiments described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art. Furthermore, while FIG. 1 depicts certain data and modules in non-persistent memory 111, some or all of these data and modules may be in persistent memory 112. For example, in various embodiments, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for implementing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various embodiments.

[0099] In some implementations, non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of system 100 that is addressable by system 100 such that system 100 can retrieve all or a portion of such data when needed.

[0100] 1A, system 100 is represented as a single computer containing all of the functionality for providing clinical support for personalized cancer therapy. However, while a single machine is illustrated, the term "system" should also be construed to include any collection of machines that individually or jointly execute a set (or sets) of instructions for implementing any one or more of the methodologies discussed herein.

[0101] For example, in some embodiments, system 100 includes one or more computers. In some embodiments, functionality for providing clinical support for personalized cancer therapy is spread across any number of networked computers and / or resident on each of several networked computers and / or hosted on one or more virtual machines at remote locations accessible over communications network 105. For example, different portions of the various modules and data stores illustrated in Figures 1A-B may be stored and / or executed on various instances of processing devices and / or processing servers / databases (e.g., processing devices 224, 234, 244, and 254, processing server 262, and database 264) within distributed diagnostic environment 210 illustrated in Figure 2B.

[0102] The system may operate in the capacity of a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment. The system may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, server, network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine.

[0103] In another embodiment, the system comprises a virtual machine including modules for executing instructions for performing any one or more of the methodologies disclosed herein. In computing, a virtual machine (VM) is an emulation of a computer system that is based on a computer architecture and provides the functionality of a physical computer. Some such embodiments may involve dedicated hardware, software, or a combination of hardware and software.

[0104] Those skilled in the art will appreciate that any of a wide variety of different computer topologies may be used in an application, and that all such topologies are within the scope of the present disclosure.

[0105] Examination patient data storage unit (120) Referring to FIG. 1B , in some embodiments, the system (e.g., system 100) includes a patient data store 120 that stores data for patients 121-1 through 121-M (e.g., cancer patients or patients being tested for cancer), including one or more of sequencing data 122, feature data 125, and clinical assessments 139. These data are used and / or generated by various processes stored in bioinformatics module 140 and feature analysis module 160 of system 100 to generate reports that ultimately provide clinical support for personalized cancer therapy for the patients. While the feature range of patient data 121 across all patients may be informationally dense, an individual patient's feature set may be sparse across the collective feature range of all features across all patients. That is, data stored for one patient may include a different set of features than data stored for another patient. Furthermore, although illustrated as a single data structure in FIG. 1B , different sets of patient data may be stored in different databases or modules spread across one or more system memories.

[0106] In some embodiments, sequencing data 122 from one or more sequencing reactions 122-i, including multiple sequence reads 123-i-1 through 123-iK, is stored in the test patient data store 120. The data store may contain different sets of sequencing data from a single subject, corresponding to different samples from the patient, e.g., tumor samples, liquid biopsy samples, tumor organoids derived from the patient's tumor, and / or normal samples, and / or samples acquired at different times, e.g., while monitoring the progression, regression, remission, and / or recurrence of cancer in the subject. The sequence reads may be in any suitable file format, e.g., BCL, FASTA, FASTQ, etc. In some embodiments, the sequencing data 122 is accessed by a sequencing data processing module 141, which performs various preprocessing, genome alignment, and demultiplexing operations, as described in detail below with reference to the bioinformatics module 140. In some embodiments, sequence data aligned to a reference construct, e.g., a BAM file 124, is stored in the test patient data store 120.

[0107] In some embodiments, the test patient data store 120 includes feature data 125 that is useful, for example, for identifying clinical support for personalized cancer therapy. In some embodiments, the feature data 125 includes patient personal characteristics 126, such as patient name, date of birth, gender, ethnicity, physical address, smoking status, alcohol consumption characteristics, anthropomorphic data, etc.

[0108] In some embodiments, the feature data 125 includes patient medical history data 127, such as cancer diagnosis information (e.g., date of initial diagnosis, date of metastasis diagnosis, cancer staging, tumor characterization, tissue of origin, previous treatments and outcomes, adverse effects of treatments, history of treatment groups, history of clinical trials, previous and current medications, history of surgery, etc.), previous or current symptoms, previous or current treatments, previous treatment outcomes, previous disease diagnoses, diabetic status, depression diagnosis, other physical or mental illness diagnoses, and family history. In some embodiments, the feature data 125 includes clinical features 128, such as pathology data 128-1, medical imaging data 128-2, and tissue culture and / or tissue organoid culture data 128-3.

[0109] In some embodiments, additional clinical features, such as previous laboratory test results, are stored in the laboratory patient data store 120. Historical data 127 and clinical features may be collected from the patient from a variety of sources, including directly from the patient's electronic medical record (EMR) or electronic health record (EHR), or curated from other sources, such as fields from various laboratory records (e.g., gene sequencing reports).

[0110] In some embodiments, the feature data 125 includes patient genomic features 131. Non-limiting examples of genomic features include allele state 132 (e.g., allele identity at one or more loci, support for wild type or variant alleles at one or more loci), support for SNV / MNV at one or more loci, support for indels at one or more loci, and / or support for genetic rearrangements at one or more loci), allele fraction 133 (e.g., the ratio of variant to reference alleles (or vice versa)), methylation state 132 (e.g., the distribution of methylation patterns at one or more loci and / or support for abnormal methylation patterns at one or more loci), genome copy number 135 (e.g., copy number value at one or more loci and / or support for abnormal (increased or decreased) copy number at one or more loci), tumor mutation burden 136 (e.g., the ratio of variant to reference alleles (or vice versa)), and the like. genomic features 131 include a measure of the number of mutations in the subject's cancer genome, and microsatellite instability status 137 (e.g., a measure of repeat unit length at one or more microsatellite loci and / or a classification of MSI status for the patient's cancer). In some embodiments, one or more of the genomic features 131 are determined by a nucleic acid bioinformatics pipeline, e.g., as described in detail below with reference to Figures 4A-4E. In particular, in some embodiments, the feature data 125 includes a variant allele fraction 133 determined using an improved method for validating somatic sequence variants. In some embodiments, one or more of the genomic features 131 are obtained from an external testing source that is not connected to the bioinformatics pipeline, e.g., as described below.

[0111] In some embodiments, feature data 125 further includes data 138 from other "-omics" research fields. Non-limiting examples of "-omics" research fields that may yield feature data useful for providing clinical support for personalized cancer therapy include transcriptomics, epigenomics, proteomics, metabolomics, metabonomics, microbiomics, lipidomics, glycomics, cellomics, and organomics.

[0112] In some embodiments, still other features may include, but are not limited to, those listed above, including features derived from machine learning approaches based at least in part on the evaluation of any relevant molecular or clinical features, considered alone or in combination. For example, in some embodiments, one or more latent features learned from the evaluation of a cancer patient training dataset improve the diagnostic and prognostic capabilities of various analysis algorithms in feature analysis module 160.

[0113] Those skilled in the art will know of other types of features that are useful for providing clinical support for personalized cancer therapy. The above list of features is merely representative and should not be construed as limiting.

[0114] In some embodiments, the test patient data store 120 includes clinical assessment data 139 for the patient, e.g., based on the feature data 125 collected for the subject. In some embodiments, the clinical assessment data 139 includes a catalog of actionable variants and features 139-1 (e.g., genomic alterations and composite metrics based on genomic features known or thought to be targetable by one or more specific cancer therapies), matched therapies 139-2 (e.g., therapies known or thought to be particularly beneficial for treating subjects with actionable variants), and / or clinical reports 139-3 generated for the subject, e.g., based on the identified actionable variants and features 139-1 and / or matched therapies 139-2.

[0115] In some embodiments, clinical assessment data 139 is generated by analysis of feature data 125 using various algorithms in feature analysis module 160, as described in further detail below. In some embodiments, clinical assessment data 139 is generated, modified, and / or verified by evaluation of feature data 125 by a clinician, e.g., an oncologist. For example, in some embodiments, a clinician (e.g., in clinical environment 220) uses feature analysis module 160 or directly accesses laboratory patient data store 120 to evaluate feature data 125 and recommend a personalized cancer treatment for the patient. Similarly, in some embodiments, a clinician (e.g., in clinical environment 220) reviews recommendations determined using feature analysis module 160 and, for example, approves, rejects, or modifies the recommendations before they are sent to a healthcare professional treating the cancer patient.

[0116] Bioinformatics Modules(140) Referring again to FIG. 1A, the system (e.g., system 100) includes a bioinformatics module 140 that includes a feature extraction module 145 and optional auxiliary data processing constructs, such as a sequence data processing module 141 and / or one or more reference sequence constructs 158 (e.g., a reference genome, exome, or target panel construct that includes reference sequences for multiple loci targeted by the sequencing panel).

[0117] In some embodiments, bioinformatics module 140 includes a sequence data processing module 141 that includes instructions for processing sequence reads, e.g., raw sequence reads 123 from one or more sequencing reactions 122, prior to analysis by various feature extraction algorithms, as described in detail below. In some embodiments, sequence data processing module 141 includes one or more preprocessing algorithms 142 that prepare the data for analysis. In some embodiments, preprocessing algorithm 142 includes instructions for converting the file format of sequence reads from the output of a sequencer (e.g., a BCL file format) into a file format compatible with downstream analysis of the sequences (e.g., a FASTQ or FASTA file format). In some embodiments, the preprocessing algorithm 142 includes instructions for assessing the quality of sequence reads (e.g., by examining quality metrics such as Phred score, base calling error probability, quality (Q) score, and the like) and / or removing sequence reads that do not meet a threshold quality (e.g., an inferred base calling accuracy of at least 80%, at least 90%, at least 95%, at least 99%, at least 99.5%, at least 99.9%, or more). In some embodiments, the preprocessing algorithm 142 includes instructions for filtering sequence reads for one or more properties, e.g., removing sequences that fail to meet a lower or upper size threshold or removing duplicate sequence reads.

[0118] In some embodiments, the sequence data processing module 141 includes one or more alignment algorithms 143 for aligning the preprocessed sequence reads 123 to a reference sequence construct 158, such as a reference genome, exome, or target panel construct. Many algorithms for aligning sequencing data to a reference construct are known in the art, such as BWA, Blat, SHRiMP, LastZ, and MAQ. One example of a sequence read alignment is the Burrows-Wheeler Alignment Tool (BWA), which uses the Burrows-Wheeler Transform (BWT) to align short sequence reads to a large reference construct and allows for mismatches and gaps. Li and Durbin, Bioinformatics, 25(14):1754-60 (2009), the contents of which are incorporated herein by reference in their entirety for all purposes. The sequence read alignment package imports raw or pre-processed sequence reads 122, for example, in BCL, FASTA, or FASTQ file format, and outputs aligned sequence reads 124, for example, in SAM or BAM file format.

[0119] In some embodiments, the sequence data processing module 141 includes one or more demultiplexing algorithms 144 to split sequence read or sequence alignment files generated from pooled nucleic acid sequencing reactions into separate sequence read or sequence alignment files, each corresponding to a different source of nucleic acid in the nucleic acid sequencing pool. For example, due to the cost of sequencing reactions, it is common practice to pool nucleic acids from multiple samples into a single sequencing reaction. Nucleic acids from each sample are tagged with sample-specific and / or molecule-specific sequence tags (e.g., UMIs) that are sequenced along with the molecule. In some embodiments, the demultiplexing algorithm 144 sorts these sequence tags within the sequence read or sequence alignment files and demultiplexes the sequencing data into separate files for each of the samples included in the sequencing reaction.

[0120] The bioinformatics module 140 includes a feature extraction module 145 that includes instructions for identifying diagnostic features, e.g., genomic features 131, from sequencing data 122 of one or more biological samples from a subject, e.g., a solid tumor sample, a liquid biopsy sample, or a normal tissue (e.g., control) sample. For example, in some embodiments, the feature extraction algorithm compares the identity of one or more nucleotides at a locus from the sequencing data 122 to the identity of the nucleotide at that locus in a reference sequence construct (e.g., a reference genome, exome, or targeted panel construct) to determine whether the subject has a variant at that locus. In some embodiments, the feature extraction algorithm evaluates data other than the raw sequence to identify genomic alterations in the subject, e.g., allele ratios, relative copy numbers, repeat unit distribution, etc.

[0121] For example, in some embodiments, feature extraction module 145 includes one or more variant identification modules that include instructions for various variant calling processes. In some embodiments, variants in a subject's germline are identified, for example, using germline variant identification module 146. In some embodiments, variants in a cancer genome, e.g., somatic variants, are identified, for example, using somatic variant identification module 150. Although separate germline and somatic variant identification modules are illustrated in FIG. 1A, in some embodiments, they are integrated into a single module. In some embodiments, the variant identification module includes instructions for identifying nucleotide variants (e.g., single nucleotide variants (SNVs) and multiple nucleotide variants (MNVs)) using one or more SNV / MNV calling algorithms (e.g., algorithms 147 and / or 151), indels (e.g., nucleotide insertions or deletions) using one or more indel calling algorithms (e.g., algorithms 148 and / or 152), and genomic rearrangements (e.g., nucleotide sequence inversions, translocations, and fusions) using one or more genomic rearrangement calling algorithms (e.g., algorithms 149 and / or 153).

[0122] For example, in some embodiments, feature extraction module 145 includes variant identification module 146, which includes a variant threshold determination module 146-a, a sequence variant data store 146-r, and a variant validation module 146-o. In some such embodiments, sequence variant data store 146-r includes one or more candidate variants for a test subject identified by aligning a plurality of sequence reads obtained from sequencing a liquid biopsy sample of the test subject to a reference sequence, where the one or more candidate variants correspond to one or more loci in the reference sequence. The plurality of sequence reads aligned to the reference sequence are used to identify a variant allele fragment count for each candidate variant. In some embodiments, sequence variant data store 146-r further includes a plurality of variants from a first set of nucleic acids obtained from a cohort of subjects (e.g., from a tumor tissue biopsy of each subject in a baseline cohort of subjects). The variant threshold determination module 146-a performs a function for each candidate variant in the one or more candidate variants, and for each corresponding locus 146-b (e.g., 146-b-1, ..., 146-bP), a dynamic variant count threshold 146-d (e.g., 146-d-1) is obtained based on the pre-test odds of a positive variant call for the locus based on the prevalence of the variant in the genomic region containing the locus using multiple variants in the baseline cohort. The variant threshold module 146-a compares the variant allele fragment count 146-c (e.g., 146-c-1) of the candidate variant against the dynamic variant count threshold 146-d of the locus corresponding to the candidate variant. In some embodiments, the variant validation module 146-o determines whether the candidate variant is validated or rejected as a somatic sequence variant based on the comparison. For example, a somatic sequence variant is validated when the variant allele fragment count of the candidate variant meets the dynamic variant count threshold for the locus, and is rejected when the variant allele fragment count of the candidate variant does not meet the dynamic variant count threshold for the locus.

[0123] In some embodiments, the dynamic variant count threshold is determined based on the distribution of variant detection sensitivity as a function of circulating variant allele fractions from a cohort of subjects (e.g., a baseline cohort). For example, in some such embodiments, the variant threshold determination module 146-a takes as input one or more variant allele fractions 133 from the genomic feature module 131. In some such embodiments, the variant allele fraction 133 comprises a plurality of variant allele fractions obtained from tumor tissue biopsies 133-t (e.g., 133-t-1, 133-t-2..., 133-tO) of the cohort of subjects. In some embodiments, the variant allele fraction comprises a plurality of variant allele fractions obtained from liquid biopsy samples 133-cf (e.g., 133-cf-1, 133-cf-2..., 133-cf-N) of the cohort of subjects. In some embodiments, the circulating variant allele fraction is obtained by comparing the liquid biopsy variant allele fraction 133-cf to the tumor biopsy variant allele fraction 133-t.

[0124] Additional embodiments for using variant allele fractions (e.g., variant allele frequencies) to identify somatic variants are detailed below (see Exemplary Methods: Variant Identification).

[0125] The SNV / MNV algorithm 147 can identify single-nucleotide substitutions occurring at specific positions in the genome. For example, at a specific base position or locus in the human genome, a C nucleotide may occur in most individuals, while in a minority of individuals, the position is occupied by an A. This means that a SNP exists at this specific position, and the two possible nucleotide variations, C or A, are said to be alleles for this position. SNPs underlie differences in human susceptibility to a wide range of diseases (e.g., sickle cell anemia, beta-thalassemia, and cystic fibrosis resulting from SNPs). Disease severity and how the body responds to treatment are also manifestations of genetic variation. For example, a single-nucleotide mutation in the APOE (apolipoprotein E) gene is associated with a lower risk of Alzheimer's disease. A single-nucleotide variant (SNV) is a variation in a single nucleotide without any frequency restriction and can occur somatically. Somatic single-nucleotide mutations (e.g., caused by cancer) may also be referred to as single-nucleotide changes. MNP (multiple nucleotide polymorphism) modules can identify substitutions of consecutive nucleotides at specific positions in the genome.

[0126] The indel calling algorithm 148 can identify insertions or deletions of bases in the genome of organisms classified among minor genetic variations. Indels typically measure 1 to 10,000 base pairs in length, while microindels are defined as indels resulting in a net change of 1 to 50 nucleotides. Indels can be contrasted with SNPs or point mutations. While indels insert and / or delete nucleotides from a sequence, point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels, which are insertions and / or deletions, can be used as genetic markers in natural populations, particularly in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except near highly repetitive regions, including homopolymers and microsatellites.

[0127] Genome rearrangement algorithms 149 can identify hybrid genes formed from two previously separated genes. This can occur as a result of translocations, interstitial deletions, or chromosomal inversions. Gene fusions can play an important role in tumorigenesis. Fusion genes can contribute to tumorigenesis because they can produce abnormal proteins that are much more active than non-fusion genes. Often, fusion genes are cancer-causing oncogenes, including BCR-ABL, TEL-AML1 (ALL with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG, an interstitial deletion on chromosome 21 that frequently occurs in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by disrupting androgen receptor (AR) signaling and inhibiting AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematologic cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a fusion gene specific and characteristic of high-grade serous ovarian cancer. Oncogenic fusion genes can lead to gene products with new or distinct functions from the two fusion partners. Alternatively, proto-oncogenes can be fused to strong promoters, thereby setting up oncogenic functions through upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where oncogenes are juxtaposed to immunoglobulin gene promoters. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events. Because chromosomal translocations play such an important role in neoplasia, a dedicated database of chromosomal aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer.

[0128] In some embodiments, feature extraction module 145 includes instructions for identifying one or more complex genomic alterations (e.g., features that incorporate more than changes in the primary sequence of the genome) in a subject's cancer genome. For example, in some embodiments, feature extraction module 145 includes copy number variation (e.g., copy number variation analysis module 153), microsatellite instability status (e.g., microsatellite instability analysis module 154), tumor mutation burden (e.g., tumor mutation burden analysis module 155), tumor ploidy (e.g., tumor ploidy analysis module 156), and homologous recombination pathway deficiency (e.g., homologous recombination pathway analysis module 157).

[0129] Feature Analysis Module(160) 1A , the present system (e.g., system 100) includes a feature analysis module 160, which includes one or more genomic variation interpretation algorithms 161, one or more optional clinical data analysis algorithms 165, an optional treatment curation algorithm 165, and an optional recommendation validation module 167. In some embodiments, feature analysis module 160 uses one or more analysis algorithms (e.g., algorithms 162, 163, 164, and 165) to evaluate feature data 125 to identify actionable variants and characteristics 139-1 and corresponding matched treatments 139-2 and / or clinical trials. The identified actionable variants and characteristics 139-1 and corresponding matched treatments 139-2, which are optionally stored in test patient data store 120, are then curated by feature analysis module 160 to generate a clinical report 139-3, which is optionally verified by a user, e.g., a clinician, before being sent to a medical professional, e.g., an oncologist, treating the patient.

[0130] In some embodiments, the genomic variation interpretation algorithm 161 includes instructions for, for example, evaluating the impact of one or more genomic features 131 of a subject identified by the feature extraction module 145 on the characteristics of the patient's cancer and / or whether one or more targeted cancer therapies may improve the patient's clinical outcome. For example, in some embodiments, the one or more genomic variant analysis algorithms 163 evaluate various genomic features 131 by querying a database, e.g., a look-up table ("LUT"), of actionable genomic alterations, targeted therapies associated with the actionable genomic alterations, and any other conditions that should be met before administering the targeted therapy to a subject with an actionable genomic alteration. For example, evidence suggests that depatuxizumab mafodotin (an anti-EGFR mAb conjugated to monomethyl auristatin F) has improved efficacy for treating recurrent glioblastoma with EGFR focal amplification. van den Bent M. et al., Cancer Chemother Pharmacol., 80(6):1209-17 (2017). Thus, the LUT of actionable genomic alterations has an entry for a focal amplification of the EGFR gene, indicating that depatuxizumab mafodotin is a targeted therapy for glioblastoma with focal gene amplification (e.g., recurrent glioblastoma). In some cases, the LUT may also include contraindications to the associated targeted therapy, such as drug interactions or personal characteristics that contraindicate administration of a particular targeted therapy.

[0131] In some embodiments, the genomic alteration interpretation algorithm 161 determines whether a particular genomic feature 131 should be reported to a healthcare professional treating a cancer patient. In some embodiments, a genomic feature 131 (e.g., genomic alterations and composite features) is reported if there is clinical evidence that the feature significantly impacts cancer biology, affects cancer prognosis, and / or impacts pharmacogenomics, for example, by indicating or contraindicating a particular treatment approach. For example, the genomic variation interpretation algorithm 161 may classify a particular CNV feature 135 as "reportable," meaning, e.g., that the CNV has been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "non-reportable," meaning, e.g., that the CNV has not been identified as affecting the cancer's characteristics, overall disease state, and / or pharmacogenomics; as "no evidence," meaning, e.g., that there is no evidence to support the CNV being "reportable" or "non-reportable"; or as "conflicting evidence," meaning, e.g., that there is evidence to support both the CNV being "reportable" and the CNV being "non-reportable."

[0132] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more pathogenic variant analysis algorithms 162 that evaluate various genomic features to identify the presence of an oncogenic pathogen associated with the patient's cancer and / or targeted therapies associated with oncogenic pathogen infection in the cancer. For example, RNA expression patterns in some cancers are associated with the presence of oncogenic pathogens that are instrumental in inducing the cancer. See, e.g., U.S. Patent Application No. 16 / 802,126, filed February 26, 2020, the contents of which are incorporated by reference in their entirety for all purposes. In some cases, recommended treatments for cancer differ when the cancer is associated with oncogenic pathogen infection than when it is not. Thus, in some embodiments, for example, if feature data 125 includes RNA abundance data for the patient's cancer, one or more pathogenic variant analysis algorithms 162 evaluate the RNA abundance data for the patient's cancer to determine whether a signature indicative of the presence of an oncogenic pathogen in the cancer is present in the data. Similarly, in some embodiments, bioinformatics module 140 includes an algorithm that searches for the presence of pathogenic nucleic acid sequences within sequencing data 122. See, e.g., U.S. Provisional Patent Application No. 62 / 978,067, filed February 18, 2020, the contents of which are incorporated by reference in their entirety for all purposes. Accordingly, in some embodiments, one or more pathogenic variant analysis algorithms 162 evaluate whether the presence of an oncogenic pathogen in a subject is associated with an actionable treatment for the infection. In some embodiments, system 100 queries a database, e.g., a look-up table (LUT), of actionable oncogenic pathogen infections, targeted therapies associated with actionable infections, and any other conditions that should be met before administering the targeted therapy to a subject infected with an oncogenic pathogen. In some cases, the LUT may also include contraindications to the associated targeted therapy, e.g., drug interactions or personal characteristics that contraindicate the administration of a particular targeted therapy.

[0133] In some embodiments, genomic alteration interpretation algorithm 161 includes one or more multi-feature analysis algorithms 164 that evaluate multiple features to classify cancers with respect to the efficacy of one or more targeted therapies. For example, in some embodiments, feature analysis module 160 includes one or more classifiers trained on feature data, one or more clinical therapies, and their associated clinical outcomes for multiple training subjects to classify cancers based on predicted clinical outcomes following one or more therapies.

[0134] In some embodiments, the classifier is implemented as an artificial intelligence engine and may include a gradient boosting model, a random forest model, a neural network (NN), a regression model, a naive Bayes model, and / or a machine learning algorithm (MLA). The MLA or NN may be trained from a training dataset that includes one or more features 125, including personal characteristics 126, medical history 127, clinical features 128, genomic features 131, and / or other "mix" features 138. MLA includes supervised algorithms (i.e., algorithms where the features / classifications in the dataset are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, nearest neighbor clustering, unsupervised algorithms (i.e., algorithms where the features / classifications in the dataset are not annotated) using Apriori, average clustering, principal component analysis, random forests, adaptive boosting, and semi-supervised algorithms (i.e., algorithms where an incomplete number of features / classifications in the dataset are annotated) using generative approaches (e.g., mixtures of Gaussian distributions, mixtures of multinomial distributions, hidden Markov models), sparse separation, graph-based approaches (e.g., min-cut, harmonic functions, manifold normalization), heuristic approaches, or support vector machines.

[0135] NNs include conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long-short-term memory networks, or other neural models where the training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample.

[0136] MLAs and neural networks identify distinct approaches to machine learning, and these terms may be used interchangeably herein. Thus, unless explicitly stated otherwise, a reference to an MLA may include a corresponding NN, or vice versa. Training may include providing an optimized dataset, labeling these features as they occur in patient records, and training the MLA to predict or classify based on new inputs. Artificial NNs are efficient computing models that have demonstrated their strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators, i.e., they can represent a wide variety of functions given appropriate parameters.

[0137] In some embodiments, system 100 includes a classifier training module that includes instructions for training one or more untrained or partially trained classifiers based on feature data from a training dataset. In some embodiments, system 100 also includes a database of training data for use in training the one or more classifiers. In other embodiments, the classifier training module accesses a remote storage device that hosts the training data. In some embodiments, the training data includes a set of training features, including, but not limited to, the various types of feature data 125 illustrated in FIG. 1B. In some embodiments, the classifier training module uses patient data 121, for example, when test patient data store 120 also stores records of treatments administered to patients and patient outcomes after treatment.

[0138] In some embodiments, feature analysis module 160 includes one or more clinical data analysis algorithms 165 that evaluate the clinical features 128 of the cancer to identify targeted therapies that may benefit the subject. For example, in some embodiments, if feature data 125 includes pathology data 128-1, one or more clinical data analysis algorithms 165 evaluate the data to determine whether an actionable therapy is indicated, for example, based on the histopathology of a tumor biopsy from the subject, which may indicate a particular cancer type and / or stage of the cancer. In some embodiments, system 100 queries a database, e.g., a look-up table (“LUT”), of actionable clinical features (e.g., pathology features), targeted therapies associated with the actionable features, and any other conditions that must be met before administering the targeted therapy to a subject associated with the actionable clinical feature 128 (e.g., pathology feature 128-1). In some embodiments, system 100 directly evaluates clinical features 128 (e.g., pathology feature 128-1) to determine whether a patient's cancer is susceptible to a particular therapeutic agent. Further details about exemplary methods, systems, and algorithms for classifying cancer and identifying targeted therapies based on clinical data, such as pathology data 128-1, imaging data 138-2, and / or tissue culture / organoid data 128-3, are discussed, for example, in U.S. Patent Application No. 16 / 830,186, filed March 25, 2020, U.S. Patent Application No. 16 / 789,363, filed February 12, 2020, and U.S. Provisional Application No. 63 / 007,874, filed April 9, 2020, the contents of which are incorporated herein in their entirety for all purposes.

[0139] In some embodiments, feature analysis module 160 includes a clinical trial module that evaluates the test patient data 121 to determine whether the patient is eligible for inclusion in clinical trials for cancer therapies, e.g., clinical trials that are currently recruiting patients, clinical trials that have not yet begun recruiting patients, and / or ongoing clinical trials that may recruit additional patients in the future. In some embodiments, the clinical trial module evaluates the test patient data 121 to determine whether clinical trial results, e.g., results of ongoing clinical trials and / or results of completed clinical trials, are relevant to the patient. For example, in some embodiments, system 100 queries a database, e.g., a look-up table (“LUT”), of clinical trials, e.g., active and / or completed clinical trials, and compares the patient data 121 with the patient inclusion criteria of the clinical trials stored in the database to identify clinical trials that have patient inclusion criteria that closely and / or exactly match the patient's data 121. In some embodiments, records of matching clinical trials, e.g., clinical trials for which the patient may be eligible and / or that may inform individualized treatment decisions for the patient, are stored in clinical evaluation database 139.

[0140] In some embodiments, the feature analysis module 160 includes a therapy curation algorithm 166 that assembles the actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials identified for the patient, as described above. In some embodiments, the therapy curation algorithm 166 evaluates certain criteria related to which actionable variants and characteristics 139-1, matched therapies 139-2, and / or relevant clinical trials should be reported, and / or whether certain matched therapies, considered alone or in combination, may be contraindicated for the patient, for example, based on the patient's personal characteristics 126 and / or known drug-drug interactions. In some embodiments, the therapy curation algorithm then generates one or more clinical reports 139-3 for the patient. In some embodiments, the therapy curation algorithm generates a first clinical report 139-3-1 that is reported to a healthcare professional treating the patient, and a second clinical report 139-3-2 that is not communicated to a healthcare professional but can be used to improve various algorithms within the system.

[0141] In some embodiments, the feature analysis module 160 includes a recommendation validation module 167 that includes an interface that allows a clinician to review, modify, and approve the clinical report 139-3 before the report is sent to a medical professional, e.g., an oncologist, treating the patient.

[0142] In some embodiments, each of one or more of the feature collection, sequencing module, bioinformatics module (including, for example, a variation module, a structural variant calling and data processing module), classification module, and outcome module is communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In some alternative embodiments, each of the feature collection, variation module, structural variant calling and feature store is communicatively coupled to each other for independent communication without sharing a data bus.

[0143] Further details about the system and example embodiments of modules and feature sets are discussed in PCT application PCT / US19 / 69149, entitled "METHOD AND PROCESS FOR PREDICTING AND ANALYZING PATIENT COHORT RESPONSE, PROGRESSION, AND SURVIVAL," filed December 31, 2019, which is incorporated herein by reference in its entirety.

[0144] Exemplary Methods For example, having disclosed details of system 100 for providing clinical support for personalized cancer therapy with improved validation of somatic sequence variants, details regarding the processes and features of the system according to various embodiments of the present disclosure are disclosed below. Specifically, exemplary processes are described below with reference to Figures 2A, 3, 4A-E, and 5A-B. In some embodiments, such processes and features of the system are performed by modules 118, 120, 140, 160, and / or 170, as illustrated in Figure 1A. With reference to these methods, the systems described herein (e.g., system 100) include instructions for validating somatic variants, which are improved over conventional methods for somatic variant detection.

[0145] Figure 2B: Distributed diagnostic and clinical environment In some aspects, the methods described herein for providing clinical support for personalized cancer therapy are implemented across a distributed diagnostic / clinical environment, for example, as illustrated in Figure 2B. However, in some embodiments, the improved methods described herein for validating somatic sequence variants are implemented at a single location, e.g., in a single computing system or environment, but ancillary procedures that support the methods described herein and / or procedures that further utilize the results of the methods described herein may be implemented across the distributed diagnostic / clinical environment.

[0146] 2B illustrates an example of a distributed diagnostic / clinical environment 210. In some embodiments, the distributed diagnostic / clinical environment is connected via a communications network 105. In some embodiments, one or more biological samples, e.g., one or more liquid biopsy samples, solid tumor biopsies, normal tissue samples, and / or control samples, are collected from a subject in a clinical environment 220, e.g., a doctor's office, hospital, or medical clinic, or in a home health care setting (not depicted). Advantageously, solid tumor samples should be collected within a clinical setting, while liquid biopsy samples can be obtained in a minimally invasive manner and are more easily collected outside of a traditional clinical setting. In some embodiments, one or more biological samples, or portions thereof, are processed within the clinical environment 220 where collection occurred using a processing device 224, e.g., a nucleic acid sequencer to obtain sequencing data, a microscope to obtain pathology data, a mass spectrometer to obtain proteomic data, etc. In some embodiments, one or more biological samples or portions thereof are sent to one or more external environments, e.g., a sequencing laboratory 230, a pathology laboratory 240, and / or a molecular biology laboratory 250, each of which includes a processing device 234, 244, and 254, respectively, for generating subject biological data 121. Each environment includes a communication device 222, 232, 242, and 252, respectively, for communicating the subject biological data 121 to a processing server 262 and / or database 264, which may be located in yet another environment, e.g., a processing / storage center 260. Thus, in some embodiments, different portions of the systems and methods described herein are performed by different processing devices located in different physical environments.

[0147] Thus, in some embodiments, a method for providing clinical support for personalized cancer therapy, e.g., with improved validation of somatic sequence variants, is implemented across one or more environments, as illustrated in FIG. 2B. For example, in some such embodiments, a liquid biopsy sample is collected in a clinical setting 220 or a home health care setting. The sample, or a portion thereof, is sent to a sequencing laboratory 230, where raw sequence reads 123 of nucleic acids in the sample are generated by a sequencer 234. The raw sequencing data 123 is communicated, for example, from a communication device 232 to a database 264 in a processing / storage center 260, where a processing server 262 extracts features from the sequence reads by performing one or more processes in a bioinformatics module 140, thereby generating a genomic feature 131 for the sample. The processing server 262 can then analyze the identified features by performing one or more processes in a feature analysis module 160, thereby generating a clinical assessment 139, including a clinical report 139-3. The clinician may access the clinical report 139-3 via the recommendation validation module 167, for example, in the processing / storage center 260 or through the communication network 105. After final approval, the clinical report 139-3 is sent to a medical professional, e.g., an oncologist, in the clinical environment 220, who uses the report to support clinical decision-making for the patient's personalized cancer treatment.

[0148] Figure 2A: Exemplary workflow for precision oncology 2A is a flowchart of an exemplary workflow 200 for collecting and analyzing data to generate a clinical report 139 for supporting clinical decision-making in precision oncology. Advantageously, the methods described herein improve this process by improving various stages within feature extraction 206, including, for example, validation of somatic sequence variants.

[0149] Briefly, the workflow begins with patient intake and sample collection 201, in which one or more liquid biopsy samples, one or more tumor biopsies, and one or more normal and / or control tissue samples are collected from a patient (e.g., in a clinical setting 220 or a home health care setting, as illustrated in FIG. 2B). In some embodiments, personal data 126 corresponding to the patient and a record of the one or more biological samples obtained (e.g., patient identifier, patient clinical data, sample type, sample identifier, cancer status, etc.) are entered into a data analysis platform, e.g., laboratory patient data store 120. Thus, in some embodiments, the methods disclosed herein include obtaining one or more biological samples from one or more subjects, e.g., cancer patients. In some embodiments, the subject is a human, e.g., a human cancer patient.

[0150] In some embodiments, one or more of the biological samples obtained from the patient are biological liquid samples, also referred to as liquid biopsy samples. In some embodiments, one or more of the biological samples obtained from the patient are selected from blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testes), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple effluent, aspirates from different parts of the body (e.g., thyroid, breast), etc. In some embodiments, the liquid biopsy sample comprises blood and / or saliva. In some embodiments, the liquid biopsy sample is peripheral blood. In some embodiments, the blood sample is collected from the patient in a commercially available blood collection container, for example, using a PAXgene® Blood DNA Tube. In some embodiments, the saliva sample is collected from the patient in a commercially available saliva collection container, for example, using an Oragene® DNA Saliva Kit.

[0151] In some embodiments, the liquid biopsy sample has a volume of about 1 mL to about 50 mL. For example, in some embodiments, the liquid biopsy sample has a volume of about 1 mL, about 2 mL, about 3 mL, about 4 mL, about 5 mL, about 6 mL, about 7 mL, about 8 mL, about 9 mL, about 10 mL, about 11 mL, about 12 mL, about 13 mL, about 14 mL, about 15 mL, about 16 mL, about 17 mL, about 18 mL, about 19 mL, about 20 mL, or more.

[0152] Liquid biopsy samples contain cell-free nucleic acids, including cell-free DNA (cfDNA).As described above, cfDNA isolated from cancer patients includes DNA derived from cancer cells, also referred to as circulating tumor DNA (ctDNA), cfDNA derived from germline (e.g., healthy or non-cancerous) cells, and cfDNA derived from hematopoietic cells (e.g., leukocytes).The relative proportion of cancerous and non-cancerous cfDNA present in liquid biopsy samples varies depending on the characteristics of the patient's cancer (e.g., type, stage, lineage, genomic profile, etc.).As used herein, the "tumor burden" of a subject refers to the percentage of cfDNA derived from cancer cells.

[0153] As described herein, cfDNA is a particularly useful source of biological data for various embodiments of the methods and systems described herein because it can be easily obtained from various bodily fluids. Advantageously, the use of bodily fluids facilitates continuous monitoring due to the ease of collection, as these fluids can be collected by non-invasive or minimally invasive methodologies. This is in contrast to methods that rely on solid tissue samples, such as biopsies, which often require invasive surgical procedures. Furthermore, because bodily fluids such as blood circulate throughout the body, cfDNA populations represent samples of many different tissue types from many different locations.

[0154] In some embodiments, the liquid biopsy sample is separated into two different samples, for example, in some embodiments, a blood sample is separated into a plasma sample containing cfDNA and a buffy coat preparation containing white blood cells.

[0155] In some embodiments, multiple liquid biopsy samples are obtained from each subject at intervals over a period of time (e.g., using serial testing). For example, in some such embodiments, the time between obtaining liquid biopsy samples from each subject is at least 1 day, at least 2 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 6 months, or at least 1 year.

[0156] In some embodiments, the one or more biological samples collected from a patient are solid tissue samples, such as solid tumor samples or solid normal tissue samples. Methods for obtaining solid tissue samples, e.g., of cancer and / or normal tissues, are known in the art and depend on the type of tissue being sampled. For example, bone marrow biopsy and isolation of circulating tumor cells can be used to obtain samples of blood cancers; endoscopic biopsy can be used to obtain samples of gastrointestinal, bladder, and lung cancers; needle biopsy (e.g., fine needle aspiration, core needle aspiration, vacuum-assisted biopsy, image-guided biopsy) can be used to obtain samples of subcutaneous tumors; skin biopsy (e.g., shave biopsy, punch biopsy, incisional biopsy, and excision biopsy) can be used to obtain samples of skin cancers; and surgical biopsy can be used to obtain samples of cancers affecting the patient's internal organs. In some embodiments, the solid tissue sample is formalin-fixed tissue (FFPE). In some embodiments, the solid tissue sample is grossly dissected formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the solid tissue sample is a fresh frozen tissue sample.

[0157] In some embodiments, a dedicated normal sample is collected from the patient for simultaneous processing with the liquid biopsy sample. Generally, the normal sample is of non-cancerous tissue and can be collected using any of the tissue collection methods described above. In some embodiments, oral cells collected from the inside of the patient's cheek are used as the normal sample. Oral cells can be collected by placing an absorbent material, such as a cotton swab, into the subject's mouth and rubbing it against the cheek for, for example, at least 15 seconds or at least 30 seconds. The swab is then removed from the patient's mouth and inserted into a tube so that the tip of the tube is immersed in a liquid that serves to extract the oral cells from the absorbent material. An example of an oral cell recovery and collection device is provided in U.S. Patent No. 9,138,205, the contents of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, oral swab DNA is used as a source of normal DNA in circulating hematologic malignancies.

[0158] Biological samples collected from patients are optionally sent to various analytical environments (e.g., sequencing lab 230, pathology lab 240, and / or molecular biology lab 250) for processing (e.g., data collection) and / or analysis (e.g., feature extraction). Wet lab processing 204 can include sample cataloging (e.g., enrollment), clinical characterization of one or more samples (e.g., pathology review), and nucleic acid sequence analysis (e.g., extraction, library preparation, capture + hybridization, pooling, and sequencing). In some embodiments, the workflow includes clinical analysis of one or more biological samples collected from a subject, for example, in pathology lab 240 and / or molecular and cell biology lab 250, to generate clinical features such as pathology features 128-3, image data 128-3, and / or tissue culture / organoid data 128-3.

[0159] In some embodiments, pathology data 128-1 collected during a clinical evaluation includes, for example, visual features identified by a pathologist's examination of a specimen (e.g., a solid tumor biopsy) on a stained H&E or IHC slide. In some embodiments, the sample is a solid tissue biopsy sample. In some embodiments, the tissue biopsy sample is formalin-fixed tissue (FFT), e.g., formalin-fixed, paraffin-embedded (FFPE) tissue. In some embodiments, the tissue biopsy sample is an FFPE or FFT block. In some embodiments, the tissue biopsy sample is a fresh-frozen tissue biopsy. The tissue biopsy sample can be prepared in thin sections (e.g., by cutting and / or mounting on slides) to facilitate pathology review (e.g., by staining with immunohistochemical stains for IHC review and / or with hematoxylin and eosin stains for H&E pathology review). For example, analysis of slides for H&E or IHC staining may reveal characteristics such as tumor infiltration, programmed death-ligand 1 (PD-L1) status, human leukocyte antigen (HLA) status, or other immunological features.

[0160] In some embodiments, liquid samples (e.g., blood) collected from patients (e.g., in EDTA-containing collection tubes) are prepared (e.g., by smearing) on ​​slides for pathology review. In some embodiments, grossly dissected FFPE tissue sections, which may be mounted on histopathology slides from solid tissue samples (e.g., tumor or normal tissue), are analyzed by a pathologist. In some embodiments, tumor samples are evaluated to determine, for example, the tumor purity of the sample, the tumor cellularity rate as a ratio of tumor to normal nuclei, etc. For each section, background tissue can be excluded or removed so that the section meets a tumor purity threshold, e.g., at least 20% of the nuclei in the section are tumor nuclei, or at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the nuclei in the section are tumor nuclei.

[0161] In some embodiments, pathology data 128-1 is extracted using a computational approach to digital pathology, in addition to or instead of visual inspection, providing, for example, morphological features extracted from digital images of stained tissue samples. A review of digital pathology methods is provided in Bera, K. et al., Nat. Rev. Clin. Oncol., 16:703-15 (2019), the contents of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, pathology data 128-1 includes features determined using machine learning algorithms to evaluate pathology data collected as described above.

[0162] Further details regarding methods, systems, and algorithms for classifying cancers and identifying targeted therapies using pathology data are discussed, for example, in U.S. Patent Application No. 16 / 830,186, filed March 25, 2020, and U.S. Provisional Application No. 63 / 007,874, filed April 9, 2020, the contents of which are incorporated herein in their entirety for all purposes.

[0163] In some embodiments, the image data 128-2 collected during clinical evaluation includes features identified by review of in vitro and / or in vivo imaging results (e.g., of a tumor site), e.g., tumor size, differences in tumor size over time (e.g., during treatment or other changes), etc. In some embodiments, the image data 128-2 includes features determined using machine learning algorithms to evaluate the image data collected as described above.

[0164] Further details regarding methods, systems, and algorithms for classifying cancer and identifying targeted therapies using medical imaging data are discussed, for example, in U.S. Patent Application No. 16 / 830,186, filed March 25, 2020, and U.S. Provisional Application No. 63 / 007,874, filed April 9, 2020, the contents of which are incorporated herein in their entirety for all purposes.

[0165] In some embodiments, the tissue culture / organoid data 128-3 collected during clinical evaluation includes features identified by evaluation of cultured tissue from a subject. For example, in some embodiments, tissue samples (e.g., tumor tissue, normal tissue, or both) obtained from a patient are cultured (e.g., in liquid culture, solid culture, and / or organoid culture) and various features, such as cell morphology, growth characteristics, genomic alterations, and / or drug sensitivity, are evaluated. In some embodiments, the tissue culture / organoid data 128-3 includes features determined using machine learning algorithms to evaluate the tissue culture / organoid data collected as described above. Examples of tissue organoid (e.g., personal tumor organoid) cultures and their feature extraction are described in U.S. Provisional Patent Application No. 62 / 924,621, filed October 22, 2019, and U.S. Patent Application No. 16 / 693,117, filed November 22, 2019, the contents of which are incorporated herein in their entirety for all purposes.

[0166] Nucleic acid sequencing of one or more samples collected from a subject is performed during wet lab processing 204, for example, in a sequencing lab 230. An exemplary workflow for nucleic acid sequencing is illustrated in Figure 3. In some embodiments, one or more biological samples obtained in the sequencing lab 230 are registered (302) to track the samples and data throughout the sequencing process.

[0167] Nucleic acids, e.g., RNA and / or DNA, are then extracted from one or more biological samples (304). Methods for isolating nucleic acids from biological samples are known in the art and depend on the type of nucleic acid being isolated (e.g., cfDNA, DNA, and / or RNA) and the type of sample from which the nucleic acid is isolated (e.g., liquid biopsy samples, leukocyte buffy coat preparations, formalin-fixed paraffin-embedded (FFPE) solid tissue samples, and fresh-frozen solid tissue samples). The selection of any particular nucleic acid isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the sample type, sample condition, type of nucleic acid being sequenced, and the sequencing technique being used.

[0168] Many techniques for DNA isolation, e.g., genomic DNA isolation, from tissue samples are known in the art, such as organic extraction, silica adsorption, and anion exchange chromatography. Similarly, many techniques for RNA isolation, e.g., mRNA isolation, from tissue samples are known in the art. For example, acid guanidine thiocyanate-phenol-chloroform extraction (see, e.g., Chomczynski and Sacchi, 2006, Nat Protoc, 1(2):581-85, incorporated herein by reference) and silica bead / glass fiber adsorption (see, e.g., Poeckh, T. et al., 2008, Anal Biochem., 373(2):253-62, incorporated herein by reference). The selection of any particular DNA or RNA isolation technique for use in conjunction with the embodiments described herein is well within the skill of one of ordinary skill in the art, taking into account the tissue type, tissue condition (e.g., fresh, frozen, formalin-fixed, paraffin-embedded (FFPE)), and the type of nucleic acid analysis to be performed.

[0169] In some embodiments where the biological sample is a liquid biopsy sample, e.g., a blood or plasma sample, cfDNA is isolated from the blood sample using commercially available reagents containing proteinase K to produce a liquid solution of cfDNA.

[0170] In some embodiments, the isolated DNA molecules are mechanically sheared to an average length using an ultrasonicator (e.g., a Covaris ultrasonicator). In some embodiments, the isolated nucleic acid molecules are analyzed to determine their fragment size, for example, through gel electrophoresis techniques and / or the use of a device such as a LabChip GX Touch. Those skilled in the art will know the appropriate range of fragment sizes based on the sequencing technique being employed, as different sequencing techniques have different fragment size requirements for robust sequencing. In some embodiments, quality control tests are performed on the extracted nucleic acids (e.g., DNA and / or RNA) to, for example, assess nucleic acid concentration and / or fragment size. DNA fragment sizing, such as determining whether DNA fragments require additional shearing prior to sequencing, provides valuable information used in downstream processing.

[0171] Wet lab processing 204 then involves preparing a nucleic acid library from the isolated nucleic acids (e.g., cfDNA, DNA, and / or RNA). For example, in some embodiments, a DNA library (e.g., a gDNA and / or cfDNA library) is prepared from the isolated DNA from one or more biological samples. In some embodiments, the DNA library is prepared using a commercially available library preparation kit, such as a KAPA Hyper Prep Kit, a New England Biolabs (NEB) kit, or a similar kit.

[0172] In some embodiments, adaptors (e.g., UDI adaptors such as Roche SeqCap double-ended adaptors, or UMI adaptors such as full-length or short Y adaptors) are ligated onto nucleic acid molecules during library preparation. In some embodiments, the adaptors comprise unique molecular identifiers (UMIs), which are short nucleic acid sequences (e.g., 3-10 base pairs) that are added to the ends of DNA fragments during adaptor ligation. In some embodiments, the UMIs are degenerate base pairs that serve as unique tags that can be used to identify sequence reads derived from specific DNA fragments. In some embodiments, patient-specific indicators are added to nucleic acid molecules, for example, when multiplex sequencing is used to sequence DNA from multiple samples (e.g., from the same or different subjects) in a single sequencing reaction. In some embodiments, patient-specific indicators are short nucleic acid sequences (e.g., 3-20 nucleotides) that are added to the ends of DNA fragments during library construction that serve as unique tags that can be used to identify sequence reads derived from specific patient samples. Examples of identifier sequences are described, for example, in Kivioja et al., Nat. Methods 9(1):72-74 (2011) and Islam et al., Nat. Methods 11(2):163-66 (2014), the contents of which are incorporated herein by reference in their entirety for all purposes.

[0173] In some embodiments, the adapters contain PCR primer landing sites designed for efficient binding of PCR or second-strand synthesis primers used during sequencing reactions. In some embodiments, the adapters contain anchor binding sites to facilitate binding of DNA molecules to anchor oligonucleotide molecules on a sequencer flow cell, serving as seeds for the sequencing process by providing a starting point for the sequencing reaction. During PCR amplification after adapter ligation, the UMI, patient index, and binding site are replicated along with the attached DNA fragment. This provides a method for identifying sequence reads derived from the same original fragment in downstream analysis.

[0174] In some embodiments, the DNA library is amplified and purified using commercially available reagents (e.g., Axygen MAG PCR cleanup beads). In some such embodiments, the concentration and / or amount of DNA molecules is then quantified using a fluorescent dye and fluorescence microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In some embodiments, library amplification is performed on a device (e.g., Illumina C-Bot2), and the resulting flow cells containing the amplified target capture DNA library are sequenced to a specific on-target depth selected by the user on a next-generation sequencer (e.g., Illumina HiSeq 4000 or Illumina NovaSeq 6000). In some embodiments, DNA library preparation is performed using an automated system that uses a liquid handling robot (e.g., SciClone NGSx).

[0175] In some embodiments in which the feature data 125 includes the methylation state 132 of one or more genomic locations, nucleic acids (e.g., cfDNA) isolated from a biological sample are processed to convert unmethylated cytosines to uracil, for example, before generating a sequencing library. Thus, when the nucleic acids are sequenced, all cytosines called in the sequencing reaction are necessarily methylated, since unmethylated cytosines are converted to uracil and would therefore be called thymidine rather than cytosine in the sequencing reaction. Commercially available kits, such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, and EZ DNA Methylation™-Lightning kits (available from Zymo Research Corp, Irvine, CA), are available for bisulfite-mediated conversion of methylated cytosines to uracil. Commercially available kits, such as the APOBEC-Seq kit (available from NEBiolabs, Ipswich, Mass.), are also available for the enzymatic conversion of methylated cytosine to uracil.

[0176] In some embodiments, the wet lab process 204 includes pooling (308) DNA molecules from multiple libraries corresponding to different samples from the same and / or different patients to form a sequencing pool of DNA libraries. When the pool of DNA libraries is sequenced, the resulting sequence reads correspond to nucleic acids isolated from the multiple samples. The sequence reads can be separated into different sequence read files corresponding to the various samples represented by the sequencing reads based on unique identifiers present in the appended nucleic acid fragments. In this manner, a single sequencing reaction can generate sequence reads from multiple samples. Advantageously, this allows for the processing of more samples per sequencing reaction.

[0177] In some embodiments, wet laboratory processing 204 includes enriching 310 a sequencing library or pool of sequencing libraries for target nucleic acids, e.g., nucleic acids encompassing loci that are beneficial for precision oncology and / or used as internal controls for sequencing or bioinformatics processes. In some embodiments, enrichment is achieved by hybridizing target nucleic acids in the sequencing library to probes that hybridize to the target sequences, and then isolating the captured nucleic acids from off-target nucleic acids that are not bound by the capture probes.

[0178] Advantageously, enriching target sequences prior to sequencing nucleic acids significantly reduces the cost and time associated with sequencing, facilitates multiplex sequencing by allowing multiple samples to be mixed together for a single sequencing reaction, and significantly reduces the computational burden of aligning the resulting sequence reads as a result of significantly reducing the total amount of nucleic acid analyzed from each sample.

[0179] In some embodiments, enrichment is performed prior to pooling multiple nucleic acid sequencing libraries, however, in other embodiments, enrichment is performed after pooling the nucleic acid sequencing libraries, which has the advantage of reducing the number of enrichment assays that need to be performed.

[0180] In some embodiments, enrichment is performed before generating a nucleic acid sequencing library. This has the advantage that fewer reagents are needed to perform both enrichment (because there are fewer target sequences at this point, before library amplification) and library production (because there are fewer nucleic acid molecules to tag and amplify after enrichment). However, in addition to the issues raised above in the Background and Introduction sections, this increases the likelihood of pull-down bias and / or small variations in the enrichment protocol may result in less consistent results.

[0181] In some embodiments, nucleic acid libraries are pooled (two or more DNA libraries can be mixed to create a pool) and treated with a reagent to reduce off-target capture, e.g., human COT-1 and / or IDT xGen Universal Blocker. The pool can be dried and resuspended in a centrifugal concentrator. The DNA library or pool can be hybridized to a probe set (e.g., a probe set specific to a panel including at least 100, 600, 1,000, 10,000, etc. loci of the 19,000 known human genes) and amplified using commercially available reagents (e.g., KAPA HiFi HotStart ReadyMix). For example, in some embodiments, the pool is incubated in an incubator, PCR machine, water bath, or other temperature-regulating device to allow the probes to hybridize. The pool can then be mixed with streptavidin-coated beads or another means to capture hybridized DNA probe molecules, such as DNA molecules representing exons of the human genome and / or genes selected for the gene panel.

[0182] Pools can be amplified and purified more than once using commercially available reagents, such as the KAPA HiFi Library Amplification Kit and Axygen MAG PCR cleanup beads, respectively. Pools or DNA libraries can be analyzed to determine the concentration or quantity of DNA molecules, for example, by using fluorescent dyes (e.g., PicoGreen pool quantification) and a fluorescent microplate reader, a standard fluorescence spectrometer, or a filter fluorometer. In one example, DNA library preparation and / or capture is performed using an automated system that uses a liquid handling robot (e.g., SciClone NGSx).

[0183] In some embodiments, multiple nucleic acid probes (e.g., probe sets) are used to enrich for one or more target sequences in a nucleic acid sample (e.g., an isolated nucleic acid sample or a nucleic acid sequencing library), e.g., where one or more target sequences are beneficial for precision oncology. For example, in some embodiments, one or more of the target sequences encompass a genetic locus associated with an actionable allele. That is, variation in the target sequence is relevant to a targeted therapeutic approach. In some embodiments, one or more of the target sequences and / or one or more characteristics of the target sequences are used in a model, e.g., a machine learning model, trained to distinguish between two or more cancer conditions. For example, in some embodiments, sequence and / or structural variants identified in one or more of the target sequences are used to estimate a patient's blood tumor mass (bTMB). Similarly, in some embodiments, the number of repetitive sequence elements in one or more microsatellite target sequences is used to determine the microsatellite stability of a sample. Similarly, in some embodiments, the copy number of one or more target sequences is used to estimate the circulating tumor fraction of a sample.

[0184] In some embodiments, the probe set includes probes that target one or more genetic loci, e.g., exon or intron loci. In some embodiments, the probe set includes probes that target one or more protein-coding genetic loci, e.g., regulatory genetic loci, miRNA genetic loci, and other non-coding genetic loci, e.g., those found to be associated with cancer. In some embodiments, the plurality of genetic loci includes at least 25, 50, 100, 150, 200, 250, 300, 350, 400, 500, 750, 1000, 2500, 5000, or more human genomic loci.

[0185] In some embodiments, the probe sets described herein target multiple genomic loci. In some embodiments, the probe sets include probes that target at least 10 genomic loci. In some embodiments, the probe sets include probes that target at least 25 genomic loci. In some embodiments, the probe sets include probes that target at least 50 genomic loci. In some embodiments, the probe sets include probes that target at least 100 genomic loci. In some embodiments, the probe sets include probes that target at least 200 genomic loci. In some embodiments, the probe sets include probes that target at least 300 genomic loci. In some embodiments, the probe sets include probes that target at least 400 genomic loci. In some embodiments, the probe sets include probes that target at least 500 genomic loci. In some embodiments, the probe sets include probes that target at least 750 genomic loci. In some embodiments, the probe sets include probes that target at least 1000 genomic loci. In some embodiments, the probe set comprises probes that target at least 2500 genomic loci. In some embodiments, the probe set comprises probes that target at least 5000 genomic loci.

[0186] In some embodiments, the probe set includes probes that target 10,000 or fewer genomic loci. In some embodiments, the probe set includes probes that target 5000 or fewer genomic loci. In some embodiments, the probe set includes probes that target 2500 or fewer genomic loci. In some embodiments, the probe set includes probes that target 1000 or fewer genomic loci. In some embodiments, the probe set includes probes that target 750 or fewer genomic loci.

[0187] In some embodiments, the probe set includes probes targeting between 10 and 10,000 genomic loci. In some embodiments, the probe set includes probes targeting between 10 and 5,000 genomic loci. In some embodiments, the probe set includes probes targeting between 10 and 1,000 genomic loci. In some embodiments, the probe set includes probes targeting between 10 and 750 genomic loci. In some embodiments, the probe set includes probes targeting between 50 and 10,000 genomic loci. In some embodiments, the probe set includes probes targeting between 50 and 5,000 genomic loci. In some embodiments, the probe set includes probes targeting between 50 and 1,000 genomic loci. In some embodiments, the probe set includes probes targeting between 50 and 750 genomic loci. In some embodiments, the probe set includes probes targeting between 100 and 10,000 genomic loci. In some embodiments, the probe set includes probes targeting between 100 and 5,000 genomic loci. In some embodiments, the probe set includes probes targeting 100 to 1000 genomic loci. In some embodiments, the probe set includes probes targeting 100 to 750 genomic loci. In some embodiments, the probe set includes probes targeting 250 to 10,000 genomic loci. In some embodiments, the probe set includes probes targeting 250 to 5000 genomic loci. In some embodiments, the probe set includes probes targeting 250 to 1000 genomic loci. In some embodiments, the probe set includes probes targeting 250 to 750 genomic loci. In some embodiments, the probe set includes probes targeting 500 to 10,000 genomic loci. In some embodiments, the probe set includes probes targeting 500 to 5000 genomic loci. In some embodiments, the probe set includes probes targeting 500 to 1000 genomic loci.In some embodiments, the probe set comprises probes targeting between 500 and 750 genomic loci.

[0188] In some embodiments, the probe sets described herein target multiple genes and / or associated non-coding regions (e.g., promoters, introns, etc.) whose sequences can be analyzed to identify, for example, sequence and / or structural variants to inform clinical treatment of cancer. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 10 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 25 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 50 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 100 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 200 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 300 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 400 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 500 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 750 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 1000 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 2500 genes. In some embodiments, the probe set includes probes that target all or a portion of the coding sequence (CDS) of at least 5000 genes.

[0189] In some embodiments, a probe set includes probes that target all or part of the coding sequences (CDS) of 10,000 or fewer genes. In some embodiments, a probe set includes probes that target all or part of the CDS of 5000 or fewer genes. In some embodiments, a probe set includes probes that target all or part of the CDS of 2500 or fewer genes. In some embodiments, a probe set includes probes that target all or part of the CDS of 1000 or fewer genes. In some embodiments, a probe set includes probes that target all or part of the CDS of 750 or fewer genes.

[0190] In some embodiments, the probe set includes probes targeting the CDSs of 10 to 10,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 10 to 5,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 10 to 1,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 10 to 750 genes. In some embodiments, the probe set includes probes targeting the CDSs of 50 to 10,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 50 to 5,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 50 to 1,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 50 to 750 genes. In some embodiments, the probe set includes probes targeting the CDSs of 100 to 10,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 100 to 5,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 100 to 1,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 100 to 750 genes. In some embodiments, the probe set includes probes targeting the CDSs of 250 to 10,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 250 to 5,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 250 to 1,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 250 to 750 genes. In some embodiments, the probe set includes probes targeting the CDSs of 500 to 10,000 genes. In some embodiments, the probe set includes probes targeting the CDSs of 500 to 5,000 genes. In some embodiments, the probe set comprises probes targeting the CDS of 500 to 1000 genes.In some embodiments, the probe set comprises probes targeting the CDS of 500 to 750 genes.

[0191] In some embodiments, the probe set includes probes that target one or more of the genes listed in List 1, provided below. In some embodiments, the probe set includes probes that target at least five of the genes listed in List 1. In some embodiments, the probe set includes probes that target at least 10 of the genes listed in List 1. In some embodiments, the probe set includes probes that target at least 25 of the genes listed in List 1. In some embodiments, the probe set includes probes that target at least 50 of the genes listed in List 1. In some embodiments, the probe set includes probes that target at least 75 of the genes listed in List 1. In some embodiments, the probe set includes probes that target at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 of the genes listed in List 1. In some embodiments, the probe set includes probes that target all of the genes listed in List 1.

[0192] Line 1(523 ingredients):ABCC3, ABL1, ABL2, ABRAXAS1, ACVR1, ACVR1B, AJUBA, AKT 1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, APLNR, AR, ARAF, ARFRP1, ARID1A. ARID1B, ARID2, ASNS, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AURKC, AXIN1, A XIN2, AXL, B2M, BAP1, BARD1, BAX, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL6, BCL AF1, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BRCA2, BRD4, BRI P1, BTG1, BTG2, BTK, CALR, CARD11, CARM1, CASP8, CBFB, CBL, CCND1, CCND2 CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12 CDK4, CDK6, CDK8, CDK9, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CH D4, CHEK1, CHEK2, CIC, CKS1B, CREBBP, CRKL, CSF1R, CSF3R, CTC1, CTCF, CTL A4, CTNNA1, CTNNB1, CUL3, CUL4A, CUX1, CXCR4, CYLD, CYP17A1, CYSLTR2, DA XX, DDB2, DDR1, DDR2, DDX3X, DDX41, DEPTOR, DICER1, DIS3, DNMT1, DNMT3A. DOT1L, DPYD, EBF1, EED, EEF2, EGFR, EGLN1, EIF1AX, ELF3, EMSY, EP300, EPCA M, EPHA2, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC2, ERCC3, ERCC4 ERCC6, ERG, ERRFI1, ESR1, ETNK1, ETV1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR. FAM46C, FANCA, FANCC, FANCD2, FANCE, FANCG, FANCI, FANCL, FANCM, FAS, FA T1, FBXW7, FCGR2A, FCGR3A, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4FGF6、FGFR1、FGFR2、FGFR3、FGFR4、FH、FHIT、FLCN、FLT1、FLT3、FLT4、FOLH1 、FOXA1、FOXL2、FOXO1、FOXO3、FOXP1、FRS2、FUBP1、GABRA6、GALNT12、GATA1 、GATA3、GATA4、GATA6、GID4、GLI2、GNA11、GNA13、GNAQ、GNAS、GPC3、GPS2、G REM1、GRIN2A、GRM3、GSK3B、GSTP1、H3F3A、HAVCR2、HDAC1、HDAC2、HGF、HIF1A 、HIST1H3B、HLA-B、HNF1A、HNF1B、HOXB13、HRAS、HSD3B1、HSP90AA1、HSPH1、ID3、IDH1、IDH2、IFNA21、IFNAR1、IFNAR2、IFNG、IFNGR1、IFNGR2、IFNW1、IG F1、IGF1R、IKBKE、IKZF1、IL10RA、IL32、IL6R、IL7R、IMPDH1、ING1、INPP4B、INSR、IRF1、IRF2、IRF4、IRS2、JAK1、JAK2、JAK3、JUN、KAT6A、KDM5A、KDM5C、K DM5D、KDM6A、KDR、KEAP1、KEL、KIT、KLF4、KLHL6、KLLN、KMT2A、KMT2C、KMT2D 、KRAS、LATS1、LCK、LMO1、LRP1B、LTK、LYN、LZTR1、MAF、MALT1、MAP2K1、MAP2 K2、MAP2K4、MAP3K1、MAP3K13、MAP3K21、MAP3K7、MAPK1、MAPK3、MAX、MC1R、M CL1、MDM2、MDM4、MED12、MEF2B、MEN1、MERTK、MET、MITF、MKNK1、MLH1、MLH3、M PL、MRE11、MS4A1、MSH2、MSH3、MSH6、MST1R、MTAP、MTHFR、MTOR、MUC16、MUTY H、MYB、MYC、MYCL、MYCN、MYD88、NBN、NCOA2、NCOR1、NF1、NF2、NFE2L2、NFKBI A、NKX2-1、NOTCH1、NOTCH2、NOTCH3、NOTCH4、NPM1、NQO1、NRAS、NRG1、NSD1、 NSD2、NSD3、NT5C2、NTRK1、NTRK2、NTRK3、NUTM1、P2RY8、PAK1、PALB2、PALLD、PARP1、PARP2、PARP3、PAX5、PBRM1、PDCD1、PDCD1LG2、PDGFRA、PDGFRB、PDK1 、PHGDH、PHLPP1、PHLPP2、PIAS4、PIK3C2B、PIK3C2G、PIK3CA、PIK3CB、PIK3CD 、PIK3CG、PIK3R1、PIK3R2、PIM1、PLCG1、PLCG2、PMS1、PMS2、POLA1、POLD1、P OLE、POLQ、POT1、PPARG、PPM1D、PPP2R1A、PPP2R2A、PPP6C、PRDM1、PREX2、PRK ACA、PRKAR1A、PRKCI、PRKN、PTCH1、PTEN、PTK2、PTPN11、PTPN13、PTPRD、PTP RO、PTPRT、QKI、RAC1、RAD21、RAD50、RAD51、RAD51B、RAD51C、RAD51D、RAD52、 RAD54L、RAF1、RARA、RASA1、RB1、RBM10、RECQL4、REL、RET、RHEB、RHOA、RICT OR、RIT1、RNF43、ROS1、RPS6KB1、RPTOR、RRM1、RSF1、RSPO2、RUNX1、RXRA、SDC 4、SDHA、SDHAF2、SDHB、SDHC、SDHD、SETBP1、SETD2、SF3B1、SGK1、SIRPA、SLC34A2、SLC9A3R1、SLFN11、SLIT2、SMAD2、SMAD3、SMAD4、SMARCA2、SMARCA4、SM ARCB1、SMC1A、SMC3、SMO、SNCAIP、SOCS1、SOS1、SOX2、SOX9、SPEN、SPOP、SRC、SRSF2、STAG2、STAT3、STAT5B、STAT6、STK11、SUFU、SUZ12、SYK、TBX3、TCF7L 2、TEK、TERC、TERT、TET2、TFEB、TGFB1、TGFBR1、TGFBR2、TIGIT、TIPARP、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TNFRSF17、TOP1、TOP2A、TP53、TP53BP1、 TP63、TRAF3、TRAF7、TSC1、TSC2、TSHR、TYMS、TYRO3、U2AF1、UGT1A1、VEGFA、V HL、VSIR、WEE1、WNK1、WRN、WT1、XBP1、XPA、XPC、XPO1、XRCC1、XRCC2、YEATS4、ZFHX3, ZMYM3, ZNF217, ZNF703, ZNF750, ZNRF3, and ZRSR2.

[0193] In some embodiments, a probe set described herein includes a first subset of probes targeting a first subset of genes used at a first average concentration in a hybridization reaction and a second subset of probes targeting a second subset of genes used at a second concentration that is 5-8 times higher than the first average concentration. In some embodiments, the first subset of genes includes at least 10 of the genes listed in List 1, and the second subset of genes includes at least 10 of the genes listed in List 1. In some embodiments, the first subset of genes includes at least 25 of the genes listed in List 1, and the second subset of genes includes at least 25 of the genes listed in List 1. In some embodiments, the first subset of genes includes at least 50 of the genes listed in List 1, and the second subset of genes includes at least 50 of the genes listed in List 1. In some embodiments, the first subset of genes comprises at least 75 of the genes listed in List 1, and the second subset of genes comprises at least 75 of the genes listed in List 1. In some embodiments, the first subset of genes comprises at least 100 of the genes listed in List 1, and the second subset of genes comprises at least 100 of the genes listed in List 1.

[0194] In some embodiments, where a first subset of probes targets a first subset of genes used at a first average concentration in a hybridization reaction and a second subset of probes targets a second subset of genes used at a second concentration that is 5-8 times higher than the first average concentration, the first subset of genes includes at least 10 of the genes listed in List 2 and the second subset of genes includes at least 10 of the genes listed in List 3. In some such embodiments, the first subset of genes includes at least 25 of the genes listed in List 2 and the second subset of genes includes at least 25 of the genes listed in List 3. In some such embodiments, the first subset of genes includes at least 50 of the genes listed in List 2 and the second subset of genes includes at least 50 of the genes listed in List 3. In some such embodiments, the first subset of genes includes at least 75 of the genes listed in List 2 and the second subset of genes includes at least 75 of the genes listed in List 3. In some such embodiments, the first subset of genes comprises at least 100 of the genes listed on List 2, and the second subset of genes comprises at least 100 of the genes listed on List 3. In some such embodiments, the first subset of genes comprises at least 200 of the genes listed on List 2, and the second subset of genes comprises at least 100 of the genes listed on List 3. In some such embodiments, the first subset of genes comprises at least 300 of the genes listed on List 2, and the second subset of genes comprises at least 100 of the genes listed on List 3. In some such embodiments, the first subset of genes comprises at least 400 of the genes listed on List 2, and the second subset of genes comprises at least 100 of the genes listed on List 3.In some such embodiments, the first subset of genes includes all of the genes listed in list 2, and the second subset of genes includes all of the genes listed in list 3.

[0195] Liquid 2 (409 ingredients): ABCC3, ABL2, ABRAXAS1, ACVR1, ACVR1B, AJUBA, AKT3, ALO X12B, AMER1, APLNR, ARFRP1, ARID1B, ARID2, ASNS, ASXL1, ATRX, AURKA, AUR KB, AURKC, AXIN1, AXIN2, AXL, BARD1, BAX, BCL2, BCL2L1, BCL2L11, BCL2L2 BCL6, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRD4, BRIP1, BTG1 TG2, CALR, CARD11, CARM1, CASP8, CBFB, CBL, CD22, CD70, CD74, CD79A, CD79 B. CDC73, CDK8, CDK9, CDKN1A, CDKN1B, CDKN2B, CDKN2C, CEBPA, CHD4, CHEK1 CIC, CKS1B, CREBBP, CSF1R, CSF3R, CTC1, CTCF, CTLA4, CTNNA1, CUL3, CUL4 A, CUX1, CXCR4, CYLD, CYP17A1, CYSLTR2, DAXX, DDB2, DDR1, DDX3X, DDX41, D EPTOR, DICER1, DIS3, DNMT1, DNMT3A, DOT1L, DPYD, EBF1, EED, EEF2, EGLN1 EIF1AX, ELF3, EMSY, EP300, EPCAM, EPHA2, EPHA3, EPHB1, EPHB4, ERBB4, ERC C2、ERCC3、ERCC4、ERCC6、ERG、ETNK1、ETV1、ETV4、ETV5、ETV6、EWSR1、EZR、F AM46C, FANCA, FANCC, FANCD2, FANCE, FANCG, FANCI, FANCL, FANCM, FAS, FAT 1, FCGR2A, FCGR3A, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6 H、FHIT、FLCN、FLT1、FLT4、FOLH1、FOXA1、FOXO1、FOXO3、FOXP1、FRS2、FUBP1 、GABRA6、GALNT12、GATA1、GATA4、GATA6、GID4、GLI2、GNA13、GPC3、GPS2、GR EM1, GRIN2A, GRM3, GSK3B, GSTP1, H3F3A, HAVCR2, HDAC1, HDAC2, HGF, HIF1A.HIST1H3B、HLA-B、HNF1B、HOXB13、HSD3B1、HSP90AA1、HSPH1、ID3、IFNA21、I FNAR1、IFNAR2、IFNG、IFNGR1、IFNGR2、IFNW1、IGF1、IGF1R、IKBKE、IKZF1、IL 10RA、IL32、IL6R、IL7R、IMPDH1、ING1、INPP4B、INSR、IRF1、IRF2、IRF4、IRS2、JUN、KAT6A、KDM5A、KDM5C、KDM5D、KDM6A、KEL、KLF4、KLHL6、KLLN、KMT2C、K MT2D、LATS1、LCK、LMO1、LRP1B、LTK、LYN、LZTR1、MAF、MALT1、MAP2K4、MAP3K 1、MAP3K13、MAP3K21、MAP3K7、MAX、MC1R、MCL1、MDM4、MED12、MEF2B、MEN1、ME RTK、MITF、MKNK1、MLH3、MRE11、MS4A1、MST1R、MTAP、MTHFR、MUC16、MUTYH、M YB、MYCL、NBN、NCOA2、NCOR1、NFKBIA、NKX2-1、NOTCH2、NOTCH3、NOTCH4、NQO1 、NRG1、NSD1、NSD2、NSD3、NT5C2、NUTM1、P2RY8、PAK1、PALLD、PARP1、PARP2、 PARP3、PAX5、PDCD1、PDK1、PHGDH、PHLPP1、PHLPP2、PIAS4、PIK3C2B、PIK3C2G 、PIK3CB、PIK3CD、PIK3CG、PIK3R2、PIM1、PLCG1、PLCG2、PMS1、POLA1、POLD1 、POLE、POLQ、POT1、PPARG、PPM1D、PPP2R1A、PPP2R2A、PPP6C、PRDM1、PREX2、P RKACA、PRKAR1A、PRKCI、PRKN、PTK2、PTPN13、PTPRD、PTPRO、PTPRT、QKI、RAC 1、RAD21、RAD50、RAD51、RAD51B、RAD51D、RAD52、RAD54L、RARA、RASA1、RBM10 、RECQL4、REL、RICTOR、RPS6KB1、RPTOR、RRM1、RSF1、RSPO2、RUNX1、RXRA、SD C4、SDHAF2、SDHB、SDHC、SDHD、SETBP1、SETD2、SF3B1、SGK1、SIRPA、SLC34A2、SLC9A3R1, SLFN11, SLIT2, SMAD2, SMAD3, SMARCA2, SMARCA4, SMARCB1, SMC1A, SMC3, SNCAIP, SOCS1, SOS1, SOX2, SOX9, SPEN, SRC, SRSF2, STAG2, STAT3, STAT5B, STAT6, SUFU, SUZ12, SYK, TBX3, TCF7L2, TEK, TERC, TET2, TFEB, TGFB1, TGFBR1, TGFBR2, TIGIT, TIPARP, TMEM127, TMPRSS2, TNFAIP3, TNFRSF14, TNFRSF17, TOP1, TOP2A, TP53BP1, TP63, TRAF3, TRAF7, TSHR, TYMS, TYRO3, U2AF1, UGT1A1, VSIR, WEE1, WNK1, WRN, WT1, XBP1, XPA, XPC, XPO1, XRCC1, XRCC2, YEATS4, ZFHX3, ZMYM3, ZNF217, ZNF703, ZNF750, ZNRF3, and ZRSR2.

[0196] List 3 (114 genes): ABL1, AKT1, AKT2, ALK, APC, AR, ARAF, ARID1A, ATM, ATR, B2M, BAP1, BRAF, BRCA1, BRCA2, BTK, CCND1, CCND2, CCND3, CCNE1, CD274, CDH1, CDK12, CDK4, CDK6, CDKN2A, CHEK2, CRKL, CTNNB1, DDR2, EGFR, ERBB2, ERBB3, ERRFI1, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FLT3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KEAP1, KIT, KMT2A, KRAS,MAP2K1, MAP2K2, MAPK1, MAPK3, MDM2, MET, MLH1, MPL, MSH2, MSH3, MSH6, MTOR, MYC, MYCN, MYD88, NF1, NF2, NFE2L2, NOTCH1, NPM1, NRAS, NTRK1, NTRK2, NTRK3, PALB2, PBRM1, PDCD1LG2, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PMS2, PTCH1, PTEN, PTPN11, RAD51C, RAF1, RB1, RET, RHEB, RHOA, RIT1, RNF43, ROS1, SDHA, SMAD4, SMO, SPOP, STK11, TERT, TP53, TSC1, TSC2, VEGFA, and VHL.

[0197] In some embodiments, in addition to including probes targeting some or all of the CDSs, the probe set includes probes directed to some and all of the introns of selected genes, e.g., genes whose sequence and / or structural variants are known to be associated with a disease or disorder, such as cancer. In some embodiments, the probe set includes probes directed to some or all of the introns of the genes listed in List 4. In some embodiments, the probe set includes probes directed to some or all of the introns of at least five genes listed in List 4. In some embodiments, the probe set includes probes directed to some or all of the introns of at least ten genes listed in List 4. In some embodiments, the probe set includes probes directed to some or all of the introns of all of the genes listed in List 4. In some embodiments, the probes directed to the introns of the genes are used in the hybridization reaction at an enriched concentration, e.g., 5-8 times the base concentration at which probes directed to the "non-enriched" genes are used in the reaction. In other embodiments, the probes directed to the introns of the genes are used in the hybridization reaction at an enriched base concentration, e.g., a non-enriched concentration. In some embodiments, probes directed to introns of some genes are used in hybridization reactions at enhanced concentrations, e.g., 5-8 times the base concentration, and probes directed to introns in other genes are used in hybridization reactions at base concentration.

[0198] List 4: ALK, BRAF, EGFR, FGFR1, FGFR2, FGFR3, NTRK1, NTRK2, NTRK3, RET, and ROS1.

[0199] In some embodiments, in addition to including probes targeting some or all of the CDSs, the probe set includes probes directed to promoter regions of genes. In some embodiments, the probe set includes a probe directed to the TERT gene. In some embodiments, the probes directed to promoter regions of genes are used in hybridization reactions at enriched concentrations, e.g., 5-8 times the base concentration at which probes directed to "non-enhanced" genes are used in the reactions. In other embodiments, the probes directed to promoter regions of genes are used in hybridization reactions at enriched concentrations, e.g., 5-8 times the base concentration at which probes directed to "non-enhanced" genes are used in the reactions. In some embodiments, probes directed to promoter regions of some genes are used in hybridization reactions at enriched concentrations, e.g., 5-8 times the base concentration, and probes directed to promoter regions in other genes are used in hybridization reactions at enriched concentrations.

[0200] In some embodiments, sequences generated from one or more target reads are evaluated for gene fusions, e.g., in addition to being evaluated for SNVs and / or MNVs and / or other variants. In some embodiments, a probe set includes probes targeting at least five genes whose sequences are to be evaluated for gene fusions. In some embodiments, a probe set includes probes targeting at least ten genes whose sequences are to be evaluated for gene fusions. In some embodiments, a probe set includes probes directed to at least five of the genes listed in List 5 whose sequences are to be evaluated for gene fusions. In some embodiments, a probe set includes probes directed to all of the genes listed in List 5 whose sequences are to be evaluated for gene fusions.

[0201] List 5: ALK, BRAF, FGFR1, FGFR2, FGFR3, NTRK1, NTRK2, NTRK3, RET, and ROS1.

[0202] In some embodiments, sequences generated from one or more target reads are evaluated for local copy number variation, e.g., in addition to being evaluated for SNV and / or MNV and / or other variants. In some embodiments, a probe set includes probes targeting at least five genes whose sequences are to be evaluated for local copy number variation. In some embodiments, a probe set includes probes targeting at least 10 genes whose sequences are to be evaluated for local copy number variation. In some embodiments, a probe set includes probes directed to at least five of the genes listed in List 6 whose sequences are to be evaluated for local copy number variation. In some embodiments, a probe set includes probes directed to all of the genes listed in List 6 whose sequences are to be evaluated for local copy number variation.

[0203] List 6: BRCA1, BRCA2, CCNE1, CD274, EGFR, ERBB2, MDM2, MET, and MYC.

[0204] Generally, probes for enrichment of nucleic acids (e.g., cfDNA obtained from a liquid biopsy sample) comprise DNA, RNA, or modified nucleic acid structures having a base sequence complementary to a locus of interest. For example, a probe designed to hybridize to a locus in a cfDNA molecule can comprise a sequence complementary to either strand, since cfDNA molecules are double-stranded. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 consecutive bases of the locus of interest. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 20, 25, 30, 40, 50, 75, 100, 150, 200, or more consecutive bases of the locus of interest.

[0205] Target panel provides several benefits for nucleic acid sequencing.For example, in some embodiments, for example, the algorithm for distinguishing between a first cancer state and a second cancer state can be trained with smaller and more informative data sets (for example, fewer genes), which leads to more computationally efficient training of the classifier for distinguishing between a first cancer state and a second cancer state.This improvement in computational efficiency due to the reduction in the size of the gene set for distinguishing can be advantageously used to accelerate classifier training, or can be used to improve the performance of such classifier (for example, through more extensive training of classifier).

[0206] In some embodiments, the probe comprises an additional nucleic acid sequence that does not share any homology with the locus of interest. For example, in some embodiments, the probe also comprises an identifier sequence, e.g., a nucleic acid sequence comprising a unique molecular identifier (UMI), that is specific to a particular sample or subject. For examples of identifier sequences, see, e.g., Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, which are incorporated herein by reference. Similarly, in some embodiments, the probe also comprises a primer nucleic acid sequence useful for amplifying the nucleic acid molecule of interest, e.g., using PCR. In some embodiments, the probe also comprises a capture sequence designed to hybridize to an anti-capture sequence to retrieve the nucleic acid molecule of interest from the sample.

[0207] Similarly, in some embodiments, each probe comprises a non-nucleic acid affinity moiety covalently bound to a nucleic acid molecule complementary to a locus of interest for retrieving the nucleic acid molecule of interest. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probes are attached to a solid surface or particle, such as a dipstick or magnetic bead, for retrieving the nucleic acid of interest. In some embodiments, the methods described herein include amplifying the nucleic acid bound to the probe set prior to further analysis, e.g., sequencing. Methods for amplifying nucleic acids, for example, by PCR, are well known in the art.

[0208] Sequence reads are then generated from the sequencing library or pool of sequencing libraries (312). Sequencing data can be obtained by any methodology known in the art, such as sequencing-by-synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent Sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing-by-ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or next-generation sequencing (NGS) technology, such as paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing-by-synthesis with reversible dye terminators. In some embodiments, sequencing is performed using next-generation sequencing technology, such as short-read technology. In other embodiments, long-read sequencing or another sequencing method known in the art is used.

[0209] Next-generation sequencing generates millions of short reads (e.g., sequence reads) for each biological sample. Thus, in some embodiments, the multiple sequence reads obtained by next-generation sequencing of cfDNA molecules are DNA sequence reads. In some embodiments, the sequence reads have an average length of at least 50 nucleotides. In other embodiments, the sequence reads have an average length of at least 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, or more nucleotides.

[0210] In some embodiments, sequencing is performed after enriching nucleic acids (e.g., cfDNA, gDNA, and / or RNA) encompassing multiple predetermined target sequences, e.g., human genes and / or non-coding sequences associated with cancer. Advantageously, sequencing a nucleic acid sample enriched for target nucleic acids, rather than all nucleic acids isolated from a biological sample, significantly reduces the average time and cost of a sequencing reaction. Thus, in some embodiments, the methods described herein include obtaining multiple sequence reads of nucleic acids hybridized to a probe set for hybrid capture enrichment (e.g., of one or more genes listed in Lists 1-6).

[0211] In some embodiments, panel targeted sequencing is performed to an average on-target depth of at least 500x, at least 750x, at least 1000x, at least 2500x, at least 500x, at least 10,000x, or more. In some embodiments, samples are further evaluated for uniformity above a sequencing depth threshold (e.g., 95% of all target base pairs at 300x sequencing depth). In some embodiments, the sequencing depth threshold is a minimum depth selected by a user or practitioner.

[0212] In some embodiments, sequence reads are obtained by whole-genome or whole-exome sequencing methodologies. In some such embodiments, whole-exome capture is performed using an automated system that uses a liquid-handling robot (e.g., SciClone NGSx). Because many loci are being sequenced, whole-genome sequencing, and to some extent whole-exome sequencing, is typically performed at a lower sequencing depth than smaller, targeted panel sequencing reactions. For example, in some embodiments, whole-genome or whole-exome sequencing is performed to an average sequencing depth of at least 3x, at least 5x, at least 10x, at least 15x, at least 20x, or more. In some embodiments, low-pass whole-genome sequencing (LPWGS) technology is used for whole-genome or whole-exome sequencing. LPWGS is typically performed to an average sequencing depth of about 0.25x to about 5x, more typically to an average sequencing depth of about 0.5x to about 3x.

[0213] Due to differences in sequencing methodology, data obtained from targeted panel sequencing may be more suitable for certain analyses than data obtained from whole genome / whole exome sequencing, and vice versa. For example, due to the higher sequencing depth achieved by targeted panel sequencing, the resulting sequence data is more suitable for identifying variant alleles present at low allele fractions in samples, for example, less than 20%. In contrast, data generated from whole genome / whole exome sequencing is more suitable for estimating whole genome metrics, such as tumor mutation burden, because the entire genome is well represented by the sequencing data. Therefore, in some embodiments, a nucleic acid sample, such as cfDNA, gDNA, or mRNA sample, is evaluated using both targeted panel sequencing and whole genome / whole exome sequencing (e.g., LPG Seq).

[0214] In some embodiments, raw sequence reads resulting from a sequencing reaction are output from the sequencer in a native file format, e.g., a BCL file. In some embodiments, the native file is passed directly to a bioinformatics pipeline (e.g., variant analysis 206), the components of which are described in detail below. In other embodiments, preprocessing is performed before passing the sequences to a bioinformatics platform. For example, in some embodiments, the format of the sequence read file is converted from the native file format (e.g., BCL) to a file format compatible with one or more algorithms used in the bioinformatics pipeline (e.g., FASTQ or FASTA). In some embodiments, the raw sequence reads are filtered to remove sequences that do not meet one or more quality thresholds. In some embodiments, raw sequence reads generated from the same unique nucleic acid molecule in a sequencing read are collapsed into a single sequence read representing the molecule, e.g., using the UMIs described above. In some embodiments, one or more of these preprocessing activities are performed within the bioinformatics pipeline itself.

[0215] In one example, a sequencer can generate a BCL file. The BCL file can contain raw image data for multiple patient samples to be sequenced. The BCL image data is an image of the flow cell for each cycle during sequencing. Cycles can be implemented by illuminating the patient sample with specific wavelengths of electromagnetic radiation to generate multiple images that can be processed into base calls via a BCL-to-FASTQ processing algorithm that identifies which base pairs are present in each cycle. The resulting FASTQ file contains the entirety of the reads for each patient sample, matched with a quality metric ranging from 0 to 64, for example, where 64 is the best quality and 0 is the worst quality. In embodiments where both liquid biopsy samples and normal tissue samples are sequenced, the sequence reads in the corresponding FASTQ files can be matched so that a liquid biopsy-normal analysis can be performed.

[0216] The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants or copy number variations are present in a sample. Each FASTQ file contains reads, which may be paired-end or single reads and may be short or long reads. Each read represents the sequence of one detected nucleotide in a nucleic acid molecule or copy of a nucleic acid molecule isolated from a patient sample, as detected by a sequencer. Each read in a FASTQ file is also associated with a quality assessment. The quality assessment may reflect the likelihood of an error occurring during the sequencing procedure that affected the associated read. In some embodiments, the paired-end sequencing results for each isolated nucleic acid sample are contained in a pair of FASTQ files, separated for efficiency. Thus, in some embodiments, the forward (read 1) and reverse (read 2) sequences for each isolated nucleic acid sample are stored separately, but in the same order and under the same identifier.

[0217] In various embodiments, the bioinformatics pipeline can filter the FASTQ data from the corresponding sequence data file for each respective biological sample. Such filtering can include correcting or masking sequencer errors, and removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, over-represented sequences, biases caused by library preparation, amplification, or capture, and other errors.

[0218] Although workflow 200 illustrates obtaining a biological sample, extracting nucleic acids from the biological sample, and sequencing the isolated nucleic acids, in some embodiments, the sequencing data used in the improved systems and methods described herein (e.g., including improved methods for validating somatic sequence variants in test subjects with cancer conditions) is obtained by receiving previously generated sequence reads in electronic form.

[0219] Referring again to Figure 2A, nucleic acid sequencing data 122 generated from one or more patient samples is then evaluated in a bioinformatics pipeline (e.g., via variant analysis 206), e.g., using bioinformatics module 140 of system 100, to identify genomic alterations and other metrics in the patient's cancer genome. An exemplary overview of a bioinformatics pipeline is described below with respect to Figures 4A-4E. Advantageously, in some embodiments, the present disclosure improves bioinformatics pipelines such as pipeline 206 by improving methods and systems for validating somatic sequence variants.

[0220] 4A shows an exemplary bioinformatics pipeline 206 for providing clinical support for precision oncology (e.g., used for feature extraction in the workflows illustrated in FIGS. 2A and 3). As shown in FIG. 4A, sequencing data 122 (e.g., sequence reads 314) obtained from wet laboratory processing 204 are input into the pipeline.

[0221] In various embodiments, the bioinformatics pipeline includes a circulating tumor DNA (ctDNA) pipeline for analyzing liquid biopsy samples. The pipeline can detect SNVs, INDELs, copy number amplifications / deletions, and genomic rearrangements (e.g., fusions). The pipeline can employ a unique molecular landmark (UMI)-based consensus-based error suppression method as well as Bayesian trinucleotide context-based position-level error suppression. In various embodiments, this can detect variants with a variant allele fraction of 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.4%, or 0.5%.

[0222] In some embodiments, the sequencing data is processed (e.g., using sequence data processing module 141) to prepare data for genomic feature identification 385. For example, in some embodiments described above, the sequencing data is in the native file format provided by the sequencer. Thus, in some embodiments, the system (e.g., system 100) applies pre-processing algorithms 142 to convert the file format into one recognized by one or more upstream processing algorithms (318). For example, BCL file output from a sequencer can be converted to FASTQ file format using bcl2fastq or bcl2fastq2 conversion software (Illumina®). The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine what genetic variants, copy number changes, etc., are present in the sample.

[0223] In some embodiments, other preprocessing functions are performed, such as filtering sequence reads 122 based on desired quality, e.g., size and / or quality of base calls. In some embodiments, quality control checks are performed to ensure that data is sufficient for variant calling. For example, entire reads, individual nucleotides, or multiple nucleotides likely to have errors can be discarded based on a quality assessment associated with the reads in the FASTQ file, the known error rate of the sequencer, and / or a comparison between each nucleotide in the read and one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools, such as Skewer. See Jiang, H. et al., BMC Bioinformatics 15(182):1-12 (2014). FASTQ files can be analyzed for quality control and rapid assessment of reads by sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC, or another similar software program. For paired-end reads, the reads can be merged.

[0224] In some embodiments, when both a liquid biopsy sample and a normal tissue sample from a patient are sequenced, two FASTQ output files are generated: one for the liquid biopsy sample and one for the normal tissue sample. A "matched" (e.g., panel-specific) workflow is implemented to jointly analyze the matched liquid biopsy-normal FASTQ files. When a matched normal sample is not available from the patient, the FASTQ file from the liquid biopsy sample is analyzed in "tumor-only" mode. See, for example, Figure 4B. When two or more patient samples, such as a liquid biopsy sample and a normal tissue sample, are processed simultaneously on the same sequencer flow cell, differences in the sequence of the adapters used for each patient sample barcode the nucleic acids extracted from both samples, associating each read with the correct patient sample and facilitating assignment to the correct FASTQ file.

[0225] For efficiency, in some embodiments, the paired-end sequencing results for each isolate are contained in a separate pair of FASTQ files. The forward (read 1) and reverse (read 2) sequences for each tumor and normal isolate are stored separately, but in the same order and under the same identifier. See, for example, Figure 4C. In various embodiments, the bioinformatics pipeline may filter the FASTQ data from each isolate. Such filtering may include correcting or masking sequencer errors, as well as removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, overrepresented sequences, biases caused by library preparation, amplification, or capture, and other errors. See, for example, Figure 4D.

[0226] Similarly, in some embodiments, sequencing (312) is performed on a pool of nucleic acid sequencing libraries prepared from different biological samples, e.g., from the same or different patients. Accordingly, in some embodiments, the system demultiplexes (320) the data (e.g., using a demultiplexing algorithm 144) to separate sequence reads into separate files for each sequencing library included in the sequencing pool, e.g., based on UMIs or patient identifier sequences added to the nucleic acid fragments during sequencing library preparation, as described above. In some embodiments, the demultiplexing algorithm is part of the same software package as one or more preprocessing algorithms 142. For example, bcl2fastq or bcl2fastq2 conversion software (Illumina®) includes instructions for both converting the native file format output from the sequencer and demultiplexing the sequence reads 122 output from the reaction.

[0227] The sequence reads are then aligned (322) to a reference sequence construct 158, such as a reference genome, reference exome, or other reference construct prepared for a particular targeted panel sequencing reaction, using, for example, an alignment algorithm 143. For example, in some embodiments, individual sequence reads 123 in electronic form (e.g., in a FASTQ file) are aligned to a reference sequence construct for the species of interest (e.g., a reference human genome) by identifying the sequence in the region of the reference sequence construct that best matches the sequence of nucleotides in the sequence read. In some embodiments, the sequence reads are aligned to the reference exome or reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information may also include the sequence read length, which may be determined from the start and end positions. The region in the reference genome may relate to a gene or a segment of a gene. Any of a variety of alignment tools can be used for this task.

[0228] For example, local sequence alignment algorithms compare different lengths of subsequences in query sequences (e.g., sequence reads) with subsequences in target sequences (e.g., reference constructs) to create the best alignment for each part of the query sequence.In contrast, global sequence alignment algorithms align the entire sequence, for example, end-to-end.Examples of local sequence alignment algorithms include the Smith-Waterman algorithm (see, for example, Smith and Waterman, J Mol. Biol., 147(1):195-97(1981) which is incorporated herein by reference), Lalign (see, for example, Huang and Miller, Adv. Appl. Math, 12:337-57(1991) which is incorporated herein by reference), and PatternHunter (see, for example, Ma B. et al., Bioinformatics, 18(3):440-45(2002) which is incorporated herein by reference).

[0229] In some embodiments, the read mapping process begins by constructing an index of either the reference genome or the read, which is then used to retrieve a set of positions in the reference sequence to which the read is more likely to align. Once this subset of possible mapping positions is identified, alignment is performed within these candidate regions using slower, more sensitive algorithms. See, for example, Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14:184, and Flicek and Birney, 2009, "Sense from sequence reads: methods for alignment and assembly," Nat Methods 6 (Suppl. 11), S6-S12, each of which is incorporated herein by reference. In some embodiments, the mapping tool methodology utilizes a hash table or the Burrows-Wheeler transform (BWT). See, for example, Li and Homer, 2010, “A survey of sequence alignment algorithms for next-generation sequencing,” Brief Bioinformatics 11, pp. 473-483, which is incorporated herein by reference.

[0230] Other software programs designed to align reads include, for example, Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and / or programs using the Smith-Waterman algorithm. Candidate reference genomes include, for example, hg19, GRCh38, hg38, GRCh37, and / or other reference genomes developed by the Genome Reference Consortium. In some embodiments, the alignment generates a SAM file, which stores the start and end positions of each read, along with its coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome.

[0231] For example, in some embodiments, each read in the FASTQ file is aligned to the location in the human genome that best matches the sequence of nucleotides in the read. Many software programs designed to align reads exist, including Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and programs using the Smith-Waterman algorithm. Alignment can be directed using a reference genome (e.g., hg19, GRCh38, hg38, GRCh37, or other reference genomes developed by the Genome Reference Consortium) by comparing the nucleotide sequence in each read with portions of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that most likely corresponds to the sequence in the read. In some embodiments, one or more SAM files are generated for the alignment, which store the start and end positions of each read according to their coordinates in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. The SAM file can be converted to a BAM file. In some embodiments, the BAM file is sorted, and duplicate reads are marked for deletion, resulting in a de-duplicated BAM file.

[0232] In some embodiments, adapter-trimmed FASTQ files are aligned to the 19th edition of the Human Reference Genome Build (HG19) using the Burrows-Wheeler Aligner (BWA, Li and Durbin, Bioinformatics, 25(14):1754-60(2009)). After alignment, reads are grouped by alignment position and UMI family and collapsed into a consensus sequence using, for example, the fgbio tool (fulcrumgenomics.github.io / fgbio / ). Bases with insufficient quality or significant discrepancies between family members (e.g., when it is uncertain whether a base is adenine, cytosine, guanine, etc.) can be replaced with N to represent the wild-type nucleotide type. PHRED scores are then scaled based on the initial base call estimates combined across all family members. After single-strand consensus generation, a double-strand consensus sequence is generated by comparing the forward and reverse PCR products with the mirrored UMI sequence. In various embodiments, consensus can be generated across read pairs. Otherwise, single-strand consensus calling is used. After consensus calling, filtering is performed to remove low-quality consensus fragments. Then, the consensus fragments are realigned to the human reference genome using BWA. After realignment, a BAM output file is generated, which is then sorted and indexed by alignment position.

[0233] In some embodiments, when both a liquid biopsy sample and a normal tissue sample are analyzed, this process generates a liquid biopsy BAM file (e.g., liquid BAM 124-1-i-cf) and a normal BAM file (e.g., germline BAM 124-1-ig), as illustrated in Figure 4A. In various embodiments, the BAM files can be analyzed to detect genetic variants and other genetic features, including single nucleotide variants (SNVs), copy number variants (CNVs), gene rearrangements, etc.

[0234] In some embodiments, sequencing data is normalized to account for, for example, pull-down, amplification, and / or sequencing bias (e.g., mappability, GC bias, etc.). See, e.g., Schwartz et al., PLoS ONE 6(1):e16685 (2011), and Benjamini and Speed, Nucleic Acids Research 40(10):e72 (2012), the contents of which are incorporated by reference in their entirety for all purposes.

[0235] In some embodiments, the SAM file generated after alignment is converted to a BAM file 124. Thus, after preprocessing the sequencing data generated for the pooled sequencing reactions, a BAM file is generated for each of the sequencing libraries present in the master sequencing pool. For example, as illustrated in FIG. 4A , a separate BAM file is generated for each of three samples obtained from subject 1 at time i (e.g., tumor BAM 124-1-it corresponding to the alignment of sequence reads of nucleic acids isolated from solid tumor samples from subject 1, liquid BAM 124-1-i-cf corresponding to the alignment of sequence reads of nucleic acids isolated from liquid biopsy samples from subject 1, and germline BAM 124-1-ig corresponding to the alignment of sequence reads of nucleic acids isolated from normal tissue samples from subject 1), and one or more samples obtained from one or more additional subjects at time j (e.g., tumor BAM 124-2-jt corresponding to the alignment of sequence reads of nucleic acids isolated from solid tumor samples from subject 2). In some embodiments, BAM files are sorted and duplicate reads are marked for removal, resulting in a de-duplicated BAM file, for example, a tool such as SamBAMBA marks and filters duplicate alignments in the sorted BAM file.

[0236] Many of the embodiments described below in conjunction with Figures 4A-4E relate to analyses performed using sequencing data derived from cancer patient cfDNA, e.g., obtained from the patient's liquid biopsy sample. Generally, these embodiments are independent and therefore do not rely on any particular sequencing data generation method, e.g., sample preparation, sequencing, and / or data pre-processing methodology. However, in some embodiments, the methods described below include one or more features 204 of generating sequencing data, as illustrated in Figures 2A and 3.

[0237] The alignment file (e.g., BAM file 124), prepared as described above, is then passed to a feature extraction module 145, where the sequences are analyzed to identify genomic alterations (e.g., SNV / MNV, indels, genomic rearrangements, copy number variations, etc.) and / or determine various characteristics of the patient's cancer (e.g., MSI status, TMB, tumor ploidy, HRD status, tumor fraction, tumor purity, methylation patterns, etc.) (324). Many software packages for identifying genomic alterations are known in the art, such as freebayes, PolyBayse, samtools, GATK, pindel, SAMtools, Breakdancer, Cortex, Crest, Delly, Gridss, Hydra, Lumpy, Manta, and Socrates. For a review of many of these variant calling packages, see, e.g., Cameron, DL et al., Nat. Commun., 10(3240):1-11 (2019), the contents of which are incorporated by reference in their entirety for all purposes. Generally, these software packages identify variants in a sorted SAM or BAM file 124 relative to one or more reference sequence constructs 158. The software package then outputs a file, e.g., a raw VCF (variant call format), that lists the called variants (e.g., genomic features 131) and identifies positions relative to the reference sequence construct (e.g., where the sequence of the sample nucleic acid differs from the corresponding sequence in the reference construct). In some embodiments, the system 100 digests the contents of the native output file and inputs feature data 125 into the test patient data store 120. In other embodiments, the native output file serves as a record of these genomic features 131 in the test patient data store 120.

[0238] In general, the systems described herein can employ any combination of available variant calling software packages and internally developed variant identification algorithms. In some embodiments, the output of a particular algorithm of a variant calling software is further evaluated, for example, to improve variant identification. Thus, in some embodiments, system 100 employs an available variant calling software package to perform some or all of the functionality of one or more of the algorithms shown in feature extraction module 145.

[0239] In some embodiments, as illustrated in FIG. 1A , separate algorithms (or the same algorithm implemented with different parameters) are applied to identify variants specific to a patient's cancer genome and variants present in the subject's germline. In other embodiments, variants are identified randomly and later classified as either germline or somatic, e.g., based on sequencing data, population data, or a combination thereof. In some embodiments, variants are classified as germline variants and / or non-actionable variants when they are represented in a population above a threshold level determined using a population database, e.g., ExAC or gnomAD. For example, in some embodiments, variants represented in at least 1% of alleles in a population are annotated as germline and / or non-actionable. In other embodiments, variants represented in at least 2%, at least 3%, at least 4%, at least 5%, at least 7.5%, at least 10%, or more of the alleles in a population are annotated as germline and / or non-actionable. In some embodiments, sequencing data from a matched sample from a patient, e.g., a normal tissue sample, is used to annotate variants identified in a cancer sample from a subject, i.e., variants present in both the cancer sample and the normal sample represent variants that were present in the germline before the patient developed cancer and can be annotated as germline variants.

[0240] In various embodiments, the detected genetic variants and genetic features are analyzed as a form of quality control. For example, a pattern of detected genetic variants or features may indicate problems related to the sample, the sequencing procedure, and / or the bioinformatics pipeline (e.g., sample contamination, mislabeling of the sample, changes in reagents, changes in the sequencing procedure and / or the bioinformatics pipeline, etc.).

[0241] FIG. 4E illustrates an exemplary workflow for genomic feature identification (324). This particular workflow is merely an example of one possible set and arrangement of algorithms for feature extraction from sequencing data 124. Generally, any combination of modules and algorithms, for example, those of feature extraction module 145 illustrated in FIG. 1A, can be used in a bioinformatics pipeline, particularly a bioinformatics pipeline for analyzing liquid biopsy samples. For example, in some embodiments, an architecture useful in the methods and systems described herein includes at least one of the modules or variant calling algorithms shown in feature extraction module 145. In some embodiments, an architecture includes at least two, three, four, five, six, seven, eight, nine, ten, or more of the modules or variant calling algorithms shown in feature extraction module 145. Furthermore, in some embodiments, feature extraction modules and / or algorithms not illustrated in FIG. 1A are utilized in the methods and systems described herein.

[0242] Variant Identification In some embodiments, variant analysis of aligned sequence reads, e.g., in a SAM or BAM format, includes identification of single nucleotide variants (SNVs), multiple nucleotide variants (MNVs), indels (e.g., nucleotide additions and deletions), and / or genomic rearrangements (e.g., inversions, translocations, and gene fusions) using a variant identification module 146, e.g., including an SNV / MNV calling algorithm (e.g., SNV / MNV calling algorithm 147), an indel calling algorithm (e.g., indel calling algorithm 148), and / or one or more genomic rearrangement calling algorithms (e.g., genomic rearrangement calling algorithm 149). An overview of an exemplary method for variant identification is shown in FIG. 4E. Essentially, the module first identifies differences (e.g., SNVs / MNVs, indels, or genomic rearrangements) between the sequence of aligned sequence reads 124 and the reference sequence to which the sequence reads are aligned, and records the variants, e.g., in a variant call format (VCF) file. For example, software packages such as freebayes and pindel are used to call variants using the sorted BAM file and the reference BED file as input. For a review of variant calling packages, see, for example, Cameron, DL et al., Nat. Commun., 10(3240):1-11(2019). A raw VCF file (variant call format) is output, indicating positions where the nucleotide base in the sample is not the same as the nucleotide base at that position in the reference sequence construct.

[0243] In some embodiments, the raw VCF data are then normalized, for example, by parsing and left alignment, as illustrated in FIG. 4E. For example, software packages such as vcfbreakmulti and vt are used to normalize the multiple nucleotide polymorphism variants in the raw VCF file, outputting a variant-normalized VCF file. See, for example, E. Garrison, "Vcflib: A C++ library for parsing and manipulating VCF files," GitHub, available on the Internet at github.com / ekg / vcflib (2012), the contents of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, the normalization algorithm is included within the architecture of a broader variant identification software package.

[0244] An algorithm is then used to annotate the variants in the (e.g., normalized) VCF file, e.g., to determine the source of the variation, e.g., whether the variant originates from the subject's germline (e.g., germline variant), cancer tissue (e.g., somatic variant), sequencing error, or is of an undeterminable source. In some embodiments, the annotation algorithm is included within the architecture of a broader variant identification software package. However, in some embodiments, an external annotation algorithm is applied to the (e.g., normalized) VCF data obtained from a conventional variant identification software package. The choice of using a particular annotation algorithm is well within the skill of one in the art and, in some embodiments, is based on the data being annotated.

[0245] A variety of methods of variant identification are contemplated for use in this disclosure.

[0246] In some embodiments, SNV / INDEL detection is achieved using VarDict (github.com / AstraZeneca-NGS / VarDictJava). Both SNVs and INDELs are called, then sorted, de-duplicated, normalized, and annotated. Annotation uses SnpEff to add transcript information, 1000 genome minor allele frequencies, COSMIC reference names and counts, ExAC allele frequencies, and Kaviar population allele frequencies. Annotated variants are then classified as germline, somatic, or equivocal using a Bayesian model based on prior expectations informed by germline and cancer variant databases. In some embodiments, equivocal variants are treated as somatic for filtering and reporting purposes.

[0247] In some embodiments, genomic rearrangements (e.g., inversions, translocations, and gene fusions) are detected after demultiplexing by aligning tumor FASTQ files to a human reference genome using a local alignment algorithm such as BWA. In some embodiments, DNA reads are sorted and duplicates can be marked using software such as SAMBlaster. Discrepancies and split reads can be further identified and separated. These data can be read into software such as LUMPY for structural variant detection. In some embodiments, structural variations are grouped by type, recurrence, and presence, stored in a database, and displayed through a fusion viewer software tool. The fusion viewer software tool can reference a database such as Ensembl to determine the gene and proximal exons surrounding the breakpoint for any possible transcript generated across the breakpoint. The fusion viewer tool can then position the breakpoint 5' or 3' relative to the subsequent exon in the direction of transcription. For inversions, this orientation can be reversed for inverted genes. After locating the breakpoints, translated amino acid sequences can be generated for both genes in the chimeric protein, and a plot can be generated containing the remaining functional domains of each protein, returned from a database, e.g., Uniprot.

[0248] For example, in an exemplary embodiment, gene rearrangements are detected using the SpeedSeq analysis pipeline. Chiang et al., 2015, "SpeedSeq: ultra-fast personal genome analysis and interpretation," Nat Methods, (12), pg. 966. Briefly, FASTQ files are aligned to hg19 using BWA. Split reads that map to multiple locations and read pairs that map to discordant locations are identified and separated, and then utilized to detect gene rearrangements using LUMPY. Layer et al., 2014, "LUMPY: a probabilistic framework for structural variant discovery," Genome Biol, (15), pg. 84. Fusions can then be filtered according to the number of supporting reads.

[0249] In some embodiments, putative fusion variants supported by fewer than a minimum number of unique sequence reads are filtered. In some embodiments, the minimum number of unique sequence reads is 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, or 20 unique sequence reads.

[0250] Additional methods and embodiments for variant identification are possible, for example, as disclosed in PCT Patent Application No. PCT / US21 / 18622, entitled "METHODS AND SYSTEMS FOR A LIQUID BIOPSY ASSAY," filed February 18, 2021, which is incorporated by reference in its entirety.

[0251] Allelic fraction determination In some embodiments, analysis of aligned sequence reads, e.g., in SAM or BAM format, includes determining 133 a variant allele fraction for one or more of the variant alleles 132 identified as described above. In some embodiments, the variant allele fraction module 151 tallies the instances where each allele is represented by unique sequence reads encompassing the variant locus of interest and generates a count for each allele represented at that locus. In some embodiments, these tallies are used to determine the ratio of variant alleles, e.g., alleles other than the most common allele in a population of subjects for each locus, relative to a reference allele. This variant allele fraction 133 can be used in several places in the feature extraction 206 workflow. For example, in some embodiments, the variant allele fraction is used during annotation of identified variants, e.g., when determining whether an allele was germline or somatically derived. In other cases, the variant allele fraction is used in processes to estimate the tumor fraction of a liquid biopsy sample or the tumor purity of a solid tumor fraction. For example, the variant allele fraction of multiple somatic alleles can be used to estimate the percentage of sequence reads that originate from one copy of a cancer chromosome.Assuming 100% tumor purity and that each cancer cell carries one copy of a variant allele, the overall purity of the tumor can be estimated.Of course, this estimation can be further corrected based on other information extracted from sequencing data, such as copy number changes, tumor ploidy abnormalities, tumor heterozygosity, etc.

[0252] Methylation determination In some embodiments, where a nucleic acid sequencing library has been processed by bisulfite treatment or enzymatic methyl-cytosine conversion, as described above, analysis of aligned sequence reads, e.g., in SAM or BAM format, includes determining the methylation state 132 of one or more loci in the patient's genome. In some embodiments, methylated sequencing data aligns to the reference sequence construct 158 ​​differently than unmethylated sequencing, because unmethylated cytosines are converted to uracil and the resulting uracil is ultimately sequenced as thymine, while methylated cytosines are not converted and are sequenced as cytosines. Therefore, a different approach, such as seeding the alignment with a shorter region of identity or converting all cytosines in the sequencing data to thymine and then aligning the data to the reference sequence construct for both the plus and minus strands of the sequence construct, must be used to align these modified sequences to the reference sequence construct. For a review of these approaches, see Zhou Q. et al., BMC Bioinformatics, 20(47):1-11 (2019), the contents of which are incorporated herein by reference in their entirety for all purposes. Algorithms for calling methylated bases are known in the art. For example, Bismark can distinguish cytosine in CpG, CHG, and CHH contexts. Krueger F. and Andrews SR, Bioinformatics, 27(11):1571-71 (2011), the contents of which are incorporated herein by reference in their entirety for all purposes.

[0253] Copy number variation analysis: In some embodiments, analysis of aligned sequence reads, for example in SAM or BAM format, includes determining copy numbers 135 for one or more loci using a copy number variation analysis module 153. In some embodiments where both patient liquid biopsy and normal tissue samples are analyzed, the duplicate-depleted BAM files and VCFs generated from the variant calling pipeline are used to calculate read depth and heterozygous germline SNV variation among the sequencing reads for each sample. In contrast, in some embodiments where only liquid biopsy samples are analyzed, a comparison between tumor samples and a pool of process-matched normal controls is used. In some embodiments, copy number analysis involves applying a circular binomial segmentation algorithm and selecting segments with highly different log2 ratios between the cancer sample and its comparison standard (e.g., a matched normal pool or normal pool). In some embodiments, approximate integer copy numbers are assessed from a combination of coverage differences in segmented regions, and estimates of stromal admixture (e.g., tumor purity, or the portion of a sample that is cancerous versus non-cancerous, such as the tumor fraction of a liquid biopsy sample) are generated by analysis of heterozygous germline SNVs.

[0254] Additional methods and embodiments for copy number variation analysis are possible, for example, as disclosed in PCT Patent Application No. PCT / US21 / 18622, filed February 18, 2021, entitled "METHODS AND SYSTEMS FOR A LIQUID BIOPSY ASSAY," which is incorporated herein by reference in its entirety.

[0255] Microsatellite instability (MSI): In some embodiments, analysis of aligned sequence reads, for example in SAM or BAM format, includes analysis of the microsatellite instability status 137 of the cancer using a microsatellite instability analysis module 154. In some embodiments, an MSI classification algorithm classifies cancers into three categories: microsatellite instability-high (MSI-H), microsatellite stable (MSS), or microsatellite indeterminate (MSE). Microsatellite instability is a clinically actionable genomic indicator for cancer immunotherapy. In microsatellite instability-high (MSI-H) tumors, defects in DNA mismatch repair (MMR) can cause a hypermutated phenotype in which alterations accumulate in repetitive microsatellite regions of DNA. MSI detection is traditionally performed by subjecting tumor tissue ("solid biopsies") to clinical next-generation sequencing or specific assays such as MMR IHC or MSI PCR.

[0256] For example, microsatellite instability status can be assessed by determining the number of repeat units present at multiple microsatellite loci, e.g., 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, 750, 1000, 2500, 5000, or more loci. In some embodiments, to avoid using reads that do not completely cover the locus, only reads encompassing the microsatellite locus with a significant number of flanking nucleotides at both ends, e.g., at least 5, 10, 15, or more nucleotides flanking each end, are used for analysis. In some embodiments, given the high incidence of polymerase slippage during replication of these repetitive sequences, a minimum number of reads, e.g., at least 5, 10, 20, 30, 40, 50, or more reads, must meet this criterion to use a particular microsatellite locus to ensure accuracy of the determination.

[0257] In some embodiments, each locus is individually tested for instability, e.g., as measured by the change or variance in the number of nucleotide base repeats in the nucleotide sequence from the cancer relative to a normal sample or standard, using, e.g., a Kolmogorov-Smirnov test. A locus is considered unstable if, for example, p≦0.05. The proportion of unstable microsatellite loci can be fed into a logistic regression classifier trained on samples from various cancer types, particularly those with clinically determined MSI status, e.g., colorectal and endometrial cohorts. For MSI tests in which only liquid biopsy samples are analyzed, the mean and variance for the number of repeats can be calculated for each microsatellite locus. A vector containing the mean and variance data can be fed into a classifier (e.g., a support vector machine classification algorithm) trained to provide the probability that the patient is MSI-H, which can be compared to a threshold. In some embodiments, the threshold for calling a patient as MSI-H is at least 60% probability, or at least 65% probability, 70% probability, 75% probability, 80% probability, or more. In some embodiments, a baseline threshold may be established for calling a patient as MSS. In some embodiments, the baseline threshold is no more than 40%, or no more than 35% probability, 30% probability, 25% probability, 20% probability, or less. In some embodiments, a patient is identified as MSE when the output of the classifier falls within the range between the MSI-H threshold and the MSS threshold.

[0258] Other methods for determining a subject's MSI status are known in the art. For example, in some embodiments, microsatellite instability analysis module 154 employs the MSI assessment methods described in U.S. Provisional Patent Application No. 62 / 881,845, filed August 1, 2019, or U.S. Provisional Patent Application No. 62 / 931,600, filed November 6, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0259] Tumor mutation burden (TMB): In some embodiments, analysis of aligned sequence reads, e.g., in a SAM or BAM format, includes determining the mutational burden of the cancer (e.g., tumor mutational burden 136) using tumor mutational burden analysis module 155. Generally, tumor mutational burden is a measure of mutations in a cancer per unit of a patient's genome. For example, tumor mutational burden may be expressed as a measure of central tendency (e.g., average) of the number of somatic variants per million base pairs in the genome. In some embodiments, tumor mutational burden refers to only a set of possible mutations, e.g., one or more of SNVs, MNVs, indels, or genomic rearrangements. In some embodiments, tumor mutational burden refers to only a subset of one or more types of possible mutations, e.g., non-synonymous mutations, meaning mutations that change the amino acid sequence of an encoded protein. In other embodiments, for example, tumor mutational burden refers to the number of one or more types of mutations that occur in a protein-coding sequence, regardless of whether they change the amino acid sequence of the encoded protein.

[0260] By way of example, in some embodiments, tumor mutation burden (TMB) is calculated by dividing the number of mutations (e.g., all variants or nonsynonymous variants) identified in the sequencing data (e.g., represented in a VCF file) by the size (e.g., in megabases) of the capture probe panel used for targeted sequencing. In some embodiments, a variant is included in the tumor mutation burden calculation only if certain criteria are met. For example, in some embodiments, a threshold sequence coverage of the locus associated with the variant must be met before the variant is included in the calculation, e.g., at least 25x, 50x, 75x, 100x, 250x, 500x, or more. Similarly, in some embodiments, a minimum number of unique sequence reads encompassing the variant allele must be identified in the sequencing data, e.g., at least 4, 5, 6, 7, 8, 9, 10, or more unique sequence reads. In some embodiments, a threshold variant allele fraction threshold, e.g., at least 0.01%, 0.1%, 0.25%, 0.5%, 0.75%, 1%, 1.5%, 2%, 2.5%, 3%, 4%, 5%, or more, must be met before a variant is included in the calculation. In some embodiments, patient selection criteria may differ for different types of variants and / or different variants of the same type. For example, variants detected at mutational hotspots within the genome may face less stringent criteria than variants detected at more stable loci within the genome.

[0261] In some embodiments, analyzing the aligned sequence reads comprises determining a hematological tumor mutation burden (bTMB) of the cancer. In some embodiments, determining bTMB comprises performing a method comprising identifying multiple variants in a liquid biological sample (e.g., a liquid biopsy sample) of the subject. In some such embodiments, the liquid biological sample of the subject is blood or a sample derived therefrom. Identifying multiple variants can include any of the methods for variant identification disclosed herein (see, e.g., the sections entitled "Variant Identification," "Allele Fraction Determination," "Methylation Determination," "Copy Number Variation Analysis," "Microsatellite Instability (MSI)," and / or "Homologous Recombination Status (HRD)").

[0262] In some embodiments, the method for determining bTMB further comprises applying to the plurality of identified variants at least a first variant filter that removes one or more variant types from the plurality of identified variants.

[0263] In some embodiments, the at least first variant filter comprises a first germline variant filter that removes germline variants from the plurality of identified variants. In some embodiments, the at least first variant filter comprises a first synonymous variant filter that removes synonymous variants from the plurality of identified variants. Thus, in some implementations, the at least first variant filter is applied to the plurality of identified variants, thereby obtaining a plurality of filtered variants (e.g., after removal of germline variants and / or synonymous variants). For example, in some embodiments, the plurality of filtered variants comprises nonsynonymous SNVs, MNVs, INDELs, and / or translocation variants.

[0264] The method further includes normalizing the plurality of filtered variants based on a plurality of nucleotide sequences corresponding to the target nucleic acid in the liquid biological sample. For example, in some embodiments, the target nucleic acid in the liquid biological sample is obtained from a target panel used for liquid biopsy sequencing, and normalizing the plurality of filtered variants includes dividing the number of filtered variants (e.g., non-synonymous SNVs, MNVs, INDELs, and / or translocation variants) by the size of the target panel. In some embodiments, the size of the target panel is the number of target genomic regions (e.g., genes) encompassed by the target panel. In some embodiments, the size of the target panel is the number of base pairs spanned by the plurality of target genomic regions (e.g., genes) encompassed by the target panel. In some embodiments, the size of the target panel is the number of target genomic regions (e.g., genes) encompassed by one or more target panels and / or a measure of central tendency (e.g., mean, median, mode, etc.) of the number of base pairs spanned by the plurality of target genomic regions (e.g., genes) encompassed by one or more target panels.

[0265] In some embodiments, the plurality of filtered variants comprises one or more of non-synonymous SNVs, MNVs, INDELs, and / or translocation variants. In some embodiments, the plurality of filtered variants comprises synonymous somatic variants (e.g., synonymous SNVs, MNVs, INDELs, and / or translocation variants). In some embodiments, the plurality of filtered variants comprises germline variants. In some embodiments, the plurality of filtered variants comprises one or more non-synonymous variants selected from the group consisting of non-synonymous SNVs, MNVs, INDELs, and translocation variants, and one or more synonymous variants selected from the group consisting of synonymous SNVs, MNVs, INDELs, and / or translocation variants. In some embodiments, the plurality of filtered variants does not comprise germline variants. In some embodiments, the plurality of filtered variants does not comprise synonymous variants. In some embodiments, the plurality of filtered variants does not comprise one or more of synonymous SNVs, synonymous MNVs, synonymous INDELs, and / or synonymous translocation variants. In some embodiments, the plurality of filtered variants does not include one or more of non-synonymous SNVs, non-synonymous MNVs, non-synonymous INDELs, and / or non-synonymous translocation variants.

[0266] An exemplary method for calculating bTMB is described below in Example 10. Other methods for calculating tumor mutation burden and / or blood tumor mutation burden in liquid biopsy and / or solid tissue samples are known in the art, for example, Fenizia F. et al., Transl Lung Cancer Res., 7(6):668-77 (2018), and Georgiadis A et al., Clin. Cancer Res., 25(23):7024-34 (2019), the disclosures of which are incorporated herein by reference for all purposes.

[0267] Homologous recombination status (HRD): In some embodiments, for example, analysis of aligned sequence reads in SAM or BAM format includes analysis of whether the cancer is homologous recombination deficient (HRD status 137-3) using the homologous recombination pathway analysis module 157.

[0268] Homologous recombination (HR) is a normal, highly conserved DNA repair process that allows the exchange of genetic information between identical or closely related DNA molecules. It is most widely used by cells to accurately repair harmful breaks (e.g., lesions) that occur on both strands of DNA. DNA damage can arise from exogenous (external) sources, such as ultraviolet light, radiation, or chemical damage, or from endogenous (internal) sources, such as errors in DNA replication or other cellular processes that generate DNA damage. Double-strand breaks are a type of DNA damage. The use of poly(ADP-ribose) polymerase (PARP) inhibitors in patients with HRD impairs both pathways of DNA repair, leading to cell death (apoptosis). The efficacy of PARP inhibitors is improved not only in ovarian cancers that exhibit germline or somatic BRCA mutations, but also in cancers where HRD is caused by other underlying etiologies.

[0269] In some embodiments, HRD status can be determined by inputting features correlated with HRD status into a classifier trained to distinguish between cancers with homologous recombination pathway defects and cancers without homologous recombination pathway defects. For example, in some embodiments, the features include one or more of: (i) the heterozygous state of a first plurality of DNA damage repair genes in the genome of the subject's cancer tissue; (ii) a measure of loss of heterozygosity across the genome of the subject's cancer tissue; (iii) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's cancer tissue; and (iv) a measure of variant alleles detected in a second plurality of DNA damage repair genes in the genome of the subject's non-cancerous tissue. In some embodiments, all four of the features described above are used as features in the HRD classifier. More detailed information about HRD classifiers using these and other features is described in U.S. Patent Application No. 16 / 789,363, filed February 12, 2020, the contents of which are incorporated herein by reference in their entirety for all purposes.

[0270] Circulating tumor fraction: In some embodiments, analysis of aligned sequence reads, for example in SAM or BAM format, includes estimating the circulating tumor fraction of a liquid biopsy sample. The tumor fraction or circulating tumor fraction is the fraction of cell-free nucleic acid molecules in a sample derived from a subject's cancer tissue but not from non-cancerous tissue (e.g., germline or hematopoietic tissue). Several open-source analysis packages have modules for calculating tumor fraction from solid tumor samples. For example, PureCN (Riester, M., et al., Source Code Biol Med, 11:13 (2016)) is designed to estimate tumor purity from targeted short-read sequencing data of solid tumor samples. Similarly, FACETS (Shen R, Seshan VE, Nucleic Acids Res., 44(16):e131 (2016)) is designed to estimate tumor fraction from sequencing data of solid tumor samples. However, estimating tumor fraction from liquid biopsy samples is generally more difficult due to the lower tumor fraction relative to solid tumor samples and the typical small size of the target panels used in liquid biopsy sequencing. Indeed, packages such as PureCN and FACETS do not perform well with sequencing data generated using small target panels at low tumor fractions.

[0271] Various methods can be used to estimate circulating tumor fraction. In some embodiments,...

Claims

1. A liquid biopsy sequencing method comprising contacting a plurality of nucleic acids in a composition with a probe set under hybridizing conditions, wherein the plurality of nucleic acids include cell-free nucleic acids derived from a first liquid biopsy of a first object, or nucleic acids prepared therefrom. The aforementioned probe set, A first set of polynucleotide probes that collectively target a first set of multiple genomic regions with an average coverage of 0.75 to 1.25 times, wherein the first set of polynucleotide probes comprises a first set of multiple polynucleotide probe species. Each of the polynucleotide probe species in the first plurality of polynucleotide probe species targets each of the genomic regions in the first plurality of genomic regions, Each of the first plurality of polynucleotide probe species is initially present in the composition at a corresponding molar concentration. The corresponding molar concentrations of each of the first plurality of polynucleotide probe species are collectively averaged to form a first average molar concentration. The first set of polynucleotide probes comprises a plurality of genomic regions, each containing at least a portion of the coding sequences of at least 25 genes listed in Figures 58A to 58BF. A second set of polynucleotide probes that collectively target a second set of multiple genomic regions with an average coverage of 0.75 to 1.25 times, wherein the second set of polynucleotide probes comprises a second set of multiple polynucleotide probe species, Each of the polynucleotide probe species in the second plurality of polynucleotide probe species targets each of the genomic regions in the second plurality of genomic regions, Each of the polynucleotide probe species in the second plurality of polynucleotide probe species is initially present in the composition at the corresponding molar concentration. The corresponding molar concentrations of each of the polynucleotide probe species in the second plurality of polynucleotide probe species are collectively averaged to obtain a second average molar concentration. The second plurality of genomic regions include at least a portion of the coding sequences of at least 25 genes selected from the genes listed in Figures 57A to 57Y, The set includes a second set of polynucleotide probes, wherein the second average molar concentration is 5 to 8 times higher than the first average molar concentration. The method described above is The recovery of the plurality of nucleic acids, wherein each nucleic acid in the plurality of nucleic acids hybridizes with each nucleic acid probe in the first set of the plurality of nucleic acid probes or the second set of the plurality of nucleic acid probes. A method comprising sequencing the plurality of nucleic acids after the recovery to obtain a plurality of sequence reads of the plurality of nucleic acids.

2. The probe set further comprises a third set of polynucleotide probes that collectively target the BRCA1 and BRCA2 genes with an average coverage of at least 1.5 times, wherein the third set of polynucleotide probes comprises a third plurality of polynucleotide probe species. Each of the polynucleotide probe species in the third plurality of polynucleotide probe species targets BRCA1 or BRCA2, Each of the three polynucleotide probe species in the above-mentioned third plurality of polynucleotide probe species is initially present in the composition at the corresponding molar concentration. The corresponding molar concentrations of each of the polynucleotide probe species in the third plurality of polynucleotide probe species are collectively averaged to form a third average molar concentration. The method according to claim 1, wherein the third average molar concentration is 5 to 8 times higher than the first average concentration.

3. The method according to claim 1 or 2, wherein each of the third polynucleotide probe species is initially present in the composition in an amount of 1.5 fmol to 3 fmol or 3 fmol to 5 fmol.

4. The method according to claim 1 or 2, wherein the probe set further comprises a fourth set of polynucleotide probes that collectively target a plurality of viral sequences, the plurality of viral sequences comprising sequences derived from the genome of at least human papillomavirus (HPV) types 16, 18, 33, human gamma herpesvirus 4 (HHV4), and Merkel cell polyomavirus isolate R17b, each of the polynucleotide probe species in the fourth plurality of polynucleotide probe species initially present in the composition at a corresponding molar concentration, and the corresponding molar concentrations of each of the polynucleotide probe species in the fourth plurality of polynucleotide probe species, collectively averaged, result in a fourth average molar concentration equal to the first average concentration.

5. The method according to claim 1 or 2, wherein the first plurality of genomic regions include at least a portion of the coding sequences of at least 50 genes, at least 100 genes, at least 200 genes, at least 300 genes, or at least 400 genes selected from the genes listed in Figures 58A to 58BF.

6. The method according to claim 1 or 2, wherein the first plurality of genomic regions include at least a portion of the introns of the genes UGT1A1, EWSR1, and TMPRSS2, respectively.

7. The method according to claim 1 or 2, wherein the first plurality of probe types are at least 100 probe types.

8. The method according to claim 1 or 2, wherein each of the first plurality of polynucleotide probe species is initially present in the composition in an amount of 200 amol to 600 amol or 500 amol to 1 fmol.

9. The method according to claim 1 or 2, wherein the second plurality of genomic regions include at least a portion of the introns of at least four, at least five, or at least ten genes selected from a list of genes consisting of ABL1, ALK, BRAF, EGFR, ERBB3, FGFR1, FGFR2, FGFR3, FLT3, MYD88, NTRK1, NTRK2, NTRK3, RET, and ROS1.

10. The method according to claim 1 or 2, wherein each of the second plurality of polynucleotide probe species is initially present in the composition in an amount of 1.5 fmol to 3 fmol or 3 fmol to 5 fmol.

11. The first portion of each polynucleotide probe species in the first plurality of polynucleotide probe species contains biotin, The first portion of each polynucleotide probe species in the second plurality of polynucleotide probe species contains biotin, The method according to claim 1 or 2, wherein recovering the plurality of nucleic acids includes binding the non-nucleotide capture portions in the first plurality of polynucleotide probe species and the second plurality of polynucleotide probe species to streptavidin or avidin.

12. The method according to claim 1 or 2, wherein each of the first plurality of polynucleotide probe species comprises a nucleic acid sequence of 50 to 250 nucleotides that targets each of the genomic regions in the first plurality of genomic regions, and each of the second plurality of polynucleotide probe species comprises a nucleic acid sequence of 50 to 250 nucleotides that targets each of the genomic regions in the second plurality of genomic regions.

13. The method according to claim 1 or 2, wherein the first plurality of genomic regions collectively targeted by the first set of polynucleotide probes encompass 1 megabase pair (Mbp) to 5 megabase pairs (Mbp), and the second plurality of genomic regions collectively targeted by the second set of polynucleotide probes encompass 200 kilobase pairs (Kbp) to 800 kilobase pairs (Kbp).

14. The method according to claim 1 or 2, further comprising associating the plurality of sequence reads with each of the cancer states in a plurality of cancer states, thereby characterizing the target cancer state.

15. The method according to claim 14, wherein associating the plurality of sequence reads with the respective cancer state includes identifying one or more single nucleotide variants (SNVs) or multiple nucleotide variants (MNVs) or indels in the plurality of sequence reads associated with the respective cancer state, wherein the one or more SNVs or MNVs or indels include SNVs or MNVs or indels in a genomic region targeted by the first set of polynucleotide probes, or the one or more SNVs or MNVs or indels include SNVs or MNVs or indels in a genomic region targeted by the second set of polynucleotide probes.