Systems and methods for tumor fraction estimation using background error rates in DNA sequencing data
Patent Information
- Application Number
- PCT/US2025/018145
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-04
- Filing Date
- 2025-03-03
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional methods for estimating tumor fractions in circulating tumor DNA (ctDNA) samples struggle to distinguish genuine tumor-derived signals from sequencing noise, particularly in samples with low tumor fractions, due to variable error rates across genomic locations and reliance on uniform background error assumptions.
A novel approach using a combined targeted sequencing panel that includes genetic variants from multiple subjects, leveraging site-specific background error rates to establish a null distribution threshold for accurate tumor fraction estimation, and identifying and addressing noisy genomic locations to enhance accuracy.
Enhances the accuracy of tumor fraction estimation by distinguishing true tumor-derived signals from sequencing noise, allowing for the detection of small amounts of cancer and improving the reliability of clinical decision-making.
Smart Images

Figure US2025018145_02102025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR TUMOR FRACTION ESTIMATION USING BACKGROUND ERROR RATES IN DNA SEQUENCING DATACROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No. 63 / 561,045, filed on March 4, 2024, which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to the field of cancer diagnostics and, more specifically, to systems and methods for estimating tumor fractions in biopsy samples using DNA sequencing data.BACKGROUND
[0003] Estimation of tumor fraction is used in various applications, including cancer screening, treatment monitoring, and minimal residual disease (MRD) assessment. Accuracy in tumor fraction estimation may thus impact the accuracy of these applications. More particularly, the ability to reliably assess the proportion of circulating tumor DNA (ctDNA) derived from tumor cells within a patient’s plasma may have significant implications for both patient care and clinical decision-making. However, samples containing low levels of ctDNA may be obscured by noise introduced during the sequencing process, thereby leading to uncertainty in determining whether tumor material is actually observed or whether the observations are just sequencing artifacts. Existing approaches struggle to distinguish genuine tumor-derived signals from potential site-specific base errors.
[0004] The present disclosure is accordingly directed to systems and methods that may leverage site-specific background error rates derived from sequencing data of nonmutated individuals to enhance the accuracy of tumor fraction estimation. The backgrounddescription provided herein is for the purpose of generally presenting context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE
[0005] According to certain aspects of the disclosure, systems and methods are described for accurate estimation of tumor fractions in ctDNA samples.
[0006] In one aspect, a computer-implemented method for estimating tumor fraction in a biological sample is provided. The computer-implemented method may contain steps including: receiving, at a computing device, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples and wherein the one or more genetic variants between each of the plurality of biological samples are different and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; and background sequencing data covering all other of the one or more genetic variants for all other of the plurality of biological samples; calculating, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establishing, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.
[0007] In another aspect, a system for estimating tumor fraction in a biological sample is provided. The system may include: one or more processors; one or more computerreadable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples and wherein the one or more genetic variants between each of the plurality of biological samples are different and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; and background sequencing data covering all other of the one or more genetic variants for all other of the plurality of biological samples; calculate, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establish, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.
[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The non-transitory computer-readable medium stores computer-executable instructions which, when executed by a system, cause the system to perform operations comprising: receiving, at a computing device, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples and wherein the one or more genetic variants between each of the plurality of biological samples are different and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; and background sequencing data covering all other of the one or more geneticvariants for all other of the plurality of biological samples; calculating, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establishing, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.
[0009] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description that follows, and in part will be apparent from the description, or may be learned by practice of the disclosed embodiments. The objects and advantages of the disclosed embodiments will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and together with the description, serve to explain the principles of the disclosure.
[0012] FIG. 1 A depicts an exemplary computer system for executing the techniques described herein.
[0013] FIG. IB depicts an exemplary software platform for executing the techniques described herein.
[0014] FIG. 2 depicts an exemplary workflow for generating a tumor fraction estimate for a biological sample, according to one or more embodiments of the present disclosure.
[0015] FIG. 3 depicts an exemplary illustration of the makeup of a combined targeted sequencing panel, according to one or more embodiments of the present disclosure.
[0016] FIGS. 4A and 4B depict graphs illustrating an exemplary null distribution threshold, according to one or more embodiments of the present disclosure.
[0017] FIGS. 5 A and 5B depict graphs illustrating tumor fraction estimate data, according to one or more embodiments of the present disclosure.
[0018] FIG. 6 depicts information associated with a contaminated assay batch, according to one or more embodiments of the present disclosure.
[0019] FIG. 7 depicts a flowchart of an exemplary method of verifying the validity of a tumor fraction estimate for a biological sample, according to one or more embodiments of the present disclosure.
[0020] FIG. 8 depicts an example computing system, according to one or more aspects of the present disclosure.DETAILED DESCRIPTION OF EMBODIMENTS
[0021] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.
[0022] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” Theterms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as “about,” “approximately,” “substantially,” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value. In addition, the term “between” used in describing ranges of values is intended to include the minimum and maximum values described herein. The use of the term “or” in the claims and specification is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.
[0023] As used herein, the term “user” generally encompasses any person or entity, such as a researcher and / or a care provider (e.g., a doctor, etc.), that may desire information, resolution of an issue, or engage in any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The term “electronic application” or “application” may be used interchangeably with other terms like “program,” or the like, and generally encompasses software that is configured to interact with, modify, override, supplement, or operate in conjunction with other software.
[0024] Tumor fraction refers to the proportion of ctDNA within a sample of circulating free (cfDNA) extracted from a patient’s bloodstream. CtDNA consists of DNA fragments that are shed into the bloodstream by cancer cells and which carry genetic mutations and alterations that are specific to the tumor cells. Estimating the tumor fraction is important in various aspects of cancer research and clinical practice. For instance, this metricprovides insights into the relative abundance of ctDNA in the overall cfDNA pool, which may be indicative of the presence, extent, and progression of tumors. Additionally, accurate tumor fraction identification is key in the post diagnostic space in which an oncologist may look to see whether there is any minimal residual disease, whether a subject is cured, or how well a subject is responding to treatment (and thus whether a subject should start, stop, or adjust treatments). Having a clear assessment of whether there is any tumor material or not, even low levels, is important to that decision-making process. Embodiments of the present disclosure may allow for achievement of a lower limit of detection and thus the accurate detection of small amounts of cancer.
[0025] Conventional techniques for estimating tumor fraction encounter various challenges that may undermine the reliability of results. For instance, one challenge is that the tumor fraction within ctDNA samples can be extremely low, making it difficult to distinguish true tumor material from background noise. Conventional techniques may lack the sensitivity to accurately estimate tumor fractions at such low levels, potentially missing critical information for diagnosis and treatment decisions. Another challenge is that traditional approaches predominately rely on direct observation of mutations in ctDNA. However, the sequencing process may introduce noise, including errors resulting from chemical changes and / or technical limitations of sequencing platforms. These errors may be mistaken for the detection of genuine tumor-derived signals, leading to false-positive or false-negative results. Yet another challenge may be that conventional methods often treat sequencing errors as uniform background across all genomic locations. However, the error rates may vary substantially between each genomic location depending on sequence context, neighboring bases, and chemical properties. This variability in error rates remains unaccounted for, further compromising the accuracy of the tumor fraction estimation.
[0026] Accordingly, the present disclosure addresses the foregoing challenges by providing a novel approach that leverages sequencing data obtained from a single combined targeted panel containing genetic variants from a plurality of subjects. Specifically, the sequencing data for each subject included on the panel may contain sequencing information about the genomic positions of their genetic variants, along with background sequencing coverage of the genomic positions of different genetic variants for all other subjects on the panel. The collective background sequencing data may be utilized to estimate error rates specific to each genomic position and, ultimately, to establish a null distribution threshold. This threshold serves as a reference for assessing whether observed tumor fractions in test samples are significantly above the expected noise level and can confidently be considered to provide an indication that tumor material is present.
[0027] In an aspect, the background sequencing data may additionally be utilized to identify certain genomic locations that exhibit unexpectedly high error rates. These “noisy” locations may arise due to various types of noise and may introduce false signals into the tumor fraction estimation, thereby compromising the accuracy of the results. Once identified, these noisy locations may be flagged and either removed or treated with a noise-reducing technique, as will be described further below.
[0028] The approach described herein involves a technological process for accurate tumor fraction estimation using sequencing data. It combines targeted sequencing panel design, somatic variant identification, deep sequencing of cfDNA, and utilization of background error rates to calculate and compare tumor fraction estimates. Specifically, the described embodiments improve upon conventional tumor fraction estimation methods by addressing specific challenges in the field, such as distinguishing true tumor-derived signals from sequencing noise. By introducing the novel concept of site-specific background error rates derived from non-mutated individuals’ sequencing data, the approach enhances theaccuracy of tumor fraction estimation, offering a practical solution that overcomes existing limitations.
[0029] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0030] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments,” or “in one aspect” or “in some aspects” as used herein does not necessarily refer to the same embodiment or aspect, and the phrase “in another embodiment” or “in another aspect” as used herein does not necessarily refer to a different embodiment or aspect. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.
[0031] FIG. 1 A depicts an exemplary system for estimating tumor fraction in samples by using patient-specific background data obtained from a combined targetedsequencing panel. Exemplary system 100 includes a data collection component 10, a database 20, and device data intelligence component 30, operably connected to each other via network 40. Alternatively, or additionally, one or more of the components may be connected with another component locally without reliance on network connection; e.g., through a wired connection. In many aspects described herein, sequencing data of cell-free nucleic acids are used to illustrate the concepts. However, one of skill in the art would understand that the current method may be applied to sequencing data of other materials as well.
[0032] As disclosed herein, data collection component 10 may include a device or machine with which sequencing data may be generated. In some embodiments, data collection component 10 may include a sequencing device or a facility that uses a sequencing device to generate nucleic acid sequence data of biological samples. In some aspects, data collection 10 may be a database that receives sequencing information generated from one or more sequencing devices. Any suitable liquid or solid biological samples may be used. In some embodiments, a biological sample may be cell-based; for example, one or more types of tissue. In some embodiments, a biological sample is a sample that includes cell-free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample (e.g., a serum sample, a plasma sample, a whole blood sample), a urine sample, a saliva sample, a tissue sample, a bone marrow sample, etc. Further, although sequencing of DNA from these samples is discussed herein, RNA from these samples may alternatively or additionally be sequenced.
[0033] Examples of sequencing data may include, but are not limited to, sequence read data of targeted genomic locations, partial or whole genome sequencing data of the genome represented by nucleic acid fragments in cell-free or cell-based samples, partial or whole genome sequencing data including one or more types of epigenetic modifications (e.g., methylation), or combinations thereof.
[0034] Data acquired by the data collection component 10 may be transferred to database 20 via network 40. In some embodiments, the collected data may be analyzed by data intelligence component 30, via a local or network connection. FIG. IB depicts exemplary functional modules that may be implemented to perform tasks of data intelligence component 30.
[0035] FIG. IB depicts an exemplary computer system 110 for processing sequencing data (e.g., from a combined targeted sequencing panel) and generating distribution thresholds against which the validity of tumor fraction estimates may be assessed. Exemplary system 110 achieves such functionalities by implementing, on one or more computer devices, user input and output (I / O) module 120, memory or database 130, data processing module 140, data analysis module 150, classification module 160, network communication module 170, and any other functional modules that may be needed for carrying out a particular task (e.g., an error correction or compensation module, a data compression module, etc.). As disclosed herein, user I / O module 120 may further include an input sub-module, such as a keyboard, and an output sub-module, such as a display (e.g., a printer, a monitor, or a touchpad). In some embodiments, all functionalities may be performed by one computer system. In some embodiments, the functionalities are performed by more than one computer system.
[0036] Also disclosed herein, a particular task may be performed by implementing one or more functional modules. In particular, each of the enumerated modules itself may, in turn, include multiple sub-modules. For example, data processing module 140 may include a sub-module for data quality evaluation (e.g., for discarding very short sequence reads or sequence reads including obvious errors), a sub-module for normalizing numbers of sequence reads that align to different regions of a reference genome, a sub-module to compensate / correct GC biases, and etc.
[0037] In some embodiments, a user may use I / O module 120 to manipulate data that is available either on a local device or can be obtained via a network connection from a remote service device or another user device. For example, I / O module 120 may allow a user, e.g., via a keyboard, a mouse, or a touchpad, to perform data analysis via a graphical user interface (GUI). In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before a user is granted access to the data being requested. In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, a user may request via user I / O module 120 input data while an existing data processing session is in process. A user may do so by selecting a menu option or type in a command discretely without interrupting the existing process. As disclosed herein, a user may use any type of input to direct and control data processing and analysis via I / O module 120.
[0038] In some embodiments, system 110 further comprises a memory or database 130. In some embodiments, database 130 comprises a local database that may be accessed via user VO module 120. In some embodiments, database 130 comprises a remote database that may be accessed by user I / O module 120 via network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real-time from internet searches. In some embodiments, database 130 may send data to and receive data from one or more of the other functional modules, including, but not limited to, a data collection module (not shown), data processing module 140, data analysis module 150, classification module 160, network communication module 170, and etc.
[0039] In some embodiments, database 130 may be a database local to the other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by the other functional modules via wired or wireless network connection (e.g.,via network communication module 170). In some embodiments, database 130 may include a local portion and a remote portion.
[0040] In some embodiments, system 110 comprises a data processing module 140. Data processing module 140 may receive the real-time data, from I / O module 120 or database 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, normalization of counts of sequence reads, correction of GC bias, etc. In some embodiments, data processing module 140 may be configured to estimate tumor fraction in a biological sample. For example, sequencing data may be received from a combined targeted sequencing panel that is designed to contain a plurality of biological samples, each of which is associated with a different subject. The sequencing data for the samples of each subject may provide sequencing coverage (“background data”) for the genomic locations associated with the genetic variants of the other subjects included on the panel. Data processing module 140 may utilize the background sequencing data to calculate site-specific error rates for each subject, which may subsequently be utilized to establish one or more thresholds against which tumor fraction estimates for each sample may be compared to assess the presence or absence of tumor material. In various embodiments, data processing module 140 may additionally utilize the background sequencing data to identify and address particular noise-generating sites that may negatively impact result accuracy.
[0041] In some embodiments, system 110 comprises a data analysis module 150. In some embodiments, data analysis module 150 includes identifying and treating systematic errors in sequencing data, as described in connection with data processing module 140.
[0042] In some embodiments, system 110 comprises a classification module 160, which analyzes data from a test sample from a test subject whose status with respect to a medical condition is unknown and subsequently classifies the unknown test sample from thetest subject based on the likelihood of the subject fitting into a particular category. In some embodiments, the one or more parameters may include a binomial probability score that is calculated based on logistic regression analysis. As disclosed herein, the binomial probability score may correspond to the likelihood of a subject having a certain medical condition, such as cancer. For example, a score of over a predefined threshold may indicate that the subject associated with a test sample is more likely to have cancer than not have cancer. In some embodiments, the one or more parameters may include a sequencing data distribution pattern correlating with the presence of cancer. A subject associated with a test sample having sequencing data with a pattern resembling the cancer pattern may be diagnosed as having cancer. In some embodiments, a sequencing data distribution pattern may be identified in connection with a specific type of cancer, thus allowing a test sample to be classified as indicative of a certain cancer type.
[0043] As disclosed herein, network communication module 170 may be used to facilitate communications between a user device, one or more databases, and any other suitable system or device through a wired or wireless network connection. Any communication protocol / device may be used, including without limitation a modem, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication facilities, etc.), a near-field communication (NFC), a Zigbee communication, a radio frequency (RF) or radio-frequency identification (RFID) communication, a PLC protocol, a 3G / 4G / 5G / LTE based communication, and / or the like. For example, a user device having a user interface platform for processing / analyzing tumor fraction data may communicate with another user device with the same platform, a regular user device without the same platform (e.g., a regularsmartphone), a remote server, a physical device of a remote loT local network, a wearable device, a user device communicably connected to a remote server, and etc.
[0044] The functional modules described herein are provided by way of example. It will be understood that different functional modules may be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement a certain utility.
[0045] Referring now to FIG. 2, an exemplary workflow 200 is provided for generating tumor fraction estimates.
[0046] At step 205, DNA from a tumor biopsy taken from a subject may be sequenced. More particularly, genetic material of a tumor biopsy may be analyzed to identify genetic mutations that are specific to the tumor cells. To facilitate this analysis, at the outset, a small piece of tissue may be extracted from the tumor site. This sample may contain a mixture of cells, including tumor cells, as well as any surrounding non-tumor cells (e.g., normal healthy cells). After the biopsy sample is obtained, DNA may be extracted from the collected tissue. The extracted DNA contains the genetic information of the cells present in the biopsy, including any mutations specific to the tumor. The extracted DNA may then be subjected to a sequencing process (e.g., whole genome bisulfite sequencing, whole genome sequencing, etc.) to produce sequencing data. This data is then analyzed to identify genetic mutations that are unique to the tumor cells. These tumor-specific mutations are known as somatic variants, which are changes in the DNA sequence that have occurred specifically in the tumor cells and are not present in the individual’s normal, healthy cells.
[0047] At step 210, matched, healthy tissue sequencing may be performed to enable distinction between germline and somatic mutations. More particularly, by comparing the genetic information of the tumor tissue with that of healthy tissue from the same individual, genetic changes unique to the tumor cells may be distinguished. To facilitate this comparison,in an aspect, a matched, healthy tissue sample may be collected from a non-tumor area of the subject’s body. This sample typically consists of normal tissue cells and does not contain any tumor cells. In an aspect, the collected sample may be a cfDNA sample, a white blood cell sample, any other healthy collected sample, etc. Therefore, any variants found in this sample would likely be germline variants. The sample collection process may be facilitated in a manner similar to the tumor biopsy collection, as described in step 205. DNA may then be extracted from the collected matched, healthy tissue samples and sequenced (e.g., using the same or different sequencing process used for sequencing the tumor biopsy sample at step 205, e.g., whole-genome bisulfite sequencing, whole genome sequencing, etc.) to produce a dataset of genetic sequences for the germline variants.
[0048] At step 215, the results from the tumor biopsy sequencing may be compared against those of the matched, healthy tissue to “call” or identify the specific genetic mutations that are unique to the tumor cells. Prior to facilitating the variant call, the raw sequencing data from both the tumor biopsy and the matched, healthy tissue may undergo preprocessing. Preprocessing may involve tasks like removing low-quality reads, trimming adaptors, and / or correcting for potential sequencing errors. The DNA fragments may then be aligned to a reference genome, which serves as a standardized template against which the sequencing reads are aligned. After alignment, the sequence data from the tumor biopsy and the matched, healthy tissue are compared base-by-base to identify genetic differences, e.g., single nucleotide changes, insertions, and deletions. These differences may represent the somatic variations between the tumor and the healthy tissue.
[0049] At step 220, the set of somatic variant calls obtained from the previous steps may then be utilized to design a combined targeted sequencing panel. In a traditional genetic sequencing assay, a separate sequencing panel may be run for each individual sample taken from a subject, and sequencing data for a given sample may be analyzed individually. Theapproach described herein, however, utilizes a single combined panel that includes samples with genetic variants from multiple subjects. Specifically, a sample from a first subject may be sequenced and analyzed in combination with a plurality of samples from other subjects. These collective samples from different individuals may all be included in the same panel and may be selected so that each sample contains genetic variants that are different than the genetic variants of the other samples on the panel. Therefore, when the panel is run, sequencing data is generated for each sample that covers its variant positions, as well as the variant positions of all the other samples included on the panel. Accordingly, the sequencing data for each sample may act as “background data” for the variant positions for the other samples in the panel. This background data may thereafter be utilized for accurate tumor fraction estimation and confidence assessment, as further described herein.
[0050] At step 225, cfDNA data obtained from samples of the subjects (e.g., plasma samples, other sample types, etc.) included on the combined panel may be utilized in the targeted sequencing process. In an aspect, each of the samples may contain cfDNA, which may include DNA fragments shed by both healthy cells and tumor cells in the bloodstream. The combined targeted sequencing of the collective cfDNA data included in the panel may enable simultaneous sequencing data generation directed to the genomic regions of interest for each panel subject, thereby leading to the identification of fragments carrying the known tumor mutations. By analyzing the results of the deep sequencing, the number of DNA fragments in the plasma that exhibit tumor-specific mutations may be quantified. This count may be compared against the total number of fragments sequenced and an approximation, conducted at step 230, of the tumor fraction may be calculated, e.g., for each subject on the panel.
[0051] Referring now to Table 1, below, various metrics associated with samples sequenced with the targeted panel are provided. For each of the samples in Table 1, deepsequencing was conducted (e.g., to approximately 1800X collapsed depth) on up to 500 putative somatic variants per participant. The “totalCount” column is representative of the total number of DNA fragments that were sequenced across all of the subject’s variants, and the “alternateAlleleCounf ’ column is representative of the number of those DNA fragments that contain a tumor specific mutation. An additional metric associated with the dataset is the expected noise rate in the sequencing assay. More particularly, in this context, noise refers to errors that may arise during the sequencing process, such as chemical changes or artifacts introduced by sequencing machines. The expected noise rate may provide a baseline expectation for the frequency of sequencing errors. For instance, the expected error rate in the assay represented by Table 1 is approximately 1 in 1 million. By comparing the observed frequency of tumor-specific mutations with the expected noise rate, it may be possible to arrive at some bright line binary observations.Table 1
[0052] In some cases, it may be clear from the observed mutation counts and the expected noise rate that tumor material is either present or absent. More particularly, for the data set associated with Table 1, an expectation may be realized that if around 3 million fragments are sequenced, around 3 errors might be expected to be found due to noise. Accordingly, in situations represented by the ccga_8500 sample (having 498 tumor-specificmutations detected out of approximately 2.7 million total sequenced fragments) and the ccga_3409-2 sample (having 672 tumor-specific mutations detected out of approximately 1.5 million total sequenced fragments), there can be high confidence tumor material is present because the detected amounts are significantly higher than the expected error rates may account for (e.g., approximately 3 in the ccga_8500 sample and approximately 1.5 in the ccga_3409-2 sample). In a similar vein, there can be high confidence that there is no tumor material in the ccga_2515_high sample because the total detected mutations are 0, which is less than even the expected noise rate (i.e., ~1).
[0053] However, contrary to the foregoing, there are instances in which it may be challenging to definitively determine whether the observed genetic variations are real tumorspecific mutations or if they might be the result of background noise or sequencing errors. More particularly, there are instances in which the observed counts fall within a range that could potentially be attributed to either tumor material or noise. For instance, the total observed mutation count in the ccga_5314 sample was 2, which is close to, but more than, the expected noise rate for the sample. In cases such as the foregoing, a quantitative method may be used to make a more informed decision about whether tumor material is confidently detected.
[0054] Referring now to FIG. 3, a graphic 300 illustrating a combined targeting sequencing panel (“combined panel”) is provided. As alluded to above, the combined panel may be designed to contain samples with variants from a plurality of subjects. It is important to note that although samples with variants from four subjects are illustrated in FIG. 3 for ease of depiction, such a designation is not limiting, and the panel may be configured to accommodate more or fewer samples from different subjects. Specifically, the maximum number of sample variants that may be included in the panel may be limited by the physical constraints of the panel and / or the capabilities of the sequencing machinery. For example,about 5 to about 30 samples from different individuals may be included on a panel, about 5 to about 20 samples, or about 10 to about 20 samples, etc., may be included on a panel.
[0055] In an aspect, the selection of the subject samples to include in the combined panel may be dependent on one or more factors. For instance, in one aspect, subjects that have many variants may be paired with subjects that have fewer variants in an attempt to create a series of panels with average total sizes that are similar, which may produce a logistical benefit when physically processing the samples (e.g., in the wet lab). In another aspect, subjects may be paired together based on characteristics of the individuals from whom the samples were obtained, such as age, sex, etc., if those pairings were thought to be beneficial to provide meaningful indications of background noise at various genomic positions. At a minimum, however, samples from subjects that share the same variants should not be included in the same panel. If, for example, two or more samples included the same variants, then the sequencing data for those variants may be intermingled, which may compromise the accuracy of the results.
[0056] As mentioned, the combined panel may contain a comprehensive set of samples with genetic variations from the selected group of sample subjects. Since the panel includes samples with genetic variants from multiple subjects, the sequencing data of a sample from one subject may cover the variants of other samples from other individuals. For instance, known variants from Subjects 1 through 4 may be included on panel A 305. For instance, a sample from Subject 1 may be associated with 500 variants, a sample from Subject 2 may be associated with 200 variants, a sample from Subject 3 may be associated with 500 variants, and a sample from Subject 4 may be associated with 100 variants. When the panel is run, sequencing data may be obtained for each of the samples from the four subjects. As a representative example, a breakdown of Subject l’s sequencing data is provided at 310. Specifically, Subject l’s sequencing data may contain data that covers thegenomic variant regions specific to Subject 1 (e.g., at 315), which may be used to estimate the tumor fraction for Subject 1, and also contains “background data” for the genomic variant regions for Subjects 2, 3, and 4 at regions 320, 325, and 330, respectively. Sequencing data from the reference samples 2 through 4 may provide information on the amount of background noise (e.g., the sequencing error rate) at a given genomic location that does not include the variant that is present in the test sample. In this way, each sample included on the panel may be considered to be both a test sample and a reference sample (e.g., a test sample when tumor fraction data is being analyzed for the sample-in-question and a reference sample when tumor fraction data is being analyzed for another sample on the panel).
[0057] Because the genetic variants in each subject sample are different than those contained in the other samples, the background data for each subject sample may contain analytical information for the genomic variant regions of the other samples. Specifically, different genomic locations may have different error rate properties due to a variety of factors. For example, an error may be more likely to occur at an adenine (“A”) base in one position of the genome versus an “A” base in another position of the genome based on the chemical properties of the surrounding sequence and their behavior at different steps in the assay and / or sequencing process. By analyzing the background data, site-specific error rates (e.g., how likely errors are to occur at specific genomic locations) may be obtained. More particularly, the background-data-generating reference samples may effectively act as negative controls, which may provide a way to assess how the tumor fraction estimation performs when no tumor material is expected (i.e., a “null” scenario).
[0058] Variants described herein may include, e.g., single nucleotide polymorphisms, methyl variants, indel variants, or copy number aberrations.
[0059] Referring now collectively to graph 400 in FIG. 4A and graph 405 in FIG.4B, the thresholds for detecting tumor material using the method described in reference toFIG. 3 are presented. In an aspect, the background data generated from the combined panel sequencing may be utilized to generate a “null” distribution of tumor fraction lower bounds. In this context, a “null” case refers to situations where there is no tumor material present at the specific positions of interest. By calculating tumor fraction lower bounds for the negative control reference samples, a distribution of expected lower bounds may be identified, which may correspondingly serve as a baseline threshold for assessing the validity of tumor fraction estimates in samples with somatic variants. Specifically, when estimating the tumor fraction for a sample, the lower bound of the tumor fraction may be compared against the null distribution threshold, such as the exemplary null distribution threshold 410 illustrated in graph 400 in FIG. 4A. If the calculated lower bound of a tumor fraction is above this threshold, then it is suggested that tumor material is present in the sample. Conversely, if the calculated lower bound of the tumor fraction is below this threshold, then it is suggested that the presence of tumor material cannot confidently be identified and that any observed variants may be attributed to sequencing errors. The positioning of this threshold may be further supported by graph 405 in FIG. 4B, which provides an indication of the limit of detection (LoD) attributable to this process. For instance, all samples with zero observed alternate alleles were classified as “undetected,” including 54 undetected despite non-zero alternate allele observations (few alternate alleles relative to total coverage). Based on this, the LoD of this process, i.e., le'5, is roughly in line with expectations. For instance, (le6coverage) x (le-5tumor fraction) x (0.3 tumor allele fraction) = 3 alternate allele fragments v1 expected from approximately le-6 noise rates.
[0060] For example, referring back to Table 1, the total observed mutation count of2 in the ccga_5314 sample may be compared to the null distribution threshold 410 of graph 400 in FIG. 4A. If the observed mutation count of 2 fell below the null distribution threshold 410, then the count of 2 for the ccga_5314 sample may be attributed to background noise. Ifthe observed mutation count of 2 fell above the null distribution threshold 410, then the count of 2 for the ccga_5314 sample may indicate the presence of tumor. In this manner, systems and methods described herein may provide more context with which to assess tumor fraction within a sample and may allow for accurate detection of smaller amounts of cancer present in the samples.
[0061] The null distribution threshold may be a specificity threshold and may be set according to a confidence interval, e.g., a 99% confidence interval. In some aspects, a threshold that is used may be consistent across panels. However, if enough reference samples are used per panel, then the threshold may be calculated for each panel and to the set of sequences being targeted in that panel.
[0062] In some aspects, the method of FIG. 3 may output a binary result, which may be the presence or absence of a disease state (e.g., cancer) in a test sample. In this aspect, the number of observed mutation counts in the test sample may be compared to the threshold, and if the number of mutation counts is above the threshold, it may be determined that the subject from which the test sample was taken has a disease state. If the number of mutation counts in the test sample is below the threshold, it may be determined that the subject from which the test sample was taken does not have a disease state. In other aspects, however, if the number of mutation counts in the test sample is within a standard deviation from the threshold or is otherwise determined as falling close to the threshold, a recommendation may be made to take a new sample or to re-sequence the sample (e.g., at a deeper depth of coverage).
[0063] One benefit to selecting a plurality of reference samples for sequencing on a panel with the test sample is that all of the samples are sequenced with the same assay, in the same lab, at the same time. Accordingly, any assay-specific error rates or environmental factors would affect each sample equally. That said, it may be possible to accumulate adatabase of error rates at specific genomic locations and to run a panel with an individual test sample and then compare the test sample to a database of previously sequenced reference samples, instead of, or in addition to, running reference samples on the same panel as the test panel. The database of reference samples may or may not be comprised of reference samples that were sequenced using the same assay type as the assay type used to sequence the test sample.
[0064] The data generated from the negative reference control samples may also be utilized to identify genomic locations that exhibit unexpectedly high error rates. These abnormal rates may arise from various sources of noise, potentially due to factors such as contamination (e.g., when foreign DNA or genetic material from sources other than the intended sample enters the sequencing process), technical artifacts (e.g., stemming from limitations in the sequencing technology and / or process), low-complexity sequences (e.g., due to the difficulty of accurately aligning these repetitive sequences with a reference genome), and the like. These sites may introduce false signals into the tumor fraction estimation and may compromise the accuracy of the results. Accordingly, in some aspects, once noisy sites are identified, they may be flagged for additional attention, as further described herein.
[0065] In an aspect, to identify noisy sites in the data, the negative control reference samples (which lack one or more targeted variants) may be processed through the tumor fraction estimation pipeline. The calculated tumor fraction lower bounds from these negative control samples may then be compared to the null distribution threshold. Genomic sites that consistently yield tumor fraction lower bounds significantly above the null distribution threshold may be identified as potential noisy sites. For instance, referring now to FIG. 5A, graph 500 is illustrated that contains a long tail of confident tumor fraction estimations. However, when this data is examined further, low level contamination across the assay batch(as presented in Table 2, below, with a focus on the top row associated with chromosome ch3, and as presented in association with FIG. 6) may be observed. Additionally, indications of outlier driven by noisy sites, as presented in Table 3 below, particularly the top two rows associated with the ccga_1544 and ccga_3008 samples, may also be observed. For instance, ccga_746 appears to have one site with high background rates across multiple samples.Table 2Table 3
[0066] Once noisy sites are identified, they may be flagged and either removed from subsequent calculations or treated with a noise reduction technique. For instance, with respect to the former, if the noisy sites are determined to be too problematic, or have clear signs ofcontamination, then they may be removed from further analysis altogether. This removal may ensure that they do not contribute to false positives in the tumor fraction estimation. With respect to the latter, a variety of different types of noise removal techniques may be employed. For instance, in one approach, the error rate estimate may specifically be adjusted for the identified noisy site. More particularly, rather than applying a uniform error rate estimate for the entire dataset, a site-specific error rate may be calculated for each noisy site based on its observed error patterns. These adjusted error rates may be used in subsequent analyses, enabling more accurate variant calling and tumor fraction estimation. In another approach, noisy sites may be subjected to stricter filtering criteria during variant calling and / or tumor fraction estimation. More particularly, sites with elevated error rates may require more stringent checks to differentiate true variants from sequencing errors. Adjusting the filtering parameters for these noisy sites may ensure that only high-confidence variants are considered in the various analytical calculations. In yet another approach, noisy sites may be assigned lower weights during tumor fraction calculation. Specifically, the contribution of each site to the overall tumor fraction estimate may be weighted based on the site’s error rate and reliability. In this way, sites with higher error rates may contribute less to the final tumor fraction, thereby reducing their impact on the results.
[0067] In one aspect, a combination of any of the foregoing noise-reduction techniques may be employed together. Additionally or alternatively, the processes for noise reduction may be iterative in nature and may involve multiple rounds of analysis, assessment, and adjustment. For instance, after applying an initial noise reduction technique, the impact on variant calling, tumor fraction estimation, and downstream analysis may be evaluated. Depending on the observations and the results, one or more additional rounds of noise reduction may be employed to effectively reduce the influence of noisy sites to a desired level.
[0068] Referring now to FIG. 7, an exemplary workflow for estimating tumor fraction in a biological sample is disclosed. At step 705, sequencing data may be received from a combined targeted sequencing panel. In an aspect, the combined targeted sequencing panel may contain a plurality of biological samples that are each associated with a different subject. Each biological sample may contain genetic variants that are unique to their respective subject and that are also not shared among the other genetic variants represented in the panel. Stated differently, the combined panel may be designed to contain no overlapping genetic variants across the biological samples. In an aspect, the sequencing data may contain sequencing information for each subject represented in the panel. Specifically, the sequencing data for each subject provides sequencing coverage for the genomic locations associated with their genetic variants along with sequencing coverage for the genomic locations associated with the genetic variants of the other subjects included on the panel. As a non-limiting example, given two subjects contained on the same panel, the sequencing information from Subject 1 may cover Subject l’s genetic variants and may also provide data that covers Subject 2’s variants, and vice versa. Because Subject 1 does not have the same genetic variants as Subject 2, the coverage of Subject 2’s variants by Subject l’s data may provide “background data” of how the genomic location associated with Subject 2’s variants may normally behave.
[0069] At step 710, site-specific error rates for each genomic location may be calculated from the sequencing data. More particularly, the background data (e.g., the portion of the sequencing data for each subject that covers the genomic locations of the other subject’s variants) from the combined panel may be treated as negative controls that may be reintroduced to the tumor fraction estimation pipeline, previously described above with respect to FIG. 2. By generating tumor fraction estimates for these negative sample controls, a distribution of expected lower bounds may be identified, which may correspondingly beused to establish, at step 715, a baseline threshold for assessing the validity of tumor fraction estimates in samples with somatic variants. Specifically, those samples having a lower tumor fraction bounds above this threshold may confidently may considered to contain tumor material, whereas those samples having a lower tumor fraction bounds below this threshold may not be confidently classified as containing tumor material.
[0070] At optional step 720, the sequencing data generated from the negative sample controls may also be utilized to identify genomic locations that exhibit unexpectedly high error rates. These abnormal rates may arise from various sources of noise, potentially due to factors such as contamination (e.g., when foreign DNA or genetic material from sources other than the intended sample enters the sequencing process), technical artifacts (e.g., stemming from limitations in the sequencing technology and / or process), low- complexity sequences (e.g., due to the difficulty of accurately aligning these repetitive sequences with a reference genome), and the like.
[0071] In certain aspects, one or more noise-reduction techniques may be implemented to address the noise at these sites. More particularly, sites with high-background error rates may be filtered out from a data pool and / or not used in various downstream processes. In an aspect, sites with high background error rates may be identified by first randomly selecting sequencing data from a subset (e.g., about 70%) of the total participants included on a combined targeted sequencing panel. The sequencing data associated with the remaining subset may be reserved for generating the null distribution of ctDNA detectability. For the randomly selected subset of participants, pileups covering their variants may be generated, which may subsequently be summed across the background participants to obtain an aggregate count of alternate and reference nucleotide observations at each site. In an aspect, an iterative filtering algorithm may then be applied to the aggregated background counts to remove sites with high background rates. In an aspect, the filtering process may befacilitated by employing the following equation, p = totalAlt / totalCov, where totalAlt corresponds to the sum of all alternate allele counts across all variants, and totalCov corresponds to the sum of total coverage across all variants. For each site, the alternate allele coverage may be represented by a variable, e.g., “k,” and the total coverage may be represented by another variable, e.g., “N ” The site may be removed if P(X > k) < T, where X is distributed according to Binomial (N,p) and T is some threshold value. For instance, a threshold value of about le'3may be utilized, whereby such a threshold may be chosen based on inspection of a small number of clear outlier sites.
[0072] In general, any process discussed in this disclosure that is understood to be computer-implementable may be performed by one or more processors of a computer system, such as system environment 100, as described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.
[0073] A computer system, such as system environment 100, may include one or more computing devices. If the one or more processors of the computer system are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If a system environment comprises a plurality of computing devices, the memory of the computer system may include the respective memory of each computing device of the plurality of computing devices.
[0074] FIG. 8 is a simplified functional block diagram of a computer system 800 that may be configured as a computing device for executing the processes described herein, according to exemplary embodiments of the present disclosure. FIG. 8 is a simplified functional block diagram of a computer that may be configured as according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 820 for packet data communication. The platform also may include a central processing unit (“CPU”) 802, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 808, and a storage unit 806 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 822, although the system 800 may receive programming and data via network communications via electronic network 825 (e.g., voice, video, audio, images, or any other data over the electronic network 825) . The system 800 may also have a memory 804 (such as RAM) storing instructions 824 for executing techniques presented herein, although the instructions 824 may be stored temporarily or permanently within other modules of system 800 (e.g., processor 802 and / or computer readable medium 822). The system 800 also may include input and output ports 812 and / or a display 810 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.
[0075] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, orassociated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non- transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0076] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0077] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example, functionality may be added or deleted from the block diagrams and operations may be interchanged among functionalblocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0078] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method for estimating tumor fraction in a biological sample, the method comprising: receiving, at a computing device, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples, and wherein the one or more genetic variants between each of the plurality of biological samples are different from one another, and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; and background sequencing data covering all other of the one or more genetic variants for all other of the plurality of biological samples; calculating, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establishing, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.
2. The computer-implemented method of claim 1, wherein the biological sample is one of: a tissue sample, a urine sample, and a blood sample.
3. The computer-implemented method of claim 1, wherein selection of the plurality of biological samples to include in the combined targeted sequencing panel is based at least in part on: biological sample size, subject age, or subject sex.
4. The computer-implemented method of claim 1, further comprising: estimating a tumor fraction for each of the plurality of biological samples based on the sequencing data; identifying a lower bound of the tumor fraction; and comparing the lower bound of the tumor fraction to the threshold.
5. The computer-implemented method of claim 4, further comprising: generating, responsive to identifying that the lower bound of the tumor fraction is greater than the threshold, a confidence indication that tumor material is present in the biological sample.
6. The computer-implemented method of claim 4, further comprising: generating, responsive to identifying that the lower bound of the tumor fraction is lower than the threshold, an indication that that tumor material cannot be confidently identified in the biological sample.
7. The computer-implemented method of claim 1, wherein the threshold includes a plurality of thresholds, and wherein each of the plurality of thresholds is associated with a different confidence rating for tumor material presence.
8. The computer-implemented method of claim 1, further comprising: utilizing the site-specific error rate to identify that at least one of the genomic locations is a noise-generating location; andapplying a noise reduction technique to mitigate noise produced from the identified noise-generating genomic location.
9. The computer-implemented method of claim 8, wherein the applying the noisereduction technique comprises treating data associated with the noise-generating genomic location differently than other data associated with other genomic locations.
10. The computer-implemented method of claim 8, wherein the applying the noisereduction technique comprises removing data associated with the noise-generating genomic location from an overarching dataset.
11. A system for estimating tumor fraction in a biological sample, the system comprising: one or more processors; one or more computer readable media storing instructions that are executable by the one or more processors to perform operations to: receive, at a computing device associated with the system, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples and wherein the one or more genetic variants between each of the plurality of biological samples are different from one another, and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; andbackground sequencing data covering all other of the one or more genetic variants for all other of the plurality of biological samples; calculate, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establish, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.
12. The system of claim 11, wherein the biological sample is one of: a tissue sample, a urine sample, and a blood sample.
13. The system of claim 11, wherein selection of the plurality of biological samples to include in the combined targeted sequencing panel is based on: biological sample size, subject age, or subject sex.
14. The system of claim 11, wherein the operations further comprise operations to: estimate a tumor fraction for each of the plurality of biological samples based on the sequencing data; identify a lower bound of the tumor fraction; and compare the lower bound of the tumor fraction to the threshold.
15. The system of claim 14, wherein the operations further comprise operations to generate, responsive to identifying that the lower bound of the tumor fraction is greater than the threshold, a confidence indication that tumor material is present in the biological sample.
16. The system of claim 14, wherein the operations further comprise operations to generate, responsive to identifying that the lower bound of the tumor fraction is lower than the threshold, an indication that tumor material cannot be confidently identified in the biological sample.
17. The system of claim 11, wherein the operations further comprise operations to: utilize the site-specific error rate to identify that at least one of the genomic locations is a noise-generating location; and apply a noise reduction technique to mitigate noise produced from the identified noise-generating genomic location.
18. The system of claim 17, wherein the operations to apply the noise reduction technique comprise operations to: treat data associated with the noise-generating genomic location differently than other data associated with other genomic locations.
19. The system of claim 17, wherein the operations to apply the noise reduction technique comprise operations to: remove data associated with the noise-generating genomic location from an overarching dataset.
20. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a system, cause the system to perform operations comprising:receiving, at a computing device, sequencing data for each of a plurality of biological samples included on a combined targeted sequencing panel, each of the plurality of biological samples being associated with one of a plurality of subjects, wherein one or more genetic variants are known for each of the plurality of biological samples and wherein the one or more genetic variants between each of the plurality of biological samples are different from one another, and wherein the sequencing data for each of the plurality of biological samples includes: genetic variant sequencing data covering the one or more genetic variants; and background sequencing data covering all other of the one or more genetic variants for all other of the plurality of biological samples; calculating, from the background sequencing data associated with each of the plurality of biological samples, a site-specific error rate for each genomic location associated with each of the one or more genetic variants; and establishing, based on the calculated site-specific error rate, a threshold for tumor fraction estimation for each of the one or more genetic variants.