Residual disease detection system and method
A system and method using algorithms and classifiers to distinguish tumor markers from noise, integrating quality metrics and subject-specific parameters, enhances the detection of minimal residual disease in low-tumor burden cancers, achieving improved sensitivity and accuracy.
Patent Information
- Application Number
- JP2024091909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-27
- Filing Date
- 2024-06-06
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2039-02-27
AI Technical Summary
Existing methods for detecting minimal residual disease (MRD) in low-tumor burden cancers, such as non-small cell lung cancer, face challenges due to limited input material and low tumor fraction, leading to reduced detection sensitivity and accuracy.
A system and method for diagnosing residual disease using algorithms and statistical classifiers to distinguish tumor-specific markers from artifactual noise, integrating quality metrics and subject-specific parameters to estimate tumor fraction, and applying integrated mathematical models for accurate detection.
The method achieves highly sensitive and accurate detection of low-abundance tumor markers, surpassing existing techniques by improving detection limits to 0.0001-0.000001 tumor fraction, enabling non-invasive monitoring of residual disease.
Smart Images

Figure 0007821415000005 
Figure 0007821415000006 
Figure 0007821415000007
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 62 / 636,150, filed February 27, 2018, the entire contents of which are incorporated herein by reference. [Technical Field] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate generally to the field of medical diagnostics. In particular, aspects of the present disclosure relate to compositions, methods, and systems for tumor detection and diagnosis. [Background technology]
[0002] Cell-free circulating DNA (cfDNA) released from dying cells allows for the time-course dynamics of the somatic genome and epigenome for clinical purposes. Biopsies can be obtained by simple blood sampling, allowing for dynamic genomic measurements in a non-invasive manner. This overcomes the spatial limitations of inaccessible tissues such as lung tissue.
[0003] Circulating tumor DNA (ctDNA), not to be confused with cell-free DNA (cfDNA), can be found and measured in the blood of cancer patients. ctDNA has been shown to correlate with changes in tumor burden and response to treatment or surgery (Non-Patent Document 1). ctDNA is also detectable in early-stage non-small cell lung cancer (NSCLC), and thus may transform the diagnosis and treatment of NSCLC (Non-Patent Documents 2-5).
[0004] One of the major promising areas for cfDNA-based cancer research is the detection of residual disease (RD) to guide clinical intervention. For example, detecting residual disease after surgical resection can lead to costly and toxic adjuvant therapy decisions for clinicians and patients. However, in low-burden tumors, such as minimal residual disease (MRD), the tumor fraction (TF) is significantly lower. To detect mutations in low-TF cfDNA, a commonly accepted paradigm is to increase the sequencing depth of a limited, high-yield target set (e.g., common cancer driver or patient-specific panels sequenced to a depth of approximately 10,000–100,000 reads / base). Furthermore, molecular and analytical approaches are integrated with ultra-deep sequencing to reduce sequencing error and improve the sensitivity of detection at low TF.
[0005] While state-of-the-art methods offer highly accurate detection in some instances, they are hampered by a fundamental limitation that reduces detection sensitivity: limited input material. In MRD, tumor burden is low, with typical plasma samples containing only 1–10 ng / ml of cfDNA. This low amount of cfDNA represents only a few hundred to a few thousand genome equivalents. Therefore, common techniques relying on ultra-deep sequencing (e.g., 100,000X) may be ineffective due to the limited number of physical fragments that can be obtained at each site in the sample (e.g., 1,000 genome equivalents in 6 ng of cfDNA). Even with extremely deep sequencing and advanced molecular error suppression, the detection limit is less than 0.1–1% tumor fraction (TF) frequency with limited input material. Thus, while detecting cancers with low tumor burden is clinically beneficial for patients and clinicians, existing methods relying on the identification of somatic mutations face significant challenges due to the low frequency of tumor-derived cfDNA samples.
[0006] Therefore, there is an urgent and unmet need for minimally invasive systems and methods that allow for tumor detection, particularly in the context of diagnosing minimal residual disease (MRD) with limited input material. Effective diagnosis of tumors in the setting of residual tumor (e.g., after surgery and / or treatment) is beneficial from an economic and clinical perspective. This is particularly true for lung cancer, as many patients are diagnosed with advanced-stage disease with poor outcomes (Non-Patent Document 6). [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] Diehl et al.,Nature medicine,14(9):985-990,2008 [Non-patent document 2] Sozzi et al., Journal of Clinical Oncology, 21(21), 3902-3908, 2003 [Non-patent document 3] Tie et al.,Science translational medicine, 8(346):346ra92-346ra92,2016 [Non-patent document 4] Bettegowda et al.,Science translational medicine,6(224):224ra24-224ra24, 2014 [Non-patent document 5] Wang et al.,Clinical Cancer Research, 16(4): 1324-1330,2010 [Non-patent document 6] Herbst et al.,N Engl J Med.,359(13):1367-80,2008 Summary of the Invention
[0008] The present disclosure relates to methods and systems for diagnosing residual tumor disease by analyzing tumor-specific markers in a subject's sample (e.g., a plasma or blood sample). The disclosed methods utilize algorithms and / or statistical classifiers to distinguish quality markers from artifactual noise based on several parameters. For example, if the marker is a single nucleotide variation (SNV), the disclosed algorithm classifies the SNV in the subject's genetic inventory as signal or noise based on qualitative characteristics of the marker, such as the SNV's base quality (BQ) and the SNV's mapping quality (MQ). Similarly, if the marker is a copy number variation (CNV), the algorithm classifies the CNV in the inventory as signal or noise based on parameters such as centromere proximity, overlap with the cfDNA coverage mask, and / or the CNV's association with low mappability (mapping quality; MQ) reads. Thus, markers likely associated with artifactual noise are removed from the subject's genetic inventory, and high-quality markers are processed through a stable, integrated mathematical model that can estimate the tumor fraction in the sample. If the estimated tumor fraction is found to be above a certain threshold, then there is a high degree of confidence in a positive diagnosis. In contrast, if the estimated tumor fraction is below the threshold, then there is no positive diagnosis at that time.
[0009] In this context, simulated testing of plasma somatic mutations calling using a synthetic mixture of tumor and normal whole-genome sequence data from lung patients, where the varying proportion of tumors ranges from 1% to 0.001% (1 / 100,000), reveals that the strength and accuracy of our method surpass existing techniques.
[0010] The present disclosure also relates to several indicators that may suggest that a variant detected by sequencing is not a true somatic mutation, but rather an artifact of the sequencing or mapping technique. In this context, previous studies have shown that sequencing errors are not random and are likely related to the context of the DNA sequence resulting from the sequencing technique and technical factors. The fidelity of sequencing is also limited by the length of each sequencing read, and the error rate increases as the read length increases. Errors may occur when reads are mapped to a reference genome. The mapping process is computationally intensive and complex due to the fact that genomes have variable regions, motifs, and repeatable elements. Short nucleotide reads may map to more than one location or may not map at all. These limitations of existing methodologies for sequencing / mapping genomic data may be corrected using the systems and methods of the present disclosure. The indicators of the present disclosure may analyze multiple factors, such as (i) low base quality; and / or (ii) low mapping quality, (iii) read mutation location, and (iv) read fragment size in the case of SNV markers, and correlation between (1) genome location score, (2) cfDNA coverage mask (blacklist), (3) low mapping quality, and (4) Log2 and read fragment size in the case of CNV markers, to call true mutations from errors.
[0011] The system and method of the present invention for detecting tumor-related biomarkers is particularly applicable to the detection of low-abundance markers. First, the model takes into account quality metrics related to the type of marker and the system / method used to detect it, as well as subject-specific parameters to calculate the estimated tumor fraction (eTF). For example, if the marker is an SNV, the integrated mathematical model takes into account process quality metrics such as estimated coverage and noise, as well as subject-specific parameters such as mutation load. In the case of a CNV, the integrated mathematical model takes into account indicator factors, along with subject-specific characteristics such as the directionality of the CNV (e.g., amplification is a positive factor, and deletion is a negative factor), to calculate the estimated tumor fraction (eTF). Thus, the analytical approach of the present disclosure integrates genome-wide mutation information to enable highly sensitive analysis of samples containing cfDNA, so that residual disease can be accurately and non-invasively diagnosed.
[0012] Accordingly, the present disclosure relates to the following non-limiting embodiments:
[0013] In various embodiments, a method for detecting residual disease in a subject in need thereof is provided. The method may include receiving a first subject-specific genome-wide list of reads associated with genetic markers from a first biological sample from the subject. The first biological sample may include a baseline sample and a normal cell sample. Each of the first read lists includes reads of a single base pair in length (e.g., SNVs or Indels), and the baseline sample may include a tumor sample or a plasma sample. The method may further include filtering artifactual sites from the first read list. The filtering may include removing repetitive sites generated across a cohort of reference healthy samples from the first list of genetic markers. And / or identifying germline mutations in peripheral blood mononuclear cells of a normal cell sample and removing the germline mutations from the first list of genetic markers. The method may further include detecting reads from a second subject-specific genome-wide list of genetic markers in a second biological sample from the subject to generate a tumor-associated genome-wide list of genetic markers in the second sample. The method may further include filtering noise from the first and second genome-wide read lists. The filtering may include using at least one error suppression protocol to generate a first filtered read list for the first genome-wide read list and a second filtered read list for the second genome-wide read list. The at least one error suppression protocol may include (a) calculating the probability that any single nucleotide variation in the first and second suppressions is an artifact and removing the variation. The probability may be calculated as a function of features selected from the group consisting of mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof. And / or, the at least one error suppression protocol may include removing artifacts using mismatch testing between independent copies of the same DNA fragment generated from a polymerase chain reaction or sequencing process. The mismatch testing and / or overlap consensus may include identifying and removing artifacts when a majority of a given overlap family does not match.The method further includes applying a background noise model to one or more integrated mathematical models to predict tumors in the first and second biological samples using the first and second filtered read sets. Fraction (eTF). The method may further include detecting residual tumor in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.
[0014] In various embodiments, a method for detecting residual disease in a subject in need thereof is provided. The method may include (A) receiving a first subject-specific genome-wide read list associated with genetic markers from a first biological sample of the subject. The biological sample may include a baseline sample. Each of the first read lists includes a single base pair long read, and the baseline sample includes a tumor sample or a plasma sample. The method may further include receiving a second subject-specific genome-wide read list associated with genetic markers from a second biological sample of the subject. The second biological sample may include a peripheral blood mononuclear cell sample (PBMC). Each of the second lists of genetic markers may include copy number variations (CNVs). The method may further include filtering artifactual sites from the first and second read lists. The filtering may include removing repetitive sites generated across a cohort of reference healthy samples from the first and second lists of genetic markers. And / or the filtering may identify CNVs shared by the first and second lists as germline mutations and remove the mutations from the first and second lists of reads. The method may further include detecting reads from a third subject-specific genome-wide list of the genetic markers in a third biological sample from the subject and generating a tumor-associated genome-wide list of the genetic markers in the third sample. The method may further include normalizing each of the first, second, and third read lists to generate a first filtered read set for the first genome-wide read list, a second filtered read set for the second genome-wide read list, and a third filtered read set for the third genome-wide read list. The method may further include using the third filtered read set to apply a background noise model to one or more integrative mathematical models to calculate an estimated tumor fraction (eTF) of the third biological sample. One or more models can be configured to generate a first eTF using the first filtered read set, or one or more models can be configured to generate a second eTF using the second filtered read set. The method can further include detecting residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold.
[0015] In some embodiments, the present disclosure relates to methods for detecting residual disease in a subject in need thereof. Preferably, detecting residual disease includes detecting minimal residual disease during treatment. In particular, the present disclosure relates to detecting residual disease during one or more of: (a) after resection surgery; (b) during or after treatment; (c) during monitoring treatment efficacy; (d) during monitoring tumor recurrence or recurrence; or (e) a combination thereof. In particular, the present disclosure relates to detecting residual disease during or after chemotherapy, immunotherapy, targeted therapy, or a combination thereof; and / or during the process of monitoring the effectiveness of such treatment.
[0016] In some embodiments, the present disclosure provides a method for detecting residual disease in a subject in need thereof, comprising the steps of: (A) receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from biological samples of the subject, the biological samples including a tumor sample and optionally a normal sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list of genetic markers in a second biological sample of the subject, thereby generating a tumor-associated genome-wide list of genetic markers in the second sample; (C) detecting a tumor-associated genome-wide list of genetic markers in the second sample; The method includes a step of filtering artificial noise markers from the list of wide genetic markers, wherein the filtering comprises statistically classifying each SNV or Indel in the list as signal or noise as a function of 1) the mapping quality (MQ) of the reads containing the SNV, 2) the fragment size length of the reads containing the SNV, 3) a consensus test within the read overlap family containing the SNV or Indel, and 4) the base quality (BQ) of the SNV or Indel, and / or statistically classifying each CNV or SV window in the list as signal or noise based on 1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and 3) overlap with a cfDNA mask (blacklist), and determining the probability of noise detection (P N(D) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (E) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by the background noise model. In some embodiments of the method, (1) for SNV markers, the estimated TF (eTF[SNV]) is calculated by integrating estimated genome coverage and sequencing noise with patient-specific parameters including mutation burden (N); and (2) for CNV markers, the estimated TF (eTF[CNV]) is calculated by integrating directional depth of coverage skewed consistent with tumor CNV directionality, where copy number gains are positively skewed and copy number losses are negatively skewed. In some embodiments, the BQ, MQ, and fragment size filters of the markers are optimized using a receiver operating characteristic curve. In some embodiments, the method includes using a combined base quality mapping quality (BQMQ) filter.
[0017] In some embodiments, the residual disease detection method of the present disclosure is performed by receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from biological samples, including tumor samples and normal samples, including non-tumor samples, from the subject. In some embodiments, the method includes generating the genome-wide list of markers using the subject's tumor sample and the subject's peripheral blood mononuclear cells (PMBCs). In particular, the genome-wide list of genetic markers is generated by whole-genome sequencing of the subject's sample (e.g., tumor sample) and a control sample (e.g., PMBCs). Preferably, the subject's tumor sample comprises a resected tumor, e.g., a solid tumor removed after surgery, such as mastectomy, prostatectomy, skin lesion resection, small bowel resection, gastrectomy, thoracotomy, adrenalectomy, colectomy, oophorectomy, thyroidectomy, hysterectomy, glossectomy, or colon polypectomy, preferably thoracotomy.
[0018] In some embodiments, the disclosure provides a method for detecting residual disease in a subject in need thereof, comprising: (A) receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from a biological sample of the subject, the biological sample comprising a tumor sample and optionally a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; detecting the subject-specific genome-wide list of genetic markers in a second biological sample of the subject to generate a tumor-associated genome-wide list of genetic markers in the second sample; and (C) filtering artifactual noise markers from the genome, wherein the noise (P) is a function of 1) mapping quality (MQ) of reads containing SNVs, 2) fragment length of reads containing SNVs, 3) consensus testing within read duplication families containing SNVs or Indels, and 4) base quality (BQ) of the SNVs or Indels. N (D) filtering the SNVs or Indels by statistically classifying each SNV or Indel as signal or noise based on its probability of detection (SNV or Indel detection probability) and / or by statistically classifying each SNV or Indel as signal or noise based on 1) its location relative to the centromere, 2) the mapping quality (MQ) of the read set containing the CNV or SV window, and 3) overlap with a cfDNA mask (blacklist); (C) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrative mathematical models; and (E) diagnosing residual disease in the subject based on an empirical threshold calculated by the estimated tumor fraction and background noise model, wherein the read set includes a read set covering a specific SNV or Indel site or a read set included in a specific CNV or SV genomic window. In some embodiments, the normal cell sample includes PMBC, a saliva sample, a hair sample, or a skin sample. In some embodiments, the subject is human, and the second biological sample from the subject includes biological material selected from blood, cerebrospinal fluid, pleural effusion, ocular fluid, stool, urine, or a combination thereof.
[0019] In some embodiments of the present disclosure, the tumor sample comprises a resected tumor or fine needle aspiration (FNA) sample, snap-frozen tissue, optimal cutting temperature compound (OCT)-embedded tissue, or formalin-fixed paraffin-embedded (FFPE) tissue.
[0020] In some embodiments of the present disclosure, the normal sample comprises peripheral blood mononuclear cells (PMBCs) or a saliva or skin sample.
[0021] In some embodiments of the present disclosure, the plurality of genetic markers is received by whole genome sequencing of a subject's biological sample and a control sample.
[0022] In some embodiments of the present disclosure, the list of tumor genetic markers includes a high mutation rate and / or a high number of SNPs, indels, CNVs or SVs, for example, at least 1, at least 2, at least 3, at least 5, at least 7, at least 10 and more, for example, about 15 SNPs or indels per megabase pair, or CNVs / SVs with a cumulative size of at least 5 megabase pairs (MBP), at least 7 MBP, at least 10 MBP or more, for example, a cumulative size of about 15 MBP.
[0023] In some embodiments, the disclosure provides a method for detecting residual disease in a subject in need thereof, comprising: (A) receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from a biological sample of the subject, the biological sample comprising a tumor sample and optionally a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list of genetic markers in a second biological sample of the subject to generate a tumor-associated genome-wide list of genetic markers in the second sample; and (C) filtering artificial noise markers from the genome, wherein the noise (PQ) is calculated as a function of 1) mapping quality (MQ) of reads containing SNVs, 2) fragment length of reads containing SNVs, 3) consensus testing within read duplication families containing SNVs or Indels, and 4) base quality (BQ) of the SNVs or Indels. N (D) filtering the SNVs or Indels by statistically classifying each SNV or Indel as signal or noise based on its probability of detection in the SV window and / or by statistically classifying each SNV or Indel as signal or noise based on 1) its location relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and 3) overlap with a cfDNA mask (blacklist); (C) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrative mathematical models; and (E) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model, wherein the empirical noise model is defined by measuring the detection error rate in normal healthy samples and converted into a base noise eTF estimate.
[0024] In some embodiments of the present disclosure, the eTF estimation noise threshold is 0.0001 (10 -4 )~0.000001(10 -6 )
[0025] In some embodiments, the disclosure provides a method for detecting residual disease in a subject in need thereof, comprising: (A) receiving a subject-specific genome-wide list of somatic genetic markers derived from a plurality of genetic markers from a biological sample of the subject, the biological sample comprising a tumor sample and a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) then detecting the subject-specific genome-wide list of genetic markers in a second biological sample comprising a plasma sample of the subject, generating a tumor-associated genome-wide list of genetic markers in the second sample; and (C) filtering artifactual noise markers from the genome, wherein the noise (PQ) is calculated as a function of 1) mapping quality (MQ) of reads containing SNVs, 2) fragment length of reads containing SNVs, 3) consensus testing within read duplication families containing SNVs or Indels, and 4) base quality (BQ) of the SNVs or Indels. N(D) filtering the SNVs or Indels by statistically classifying each SNV or Indel as signal or noise based on its probability of detection (SNV or Indel window) and / or by statistically classifying each SNV or Indel as signal or noise based on 1) its location relative to the centromere, 2) the mapping quality (MQ) of the reads containing the SNV or Indel window, and 3) overlap with a cfDNA mask (blacklist); (C) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrative mathematical models; and (E) diagnosing the subject's residual disease based on the estimated tumor fraction and an empirical threshold calculated by the background noise model. In some embodiments, the normal cell sample comprises PMBC, a saliva sample, a hair sample, or a skin sample. In some embodiments, the subject is human, and the second biological sample from the subject comprises biological material selected from blood, cerebrospinal fluid, pleural effusion, ocular fluid, stool, urine, or a combination thereof. In some embodiments, the BQ, MQ, and fragment size filter of the markers are optimized using a receiver operating characteristic curve. In some embodiments, the method includes using a combined base quality mapping quality (BQ MQ) filter.
[0026] In some embodiments, detecting minimal residual disease includes quantitatively estimating a patient's minimal residual disease burden during patient treatment, observation, or monitoring. In particular, detecting minimal residual disease includes detecting residual disease after resective surgery; detecting residual disease during or after treatment; detecting residual disease in monitoring treatment efficacy; detecting residual disease in monitoring cancer recurrence or recurrence; or a combination thereof. In certain embodiments, detecting minimal residual disease includes detecting residual disease after resective surgery, including lymph node biopsy; head and neck surgery; uterine or endometrial biopsy; bladder biopsy; mastectomy; prostatectomy; removal of skin lesions; small bowel resection; gastrectomy; thoracotomy; adrenalectomy; colectomy; oophorectomy; thyroidectomy; hysterectomy; glossectomy; or colon polypectomy. In certain embodiments, detecting minimal residual disease includes detecting residual disease after treatment, including chemotherapy, immunotherapy, targeted therapy, radiation therapy, or a combination thereof.
[0027] In some embodiments of the present disclosure, the disease detection method includes receiving a plurality of genetic markers from a biological sample of a subject, the biological sample including a tumor sample and a normal cell sample, and further includes generating a subject-specific genome-wide list of genetic markers from the received plurality of genetic markers.
[0028] In some embodiments of the present disclosure, the disease detection method further comprises detecting a subject-specific genome-wide list of genetic markers in a second biological sample, e.g., a plasma sample. In some embodiments, the second biological sample is detected in the subject over time (e.g., 2 days, 1 week, 2 weeks, 1 month, 2 months, 3 months, 4 months, 6 months, 1 year, 18 months, 2 years, 30 months, 3 years, 42 months, 4 years, 4 years, 5 years, 7 years, 10 years, or more, e.g., 15 or 20 years) for generating a temporally updated list of tumor genome-wide genetic markers in patient plasma.
[0029] In some embodiments of the present disclosure, the disease detection method includes empirically determining a background noise threshold, where a tumor fraction above the background noise threshold provides a quantitative estimate of tumor burden, and in particular, a tumor fraction below the noise threshold is considered not detected (ND).
[0030] In some embodiments of the present disclosure, the disease detection method involves quantitative monitoring of tumor disease (e.g., tumor fraction) over time. In certain embodiments, the tumor is brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, gastric cancer, osteosarcoma, or a solid tumor, heterogeneous or homogeneous in nature. Preferably, the tumor is lung cancer, breast cancer, melanoma, bladder cancer, or osteosarcoma, such as lung adenocarcinoma, ductal adenocarcinoma, non-small cell lung cancer (NSCLC LUAD), cutaneous melanoma, urothelial carcinoma, or osteosarcoma.
[0031] In some embodiments, the residual disease detection method of the present disclosure further comprises calculating the eTFs of the SNV or indel markers by integrating a probabilistic model including: 1) the integrated signal of plasma SNV or indel detection; 2) process quality metrics including estimated genome coverage and a sequencing noise model; and 3) patient-specific parameters including mutational burden (N); and / or calculating the eTFs of the CNV or SV markers using a stochastic dilution model including: 1) integrating the directional depth of skewed coverage between plasma and normal patient samples, consistent with tumor CNV or SV orientation where copy number amplifications are positively skewed and copy number deletions are negatively skewed; 2) integrating the cumulative depth of skewed coverage between tumor and normal (PBMC) patient samples; and 3) finding the dilution ratio between the signals.
[0032] In some embodiments, the residual disease detection method of the present disclosure includes the steps of: (A) receiving a plurality of genetic markers comprising single nucleotide variations (SNVs) or copy number variations (CNVs), or a combination thereof, in a biological sample from a subject and a normal cell sample from the subject, and creating a subject-specific genome-wide list of genetic markers; (B) identifying and filtering artificial noise markers from the genome-wide marker list, wherein: (1) the noise SNVs are determined by filtering each SNV in the list as a noise (P) as a function of the base quality (BQ) of the SNV and the mapping quality (MQ) of the SNV; Nand / or (2) noise CNVs are identified by statistically classifying each CNV in the list as signal or noise based on its relative position from the centromere and overlapping the cfDNA mask blacklist within a predetermined coverage depth and read mappability; (C) calculating an estimate of the tumor fraction (eTF) of the sample based on one or more integrative mathematical models, wherein for an SNV marker, the estimated TF value (eTF[SNV]) is calculated by the formula eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific group detections in the patient sample. where σ is an empirically estimated measure of noise, R is the total number of distinct reads in a region of interest (ROI), N is the tumor mutation load, cov is the total number of distinct reads per site in the ROI, and / or for CNV markers, eTF[CNV] is calculated by eTF[eTF[CNV]=(sum_{i]=(P(i)-N(i)]*sign[T(i)-N(i)]]-E(sigma)] / (sum_{i}[abs])(T)(i)-N(i))-E(σ)), where P is the median genomic window depth where {i} represents plasma, T is the median genomic window depth where {i} represents tumor, and N is the median genomic window depth where {i} represents normal depth coverage. In particular, under this embodiment, the genomic window for estimating tumor fraction based on detection of one or more CNV markers is about 500 base pairs (bp).
[0033] In some embodiments, the disclosure provides a method of diagnosing minimal residual disease in a subject, comprising: (A) receiving a genome-wide list of reads from sequenced genetic data received from multiple biological samples from the subject, the biological samples including a tumor sample, a normal sample, and a plasma sample; (B) performing mutation calling on the tumor and PBMC samples from the subject, including MUTECT, LOFREQ, and / or STRLKA mutation calling, to generate subject-specific reads of somatic SNVs (sSNVs) or indels as an individualized reference set; (C) collecting and filtering reads from subject-specific mutation sites, including: (1) removing reads with low mapping quality (e.g., <29, ROC optimized); (2) constructing duplicate families (representing multiple PCR / sequencing copies of the same DNA fragment) and generating corrected reads based on a consensus test; and (3) filtering low base quality reads (e.g., <21, ROC optimized). (D) calculating the number of subject-specific mutation sites with at least one supporting read (in the filtered set) that has the exact same substitution as in the tumor; and (F) estimating the tumor fraction of SNVs based on the mathematical model eTF[SNV]=1-[1-(ME(σ)*R] / N]^(1 / cov) (Equation 1), where M is the number of tumor-specific group detections in the patient sample and σ is the empirically estimated No. where R is a measure of the size of the TFs, R is the total number of distinct reads in the region of interest (ROI), N is the tumor mutation burden, and cov is the average number of distinct reads per site in the ROI; (G) comparing the eTF[SNV] to a detection threshold comprising an empirically determined baseline noisy TF estimate from healthy samples, where if the eTF[SNV] is above a threshold level (e), e.g., 2 standard deviations of the noisy TF distribution (FPR<2.5%), indicates a positive detection; and (K) diagnosing residual disease in the subject based on the eTFs.
[0034] In some embodiments, the disclosure provides a method of diagnosing minimal residual disease in a subject, comprising: (A) receiving a genome-wide list of reads in sequenced genetic data from multiple biological samples received from the subject, the biological samples including a tumor sample, a normal sample, and a plasma sample; (B) calling the subject-derived tumor and PBMC samples to generate reference segments for multiple CNV segments above a threshold length (e.g., >2 Mbp, preferably >5 Mbp) along with annotation of segment directionality, where amplifications are annotated as positive and deletions are annotated as negative; (C) collecting single bp depth coverage information for the plasma, tumor, and PBMC samples covering a patient-specific CNV segmentation region of interest (ROI); (D) dividing the patient-specific CNV or SV segmentation ROI into 500 bp windows and calculating the median value per window (artificial suppression) across all samples and windows; and (E) performing a 500 bp-based analysis using (a) stable z-score normalization per sample; and / or (b) stable principal component analysis (RPCA). (F) filtering reads / windows derived from patient-specific segmentation, wherein the filtering includes: (1) removing low mapping quality reads (e.g., <29, ROC optimized); and / or (2) removing centromeric regions (e.g., removing windows with normalized normals >10); and / or (3) removing non-representative regions in cfDNA (e.g., removing a cfDNA representation mask composed of multiple cfDNA samples). (G) integrating the directional depth of skewed coverage between plasma and normal (PBMC) patient samples using the mathematical model sum[(P(i)-N(i)]*[T(i)-N(i)]sign]-E(σ) (Equation 2), where P is the median genomic window depth indexed by {i} and represents the stable z-score or stable PCA normalized plasma depth coverage compared to a cohort of normal samples; and E(σ) is an empirically estimated measure of error rate;T is the median genomic window depth indexed by {i}, which represents tumor depth coverage normalized by stable z-score or stable PCA methods, and N is the median genomic window depth indexed by {i}, which represents normal depth coverage normalized by stable z-score or stable PCA methods compared to a cohort of normal samples; (H) integrating the cumulative coverage depth of tumor and normal (PBMC) patient samples using the mathematical model sum[abs(T(i)-N(i)]-E(σ)) (Equation 3), where E(σ) is an empirically estimated measure of error rate. where T is the median genomic window depth indexed by {i}, representing the tumor depth, normalized by the stable z-score or stable PCA method; N is the median genomic window depth indexed by {i}, representing the normal depth coverage normalized by the stable z-score or stable PCA method compared to a cohort of normal samples; (I) is the estimated tumor for CNV or SV (eTF[CNV]) = (Sumi[(P(i)-N(i)-N(i)]*sign[T(i)]-E(σ)] / (sumi[abs[T(i)-N(i)]]-E(σ)]-E(σ)) (Equation 4); Fraction (J) comparing the eTF[CNV] to a detection threshold comprising empirically measured baseline noisy TF estimates from healthy samples, indicating a positive detection if the eTF[CNV] is higher than the threshold level (e.g., 2 standard deviations of the noisy TF distribution (FPR<2.5%)); and (K) diagnosing residual disease in the subject based on the eTF.
[0035] In some embodiments, the present disclosure relates to a system for detecting residual disease in a subject in need thereof, comprising: (A) an analysis unit configured and arranged to filter artificial noise markers from a genome-wide marker list, wherein the genome-wide marker list is generated from a plurality of genetic markers from biological samples of the subject, the biological samples comprising a tumor sample and a normal cell sample, and wherein the genetic marker list is selected from the group consisting of single nucleotide variations (SNVs), indels, copy number variations, SVs, and combinations thereof; and wherein the analysis unit further comprises detecting a subject-specific genome-wide list of genetic markers in a second biological sample comprising a plasma sample of the subject to generate a tumor genome list, wherein the analysis unit further comprises an SNV and indel classification engine, a CNV and SV classification engine, and a combination thereof. (B) an engine selected from the group consisting of a classification engine, a classification engine, and a combination thereof, wherein the SNV and indel classification engine statistically classifies each SNV in the list as signal or noise as a function of 1) the mapping quality (MQ) of the reads comprising the SNV or indel, 2) the fragment size length of the reads containing the SNV or indel, 3) a consensus test within the read duplication family containing the particular SNV, and 4) the base quality (BQ) of the SNV or indel, and the CNV and SV classification engine statistically classifies each CNV or SV window in the list as signal or noise based on 1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and 3) the list of CNVs or SV windows in the cfDNA data; and (B) a putative tumor of the sample based on one or more integrative mathematical models. Fraction (C) a display unit configured and arranged to calculate (eTF), and (D) a disease profile of the subject based on the estimated tumor fraction.
[0036] In some embodiments of the disclosed system, the eTF unit is further configured and arranged to integrate the probabilistic model to calculate eTFs for SNV or indel markers using a stochastic mixture model including: 1) an integrated signal of plasma SNV or indel detection; 2) process quality metrics including estimated genome coverage and a sequencing noise model; 3) patient-specific parameters including mutation burden (N); and / or 1) an integration of the directional depth of skewed coverage between plasma and normal patient samples, consistent with tumor CNV or SV direction in which copy number amplifications are positively skewed and copy number deletions are negatively skewed; 2) an integration of the cumulative depth of skewed coverage between tumor and normal patient samples; and 3) finding a dilution ratio between the above signals.
[0037] In some embodiments of the disclosed system, the tumor fraction estimation unit (B) comprises a processor, wherein the processor is configured to execute computer-readable instructions, which, when executed, calculates the tumor fraction estimation unit (B) by using the following integrated mathematical model: (1) eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific SNVs detected in a patient plasma sample, σ is an empirically estimated measure of error rate, R is the total number of distinct reads in the SNV enumeration region of interest (ROI) of the subject of interest, and N is the tumor mutation burden; and / or (2) eTF[CNV]=(sum__{i(P(i)-N(i)]*sign)]*T(i)-N(i)]-E(sigma) / (sum_{i}[abs(T(i)-N(i)]]-E(σ)), where where P is the median genomic window depth coverage indexed by {i}, representing plasma depth coverage, normalized by either stable z-score or stable PCA compared to a cohort of normal samples; T is the median genomic window depth indexed by {i}, representing tumor depth coverage, normalized by either stable z-score or stable PCA compared to a cohort of normal samples; N is the median depth indexed by {i}, relative to a cohort of normal samples, normalized by either stable z-score or stable PCA, and {i} is an individual indexation that counts all genomic windows that cover the patient's tumor-specific amplified and deleted genomic segments.
[0038] In some embodiments, the disclosure provides a method for detecting residual disease, or a computer-readable medium comprising computer-executable instructions that cause a processor to execute a series of steps, including: (A) receiving a subject-specific genome-wide list of somatic genetic markers derived from a plurality of genetic markers from a biological sample of a subject, the biological sample comprising a tumor sample and a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list in a second biological sample of the subject, and generating a tumor-associated genome-wide list of genetic markers in the second sample; and (C) filtering artificial noise markers from the genome, wherein the noise (PQ) is calculated as a function of: 1) mapping quality (MQ) of reads containing SNVs; 2) fragment length of reads containing SNVs; 3) consensus testing within read duplication families containing SNVs or Indels; and 4) base quality (BQ) of the SNVs or Indels. N (D) filtering the SNVs or Indels by statistically classifying each SNV or Indel as signal or noise based on its probability of detection in the SV window and / or by statistically classifying each SNV or Indel as signal or noise based on 1) its location relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and 3) overlap with a cfDNA mask (blacklist); (C) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrative mathematical models; and (E) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by the background noise model.
[0039] The present disclosure further relates to a cancer stratification method comprising detecting minimal residual disease (MRD) in cancer patients, the method comprising identifying low-abundance MRD-specific markers according to the above-described methods and detecting markers diagnostic of MRD. The cancer stratification method may further comprise detecting tumors by methods such as RT-PCR and / or molecular imaging using probes for lung cancer-specific markers. The details of one or more embodiments of the disclosure are set forth in the accompanying drawings / tables and the description below. Other features, objects, and advantages of the disclosure will be apparent from the drawings / tables and detailed description, and from the claims. [Brief explanation of the drawings]
[0040] [Figure 1] A is a schematic diagram of various embodiments of the disclosed diagnostic methods, e.g., detecting minimal residual tumor disease. B shows a representative workflow for detecting residual disease in a subject, according to various embodiments. C shows a representative workflow for detecting residual disease in a subject, according to various embodiments. D shows a representative workflow for diagnosing minimal residual disease (MRD) in a subject based on the measurement of single nucleotide polymorphisms or indels, according to the disclosed methods. E shows a representative workflow for diagnosing minimal residual disease (MRD) in a subject based on the measurement of copy number variations or structural variations, according to the disclosed methods.
[0041] [Figure 2] Figures A-B show charts of detection probability based on extrinsic or intrinsic parameters. Figure A shows the detection probability for various tumor fractions and coverages (up to the genome equivalent limit: ∼1000 molecules) based on the Bernoulli model. Figure B shows the detection probability for genome-wide SNV integration (binomial model) assuming an integration of 20,000 point mutations.
[0042] [Figure 3]Figures A-K show the effect of applying various filters according to various embodiments and the tumor fraction estimates provided by the method. Figure A shows the effect of applying a base quality (BQ) filter. Figure B shows the effect of optimizing base quality filtering by receiver operating curve (ROC). Figure C shows the effect of applying a combined base quality (BQ) and mapping quality (MQ) optimized filter when assessing the error rate distribution across multiple replicates using a control sample, which suppresses sequencing errors by approximately 7-fold change (FC). Pre-filter noise is ~2x10-3 for both lung cancer and melanoma, while post-filter noise is reduced to ~2x10-4 for both. Figure D shows the effect of applying a combined base quality (BQ) and mapping quality (MQ) optimized filter with relaxed 35x coverage. This filter allows marker detection in samples with TF as low as 1 / 20,000. The red line represents the theoretical (binomial model) expected value, the empirical measurements are shown in black (mean & confidence interval of 5 independent replicates), and the noise level is represented by the gray area due to the detection distribution of TF = 0. E shows in silico validation of TF estimation for melanoma samples. TFs estimated from the input mixture TFs (x-axis) versus mutation patterns (y-axis) showed a high correlation (R2 = 0.999). Accurate and specific estimates were obtained for all TFs above 5 × 10 −5 . F and G show diagnostic methods according to various embodiments, enabling the detection of genetic biomarker signatures for other types of solid tumors, such as lung tumor fraction (F) and breast cancer patients (G), even at tumor fractions as low as 1 / 10,000 of the tumor fraction (TF). H shows tumor fraction estimation based on low-confidence sSNVs with a tumor fraction (TF) of 5 × 10 −5 . I shows tumor fraction estimation based on reliable sCNVs with a tumor fraction (TF) of 5 × 10 −5 , preferably TF > 10 −4 . J indicates a strong correlation between TF estimation using SNV-based estimation (x-axis) and CNV-based estimation (y-axis). The gray quadrants indicate that the correlation between SNV-based estimation and CNV-based estimation becomes weaker when TFs fall below the 5 × 10-5 threshold. K shows a boxplot comparing our method with the ICHOR-CNA method.
[0043] [Figure 4] 1 shows SNV detection rates in a background noise model (healthy PBMC and cfDNA samples) for cfDNA samples collected from two cancer patients (BB1122, BB1125) before (pre-op) and after (post-op) resection surgery and two healthy control cfDNA samples (BB600 and BB601), according to various embodiments.
[0044] [Figure 5] Figures A and B show clinical evaluation of patient samples using the disclosed system and method. Figure A shows an exemplary evaluation of the disclosed system and method using clinical samples obtained from subjects with early-stage lung cancer and / or minimal residual disease (MRD), according to various embodiments. The data show tumor fraction (TF) estimates for pre- and post-operative plasma samples from all patients analyzed. Only two cases showed post-operative TFs above the noise threshold of 5 x 10-5. However, all healthy control samples had TFs below the detection threshold. "ND" indicates non-detection. The data were consistent with the results of the SNV method for plasma detection and TF correlation. Figure B shows the calculation of z-scores for 11 samples obtained from adenocarcinoma patients. The data show that the z-scores for healthy controls are below the threshold level (e.g., a z-score of 2, indicated by the horizontal dotted line). Figure C shows the calculation of z-scores for 11 samples obtained from adenocarcinoma patients compared to cross-patient negative controls. The data show that the z-scores of healthy controls are below the threshold level (e.g., a z-score of 2, indicated by the horizontal dotted line). Concordance between the sSNV-based and sCNV-based detection methods was observed (D).
[0045] [Figure 6]Panels A–E show an analytical approach to integrate multiple directional depth coverage distortions across large genomic CNV segments. Panel A shows the integration of sparse CNV skew at TF = 0.001. The top panel shows a comparison of single-bp depth coverage between synthetic plasma (TF = 10-3) and matched PBMCs in a 10-kbp segment of amplification. The middle panel shows the residuals between plasma and PBMCs, and the bottom panel shows the sum of the residuals. Note the sparse but positive bias in the residuals in the middle panel, and the accumulation of signal in the bottom panel, where the sum of the residuals is amplified and integrated into the genome, in part due to the positive bias in amplification. Panel B shows profiles of tumor read depth (red), germline read depth (pink), and pre-operative plasma cfDNA read depth (blue) in a representative amplified segment. Pre-operative plasma exhibits read depth comparable to germline DNA but also exhibits amplification depth skew at the telomeric end of the amplified segment. The mathematical method integrates read depth skew across the genome as described. C shows the signal-to-noise (SNR) for each TF, where all TFs above 10-6 exhibit positive (>0) SNR detection (indicating high sensitivity). D shows that CNV plasma SNR is linear with respect to TF (dilution model), exhibiting similar dynamics for lung, melanoma, and breast patients. E shows a chart of skew versus tumor fraction (TF) when neutral regions of the genome (e.g., regions without amplifications and / or deletions) are sampled. Thus, in these regions, the depth coverage skew between plasma and PBMCs is unbiased, and the probability of positive and negative skew is similar. Therefore, regardless of the TF (x-axis), there is no signal and SNR = 0.
[0046] [Figure 7] 1A-C provide schematic diagrams of the systems of the present disclosure, according to various embodiments.
[0047] [Figure 8] 1 provides an exemplary flow chart outlining the identification and / or classification of post-operative cancer subjects as candidates for adjuvant therapy, according to various embodiments.
[0048] [Figure 9] A comparison of patient-specific sSNV integration of various embodiments herein with ICHOR (Broad Institute) is shown. In particular, the detection sensitivity is increased by approximately 100-fold compared to the MIT-Broad Institute's ICHOR detection method.
[0049] [Figure 10] Figures A-E show the use of orthogonal features, such as fragment size, in the diagnostic methods of the present disclosure and the concomitant effects of applying such orthogonal features in SNV-based methods. Figure A shows the fragment size distribution exhibited by healthy normal cfDNA samples. Figure B shows the fragment size shift in breast tumor cfDNA (red and purple) compared to normal cfDNA samples. Figure C shows that in a mouse xenograft (PDX) model, tumor-derived circulating DNA is significantly shorter than normal-derived circulating DNA. Figure D shows a line graph of fragment DNA size (x-axis; number of bases) plotted against the frequency of observing fragments of that length across tumor and normal samples. Figure E shows patient-specific mutation detection using orthogonal features, such as the correspondence between DNA fragments and tumor origin, based on fragment size distribution (x-axis) and GMM-bound log-odds ratio (y-axis).
[0050] [Figure 11]Figures A and J show the use of orthogonal features such as fragment size in the diagnostic methods of the present disclosure, and the concomitant effects of applying such orthogonal features in CNV-based methods. Figure A shows line graphs of genomic region (bp) versus cumulative plasma depth coverage skew (lower panel), plasma versus vertical depth coverage skew (middle panel), and coverage (upper panel). Figure B shows the relationship between the log2 of depth coverage (log2 > 0.5 = amplification, log2 < -0.5 = deletion) and the local fragment size center of mass (COM) for that segment. Figure C shows the relationship between depth coverage-based CNV detection and fragment size center of mass-based CNV detection in patient samples. Figure D shows the lack of relationship between depth coverage-based CNV detection and fragment size center of mass (COM)-based CNV detection in normal (healthy) plasma samples. Figures E and F show the change in COM, absolute slope value, and R2 for two patients during treatment. Values are shown for baseline (day 0), 21 days after treatment, and 42 days after treatment. G shows the relationship between the slope of the log2 fragment size of a patient and tumor fraction. H shows the results of a clinical study of cancer patients examining the association between recurrence-free time and post-surgery (2 weeks post-surgery) tumor DNA detection (z-score). I shows a bar graph of tumor fraction for four patients at baseline (day 0), midpoint (day 21), and end (day 42) of treatment. J shows a bar graph of normalized CNV score for four patients at baseline (day 0), midpoint (day 21), and end (day 42) of treatment. DETAILED DESCRIPTION OF THE INVENTION
[0051] The following description of various embodiments is exemplary and explanatory only and should not be construed as limiting or restrictive in any sense. Other embodiments, features, objects, and advantages of the present teachings will be apparent from the description and accompanying drawings, and from the claims.
[0052] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those of ordinary skill in the art. The terms used in describing the disclosure herein are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. Furthermore, unless the context otherwise requires, singular terms include plural terms and plural terms include singular terms. In general, the nomenclature utilized in connection with molecular biology and the protein and oligo- or polynucleotide chemistry and hybridization described herein is well known and commonly used in the art. Standard techniques are used, for example, for nucleic acid purification and preparation, chemical analysis, recombinant nucleic acid, and oligonucleotide synthesis. Enzymatic reactions and purification techniques are performed according to manufacturer's specifications or as commonly accomplished in the art or as described herein. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and described in the various general and more specific references cited and discussed throughout this specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). The nomenclature used in connection with the experimental procedures and techniques described herein is well known and commonly used in the art.
[0053] Various embodiments of the present disclosure are described in further detail in the following paragraphs.
[0054] As used in the description of this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, "and / or" encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combination in the alternative ("or") interpretation.
[0055] The term "about" refers to a range of plus or minus 10% of the value; for example, "about 5" means 4.5 to 5.5, "about 100" means 90 to 100, etc., unless the context of the disclosure dictates otherwise; for example, in a list of values such as "about 49, about 50, about 55," "about 50" refers to a range of less than half the interval between the preceding and succeeding values, e.g., greater than 49.5 or less than 52.5. Furthermore, the terms "less than about" or "greater than about" should be understood in light of the definition of the term "about" provided herein.
[0056] When a range of values is provided in this disclosure, each intervening value between the upper and lower limits of that range, and any other stated or intervening value within that stated range, is intended to be included within the scope of the disclosure. For example, if a range of 1 μM to 8 μM is recited, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, and 7 μM are also intended to be expressly disclosed.
[0057] As used herein, the term "plurality" can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.
[0058] As used herein, the term "detecting" refers to the process of determining a value or set of values associated with a sample by measuring one or more parameters in the sample, and may further include comparing the test sample to a reference sample. According to the present disclosure, detecting a tumor includes identifying, assaying, measuring, and / or quantifying one or more markers.
[0059] As used herein, the term "diagnosis" refers to a method that can determine whether a subject is likely to suffer from a given disease or condition, including, but not limited to, diseases or conditions characterized by genetic mutations. Those skilled in the art often base their diagnosis on one or more diagnostic indicators, e.g., markers, the presence, absence, amount, or change in amount, the amount of which indicates the presence, severity, or absence of a disease or condition. Other diagnostic indicators include a patient's medical history, physical symptoms (e.g., unexplained weight loss, fever, fatigue, pain, or skin abnormalities), phenotype, genotype, or environmental or genetic factors. Those skilled in the art will understand that the term "diagnosis" refers to an increased likelihood of a particular course or outcome, i.e., an increased likelihood of a course or outcome in patients who exhibit a given characteristic, e.g., the presence or level of a diagnostic indicator, compared to individuals who do not exhibit that characteristic. The diagnostic methods of the present disclosure can be used independently or in combination with other diagnostic methods to determine whether a course or outcome is more likely to occur in patients who exhibit a given characteristic.
[0060] The term "normal," when used in the context of "normal cells," refers to cells of an untransformed phenotype or cells exhibiting the morphology of non-transformed cells of the tissue type being examined (e.g., PBMCs). In some embodiments, a "normal sample," as used herein, includes a non-tumor sample, e.g., a saliva sample, a skin sample, a hair sample, etc. It should be noted that the methods of the present disclosure can be performed without the use of a normal sample.
[0061] The term "abnormality," as used herein, generally refers to a state of a biological system that deviates to some degree from normal (e.g., wild-type). An abnormal state can occur at the physiological or molecular level. Representative examples include, for example, physiological conditions (diseases, pathologies) or genetic abnormalities (mutations, single nucleotide variants, copy number variants, gene fusions, indels, etc.). A pathological condition can be cancer or a precancerous condition. An abnormal biological state can be associated with some degree of abnormality (e.g., a quantitative measure indicating distance from a normal state).
[0062] The term "likelihood," as used herein, generally refers to probability, relative probability, presence or absence, or degree.
[0063] The term "tumor" as used herein includes any cell or tissue that may have undergone transformation at a genetic, cellular, or physiological level compared to normal or wild-type cells. The term generally refers to a neoplastic growth that may be benign (e.g., a tumor that does not form metastases and destroys adjacent normal tissue) or malignant / cancerous (e.g., a tumor that invades surrounding tissue and can usually give rise to metastases), and may be lethal to the host unless properly treated. Steadman's Medical Dictionary, 28 th Ed Williams & Wilkins, Baltimore, MD (2005).
[0064] The term "cancer" (used interchangeably with "tumor") refers to human cancers and carcinomas, sarcomas, adenocarcinomas, lymphomas, leukemias, solid and lymphatic cancers, etc. Examples of various types of cancer include, but are not limited to, lung cancer, pancreatic cancer, breast cancer, stomach cancer, bladder cancer, oral cancer, ovarian cancer, thyroid cancer, prostate cancer, uterine cancer, testicular cancer, neuroblastoma, squamous cell carcinoma of the head, cervix, cervix and vagina, multiple myeloma, soft tissue and osteogenic sarcoma, colon cancer, colorectal cancer, renal cancer (e.g., RCC), pleural cancer, cervical cancer, anal cancer, bile duct cancer, gastrointestinal carcinoid tumor, esophageal cancer, gallbladder cancer, small intestine cancer, central nervous system cancer, skin cancer, choriocarcinoma; osteogenic sarcoma, fibrosarcoma, glioma, melanoma, etc. In certain aspects, "liquid" cancers, e.g., blood cancers, e.g., lymphoma and / or leukemia, are excluded.
[0065] Examples of cancer include adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, anorectal cancer, anal canal cancer, appendix cancer, childhood cerebellar astrocytoma, childhood cerebral astrocytoma, basal cell carcinoma, skin cancer (non-melanoma), biliary tract cancer, extrahepatic bile duct cancer, intrahepatic bile duct cancer, bladder cancer, bone and joint cancer, osteosarcoma and malignant fibrous histiocytoma, brain cancer, brain tumor, brain glioma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive extraneural tumor, visual pathway and hypothalamic glioma, breast cancer, bronchial adenoma / carcinoid, carcinoid, gastrointestinal cancer, nervous system cancer, nervous system lymphoma lymphoma, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myeloproliferative disorders, colon cancer, colorectal cancer, cutaneous T-cell lymphoma, lymphoma, mycosis fungoides, Thesia syndrome, endometrial cancer of the esophagus, extracranial germ cell tumor, cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, eye cancer, intraocular melanoma, retinoblastoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid, gastrointestinal stromal tumor (GIST), germ cell tumor, ovarian germ cell tumor, gestational trophoblastic tumor glioma, head and neck cancer, hepatocellular (liver) cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, eye cancer, pancreatic islet cancer (endocrine pancreas), Kaposi's sarcoma, Renal cancer, renal cancer, laryngeal cancer, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, hairy cell leukemia, lip and oral cavity cancer, liver cancer, lung cancer, non-small cell lung cancer, AIDS-related lymphoma, non-Hodgkin's lymphoma, primary central nervous system lymphoma, Waldenstrom's macroglobulinemia, medulloblastoma, melanoma, intraocular melanoma, Merkel cell carcinoma, malignant mesothelioma, mesothelioma, metastatic squamous cell carcinoma, oral cavity cancer, tongue cancer, multiple endocrine neoplasia, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative disorder, chronic myeloid leukemia, acute myeloid leukemia disease, multiple myeloma, chronic myeloproliferative disorders, nasopharyngeal cancer, neuroblastoma, oral cavity cancer, oral cancer, oropharyngeal cancer, ovarian cancer, ovarian epithelial cancer, ovarian low malignant potential tumor, pancreatic cancer, pancreatic islet cell cancer, cancer of the paranasal sinuses and nasal cavity, parathyroid cancer, pharyngeal cancer, pheochromocytoma, pineoblastoma and supratentorial primitive neuroectodermal tumor, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma, prostate cancer, rectal cancer, renal pelvis and ureter cancer, transitional cell carcinoma, retinoblastoma, salivary gland cancer, Ewing's sarcoma, Kaposi's sarcoma, uterine cancer, uterine sarcoma, skin cancer (non-melanoma), skin cancer, Merkel cell carcinoma, small intestine cancer,These include, but are not limited to, soft tissue sarcoma, squamous cell carcinoma, gastric cancer, supratentorial primitive neuroectodermal tumor, testicular cancer, thymoma, thymic carcinoma, thyroid cancer, transitional cell carcinoma, renal pelvis and ureter and other urinary tract cancers, gestational trophoblastic tumor, urethral cancer, endometrial cancer, uterine sarcoma, uterine cancer, vaginal cancer, vulvar cancer, and Wilms' tumor.
[0066] As used herein, the term "non-small cell lung cancer" or NSCLC refers to all lung cancers other than small cell lung cancer, including several subtypes, including, but not limited to, large cell carcinoma, squamous cell carcinoma, and adenocarcinoma, and includes all stages and metastases. Squamous cell carcinoma, which accounts for 25% of lung cancers, usually begins near the central bronchi. Tumors usually contain cavities and associated necrosis in the center. Well-differentiated squamous cell carcinomas often grow more slowly than other types of cancer. Adenocarcinoma accounts for 40% of non-small cell lung cancers and usually arises in peripheral lung tissue. While most cases of adenocarcinoma are associated with smoking, it is the most common type of lung cancer among never-smokers. See Rosell et al., Lung Cancer, 46(2), 135-48, 2004; Coate et al., Lancet Oncol, 10, 1001-10, 2009.
[0067] As used herein, the term "residual disease" refers to the persistence of remaining neoplastic cells even after intervention, such as surgical intervention, radiological resection, chemotherapy, etc., and the term "minimal residual disease (MRD)" refers to the situation in which morphologically normal tissue (e.g., lung tissue) may still harbor a significant amount of residual malignant cells after tumor treatment (e.g., chemotherapy, immunotherapy, or targeted therapy). Detection of minimal residual disease (MRD) is a novel and practical means to more accurately measure remission induction during treatment. In the context of liquid tumors (e.g., lymphoma or myeloma), the term MRD refers to the persistence of neoplastic cells after tumor treatment (e.g., chemotherapy, immunotherapy, or targeted therapy). -4 Less than, e.g., 10 -5 Less than or equal to 10 -6In the context of solid tumors, the term "minimal residual disease" can refer to a situation in which tumor markers are below what can be detected using conventional detection means, such as ctDNA detection or plasma DNA analysis. In some embodiments, MRD refers to a situation in which less than 100 copies, preferably less than 40 copies, and particularly less than 10 copies of ctDNA are detected per 5 ml of plasma (Bettegowda et al., Sci Transl Med., 6(224), 224ra24, 2014).
[0068] As used herein, the term "subject" refers to a mammal, including humans, veterinary or farm animals, livestock or pets, and animals commonly used in clinical research. In particular, the subject is a human subject, e.g., a human patient diagnosed with or suspected of having a tumor. The subject can have, potentially have, or suspected of having one or more characteristics selected from cancer, a cancer-related symptom, or be asymptomatic or undiagnosed (e.g., have not been diagnosed with cancer) for cancer. The subject can have cancer, the subject can exhibit cancer-related symptoms, the subject can be free of cancer-related symptoms, or the subject can not be diagnosed with cancer. In some embodiments, the subject is a human.
[0069] As used herein, the term "single nucleotide polymorphism" or "single nucleotide variation" ("SNP" or "SNV"), with respect to a mutation, refers to a difference of at least one nucleotide in a sequence compared to another sequence.
[0070] The term "copy number variation" or "CNV" refers to a comparative numerical change in the presence / absence / insertion or deletion of a gene segment with identical nucleotide sequence. In the human genome, copy number variants can include homozygous or heterozygous duplications or multiplications of one or more segments of DNA, or homozygous or heterozygous deletions of one or more segments of DNA. The directionality of a CNV is usually designated positive for a CNV duplication / multiplication and negative for a CNV deletion.
[0071] As used herein, the term "indel" refers to a genomic location where one or more bases are present in one allele and no bases are present in the other allele. Although insertions and deletions are distinct from an evolutionary perspective, the analyses described herein often do not distinguish between an insertion in one allele as equivalent to a deletion in the other allele. Thus, the term "indel" refers to the location of an insertion / deletion between two alleles.
[0072] The term "structural variant" as used herein refers to a change in some portion of a chromosome instead of a change in the number of chromosomes or sets of chromosomes in the genome. There are four general types of mutations that result in structural variations: deletions and insertions, e.g., duplications (changes in the amount of DNA in a chromosome, deletions and gains of genetic material), inversions (changes in the arrangement of chromosomal fragments), and translocations (changes in the position of chromosomal fragments that can result in gene fusions). The term "structural variant" of the present invention includes losses of genetic material, gains of genetic material, translocations, gene fusions, and combinations thereof.
[0073] As used herein, the term "sample" refers to a composition obtained or derived from a subject, including cells and / or other molecular entities to be characterized and / or identified, e.g., based on physical, biochemical, chemical, and / or physiological characteristics. Preferably, the sample is a "biological sample," meaning, for example, cells, tissues, organs, or other samples of biological origin. In certain embodiments, the source of a tissue sample can be blood or any blood component; body fluids; fresh, frozen, and / or preserved organ or tissue samples, or solid tissue from biopsies or aspirates; and cells from any point during pregnancy or development of a subject or plasma. Samples include, but are not limited to, primary cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous humor, ocular fluid, lymphatic fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, urine, cerebrospinal fluid (CSF), saliva, sputum, tears, sweat, mucus, tumor lysates, and tissue culture media, as well as tissue extracts such as homogenized tissues, tumor tissues, and cell extracts. Samples also include biological samples that have been manipulated in some way after their procurement, e.g., by reagents, solubilization, or enrichment for certain components, such as proteins or nucleic acids, or by embedding in semi-solid or solid matrices for sectioning, such as thin tissue sections or cells in histological samples. Samples may also include environmental components, such as water, soil, mud, air, resins, minerals, and the like. In certain embodiments, the sample may include a biological sample containing DNA (e.g., gDNA), RNA (e.g., mRNA, tRNA), protein, or a combination thereof obtained from a subject (e.g., a human or other mammalian subject).
[0074] As used herein, the term "cell" is used interchangeably with "biological cell." Non-limiting examples of biological cells include animal cells such as eukaryotic cells, plant cells, mammalian cells, reptilian cells, avian cells, and fish cells; prokaryotic cells, bacterial cells, fungal cells, and protozoan cells; cells dissociated from tissues such as muscle, cartilage, fat, skin, liver, lung, and nervous tissue; immunological cells such as T cells, B cells, natural killer cells, and macrophages; embryos (e.g., zygotes), oocytes, eggs, sperm cells, hybridomas, cultured cells, cells derived from cell lines, cancer cells, infected cells, transfected and / or transformed cells, reporter cells, and the like. Mammalian cells can be obtained, for example, from humans, mice, rats, horses, goats, sheep, cattle, primates, and the like.
[0075] As used herein, the term "marker" refers to a characteristic that can be objectively measured as an indicator of normal biological processes, pathogenic processes, or therapeutic intervention, such as the pharmacological response to anti-cancer drug treatment. Representative types of markers include molecular changes in the structure (e.g., sequence) or number of markers, including multiple differences such as gene mutation, gene duplication, or somatic mutation of cfDNA, copy number variation, tandem repeat, or combinations thereof.
[0076] The term "genetic marker" as used herein refers to a sequence of DNA having a specific location on a chromosome that can be measured in a laboratory. The term "genetic marker" can also be used to refer, for example, to cDNA and / or mRNA encoded by a genomic sequence, as well as the genomic sequence itself. A genetic marker can include two or more alleles or variants. A genetic marker can be a direct marker (e.g., a marker located within a subject gene or subject locus (e.g., a candidate gene)) or an indirect marker (e.g., a marker closely associated with a subject gene or subject locus because it is adjacent to the subject gene or subject locus but not within the subject gene or subject locus). Furthermore, a genetic marker can also be unrelated to a gene or locus that resides in a non-coding region of the genome, such as an SNV, CNV, indel, SV, or tandem repeat. A genetic marker includes nucleic acid sequences that may or may not encode a gene product (e.g., a protein). In particular, the genetic markers include single nucleotide polymorphisms / mutations (SNPs / SNVs) or copy number variations (CNVs) or combinations thereof. Preferably, the genetic markers include somatic mutations in DNA, such as sSNVs or sCNVs, indels, SVs or combinations thereof compared to a reference sample.
[0077] As used herein, the term "cell-free DNA" or "cfDNA" refers to cell-free deoxyribose nucleic acid (DNA) strands, such as those extracted or isolated from circulating blood plasma / serum, lymph, cerebrospinal fluid (CSF), urine, or other bodily fluids. The term "cfDNA" is in contrast to "circulating tumor DNA" or "ctDNA." Cell-free DNA (cfDNA) is a broader term that describes DNA that circulates freely in the bloodstream, but is not necessarily derived from a tumor.
[0078] As used herein, the term "germline DNA" or "gDNA" refers to DNA isolated or extracted from a patient's peripheral mononuclear cells, including lymphocytes, which in turn are obtained from circulating blood.
[0079] The term "mutation" as used herein refers to a change or deviation. With respect to nucleic acids, mutation refers to a difference(s) or alteration between DNA nucleotide sequences, including copy number differences (CNVs). This actual difference in nucleotides between DNA sequences can be SNPs and / or changes in the DNA sequence, such as fusions, deletions, additions, repeats, etc., observed when comparing the sequence with a reference, such as germline DNA (gDNA) or the reference human genome HG38 sequence. Preferably, mutation refers to a difference between a cfDNA sequence and a control DNA sequence not derived from tumor cells, such as when cfDNA is compared with the reference HG38 sequence; when cfDNA is compared with gDNA. Differences identified in both gDNA and cfDNA may be considered "constitutional" and disregarded.
[0080] The term "control," as used herein, refers to a reference for a test sample, such as control DNA isolated from peripheral blood mononuclear cells and lymphocytes (where the cells are not cancer cells), and a "reference sample" refers to a sample of tissue or cells that may or may not have cancer, used for comparison. Thus, a "reference" sample provides a basis against which another sample, such as a plasma sample containing cfDNA, can be compared. In contrast, a "test sample" refers to a sample that is compared to a reference or control sample. The reference sample does not need to be cancer-free, as would be the case if the reference and test samples were obtained from the same patient separated in time.
[0081] In some embodiments, a reference sample or control may comprise a reference assembly. The term "reference assembly" refers to a digital nucleic acid sequence database, such as the Human Genome (HG38) database (assembled December 2013), which includes the HG38 assembly sequence. The gateway may be accessed via the Human (Homo sapiens) University of California, Santa Cruz (UCSC) Genome Browser Gateway at the world-wide web URL GENOME(dot)UCSC(dot)EDU at GENOME(dot)UCSC(dot)EDU. Alternatively, the reference assembly may refer to the Genome Reference Consortium's Human Genome Assembly (Build #38; assembled June 2017), accessible on the internet via the National Center for Biotechnology Information (NCBI) web site.
[0082] As used herein, the term "sequencing" or "sequencing" as a verb refers to the process by which the nucleotide sequence of DNA, or the order of nucleotides, is determined, such as the nucleotide order AGTCC. The term "sequence" as a noun refers to the actual nucleotide sequence obtained from sequencing, e.g., DNA having the sequence AGTCC. While "sequencing" may be provided and / or received in digital form, e.g., on a disk or remotely via a server, "sequencing" refers to a collection of DNA that is grown, manipulated, and / or analyzed using the methods and / or systems of the present disclosure.
[0083] The term "DNA sequence," as used herein, generally refers to a "raw sequence read" and / or a "consensus sequence." A raw sequence read is the output of a DNA sequencer and typically contains redundant sequences of the same parent molecule, e.g., after amplification. A "consensus sequence" is a sequence derived from overlapping sequences of parent molecules that is intended to represent the sequence of the original parent molecule. A consensus sequence can be generated by voting (where each majority nucleotide, e.g., the most commonly observed nucleotide at a given base position in a sequence, is the consensus nucleotide) or other approaches, such as comparison to a reference genome. A consensus sequence can be generated by tagging the original parent molecule with a unique or non-unique molecular tag (e.g., a barcode) that allows tracking of the progeny sequence (e.g., after PCR).
[0084] The sequencing method can be a first-generation sequencing method such as Maxam-Gilbert or Sanger sequencing, or a high-throughput sequencing (e.g., next-generation sequencing or NGS) method. High-throughput sequencing methods can simultaneously (or substantially simultaneously) sequence at least 10,000, 100,000, 1 million, 10 million, 100 million, 1 billion, 1 billion, or more polynucleotide molecules. Sequencing methods can include, but are not limited to, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-hybridization, digital gene expression (Helicos), massively parallel sequencing (e.g., Helicos, clonal single molecule array (Solexa / Illumina)), sequencing using PACBIO, SOLID, Ion Torrent, or NANOPORE platforms.
[0085] The term "whole genome sequencing" refers to the experimental process of determining the DNA sequence of each DNA strand in a sample, and the resulting sequence can be referred to as "raw sequencing data" or "read". As used herein, a read is "mappable" if its sequence is similar to that of a region of a reference chromosomal DNA sequence. The term "mappable" refers to a region that shows similarity to a reference sequence and is therefore "mapped", for example, a segment of cfDNA that shows similarity to a reference sequence in a database, for example, a cfDNA that has a high proportion of similarity with the human chromosomal region 8q248q24.3 in the Human Genome (HG38) database is a "mappable read".
[0086] "Deep sequencing" refers to the general concept of taking multiple replicate reads of each region of a sequence.
[0087] As used herein, the term "mapping" generally refers to aligning a DNA sequence with a reference sequence based on sequence homology. Alignment can be performed using an alignment algorithm, such as the Needleman-Wunsch algorithm, BLAST, or EMBOSS.
[0088] In addition to "WGS," genome catalogs can be obtained using targeted sequencing. In contrast to WGS, the term "targeted sequencing," as used herein, refers to an experimental process of determining the DNA sequence of one or more selected DNA loci in a sample, e.g., determining the sequence of a selected group (e.g., targets) of cancer-related genes or markers. In this context, the term "target sequence" refers to a selected target polynucleotide, e.g., a sequence present in a cfDNA molecule whose presence, quantity, and / or nucleotide sequence, or alteration, is desired to be determined. The target sequence is examined for the presence or absence of somatic mutations. The target polynucleotide can be a region of a gene associated with a disease, e.g., cancer. In some embodiments, the region is an exon.
[0089] As used herein, the term "low abundance" with respect to cfDNA refers to less than about 20 ng / mL, e.g., about 15 ng / mL, about 10 ng / mL, or less, e.g., about 9 ng / mL, 8 ng / mL, 7 ng / mL, 6 ng / mL, 5 ng / mL, 4 ng / mL, 3 ng / mL, 2 ng / mL, 1 ng / mL, 0.7 ng / mL, 0.5 ng / mL, 0.3 ng / mL, or less, e.g., 0.1 ng / mL or 0.05 ng / mL. In some embodiments, the term "low abundance" can be understood in the context of the uniqueness of a marker, e.g., length or base composition. For example, a subject's sample may contain an abundant amount of cfDNA (e.g., >20 ng / mL), but the actual number of unique genetic markers (e.g., sSNVs, sCNVs, indels, SVs) contained in the cfDNA may be very low. Typically, this parameter is expressed as genome equivalent (GE) or coverage, as described below. In some embodiments, the term "low abundance" can be understood in the context of the tumor specificity of the marker. For example, a subject's sample may contain an abundant amount of cfDNA (e.g., >20 ng / mL), but most of the genetic markers (e.g., sSNVs, sCNVs, indels, SVs) contained in the cfDNA may be redundant and / or may be related to a reference (e.g., PBMC gDNA). Typically, this parameter is expressed as a tumor fraction, as described below.
[0090] As used herein, the terms "tumor-specific" or "tumor-associated" with respect to cfDNA refer to differences in the DNA sequence of cfDNA in a subject with cancer, such as a lung cancer patient, when the cfDNA is compared to reference DNA, such as when the cfDNA is compared to control DNA (gDNA) from cells that are not tumorous, as described herein.
[0091] As used herein, the term "read overlap family" includes PCR and sequencing overlaps. Generally, these are independent copies of the same unique fragment and can therefore be used in statistical tests (consensus tests) to correct for low-frequency PCR and sequencing errors.
[0092] The terms "coverage" or "read depth" refer to sequencing effort. For example, 20X coverage refers to a moderate sequencing effort, 35X or more coverage refers to a high sequencing effort, and 5X coverage refers to a low sequencing effort. In embodiments of the present disclosure, coverage is typically about 5X to about 100X, particularly 15X to about 40X, e.g., 20X, 30X, 35X, 40X, 50X, 70X, or more.
[0093] As used herein, "depth coverage" refers to the number of unique reads whose mapping overlaps at or on a particular genomic coordinate.
[0094] As used herein, the term " cfDNA coverage mask " refers to the mask that represents the genome region that is covered by cfDNA in normal cfDNA cohort.As known in the art, the coverage of cfDNA is not completely uniform (accessible chromatin genome region is not often represented), therefore, blacklist or mask can be implemented to remove bias and selectively analyze well-covered region.
[0095] The term "read mappability" as used herein relates to a numerical value (eg, percent identity) or statistical measure (eg, confidence estimate) of the accuracy of mapping of a read to the genome.
[0096] As used herein, the term "mutation burden" or "N" refers to the level, e.g., number, of alterations (e.g., one or more genetic alterations, particularly one or more somatic alterations) per preselected unit (e.g., per megabase pair) in a given genomic window. Mutational burden may be measured, for example, on a whole-genome or exome basis, or based on a subset of the genome or exome. In certain embodiments, a mutational burden measured based on a subset of the genome or exome may be extrapolated to determine a whole-genome or exome mutational burden. In certain embodiments, mutational burden is measured in a sample from a subject, e.g., a subject described herein, such as a tumor sample (e.g., a lung tumor sample, or an obtained or derived sample). Preferably, mutational burden is a measure of the number of mutations per megabase pair (1,000,000 bp or MBP) of cfDNA. As is known in the art, mutational burden may vary depending on tumor type, genetic lineage, and other subject-specific characteristics, such as age, sex, and tobacco consumption. For tumor diagnosis, the mutational burden can be about 1,000 to about 10,000 mutations per MBP, e.g., about 1,000, 2,000, 4,000, 6,000, 8,000, 10,000, 12,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 10,000, or more, e.g., about 200,000 mutations per MBP. Typically, the mutational burden is about 8,000 / MBP in non-smokers and greater than 40,000 / MBP in subjects with melanoma.
[0097] The term "genomic window," as used herein, refers to a region of DNA within selected nucleotide sequence boundaries. Windows may be separate from one another or may overlap one another.
[0098] The term "tumor fraction" or "TF" as used herein refers to the level, e.g., amount, of tumor DNA molecules relative to normal DNA molecules. In some embodiments, "tumor fraction" refers to the ratio of circulating cell-free tumor DNA (cfDNA) to the total amount of cell-free DNA. The tumor fraction is considered to indicate the size of the tumor. Typically, the tumor fraction (TF) is about 0.001% to about 1%, e.g., about 0.001%, 0.05%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1% or more, e.g., 2%.
[0099] The term "abundance" can be a binary (e.g., absent / present), qualitative (e.g., absent / low / medium / high), or quantitative (e.g., a value proportional to number, frequency, or concentration) value indicating the presence of a particular molecular species. In this context, mutations present at higher relative concentrations are associated with a greater number of malignant cells, e.g., cells transformed early in the tumorigenesis process, compared to other malignant cells in the body (Welch et al., Cell, 150:264-278, 2012). Due to their high relative abundance, such mutations are expected to have higher diagnostic sensitivity for detecting cancer DNA than mutations with lower relative abundance.
[0100] As used herein, "sequencing noise" refers to noise introduced by the sequencing instrument, software, or other artificial means during "run." There are at least two sources of noise in a sequencing pipeline. First, the DNA mixture generated from the input pellet (DNA or cell pellet) is a complex mixture of cells; therefore, any useful signal is diluted by DNA lacking information content. The second noise source results from the specific sequencing technology used. For example, sequencing noise or "machine" noise can derive from the ion-base sequencing process, e.g., the IONTORENT PGM™ platform. For example, ion-detection sequencing, which reads bases based on pH detection, is sensitive to homopolymers and may sometimes read homopolymer chains as one base too long or too short.
[0101] As used herein, "sequencing error rate" refers to the inaccurate proportion of sequenced nucleotides. For example, in the context of whole genome sequencing, sequencing error rates of approximately 1 / 1000 bases are reported in the literature (range: error rates on the order of 0.1-1% per base call; see Wu et al., Bioinformatics, 33(15):2322-2329, 2017).
[0102] The term " sequencing depth " used herein refers to the number of times that sequencing area is covered by sequence reading.For example, the average sequencing depth of 10 times means that each nucleotide in the sequencing area is covered by 10 sequence readings on average.It is expected that the greater the sequencing depth, the greater the probability of detecting cancer-related mutations.However, in reality, the probability of detection does not increase linearly with sequencing depth, as evidenced by the fact that even at a median depth of 42,000X, the fundamental limitations of cfDNA abundance only lead to the positive detection of early-stage lung adenocarcinoma at only 19% (Abbosh et al., Nature, 545(7655):446-451, 2017).
[0103] As used herein, the term "noise" in its broadest sense refers to unwanted disturbances (e.g., signals not directly related to a true event) that may nevertheless be processed or received as a true event. Noise is the sum of unwanted or disruptive energy introduced into a system from artificial and natural sources, and may distort a signal in such a way that the information carried by the signal is degraded or unreliable. Noise is in contrast to a "signal," which is a function that conveys information about the behavior or characteristics of some phenomenon, such as the probabilistic association between a marker (SNV, CNV, indel, SV) and a tumor.
[0104] The term "signal-to-noise ratio" as used herein refers to the ability to resolve true signal from system noise. The signal-to-noise ratio is calculated by obtaining the ratio of the level of the desired signal to the level of noise present in the signal. Phenomena that affect the signal-to-noise ratio include, for example, detector noise, system noise, and background artifacts. The term "detector noise" as used herein refers to unwanted disturbances (i.e., signals that are not directly attributable to the intended energy of the detector) that occur within the detector. Detector noise includes dark current noise and shot noise. Dark current noise in optical detector systems such as sequencers can arise from various thermal emissions from the photodetector. Shot noise in an optical system is the product of fundamental particle properties (i.e., Poisson distribution energy fluctuations) of incident photons as they pass through the photodetector.
[0105] The term "filter" is used in many ways by those skilled in the art to mean discarding or removing undesired data, retaining desired data, or both.
[0106] The term "base quality" (BQ) score relates to the confidence in the sequencing quality at each nucleotide base in a polynucleotide. In some embodiments, base quality (BQ) includes variable base quality (VBQ) or mean read base quality (MRBQ), both of which are variations of the base quality metric.
[0107] The term "mapping quality" (MQ) score relates to a confidence estimate of the accuracy of the mapping of the marker to the genome.
[0108] The term "read position" or "read position (PIR)" refers to the position of a read (e.g., a marker) in a nucleotide sequence. As understood in genomics, many sequencing protocols are prone to various types of amplification-induced bias and errors, which can be reduced by implementing filters such as the "read direction" and "read position" filters. The read direction filter removes variants that are present almost exclusively in either the forward or reverse read. In many sequencing protocols, such variants are most likely the result of amplification-induced errors. The read position filter is implemented in a similar manner to the "read direction filter" to remove systematic errors, but is also suitable for hybridization-based data. It removes variants located in a read that differ from those expected based on the general position of the read covering the mutation site. This is done by sorting each sequenced nucleotide (or gap) by the mapping direction of the read and where the nucleotide is found in the read; each read is divided into portions (e.g., 5 portions) along its length, and the portion number of the nucleotide is recorded. This results in a total of 10 categories for each sequenced nucleotide, and a given site will be distributed among these 10 categories for the reads that cover that site. If a variant is present at this site, the variant nucleotides are expected to follow the same distribution. The read position filter performs a test to measure the significance of the read position, for example, whether the read position distribution of the variant differs from that of the entire set of reads that cover the site.
[0109] As used herein, the term "location attribute" of a marker (e.g., a CNV) refers to the spatial location of the marker in a chromosome or gene sequence. For example, the location attribute of a marker can be determined based on whether it is at least 1000 kilobases (kb), at least 400 kb, at least 100 kb, or at least 20 kb or less, e.g., 1 kb from a telomere, centromere, or heterochromatic region of a chromosome. CNVs mapped to subtelomeric or pericentromeric regions characterized by hotspots of chromosomal rearrangements may be undesirable. As used herein, the term "representative" with respect to a marker (e.g., a CNV) refers to its association with a phenotype or disease. For example, previous studies have found that the calling of CNVs in the immunoglobulin region is not representative of gDNA and tends to depend substantially on the DNA source—e.g., saliva versus blood or lymphoblastoid cell lines versus blood—(Need et al., 2009; Wang et al., 2007; Sebat et al., 2004).
[0110] As used herein, the term "coverage" or "depth" in DNA sequencing refers to the number of reads that contain a given nucleotide in the reconstructed sequence. Coverage histograms are generally used to show the range and uniformity of sequencing coverage across a data set, and they show the overall coverage distribution by listing the number of reference bases covered by sequencing reads mapped at various depths. Mapped "read depth" refers to the total number of bases that are sequenced and aligned at a given reference base position. Typically, in sequencing coverage histograms, read depth is listed in bins on the x-axis, while the total number of reference bases that occupy each read depth bin is listed on the y-axis. These can also be described as the ratio of reference bases.
[0111] As used herein, "depth coverage" refers to the number of unique reads whose mapping overlaps with a particular genomic coordinate.
[0112] As used herein, the term "read mappability" refers to a confidence estimate of the accuracy of mapping of a read associated with a CNV to the genome.
[0113] As used herein, the term "unique read" refers to a distinctive feature, e.g., a read that occurs uniquely in a reference genome; in contrast, a "non-unique read" refers to a distinctive feature, e.g., a read that has no or very few occurrences (i.e., repeats) among the reads.
[0114] As used herein, a genomic "region of interest" or ROI can be any genomic region from which genetic information is desired. The genomic region of a subject of interest can include a region of a chromosome. The genomic region of interest can include an entire chromosome. A chromosome is a diploid chromosome. In the human genome, for example, a diploid chromosome can be any of chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. In some cases, the chromosome can be the X or Y chromosome. In some cases, the genomic region of interest includes a portion of a chromosome. The genomic region of a subject of interest can be of any length. The length of the subject genomic region can be, for example, about 1 to about 10 bases, about 5 to about 50 bases, about 10 to about 100 bases, about 70 to about 300 bases, about 200 to about 1000 bases (1 kb), about 700 to about 2000 bases, about 1 to about 10 kb, about 5 to about 50 kb, about 20 to about 100 kb, about 50 to about 500 kb, about 100 to about 2000 kb (2 Mb), about 1 Mb to about 50 Mb, about 10 to about 100 Mb, or about 50 to about 300 Mb. For example, the genomic region of interest may be greater than 1 base, greater than 10 bases, greater than 20 bases, greater than 50 bases, greater than 100 bases, greater than 200 bases, greater than 400 bases, greater than 600 bases, greater than 800 bases, greater than 1000 bases, greater than 1.5 kb, greater than 2 kb, greater than 3 kb, greater than 4 kb, greater than 5 kb, greater than 10 kb, greater than 20 kb, greater than 30 kb, greater than 40 kb, greater than 50 kb, greater than 60 kb, greater than 70 kb, greater than 80 kb, greater than 90 kb, greater than 100 kb, greater than 200 kb, The genomic region of interest may be greater than 1 kb, greater than 300 kb, greater than 400 kb, greater than 500 kb, greater than 600 kb, greater than 700 kb, greater than 800 kb, greater than 900 kb, greater than 1000 kb, greater than 1 Mb, greater than 2 Mb, greater than 3 Mb, greater than 4 Mb, greater than 5 Mb, greater than 6 Mb, greater than 8 Mb, greater than 9 Mb, greater than 10 Mb, greater than 20 Mb, greater than 30 Mb, greater than 40 Mb, greater than 50 Mb, greater than 60 Mb, greater than 70 Mb, greater than 80 Mb, greater than 90 Mb, greater than 100 Mb, or greater than 200 Mb. The genomic region of interest may comprise one or more informative loci. An informative locus may be, for example, a polymorphic locus comprising two or more alleles. In some cases, the two or more alleles constitute minor alleles.
[0115] As used herein, the term "directionality" in reference to reading refers to the direction or manner in which reading is performed. For example, in single-end reading, a sequencer reads a fragment from one end to the other to generate a base-paired sequence. In paired-end reading, a sequencer starts with one read, finishes this direction at a specified read length, and then starts the next read from the opposite end of the fragment. Paired-end reading improves the ability to identify the relative positions of various reads in the genome and is much more effective than single-end reading in unraveling structural rearrangements such as gene insertions, deletions, and inversions. It can also improve the cataloging of repetitive regions. However, paired-end reading is more expensive and time-consuming to perform than single-end reading.
[0116] As used herein, the term "CNV directionality" refers to the direction of copy number change. For example, copy number increases (e.g., expansions and amplifications) are positive values, while decreases (e.g., losses and fragmentations) are negative values.
[0117] As used herein, the term "bin" refers to a group of DNA sequences grouped together, such as a "genomic bin." In certain cases, a bin may include a group of DNA sequences binned based on a "genomic bin window," which involves grouping DNA sequences using a genomic window.
[0118] As used herein, the term "estimate" in relation to marker levels is used broadly, and the term "estimate" can be an actual value (e.g., 1 / mbp), a range of values, a statistical value (e.g., mean, median, etc.), or other means of estimation (e.g., probabilistic).
[0119] As used herein, "substantially" means sufficient to function for its intended purpose. Thus, the term "substantially" allows for small, insignificant variations from an absolute or perfect state, dimension, measurement, result, etc., that would be expected by one of ordinary skill in the art, but that do not affect overall performance. When used in reference to a number or a parameter or characteristic that can be expressed as a number, "substantially" means within 10%.
[0120] As used herein, the term "substantially purified" refers to cfDNA molecules that have been removed from their natural environment, isolated or separated or extracted, and are at least 60% free, preferably 75% free, more preferably 90% free, and most preferably 99% free from other components that are naturally associated with them.
[0121] All publications mentioned herein are incorporated by reference for the purpose of describing and disclosing the devices, compositions, formulations and methods that are described in the publications and that might be used in connection with the present disclosure.
[0122] As used herein, the terms "comprise," "include," "contain," "are," "have," and "comprise" are not intended to be limiting, inclusive, or open-ended, and do not exclude additional, unrecited additives, components, integers, elements, or method steps. For example, a process, method, system, composition, kit, or device that includes a list of features is not necessarily limited to those features and may include other features that are not expressly recited or inherent in the process, method, system, composition, kit, or device.
[0123] The practice of the subject matter may employ, unless otherwise indicated, conventional techniques and procedures of organic chemistry, molecular biology (including recombinant techniques), cell biology, and biochemistry, which are within the skill of the art.
[0124] 〔method〕 The present disclosure relates to methods and systems for the detection and / or diagnosis of residual tumor that analyze markers present in cell-free DNA (cfDNA), which, alone or in combination with existing techniques, can be used to determine the presence or absence of residual tumor, predict the likelihood of developing the disease, and develop therapeutic or preventative interventions for the disease.
[0125] In some embodiments, the methods of the present disclosure are performed on a sample obtained from a subject. Preferably, the sample comprises blood (including whole blood), plasma, blood serum, hemolysate, lymph, synovial fluid, spinal fluid, urine, cerebrospinal fluid, stool, sputum, mucus, amniotic fluid, tears, cyst fluid, sweat gland secretions, bile, milk, tears, saliva, or ear wax. The sample can be processed to remove specific cells using various methods, such as centrifugation, affinity chromatography (e.g., immunoabsorption), immunoselection, and filtering. Thus, in examples, the sample can contain a specific cell type or mixture of cell types isolated directly from the subject or purified from a sample obtained from the subject (e.g., purifying T cells from whole blood). In one example, the biological sample is peripheral blood mononuclear cells (PBMCs). In other examples, the sample may be selected from the group consisting of B cells, dendritic cells, granulocytes, innate lymphocytes (ILCs), megakaryocytes, monocytes / macrophages, natural killer (NK) cells, platelets, red blood cells (RBCs), T cells, thymocytes, etc. In certain embodiments, the sample may include skin cells, hair follicle cells, sperm, etc.
[0126] A representative, non-limiting outline of the diagnostic method is shown in FIGS.
[0127] [Workflow]
[0128] FIG. 1A is a flowchart illustrating a method 100 for detecting residual disease, e.g., post-surgery tumor disease or post-treatment (e.g., post-chemotherapy, immunotherapy, targeted therapy, radiation therapy), according to various embodiments of the present disclosure. Method 100 is merely exemplary, and embodiments may employ variations of method 100. Method 100 may include receiving a list of markers, filtering marker-related noise based on multiple features, and removing artificial noise markers from the list to generate subject-specific markers, which are then used to estimate tumor fraction for use in diagnosing residual disease. It should be noted that TF refers to the proportion of tumor DNA (ctDNA) in total plasma DNA (cfDNA). Thus, the term "ctDNA abundance" in this disclosure and elsewhere may be used synonymously with the term "tumor fraction."
[0129] In step 110 of method 100 of FIG. 1A, a subject-specific genome-wide list of associated genetic markers (e.g., SNVs, CNVs, SVs, indels) in a biological sample (tumor sample and, optionally, normal sample) is received from a subject. In some embodiments, the list of genetic markers is received in a variant call format (VCF) file. As understood in the art, VCF files are used in bioinformatics to store genetic sequence variations. The VCF format was developed with the advent of large-scale genotyping and DNA sequencing projects, such as the 1000 Genomes Project. Alternatively, the list can be provided in a general feature format that includes all of the genetic data. Generally, GFFs are shared genome-wide, providing redundant features. In contrast, VCF only stores variations along with the reference genome. In some embodiments, the subject's sample is sequenced, e.g., using whole genome sequencing (WGS), and the sequence file is processed using a tool such as genomic VCF (gVCF).
[0130] In step 120 of method 100 of FIG. 1A, a subject-specific genome-wide list of genetic markers is detected in a second sample (e.g., plasma or blood) from the subject to generate a list of tumor-associated genome-wide genetic markers in the patient sample (e.g., plasma or blood sample).
[0131] In step 130 of the method 100 of FIG. 1A, the noise probability of each marker is analyzed. For example, if the marker is an SNV or an indel, P Ncan be analyzed as a function of 1) the MQ of the SNV / indel; 2) the fragment length of the read containing the SNV / indel; 3) a consensus test within the read duplication family containing the SNV or indel; and / or 4) the BQ of the SNV / indel. Similarly, if the marker is a CNV or SV, the probability that the marker is noise-associated can be analyzed by statistically classifying each CNV or SV window in the list as signal (S) or noise (N) based on (1) its position relative to the centromere, (2) the MQ of the read group containing the CNV / SV, and / or (3) a list of CNV windows in the artificially read cfDNA data. The noise removal step 130 can include implementing a best-fit receiver operating characteristic curve that involves probabilistic classification of the genetic markers in the list based on the combined base quality score and mapping quality score. Typically, the combined BQMQ score is provided as a matrix (x, y), where x is the BQ score and y is the MQ score. In exemplary embodiments, a combined BQMQ score of 10-50 (for each parameter) is typically used, e.g., BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), and (30, 40). In some embodiments, marker classification involves measuring the area under the ROC curve (AUC), which typically represents the probability that a randomly selected candidate marker from among potential markers will exhibit a higher value than a randomly selected control marker. For completely uninformative markers, the ROC curve approaches the ascending diagonal (referred to as the "chance diagonal" or "chance line"), and the AUC is 0.5 (i.e., the expected probability of classification by chance alone). Conversely, for perfect classification, the ROC curve peaks at theoretical accuracy (both sensitivity and specificity are 100%), and the AUC tends to be 1, i.e., the highest probability value. A representative ROC is shown in Figure 3B. The pre-filtration error model and post-filtration effect of the base quality filter are shown in Figure 3A. FIG. 3C shows that application of base quality (BQ) and mapping quality filters reduces sequencing errors by approximately 7-fold.
[0132] In step 140 of method 100 of FIG. 1A, an estimated tumor fraction (eTF) of a biological sample is calculated based on one or more integrative mathematical models. Depending on the marker (e.g., SNVs / indels vs. CNVs / SVs), the mathematical model integrates multiple process quality criteria as well as patient-specific attributes to estimate the tumor fraction (TF). The disclosed systems and methods involve using marker-specific mathematical algorithms to estimate the tumor fraction, recognizing the fundamental differences between SNVs / indels and CNVs / SVs in terms of frequency and association characteristics with traits (e.g., cancer). In each case, the mathematical inference model outputs an estimated fraction of tumor DNA in the biological sample (e.g., plasma) based on the number / frequency of markers, estimated noise, reads, mutation load, and / or coverage or depth.
[0133] In some embodiments, the disclosed methods include estimating TFs based on the detection of multiple SNV / indel markers. Here, estimated TFs (eTFs[SNVs]) are calculated by integrating process-quality criteria, including estimated genome coverage and sequencing noise, with patient-specific parameters, including mutational burden (N). Preferably, the method includes calculating the estimated tumor fraction (eTF) for an SNV / indel marker, where eTFs[SNVs] = 1-[1-(ME(σ)R) / N]^(1 / cov), where M is the number of tumor-specific common detections in the patient sample, σ is an empirically estimated measure of noise, R is the total number of unique reads in the region of interest (ROI), N is the tumor mutational burden, and cov is the average number of unique reads per site in the ROI.
[0134] In some embodiments, the methods of the present disclosure include estimation of TFs based on the detection of multiple CNV / SV markers, where estimated TFs (eTFs[CNV]) are calculated by integrating the directionality of coverage depth skewed consistent with tumor CNV / SV directionality, where copy number gains are positively skewed and copy number losses are negatively skewed. Preferably, the method comprises calculating an estimated tumor fraction (eTF) for a CNV marker, where eTF[CNV]=(sum_{i]=[(P(i)-N(i)]*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i)]-E(σ)], where P is the median depth in the genomic window indexed by {i}, representing plasma depth coverage; T is the median depth in the genomic window indexed by {i}, representing tumor depth coverage; and N is the median depth in the genomic window indexed by {i}, representing normal depth coverage.
[0135] In step 150 of method 100 of FIG. 1A, residual disease is diagnosed in the subject based on the eTF (calculated in step 140) and an empirical threshold calculated by a background noise model. In some embodiments, the detection threshold comprises a baseline noise TF estimate empirically measured from healthy samples. In such embodiments, any eTF above the threshold (e.g., at least 2 standard deviations of the noise TF distribution (FPR<2.5%); preferably, greater than 3 STDs or greater than 5 STDs) is defined as a positive detection.
[0136] Additionally, various embodiments provide a method for detecting residual disease in a subject in need thereof, as provided by the exemplary workflow 100 shown in FIG. 1B. As provided in step 110 of method 100 of FIG. 1B, the workflow may include receiving a first subject-specific genome-wide list of reads associated with genetic markers from a first biological sample of the subject. The first biological sample may include a baseline sample. The first list may include reads each one base pair in length. The baseline sample may include a tumor sample or a plasma sample. The first biological sample may also include a normal cell sample.
[0137] As provided in step 120 of method 100 of FIG. 1B, the method may include filtering real sites from the first read list. Filtering may include removing repetitive sites generated across a cohort of reference healthy samples from the read list. Alternatively, or in combination, filtering may include identifying germline mutations in the biological sample and / or identifying shared mutations between the tumor sample and peripheral blood mononuclear cells of the normal cell sample as germline mutations and removing the germline mutations from the read list. As provided in step 120 of method 100 of FIG. 1B, the workflow may include filtering artificial sites from the first list, where the filtering includes removing repetitive sites generated across a cohort of reference healthy samples from the first list of genetic markers. And / or the filtering may include identifying germline mutations in peripheral blood mononuclear cells of the normal cell sample and removing the germline mutations from the first list of genetic markers.
[0138] As provided in step 130 of method 100 of FIG. 1B, the workflow may include detecting reads from a second subject-specific genome-wide list of genetic markers in a second biological sample from the subject and generating a tumor-associated genome-wide list of genetic markers in the second sample.
[0139] As provided in step 140 of method 100 of FIG. 1B , the workflow may include using at least one error suppression protocol to filter noise from the genome-wide lists of first and second reads to generate a first filtered read set for the genome-wide list of first reads and a second filtered read set for the genome-wide list of second reads. The at least one error suppression protocol may include calculating the probability that any single nucleotide variation in the first and second lists is an artifact and removing the variation. The probability may be calculated as a function of features selected from the group consisting of mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof. Alternatively, or in combination, the at least one error suppression protocol may include removing artifacts using discrepancy testing between independent copies of the same DNA fragment generated from polymerase chain reaction or sequencing processing. and / or may involve removing artificial variations using overlap consensus, whereby artificial variations are identified and removed when there is no match in the majority of a given overlap family.
[0140] As provided in step 150 of method 100 of FIG. 1B, the workflow applies a background noise model to one or more integrated mathematical models to generate putative tumor profiles for the first and second biological samples using the first and second filtered read sets. Fraction This may include the calculation of (eTF).
[0141] As provided in step 160 of method 100 of FIG. 1B, the workflow can include detecting residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.
[0142] Additionally, various embodiments provide a method for detecting residual disease in a subject in need thereof, as provided by the exemplary workflow 100 shown in FIG. 1C. As provided in step 110 of method 100 of FIG. 1C, the workflow can include receiving a first subject-specific genome-wide list of reads associated with genetic markers from a first biological sample of the subject. The first biological sample can include a baseline sample. Each of the first list of reads can include copy number variations (CNVs). The baseline sample can include a tumor sample or a plasma sample.
[0143] As provided in step 120 of method 100 of FIG. 1C, the workflow can include receiving a second subject-specific genome-wide list of reads associated with genetic markers from a second biological sample from the subject. The second biological sample can include a peripheral blood mononuclear cell sample (PBMC). The second list of genetic markers can each include copy number variations (CNVs).
[0144] As provided in step 130 of method 100 of Figure 1C, the workflow may include filtering artificial sites from the first and second read lists, which filtering may include removing repetitive sites generated on a cohort of reference healthy samples from the first and second read lists. Alternatively, or in combination, filtering may include identifying CNVs shared by the first and second lists as germline mutations and removing the mutation reads from the first and second lists.
[0145] As provided in step 140 of method 100 of FIG. 1C, the workflow may include detecting reads from a third subject-specific genome-wide list of genetic markers in a third biological sample from the subject and generating a tumor-associated genome-wide representation of the genetic markers in the third sample.
[0146] As provided in step 150 of method 100 of FIG. 1C, the workflow may include normalizing each of the first, second, and third read lists to generate a first filtered read set for the first genome-wide read list, a second filtered read set for the second genome-wide read list, and a third filtered read set for the third genome-wide read list.
[0147] As provided in step 160 of method 100 of FIG. 1C, the workflow uses the third filtered read set to apply a background noise model to one or more integrated mathematical models, one or more models using the first filtered read set to generate a first eTF, and / or one or more models using the second filtered read set to generate a second eTF, to generate a putative tumor profile for the third biological sample. Fraction This may include the calculation of (eTF).
[0148] As provided in step 170 of method 100 of FIG. 1C, the workflow can include detecting residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold.
[0149] [Scheme]
[0150] 1D and 1E show schematic workflows for implementing the methods of the present disclosure. FIG. 1D outlines a workflow typically used when the markers of a subject of interest include SNVs / indels, and FIG. 1E outlines a workflow typically used when the markers of a subject of interest include CNVs / CVs. Note that while the separate workflows are provided for illustrative purposes, they are not required to be separate for implementing the methods of the present disclosure. For example, certain features / elements of the workflows may be used in combination to generate an output (e.g., a combined estimated tumor fraction based on SNVs / indels and CNVs / SVs) related to an outcome of interest (e.g., whether a subject develops cancer).
[0151] As shown in Figure 1D, MRD detection based on SNV / indel markers typically utilizes the steps of receiving data; generating patient-specific patterns of SNVs / indels; removing / filtering artifactual sites; detecting reads / sites in follow-up samples; machine learning; read correction; error suppression using specific algorithms including detection of sites that provide an estimate of tumor fraction; and, in some cases, orthogonally integrating analysis of secondary features in the genomic data (e.g., analysis of fragment size shifts) to improve the sensitivity, specificity, and / or reliability of detection.
[0152] In the first step of Figure 1D, genetic data from a baseline sample (typically a tumor sample, but may include pre-treatment plasma alone or in combination with a tumor sample) and a normal sample (typically PBMC, but may include adjacent normal tissue or buccal swabs) are received to generate a patient-specific marker pattern (e.g., including SNVs / indels). Next, artificial sites are filtered to call a reference list of somatic mutations from the baseline sample, where germline mutations are removed from the sample. Somatic mutation calling is also performed independently using multiple callers (e.g., MUTECT, STRELAKA) using the intersection of the callers to generate a list of high-confidence mutations. Serially or in parallel, recurrent artificial sites are generated across a cohort of normal plasma samples (a panel of normal (PON) blacklist or masks) to remove patient-detected mutations and eliminate common sequencing or alignment artifacts. The filtered high-confidence patient-specific mutation dataset is then used to detect mutations in follow-up plasma samples. Typically, follow-up plasma is collected after surgery, during or after treatment (e.g., during chemotherapy), or at follow-up (e.g., to check for recurrence or recurrence).
[0153] Next, a highly sensitive method capable of detecting a single mutant fragment is employed. This process employs one or more error suppression steps. The first error suppression step uses a filtering scheme to analyze single read bases and quantify the probability that the read represents an artifact. An exemplary method involves a multidimensional classification framework using a support vector machine (SVM) classifier with a linear kernel. The classification engine is trained on germline SNPs compared to low variant allele fraction (VAF) sequencing artifacts in normal PBMC samples. Here, classification decision boundaries are defined in a multidimensional space, including variant base quality (VBQ), mapping quality (MQ), read position (PIR), and / or mean read base quality (MRBQ). To evaluate the classification scheme, validation criteria for the SVM classification scheme were compared with random forests under the same protocol after 10-fold cross-validation. The SVM classifier demonstrated high classification performance, slightly outperforming the random forest model. SVM achieved an average sensitivity of 90.7% and specificity of 83.9% across all patients (N = 10 samples, F1 = 87.7%, PPV = 84.9%).
[0154] In the second error suppression step, artifacts introduced by PCR or sequencing were corrected using comparison of independent copies of the same original DNA fragment. In cfDNA samples, 150 bp of the paired ends were typically sequenced, and given the short size of typical cfDNA fragments (approximately 165 bp), duplicated paired reads (overlapping R1 and R2 sequences) were obtained. Therefore, discrepancies between R1 and R2 pairs were considered potential sequencing artifacts that were matched back to the corresponding reference genome. Furthermore, to recognize the potential generation of independent duplicates due to any DNA molecules copied multiple times during sequencing and PCR, duplicate families were identified by 5' and 3' similarity as well as alignment position. Each duplicate family was then used to check for consensus of specific mutations across the independent replicates, and artifacts that showed no match across the majority of the duplicate family were corrected.
[0155] Next, we estimate the proportion of patient-specific mutations appearing in plasma. This parameter follows a binomial distribution over N independent Bernoulli experiments, where N is the patient's mutational burden. Each such experiment involves multiple rounds of random sampling, where the probability of sampling a mutant fragment in each round depends on the local coverage, which is the tumor fraction. Therefore, we estimate the proportion of patient-specific mutations appearing in plasma by the following formula: M = N(1-(1-TF) cov ) + μ * R, where M is the number of mutations detected in the follow-up plasma samples, N is the mutation burden in the patient-specific mutation pattern, TF is the tumor fraction, cov is the local coverage at the patient's mutation site, and μ is the noise rate corresponding to the mutation site for a particular patient. This relationship allows the calculation of patient tumor fraction from mutation detection rate even at very low allele fractions, where the mutant allele fraction itself is not informative (it primarily represents random sampling between 0 and 1 for valid coverage-only reads).
[0156] To address noise variations between patients with different mutation patterns, patient-specific mutation patterns are used to calculate the expected noise distribution across a cohort of healthy plasma samples (panel of normals, PON). The same procedure as above is then followed to detect patient-specific patterns in healthy samples (PON) or other patients (inter-patient analysis). This detection represents a background noise model, which calculates the mean and standard deviation (μ, σ) of the artificial mutation detection rate. Tumor detection and tumor fraction estimation are achieved when the tumor fraction detected in a patient is more reliable than the artificial tumor fraction, with an error rate above the mean, equivalent to 1.5 × σ.
[0157] In some cases, the workflow may then include orthogonal integration of calculations based on fragment size shifts. Here, read-based features, such as DNA fragment size shifts, may be orthogonally integrated into the model to make the prognostic / diagnostic method more robust, accurate, and / or sensitive. The significance of orthogonal features (in determining MRD) may be determined using statistical approaches or probabilistic mixture models (e.g., Gaussian models). See Example 3A for a detailed listing.
[0158] In the illustrated method, highly reliable tumor-specific detections in plasma samples are aggregated and converted into estimates of tumor DNA (TF) fractions based on a stochastic dilution model. The entire detection protocol (detection, error suppression, and tumor fraction estimation) is also performed on a panel of healthy plasma samples (PONs) using the patient-specific mutation list, and the same features are used to calculate the distribution of noisy TF values in the healthy samples. Tumor detection and estimation are then performed only on samples that exhibit a significantly higher tumor fraction than the PON noisy TF values, using a statistical significance framework (z-score) that ensures a low false positive rate (high specificity). To orthogonally confirm the presence of tumor DNA in mutation detections in plasma, statistical methods (significance tests or GMMs) are used to quantify the within-patient fragment size shift between the tumor-specific detection list and other random mutation detection lists.
[0159] Alternatively, or in combination with the above workflow, the present disclosure also relates to the detection of residual disease (or monitoring therapy) using CNV / SV markers. As shown in Figure 1E, MRD detection based on CNV / SV markers typically involves receiving data; generating baseline sample-specific and / or normal sample-specific CNV / SV signatures; removing germline CNV events; filtering artificial windows; detecting window-based depth coverage in follow-up samples; normalizing, for example, using guanine-cytosine (GC) normalization and / or z-score normalization; detecting tumor CNV signals to provide an estimate of tumor fraction; and optionally orthogonally integrating secondary feature analysis (e.g., fragment size shift analysis) in genomic data to improve the sensitivity, specificity, and / or reliability of detection.
[0160] In the first step of Figure 1E, genetic data from a baseline sample (usually a tumor sample, but may include pre-treatment plasma alone or together with a tumor sample) and a normal sample (usually a PBMC, but may include adjacent normal tissue or a buccal swab) are received to generate tumor-specific and normal marker patterns (e.g., patterns including CNVs / SVs). Next, tumor copy number variations (T_CNVs) are called using the baseline to normal panel (PON). PBMC copy number variations (P_CNVs) are called using the PBMC sample against the PON-of-normal (PON). Shared copy number variations are considered germline. Tumor somatic events (T_CNVs detected only in tumor tissue) and PBMC somatic events (P_CNVs, P_CNVs detected only in PBMC tissue) can be used to detect and estimate tumor fraction.
[0161] Next, germline mutations (e.g., CNV / SV events) are removed from the CNV / SV reference list to generate baseline sCNV / SV and / or normal sCNV / SV. Also, windows with low mappability and / or coverage are filtered. Serially or in parallel, recurrent artificial sites are generated across a cohort of healthy plasma samples (panel of normal (PON) blacklist or mask). The samples are removed from the windows to filter out the artificial windows. The filtered high-confidence reference CNV / SV segments are used to detect mutations in follow-up plasma samples. Follow-up plasma is typically collected after surgery, during or after treatment (e.g., during chemotherapy), or at follow-up (e.g., for recurrence or relapse checks).
[0162] Currently, CNV sites with artifacts are generated on a cohort of healthy plasma samples (normal PON blacklist panel) and removed from patient-detected mutations to remove common sequencing or alignment artifacts such as centromere and repeat regions.
[0163] Next, regions of interest (ROIs) containing all genomic segments of sT_CNV and sP_CNV were binned into windows (≥500 bp). The depth coverage (read counts) of each window was estimated from plasma samples at follow-up (post-surgery, during treatment, and follow-up for recurrence). The median depth coverage per window was calculated and divided by the mean sample coverage.
[0164] Depth coverage values were then normalized and corrected for GC content bias and mappability bias by performing two LOESS regression curve fits on the bin-wise GC fraction and mappability scores.
[0165] Further batch effect correction is performed using stable z-score normalization applied to each sample separately. Briefly, the median and median absolute deviation (MAD) are calculated based on the neutral region of each sample, and then all CNV bins are normalized by (B(i)-Median) / MAD.
[0166] For each bin, depth coverage skew and fragment size center of mass (COM) skew were calculated relative to a panel of normal (PON) healthy plasma samples. Here, low tumor fraction samples exhibit sparse depth coverage skew biased by the directionality of CNV segment amplifications, while deletions exhibit a bias toward negative depth coverage skew. Neutral regions, on the other hand, exhibit random distortion with unfavorable directionality. Therefore, multiplying the differential (plasma PON) depth coverage distortion by the directionality of CNV segments (amplifications multiplied by +1, deletions multiplied by -1) sums up the genome-wide CNV signal, while neutral region noise is canceled out due to random directionality.
[0167] This step is performed using the following formula, where M is the number of windows covering the ROI:
number
[0168] The tumor fraction can then be calculated by ascertaining the linear dilution ratio between the cumulative signal detected in the plasma sample compared to the cumulative signal detected in the tumor. This procedure is carried out using the following formula:
number
[0169] where N(i), P(i), and T(i) represent the patient PBMC, plasma, and tumor depth coverage in window I, respectively.
[0170] To address noise variations between patients with different CNV patterns, patient-specific CNV patterns are used to calculate the expected noise distribution across a cohort of healthy plasma samples (panel of normals, PON). The process is essentially the same as for SNV marker analysis, allowing for the detection of patient-specific patterns in healthy plasma samples (PON) or other patients (inter-patient analysis). This detection represents a background noise model that calculates the mean and standard deviation (μ, σ) of the artificial mutation detection rate. Tumor detection and tumor fraction estimation are achieved when the tumor fraction detected by a patient is more reliable than the artificial tumor fraction, with an error rate above the mean, equivalent to 1.5 × σ.
[0171] We can also infer tumor fraction from the directional genome-wide depth coverage skew in sP_CNVs. Here, PBMC-specific CNV events are expected to show a decrease in signal with increasing tumor DNA fraction (because tumor DNA does not contain this CNV event). Therefore, a negative correlation is expected between tumor fraction and P_CNV detection signal in plasma. Therefore, we multiply the differential (PBMC-plasma) depth coverage skew by the directionality of the PBMC CNV segments (amplifications are multiplied by +1, deletions by -1) to sum the PBMC CNV signal across the genome (Figure 11A).
[0172] The percentage of loss of PBMC CNV signal is then calculated, for example, using the following formula:
number
[0173] As in the case of MRD estimation using SNV / indel markers, secondary features can be orthogonally integrated into the final calculation. Here, to improve the stability, accuracy, and / or sensitivity / specificity of the detection method, read-based features, such as DNA fragment size shift, can be orthogonally incorporated into the model. The significance of orthogonal features (in determining MRD) can be determined using a generalized linear model (GLM) to orthogonally determine tumor fraction based on the relationship between CNV depth coverage and fragment size shift. For a detailed list, see Example 3B.
[0174] It should be understood that the workflow disclosed herein may also be used generally, with some modifications, to detect residual disease during or after chemotherapy, immunotherapy, targeted therapy, or a combination thereof, and / or in the process of monitoring the effectiveness of such treatment.
[0175] The exemplary method is based, in part, on the recognition that genome-wide CNV signals in plasma samples accumulate only if the coverage skew in plasma follows the same direction as copy number variations (amplifications and deletions) in the baseline tissue (e.g., tumor). Therefore, the tumor DNA ratio can be calculated from the signal gain in the plasma sample from CNV events specific to the patient's tumor, e.g., using a linear dilution ratio of cumulative CNV signal in plasma divided by cumulative CNV signal in tumor. The tumor fraction can be orthogonally estimated using a similar mixture dilution model based on signal loss from CNV events specific to only the patient's PBMCs (hematopoietic cell body CNV events). Additionally, the entire CNV detection protocol is performed on a panel of healthy plasma samples (PONs) using the patient-specific copy number variation list, and the distribution of noisy TF values in the healthy samples is calculated using the same CNV patterns. Then, tumor detection and estimation are performed only for samples that exhibit a tumor fraction significantly higher than the PON noisy TF values, using a statistical significance framework (z-score) that ensures a low false positive rate (high specificity). Orthogonal confirmation of the presence of tumor DNA in plasma was performed by confirming the relationship (negative correlation) between CNV log2 values across patient-specific CNV segments and center-of-mass (COM) values of fragment sizes, which can be converted into orthogonal estimates of CNV-based TF estimation based on generalized linear models (GLMs).
[0176] Machine Learning Without being bound to a single embodiment, and purely for illustrative purposes, machine learning (ML) algorithms have been integrated into existing methodologies, either individually or in combination with individual steps, according to various embodiments herein. ML can be incorporated to optimize the results output from an algorithm (e.g., a neural network, an ML algorithm, etc.) by utilizing an input training dataset, cross-referencing the output to known answers, backpropagation, and adjusting weighting coefficients and parameters associated with a given ML algorithm in an iterative loop to reach a threshold quality of data output. In a subsequent step, a probabilistic model (e.g., optimized, combined, or alternatively trained), such as logistic regression, can be used to validate the model's predictive ability on a test dataset. In some cases, resampling can be performed to obtain an unbiased assessment of the model's expected future performance. Features of the ROC curve, such as the area under the curve (also known as c-indexation), or the probability of agreement from a statistical test such as the Wilcoxon-Mann-Whitney test, can provide a good summary measure of pure predictive discrimination.
[0177] Preferably, the ML algorithm adaptively and / or systematically filters sequencing noise associated with each read in the list based on one or more quality filters or read functions. In some embodiments, the ML algorithm implements a base quality (BQ) filter (more specifically, variable base quality (VBQ) or mean read base quality (MRBQ)) to filter noise. In some embodiments, the ML algorithm implements a mapping quality filter to filter noise. In some embodiments, the ML algorithm implements a position in read (PIR) filter to filter noise. In some embodiments, the ML algorithm implements a combination of filters.
[0178] In some embodiments, the machine learning (ML) method used in the systems and / or methods of the present disclosure includes a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), discriminant analysis, nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof, preferably a support vector machine (SVM). In some embodiments, the ML is trained to distinguish between cancer-altered sequencing reads and reads altered by sequencing or PCR errors. In some embodiments, the ML is trained on a large whole genome sequencing (WGS) cancer dataset containing billions of reads across tumor mutations and normal sequencing errors. In some embodiments, the ML (a) identifies sequencing or PCR artifacts with high accuracy and (b) integrates sequence context to identify specific signature reads.
[0179] The present disclosure further relates to systems and programs that utilize ML, e.g., engines, to adaptively and / or systematically filter ordering noise. The present disclosure also relates to a computer-readable storage medium containing a program for detecting tumor markers, including somatic mutations, in genomic reads, the program utilizing ML, e.g., support vector machines.
[0180] Convolutional neural networks, as known in the art, generally achieve advanced forms of processing and classification / detection by first looking for low-level features, such as repetitive sequences in the reads, and then progressing to more abstract concepts through a series of convolutional layers. CNNs may do this by passing the data through a series of convolutional, nonlinear, pooling (or downsampling, as described below), and fully connected layers to obtain an output. Again, the output may be a single class or class probabilities that best describe the data, or detect objects in the data.
[0181] In a CNN, the first layer is typically a convolutional layer (conv). This first layer processes a representative array of reads using a set of parameters. Rather than processing the entire data, a CNN uses filters (or neurons or kernels) to analyze a list of data subsets. The subsets include the focal point and surrounding points in the array. For example, a filter might examine a series of 5x5 regions (or areas) in a 32x32 representation. These regions are called the receptive field. The filter is typically the same depth as the input; a filter for a representation with dimensions of 32x32x3 would have the same depth (e.g., 5x5x3). Using these exemplary dimensions, the actual convolution process involves sliding the filter along the input data, multiplying the filter value by the original representation of the data, calculating the element-wise multiplication, and adding the values together to arrive at a single number for the examined region of the representation.
[0182] Using a 5x5x3 filter, an activation map (or filter map) with dimensions of 28x28x1 is obtained after this convolution process is completed. For each additional layer used, the spatial dimensions are better preserved, as using two filters results in a 28x28x2 activation map. Each filter typically has a unique signature that together indicate the feature identifier required for the final data output. Using such filters in combination, the CNN can process the data input and detect the features present in each representation. Thus, if a filter functions as a curve detector, convolution of the filter along the data input generates an array of numbers in the activation map corresponding to a high probability of a curve (high summation element-wise multiplication), a low probability of a curve (low summation element-wise multiplication), or a zero value if the input volume at a particular point does not provide anything to activate the curve detector filter. Thus, the more filters (also called channels) in a Conv, the more depth (or data) is provided on the activation map, and therefore, more information about the input, leading to a more accurate output.
[0183] The balance between accuracy of CNN is the processing time and power required to produce the results. In other words, the more filters (or channels) there are, the more time and processing power is required to perform the convolution. Therefore, the selection and number of filters (or channels) that satisfy the requirements of a CNN method should be specifically chosen to produce the most accurate output possible while taking into account the available time and power.
[0184] Furthermore, to enable the CNN to detect more complex features, additional convs can be added to analyze the output (e.g., activation maps) from previous convs. For example, if the first conv looks for basic features such as curves and edges, the second conv can search for more complex features, which can be combinations of individual features detected in previous conv layers. By providing a series of convs, the CNN can detect increasingly higher-level features, eventually reaching a detection probability of a particular desired object. Furthermore, because conv stacks overlap each other and each conv level in the stack shrinks due to analysis of previous activation map outputs, each conv naturally analyzes a wider receptive field, thereby allowing the CNN to accommodate an expanded representation space when detecting the desired object.
[0185] A CNN structure generally consists of a set of processing blocks, including at least one processing block for convolution of an input volume (data) and at least one processing block for deconvolution (or deconvolution). Furthermore, the processing blocks may include at least one pooling block and a non-pooling block. Pooling blocks can be used to reduce the resolution of data to generate an output usable by Conv. This provides computational efficiency (time and power efficient) and may improve the actual performance of the CNN. The pooling, or subsampling, blocks make filters smaller and keep computational requirements reasonable. They may coarsen the output (which may lose spatial information within the acceptable field) and reduce the size of the input by only a certain factor.
[0186] A deconvolution block may be used to reconstruct this coarse output, generating an output volume of the same dimensions as the input volume. The deconvolution block may be viewed as the inverse of the convolution block, which returns the activation output to the original input volume dimensions. However, the deconvolution process generally simply spreads the coarse output into a sparse activation map. To avoid this result, the deconvolution block densifies this sparse activation map, ultimately generating an enlarged and dense activation map that, after any necessary processing, produces a final output volume whose size and density more closely resembles the input volume. Rather than reducing multiple array points in the receiving region to a single number as the inverse of the convolution block, the deconvolution block associates a single activation output point with multiple outputs, enlarging and densifying the resulting activation output.
[0187] It should be noted that although pooling blocks can be used to reduce data and non-pooling blocks can be used to expand the reduced activation map, convolution and deconvolution blocks can structure both convolution / deconvolution and reduction / expansion without separate pooling and non-pooling blocks.
[0188] Pooling and unpooling processes can suffer from subject-object dependency found in the data input. Pooling generally shrinks the data by looking at sub-data windows without overlapping windows, so loss of spatial information becomes apparent as the shrinking occurs.
[0189] A processing block may include other layers packaged with the convolutional or deconvolutional layers. These may include, for example, rectified linear unit layers or exponential linear unit layers, which are activation functions that examine the output from the Conv in that processing block. The ReLU or ELU layer acts as a gating function that advances only values that correspond to positive detection of the feature of interest specific to the Conv.
[0190] After the CNN is given a basic structure, it is prepared for a training process to improve its accuracy in data classification / detection (of the subject of interest). This involves a process called backpropagation. In this process, a training dataset, or sample data for training the CNN, is used to update parameters to reach an optimal, or threshold, accuracy. Backpropagation involves a series of iterative steps (training iterations), which train the CNN slowly or quickly depending on the backpropagation parameters. The backpropagation process generally includes a forward pass, a loss function, a backward pass, and parameter (weight) updates, depending on the learning rate. The forward pass involves passing training data through the CNN. The loss function is a measure of the error in the output. The backward pass determines the contribution factor of the loss function. The weight updates involve updating the filter parameters, which moves the CNN toward the optimal. The learning rate determines the extent of the weight updates at each iteration to reach the optimal. If the learning rate is too low, training may take too long and require high processing power. If the learning rate is too fast, each weight update may be too large and may not accurately achieve a predetermined optimum value or threshold.
[0191] The backpropagation process can complicate training, resulting in a slower learning rate and requiring more specific and carefully determined initial parameters at the beginning of training. One such complication is the deep amplification of the network due to changes in the parameters of the convs, with weight updates at the end of each iteration. For example, as mentioned above, if a CNN has multiple convs capable of higher-level characterization, parameter updates to the first conv are multiplied by each subsequent conv. The net effect is that, depending on the depth of a given CNN, even the smallest changes to the parameters have a large impact. This phenomenon is called internal covariate shift.
[0192] In general, the CNN of the present disclosure can adaptively and / or systematically filter ordering noise. In some embodiments, the CNN structure is designed based on the inventors' recognition that trinucleotide contexts contain distinct features involved in mutagenesis. Therefore, the CNN uses a perceptual field of view of size 3 to cover all features (columns) at a given position. After two consecutive convolutional layers, downsampling is applied using max pooling with a receptive field of 2 and a step count of 2, forcing the engine model to retain only the most important features in a narrow spatial region. The resulting structure maintains spatial invariance when convolved across a 3-nucleotide window and captures a "quality map" by collapsing read fragments into 25 segments corresponding to regions of approximately 8 nucleotides. Final classification is performed by directly applying the output of the last convolutional layer to a sigmoidally fully connected layer. The CNN employs a simple logistic regression layer, rather than a multilayer perceptron or global average pooling, to preserve position-related features in genomic reads.
[0193] To train the engine, various lung cancer patients and their corresponding systemic error profiles are first sampled. The goal of training is to use a training scheme that enables highly sensitive detection of true somatic mutations and rejects candidate mutations caused by systemic errors. For example, samples from subjects with or suspected of having cancer, such as a mixture of whole tumor samples and healthy tissue samples, can be used for training.
[0194] Upstream process: [Receiving genetic data] In some embodiments, genetic data is received in situ from a subject's biological sample (e.g., a tumor sample or a normal cell sample, including PBMCs). This is typically accomplished by sequencing. In some embodiments, the sample can be purified using conventional methods to obtain cell subpopulations. For example, PBMCs can be purified from whole blood using various known Ficoll-based centrifugation methods (e.g., Ficoll-Hypaque density gradient centrifugation). Other cells, such as T cells, can also be purified by selecting for appropriate phenotypes using techniques such as immunomagnetic cell sorting (e.g., DYNABEADS, Invitrogen, Carlsbad, CA, USA). For example, T cells can be purified using a two-step selection process that first removes CD8+ cells and then selects for CD4+ cells. The purity of the cell population can be confirmed by assessing appropriate markers, such as CD19-FITC, CD3-PE, CD8-PerCP, CD11c-PE Cy7, CD4-APC, and CD14-APC Cy7, using commercially available antibodies (e.g., BD Biosciences).
[0195] After sample preparation, DNA is extracted from the sample and marker analysis is performed. In this example, the DNA is genomic DNA. Various methods for isolating DNA, particularly genomic DNA, are known to those skilled in the art. Generally, known methods involve disruption and dissolution of the starting material, followed by removal of proteins and other contaminants, and finally recovery of the DNA. For example, techniques including alcohol precipitation; organic phenol / chloroform extraction, and salting out have been used for many years to extract and isolate DNA. An example of DNA isolation is illustrated below (e.g., Qiagen ALL-PREP Kit). However, various other commercially available kits for genomic DNA extraction exist (Thermo-Fisher, Waltham, MA; Sigma-Aldrich, St. Louis, MO). DNA purity and concentration can be assessed by various methods, such as spectrophotometry.
[0196] In some embodiments, the list of genetic markers comprises a list of genetic markers compiled in a variant call format (VCF) file. As understood in the art, VCF files are used in bioinformatics to store genetic sequence variations. The VCF format was developed with the advent of large-scale genotyping and DNA sequencing projects, such as the 1000 Genomes Project. Alternatively, the list can be provided in a general feature format (GFF), which contains all of the genetic data. GFFs generally provide redundant features that are shared genome-wide. In contrast, VCF only requires that variations be stored with the reference genome.
[0197] Microarray technology is widely used to detect the disclosed markers, such as SNVs / indels and CNVs / SVs. For example, array comparative genomic hybridization (array CGH) and single nucleotide polymorphism (SNP) microarrays can be used. In traditional array CGH, reference and test DNAs are fluorescently labeled and hybridized to an array, and the signal ratio is used as an estimate of the copy number (CN) ratio. SNP microarrays can also be based on hybridization, but a single sample is processed on each microarray, and an intensity ratio is formed by comparing the intensity of the sample under investigation with a collection of reference samples or all other samples tested. Microarrays / genotyping arrays are efficient for large-scale CNV detection, but have low sensitivity for detecting CNVs of short genes or DNA sequences (e.g., less than about 50 kilobases (kb) in length).
[0198] In some embodiments, the markers of the present disclosure can be detected using next-generation sequencing (NGS). By providing a base-by-base view of the genome, NGS can detect small or novel CNVs that may not be detected by arrays. Examples of suitable NGS methods can include whole genome, whole exome sequencing, or targeted exome sequencing. Preferably, the sequencing method uses WGS.
[0199] In some embodiments, a subject's sample is sequenced, for example, using whole genome sequencing (WGS), and called (for SNVs / indels and / or CNVs / CV markers) using standard methods. For example, SNV calling from NGS data utilizes computational methods to identify the presence of single nucleotide variants (SNVs) from the results of next-generation sequencing (NGS) experiments. With the increase in NGS data, this technology is increasingly popular for performing SNP genotyping, using a wide variety of algorithms designed for specific experimental designs and applications. Similarly, there are several bioinformatics approaches for detecting CNVs from next-generation sequencing data (Pirooznia et al., Front Genet., 6:138, 2015). In some embodiments, the sample is processed and sequenced to obtain a sequence file, which is then processed using tools such as genome VCF or exome VCF (eVCF).
[0200] In some embodiments, the methods of the present disclosure can include generating a genetic marker inventory. A typical inventory includes genetic data from a tumor sample that has been whole-genome sequenced, as well as a control (e.g., PMBC). The tumor sample preferably includes a resected tumor or FNA, e.g., a lung adenocarcinoma or a skin melanoma. The control sample preferably includes PMBC obtained using Ficoll separation, as described above. An admixture is then generated, and the markers therein are analyzed using the computational methods of the present disclosure.
[0201] In certain embodiments, the methods of the present disclosure may include classifying genetic data into distinct components based on the markers contained therein, e.g., SNVs, CNVs, indels, SVs, mutations, deletions, fusions, etc. In preferred embodiments, the classifying step may include separate binning of somatic SNV (sSNV) markers and somatic CNV (sCNV) markers, which are noise filtered and analyzed separately according to the computational methods of the present disclosure. Here, the computational methods for analyzing SNV markers for noise and uniqueness may differ from the methods for analyzing CNVs. In some embodiments, the computational analysis of SNVs or indels may be performed sequentially with the computational analysis of CNVs or SVs. In some embodiments, the analyses may be performed together.
[0202] The present disclosure provides the use of mathematical algorithms and computational methods to (a) filter artificial noise and (b) screen for true markers.
[0203] For noise cancellation where the marker is an SNV or indel, artificial noise is cancelled based on multiple parameters, including base quality and / or mapping quality. Typically, base quality (BQ) relates to the reliability of the sequencing quality of each base, and mapping quality (MQ) score relates to a confidence estimate of the accuracy of the marker's mapping to the genome. In the context of sSNV markers, the base quality (BQ) score is a measure of the quality of the nucleobase identification generated by automated DNA sequencing. It can be determined using conventional methods, such as the Phred quality score assigned to each nucleotide base call in an automated sequencer trace. The Phred quality score (Q) is defined as a property logarithmically related to the base call error probability (P). For example, if Phred assigns a quality score of 30 to a base, the probability that this base will be called incorrectly is 1 / 1000. Typically, the BQ of a sequencing read is between 10 and 50, for example a BQ score of 10, 15, 20, 25, 30, 35 or 40.
[0204] Also, in the context of sSNV markers, a mapping quality (MQ) score is a measure of the confidence that a read is actually derived from a position aligned by a mapping algorithm. This can be determined using conventional methods, such as the mapping quality score (see Li et al., Genome Research 18:1851-8, 2008). Typically, the MQ of a read is between 10 and 50, e.g., an MQ score of about 10, 15, 20, 25, 30, 35, or 40.
[0205] In some embodiments, the denoising step involves performing an optimal receiver operating characteristic (ROC) curve, which involves a probabilistic classification of the genetic markers in the list based on combined base quality (BQ) and mapping quality (MQ) scores. Typically, the combined BQMQ scores are provided as a matrix (x, y), where x is the BQ score and y is the MQ score. In exemplary embodiments, combined BQMQ scores of 10-50 (for each parameter) are typically used, such as BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), and (30, 40).
[0206] Without being bound by any particular theory, in some aspects, the removal step filters "noise" markers with low base quality and / or mapping quality from a list of markers initially identified as being strongly associated with disease. In some embodiments, the removal step may involve taking each marker that meets a threshold probability of detection (PD), classifying the marker as signal or noise based on the marker's ROC curve, and removing the marker from the list if classified as noise. Alternatively, for example, the detection probability (PD) versus noise probability (P N A scoring system including a ratio of ) can be used to eliminate markers that do not meet a pre-set threshold score.
[0207] In addition to the BQ and MQ, the read position (RP) can also affect signal quality. Therefore, other factors, such as the intra-read position (RP or PIR), can be used to filter out artificial noise. In the context of sSNV or indel markers, the RP can be mapped, for example, by mapping the first base position of the sequencing read. Other factors that affect marker quality include specific sequence contexts, which are associated with a higher probability of sequencing errors (Chen et al., Science, 355(6326):752-756, 2017). In this regard, true mutations are often mappable to their own specific sequence contexts, while errors are not. For example, tobacco-related mutations tend to occur in CC contexts, while mutations associated with the activity of APOBEC enzymes prefer TpC contexts to insert somatic mutations (see Greenman et al., Nature, 446(7132):153-158, 2007). Thus, sequence context helps identify changes that are likely to be due to sequencing artifacts and changes that are likely to be due to dominant mutational processes.
[0208] For noise cancellation where the marker is a CNV, artificial noise is cancelled based on multiple parameters specific to the CNV. In some embodiments, the CNV-specific noise parameters include the "location attributes" of the CNV. Typically, centromeres, telomeres, and / or heterochromatic regions of chromosomes have extensive diversity due to their involvement in rearrangements. CNVs located in or near these regions (even detected via in situ methods via computer software) may be undesirable. In some embodiments, the location attributes of a CNV can be measured based on whether it is at least 1,000 kilobases (kb), at least 400 kb, at least 100 kb, or at least 20 kb or less, e.g., 1 kb from the telomere, centromere, or heterochromatic region of a chromosome. In some embodiments, CNVs located in subtelomeric or pericentromeric regions, which are characterized by chromosomal rearrangement hotspots, are undesirable. One additional feature that can be used in the methods of the present disclosure includes read position or read location. Read position information may be obtained by a variety of techniques using different location measurements, such as the genomic coordinates of the read, the position on a reference sequence, or the chromosomal location. In further embodiments, unique molecular indexing (UMI) and read position may be combined to generate folded reads.
[0209] In some embodiments, the CNV-specific noise parameter includes an assessment of the "representativeness" of a diseased CNV. For example, previous studies have found that calling CNVs in the immunoglobulin region tends to be unrepresentative of gDNA and substantially dependent on the DNA source—e.g., saliva vs. blood or lymphoblastoid cell lines vs. blood—(Need et al., 2009; Wang et al., 2007; Sebat et al., 2004). Such unrepresentative CNVs are undesirable.
[0210] In some embodiments, CNV-specific noise parameters include an assessment of the "depth coverage" of a CNV, which refers to the number of unique reads whose mapping overlaps with a particular genomic coordinate in the CNV genome segment.
[0211] Once the noisy markers have been filtered, the next step in the diagnostic method involves integrating the genome-wide inventory signals from the plasma sample into a mathematical inference model that outputs an estimated fraction of tumor DNA in the biological sample (e.g., plasma). Depending on the marker, the mathematical model integrates multiple process quality criteria as well as patient-specific attributes to estimate the tumor fraction (TF). Recognizing the fundamental differences between SNVs (or indels) and CNVs (SVs) in terms of frequency and association properties with traits (e.g., cancer), the disclosed systems and methods involve the use of marker-specific mathematical algorithms to estimate the tumor fraction.
[0212] From a workflow perspective, CNV-based detection methods may implement variations of the aforementioned SNV-based detection methods. In some embodiments, baseline samples (e.g., plasma samples and / or tumor samples) and normal cell samples (e.g., PBMCs) are processed and analyzed separately. In the final analysis step, tumor signals are binned separately from PBMC signals, for example, based on directional coverage skew and local fragment size skew. If a signal is identified as being from the tumor (tumor CNV / SV), the mathematical model used to estimate the tumor fraction is forward-directional; conversely, if a signal is identified as being from the PBMC, the mathematical model used to estimate the tumor fraction is backward-directional. While tumor fraction can be estimated using only the tumor sample (i.e., without using a PBMC sample), the method preferably integrates bidirectionally (i.e., both tumor-based and PBMC-based tumor fraction estimates are integrated).
[0213] As with SNV-based detection methods, CNV-based detection methods also allow for orthogonal integration of secondary features, e.g., fragment size shifts, where a mathematical formula incorporating directional features is used to estimate the putative tumor FractionThe main methods for determining (eTF) have been covered with preliminary applications, particularly tumor-based eTF estimation using CNVs. However, to make prognostic / diagnostic methods more robust, accurate, and / or sensitive, read-based features, such as DNA fragment size shifts, may be orthogonally integrated into the model. The significance of orthogonal features (in determining MRD) may be determined using a generalized linear model (GLM) to orthogonally determine tumor fraction based on the relationship between CNV depth coverage and fragment size shifts.
[0214] In some embodiments, CNV-based methods are performed to remove germline markers from baseline samples (usually plasma samples containing tumor samples) and normal samples (usually PBMCs). Next, artificial CNV sites are generated across a cohort of healthy plasma samples (a panel of normal PON blacklists), and mutations detected in patients are removed to remove common sequencing or alignment artifacts, such as centromeres and repetitive regions. Regions of interest (ROIs) encompassing all genomic segments of the tumor (sT_CNV) and PMBC (sP_CNV) are then binned into discrete windows (≥500 bp), and the depth coverage (read count) in each window is estimated from plasma samples at follow-up (post-surgery, during treatment, and recurrence follow-up). The median depth coverage per window is calculated and divided by the mean sample coverage.
[0215] Next, depth coverage values were normalized, and two LOESS regression curve fits were performed on the bin-wise GC fraction and mappability scores to correct for GC content bias and mappability bias. Further batch effect correction was performed using stable z-score normalization, applied separately to each sample. Briefly, the median and median absolute deviation (MAD) were calculated based on the neutral region of each sample, and then all CNV bins were normalized by (B(i)-Median) / MAD. Next, for each bin, depth coverage skew and fragment size center of mass (COM) skew were calculated compared to a panel of normal (PON) healthy plasma samples. Here, low tumor fraction samples exhibit sparse depth coverage skew, biased by the directionality of CNV segment amplification segments, while deletions exhibit a bias toward negative depth coverage skew. On the other hand, neutral regions exhibit random distortion with no preferred directionality, and therefore multiplying the differential (plasma PON) depth coverage distortion by the directionality of the CNV segment (amplifications multiplied by +1, deletions multiplied by -1) sums up the CNV signal across the genome, while neutral region noise is canceled out due to random directionality.
[0216] This process is performed mathematically, and tumor fraction is estimated by determining the linear dilution ratio between the cumulative signal detected in the plasma sample compared to the cumulative signal detected in the tumor. To address noise variability between patients with different CNV patterns, patient-specific CNV patterns are used to calculate the expected noise distribution across a cohort of healthy plasma samples (panel of normals, PON). Essentially, a similar process to that used for SNV marker analysis can be used to detect patient-specific patterns in healthy plasma samples (PON) or other patients (inter-patient analysis). This detection represents a background noise mode, where the mean and standard deviation (μ, σ) of the artifactual mutation detection rate are calculated. If the patient's detected tumor fraction (e.g., an artifactual tumor fraction corresponding to an error rate above the mean of 1.5 × σ) is higher than a threshold, reliable tumor detection and tumor fraction estimation are achieved.
[0217] It may also be possible to infer tumor fraction from directional genome-wide depth coverage skew in sP_CNVs, for example, using the reverse method in a workflow. Finally, orthogonal features may be integrated into this computational model to improve the stability, accuracy, sensitivity, or specificity of the algorithm and method. In some embodiments, the disclosed methods include TF estimation based on the detection of multiple SNV markers. Here, estimated TFs (eTFs[SNV]) were calculated by integrating process-quality criteria, including estimated genome coverage and sequencing noise, with patient-specific parameters, including mutation burden (N). Preferably, the method comprises calculating the estimated tumor fraction (eTF) for an SNV marker, where eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the total tumor-specific detections in the patient sample, σ is an empirically estimated measure of noise, R is the total number of unique reads in the region of interest (ROI), N is the tumor mutation burden, and cov is the average number of unique reads per site in the ROI.
[0218] In some embodiments, the disclosed methods include estimating TFs based on the detection of multiple CNV markers, where estimated TFs (eTFs[CNV]) were calculated by integrating the directionality of coverage depth skewed consistent with tumor CNV directionality, which is positively skewed for copy number gains and negatively skewed for copy number losses. Preferably, the method comprises calculating an estimated tumor fraction (eTF) for a CNV marker, where eTF[CNV]=(sum_{i]=[(P(i)-N(i)]*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i)]-E(σ)], where P is the median depth in the genomic window indexed by {i}, representing plasma depth coverage; T is the median depth in the genomic window indexed by {i}, representing tumor depth coverage; and N is the median depth in the genomic window indexed by {i}, representing normal depth coverage.
[0219] In one embodiment, determining the TF score may involve constructing an optimized base / mapping quality filter, using an optimal receiver operating point to filter SNV noise, and analyzing the filtered SNV signal using the integrated mathematical model described above. A representative method is shown in Example 2, with results shown in Figure 2. Error rate distributions may be evaluated across multiple replicates using control and tumor samples. A theoretical threshold cutoff value may be established using a statistical model (e.g., a binomial model) against which empirical measurements are plotted and the mean / confidence interval for each measurement calculated. A noise level is identified within the distribution using statistical modeling. A baseline tumor fraction (TF) that may be diagnostic of tumors is established based on statistical measurements. As seen in the data in Figures 3D-3G, a baseline TF value of approximately 1 x 10 -5 A tumor fraction greater than 100% represents minimal residual disease in most solid tumors, including melanoma, lung, and breast tumors.
[0220] In one embodiment, determining the TF score may involve constructing an appropriate filter for filtering CNV noise and analyzing the filtered CNV signal using the integrated mathematical model described above. A representative method is described in Example 3, and the results are shown in Figure 5. First, genetic data from resected tumors, germline (e.g., PBMCs), and pre-operative biological samples (preferably cfDNA) are obtained. Profiles of tumor read depth, germline read depth, and pre-operative plasma cfDNA read depth in a representative amplified segment (e.g., 500 kb; preferably 100 kb) are generated. The depth coverage is normalized across all samples to minimize bias. As described above, an integrated mathematical model that integrates genome-wide read depth distortion is used to evaluate the differences between the three sample genomes. The results show that the above method has high detection sensitivity when integrating genome-wide CNV patterns. More specifically, the above method can demonstrate the surprising and unexpected ability to detect tumors with TFs as low as approximately 1 / 100,000. This feature is evident from the signal-to-noise (SNR) for each TF, which is 10 -5All the above TFs show positive (>0) detection of signal compared to noise.
[0221] An exemplary system using the methods of the present disclosure is shown in Figures 7A-C. Here, a list of genetic markers is received from a subject (e.g., a cancer patient). The list of genetic markers includes, for example, tumor DNA (e.g., obtained from a resected tumor) and control DNA (e.g., PMBC). Mutation calling is used to analyze the genetic data, and somatic SNVs (sSNVs) are set as references for downstream analysis. In some embodiments, this reference standard can be individualized, for example, for a particular subject. In some aspects, this reference standard can be used with a cohort of additional reference standards.
[0222] Preferably, the outputs of three different variant callers, MUTECT, LOFREQ, and STRELIKA, are crossed to utilize a highly clean and high-quality reference set. MUTECT provides reliable and accurate identification of somatic point mutations in next-generation sequencing data of cancer genomes (Cibulskis et al., Nature Biotechnology, 31, 213-219, 2013); the LOFREQ model determines the run-specific error rate for accurate calling of variants occurring in <0.05% of the population (Wilm et al., Nucleic Acids Res., 40(22):11189-11201, 2012); and STRELIKA is an analytical package designed to detect somatic SNVs and small indels from aligned sequence reads of matched tumor-normal samples (Saunders et al., Bioinformatics, 28(14):1811-7, 2012).
[0223] Mutation calling intersections typically involve the use of multiple art-known calls. In some embodiments, three mutation calling methods (MUTECT, LOFREQ, and STRELIKA) are used on patient tumor and normal sequencing reads, and the intersection mutant list is defined as mutants that show detection of the exact same substitution (same genomic coordinate and nucleotide change) in all calls.
[0224] Next, reads from the patient-specific mutation sites are collected and filtered. In some embodiments, the collecting and / or filtering steps include removing reads with low mapping quality. For example, any reads with a mapping quality score less than 29 (ROC optimized) are filtered. Additionally or alternatively, filtering can include constructing duplicate families. For example, duplicates can include multiple PCR / sequencing copies of the same DNA fragment (i.e., overlapping non-unique marker and subject regions). Finally, corrected reads can be generated based on consensus testing. The filtering step can include removing reads with low base quality. For example, any reads with a base quality score less than 21 (ROC optimized) can be filtered. Finally, the filtering step can include removing reads with high fragment sizes. For example, any reads with a fragment size greater than 160 (ROC optimized) can be filtered. The rationale for this is that tumor DNA tends to be shorter than normal DNA, so filtering low fragment sizes enriches for tumor DNA. See Jiang et al., PNAS USA, 112.11(2015):E1317-E1325; and Mouliere et al., bioRxiv, 134437, 2017.
[0225] The next step is to calculate the number of patient-specific mutation sites with at least one supporting read (in the filtered set) using the exact same substitution as the tumor. In situations where the marker is an SNV, the calculation step may include integrating a probabilistic model including: 1) the integrated signal of plasma SNV detection; 2) process quality measures including estimated genome coverage and a sequencing noise model; and 3) patient-specific parameters including mutation burden (N). More specifically, the integrated mathematical model may include calculating the estimated eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific SNVs detected in the patient plasma sample, σ is an empirically estimated measure of error rate, R is the total number of unique reads in the SNV enumeration region of interest (ROI), N is the tumor mutation burden, and cov is the average number of unique reads per site in the SNV enumeration ROI. The estimated TFs are then checked against a detection threshold defined by empirically measured baseline noise TF estimates from healthy samples. In some embodiments, a TF is defined as detected if it exceeds a threshold, eg, two standard deviations of the noise TF distribution (eg, FPR<2.5%).
[0226] In some embodiments where the marker is a CNV, the filtering step may include calling the CNV (e.g., analyzing amplifications and / or deletions) on tumor and patient-derived normal (e.g., PBMC) samples, and generating reference segments for all CNV segments that meet a threshold characteristic (e.g., greater than 5 megabase pairs in length) along with the directionality of the change (where amplifications are positive factors, e.g., +1, and deletions are negative factors, e.g., -1). Next, single-base-pair depth coverage information was collected for plasma, tumor, and PBMC samples covering the patient-specific CNV segmentation ROI. Next, the patient-specific CNV segmentation ROI was normalized to 500-bp windows, and the median value per window was calculated for all samples and windows (artificial suppression). Next, normalized depth coverage information for all 500-bp windows was generated.
[0227] In some embodiments, normalization can be performed using (1) per-sample stable z-score normalization and / or (2) stable principal component analysis (RPCA). For example, the z-score method can include using the algebraic function preop_median=(preop_median-median(preop_median)) / (1.4826*mad(preop_median,1)). Alternatively, the stable principal component analysis (RPCA) method can include solving an optimization problem for M=L+S to remove noisy high-frequency artifacts (the S matrix). A combination of such methods can also be used.
[0228] Next, the reads / windows derived from the patient-specific segmentation are filtered. In some embodiments, filtering may include removing reads with low mapping quality (e.g., <29, ROC optimized); removing reads close to centromeric regions, e.g., removing windows with normalized normal values above a threshold (e.g., 10). With the centromeric proximity filter, it has been determined that ~70%-80% of CNV noise co-localizes with centromeric regions and can be detected by abnormally high depth coverage in PBMC samples. Such centromeric hotspots may be removed in the filtering step.
[0229] Next, non-expressed regions in the cfDNA are removed. For example, windows not included in the cfDNA list mask constructed from multiple cfDNA samples can be removed. The rationale for this filtering step is that if cfDNA is biased to only represent nucleosome-protected genomic regions and to represent non-listed gaps in accessible chromatin genomic regions, including these non-listed regions in the calculation is likely to be a source of bias and error. Therefore, a mask of regions represented (>0 reads) in the cfDNA cohort is generated using the cohort of cfDNA samples.
[0230] Computational methods are then used to integrate coverage parameters across plasma and normal samples. Thus, the directional depth of skewed coverage between plasma and normal (PBMC) patient samples can be integrated using the equation [(P(i)-N(i)*sign[T(i)-N(i)]-E(sigma)]. Similarly, the cumulative depth of skewed coverage between tumor and normal (PBMC) patient samples can be integrated using the equation [abs(T(i)-N(i)]-E(σ)].
[0231] Next, the dilution ratio between the signals, i.e., the dilution ratio for directional depth and cumulative coverage depth, is calculated, which corresponds to the estimated tumor fraction (eTF). In some aspects, the calculating step may include calculating the eTF of the CNV marker using a stochastic dilution model including: 1) integrating the direction of skewed coverage depth between plasma and normal (PBMC) patient samples, consistent with the tumor CNV direction in which copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative product of skewed coverage depth between tumor and normal (PBMC) patient samples; and 3) determining the dilution ratio between the signals. More specifically, the integrated mathematical model involves calculating an estimated eTF[CNV] = (sum_{i}[(P(i)-N(i)]*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i)]]-E(σ)), where P is the median depth coverage in the genomic window indexed by {i}, representing plasma depth coverage, normalized by either a stable z-score method or a stable PCA method compared to a cohort of normal samples, and T is the median depth coverage in the genomic window indexed by {i}, representing tumor depth coverage, normalized by either a stable z-score method or a stable PCA method. where N is the median depth in the genomic window indexed by {i}, normalized by either the z-score method or stable PCA compared to a cohort of normal samples, and N is the median depth in the genomic window indexed by {i}, normalized by either the stable z-score method or stable PCA compared to a cohort of normal samples. The estimated TFs (CNVs) are then checked against a detection threshold defined by empirically measured baseline noise TF estimates from healthy samples. In some embodiments, an eTF (CNV) is defined as detected if it exceeds a threshold, e.g., 2 standard deviations of the noise TF distribution (e.g., FPR<2.5%).
[0232] In some embodiments, a probabilistic model is used to calculate the effective coverage per genomic site based on the mathematical operation A*PBMC_cov+B*tumor_cov, where if a particular site is associated with an amplification or deletion, the PBMC coverage and the tumor coverage are not the same, and A+B=1. In certain embodiments, A and B for various samples are as follows: control (e.g., PBMC sample) A=1 and B=0; tumor sample B=purity and A=1 purity; plasma sample B=TF and A=1-TF. In some embodiments, the relationship between signal in plasma and tumor is linearly related to the dilution (or change in the mixing ratio) between purity and TF. As is known in the art, the model is also subject to noise, which may be included in probabilistic models.
[0233] Use of this method in treating post-operative patients The prognosis of cancer patients who have undergone surgical removal of their tumor (e.g., removal of a breast tumor by mastectomy; removal of a lung tumor by pneumonectomy or lobectomy; or prostatectomy for prostatectomy) is extremely important. For example, in the case of breast cancer, it is reported that the majority of women considering adjuvant therapy would prefer to be informed of their prognosis without adjuvant therapy (Ravdin et al., J Clin Oncol., 16(2):515-521, 1998). Adjuvant therapy is unpleasant, inconvenient, and undesirable (Ravdin et al., J Clin Oncol., 16(2):515-521, 1998). In some cases, it provides only marginal benefit (Simes et al., J Natl Cancer Inst Monogr., 30, 146-152, 2001). The decision to perform it is legitimate (Duric et al., supra). This includes the trade-offs described by Wouters et al. (Ann Oncol., 24(9):2324-2329, 2013) and calls for refinement of cancer risk determinations (Kratz et al., Transl Lung Cancer Res., 2(3):222-225, 2013).
[0234] Many studies have indicated that tumor size is an important prognostic variable. However, in the setting of MRD, tumors are generally not detectable using conventional diagnostic tools such as CT scans, making tumor size inappropriate. Therefore, tumor size cutoff values are problematic.
[0235] Thus, computational prediction models provide an important step in this direction and may be the most accurate prediction method currently available. Figure 7 shows model predictions in post-surgical patients based on estimated tumor fraction. For example, estimated tumor fractions above a threshold (e.g., SNV markers approximately 10 -4 , and / or SNV markers of approximately 10 -5 ) indicates that the subject requires adjuvant therapy.
[0236] This model is useful not only for patient counseling but also for physician decisions regarding adjuvant therapy. Thus, the disclosed method provides physicians and clinicians with a tool to predict outcomes (e.g., metastasis or death) in the absence of adjuvant therapy. Presumably, patients with very low baseline risk as a function of estimated tumor fraction (eTF) will want to avoid the toxicity associated with adjuvant therapy. Thus, the predictive tool can be an effective decision support. This predictive tool may also be useful as a benchmark for determining the predictive ability of new treatments (e.g., the use of investigational drugs), such as chemotherapy, immunotherapy, and targeted therapy.
[0237] 〔system〕 The present disclosure further relates to systems for implementing the methods of the present disclosure. A representative system is provided in the schematic diagram of FIG. 7A, which shows an exemplary system for implementing the diagnostic methods of the present disclosure. As shown herein, a system 500 is provided that may include an analysis unit 510, a classification unit 520, a computing unit 530, and a display 540 that outputs data and receives user input via associated input devices (not shown). The analysis unit 510 typically receives input of genetic data, e.g., a VCF file containing reads from a subject's tumor sample, optionally a normal (e.g., PBMC) sample, and a second biological sample, e.g., a plasma sample from the same subject (Note: The first and second sample collections may be performed simultaneously or sequentially, i.e., temporally separated). The classification unit 520 may include one or more engines for classifying various types of markers, e.g., CNV / SV versus SNP / indel. Note that FIG. 7A shows one configuration of the system. The orientation and configuration of the components may be changed as needed. Furthermore, additional components may be added to this system. The various components, their various operations, their various orientations, and their various relationships with one another are discussed in detail below.
[0238] In some embodiments, the present disclosure relates to a system for detecting residual disease in a subject in need thereof. The system may include an analysis unit 510 configured and arranged to filter genome-wide noise markers from a genome-wide list of markers generated from a plurality of genetic markers from a biological sample of the subject, the biological samples including a tumor sample and a normal cell sample, the list of genetic markers being selected from the group consisting of single nucleotide variations (SNVs), indels, copy number variations (CNVs), structural variations (SVs), and combinations thereof, the analysis unit further including detecting the list of genome-wide genetic markers in a second biological sample to generate a list of tumor genome-wide genetic markers in the second sample, and the analysis unit further includes a classification engine 520. In some embodiments, the classification engine 520 statistically classifies each marker in the list as signal or noise. For example, if the marker is an SNV or an indel (grouped together due to similar structural features, but not necessarily using the same classification scheme), the classification engine can use noise (P N Similarly, if a marker is an SNV or indel (grouped due to similar structural features, but not necessarily using the same classification scheme), the classification engine classifies the SNV or indel as signal or noise based on 1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads that include the CNV or SV window, or 3) the representation of the CNV or SV window in the cfDNA data.
[0239] In some embodiments, the SNV / indel classification unit 520 calculates the noise (P N) based on the probability of detection. In some embodiments, the CNV / SV classification unit 520 statistically classifies each CNV / SV in the list as signal or noise based on its location relative to the centromere, its non-listing at a given coverage depth, and its readability. In some embodiments, the classification unit 520 classifies both SNV / indel markers and CNV / SV markers based on one or more of the aforementioned parameters.
[0240] In some embodiments, the disclosed system generates a predicted tumor profile for a sample based on one or more integrated mathematical models. Fraction For example, the computing unit may include a computing unit 530 configured and arranged to calculate a putative tumor function (eTF) of the sample based on one or more integrated mathematical models that are specific for SNV / indel markers or specific for CNV / SV markers. Fraction The computing unit may be configured and arranged to calculate the eTF (eTF). In such embodiments, if the marker is an SNV / indel, the computing unit may integrate process-quality metrics including estimated genome coverage and sequencing noise with patient-specific parameters including mutation burden (N). Similarly, if the marker is a CNV or SV, the computing unit may calculate the eTF of the CNV marker by integrating the directional depth of coverage skewed consistent with tumor CNV directionality, where copy number gains are positively skewed and copy number losses are negatively skewed.
[0241] The system of the present disclosure further includes a listing unit 540 that outputs a residual disease profile of the subject based on the estimated tumor fraction, and if the estimated tumor fraction exceeds an empirical threshold calculated by the background noise model, the residual disease profile of the subject is output in the listing unit 540. In some embodiments, in the system of the present disclosure, the classification engine unit and / or the arithmetic unit may be separately or collectively coupled to the listing unit that outputs a residual disease profile of the subject based on the estimated tumor fraction.
[0242] In some embodiments, the system 500 of the present disclosure comprises an analysis unit 510 comprising a classification unit 520. The classification unit 520 comprises at least one engine selected from the group consisting of an SNV classification engine 520-1, a CNV classification engine 520-2, an indel classification unit 520-3, a structural variant (SV) classification unit 520-4, or a combination thereof, wherein the SNV / indel classification engine is configured to classify noise (P N ) based on the detection probability of noise (P N ) statistically classifies each SNV in the list as signal or noise as a function of the SNV's base quality (BQ) and the SNV's mapping quality (MQ), and / or the CNV / SV classification engine statistically classifies each CNV / SV in the list as signal or noise based on its position relative to the centromere, the predetermined coverage, and the readability. The system 500 further calculates a putative tumor profile for the sample based on one or more marker type-specific integrated mathematical models. FractionFor example, if the marker is an SNV, the computing unit 530 may be configured to calculate the eTF based on the mathematical model eTF[SNV]=1-[1-(ME(σ)R] / N]^(1 / cov), where M is the known number of tumor-specific detections in the patient sample, σ is an empirically estimated measure of noise, R is the total number of unique reads in a region of interest (ROI), N is the tumor mutation burden, and cov is the average number of unique reads per site in the ROI. Similarly, if the marker is a C ... CNV] = (sum_{i}[(P(i)-N(i)]*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i)]]-E(σ)), where P is the median depth in the genomic window indexed by {i} representing plasma depth coverage, T is the median depth in the genomic window representing tumor depth coverage indexed by {i}, and N is the median depth in the genomic window indexed by {i}.
[0243] In some embodiments, the computing unit 530 may be configured to calculate eTFs based on an indel-specific mathematical model (generally similar or identical to the mathematical model for calculating eTFs for SNPs). In some embodiments, the computing unit 530 may be configured to calculate eTFs based on an SV-specific mathematical model (generally similar or identical to the mathematical model for calculating eTFs for CNVs). In some embodiments, the computing unit 530 may be configured to calculate eTFs based on an SNP-specific mathematical model comprising the formula eTF[SNV]=1-[1-(ME(σ)R) / N]^(1 / cov), where M is the number of tumor-specific list detections in the patient sample, σ is an empirically estimated measure of noise, R is the total number of unique reads in a region of interest (ROI), N is the number of unique reads per site in the ROI, and Cov is the average number of unique reads per site in the ROI. , a mathematical model specific to CNVs that includes the formula eTF[CNV] = (sum_{i}[(P(i)-N(i)-N(i)]*[T(i)-N(i)]-E(sigma)] / (sum_{i}[abs(T(i)-N(i)]]-E(sigma)]), where P represents the median genomic window depth {i}, T represents the range of tumor depth {i}, and N represents the range of normal depth {i}.
[0244] In some embodiments, the computing unit 530 is configured to integrate a probabilistic model to calculate the eTFs of the SNV or indel markers, the probabilistic model including: 1) the integrated signal of plasma SNV or indel detection, 2) process quality metrics including estimated genome coverage and a sequencing noise model, and / or 3) patient-specific parameters including mutation burden (N); and / or to calculate the eTFs of the CNV or SV markers using a probabilistic mixture model, the probabilistic dilution model including: 1) integrating the directional depth of skewed coverage between plasma and normal patient samples, consistent with tumor CNV or SV directionality, where copy number amplifications are positively skewed and copy number deletions are negatively skewed; 2) integrating the cumulative depth of skewed coverage between tumor and normal patient samples; and / or 3) finding a dilution ratio between the signals.
[0245] In various embodiments herein, a computer-readable medium is provided, the computer-readable medium comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for filtering noise in a list of genetic markers received from a sample of a subject, the genetic markers comprising SNVs (preferably sSNVs), CNVs (preferably sCNVs), indels, and / or SVs (preferably translocations, gene fusions or combinations thereof) in a genomic read. Preferably, the filter removes artificial noise markers from the genome-wide marker list by statistically classifying each noise SNV or Indel based on the probability of detection as a function of: 1) the mapping quality (MQ) of the reads containing the SNV, 2) the fragment size length of the reads containing the SNV, 3) a consensus test within the read duplication family containing the SNV or Indel, 4) the base quality (BQ) of the SNV or Indel and / or its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and 3) the list of CNV windows in the cfDNA data. The computer-readable medium may further include computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for calculating an estimated tumor fraction (eTF) of a biological sample based on one or more integrated mathematical models; and then diagnosing residual disease in a subject based on the estimated tumor fraction and an empirical threshold calculated by the background noise model.
[0246] In some embodiments, the system comprises a computation unit 530 comprising computer-executable instructions that, when executed by the processor, cause the processor to perform a method or sequence of steps for estimating tumor fraction (eTF) based on one or more of the above-described mathematical models for calculating eTF, and a diagnosis unit that makes a qualifying diagnosis based on the calculated eTF (e.g., a positive diagnosis is made if eTF≧2 std exceeds a noise threshold). The system may further comprise a display 540 that outputs data and receives user input via an associated input device (e.g., a mouse). In some embodiments, results may be listed on the display 540 in the form of a binary output (i.e., "+ve for MRD" or "-ve for MRD") or an ordinal score (e.g., on a scale of 1 to 5), where a score of 1 indicates that the subject is unlikely to have MRD and a score of 5 indicates that the subject is likely to have MRD.
[0247] As shown in Figure 7B, an exemplary system 100 is configured and arranged to detect residual disease in a subject in need thereof. Referring to Figure 7B, the system 100 may include an analysis unit 110 and a computing unit 150. The analysis unit 110 may include a pre-filter engine 120 and a correction engine 130. The system components and associated engines are described in further detail below.
[0248] 7B , the pre-filter engine 120 of the analysis unit 110 may be configured and arranged to receive a first subject-specific genome-wide read list associated with a plurality of genetic markers from a first biological sample of the subject. As discussed with respect to the workflow herein, according to various embodiments, the first biological sample may include a baseline sample, and the first read list may each include single base pair long reads, and the baseline sample may include a tumor sample or a plasma sample.
[0249] 7B may also be constructed and arranged to filter artificial sites from the first list of reads. As described in the workflows herein, according to various embodiments, filtering may include removing repetitive sites generated across a cohort of reference healthy samples from the first list of genetic markers, and / or identifying germline mutations in peripheral blood mononuclear cells of the normal cell sample and removing said germline mutations from the first list of genetic markers.
[0250] In FIG. 7B, the correction engine 130 of the analysis unit 110 may be configured and arranged to receive output from the engine 120. The correction engine 130 may also be configured and arranged to receive reads from a second subject-specific genome-wide list of genetic markers in a second biological sample from the subject and generate a tumor-associated genome-wide representation of the genetic markers in the second sample. As shown in FIG. 7B, the reads from the second biological sample may be detected using a detection unit 140. The detection unit 140 may or may not be part of the system 100, in which case the reads may simply be received by the correction engine 130 from an external system 100. Furthermore, the reads may be received by the analysis unit 110 at any point in the system prior to noise filtering, as described below. Furthermore, the reads may also be received after noise filtering, in which case the reads are provided to the system 110 with noise already filtered. Furthermore, the detection unit 140 may be integrated into the analysis unit 110, as shown in FIG. 7B, or may be separate from the analysis unit 110.
[0251] The correction engine 130 may also be configured and arranged to filter noise from the first and second lists of genome-wide reads using at least one error suppression protocol, generating a first filtered read set for the first list of genome-wide reads and a second filtered read set for the second list of genome-wide reads.
[0252] As described in the workflow herein, according to various embodiments, the at least one error suppression protocol can include calculating the probability that any single nucleotide variation in the first and second lists is an artificial variation and removing the variation.
[0253] As described in the workflows herein, in various embodiments, the probability may be calculated as a function of features selected from the group consisting of mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof.
[0254] As described in the workflow herein, and according to various embodiments, at least one error suppression protocol may include removing artifacts by testing for mismatches between independent copies of the same DNA fragment generated from polymerase chain reaction or sequencing processing, and / or by using overlap consensus whereby artifacts are identified and removed when a majority of a given overlap family is mismatched.
[0255] The computational unit 150 of the system 100 receives the output from the correction engine 130 and applies the background noise model to one or more integrated mathematical models to generate a putative tumor profile for the first and second biological samples using the first and second filtered read sets. Fraction The computing unit 150 may be configured and arranged to calculate: ##EQU1## The computing unit 150 may be further configured and arranged to detect residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold. The background noise model, the integral mathematical model, and the empirical threshold are discussed in detail herein.
[0256] The computing unit 140 of the system 100 may also include a display 160, as shown in FIG. 7B. The display may be configured and arranged to receive output from the computing unit 150. The output may include data related to the detection of residual disease in the subject / user. Alternatively, the system 100 may omit a display and instead transmit data output from the computing unit 150 to any form of storage or display device or location external to the system 100. Also, as described herein, the components of the system 100 may be integrated into one single unit or may be separated into more separate physical units than those shown in FIG. 7B. Furthermore, the system 100 may be part of a distributed network of systems, each performing substantially similar tasks and transmitting data from each system to a hub.
[0257] As shown in Figure 7C, an exemplary system 100 is configured and arranged to detect residual disease in a subject in need thereof. Similar to the exemplary system of Figure 7C, the system 100 may include an analysis unit 110 and a computing unit 150. In contrast to the system of Figure 7B, the analysis unit 110 of Figure 7C may include a pre-filter engine 120 and a normalization engine 130. Such system components and associated engines are described in further detail below.
[0258] 7C , the pre-filter engine 120 of the analysis unit 110 may be configured and arranged to receive a first subject-specific genome-wide read list associated with the genetic markers from a first biological sample of the subject. As discussed with respect to the workflow herein, according to various embodiments, the first biological sample may include a baseline sample, the first read list may each include reads of a single base pair in length, and the baseline sample may include a tumor sample or a plasma sample.
[0259] The pre-filter engine 120 may also be configured and arranged to receive a second subject-specific genome-wide list of reads associated with the genetic markers from a second biological sample of the subject. As discussed with respect to the workflows herein, according to various embodiments, the second biological sample may include a peripheral blood mononuclear cell sample (PBMC), and the second list of genetic markers may each include copy number variations (CNVs).
[0260] The pre-filter engine 120 may also be constructed and arranged to filter artificial sites from the first and second read lists. As discussed with respect to the workflow herein, according to various embodiments, filtering may include removing repetitive sites from the first and second read lists generated on a cohort of reference healthy samples; identifying shared CNVs between the first and second lists as germline mutations, and removing said mutations from the first and second read lists.
[0261] The normalization engine 130 of the analysis unit 110 may be configured and arranged to receive output from engine 120. The normalization engine 130 may also be configured and arranged to receive reads from a third subject-specific genome-wide list of genetic markers in a third biological sample from the subject to generate a tumor-associated genome-wide representation of the genetic markers in the second sample.
[0262] As shown in FIG. 7C, a readout of the third biological sample may be detected using a detection unit 140. The detection unit 140 may or may not be part of the system 100, in which case the readout may simply be received by the normalization engine 130 from an external system 100. Furthermore, the readout may be received by the analysis unit 110 at any point in the system prior to noise filtering, as described below. Furthermore, the readout may also be received after noise filtering, in which case the readout is provided to the system 110 with noise already filtered. Furthermore, the detection unit 140 may be integrated into the analysis unit 110 or separate from the analysis unit 110, as shown in FIG. 7C.
[0263] The normalization engine 130 may also be configured and arranged to normalize each of the first, second, and third lists of reads to generate a first filtered set of reads for the first genome-wide list of reads, a second filtered set of reads for the second genome-wide list of reads, and a third filtered set of reads for the third genome-wide list of reads. Normalization methods are discussed in detail herein and may be used in any combination contemplated to normalize reads as discussed.
[0264] The computing unit 150 of the system 100 in FIG. 7C receives the output from the normalization engine X30 and calculates the estimated tumor density of the third biological sample. Fraction The computing unit 150 may be configured and arranged to calculate an eTF (eTF) using a third filtered read set, for example, by applying a background noise model to one or more models that generate a first eTF using the first filtered read set and / or one or more models that generate a second eTF using the second filtered read set. The computing unit 150 may be further configured and arranged to detect residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold. Background noise models, integral mathematical models, and empirical thresholds are discussed in detail herein.
[0265] System 100 may also include a display 160, as shown in FIG. 7C. The display may be constructed and arranged to receive output from computing unit 150. The output may include data related to the detection of residual disease in the subject / user. Alternatively, system 100 may omit a display and instead transmit data output from computing unit 150 to any type of storage or display device or location external to system 100. Also, as described herein, the components of system 100 may be integrated into one single unit or separated into separate physical units other than those shown in FIG. 7C. Furthermore, system 100 may be part of a distributed network of systems, each performing substantially similar tasks and transmitting data from each system to a hub.
[0266] Other Related Embodiments
[0267] [Estimation of transplant rejection] The present disclosure further relates to predicting transplant rejection using the above systems, methods and algorithms. Preferably, transplant rejection may be predicted using the SNV / indel-based workflow outlined in Figures 1B and 1D.
[0268] In some embodiments, the prediction of transplant rejection is based on a protocol that utilizes reference SNPs that are specific only to the donor (and not present in the recipient). Based on the detection rate of the donor-specific SNPs in the recipient's blood (e.g., post-transplant), the donor-DNA fraction can be calculated using the disclosed methods and systems.
[0269] The donor-DNA fraction is expected to correlate with the apoptosis or rejection rate of the transplanted tissue, e.g., a high donor-DNA fraction is associated with a high rejection phenotype, and a low donor-DNA fraction is associated with a low rejection phenotype.
[0270] In some embodiments, the differential SNPs between donors and recipients measured using the methods of the present disclosure can be used to estimate the proportion of donor DNA (eDF) in the recipient's blood sample. The probability / likelihood of transplant rejection is calculated based on the eDF. For example, if the eDF is greater than a certain threshold, it indicates that the transplanted tissue will be rejected by the host or is incompatible with the host. Conversely, if the eDF is below a threshold level, it indicates that the transplanted tissue will be accepted by the host or is compatible with the host.
[0271] Non-invasive prenatal testing (NIPT) for chromosomal abnormalities
[0272] The present disclosure further relates to non-invasive prenatal testing for chromosomal abnormalities using the above-described systems, methods, and algorithms. Preferably, NIPT can be performed using the CNV / SV-based workflow outlined in Figures 1C and 1E. Herein, known amplifications and deletions are used as the CNV reference set against which a subject's sample (e.g., amniotic fluid or blood from a pregnant woman carrying a fetus with a suspected chromosomal abnormality) is measured. The workflows in Figures 1C and 1E are designed to detect copy number variation changes, even when the signal is low and sparse, assuming the subject's segment and direction (amplification, deletion) are known. In the context of NIPT, assume that testing for trisomy 21 in maternal blood is of interest, and both the region of interest (chromosome 21) and the direction of the variation (amplification) are known. [Example]
[0273] It will be understood that the structures, materials, compositions, and methods described herein are intended to be representative examples of the present disclosure, and that the scope of the present disclosure is not limited by the scope of the examples. Those skilled in the art will understand that the present disclosure can be practiced with variations on the disclosed structures, materials, compositions, and methods, and that such variations are considered to be within the scope of the present disclosure.
[0274] Example 1: Methods and systems for the detection and validation of tumor-specific low-abundance tumor markers and their use in cancer diagnosis
[0275] The disclosed systems and methods are useful in detecting minimal residual disease. As is known in the art, in contrast to metastatic cancer (characterized by high disease burden and significantly elevated ctDNA), the abundance of ctDNA limits the use of targeted sequencing techniques in the context of residual disease detection. Given the known limited amount of cfDNA in settings with low tumor burden, we first investigated the possibility of optimizing cfDNA extraction. First, to reduce variability arising from sample acquisition and interindividual variability, commercially available extraction kits and methods were compared using homogenous cfDNA material generated through large-volume plasma collection (approximately 300 cc) via plasmapheresis of healthy subjects and cancer patients undergoing hematopoietic stem cell collection. This large volume of plasma allowed multiple method and protocol parameters to be tested on the same cfDNA input, allowing subtle differences in yield and quality to be accurately measured.
[0276] Extraction was performed on large 1 ml plasma samples using kits and reagents from Capital Biosciences (Gaithersburg, MD, USA; Catalog #CFDNA-0050), Qiagen (Germantown, MD, USA), Zymo (Irvine, CA, USA; Catalog #D4076), Omega BIO-TEK (Norcross, GA, USA; Catalog #M3298), and NEOGENESTAR (Somerset, NJ, USA; Catalog #NGS-cfDNA-WPR) according to the manufacturer's instructions. Multiple plasma aliquots were processed in parallel to assess inter- and intra-method variability. The yield and purity of each recovered cfDNA sample were measured using fluorimetry (total mass), UV absorbance (detection of salt and protein contaminants), and on-chip electrophoresis (size distribution and gDNA contamination).
[0277] Results demonstrated that the Omega BIO-TEK MAG-BIND cfDNA extraction kit outperformed all other tested methods. Further systematic optimization of each step in the manufacturer's protocol reduced contaminant carryover and improved cfDNA recovery. Nevertheless, cfDNA yields in early-stage NSCLC (n=21) were low and highly variable (median 5 ng / ml (<1000 genome equivalents); range 3-30 ng / ml).
[0278] The above data support the hypothesis that the detection of single point mutations in patient plasma samples results from two sequential statistical sampling processes: (i) the probability that a mutant fragment is sampled in the limited number of genome equivalents present in a typical plasma sample, and (ii) the probability that a mutant fragment is detected in the sample based on its abundance, sequencing depth, and sequencing error (signal-to-noise). While the latter process has been the focus of intensive research and technological development by the scientific community (e.g., ultra-deep error-free sequencing protocols), the former stochastic process has rarely been addressed. Nevertheless, in low-disease burden ctDNA detection, both processes play important roles, as shown in Figure 2. In the absence of a physical fragment containing the target point mutation, even ideal ultra-deep targeted sequencing will fail to discover a cancer signal. In practice, this problem is further complicated by the fact that a single observation (mutation sequencing read) is rarely sufficient for reliable detection.
[0279] Thus, the genome equivalents present in a plasma sample constitute a random sampling of the entire pool of cfDNA fragments circulating in a patient, which can be formulated by the Bernoulli trial random sampling model. This model predicts that the detection probability (TF<1%) among TFs associated with early cancer regimens drops rapidly for low TFs. Even at a frequency of 0.1% (1 / 1000), the detection probability is predicted to be lower than 0.65 (Figure 2A). However, by implementing extensive sequencing, repeated Bernoulli trials at multiple sites can compensate for the limited coverage per site (a function of the limited genome equivalents). Using this model, we found that integrating more than 20,000 point mutations (approximately 10 mutations / mb found in 17% of human cancers), as easily achievable with standard whole genome sequencing (WGS), yields a high detection probability (up to 0.98) even at a TF of 1:100,000 (e.g., 20x coverage in Figure 2B).
[0280] The optimized extraction protocol was then applied to patient samples. This cohort included six postoperative (~14 days) plasma samples collected from the same patients for minimal residual disease (MRD) estimation and four plasma samples collected from benign patients (controls). Despite optimal extraction, cfDNA yields in low-disease burden samples were low and showed high variability between patients, ranging from 0.13 ng / mL to 1.6 ng / mL. The data confirm that the number of DNA molecules available for cfDNA sequencing is low and variable.
[0281] Taken together, the present results demonstrate that in the context of MRD detection, limited input material constitutes a major barrier to the effective application of ultra-deep targeted sequencing, given that the number of genome equivalents is much lower than the applied sequencing depth (minimal ctDNA frequency of 0.1–1%).
[0282] Example 2: Genome-wide integration enables sensitive WGS-based NSCLC ctDNA detection of post-surgical residual disease, enabling adjuvant therapy stratification and treatment optimization . Ultrasensitive identification of MRD in cfDNA may have fundamental prognostic significance and enable patient stratification for follow-up adjuvant chemotherapy. Current approaches primarily aim to extend the paradigm of read-driver hotspot mutation detection by increasing depth sequencing to counter the low fraction of ctDNA in cfDNA. Nevertheless, such approaches are inherently limited by the upper limit of genome equivalents. To overcome this limitation, genome-wide information has been integrated. This is based on the rationale that pooling information across the genome could take advantage of the high mutation rate in lung cancer. Thus, mutation detection has been broadened genome-wide, increasing sensitivity without relying on deeper sequencing of a small number of sites. Therefore, WGS has been applied to base-sensitive detection of the cumulative signal provided by the 10,000–30,000 somatic mutations observed in a significant proportion of NSCLC. Notably, the majority of these mutations are thought to occur before transformation and are therefore likely to be present even in early-stage NSCLC. To evaluate this approach for residual disease detection in NSCLC patients after curative surgery, samples from five patients with early-stage lung cancer were analyzed (full clinical details are shown in Table 1). [Table 1]
[0283] Initial WGS was performed to generate a patient-specific genome-wide sSNV inventory using matched tumor and germline DNA from peripheral blood mononuclear cells (PBMCs). Plasma samples were collected preoperatively and approximately 14 days after surgical resection. cfDNA was extracted using an optimized MAG-BIND cfDNA Extraction Kit, and libraries were prepared using as little as 1 ng of patient cfDNA.
[0284] Next, point mutation pattern matching was used to detect MRD. To this end, a robust mathematical model was constructed to estimate the tumor fraction of SNV and CNV markers. The mathematical model indicates that increasing the number of sites results in a significant increase in the probability of detection. To validate this prediction, we simulated cfDNA detection using in silico mixtures of tumor and normal WGS data from multiple lung adenocarcinoma patients, mixing tumor and normal WGS reads in various ratios and acquiring virtual plasma samples (10-2 to 10-6, n = 5 replicates each) for different TFs. To simulate noise and potential false positives, a complementary dataset of sequencing reads was created from matched normal germline WGS data without tumor read mixing (TF = 0, n = 20 replicates). To simulate detection in the context of residual disease, somatic mutation calling was performed on the original tumor and germline WGS data to obtain a patient-specific list of somatic SNVs. The number of tumor-associated mutation sites in the in silico plasma simulation mixture was then measured through the detection of at least one support for the patient-specific SNV list. By analyzing simulated plasma samples with and without ctDNA, we identified sequencing noise as a major barrier to highly sensitive detection. To reduce the impact of sequencing artifacts, we filtered out errors associated with low base quality (BQ) and mapping quality (MQ) markers. A combined BQ and MQ optimized filter was developed using receiver operating characteristic (ROC) analysis (Figure 3A), which reduced the measurement error rate by 10-fold (to approximately 2 / 10,000 of Figure 3B). In summary, this optimized SNV detection method demonstrated high agreement between the proposed mathematical method (red line, Figure 3C) and the measured empirical data (mean + / - confidence interval, Figure 3C), demonstrating high sensitivity approaching TF = 1 / 100,000. Furthermore, the high agreement between the experimental results and the mathematical model allowed us to accurately convert empirical SNV detection into TF estimates (Figure 3D), enabling quantitative MRD monitoring. Furthermore, in silico validation of TF estimates demonstrated a 5 × 10 -5We show that accurate and specific estimations were obtained for all TFs over the 1000-fold increase in the number of TFs (Figure 3E, F, and G). Here, in three different samples, e.g., melanoma (Figure 3E), lung (Figure 3F), and breast (Figure 3G) tumor samples, a high correlation (R2 = 0.999) was observed between the input mixture TFs (x-axis) and the TFs estimated from the mutation patterns (y-axis).
[0285] The data showed that the filters reduced noise in the samples. For example, the pre-filter noise was ~2 × 10 for both lung and melanoma cancers. -3 The velocity after filtering noise is ~2×10 for both cancers. -4 (Figure 3C). Using a relaxed 35x coverage filter optimized for base quality (BQ) and mapping quality (MQ), we were able to detect markers in samples with as few as 20,000 times fewer TFs. Here, the red line represents the theoretical (binomial model) expectation, while the empirical measurements are shown in black (mean & confidence interval of 5 independent replicates (Figure 3D)). The noise level is represented by the grey area in the detection distribution for TF = 0. Furthermore, in silico validation of TF estimates in melanoma samples revealed a 5x10 -5 Accurate and specific estimates were obtained for all TFs (Figure 3E).
[0286] Analytical validation of the marker using synthetic plasma mixtures demonstrated a total TF >5 × 10 -5 , especially TF>5×10 -4 We further demonstrate the relevance of somatic SNVs and cCNVs for tumor fraction estimation in 2016. The data are shown in Figures 3H and 3I.
[0287] Further analytical validation of the method using synthetic samples showed very good correlation (R2 = 83.5%) between the SNV and CNV detection methods (see Figure 3J).
[0288] A comparative evaluation of the disclosed method compared to ICHOR showed that the ICHOR method achieved a TF > 5 × 10 -3 We show that only in the case of α = 0.05 provides a correlation between input and output tumor fractions (Figure 3K).
[0289] A graph showing SNV detection rates in silico or in ctDNA samples from control subjects (BB601) or cancer patients (BB1122 or BB1125) using the methods and systems of the present disclosure is shown in Figure 4.
[0290] To evaluate approaches for detecting residual disease in NSCLC patients after surgery with curative intent, five early-stage lung cancer specimens were collected (Table 1). Initial WGS was performed on matched tumor and germline DNA (PBMCs) to generate a patient-specific genome-wide SNV catalog. Additionally, plasma samples were collected from the subjects before surgery and approximately 14 days after surgical resection. CfDNA was extracted and sequenced through an optimized WGS protocol, after which SNV detection in all plasma samples was analyzed based on the patient-specific genome-wide SNV catalog.
[0291] The results are shown in Figure 5A. The data show that genome-wide SNV detection exceeded the noise threshold in all five preoperative plasma samples of early-stage NSCLC adenocarcinoma cases (Figure 5A). Furthermore, SNVs were detected in postoperative plasma in two of the five cases, and correlated with the patient's clinical outcome (recurrence or death) (Figure 5A). Specifically, postoperative TFs were detected above the noise threshold of 5 × 10 -5 Only two cases exceeded the threshold. However, all healthy control samples had TFs below the detection threshold. "ND" indicates non-detection. The data showed consistent results with the SNV method for plasma detection and TF correlation.
[0292] To clinically validate this innovative approach and facilitate its implementation in clinical settings, we apply the method to 30 cases of early-stage lung cancer (stage I and II). Initial WGS is performed on the patient's matched previously collected tumor and PBMC DNA, as well as pre- and post-operative plasma samples. An SNV-based detection algorithm is used to quantify pre- and post-operative TFs. Clinical variables (e.g., disease stage, lymph node metastasis, pathological features, and patient demographics) associated with elevated pre- or post-operative plasma TF levels are identified. The impact of a positive post-operative plasma sample on the patient's progression-free survival is specifically examined. Data from a representative cohort of 11 patients are shown in Figure 5B (adenocarcinoma versus healthy plasma controls) and Figure 5C (adenocarcinoma versus negative controls between patients), demonstrating a sensitivity of over 60% and a specificity of over 85%. The concordance between sSNV and sCNV detection is shown in Figure 5D.
[0293] Postoperative tumor DNA detection can be used as a prognostic marker for aggressive disease requiring adjuvant therapy. For example, in a postoperative analysis of outcomes in 11 patients (plasma collected 2 weeks after surgery), recurrence-free time was found to be inversely correlated with sSNV-based z-score detection (Figure 11H).
[0294] Example 3A: Orthogonal integration of fragment size features in SNV-based methods
[0295] cfDNA fragment distribution has a unique profile for DNA degradation in the blood circulation. Normal cfDNA samples exhibit the fragment size distribution shown in Figure 10A. Circulating DNA fragments derived from tumors are shorter than "normal" DNA fragments, primarily derived from apoptosis of hematopoietic (immune) cells. Breast tumor cfDNA (red and purple) exhibits a fragment size shift compared to normal cfDNA samples (Figure 10B). Calculation of the center of mass (COM) of the first nucleosome (peak at approximately 170 bp) shows a shift to a lower COM, which corresponds linearly to TFs. Using human tumor xenograft models (PDX) in mice, tumor-derived circulating DNA (red, aligned to human) was shown to be significantly shorter than normal-derived circulating DNA (black, aligned to mouse). See Figure 10C.
[0296] To create a stable model that can quantify the probability that a single DNA fragment is derived from tumor or normal origin, we used a joint Gaussian mixture model (GMM) to characterize the fragment size distribution of circulating DNA. The circulating tumor DNA model (red dashed line) was estimated by applying GMM analysis to circulating tumor DNA extracted from our PDX samples using only circulating DNA aligned to the human genome. The circulating normal DNA model (gray dashed line) was estimated by applying GMM analysis to circulating DNA from plasma samples of healthy human volunteers. The joint log-odds ratio (yellow line) was then used to estimate the probability that a particular circulating DNA fragment size is of tumor or normal origin. The data are shown in Figure 10D.
[0297] Patient-specific mutation detection can be used to confirm whether the DNA fragment is tumor-derived or not based on its fragment size distribution and GMM combined log odds ratio. To increase reliability and reduce batch effect bias, intra-patient controls were developed using inter-patient cross-detection. For example, specific patients shown below the detected tumor mutation (gray, concordant detection) show a tendency for fragment size to shift to smaller sizes. Mutations associated with other patients in the same patient sample (red inter-patient detection) share the same tobacco pattern context information, but these artifactual detections are not true detections. Interestingly, these inter-patient detections did not show a tendency for a lower fragment size shift, and their fragment size distributions were significantly different from those of true tumor detections (Wilcoxon rank sum, P value 3 × 10-9). Using the GMM combined log odds ratio, the patient-specific mutation detection was confirmed to be tumor-derived (combined log odds ratio = 0.3), while the artifactual mutation from the same patient sample was confirmed to be normal-derived (combined log odds ratio = -0.35). Representative data from three patients are shown in Figure 10E.
[0298] Example 3B: Orthogonal integration of fragment sizes associated with CNV marker cfDNA fragment distribution has a unique profile due to DNA degradation in the blood circulation. Normal cfDNA samples show a shift in fragment size distribution (see Figures 10A and 10B above). Here, in the context of analyzing the center of mass distribution (COM), calculation of the COM of the first nucleosome (peak at approximately 170 bp) shows a shift to a lower COM that linearly corresponds to TF.
[0299] Comparative analysis of fragment size center of mass (COM) between patients may have limited sensitivity and may be prone to batch effects. The local fragment size COM within a patient may vary due to epigenetic patterns and copy number events. Indeed, at amplification segments, the tumor fraction increases locally (due to an increase in the proportion of tumor DNA), resulting in a decrease in the local fragment size center of mass (COM). On the other hand, at deletion sites, the tumor fraction decreases locally (due to a decrease in the proportion of tumor DNA), resulting in an increase in the local fragment size center of mass (COM).
[0300] We validated this concept in plasma samples from cancer patients, revealing a clear negative correlation between the log2 of depth coverage (log2 > 0.5 = amplification, log2 < -0.5 = deletion) and the local fragment size center of mass (COM) of that segment (see Figure 11B). Further validation across plasma samples from 12 different cancer patients showed a clear relationship between CNV detection based on depth coverage and that based on fragment size center of mass (COM) (Figure 11C), a relationship that was not evident in normal (healthy) plasma samples (Figure 11D).
[0301] Several quantitative features can be extracted from this relationship between depth coverage (Log2) and fragment size per sample (COM). More specifically, the center of mass of the neutral region (Log2=0), the slope of the Log2 / COM relationship, and the R2 of the Log2 / COM relationship. These features indicate the dynamic response of patients to changes in tumor fraction after surgery or during treatment. For example, the following cancer patient progressing during treatment shows a decrease in COM and an increase in absolute slope value, and an increase in R2 (Figures 11E and 11F).
[0302] Using multiple linear regression or GLM, the log2 / COM features can be converted into tumor fractions to monitor patients after surgery and during treatment (Figure 11G). For example, the outcomes of patients undergoing treatment were monitored over a 6-week (42-day) period. The estimated tumor fractions (Figure 11I) and normalized CNV scores (Figure 11J) were summarized and presented in a comparative bar graph for residual disease monitoring. The data indicate that patient 4, but not patients 1-3, responded to treatment, as evidenced by the significantly lower eTF at 42 days posttreatment compared with the eTF at the time of treatment (Figure 11I). Analysis of the normalized CNV scores also revealed a positive response in patient 4 receiving immunotherapy and chemotherapy, in contrast to patients 1-3 receiving monotherapy (either chemotherapy or immunotherapy alone). Treatment response outcomes were confirmed by imaging studies and long-term clinical follow-up and were shown to be consistent with the eTF predictions.
[0303] Example 4: Highly sensitive ctDNA detection using genome-wide integration of large somatic copy number variations (sCNVs)
[0304] In addition to somatic point mutations, cancer genomes are characterized by significant aneuploidy. Through this process, large swaths of the genome undergo amplification and deletion, which can generate a strong signal for ctDNA detection. This is primarily because the coverage depth of WGS is a function of the DNA content of each site. Another notable feature is the short fragment length of ctDNA compared to normal cfDNA and nucleosome positioning information.
[0305] Therefore, WGS offers an additional advantage over targeted sequencing, as it is rich in orthogonal information sources that enhance detection. To take advantage of this orthogonal genome-wide signal provided by WGS, a similar approach has been developed to exploit differential read depth coverage in large amplified and deleted genomic segments. This read depth detection method is designed to integrate millions of small genomic windows to sensitively detect subtle depth changes in regions of patient-specific sCNVs, which can sensitively distinguish between low TF plasma and healthy (TF = 0) controls.
[0306] Thus, this disclosure provides an analytical approach that integrates multiple directional depth coverage skews across large genomic CNV segments (Figure 6A). When tested on our NSCLC hypothetical plasma samples, integration of genome-wide CNV patterns achieved high detection sensitivity of up to 1 / 100,000 TFs (Figure 6B). Furthermore, comparison between detected signals and TFs was linear (R2 = 1, P value = 2 × 10). -24 ) relationship and demonstrated appropriate modeling by a simple dilution model. Here, local depth coverage differences (amplifications, deletions) in tumors are diluted by proportional mixing with normal reads. This clear relationship allows TFs to be calculated from empirical patient measurements. This approach, similar to the SNV approach, was validated in parallel in the same patient cohort described above, and helps to build a joint classification model to synergistically improve sensitivity by integrating these orthogonal signals.
[0307] It should be noted that this method provides complementary and sensitive detection for patients with low SNV mutation burden but high CNV burden. Alternatively, the method described herein can be integrated with SNV-based methods to further improve detection regardless of cfDNA abundance. The integration of the two methods on exemplary samples demonstrates the detectability of minimal residual disease. The data demonstrate that even without matched tumor samples, genome-wide sSNV integration provides highly sensitive MRD detection through the application of mutation inference patterns.
[0308] The methods of the present disclosure are not limited to the types of markers exemplified herein. For example, residual disease detection / diagnosis can be performed by analyzing insertions or deletions (indels) in the genomic list of reads in a manner similar to SNV analysis (exemplified in Example 2). Similarly, residual disease detection / diagnosis can be performed by analyzing structural variants (SVs) in the genomic list of reads in a manner similar to CNV analysis (exemplified in Example 3).
[0309] While several exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, permutations, additions, and subcombinations thereof. Accordingly, the appended claims, and any claims hereafter introduced, are intended to include all such variations, permutations, additions, and subcombinations as fall within their true spirit and scope.
[0310] Example 5: Comparative Evaluation
[0311] The systems and methods of the present disclosure were compared to prior art calls.
[0312] Current mutation calling methods do not work with low-TF regimens. More specifically, MUTECT does not work below 1% TF. Applicable alternative methods for identifying ctDNA markers include high-coverage targeted sequencing with error suppression (e.g., double-stranded sequencing). An example of a technical method is presented in Phallen et al.'s "Direct Detection of Early Stage Cancers Using Circulating Tumor DNA" (Science Translational Medicine, 9, 203, 2017). The method described by Phallen et al. has limited sensitivity at low TF (i.e., detection of <1 / 1000 TF is rare). A second technical method from the Broad Institute (ICHOR) has similar limitations. ICHOR (see Adalsteinsson et al. "Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors," Nature Communications, 8.1, 1324, 2017) shows high concordance with metastatic tumors. As can be seen from the comparative results shown in Figure 9, the broad ICHOR method has significantly less sensitivity than the method of the present invention. In particular, the 100-fold increase in sensitivity achieved by the disclosed method and system is significantly superior to and unexpectedly advantageous over the ICHOR method.
[0313] Accordingly, the present disclosure relates to the following non-limiting embodiments.
[0314] Embodiment 1: A method of detecting residual disease in a subject in need thereof, comprising the steps of: (A) receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from a first biological sample of the subject, wherein the biological sample comprises a tumor sample and optionally a normal cell sample, and wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list of genetic markers in a second biological sample of the subject to generate a whole genome-wide representation of tumor-associated genetic markers in the second sample; (C) filtering artificial noise markers from the genome-wide list of markers in the first and second biological samples, wherein the filtering comprises: (a) filtering each SNV or Indel in the list as a noise (P N (b) statistically classifying each CNV or SV window in the list as signal or noise based on (1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and / or (3) overlap with a cfDNA mask (blacklist); (D) calculating an estimated tumor fraction (eTF) of the first and second biological samples based on one or more integrative mathematical models; and (E) detecting residual disease in the subject if the estimated tumor fraction exceeds an empirical threshold calculated using a background noise model.
[0315] Embodiment 2: The method of embodiment 1, wherein step (A) comprises receiving a subject-specific genome-wide list of genetic markers from a plurality of genetic markers from a biological sample, the biological sample comprising a tumor sample and a normal cell sample from the patient.
[0316] Embodiment 3: The method of embodiment 1 or 2, wherein the read set comprises a read set covering a specific SNV or indel site, or a read set falling within a specific CNV or SV genomic window.
[0317] Embodiment 4: The method of any one of embodiments 1 to 3, wherein the tumor sample comprises a resected tumor or FNA, including snap-frozen tissue, OCT-embedded tissue, or FFPE.
[0318] Embodiment 5: The method of any one of embodiments 1 to 4, wherein the normal sample comprises peripheral blood mononuclear cells (PMBCs), or a saliva or skin sample.
[0319] Embodiment 6: The method of any one of embodiments 1 to 5, wherein the plurality of genetic markers is obtained by whole genome sequencing of a biological sample of the subject.
[0320] Embodiment 7: The method of any one of embodiments 1 to 6, wherein the list of genetic markers from the plurality of genetic markers from the first biological sample of the subject comprises a high mutation rate and / or a high number of CNVs or SVs.
[0321] Embodiment 8: The method of embodiment 7, wherein the high mutation rate comprises at least one somatic single nucleotide polymorphism or indel / megabase pair mutation rate, and the high copy number mutations comprise somatic CNVs or SVs with a cumulative size of at least 5 megabase pairs.
[0322] Embodiment 9: The method of any one of embodiments 1 to 8, wherein the background noise model comprises measuring an error rate of detection in normal healthy samples and converting the error rate into a base noise eTF estimation model.
[0323] Embodiment 10: The threshold calculated by the eTF estimation model is 10 -4 ~10 -6 10. The method of embodiment 9, wherein
[0324] Embodiment 11: The method of any one of embodiments 1 to 11, wherein step (A) comprises receiving a subject-specific genome-wide list of somatic genetic markers derived from a plurality of genetic markers from a biological sample of a subject, the biological sample comprising a tumor sample and a normal cell sample, and step (B) comprises subsequently detecting the subject-specific genome-wide list of genetic markers in a second biological sample comprising a plasma sample of the subject to generate a temporarily updated tumor-associated genome-wide list of genetic markers in the patient's plasma.
[0325] Embodiment 12: The method of any one of embodiments 1 to 11, wherein the normal cell sample comprises PMBC, a saliva sample, a hair sample, or a skin sample.
[0326] Embodiment 13: The method of any one of embodiments 1 to 12, wherein the subject is a human and the second biological sample from the subject is a biological material selected from the group consisting of blood, cerebrospinal fluid, pleural effusion, ocular fluid, stool, urine, and combinations thereof.
[0327] Embodiment 14: A method of quantitatively estimating minimal residual disease burden in a patient during treatment, observation, or follow-up, comprising: (A) receiving a subject-specific genome-wide list of genetic markers derived from a plurality of genetic markers from a first biological sample of the subject, wherein the biological sample comprises a tumor sample and optionally a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list of genetic markers from a second biological sample of the subject to generate a tumor-associated genome-wide list of genetic markers in the second sample; (C) filtering artificial noise markers from the genome-wide lists of markers in the first and second biological samples, wherein the filtering comprises: (a) filtering each SNV or Indel in the list as a noise (P N(b) statistically classifying each CNV or SV window in the list as signal or noise based on (1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and / or (3) its overlap with a cfDNA mask (blacklist); (D) calculating an estimated tumor fraction (eTF) of the first and second biological samples based on one or more integrative mathematical models; and (E) detecting residual disease in the subject if the estimated tumor fraction exceeds an empirical threshold calculated using a background noise model.
[0328] Embodiment 15: The method of embodiment 14, wherein (E) further comprises detecting residual disease in the subject after resection surgery; detecting residual disease during or after treatment; detecting residual disease to monitor the effectiveness of treatment; detecting residual disease to monitor for recurrence or recurrence of cancer; or a combination thereof.
[0329] Embodiment 16: The method of embodiment 15, wherein the resective surgery comprises lymph node biopsy, head or neck surgery, uterine or endometrial biopsy, bladder biopsy, mastectomy, prostatectomy, skin lesion resection, small bowel resection, gastrectomy, thoracotomy, adrenalectomy, colectomy, oophorectomy, thyroidectomy, hysterectomy, glossectomy, or colon polypectomy.
[0330] Embodiment 17: The method of embodiment 15, wherein the treatment comprises chemotherapy, immunotherapy, targeted therapy, radiation therapy, or a combination thereof.
[0331] Embodiment 18: The method according to any one of embodiments 14 to 17, wherein the BQ, MQ and fragment size parameters of the marker are optimized using a ROC curve.
[0332] Embodiment 19: The method of any one of embodiments 14 to 18, comprising using combined base quality mapping quality (BQ MQ) parameters.
[0333] Embodiment 20: The method of any one of embodiments 14 to 19, further comprising receiving a plurality of genetic markers from a biological sample of the subject, the biological sample comprising a tumor sample and a normal cell sample, and generating a subject-specific genome-wide list of genetic markers from the received plurality of genetic markers.
[0334] Embodiment 21: The method of any one of embodiments 14 to 20, further comprising detecting a subject-specific genome-wide list of genetic markers in a third biological sample from the subject and comparing the detected list to the subject-specific genome-wide list of genetic markers generated in the first biological sample from the subject.
[0335] Embodiment 22: The method of embodiment 21, wherein the third biological sample is a plasma sample of the subject obtained to generate a temporarily updated list of tumor genome-wide genetic markers in patient plasma.
[0336] Embodiment 23: The method of any one of embodiments 14 to 22, further comprising empirically determining a background noise threshold, wherein the tumor fraction above said background noise threshold provides a quantitative estimate of tumor burden.
[0337] Embodiment 24: The method of any one of embodiments 14 to 23, wherein a tumor fraction below the noise threshold is considered not detected (ND).
[0338] Embodiment 25: The method of any one of embodiments 14 to 24, wherein said detecting comprises quantitative monitoring over time.
[0339] Embodiment 26: The method of any one of embodiments 14 to 25, wherein the tumor is a brain tumor, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colon cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, melanoma, osteosarcoma, or a solid tumor, and is heterogeneous or homogeneous in nature.
[0340] Embodiment 27: The method of any one of embodiments 14 to 26, wherein the tumor is lung adenocarcinoma, ductal adenocarcinoma, non-small cell lung cancer (NSCLC LUAD), cutaneous melanoma, urothelial carcinoma, or osteosarcoma.
[0341] Embodiment 28: The method of any one of embodiments 14 to 27, wherein the calculating step further comprises: calculating eTFs for SNV or indel markers by integrating a probabilistic model comprising patient-specific parameters including: 1) integrated signal of plasma SNV or indel detection, 2) estimated genome coverage and a sequencing noise model, and / or 3) mutational burden (N); and calculating eTFs for CNV or SV markers utilizing a stochastic dilution model, wherein the stochastic dilution model 1) integrates the directional depth of skewed coverage between plasma and normal patient samples to be consistent with tumor CNV or SV directionality, where copy number amplifications are positively skewed and copy number deletions are negatively skewed; 2) integrates the cumulative depth of skewed coverage between tumor and normal (PBMC) patient samples; and / or 3) determining a dilution ratio between the signals.
[0342] Embodiment 29: A system for detecting residual disease in a subject in need thereof, comprising: (A) an analysis unit constructed and arranged to filter artificial noise markers from a genome-wide list of markers, wherein the genome-wide list of markers is generated from a plurality of genetic markers from biological samples of the subject, wherein the biological samples comprise tumor samples and normal cell samples, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), indels, copy number variations, SVs, and combinations thereof; the analysis unit further comprises detecting a subject-specific genomic list of genetic markers in a second biological sample to generate a tumor genomic list; the analysis unit further comprises a classification engine, wherein the classification engine performs the following: (a) classifying each SNV or Indel in the list as noise (P N (b) statistically classifying each CNV or SV window in the list as signal or noise based on (1) its position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and (3) the representation of the CNV or SV window in the cfDNA data; (B) a computing unit configured and arranged to calculate an estimated tumor fraction (eTF) of the sample based on one or more integrative mathematical models; and (C) a display unit that outputs a residual disease profile for the subject based on the estimated tumor fraction, wherein the display unit outputs a residual disease profile for the subject if the estimated tumor fraction exceeds an empirical threshold calculated by a background noise model.
[0343] Embodiment 30: The computing unit is further configured to calculate eTFs for SNV or Indel markers by integrating a probabilistic model, wherein the probabilistic model includes: 1) an integrated signal of plasma SNV or Indel detection, 2) process quality metrics including estimated genome coverage and a sequencing noise model, and / or 3) patient-specific parameters including mutation burden (N); and / or calculating eTFs for CNV or SV markers using a probabilistic mixture model, wherein the probabilistic dilution model includes: 1) integrating the directional depth of skewed coverage between plasma and normal patient samples consistent with tumor CNV or SV directionality, where copy number amplifications are positively skewed and copy number deletions are negatively skewed; 2) integrating the cumulative depth of skewed coverage between tumor and normal patient samples; and / or 3) finding a dilution ratio between the signals.
[0344] Embodiment 31: The arithmetic unit (B) includes a processor, and the processor is configured to execute the computer-readable instructions, which, when executed, generate the following integrated mathematical model (1)(1): (1) eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov) to estimate the tumor fraction (eTF) of the sample, where M is the number of tumor-specific SNVs detected in the patient plasma sample, σ is an empirically estimated error rate measure, R is the total number of unique reads in the SNV group region of interest (ROI), N is the tumor mutation burden, and cov is the average number of unique reads per site in the SNV group ROI; and / or (2) eTF[CNV]=(sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth coverage in the genomic window indexed by {i}, representing the depth coverage of the plasma and cov of the normal samples. wherein {i} is a discrete index value that counts all genomic windows that cover patient-specific amplified and deleted genomic segments; T is the median depth in the genomic window indexed by {i}, representing tumor depth coverage, normalized by either stable z-score or stable PCA compared to a cohort of normal samples; N is the median depth in the genomic window that represents normal depth coverage, indexed by either stable z-score or stable PCA, normalized by either stable z-score or stable PCA compared to a cohort of normal samples; and {i} is a discrete index value that counts all genomic windows that cover patient-specific amplified and deleted genomic segments.
[0345] Embodiment 32: A computer-readable medium comprising computer-executable instructions which, when executed by a processor, cause the processor to perform a method or set of steps for residual disease detection, the method and set of steps comprising: (A) receiving a subject-specific genome-wide list of genetic markers from a plurality of genetic markers from a biological sample of a subject, the biological sample comprising a tumor sample and optionally a normal cell sample, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (B) detecting the subject-specific genome-wide list of genetic markers in a second biological sample of the subject to generate a list of tumor-associated genome-wide genetic markers in the second sample; (C) filtering each SNV or Indel in the list by noise (P N (B) statistically classifying markers based on the probability of detection of a CNV or indel as a function of (1) the mapping quality of the reads containing the SNV, (2) the fragment size length of the reads containing the SNV, (3) a consensus test within the read duplication family containing the SNV or indel, and / or 4) the base quality (BQ) of the SNV or indel, and / or filtering artificial noise markers from the genome-wide list of markers based on (1) their position relative to the centromere, 2) the mapping quality (MQ) of the reads containing the CNV or SV window, and / or (3) overlap with a cfDNA mask (blacklist); (D) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrative mathematical models; and (E) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by the background noise model.
[0346] Embodiment 33: A method of detecting minimal residual disease in a subject, comprising: (A) receiving a genome-wide list of reads in genetic data sequenced from multiple biological samples received from the subject; and (B) performing mutation calling on tumor and peripheral blood mononuclear cell (PBMC) samples from the subject, wherein the calling mutations include MUTECT, LOFREQ, and / or STRELKA mutations to generate subject-specific reads of somatic SNVs (sNVs) or indels as an individualized reference set; C) collecting and filtering the reads from the subject-specific somatic SNVs (sNVs) or indels, (1) removing reads with low mapping quality (e.g., ROC<29, optimized); (2) constructing multiple PCR / sequencing copies of the same DNA fragment; (3) constructing duplicate families (representing multiple PCR / sequencing copies of the same DNA fragment) and generating corrected reads based on a consensus test; (4) removing reads with low base quality (e.g., <21, ROC optimized). and (4) removing reads with large fragment sizes (e.g., >160, ROC optimized); (D) calculating the number of subject-specific mutation sites with at least one supporting read (within the filtered set) with an identical substitution in the tumor; (F) estimating the tumor prevalence of SNVs based on the mathematical model eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov) (Equation 1), where M is the number of tumor-specific list detections in the patient sample, σ is an empirically estimated measure of noise, and R is the region of interest ( where N is the total number of unique reads in the ROI, N is the tumor mutation burden, and cov is the average number of unique reads per site in the ROI; G) comparing eTF[SNV] to a detection threshold consisting of empirically measured baseline noisy TF estimates from healthy samples, where eTF[SNV] above the threshold level (e.g., eTF[SNV] above 2 standard deviations of the noisy TF distribution (FPR<2.5%) indicates a positive detection); and (K) detecting residual disease in the subject based on eTF estimates above the detection threshold level.
[0347] Embodiment 34: A method for detecting minimal residual disease in a subject, comprising: (A) receiving a genome-wide inventory of sequenced sequences from a plurality of biological samples received from the subject, wherein the plurality of biological samples comprises a tumor sample, a normal sample, and a plasma sample; (B) performing CNV or SV calling on the tumor and peripheral blood mononuclear cell (PBMC) samples from the subject to generate a plurality of reference segmentations of CNV or SV segments or SVs above a threshold length (e.g., >2 Mbp, preferably >5 Mbp) and annotating the directionality of the segments, wherein amplifications are annotated as positive and deletions are annotated as negative; and (C) single bp depth coverage for the plasma, tumor, and PBMC samples covering patient-specific CNV or SV segmentation regions of interest (ROIs). collecting information; D) dividing the patient-specific CNV or SV segmentation ROI into 500 bp windows and calculating the median (artificial suppression) across all samples and windows; E) generating normalized depth coverage information for all 500 bp windows using (a) sample-wise stable z-score normalization, and / or (2) stable principal component analysis (RPCA); (F) filtering windows from the patient-specific segmentation, where the filtering includes: (1) removing reads with low mapping quality (e.g., ROC<29, optimized); and / or (2) removing centromeric regions (e.g., removing windows with normalized normals greater than 10); (3) removing regions not represented in cfDNA (e.g., removing windows not included in a cfDNA representation mask that includes multiple cfDNA samples);(G) Integrating coverage depth between plasma and normal (PBMC) patient samples using the mathematical model sumi[(P(i)-N(i)*sign[T(i)-N(i)]]-E(σ) (Equation 2), where P is the median depth coverage within the genomic window indexed by {i}, representing plasma depth coverage normalized by either stable z-score or stable PCA methods compared to a cohort of normal samples; E(sigma) is an empirically estimated measure of error rate; T is the median depth within the genomic window indexed by {i}, representing tumor depth coverage normalized by either stable z-score or stable PCA methods compared to a cohort of normal samples; and N is the median depth within the genomic window indexed by {i}, representing normal depth coverage normalized by either stable z-score or stable PCA methods compared to a cohort of normal samples; (H) Integrating the skewed cumulative coverage depth between tumor and normal (PBMC) patient samples using the mathematical model sumi[abs(T(i)-N(i)]-E(σ)] (Equation 3), where T, N, and E(σ) are as described above; (I) a putative tumor for CNV or SV; Fraction Calculating the dilution ratio between the directional depth coverage of (G) and the cumulative depth coverage of (H) corresponding to (eTF[CNV]) = (sumi[(P(i)-N(i)*sign[T(i)-N(i)]]-E(σ)) / (sumi[abs(T(i)-N(i))]-E(σ)) (Equation 4); (J) comparing eTF[CNV] to a detection threshold consisting of empirically determined baseline noisy TF estimates from healthy samples, where an eTF[CNV] above the threshold level (e.g., 2 standard deviations of the noisy TF distribution (FPR<2.5%)) indicates a positive detection; and (K) detecting residual disease in the subject based on eTF estimates above the detection threshold level.
[0348] Embodiment 35: A method for detecting residual disease in a subject in need thereof, comprising: (A) receiving a first subject-specific genome-wide list of reads associated with genetic markers from a first biological sample of the subject, the first biological sample comprising a baseline sample and a normal cell sample, each comprising a list of reads of a single base pair in length, the baseline sample comprising a tumor sample or a plasma sample; (B) filtering artificial sites from the first list, the filtering comprising removing repetitive sites generated on a cohort of reference healthy samples from the first list of genetic markers, and / or identifying germline mutations in peripheral blood mononuclear cells of the normal cell sample and removing the germline mutations; (C) detecting a second subject-specific genome-wide list of genetic markers in a second biological sample of the subject, generating a tumor-associated genome-wide list of genetic markers in the second sample; (D) filtering noise from the first and second genome-wide lists of reads using at least one error suppression protocol, and generating a first genome-wide list of genetic markers. generating a first filtered read list for the genome-wide read list and a second filtered read list for the second genome-wide read list, wherein at least one error suppression protocol (a) calculates the probability that any single nucleotide variation in the first and second suppression lists is an artifact and removes said variation, wherein the probability is calculated as a function of features selected from the group consisting of mapping quality (MQ), variant base quality (MBQ), positional read quality (PIR), mean read quality (MRBQ), and combinations thereof; and / or (b) removes accidental variations using discordance testing and / or overlap consensus between independent replicates of the same DNA fragment generated from polymerase chain reaction or sequencing, wherein artifactual variations are identified and removed when there is no match across a majority of a given replicate family; (E) applying a background noise model to one or more integrative mathematical models to generate putative tumor profiles for the first and second biological samples using the first and second filtered read sets. Fraction(F) calculating (eTF); and (F) detecting residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.
[0349] Embodiment 36: A method of detecting residual disease in a subject in need thereof, comprising: (A) receiving a first subject-specific genome-wide list of reads associated with genetic markers from a first biological sample of the subject, the first biological sample comprising a baseline sample, wherein the first list of reads each comprises a copy number variation (CNV) or a structural variation (SV), the baseline sample comprising a tumor sample or a plasma sample; (B) receiving a second subject-specific genome-wide list of reads associated with genetic markers from a second biological sample of the subject, the second biological sample comprising a peripheral blood mononuclear cell sample (PBMC), wherein the second list of genetic markers each comprises a CNV or a SV; and (C) filtering artifactual sites from the first and second lists of reads, wherein the filtering is performed on a cohort of reference healthy samples from the first and second lists of reads. (D) detecting reads from the subject-specific genome-wide list of third genetic markers in a third biological sample of the subject to generate a list of tumor-associated genome-wide genetic markers in the third sample; (E) normalizing each of the first, second, and third lists of reads to generate a first filtered read set for the first genome-wide list of reads, a second filtered read set for the second genome-wide list, and a third filtered read set for the third genome-wide list of reads; (F) applying a background noise model to one or more integrative mathematical models to generate a putative tumor-associated gene of the third biological sample using the third filtered read set. Fraction(G) calculating (eTF), wherein one or more models use the first filtered read set to generate a first eTF, and / or one or more models use the second filtered read set to generate a second eTF; and Fraction detecting residual disease in said subject if β exceeds an empirical threshold.
[0350] Embodiment 37: A system for detecting residual disease in a subject in need thereof, comprising: an analysis unit, the analysis unit receiving a first subject-specific genome-wide read list associated with genetic markers from a first biological sample of the subject, wherein the first biological sample comprises a baseline sample and a normal sample, and wherein the first read list each comprises reads of a single base pair in length, and the baseline sample comprises a tumor sample or a plasma sample; and a pre-filter engine constructed and arranged to: filter artificial sites from the first read list, wherein the filtering comprises removing, from the first list of genetic markers, repetitive sites generated on a cohort of reference healthy samples, and / or identifying germline mutations in peripheral blood mononuclear cells of the normal cell sample and removing the germline mutations from the germline from the first list of genetic markers; and receiving reads from a second subject-specific genome-wide list of genetic markers in a second biological sample of the subject, and generating a tumor-associated genome-wide representation of the genetic markers in the second sample, and and a correction engine configured and arranged to filter noise from the first and second genome-wide read lists using at least one error suppression protocol to generate a first filtered read set for the first genome-wide read list and a second filtered read set for the second genome-wide read list, wherein the at least one error suppression protocol (a) calculates the probability that any single nucleotide variation in the first and second suppression lists is an artificial variation and removes said variation, wherein the probability is calculated as a function of features selected from the group consisting of mapping quality (MQ), mutant base quality (MBQ), positional read base quality (PIR), average read base quality (MRBQ), and combinations thereof; and / or (b) removes accidental variations using discrepancy testing and / or overlap consensus between independent replicates of the same DNA fragment generated from polymerase chain reaction or sequencing, wherein artificial variations are identified and removed when there is no match across a majority of a given replicate family; and applying a background noise model to one or more integrated mathematical models; and using the first and second filtered read sets to generate estimated tumors for the first and second biological samples. Fraction (eTF) was calculated and the estimated tumor Fraction a computing unit that detects residual disease in the subject if β exceeds an empirical threshold.
[0351] Embodiment 38: A system for detecting residual disease in a subject in need thereof, comprising: receiving a first subject-specific genome-wide list of associated genetic markers from a first biological sample of the subject; receiving a second subject-specific genome-wide list of associated genetic markers from a second biological sample of the subject, said second biological sample comprising a peripheral blood mononuclear cell sample (PBMC), wherein said second list of genetic markers each comprises copy number variations (CNVs); and a pre-filter engine configured and arranged to filter artificial sites from the first and second lists of genetic markers, said filtering comprising removing repetitive sites generated on a cohort of reference healthy samples from the first list of genetic markers and / or removing repetitive sites generated on peripheral blood mononuclear cells of said normal cell sample. and identifying germline mutations in the first list of genetic markers and removing the germline mutations from the germline from the first list of genetic markers; and a correction engine configured and arranged to receive reads from a subject-specific genome-wide list of third genetic markers in a second biological sample of the subject and generate a representation of tumor-associated genome-wide genetic markers in the third sample; and normalize each of the first, second, and third lists to generate the reads of the first genome-wide list, a second filtered read set for the second genome-wide list of reads, and a third filtered read set for the third genome-wide list of reads; and applying a background noise model to one or more integrative mathematical models to generate a putative tumor associated with the third biological sample. Fractiona computing unit configured and arranged to calculate (eTFs), wherein the one or more models generate a first eTF using a first filtered read set, and / or the one or more models generate a second eTF using a second filtered read set, and a putative tumor in the third biological sample; Fraction Residual disease in the subject is detected if β exceeds an empirical threshold.
[0352] Embodiment 39: The method of embodiment 35, wherein the marker comprises a single nucleotide variation (SNV) or an insertion / deletion (indel); preferably an SNV.
[0353] Embodiment 40: The method of embodiments 35 and 39, wherein filtering the repetitive sites generated on a cohort of reference healthy samples comprises generating a panel of normal (PON) blacklist or mask.
[0354] Embodiment 41: The method of any one of embodiments 35 and 39 to 40, wherein the normal sample comprises peripheral blood mononuclear cells (PBMCs), and germline mutations in PBMCs are removed in the artificial site filtering step (B).
[0355] Embodiment 42: The method of any of embodiments 35 and 39 to 41, wherein in step (A), the first biological sample comprises a plasma sample obtained from the subject before surgery or treatment.
[0356] Embodiment 43: The method of any of embodiments 35 and 39 to 42, wherein in step (C), the second biological sample comprises a plasma sample obtained from the same subject after treatment or surgery.
[0357] Embodiment 44: The method of any of embodiments 35 and 39 to 43, wherein step (D) comprises filtering artificial noise using a machine learning (ML) algorithm, such as a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), discriminant analysis, nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof; preferably a support vector machine (SVM).
[0358] Embodiment 45: The method of any of embodiments 35 and 39 to 44, wherein in step (D), the second error suppression step comprises correction of artificial mutations generated by PCR or sequencing using comparison of independent copies of the same original nucleic acid fragment.
[0359] Embodiment 46: The method of embodiment 45, wherein in step (D), the second error suppression step includes correction of artificial mutations generated by paired-end 150 bp sequencing, resulting in overlapping paired reads (R1 and R2), and mismatches between R1 and R2 pairs are returned to the corresponding reference genome.
[0360] Embodiment 47: In step (D), the second error suppression step comprises correcting overlapping families generated during sequencing and / or PCR amplification, wherein the overlapping families are recognized by 5' and 3' similarity and alignment position, and each overlapping family is used to check the consensus of a particular mutation across independent replicates, thereby correcting mutations in arch families that do not match in the majority of the overlapping families.
[0361] Embodiment 48: The method of any of embodiments 35 and 39 to 47, wherein in step (E), the mathematical model integrates the relationship between coverage, mutation burden, number of detected mutations and tumor fraction (TF).
[0362] Embodiment 49: The method of any of embodiments 35 and 39 to 48, wherein in step (E), calculating the background noise comprises using the patient-specific mutation pattern to calculate (1) an expected noise distribution across a cohort of healthy plasma samples (Panel of Normals or PON), or (2) an expected noise distribution across other patients (inter-patient analysis).
[0363] Embodiment 50: The method of embodiment 49, wherein the background noise model provides an estimated mean and standard deviation (μ, σ) of the artifactual mutation detection rate.
[0364] Embodiment 51: The method of any of embodiments 35 to 50, further comprising orthogonal integration of secondary features including fragment size shifts.
[0365] Embodiment 52: The method of embodiment 51, wherein the intra-patient fragment size shift in the list of tumor-specific markers and random markers is analyzed using statistical methods, such as significance or Gassan mixture model (GMM).
[0366] Embodiment 53: The method of embodiment 36, wherein the marker comprises a copy number variation (CNV).
[0367] Embodiment 54: The method of any one of embodiments 36 and 37, wherein filtering the repetitive sites generated on the cohort of reference healthy samples comprises generating a panel of normal (PON) blacklist or mask.
[0368] Embodiment 55: The method of any one of embodiments 36 and 53 to 54, wherein germline events in the PBMCs are removed in the artificial site filtering step (C).
[0369] Embodiment 56: The method of any of embodiments 36 and 53 to 55, wherein in step (A), the first biological sample comprises a plasma sample obtained from the subject before surgery or treatment, and the second biological sample comprises PBMCs obtained from the same subject before surgery or treatment.
[0370] Embodiment 57: The method according to any one of embodiments 36 and 53 to 56, wherein in step (C), the third biological sample comprises a plasma sample obtained from the same subject after treatment or surgery.
[0371] Embodiment 58: The method of any of embodiments 36 and 53 to 57, wherein step (C) comprises binning regions of interest (ROIs) that include all genomic segments of somatic tumor CNVs (sT_CNVs) and somatic PBMC_CNVs (sP_CNVs), estimating depth coverage (read counts) in each window from follow-up plasma samples, and calculating the median depth coverage per window.
[0372] Embodiment 59: The method of any of embodiments 36 and 53 to 58, wherein the follow-up plasma sample is obtained after surgery, during treatment, or at follow-up.
[0373] Embodiment 60: The method of any of embodiments 36 and 53 to 59, wherein the normalization step comprises normalizing depth coverage values to correct for GC content bias and mappability bias by performing two LOESS regression curve fittings on the bin-wise GC fraction and mappability scores.
[0374] Embodiment 61: The method described in any of embodiments 36 and 53 to 60, wherein the normalization step comprises batch effect correction using stable z-score normalization applied to each sample separately.
[0375] Embodiment 62: The method of Example 62, wherein the normalization of the z-scores comprises calculating the median and median absolute deviation (MAD) based on the neutral region of each sample, and normalizing all CNV bins is normalized by subtracting the median and dividing the difference by the MAD.
[0376] Embodiment 63: The method of any of embodiments 36 and 53 to 62, wherein step (E) comprises calculating depth coverage skew and / or fragment size center of mass (COM) skew in the third sample compared to a panel of normal (PON) healthy plasma samples.
[0377] Example 64: The method of any of embodiments 36 and 53-63, wherein step (E) comprises calculating the tumor fraction by checking the linear dilution ratio between the cumulative signal detected in the follow-up plasma sample compared to the cumulative signal detected in the tumor sample.
[0378] Example 65: The method of any of embodiments 36 and 53-64, wherein in step (F), calculating the background noise comprises using the patient-specific CNV / SV patterns to calculate (1) the expected noise distribution across a cohort of healthy plasma samples (normal panel or PON), or (2) the expected noise distribution across other patients (inter-patient analysis).
[0379] Embodiment 66: The method of Example 65, wherein the background noise model provides an estimated mean and standard deviation (μ, σ) of the artifactual SNV / SV detection rate.
[0380] Embodiment 67: The method of any of embodiments 36 and 53 to 66, further comprising orthogonal integration of secondary features including fragment size shifts.
[0381] Embodiment 68: The method of Example 67, wherein the correlation between depth coverage skew and fragment size skew in CNV segments is analyzed to infer tumor fraction, for example, using a generalized linear model.
[0382] For convenience, certain terms employed in the specification, examples, and claims are collected here. Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0383] Throughout this disclosure, various patents, patent applications, and publications are referenced. The disclosures of such patents, patent applications, and accessions (e.g., those identified by PUBMED, PUBCHEM, NCBI, UNIPROT, or EBI accession numbers) and their entire publications are incorporated by reference into this disclosure to more fully describe the state of the art known to those skilled in the art as of the date of this disclosure. The present disclosure governs in the event of a conflict between the cited patents, patent applications, and publications and this disclosure.
Claims
1. 1. A method for detecting residual disease in a subject, comprising: receiving a first subject-specific list of genome-wide genetic markers from a first biological sample of the subject, wherein the first subject-specific list of genome-wide genetic markers is selected from the group consisting of single nucleotide variations (SNVs) or short insertions / deletions (indels), copy number variations (CNVs), structural variations (SVs), and combinations thereof; detecting a second subject-specific genome-wide genetic marker list in a second biological sample from the subject to generate a tumor-associated genome-wide genetic marker representation in the second biological sample; filtering artificial noise markers from the first and second subject-specific genome-wide genetic marker lists, wherein the filtering comprises classifying each SNV, indel, CNV, and / or SV in the subject-specific genome-wide genetic marker lists as signal or noise. ;and, detecting whether residual tumor is present in the subject based on one or more models that integrate coverage, mutation burden, detected mutations, and tumor fraction across subject-specific genome-wide genetic marker inventories in the first and second biological samples; A method comprising:
2. 10. The method of claim 1, wherein the first and second biological samples comprise a tumor sample and a normal cell sample from the subject.
3. 2. The method of claim 1, wherein the subject-specific genome-wide genetic marker list comprises SNVs and / or indels, and wherein the filtering comprises classifying the SNVs and / or indels in the subject-specific genome-wide genetic marker list as signal or noise.
4. The method described in claim 3, wherein each of the subject-specific genome-wide genetic marker lists has a mutation rate of at least one somatic SNV per megabase pair and / or at least one somatic indel per base pair.
5. The method described in claim 1, wherein the subject-specific genome-wide genetic marker list includes CNVs and / or SVs, and the filtering includes classifying each CNV and / or SV in the subject-specific genome-wide genetic marker list as signal or noise.
6. The method described in claim 5, wherein the cumulative size of the subject-specific genome-wide genetic marker list is at least 5 megabase pairs of somatic CNVs and / or SVs.
7. The method of claim 1, wherein the subject-specific genome-wide genetic marker list is obtained from a biological sample via whole genome sequencing.
8. Detecting whether the subject has residual tumor comprises: calculating an estimated tumor fraction (eTF) of the first and second biological samples based on one or more models integrating coverage, mutation burden, detected variables, and tumor fraction across the subject-specific genome-wide genetic marker inventory in the first and second biological samples; and detecting residual tumor in said subject if said eTF exceeds an empirical threshold. The method of claim 1 , comprising:
9. The method described in claim 8, wherein the empirical threshold is calculated using a background noise model based on the error rate of detection in normal healthy samples.
10. The method of claim 8, wherein the empirical threshold is between 10 −4 and 10 −6 .
11. The method described in claim 1, wherein receiving the subject-specific genome-wide genetic marker list includes receiving the subject-specific genome-wide somatic genetic marker list from the biological sample of the subject, and the biological sample includes a tumor sample and a normal cell sample.
12. The second biological sample comprising a plasma sample from the subject, wherein: detecting the subject-specific genome-wide genetic marker inventory in the second biological sample comprises generating a temporally updated tumor-associated genome-wide genetic marker representation inventory of genetic markers in a plasma sample from the subject; and the subject-specific genome-wide genetic marker inventory in the second biological sample comprises a temporally updated tumor-associated genome-wide representation of genetic markers in a plasma sample from the subject; The method of claim 11.
13. The method of claim 1, wherein the biological sample comprises a tumor sample and a normal cell sample, and the normal cell sample comprises a PMBC, a saliva sample, a hair sample or a skin sample.
14. The method of claim 1, wherein the subject is a human and the second biological sample is a biological material selected from the group consisting of blood, cerebrospinal fluid, pleural effusion, ocular fluid, feces, urine, and combinations thereof.
15. The method of claim 1, comprising determining whether or not to administer adjuvant therapy to the subject based on the detected residual lesions.
16. 16. The method of claim 15, wherein determining whether to administer adjunctive therapy comprises predicting a treatment outcome based on the adjunctive therapy.
17. The step of filtering artificial noise markers comprises filtering SNVs and / or indels in the subject-specific genome-wide genetic marker list according to one of the following: 1) Mapping quality (MQ) of the read set containing the SNV; 2) the fragment length of the read group containing the SNV; 3) Consensus testing within read duplication families containing SNVs and / or indels and / or 4) Base Quality (BQ) of SNVs and / or Indels 2. The method of claim 1, comprising classifying as signal or noise based on a probability of detection of noise (P N ) as a function of .
18. The process of filtering artificial noise markers, comprising filtering each CNV and / or SV in the subject-specific genome-wide genetic marker list according to one of the following: (1) Position relative to the centromere, (2) the mapping quality (MQ) of the read set containing the CNV or SV window; and / or (3) Overlap with cfDNA mask 2. The method of claim 1, comprising classifying the signal as signal or noise based on:
19. The step of detecting whether a residual tumor is present in the subject comprises the steps of: determining a z-score based on one or more models that integrate coverage, mutation burden, detected mutations, and tumor fraction across the subject-specific genome-wide genetic marker inventory in the first and second biological samples; and detecting whether the subject has residual tumor based on the z-score; The method of claim 1 , comprising:
20. A system comprising a memory for storing computer-readable instructions and a processor connected to the memory and configured to execute a method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Identification and Use of Circulating Nucleic Acid Tumor Markers
US20160032396A1
Methods for fragmentome profiling of cell-free nucleic acids
WO2018009723A1