Systems and methods for detecting residual disease

By analyzing tumor-specific markers in plasma samples, using algorithms and statistical classifiers to distinguish quality markers from noise, integrating mathematical models to process high-quality markers, and correcting sequencing and mapping errors, this approach solves the problem of low sensitivity in detecting residual diseases with low tumor burden in existing technologies, and achieves accurate diagnosis of residual diseases.

CN121439155APending Publication Date: 2026-01-30CORNELL UNIVERSITY +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511440082.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-02-27
Filing Date
2019-02-27
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing technologies face the problem of low detection sensitivity when detecting residual disease with low tumor burden, especially when using limited input materials, making it difficult to accurately diagnose minimal residual disease, particularly in diseases such as lung cancer.

Method used

By analyzing tumor-specific markers in plasma samples, algorithms and statistical classifiers are used to distinguish quality markers from artificial noise based on multiple parameters. A mathematical model is integrated to process high-quality markers, eliminate markers that may be associated with artificial noise, and multiple indicators are used to correct errors in sequencing and mapping techniques. The estimated tumor score is then calculated to achieve accurate diagnosis of residual disease.

Benefits of technology

It achieves sensitive and accurate detection of residual disease under low tumor burden conditions, and can effectively diagnose residual disease with limited input materials, especially improving the accuracy of detection in diseases such as lung cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121439155A_ABST
    Figure CN121439155A_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems, software and methods for detecting residual diseases, such as residual neoplastic diseases, in a subject, such as a human cancer patient.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201980027654.4, filed on February 27, 2019, entitled "System and method for detecting residual diseases".

[0002] Cross-references to related applications

[0003] This application claims the benefit of U.S. Provisional Application 62 / 636,150, filed February 27, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0004] Embodiments of this disclosure generally relate to the field of medical diagnostics. In particular, embodiments of this disclosure relate to compositions, methods, and systems for tumor detection and diagnosis.

[0005] introduction

[0006] Cell-free circulating DNA (cfDNA) released from dying cells allows for dynamic investigation of somatic and epigenomic genomes over time for clinical purposes. The ability to obtain a biopsy via a simple blood draw enables dynamic genomic measurements in a non-invasive manner. It overcomes spatial limitations, such as difficulty accessing lung tissue.

[0007] Circulating tumor DNA (ctDNA) can be detected and measured in the blood of cancer patients, and should not be confused with cell-free DNA (cfDNA). ctDNA has been shown to be associated with changes in tumor burden and response to treatment or surgery (Diehl et al., Nature Medicine, 14(9):985–990, 2008). ctDNA can even be detected in early-stage non-small cell lung cancer (NSCLC), and therefore has the potential to transform the diagnosis and treatment of NSCLC (Sozzi et al., Journal of Clinical Oncology, 21(21), 3902–3908, 2003; Tie et al., Science translational medicine, 8(346):346ra92–346ra92, 2016; Bettegowda et al., Science translational medicine, 6(224): 224ra24–224ra24, 2014; Wang et al., Clinical Cancer Research, 16(4): 1324–1330, 2010).

[0008] One of the key areas of future potential for cfDNA-based cancer research is the detection of residual disease (RD) to guide clinical interventions. For example, detecting RD after surgical resection can help clinicians and patients make decisions regarding expensive and toxic adjuvant therapies. However, in cases with low tumor burden (e.g., minimal residual disease (MRD)), the tumor fraction (TF) is very low. To enable the detection of mutations in low-TF cfDNA, a popular paradigm is to increase the sequencing depth of limited high-yield target sets (e.g., common cancer drivers or patient-specific groups sequenced to depths of approximately 10,000 to 100,000 reads / base). Furthermore, molecular and analytical methods have been integrated with ultra-deep sequencing to reduce sequencing errors and improve detection sensitivity in low tumor fractions (TF).

[0009] While these existing methods offer high-precision detection in some cases, they are hampered by fundamental limitations that reduce detection sensitivity (limited input material). In MRDs, tumor burden is low, with typical plasma samples containing only 1–10 ng / ml of cfDNA. This small amount of cfDNA translates to only a few hundred to a few thousand genomic equivalents. Therefore, the limited number of physical fragments covering each site present in the sample (e.g., 1000 genomic equivalents in 6 ng of cfDNA) can render popular techniques relying on ultra-deep sequencing (e.g., 100,000X) ineffective. Even with ultra-deep sequencing and advanced molecular error suppression, the limited input material imposes a detection limit at tumor fraction (TF) frequencies below 0.1–1%. Thus, while detecting cancers with low tumor burdens is clinically beneficial for patients and clinicians, existing methods relying on somatic mutation identification face significant challenges due to the low frequency of tumor-derived cfDNA samples.

[0010] Therefore, there is an urgent but unmet need for minimally invasive systems and methods that allow for the detection of tumors, particularly in cases where minimal residual disease (MRD) is diagnosed using limited input materials. Effective diagnosis of tumors is advantageous from both an economic and clinical perspective, especially in cases of residual disease (e.g., after surgery and / or treatment). This is particularly true in the case of lung cancer, where most patients are diagnosed with advanced disease and depressing outcomes (Herbst et al., N Engl J Med., 359(13):1367-80, 2008). Invention Overview

[0011] This disclosure relates to methods and systems for diagnosing residual neoplastic disease by analyzing tumor-specific markers in a subject's sample (e.g., a plasma or blood sample). The methods of this disclosure utilize algorithms and / or statistical classifiers to distinguish quality markers from artificial noise based on multiple parameters. For example, when the marker is a single nucleotide variant (SNV), the algorithm of this disclosure classifies such SNVs in the subject's genetic profile as signal or noise based on qualitative characteristics of the marker, such as, for example, the base quality (BQ) and mapping quality (MQ) of the SNV. Similarly, where the marker is a copy number variant (CNV), the algorithm classifies CNVs in the profile as signal or noise based on parameters such as centromere proximity, overlap with cfDNA overlay mask, and / or association of CNVs with low mapping ability (mapping quality; MQ) readings. Thus, markers that may be associated with artificial noise are eliminated from the subject's genetic profile, and high-quality markers are processed by a robust, integrated mathematical model that allows for the assessment of tumor fractions in the sample. If the estimated tumor fraction is found to be above a certain threshold, a positive diagnosis can be made with high confidence. Conversely, if the estimated tumor score is below the threshold, a positive diagnosis is not made at this time.

[0012] In this context, the use of a synthetic mixture of tumor and normal whole-genome sequencing data from lung patients with variable tumor read fractions ranging from 1% to 0.001% (1 / 100,000) to simulate plasma somatic mutation calling reveals the strength and accuracy of the method of the present invention relative to the prior art.

[0013] This disclosure also relates to several indicators that can suggest that variants detected by sequencing are not true somatic mutations but rather artifacts of the sequencing or mapping technology. In this context, previous studies have demonstrated that sequencing errors are not random and may be related to the DNA sequence background and the corresponding technical factors of the sequencing technology. Sequencing fidelity is also limited by the length of each sequencing read, with the error rate increasing with read length. Errors can be imposed when reads are mapped to a reference genome. Due to the presence of variable regions, motifs, and repetitive elements in the genome, the mapping process is computationally intensive and complex. Short nucleotide reads may be mapped to more than one location or not at all. The systems and methods of this disclosure can overcome these limitations of existing sequencing / mapping methods for genomic data. The indicators disclosed herein are capable of recalling true mutations from errors by analyzing multiple factors, such as (i) low base quality; and / or (ii) low mapping quality, (iii) mutation location in the reads, and (iv) read fragment size in the case of SNV labeling, and (1) genomic location score, (2) cfDNA coverage mask (blacklist), (3) low mapping quality, and (4) correlation between Log2 and read array fragment size in the case of CNV labeling.

[0014] This system and method for detecting tumor-associated biomarkers are particularly well-suited for detecting low-abundance markers. First, the model considers both quality metrics associated with the marker type and the system / method used in its detection, as well as subject-specific parameters, to calculate an estimated tumor fraction (eTF). For example, where the marker is a SNV, the integrated mathematical model considers process quality metrics such as estimated coverage and noise, as well as subject-specific parameters such as mutational burden. In the case of CNVs, the integrated mathematical model considers exponential factors and subject-specific characteristics such as CNV directionality (e.g., positive decomposition for amplification; negative decomposition for deletion) to calculate the estimated tumor fraction (eTF). Therefore, the analytical method of this disclosure integrates genome-wide mutational information to allow for sensitive analysis of samples containing cfDNA, enabling accurate and non-invasive diagnosis of residual disease.

[0015] Therefore, this disclosure relates to the following non-restrictive embodiments:

[0016] In various embodiments, a method for detecting residual disease in a subject in need of it is provided. The method may include receiving a first subject-specific whole-genome profile of readings associated with genetic markers in a first biological sample from the subject. The first biological sample may include a baseline sample. The first reading profile may each contain readings of a single base pair length (e.g., SNV or Indel), and wherein the baseline sample includes a tumor sample or a plasma sample. The method may further include filtering artificial sites from the first reading profile. Filtering may include removing repetitive sites generated in a reference healthy sample cohort from the first profile of genetic markers. Alternatively or additionally, filtering may include identifying germline mutations in peripheral blood mononuclear cells of a normal cell sample and removing said germline mutations from the first genetic marker profile. The method may further include detecting readings in a second subject-specific whole-genome profile of genetic markers in a second biological sample from the subject to generate a tumor-associated whole-genome representation of the genetic markers in the second sample. The method may further include filtering noise from the first and second reading whole-genome profiles. Noise filtering may include using at least one error suppression scheme to generate a first set of filtered readings for the first reading whole-genome profile and a second set of filtered readings for the second reading whole-genome profile. At least one error suppression protocol may include calculating the probability that any single nucleotide variant in the first and second profiles is an artificial mutation, and removing said mutation. The probability may be calculated as a function of features selected from: mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof. Alternatively or in combination, at least one error suppression protocol may include removing artificial mutations using a dissimilarity test between independent repeats of the same DNA fragment generated by polymerase chain reaction or sequencing processing. In addition to or as an alternative to dissimilarity testing, repeat consistency may be included, wherein artificial mutations are identified and removed when consistency is lacking in a large portion of a given repeat family. The method may further include calculating an estimated tumor fraction (eTF) for the first and second biological samples using first and second filtered read sets by applying a background noise model to one or more integrated mathematical models. The method may further include detecting residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.

[0017] In various embodiments, a method for detecting residual disease in a subject in need of it is provided. The method may include receiving a whole-genome summary of first subject-specific readings associated with a genetic marker in a first biological sample from the subject. The biological sample may include a baseline sample. The first reading summary may each contain copy number variations (CNVs), and wherein the baseline sample includes a tumor sample or a plasma sample. The method may further include receiving a second subject-specific whole-genome summary of readings associated with a genetic marker in a second biological sample from the subject. The second biological sample may include a peripheral blood mononuclear cell (PBMC) sample. The second summary of the genetic marker may each contain copy number variations (CNVs). The method may further include filtering artificial loci from the first and second summaries of readings. Filtering may include removing duplicate loci generated in a reference healthy sample cohort from the first and second reading summaries. Alternatively or in combination, filtering may include identifying CNVs shared between the first and second summaries as germline mutations and removing said mutations from the first and second reading summaries. The method may further include detecting readings in a third subject-specific whole-genome summary of the genetic marker in a third biological sample from the subject to generate a tumor-associated whole-genome representative of the genetic marker in the third sample. The method may further include normalizing each of the first, second, and third readout summaries to generate a first filtered readout set for the first readout genome-wide summary, a second filtered readout set for the second readout genome-wide summary, and a third filtered readout set for the third readout genome-wide summary. The method may further include using the third filtered readout set to compute an estimated tumor fraction (eTF) of the third biological sample by applying a background noise model to one or more integrated mathematical models. One or more models may be configured to generate a first eTF using the first filtered readout set, and / or one or more models may be configured to generate a second eTF using the second filtered readout set. The method may further include detecting residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold.

[0018] In some embodiments, this disclosure relates to methods for detecting residual disease in subjects who require it. Preferably, residual disease detection includes detecting minimal residual disease during treatment. In particular, this disclosure relates to detecting residual disease in one or more of the following situations: (a) after resection surgery; (b) during or after treatment; (c) while monitoring treatment efficacy; (d) while monitoring tumor recurrence or relapse; or (e) any combination thereof. In particular, this disclosure relates to detecting residual disease during or after chemotherapy, immunotherapy, targeted therapy, or combinations thereof and / or in the process of monitoring the effectiveness of such therapies.

[0019] In some embodiments, this disclosure relates to a method for detecting residual disease in a subject in need of it, comprising: (A) receiving a subject-specific whole-genome profile of genetic markers from a plurality of genetic markers in a biological sample of the subject, said biological sample including a tumor sample and optionally a normal cell sample, said genetic marker profile being selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome profile of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; and (C) filtering artificial noise markers from the whole-genome profile of the markers by: based on the probability of detecting noise (P) N (a) as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and 4) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; or based on 1) the position relative to the centromere; 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and 3) the overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (d) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (e) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model. In some embodiments of the aforementioned method, (1) for SNV markers, the estimated TF (eTF[SNV]) is calculated by integrating a process quality metric including estimated genomic coverage and sequencing noise with a patient-specific parameter including mutational burden (N); (2) for CNV markers, the estimated TF (eTF[CNV]) is calculated by integrating a skewed directional coverage depth consistent with the tumor CNV orientation, where copy number amplification is positively skewed and copy number deletion is negatively skewed. In some embodiments, ROC curves are used to optimize the BQ, MQ, and fragment size filters for the markers. In some embodiments, the method includes employing a combined base quality mapping quality (BQ MQ) filter.

[0020] In some embodiments, the residual disease detection method of this disclosure is performed by receiving a subject-specific whole-genome summary from multiple genetic markers, wherein the multiple genetic markers are derived from a biological sample containing a tumor sample of the subject and a normal sample containing a non-tumor sample. In some embodiments, the method includes generating a whole-genome summary of the markers using the subject's tumor sample and the subject's peripheral blood mononuclear cells (PMBCs). In particular, the whole-genome summary of the genetic markers is generated by whole-genome sequencing of the subject sample (e.g., the tumor sample) and a control sample (e.g., PMBCs). Preferably, the subject's tumor sample includes a resected tumor, such as a solid tumor removed after a surgical procedure; said surgical procedure includes, for example, gastrectomy, mastectomy, prostatectomy, skin lesion removal, small bowel resection, gastrectomy, thoracotomy, adrenalectomy, colectomy, oophorectomy, thyroidectomy, hysterectomy, tongue resection, or colon polypectomy, preferably thoracotomy.

[0021] In some embodiments, this disclosure relates to a method for detecting residual disease in a subject in need of it, comprising: (A) receiving a subject-specific whole-genome profile of genetic markers from a plurality of genetic markers in a biological sample of the subject, said biological sample including a tumor sample and optionally a normal cell sample, said genetic marker profile being selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome profile of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; and (C) filtering artificial noise markers from the whole-genome profile of the markers by: based on the probability of detecting noise (P) NThe following are considered as functions of: (a) the mapping quality (MQ) of the readset containing SNVs; (b) the fragment size length of the readset containing SNVs; (c) the consistency test within the read repeat family containing SNVs or Indels; and (d) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; and / or based on: (a) position relative to the centromere; (b) the mapping quality (MQ) of the readset containing CNV or SV windows; and (c) the overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (d) calculating an estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (e) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model, wherein the readset comprises a readset covering a specific SNV or indel site, or a readset contained within a specific CNV or SV genomic window. In some embodiments, normal cell samples include PMBCs, saliva samples, hair samples, or skin samples. In some implementations, the subject is a human being, and the subject's second biological sample comprises biological material selected from blood, cerebrospinal fluid, pleural fluid, ocular fluid, feces, urine, or combinations thereof.

[0022] In some embodiments of this disclosure, the tumor sample includes excised tumor or fine needle aspiration (FNA) sample, frozen tissue, optimal cutting temperature compound (OCT) embedded tissue, or formalin-fixed paraffin-embedded (FFPE) tissue.

[0023] In some embodiments of this disclosure, a normal sample comprises peripheral blood mononuclear cells (PMBC) or a saliva or skin sample.

[0024] In some embodiments of this disclosure, multiple genetic markers are received by sequencing the biological samples of the subject and control samples through whole-genome sequencing.

[0025] In some embodiments of this disclosure, the tumor genetic marker profile includes a high mutation rate and / or a high number of SNPs, indels, CNVs, or SVs, such as at least 1, at least 2, at least 3, at least 5, at least 7, at least 10, or more, for example, about 15 SNPs or indels per megabase pair, or CNVs / SVs with a cumulative size of at least 5 megabase pairs (MBP), at least 7 MBP, at least 10 MBP, or greater (e.g., a cumulative size of about 15 MBP).

[0026] In some embodiments, this disclosure relates to a method for detecting residual disease in a subject in need of it, comprising: (A) receiving a subject-specific whole-genome profile of genetic markers from a plurality of genetic markers in a biological sample of the subject, said biological sample including a tumor sample and optionally a normal cell sample, said genetic marker profile being selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome profile of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; and (C) filtering artificial noise markers from the whole-genome profile of the markers by: based on the probability of detecting noise (P) N (a) as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and 4) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; and / or based on 1) position relative to the centromere; 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and 3) overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (d) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (e) diagnosing residual disease in subjects based on the estimated tumor fraction and an empirical threshold calculated by a background noise model, wherein the empirical noise model is defined by measuring the detection error rate in normal healthy samples and converting it into a baseline noise eTF estimate.

[0027] In some embodiments of this disclosure, the eTF estimates the noise threshold at 0.0001 (10^6) -4 ) and 0.000001(10 -6 )between.

[0028] In some embodiments, this disclosure relates to a method for detecting residual disease in a subject in need of it, comprising: (A) receiving a subject-specific whole-genome summary of somatic genetic markers from a plurality of genetic markers in a biological sample of the subject, said biological sample including tumor samples and normal cell samples, wherein said genetic marker summary is selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) subsequently detecting the subject-specific whole-genome summary of the genetic markers in a second biological sample containing a plasma sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; (C) filtering artificial noise markers from the whole-genome summary of the markers by: based on the probability of detecting noise (P N The method is as follows: (a) the mapping quality (MQ) of the read group containing SNVs, (b) the fragment size length of the read group containing SNVs, (c) the consistency test within the read repeat family containing SNVs or Indels, and (d) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; and / or based on (c) the position relative to the centromere, (d) the mapping quality (MQ) of the read group containing CNVs or SV windows, and (e) the overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (f) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (f) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model. In some embodiments, normal cell samples include PMBCs, saliva samples, hair samples, or skin samples. In some embodiments, the subject is a human, and the subject's second biological sample contains biological material selected from blood, cerebrospinal fluid, pleural fluid, eye fluid, feces, urine, or combinations thereof. In some implementations, ROC curves are used to optimize the labeled BQ, MQ, and fragment size filters. In some implementations, the method includes employing a combined base quality mapping quality (BQ MQ) filter.

[0029] In some implementations, residual disease detection includes quantitatively estimating the patient's minimum residual disease burden during treatment, observation, or follow-up. Specifically, minimum residual disease detection includes detection of residual disease after resection surgery; detection of residual disease during or after treatment; detection of residual disease to monitor treatment effectiveness; detection of residual disease to monitor cancer recurrence or recurrence; or combinations thereof. In some implementations, minimum residual disease detection includes detection of residual disease after resection surgery including lymph node biopsy; head or neck surgery; uterine or endometrial biopsy; bladder biopsy; mastectomy; prostatectomy; removal of skin lesions; small bowel resection; gastrectomy; thoracotomy; adrenalectomy; colectomy; oophorectomy; thyroidectomy; hysterectomy; tongue resection; or colon polyp removal. In some implementations, minimum residual disease detection includes detection of residual disease after therapies including chemotherapy, immunotherapy, targeted therapy, radiotherapy, or combinations thereof.

[0030] In some embodiments of this disclosure, the disease detection method further includes receiving multiple genetic markers from a subject's biological sample, said biological sample including tumor samples and normal cell samples, and generating a subject-specific whole genome profile from the received multiple genetic markers.

[0031] In some embodiments of this disclosure, the disease detection method further includes detecting a subject-specific whole-genome profile of genetic markers in a second biological sample, such as a plasma sample. In some embodiments, a second biological sample from a subject is detected in a process (e.g., 2 days, 1 week, 2 weeks, 1 month, 2 months, 3 months, 4 months, 6 months, 1 year, 18 months, 2 years, 30 months, 3 years, 42 months, 4 years, 5 years, 7 years, 10 years, or longer, such as 15 years or 20 years) to generate a time-updated representative of tumor whole-genome genetic markers in the patient's plasma.

[0032] In some embodiments of this disclosure, the disease detection method includes empirically determining a background noise threshold, wherein tumor fractions above the background noise threshold provide a quantitative estimate of the tumor burden. Specifically, tumor fractions below the noise threshold are considered undetectable (ND).

[0033] In some embodiments of this disclosure, the disease detection method includes quantitative monitoring of tumor disease over time (e.g., tumor score). In some embodiments, the tumor is brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, osteosarcoma, or a solid tumor that is inherently heterogeneous or homogeneous. Preferably, the tumor is lung cancer, breast cancer, melanoma, bladder cancer, or osteosarcoma, such as lung adenocarcinoma, ductal adenocarcinoma, non-small cell lung cancer adenocarcinoma (NSCLC LUAD), skin melanoma, urothelial carcinoma, or osteosarcoma.

[0034] In some embodiments, the residual disease detection method of this disclosure further includes: calculating the eTF of SNV or indel markers using an integral probability model, the probability model including: 1) an integral signal of plasma SNV or indel detection; 2) a process quality metric including an estimated genome coverage and sequencing noise model; and 3) a patient-specific parameter including mutational burden (N); and / or calculating the eTF of CNV or SV markers using a probabilistic dilution model, the probabilistic dilution model including: 1) integrating the skewed directional coverage depth between plasma and normal patient samples based on the directionality of tumor CNV or SV, wherein copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative coverage depth of the skewed skew between tumor and normal (PBMC) patient samples; and 3) obtaining the dilution ratio between the above signals.

[0035] In some embodiments, the residual disease detection method of the present invention includes: (A) receiving multiple genetic markers, including single nucleotide variants (SNVs) or copy number variants (CNVs) or combinations thereof, in a biological sample of a subject and a normal cell sample of the subject, to generate a subject-specific whole-genome profile of the genetic markers; (B) identifying and filtering artificial noise markers from the marker whole-genome profile, wherein (1) noisy SNVs are identified by statistically classifying each SNV in the profile as a signal or noise based on the probability (PN) of detecting noise as a function of the base quality (BQ) and the mapping quality (MQ) of the SNV; and / or (2) noisy CNVs are identified based on their position relative to the centromere, within a given coverage area. The overlap of its cfDNA mask blacklist and its reading mapping in depth are identified by statistically classifying each CNV in the profile as signal or noise; (C) the estimated tumor fraction (eTF) of the sample is calculated based on one or more integrated mathematical models, where, for SNV markers, the estimated TF (eTF[SNV]) is calculated by the mathematical formula eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific profile detections in the patient sample, σ is the empirically estimated measure of noise, R is the total number of unique readings in the target region (ROI), N is the tumor mutation burden, and cov is the average of the unique readings at each site in the ROI; and / or for CNV markers, eTF [CNV] is calculated using the mathematical formula eTF[CNV]=(sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth value in a genomic window indexed by {i} representing plasma, T is the median depth value in a genomic window indexed by {i} representing tumor, and N is the median depth value in a genomic window indexed by {i} representing normal depth coverage. Specifically, in these implementations, the genomic window for estimating tumor fractions based on the detection of one or more CNV markers is approximately 500 base pairs (bp).

[0036] In some embodiments, this disclosure relates to a method for diagnosing a subject with minimal residual disease, comprising (A) receiving a genome-wide summary of readings from genetic data sequenced from multiple biological samples received from the subject, said biological samples including tumor samples, normal samples, and plasma samples; (B) performing mutation calls (including MUTECT, LOFREQ, and / or STRELKA mutation calls) on tumor and PBMC samples from the subject to generate subject-specific readings of somatic SNVs (sSNVs) or indels as a personalized reference set; and (C) collecting and filtering readings from subject-specific mutation sites, including (1) (i) Remove low-mapped quality reads (e.g., <29, ROC optimized); (ii) Establish replication families (representing multiple PCR / sequencing copies of the same DNA fragment) and generate corrected reads based on a consistency test; (iii) Remove low-base-quality reads (e.g., <21, ROC optimized); (iv) Remove high-fragment-size reads (e.g., >160, ROC optimized); (d) Calculate the number of subject-specific mutation sites with at least one supporting read (in the filtered set) and replace them with the exact same ones found in the tumor; (f) Based on the mathematical model eTF[SNV]=1-[1-(ME(σ)*R) / N]^ (1 / cov)…(Equation 1) estimates the tumor fraction of SNV, where M is the number of tumor-specific profile detections in the patient sample, σ is a measure of noise estimated empirically, R is the total number of unique readings within the target region (ROI), N is the tumor mutational burden, and cov is the average number of unique readings at each site in the ROI; (G) compares eTF[SNV] to a detection threshold that includes an empirically measured baseline noise TF estimate from healthy samples, where an eTF[SNV] higher than the threshold level (e.g., 2 standard deviations of the noise TF distribution (FPR < 2.5%)) indicates a positive detection; and (K) diagnoses residual disease in the subject based on eTF.

[0037] In some embodiments, this disclosure relates to a method for diagnosing minimal residual disease in a subject, comprising (A) receiving a read genome-wide summary from genetic data sequenced from multiple biological samples received from the subject, said biological samples including tumor samples, normal samples, and plasma samples; (B) performing CNV or SV calls on tumor and PBMC samples from the subject and generating reference segmentations of multiple CNV fragments exceeding a threshold length (e.g., >2 Mbp, preferably >5 Mbp) along the directional annotation of the fragments, wherein amplifications are positively annotated and deletions are negatively annotated; (C) collecting single-bp depth coverage information of plasma, tumor, and PBMC samples covering the patient-specific CNV segmentation target region (ROI); and (D) converting the patient-specific CNV or SV into a reference segmentation. (V) The ROI in the V partition is divided into 500bp windows, and the median (artificially suppressed) of each window for all samples is calculated; (E) Normalized depth coverage information for all 500bp windows is generated using the following methods: (a) robust z-score normalization for each sample; and / or (2) robust principal component analysis (RPCA); (F) Readings / windows are filtered from the patient-specific partition, wherein the filtering includes: (1) removal of low-map quality readings (e.g., <29, ROC optimized); and / or (2) removal of centromere regions (e.g., removal of windows with normalized normal values ​​greater than 10); and / or (3) removal of regions not represented in cfDNA (e.g., removal of windows not included in a cfDNA representation mask consisting of multiple cfDNA samples); and (G) using a mathematical model sum i [(P(i)-N(i))*sign[T(i)-N(i)]]-E(σ)…(Equation 2) integrates the skewed coverage direction depth between plasma and normal (PBMC) patient samples, where P is the median depth coverage value in the genomic window indexed by {i} representing plasma depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; E(sigma) is a measure of the error rate estimated empirically; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; and N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; (H) uses the mathematical model sum i[abs(T(i)-N(i))]-E(σ)) …(Equation 3), integrates the skewed cumulative depth of coverage between tumor and normal (PBMC) patient samples, where E(σ) is a measure of the empirically estimated error rate; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; (I) calculates the dilution rate between the directional depth coverage and the cumulative depth coverage (H) of (G), which corresponds to the estimated tumor score of CNV or SV ((eTF[CNV])=(sum i [(P(i)-N(i))*sign[T(i)-N(i)]]-E(σ)) / (sum i [abs(T(i)-N(i))]-E(σ))…(Equation 4); (J) compare the eTF [CNV] with a detection threshold that includes an empirically measured baseline noise TF estimate from healthy samples, where a positive detection is indicated by an eTF [CNV] higher than the threshold level (e.g., 2 standard deviations of the noise TF distribution (FPR < 2.5%)); and (K) diagnose residual disease in subjects based on eTF.

[0038] In some embodiments, this disclosure relates to a system for detecting residual disease in subjects in need of it, comprising, (A) an analysis unit configured and arranged to filter artificial noise markers from a tagged genome profile, wherein the tagged genome profile is generated from a plurality of genetic markers from a subject's biological sample, including tumor samples and normal cell samples, wherein the genetic marker profile is selected from the group consisting of single nucleotide variants (SNVs), indels, copy number variants, SVs, and combinations thereof; the analysis unit further comprises detecting a subject-specific genome profile of the genetic markers in a second biological sample containing a subject's plasma sample to generate a representative of tumor genome genetic markers in the patient's plasma; the analysis unit further comprises an engine selected from the group consisting of SNV and indel classification engines, CNV and SV classification engines, and combinations thereof, wherein: based on the probability of detecting noise (PN) (A) As a function of 1) the mapping quality (MQ) of the read group containing SNVs or Indels, 2) the fragment size length of the read group containing SNVs or Indels, 3) the consistency test within the read repeat family containing specific SNVs, and 4) the base quality (BQ) of SNVs or Indels, the SNV and Indel classification engine statistically classifies each SNV in the summary as signal or noise, and based on 1) the position relative to the centromere, 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and 3) the representation of CNVs or SV windows in the cfDNA data, the CNV and SV classification engine statistically classifies each CNV or SV in the summary as signal or noise; (B) an eTF unit configured and arranged to calculate the estimated tumor score (eTF) of the sample based on one or more integrated mathematical models; and (C) a display unit that outputs the subject's residual disease profile based on the estimated tumor score.

[0039] In some embodiments of the aforementioned system disclosed herein, the eTF unit is further configured and arranged to: calculate the eTF of SNV or Indel markers by integrating a probabilistic model that includes: 1) the integrated signal of plasma SNV or indel detection, 2) a process quality metric including an estimated genome coverage and sequencing noise model, and 3) a patient-specific parameter including mutational burden (N); and / or calculate the eTF of CNV or SV markers by utilizing a probabilistic mixture model that includes: 1) integrating the skewed directional coverage depth between plasma and normal patient samples based on the directionality of tumor CNV or SV, where copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative skewed coverage depth between tumor and normal patient samples; and 3) identifying the dilution ratio between the aforementioned signals.

[0040] In some embodiments of the aforementioned system disclosed herein, the tumor score estimation unit (B) includes a processor configured to execute computer-readable instructions that, when executed, perform a method for estimating the tumor score (eTF) of a sample based on one or more integrated mathematical models: (1) TF[SNV] = 1 - [1 - (ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific SNV profiles detected in the patient's plasma sample, σ is a measure of the error rate estimated empirically, R is the total number of unique readings in the target region of the SNV profile, N is the tumor mutational burden, and cov is the average number of unique readings at each site in the SNV profile ROI; and / or (2) eTF[CNV] = (sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]) ]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth coverage value in the genomic window indexed by {i} representing plasma depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort, where {i} is a discrete index used to count all genomic windows covering patient tumor-specific amplified and deleted genomic regions.

[0041] In some embodiments, this disclosure relates to a computer-readable medium including computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for detecting residual disease, the method or steps comprising: (A) receiving a subject-specific whole-genome summary of genetic markers from a plurality of genetic markers in a biological sample of a subject, the biological sample including a tumor sample and optionally a normal cell sample, wherein the genetic marker summary is selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome summary of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; and (C) filtering artificial noise markers from the whole-genome summary of the markers by: based on the probability of detecting noise (P) N(a) as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and 4) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; and / or based on 1) position relative to the centromere; 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, 3) overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (d) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (e) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model.

[0042] This disclosure also relates to methods for cancer stratification, including the detection of minimal residual disease (MRD) in cancer patients. The stratification method includes identifying low-abundance MRD-specific markers according to the methods described above; and detecting the markers to diagnose MRD. The cancer stratification method may further include methods for detecting tumors by RT-PCR using lung cancer-specific markers and / or molecular imaging using probes. Brief description of the attached diagram

[0043] Details of one or more embodiments of this disclosure are set forth in the accompanying drawings / tables and the following description. Further features, objects, and advantages of this disclosure will be apparent from the drawings / tables, the detailed description, and the claims. Figure 1A A schematic diagram of a diagnostic method of the present disclosure according to various embodiments is shown, which is used, for example, for detecting minimal residual tumor diseases. Figure 1B Representative workflows for detecting residual diseases in subjects are shown according to various implementation schemes. Figure 1C Representative workflows for detecting residual diseases in subjects are shown according to various implementation schemes. Figure 1D A representative workflow of this disclosure is shown for diagnosing minimal residual disease (MRD) in subjects based on single nucleotide polymorphism or indel measurements. Figure 1E A representative workflow of this disclosure is shown for diagnosing minimal residual disease (MRD) in subjects based on measurements of copy number variation or structural variation.

[0044] Figure 2A-2B A graph showing the detection probability based on external or internal parameters is presented. Figure 2A The detection probabilities for various tumor scores and coverage (up to the genomic equivalent limit: ~1000 molecules) based on the Bernoulli model are shown. Figure 2BThe detection probability of whole-genome SNV integration (binomial model) is shown, assuming integration of 20,000 point mutations.

[0045] Figure 3A-3K The effects of applying various filters according to various implementation schemes and the estimation of tumor scores provided by the method of the present invention are illustrated. Figure 3A The effect of applying a base quality (BQ) filter is shown. Figure 3B The effect of optimizing base quality filtration by receiver operating curve (ROC) is shown. Figure 3C The study demonstrates the effectiveness of a filter optimized with combined base quality (BQ) and mapping quality (MQ) in assessing the error rate distribution across multiple replicates using control samples, providing approximately 7 fold change (FC) suppression in sequencing errors. For lung cancer and melanoma cancer types, the pre-filter noise showed ~2 x 10⁻⁶. -3 For both cancer types, the post-filter noise rate decreased to ~2 x 10⁻⁶. -4 . Figure 3D The effect of applying a filter optimized with combined base quality (BQ) and mapping quality (MQ) is shown, with reduced 35X coverage. This filter allows for the detection of markers in samples with TF as low as 1 / 20,000. The red line represents the theoretical (binomial model) expected value, and the empirical measurements (means and confidence intervals of 5 independent replicates) are shown in black. Noise levels are represented by gray areas based on the detection distribution according to TF = 0. Figure 3E Computer validation of TF estimation in melanoma samples is shown. Inputting mixture TF (x-axis) against TF estimated from mutation patterns (y-axis) indicates a high correlation (R²). 2 =0.999), for values ​​higher than 5 x 10 -5 All TFs were accurately and specifically estimated. Figure 3F and Figure 3G Diagnostic methods according to various implementation schemes are illustrated, which allow for the detection of other types of solid tumors, such as lung tumor fractions. Figure 3F ) and breast cancer patients ( Figure 3G The characteristics of genetic biomarkers in tumors, even in cases where the tumor score (TF) is as low as 1 / 10000. Figure 3H Demonstrates reliable sSNV-based tumor score estimation with tumor scores (TF) as low as 5 x 10⁻⁶. -5 . Figure 3I Demonstrates reliable sCNV-based tumor score estimation, with tumor scores (TF) as low as 5 x 10⁻⁶. -5 Preferred TF > 10 -4 . Figure 3JThe diagram shows a strong correlation between TF estimates using SNV-based estimation (x-axis) and CNV-based estimation (y-axis). The gray quadrant shows the correlation below 5 x 10. -5 At the threshold TF, there is a weak correlation between the SNV-based estimate and the SNV-based estimate. Figure 3K A box plot is shown, illustrating a comparison of the present invention's method with the ICHOR-CNA method.

[0046] Figure 4 The SNV detection rates in background noise models (healthy PBMCs and cfDNA samples) of two cancer patients (BB1122, BB1125) and two healthy controls (BB600 and BB601) before (preoperative) and after (postoperative) resection are shown according to various implementation schemes.

[0047] Figure 5A and Figure 5B Clinical evaluation of patient samples using the systems and methods of this disclosure is illustrated. Figure 5A Exemplary evaluations of the systems and methods of this disclosure using clinical samples obtained from subjects with early-stage lung cancer and / or patients with minimal residual disease (MRD) are illustrated according to various embodiments. Data show tumor fraction (TF) estimates from preoperative and postoperative plasma samples of all patients analyzed. Only two patients showed a postoperative TF greater than 5 x 10⁻⁶. -5 The noise threshold was determined. However, the TF in all healthy control samples was below the detection threshold. ND indicates not detected. The data showed results consistent with the SNV method in terms of plasma detection and TF correlation. Figure 5B The calculation of z-scores is shown for 11 different samples obtained from adenocarcinoma patients. The data shows that the z-scores of healthy controls were below the threshold level (e.g., a z-score of 2, as indicated by the horizontal dashed line). Figure 5C The calculation of z-scores in 11 different samples obtained from adenocarcinoma patients is shown compared to negative controls across patients. Data shows that the z-scores of healthy controls were below a threshold level (e.g., a z-score of 2, as indicated by the horizontal dashed line). Consistency was observed between the sSNV-based and sCNV-based detection methods. Figure 5D ).

[0048] Figures 6A-6E This paper presents a method for analyzing large-scale directional depth coverage skewness across large genome CNV regions. Figure 6A The integral of the sparse CNV skewness at TF=0.001 is shown, with the upper figure showing the synthetic plasma (TF=10) in the amplified 10Kbp region. -3A comparison of single-bp depth coverage between plasma and matched PBMCs; the middle plot shows the residuals between plasma and PBMCs, and the lower plot shows the sum of the residuals. In the middle plot, note the sparse but positive bias of the residuals, while in the lower plot, the sum of the residuals (signals) accumulates, partly due to the positive bias of amplification, when integrated onto the genome. Figure 6B This diagram shows an overview of tumor read depth (red), germline read depth (pink), and preoperative plasma cfDNA read depth (blue) in representative amplified regions. Preoperative plasma shows read depths comparable to germline DNA, but also exhibits amplification depth skewness at the telomere ends of the amplified regions. As described above, this mathematical method integrates the read depth skewness across the entire genome. Figure 6C The signal-to-noise ratio (SNR) for each TF is shown, where 10 -6 All of the above TFs showed positive (>0) SNR detection (indicating high sensitivity). Figure 6D The study showed a linear relationship between plasma SNR and TF in CNV (dilution model), with similar dynamics observed in lung / melanoma / breast cancer patients. Figure 6E This plot shows the skewness relative to tumor fraction (TF) when acquiring neutral regions of the genome (e.g., regions excluding amplification and / or deletion). It can be seen that in these regions, there is no bias in depth coverage skewness between plasma and PBMCs, and the probabilities of positive and negative skewness are similar. Therefore, there is no signal and SNR = 0, independent of TF (x-axis).

[0049] Figures 7A-7C Schematic diagrams of the systems of this disclosure according to various implementation schemes are provided.

[0050] Figure 8 Representative flowcharts based on various implementation schemes are provided, outlining the identification and / or classification of postoperative cancer subjects as candidates for adjuvant therapy.

[0051] Figure 9 The patient-specific sSNV scores of the various embodiments described in this paper are compared with those of ICHOR (Broad Institute). In particular, the detection sensitivity is improved by approximately 100-fold compared to the ICHOR detection method from the MIT-Broad Institute.

[0052] Figures 10A-10E The use of orthogonal features, such as fragment size, in the diagnostic methods of this disclosure is shown, as well as the accompanying effects of applying such orthogonal features in SNV-based methods. Figure 10A The distribution of fragment sizes is shown in a healthy, normal cfDNA sample. Figure 10BThe images show the fragment size shift in breast tumor cfDNA (red and purple) compared to normal cfDNA samples. Figure 10C The results showed that in a mouse xenograft (PDX) model, circulating DNA from tumor origin was significantly shorter than circulating DNA from normal origin. Figure 10D A linear graph showing the size of a DNA fragment (x-axis; number of bases) relative to the frequency of fragments of that length observed in tumor and normal samples. Figure 10E Patient-specific mutation detection using orthogonal features, such as the correspondence between DNA fragments of tumor origin based on fragment size distribution (x-axis) and GMM combined log-ratio ratio (y-axis), is shown.

[0053] Figures 11A-11J The use of orthogonal features, such as fragment size, in the diagnostic methods of this disclosure is shown, as well as the accompanying effects of applying such orthogonal features in CNV-based methods. Figure 11A Linear plots showing genomic region (bp) skewness relative to cumulative plasma depth coverage (bottom plot), plasma skewness relative to normal depth coverage (middle plot), and coverage (top plot) are displayed. Figure 11B The relationship between log2 of depth coverage (log2 > 0.5 = amplification, log2 < -0.5 = deletion) and the local fragment size centroid (COM) in the segment is shown. Figure 11C The relationship between depth-coverage-based CNV detection and fragment-size-based center of mass (COM)-based CNV detection in patient samples was shown. Figure 11D This study demonstrates a lack of relationship between depth-coverage-based CNV detection and fragment-size-based CNV detection in normal (healthy) plasma samples. Figures 11E and 11F show the COM, absolute slope values, and R values ​​for two treated patients. 2 Changes were observed at baseline (day 0) and at days 21 and 42 post-treatment. Figure 11G The relationship between the fragment size log2 slope and the patient's tumor score is shown. Figure 11H Clinical study results from cancer patients show the association between time to recurrence-free period and tumor DNA testing (zscore) 2 weeks post-surgery. Figure 11I Bar graphs showing tumor scores for four patients at baseline (day 0), midpoint (day 21), and end (day 42) of treatment. Figure 11J Bar graphs showing the normalized CNV scores of four patients at baseline (day 0), midpoint (day 21), and end (day 42) of treatment. Detailed description

[0054] The following descriptions of the various embodiments are merely exemplary and illustrative and should not be construed as limiting or restrictive in any way. Other embodiments, features, purposes, and advantages of this teaching will be apparent from the specification, drawings, and claims.

[0055] Unless otherwise defined, scientific and technical terms used in conjunction with the teachings described herein shall have the meanings commonly understood by one of ordinary skill in the art. The terminology used in the description of the disclosure herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. Furthermore, unless the context requires otherwise, singular terms shall include plural terms, and plural terms shall include singular terms. Generally, the terms used in connection with the following and the techniques described herein are well-known and commonly used in the art: molecular biology, protein and oligonucleotide or polynucleotide chemistry and hybridization. The use of standard techniques, e.g., for the purification and preparation of nucleic acids, chemical analysis, recombinant nucleic acid and oligonucleotide synthesis. Enzymatic reactions and purification techniques performed according to the manufacturer's instructions (or as commonly performed in the art or as described herein). The techniques and procedures described herein are generally performed according to conventional methods well known in the art and described in various general and more specific references cited and discussed throughout this specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (3rd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). The terminology, laboratory procedures, and techniques used in connection with this document are well-known and commonly used in the field.

[0056] The various embodiments of this disclosure are described in further detail in the following paragraphs.

[0057] As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise. Similarly, as used herein, “and / or” means and covers any one or more of the associated listed items and all possible combinations thereof, as well as any lack of combinations when interpreted in an alternative manner (“or”).

[0058] The word “about” indicates a range of plus or minus 10% of the value; for example, “about 5” means 4.5 to 5.5, “about 100” means 90 to 100, and so on, unless the context of this disclosure indicates otherwise or is inconsistent with this interpretation. For example, in a list of values ​​such as “about 49,” “about 50,” “about 55” means a range extending to less than half the interval between the previous and subsequent values, such as greater than 49.5 to less than 52.5. Furthermore, the phrases “less than about” or “greater than about” should be understood based on the definition of the term “about” provided herein.

[0059] Where a range of values ​​is provided in this disclosure, it is intended that every intermediate value between the upper and lower limits of that range, as well as any other declared value or intermediate value within that range, be included in this disclosure. For example, if the declared range is 1 μM to 8 μM, it is intended that 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, and 7 μM be also explicitly disclosed.

[0060] As used in this article, the term "multiple" can be 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.

[0061] As used herein, the term "detection" refers to the process of determining a value or set of values ​​relating to a sample by measuring one or more parameters in the sample, and may further include comparing the test sample with a reference sample. According to this disclosure, tumor detection includes identifying, determining, measuring, and / or quantifying one or more markers.

[0062] As used herein, the term "diagnosis" refers to a method that determines whether a subject is likely to have a given disease or condition, including but not limited to diseases or conditions characterized by genetic variations. A technician typically makes a diagnosis based on one or more diagnostic indicators (e.g., markers) whose presence, absence, quantity, or variation in quantity indicates the presence, severity, or absence of a disease or condition. Other diagnostic indicators may include patient history; physical symptoms such as unexplained weight loss, fever, fatigue, pain, or skin abnormalities; phenotype; genotype; or environmental or genetic factors. A technician will understand that the term "diagnosis" refers to an increased likelihood of a particular course or outcome occurring; that is, a course or outcome is more likely to occur in patients exhibiting a given characteristic (e.g., the presence or level of a diagnostic indicator) compared to individuals who do not exhibit that characteristic. The diagnostic methods of this disclosure may be used alone or in combination with other diagnostic methods to determine whether a course or outcome is more likely to occur in patients exhibiting a given characteristic.

[0063] In the context of “normal cells,” the term “normal” refers to cells with an untransformed phenotype or cells exhibiting the untransformed cell morphology of the tissue type being examined (e.g., PBMCs). In some embodiments, “normal sample” as used herein includes non-tumor samples, such as saliva samples, skin samples, hair samples, etc. It should be noted that the methods of this disclosure can be implemented without using normal samples.

[0064] As used herein, the term "abnormal" generally refers to a state of a biological system that deviates to some extent from the normal (e.g., wild-type) state. Abnormal states can occur at the physiological or molecular level. Representative examples include, for example, physiological states (diseases, pathologies) or genetic aberrations (mutations, single nucleotide variants, copy number variants, gene fusions, indels, etc.). Disease states can be cancer or precancerous conditions. Abnormal biological states may be associated with the degree of aberration (e.g., quantitative measurements indicating the distance from the normal state).

[0065] As used in this article, the term "possibility" generally refers to probability, relative probability, presence or absence, or degree.

[0066] As used herein, the term "tumor" includes any cell or tissue that may have undergone transformation at the genetic, cellular, or physiological level compared to normal or wild-type cells. The term generally refers to tumor growth, which can be benign (e.g., a tumor that does not metastasize and destroys adjacent normal tissue) or malignant / cancerous (e.g., a tumor that invades surrounding tissue and is often capable of metastasizing, which may recur after attempted removal and is likely to cause host death unless properly treated). See Steadman's Medical Dictionary, 28. th Ed Williams & Wilkins, Baltimore, MD (2005).

[0067] The term "cancer" (used interchangeably with "tumor") refers to human cancers and carcinomas, sarcomas, adenocarcinomas, lymphomas, leukemias, solid tumors, and lymphomas. Examples of different types of cancer include, but are not limited to, lung cancer, pancreatic cancer, breast cancer, stomach cancer, bladder cancer, oral cancer, ovarian cancer, thyroid cancer, prostate cancer, uterine cancer, testicular cancer, neuroblastoma, squamous cell carcinoma of the head, neck, cervix, and vagina, multiple myeloma, soft tissue and osteosarcoma, colorectal cancer, liver cancer, kidney cancer (e.g., RCC), pleural cancer, cervical cancer, anal cancer, bile duct cancer, gastrointestinal carcinoid tumors, esophageal cancer, gallbladder cancer, small bowel cancer, central nervous system cancers, skin cancer, choriocarcinoma, osteosarcoma, fibrosarcoma, glioma, melanoma, etc. In some implementations, "liquid" cancers, such as blood cancers, such as lymphomas and / or leukemias, are excluded.

[0068] Exemplary cancers include, but are not limited to, adrenocortical carcinoma, AIDS-related cancers, AIDS-related lymphomas, anal cancer, colorectal cancer, anal canal cancer, appendix cancer, pediatric cerebellar astrocytoma, pediatric brain astrocytoma, basal cell carcinoma, skin cancer (non-melanoma), bile duct cancer, extrahepatic bile duct cancer, intrahepatic bile duct cancer, bladder cancer, and urinary bladder cancer. Cancer, bone and joint cancer, osteosarcoma and malignant fibrous histiocytoma, brain cancer, brain tumor, brainstem glioma, cerebellar astrocytoma, brain astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma, breast cancer, bronchial adenoma / carcinoid, carcinoid, gastrointestinal, nervous system cancer, nervous system lymphoma, central nervous system cancer, central nervous system lymphoma, cervical cancer, childhood cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorders, colon cancer, colorectal cancer, cutaneous T-cell lymphoma, lymphoid tumor, mycosis fungoides, Seziary syndrome, endometrial cancer, esophageal cancer, extracranial germ cell tumor. Tumors, extragonadal germ cell tumors, extrahepatic bile duct carcinomas, ocular cancer, intraocular melanoma, retinoblastoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), germ cell tumors, ovarian germ cell tumors, gestational trophoblastic tumors, gliomas, head and neck cancer, hepatocellular carcinoma, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma, ocular cancer, islet cell tumors (endocrine pancreas), Kaposi's sarcoma, kidney cancer, renal cell carcinoma. Cancer, laryngeal cancer, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, hairy cell leukemia, lip and oral cancer, liver cancer, lung cancer, non-small cell lung cancer, small cell lung cancer, AIDS-related lymphoma, non-Hodgkin lymphoma, primary central nervous system lymphoma, Waldenstram macroglobulinemia, medulloblastoma, melanoma, intraocular (ocular) melanoma, Merkel cell carcinoma, malignant mesothelioma, mesothelioma, metastatic squamous neck cancer, oral cancer, tongue cancer, multiple endocrine neoplasia syndrome, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative disorders, chronic myeloid leukemia, acute myeloid leukemia, multiple myeloma, chronic myeloproliferative disorders, nasopharyngeal carcinoma, neuroblastoma, oral cancer, oral cancerCavity cancer, oropharyngeal cancer, ovarian cancer, ovarian epithelial cancer, low-grade malignant potential ovarian tumors, pancreatic cancer, islet cell pancreatic cancer, sinus and nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal cell tumor and supratentorial primitive neuroectodermal tumor, pituitary adenoma, plasma cell tumor / multiple myeloma, pleural pulmonary blastoma, prostate cancer, rectal cancer, renal pelvis and ureter cancer, transitional cell carcinoma, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, Ewing's sarcoma family, Kaposi's. Sarcoma, uterine cancer, uterine sarcoma, skin cancer (non-melanoma), skin cancer (melanoma), Merkel cell skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, gastric cancer, supratentorial primitive neuroectodermal tumor, testicular cancer, laryngeal cancer, thymoma, thymoma and thymic carcinoma, thyroid cancer, transitional cell carcinoma of the renal pelvis and ureter and other urinary organs, gestational trophoblastic tumor, urethral cancer, endometrial uterine cancer, uterine sarcoma, uterine corpus cancer, vaginal cancer, vulvar cancer, and Wilm's tumor.

[0069] As used herein, the term “non-small cell lung cancer” or NSCLC refers to all lung cancers that are not small cell lung cancers and includes several subtypes, including but not limited to large cell carcinoma, squamous cell carcinoma, and adenocarcinoma. It includes all stages and metastases. Squamous cell carcinoma accounts for 25% of lung cancers and typically begins near the central bronchus. Hollow cavities and associated necrosis are often found in the center of the tumor. Well-differentiated squamous cell carcinomas generally grow more slowly than other types of cancer. Adenocarcinoma accounts for 40% of non-small cell lung cancers. It typically originates in the peripheral lung tissue. Most adenocarcinoma cases are associated with smoking; however, in never-smokers, adenocarcinoma is the most common form of lung cancer. See, Rosell et al., Lung Cancer, 46(2), 135-48, 2004; Coate et al., Lancet Oncol, 10, 1001-10, 2009.

[0070] As used herein, the term "residual disease" refers to the persistence of residual neoplastic cells even after interventions such as surgical intervention, radiation ablation, or chemotherapy. The term "minimal residual disease" (MRD) describes a situation where morphologically normal tissue (e.g., lung tissue) can still have a relevant number of residual malignant cells after treatment of the tumor (e.g., chemotherapy, immunotherapy, or targeted therapy). The detection of minimal residual disease (MRD) is a novel and practical tool for more precisely measuring remission induction during treatment. In the case of liquid tumors (e.g., lymphoma or myeloma), the term MRD can refer to less than 10 -4 For example, 10 -5 Or even 10 -6The detection limit. In the case of solid tumors, the term “minimal residual disease” may refer to a situation where tumor markers are below what can be detected using conventional detection methods such as ctDNA detection or plasma DNA analysis. In some embodiments, MRD refers to a situation where less than 100 copies, preferably less than 40 copies, and particularly less than 10 copies of ctDNA are detected per 5 ml of plasma (Bettegowda et al., Sci Transl Med., 6(224), 224ra24, 2014).

[0071] As used herein, the term "subject" refers to a mammal, including humans, veterinary or farm animals, livestock or pets, and animals commonly used in clinical research. Specifically, a subject is a human subject, such as a human patient diagnosed with or suspected of having a tumor. A subject may have, potentially have, or be suspected of having one or more characteristics selected from cancer, cancer-related symptoms, or be asymptomatic or undiagnosed regarding cancer (e.g., not diagnosed with cancer). The subject may have cancer, the subject may exhibit cancer-related symptoms, the subject may not have cancer-related symptoms, or the subject may not have been diagnosed with cancer. In some embodiments, the subject is a person.

[0072] As used herein, the term “single nucleotide polymorphism” or “single nucleotide variation” (“SNP” or “SNV”) in relation to mutation refers to a difference of at least one nucleotide in a sequence compared to another sequence.

[0073] The term "copy number variation" or "CNV" refers to a comparative numerical change in the presence or absence / gain or loss of gene segments with the same nucleotide sequence. In the human genome, copy number variations can involve homozygous or heterozygous duplications or duplications of one or more DNA segments, or homozygous or heterozygous deletions of one or more DNA segments. For duplications / duplications of CNVs, the directionality of the CNV is typically represented as positive, while for deletions of CNVs, it is typically represented as negative.

[0074] As used herein, the term "indel" refers to a location on the genome where one allele contains one or more bases while another allele does not. From an evolutionary perspective, insertions and deletions are distinct, but in analyses such as those described herein, they are often not distinguished because an insertion in one allele is equivalent to a deletion in another. Therefore, the term indel refers to the insertion / deletion location between two alleles.

[0075] As used herein, the term "structural variant" refers to an alteration of a portion of a chromosome, rather than a change in the number of chromosomes or a set of chromosomes in the genome. There are four common types of mutations that lead to structural variants: deletions and insertions, such as duplications (involving changes in the amount of DNA in a chromosome, loss and gain of genetic material, respectively), inversions (involving changes in the arrangement of chromosome segments), and translocations (involving changes in the position of chromosome segments, which can cause gene fusions). In this invention, the term "structural variant" includes loss of genetic material, addition of genetic material, translocations, gene fusions, and combinations thereof.

[0076] As used herein, the term "sample" refers to a composition obtained from or derived from a subject of interest, comprising, for example, cells and / or other molecular entities that will be characterized and / or identified based on physical, biochemical, chemical, and / or physiological characteristics. Preferably, the sample is a "biological sample," which refers to a sample derived from a living organism, such as cells, tissues, organs, etc. In some embodiments, the source of the tissue sample may be blood or any blood component; body fluids; organ or tissue samples or solid tissues from fresh, frozen, and / or preserved sources, or from biopsies or aspirations; and cells or plasma at any time during the subject's pregnancy or development. Samples include, but are not limited to, primary or cultured cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, eye fluid, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, urine, cerebrospinal fluid (CSF), saliva, sputum, tears, sweat, mucus, tumor lysates, and tissue culture media, as well as tissue extracts (e.g., homogenized tissue), tumor tissue, and cell extracts. The samples further include biological samples that have been processed in any way after acquisition, such as by reagent treatment, solubilization or enrichment of certain components (e.g., proteins or nucleic acids), or embedded in a semi-solid or solid matrix for sectioning purposes, such as thin sections of tissue or cells in a histological sample. Samples may contain environmental components such as water, soil, mud, air, resin, minerals, etc. In some embodiments, the samples may include biological samples containing DNA (e.g., gDNA), RNA (e.g., mRNA, tRNA), proteins, or combinations thereof obtained from a subject (e.g., a human or other mammalian subject).

[0077] As used herein, the terms "cell" and "biological cell" are used interchangeably. Non-limiting examples of biological cells include eukaryotic cells, plant cells, animal cells (e.g., mammalian cells, reptile cells, avian cells, fish cells, etc.), prokaryotic cells, bacterial cells, fungal cells, protozoan cells, cells isolated from tissues (e.g., muscle, cartilage, fat, skin, liver, lung, nerve tissue, etc.), immune cells (e.g., T cells, B cells, natural killer cells, macrophages, etc.), embryos (e.g., fertilized eggs), oocytes, egg cells, sperm cells, hybridomas, cultured cells, cells derived from cell lines, cancer cells, infected cells, transfected and / or transformed cells, reporter cells, etc. Mammalian cells may be derived, for example, from humans, mice, rats, horses, goats, sheep, cattle, primates, etc.

[0078] As used herein, the term "marker" refers to a characteristic that can be objectively measured as an indicator of a normal biological process, a pathogenic process, or a pharmacological response to a therapeutic intervention (e.g., treatment with anticancer drugs). Representative types of markers include molecular variations, such as the structure (e.g., sequence) or number of markers, including, for example, gene mutations, gene duplications, or multiple differences, such as somatic alterations, copy number variations, tandem duplications, or combinations thereof in cfDNA.

[0079] As used herein, the term "genetic marker" refers to a DNA sequence located at a specific position on a chromosome that can be measured in a laboratory. The term "genetic marker" can also be used to refer, for example, cDNA and / or mRNA encoded by a genomic sequence, as well as the genomic sequence itself. A genetic marker may include two or more alleles or variants. A genetic marker may be direct (e.g., located within a target gene or locus (e.g., a candidate gene)) or indirect (e.g., closely linked to a target gene or locus due to proximity but not within the target gene or locus). Furthermore, a genetic marker may also be independent of genes or loci (e.g., SNVs, CNVs, indels, SVs, or tandem repeats) present in non-coding regions of the genome. Genetic markers include nucleic acid sequences that encode or do not encode gene products (e.g., proteins). In particular, genetic markers include single nucleotide polymorphisms / variations (SNPs / SNVs) or copy number variations (CNVs) or combinations thereof. Preferably, compared to a reference sample, genetic markers include somatic variations in DNA, such as sSNVs or sCNVs, indels, SVs, or combinations thereof.

[0080] As used herein, the term “cell-free DNA” or “cfDNA” refers, for example, a cell-free strand of deoxyribonucleic acid (DNA) extracted or isolated from circulating blood plasma / serum, or from lymph, cerebrospinal fluid (CSF), urine, or other bodily fluids. The term “cfDNA” is distinctly different from “circulating tumor DNA” or “ctDNA.” Cell-free DNA (cfDNA) is a broader term that describes DNA that circulates freely in the bloodstream but is not necessarily derived from a tumor.

[0081] As used herein, the term “germ DNA” or “gDNA” refers to DNA isolated or extracted from a patient’s peripheral mononuclear blood cells (including lymphocytes obtained from circulating blood).

[0082] As used herein, the term "variation" refers to a change or deviation. In the context of nucleic acids, variation refers to differences or changes between DNA nucleotide sequences, including copy number differences (CNVs). This actual difference in nucleotides between DNA sequences can be SNPs and / or variations in the DNA sequence, such as changes observed when comparing a sequence to a reference (e.g., germline DNA (gDNA) or a reference human genome HG38 sequence), such as fusions, deletions, additions, duplications, etc. Preferably, variation refers to differences between cfDNA sequences and control DNA sequences not derived from tumor cells, such as when cfDNA is compared to a reference HG38 sequence, or when cfDNA is compared to gDNA. Differences identified in gDNA and cfDNA are considered "constitutional" and can be disregarded.

[0083] As used herein, the term "control" refers to a reference for a test sample, such as control DNA isolated from peripheral mononuclear blood cells and lymphocytes, where these cells are not cancer cells, etc. As used herein, a "reference sample" refers to a sample of tissue or cells that may or may not have cancer for comparison. Thus, a "reference" sample provides a basis for comparison with another sample, such as a plasma sample containing cfDNA. Conversely, a "test sample" is a sample compared to a reference or control sample. A reference sample does not necessarily have to be cancer-free, for example, when reference and test samples are obtained from the same patient at different times.

[0084] In some implementations, the reference sample or control may include a reference assembly. The term "reference assembly" refers to a digital nucleic acid sequence database, such as the Human Genome (HG38) database containing the HG38 component sequence (assembled in December 2013). This is accessible via the Human Genome Explorer Gateway at the University of California, Santa Cruz (UCSC) via the URL GENOME.UCSC.EDU. Alternatively, a reference assembly may also refer to the Human Genome Assembly (version 38; assembled in June 2017) from the Genome Reference Association, which is accessible on the Internet through the website of the US National Center for Biotechnology Information (NCBI).

[0085] As used herein, the term “sequencing” or “sequencing determination” as a verb refers to the process of determining the nucleotide sequence or order of nucleotides in DNA, such as the nucleotide sequence AGTCC. The term “sequence” as a noun refers to the actual nucleotide sequence obtained from sequencing; for example, DNA having the sequence AGTCC. Wherein “sequence” is provided and / or received in digital form, such as on a disk or remotely via a server, “sequencing” can refer to a collection of DNA that has been propagated, manipulated, and / or analyzed using the methods and / or systems of this disclosure.

[0086] As used herein, the term “DNA sequence” generally refers to “raw sequence reading” and / or “consistency”. A raw sequence reading is the output of a DNA sequencer and typically contains redundant sequences of the same parent molecule, such as after amplification. A “shared sequence” is a sequence of redundant sequences derived from a parent molecule intended to represent the original parent molecule. Consistency can be generated by voting (where each majority nucleotide, e.g., the nucleotide most frequently observed at a given base position, is a shared nucleotide in the sequence) or by other methods such as comparison with a reference genome. Consistency can be generated by tagging the original parent molecule with a unique or non-unique molecular tag (e.g., a barcode) that allows for tracking of progeny sequences (e.g., after PCR).

[0087] Sequencing methods can be first-generation sequencing methods, such as Maxam-Gilbert or Sanger sequencing, or high-throughput sequencing methods (e.g., next-generation sequencing or NGS). High-throughput sequencing methods can simultaneously (or substantially simultaneously) sequence at least 10,000, 100,000, 1,000,000, 10 million, 100 million, 1 billion, or more polynucleotide molecules. Sequencing methods can include, but are not limited to: pyrosequencing, sequencing by synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, digital gene expression (Helicos), massively parallel sequencing, such as Helicos, cloning single-molecule arrays (Solexa / Illumina), and sequencing using PACBIO, SOLID, Ion Torrent, or NANOPORE platforms.

[0088] The term “whole genome sequencing” refers to the laboratory procedure of determining the DNA sequence of each DNA strand in a sample. The resulting sequence may be referred to as “raw sequencing data” or “readings.” As used herein, a reading is a “mappable” reading when the sequence shows similarity to a region of a reference chromosomal DNA sequence. The term “mappable” can refer to a region that shows similarity to a reference sequence and is therefore “mapped,” for example, a cfDNA segment that shows similarity to a reference sequence in a database, such as cfDNA with a high percentage of similarity to the human chromosome region 8q248q24.3 in the Human Genome (HG38) database is a “mappable reading.”

[0089] "Deep sequencing" is a general concept that involves a large number of repeated reads of each region of a sequence.

[0090] As used in this article, the term "mapping" generally refers to aligning a DNA sequence with a reference sequence based on sequence homology. Alignment algorithms can be used for this purpose, such as the Needleman-Wunsch algorithm, BLAST, or EMBOSS.

[0091] In addition to “WGS”, targeted sequencing can be used to obtain a genome profile. Compared to WGS, the term “targeted sequencing” as used herein refers to a laboratory procedure for determining the DNA sequence of selected DNA loci or genes in a sample, such as sequencing selected groups of cancer-related genes or markers (e.g., targets). In this context, the term “target sequence” refers to a selected target polynucleotide, such as a sequence present in a cfDNA molecule, whose presence, quantity, and / or nucleotide sequence or variations therein need to be determined. The target sequence is queried to determine the presence of somatic mutations. The target polynucleotide may be a gene region associated with a disease such as cancer. In some embodiments, this region is an exon.

[0092] As used herein, the term "low abundance" for cfDNA refers to a sample containing less than about 20 ng / mL, such as about 15 ng / mL, about 10 ng / mL or less, such as about 9 ng / mL, 8 ng / mL, 7 ng / mL, 6 ng / mL, 5 ng / mL, 4 ng / mL, 3 ng / mL, 2 ng / mL, 1 ng / mL, 0.7 ng / mL, 0.5 ng / mL, 0.3 ng / mL or less, such as 0.1 ng / mL or even 0.05 ng / mL. In some implementations, the term "low abundance" can be understood in the context of the uniqueness of the marker, such as length or base composition. For example, although a subject's sample may contain a large amount of cfDNA (e.g., >20 ng / mL), the actual number of unique genetic markers contained in the cfDNA (e.g., sSNV, sCNV, indel, SV) may be very low. Typically, this parameter is expressed as genomic equivalent (GE) or coverage, as described below. In some implementations, the term "low abundance" can be understood in the context of tumor specificity of the marker. For example, although a subject's sample may contain a large amount of cfDNA (e.g., >20 ng / mL), the vast majority of genetic markers contained in the cfDNA (e.g., sSNV, sCNV, indel, SV) may be redundant and / or relevant to a reference (e.g., PBMC gDNA). Typically, this parameter is expressed as a tumor fraction (TF), as described below.

[0093] As used herein, the terms "tumor-specific" or "tumor-associated" with respect to cfDNA refer to the difference in the DNA sequence of cfDNA in a subject with cancer-forming tumors (e.g., a lung cancer patient) when compared to reference DNA, such as when cfDNA is compared to control DNA (gDNA) from non-tumor cells, as described herein. Alternatively, "tumor-specific" may refer to cfDNA collected during or after treatment, in the context of comparisons with cfDNA collected during or after treatment.

[0094] As used in this article, the term "read repeat family" includes PCR and sequencing repeats. Typically, these are independent repeats of the same unique fragment and can therefore be used in statistical tests (consistency tests) to correct low-frequency PCR and sequencing errors.

[0095] The terms "coverage" or "reading depth" refer to sequencing operations. For example, 20X coverage indicates moderate sequencing operations, while 35X or higher coverage indicates high sequencing operations, and 5X coverage indicates low sequencing operations. In embodiments of this disclosure, coverage is typically between about 5X and about 100X, particularly between 15X and about 40X, such as 20X, 30X, 35X, 40X, 50X, 70X, or greater.

[0096] As used in this article, “deep coverage” refers to the number of unique readings that overlap at or above a specific genome coordinate.

[0097] As used herein, the term "cfDNA coverage mask" refers to a mask of a genomic region representing the coverage of cfDNA reads in a normal cfDNA cohort. As is known in the art, cfDNA coverage is not perfectly uniform (less accessible chromatin genomic regions are represented), therefore, to eliminate bias, blacklists or masks may be implemented to allow selective analysis of well-covered regions.

[0098] As used herein, the term “reading mappability” refers to a numerical (e.g., percentage identity) or statistical measure (e.g., confidence estimate) of the accuracy of a reading’s mapping to the genome.

[0099] As used herein, the term “mutation load” or “N” refers to the level (e.g., number) of changes (e.g., one or more genetic alterations, particularly one or more individual cell alterations) per preselected unit (e.g., per megabase pair) within a predetermined genomic window. Mutation load can be measured, for example, based on the entire genome or exome, or based on a subset of the genome or exome. In some embodiments, mutation load measured based on a subset of the genome or exome can be extrapolated to determine the mutation load of the entire genome or exome. In some embodiments, mutation load is measured in a sample from a subject, such as a subject described herein, for example, a tumor sample (e.g., a lung tumor sample or a sample obtained from or derived from a lung tumor). Preferably, mutation load is a measure of the number of mutations per megabase pair (1,000,000 bp or MBP) of cfDNA. As is known in the art, mutation load can vary depending on tumor type, genetic lineage, and other subject-specific characteristics (e.g., age, sex, tobacco consumption, etc.). In the context of tumor diagnosis, the mutational burden can range from approximately 1,000 to approximately 10,000 mutations per MBP, for example, approximately 1,000, 2,000, 4,000, 6,000, 8,000, 10,000, 12,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 60,000, 70,000, 80,000, 90,000, 10,000, 10,000 or more, such as approximately 200,000 per MBP. Typically, the mutational burden is approximately 8,000 per MBP in non-smokers, while in subjects with melanoma, the mutational burden is more than 40,000 per MBP.

[0100] As used in this article, the term "genomic window" refers to a region of DNA within the boundaries of a selected nucleotide sequence. Windows may be separate from or overlap each other.

[0101] As used herein, the term "tumor fraction" or "TF" refers to the level (e.g., number) of tumor DNA molecules relative to normal DNA molecules. In some embodiments, "tumor fraction" refers to the proportion of circulating cell-free tumor DNA (ctDNA) to the total amount of cell-free DNA (cfDNA). The tumor fraction is believed to indicate the size of the tumor. Typically, the tumor fraction (TF) is between about 0.001% and about 1%, for example, about 0.001%, 0.05%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1% or higher, such as 2%.

[0102] The term "abundance" can refer to binary (e.g., absent / present), qualitative (e.g., absent / low / moderate / high), or quantitative information (e.g., a value proportional to quantity, frequency, or concentration) indicating the presence of a particular molecular species. In this context, mutations present at higher relative concentrations are associated with a large number of malignant cells, such as cells that transform earlier in the process of carcinogenesis compared to other malignant cells in the body (Welch et al., Cell, 150:264-278, 2012). Due to the higher relative abundance of these mutations, they are expected to exhibit greater diagnostic sensitivity for detecting cancer DNA compared to those with lower relative abundance.

[0103] As used herein, “sequencing noise” refers to noise introduced during “run” by sequencing instruments, software, or other human intervention. There are at least two sources of noise in a sequencing pipeline. First, the DNA mixture produced by the input precipitate (DNA or cell precipitate) is a complex mixture of cells, so any useful signal is diluted by DNA that lacks informational content. The second source of noise is due to the specific sequencing technology employed. For example, sequencing noise, or “machine” noise, can originate from ion-pair-base sequencing processes, such as those using the IONTORENT PGM™ platform. For instance, ion detection sequencing, which reads bases based on pH detection, is sensitive to homopolymers and may sometimes read a homopolymer chain as a single base that is too long or too short.

[0104] As used in this article, “sequencing error rate” refers to the proportion of incorrectly sequenced nucleotides. For example, in the context of whole-genome sequencing, a sequencing error rate of approximately 1 per 1000 bases has been reported in the literature (range: error rate per base-call is approximately 0.1–1%; Wu et al., Bioinformatics, 33(15):2322-2329, 2017).

[0105] As used in this paper, the term “sequencing depth” refers to the number of times a sequencing region is covered by sequence reads. For example, an average sequencing depth of 10x means that each nucleotide in the sequencing region is covered by an average of 10 sequence reads. As sequencing depth increases, the chances of detecting cancer-related mutations are expected to increase. However, in reality, the probability of detection does not increase linearly with sequencing depth, as evidenced by the fact that even at a median depth of 42,000x, the fundamental limitation of cfDNA abundance only results in a positive detection rate of approximately 19% for early-stage lung adenocarcinoma (Abbosh et al., Nature, 545(7655): 446-451, 2017).

[0106] As used herein, the term "noise" in its broadest sense refers to any unintended interference (e.g., signals not directly related to real events) that can still be processed or received as real events. Noise is the sum of unwanted or interfering energy introduced into a system from both human and natural sources. Noise distorts signals, thereby degrading or reducing the reliability of the information they carry. This term is the opposite of "signal," which is the function of conveying information about the behavior or properties of a phenomenon (e.g., the probabilistic association between markers (SNV, CNV, indel, SV) and tumors).

[0107] As used herein, the term "signal-to-noise ratio" (SNR) refers to the ability to resolve a true signal from noise in a system. SNR is calculated as the ratio of the desired signal level to the level of noise present in the signal. Phenomena affecting SNR include, for example, detector noise, system noise, and background artifacts. As used herein, the term "detector noise" refers to unwanted interference originating within the detector (i.e., a signal not directly generated by the expected detection energy). Detector noise includes dark current noise and shot noise. Dark current noise in optical detector systems, such as sequencers, can be caused by various thermal radiations from the photodetector. Shot noise in optical systems is a product of the fundamental particle properties (i.e., energy fluctuations in a Poisson distribution) of incident photons as they pass through the photodetector.

[0108] Those skilled in the art use the term "filter" in various ways to indicate discarding or removing unwanted data, retaining desired data, or both. In this disclosure, the term "filter" is primarily used to imply retaining desired data, such as signals.

[0109] The term "base quality" (BQ) score refers to the confidence level of sequencing quality at each nucleobase in a polynucleotide. In some implementations, base quality (BQ) includes variable base quality (VBQ) or average reading base quality (MRBQ), both of which are variations of base quality metrics.

[0110] The term “mapping quality” (MQ) score refers to a confidence estimate of the accuracy of the mapping between markers and the genome.

[0111] The term "readout position" or "position in readout (PIR)" refers to the location on a readout (e.g., a marker) within a nucleotide sequence. As understood in genomics, many sequencing protocols are prone to various types of amplification-induced biases and errors, which can be reduced by implementing filters such as "readout orientation" and "readout position" filters. Readout orientation filters remove variants that are almost exclusively present in forward or reverse readouts. For many sequencing protocols, such variants are most likely the result of amplification-induced errors. Readout position filters are implemented in a similar manner to "readout orientation filters" to eliminate systematic errors, but are also applicable to hybridization-based data. It removes variants with different localizations from the readout carrying the variant site, compared to the expected general position of the readout covering the variant site. This is done by classifying each sequenced nucleotide (or gap) according to the mapping direction of the readout and the location where the nucleotide is found within the readout; each readout is divided into several parts along its length (e.g., 5 parts), and the part number of the nucleotide is recorded. For each sequenced nucleotide, a total of ten categories are given, and a given site will be distributed among these ten categories in the readouts covering that site. If a variant is present at that site, it can be expected that the variant nucleotides follow the same distribution. The read position filter performs tests to measure the significance of the read positions, for example, measuring whether the read position distribution of the variant carrying the read differs from the read position distribution of the total set of reads covering that site.

[0112] As used herein, the term “location attribute” for a marker (e.g., CNV) refers to the spatial location of the marker within a chromosome or gene sequence. For example, the location attribute of a marker can be measured based on whether it is at least 1000 kb, at least 400 kb, at least 100 kb, at least 20 kb, or less than 1 kb, such as 1 kb, from a telomere, centromere, or heterochromatin region of the chromosome. CNVs characterized by chromosomal rearrangement hotspots mapped to subtelomere or centromere-like regions may be undesirable. As used herein, the term “representative” for a marker (e.g., CNV) refers to its association with a phenotype or disease. For example, previous studies have found that CNV calls in immunoglobulin regions do not represent gDNA and tend to depend primarily on the DNA source—e.g., saliva versus blood or lymphoid stem cell lines versus blood (Need et al., 2009; Wang et al., 2007; Sebat et al., 2004).

[0113] As used in this article, the terms "coverage" or "depth" in DNA sequencing refer to the number of reads that include a given nucleotide in the reconstructed sequence. Coverage histograms are commonly used to describe the range and uniformity of sequencing coverage across an entire dataset. They illustrate the overall coverage distribution by showing the number of reference bases mapped across sequencing read coverage at various depths. The mapped "read depth" refers to the total number of bases sequenced and aligned at a given reference base position. Typically, in a sequencing coverage histogram, read depths are binned and displayed on the x-axis, while the total number of reference bases occupying each read depth bin is displayed on the y-axis. These can also be written as a percentage of reference bases.

[0114] As used in this article, “deep coverage” refers to the number of unique readings whose mapping overlaps with specific genome coordinates.

[0115] As used in this article, the term “reading mappability” for CNVs refers to a confidence estimate of the accuracy of mapping of reads associated with that CNV to the genome.

[0116] As used herein, the term "unique read" refers to a read that has unique characteristics, such as a read that appears uniquely in a reference genome. Conversely, "non-unique read" refers to a read that has little or no unique characteristics, such as a read that occurs more than once (i.e., a repetition).

[0117] As used herein, a genomic “target region” or ROI can be any genomic region from which genetic information is desired. The target genomic region can contain a region of a chromosome. The target genomic region can contain an entire chromosome. The chromosome can be a diploid chromosome. For example, in the human genome, a diploid chromosome can be any of chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. In some cases, the chromosome can be an X or Y chromosome. In some cases, the target genomic region includes a portion of a chromosome. The target genomic region can be of any length. The target genomic region may have the following lengths: for example, about 1 to about 10 bases, about 5 to about 50 bases, about 10 to about 100 bases, about 70 to about 300 bases, about 200 bases to about 1000 bases (1 kb), about 700 bases to about 2000 bases, about 1 kb to about 10 kb, about 5 kb to about 50 kb, about 20 kb to about 100 kb, about 50 kb to about 500 kb, about 100 kb to about 2000 kb (2 Mb), about 1 Mb to about 50 Mb, about 10 Mb to about 100 Mb, about 50 Mb to about 300 Mb. For example, the target genomic region can be more than 1 base, more than 10 bases, more than 20 bases, more than 50 bases, more than 100 bases, more than 200 bases, more than 400 bases, more than 600 bases, more than 800 bases, more than 1000 bases (1 kb), more than 1.5 kb, more than 2 kb, more than 3 kb, more than 4 kb, more than 5 kb, more than 10 kb, more than 20 kb, more than 30 kb, more than 40 kb, more than 50 kb, more than 60 kb, more than 70 kb, more than 80 kb, more than 90 kb, more than 100 kb, more than 200 kb, more than 300 kb, more than 400 kb, more than 500 kb, more than 600 kb, more than 700 kb, more than 800 kb, more than 900 kb, more than 1000 kb (1 kb). The target genomic region can contain one or more information loci. Information loci can be polymorphic, such as those containing two or more alleles. In some cases, the two or more alleles include minor alleles.

[0118] As used herein, the term "directionality" regarding reads refers to the direction or manner in which reads are performed. For example, in single-end reads, the sequencer reads fragments only from one end to the other, generating a base-pair sequence. In paired-end reads, it starts with one read, completes the direction with a specified read length, and then begins another round of reads from the opposite end of the fragment. Paired-end reads improve the ability to identify the relative positions of individual reads in the genome, making them more efficient than single-end reads in resolving structural rearrangements such as gene insertions, deletions, or inversions. It can also improve the assembly of repetitive regions. However, paired-end reads are more expensive and time-consuming than single-end reads.

[0119] As used herein, the term “CNV directionality” refers to the direction of copy number change. For example, an increase in copy number (e.g., increase or doubling) is positive, while a decrease (e.g., loss or fragmentation) is negative.

[0120] As used herein, the term "box" refers to a group of DNA sequences grouped together, for example, in a "genome box". In certain cases, a box may include groups of DNA sequences binned based on a "genome box window", which includes grouping DNA sequences using a genome window.

[0121] As used herein, the term “estimate” is used broadly when the level is marked. Thus, the term “estimate” can refer to an actual value (e.g., 1 / mbp), a range of values, a statistical value (e.g., mean, median, etc.), or other estimation methods (e.g., probabilistically).

[0122] As used herein, “substantially” means sufficient for the intended purpose. Therefore, the term “substantially” allows for minute, negligible variations from an absolute or perfect state, size, measurement, result, etc., that would be expected by one of ordinary skill in the art without significantly affecting overall performance. When used for numerical values, parameters, or characteristics that can be expressed as numerical values, “substantially” means within ten percent.

[0123] As used herein, the term “substantially purified” means cfDNA molecules that have been removed from their natural environment, isolated or separated or extracted, and that contain at least 60%, preferably 75%, more preferably 90%, and most preferably 99% free of other components naturally associated with them.

[0124] All publications mentioned herein are incorporated herein by reference and are intended to describe and disclose the apparatus, compositions, formulations, and methods described in the publications and that may be used in connection with this disclosure.

[0125] As used herein, the terms “comprise,” “comprises,” “comprising,” “contain,” “contains,” “containing,” “have,” “having,” “include,” “includes,” and “including,” and variations thereof, are not intended to be restrictive; they are inclusive or open-ended and do not exclude additional, undescribed additives, components, integers, elements, or method steps. For example, a process, method, system, component, kit, or apparatus that comprises a list of features is not limited to those features but may also include other features not expressly listed or inherent to such process, method, system, component, kit, or apparatus.

[0126] Unless otherwise indicated, the practice of this subject matter may employ conventional techniques within the scope of the art and descriptions of organic chemistry, molecular biology (including recombinant techniques), cell biology, and biochemistry.

[0127] method

[0128] This disclosure relates to methods and systems for detecting and / or diagnosing residual tumors by analyzing markers present in cell-free DNA (cfDNA). The detection can be used alone or in combination with existing technologies to determine the presence of residual tumors, predict the likelihood of having such disease, and develop therapeutic or preventative interventions for such disease.

[0129] In some embodiments, the methods of this disclosure are performed on samples obtained from a subject. Preferably, the samples include blood (including whole blood), plasma, serum, hemolysin products, lymph, synovial fluid, cerebrospinal fluid, urine, cerebrospinal fluid, feces, sputum, mucus, amniotic fluid, tears, cyst fluid, sweat gland secretions, bile, milk, tears, saliva, or earwax. Various methods can be used to process the samples to remove specific cells, such as centrifugation, affinity chromatography (e.g., immunoadsorption), immunoselection, and filtration. Thus, in one instance, the sample may contain a specific cell type or mixture of cell types (e.g., T cells purified from whole blood) isolated directly from the subject or purified from samples obtained from the subject. In one instance, the biological sample is peripheral blood mononuclear cells (PBMCs). In other instances, the sample may be selected from the group consisting of: B cells, dendritic cells, granulocytes, innate lymphoid cells (ILCs), megakaryocytes, monocytes / macrophages, natural killer (NK) cells, platelets, erythrocytes (RBCs), T cells, and thymocytes. In some embodiments, the sample may contain skin cells, hair follicle cells, sperm, etc.

[0130] In Figure 1 and Figure 8 The diagram provides a representative, non-limiting illustration of diagnostic methods.

[0131] Workflow

[0132] Figure 1A This is a flowchart illustrating method 100 for detecting residual disease according to various embodiments of the present disclosure, said residual disease being, for example, a tumor disease invented after surgery or treatment (e.g., chemotherapy, immunotherapy, targeted therapy, radiotherapy). Method 100 is merely illustrative, and embodiments may use variations of method 100. Method 100 may include the steps of: receiving a marker profile; filtering noise associated with the markers based on multiple features; removing artificial noise markers from the profile to generate subject-specific markers, then using said subject-specific markers to assess a tumor score (eTF), and then using the tumor score to diagnose residual disease. It should be noted that TF refers to the fraction of tumor DNA (ctDNA) in total plasma DNA (cfDNA). Therefore, in this disclosure and elsewhere, the term "ctDNA abundance" may be used interchangeably with the term tumor score.

[0133] exist Figure 1A In step 110 of method 100, a subject-specific whole-genome summary of readings associated with multiple genetic markers (e.g., SNV, CNV, SV, indel) in a biological sample (tumor sample and optionally normal sample) is received from the subject. In some embodiments, the summary of genetic markers is received as a variant call format (VCF) file. As understood in the art, VCF files are used in bioinformatics to store gene sequence variations. The VCF format was developed with the advent of large-scale genotyping and DNA sequencing projects (e.g., the 1000 Genomes Project). Alternatively, the summary can be provided in a generic feature format (GFF) containing all genetic data. Typically, GFFs provide redundant functionality because they are shared between genomes. In contrast, with VCF, only variations need to be stored along with a reference genome. In some embodiments, the subject's sample is sequenced (e.g., using whole-genome sequencing (WGS)), and the sequence file is processed, for example, using a tool such as genomic VCF (gVCF).

[0134] exist Figure 1A In step 120 of method 100, a subject-specific whole-genome profile of genetic markers in a second sample (e.g., plasma or blood) of the subject is detected to generate a representative of tumor-associated whole-genome genetic markers in the patient sample (e.g., plasma or blood sample).

[0135] exist Figure 1AIn step 130 of method 100, the noise probability (P) of each marker is analyzed. N For example, in the case of a tag being SNV or indel, P N Analysis can be performed as a function of: 1) the MQ of the SNV / indel; 2) the fragment length of the reads containing the SNV / indel; 3) the consistency test in the repeat family of reads containing the SNV or Indel; and / or 4) the BQ of the SNV / indel. Similarly, in the case of a marker being a CNV or SV, the probability that the marker is noise-related can be analyzed based on: (1) its position relative to the centromere; 2) the MQ of the read group containing the CNV / SV; and / or 3) the representation of the CNV window in the cfDNA data of the artifact reads. The noise removal step 130 may include performing an optimal receiver operating characteristic (ROC) curve based on the combined base quality (BQ) and mapping quality (MQ) scores, where the curve includes the probability classification of the genetic marker in the profile. Typically, the combined BQMQ scores are provided in the form of a matrix (x, y), where x is the BQ score and y is the MQ score. In exemplary embodiments, a joint BQMQ score between 10 and 50 (for each parameter) is typically used, for example, BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), (30, 40). In some embodiments, the classification of markers includes a measure of the area under the ROC curve (AUC), which typically represents the probability that a candidate marker randomly selected from potential markers will show a value higher than that of a randomly extracted control marker. For completely non-informative markers, the ROC curve will approach an ascending diagonal (called the “randomness diagonal” or “randomness line”), and the AUC will tend to 0.5, the expected probability of classification due to chance alone. Conversely, in the case of perfect classification, the ROC curve will reach the point of highest theoretical accuracy (100% sensitivity and specificity), and the AUC will tend to 1, the highest probability value. Figure 3B The document provides a representative ROC model. The pre-filtration error model and post-filtration effect of the base quality filter are shown below. Figure 3A As shown. Figure 3C The application of base quality (BQ) and mapping quality filter (MQ) showed that sequencing errors were suppressed by about seven times.

[0136] exist Figure 1AIn step 140 of method 100, an estimated tumor score (eTF) of the biological sample is calculated based on one or more integrated mathematical models. Depending on the marker (e.g., SNV / indel vs. CNV / SV), the mathematical model integrates multiple process quality metrics and patient-specific attributes to estimate the tumor score (TF). Recognizing the fundamental differences in frequency between SNV / indel and CNV / SV and their association with traits (e.g., cancer), the systems and methods of this disclosure involve using marker-specific mathematical algorithms to estimate the tumor score. In each case, the mathematical inference model outputs an estimated score of tumor DNA in the biological sample (e.g., plasma) based on the number / frequency of markers, estimated noise, readouts, mutational burden, and / or coverage or depth.

[0137] In some embodiments, the method of this disclosure includes estimating TF based on the detection of multiple SNV / indel markers. Herein, the estimated TF (eTF[SNV]) is calculated by integrating a process quality metric including estimated genomic coverage and sequencing noise with a patient-specific parameter including mutational burden (N). Preferably, the method includes calculating an estimated tumor fraction (eTF) of SNV / indel markers, where eTF[SNV] = 1 - [1 - (ME(σ)]] R ) / N]^(1 / cov), where M is the number of tumor-specific profile detections in the patient sample, σ is the noise measurement estimated based on experience, R is the total number of unique readings in the target region (ROI), N is the tumor mutation burden, and cov is the average number of unique readings at each site in the ROI.

[0138] In some embodiments, the method of this disclosure includes estimating TF based on the detection of multiple CNV / SV markers. Herein, the estimated TF (eTF[CNV]) is calculated by integrating the directional coverage depth with respect to a skewness consistent with the tumor CNV / SV orientation, where copy number amplification is positively skewed and copy number loss is negatively skewed. Preferably, the method includes calculating an estimated tumor score (eTF) for the CNV markers, where eTF[CNV] = (sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth value in a genomic window indexed by {i} representing plasma depth coverage, T is the median depth value in a genomic window indexed by {i} representing tumor depth coverage, and N is the median depth value in a genomic window indexed by {i} representing normal depth coverage.

[0139] exist Figure 1AIn step 150 of method 100, residual disease in the subject is diagnosed based on the eTF (calculated in step 140) and an empirical threshold calculated by a background noise model. In some embodiments, the detection threshold includes an empirically measured baseline noise TF estimate from a healthy sample. In such embodiments, any eTF above the threshold (e.g., at least 2 standard deviations (FPR < 2.5%) of the noise TF distribution; preferably greater than 3 or 5 standard deviations) is defined as a positive detection.

[0140] like Figure 1B The exemplary workflow 100 shown further provides, according to various embodiments, a method for detecting residual disease in subjects in need. As in Figure 1B The workflow provided in step 110 of method 100 may include receiving a whole-genome summary of first subject-specific readings associated with genetic markers from a first biological sample from a subject. The first biological sample may include a baseline sample. Each of the first reading summaries may contain readings of single base pair lengths. The baseline sample may include a tumor sample or a plasma sample. The first biological sample may also include a normal cell sample.

[0141] As in Figure 1B The workflow provided in step 120 of method 100 may include filtering artificial loci from a first reading profile, wherein the filtering includes removing repetitive loci generated in a reference healthy sample cohort from a first genetic marker profile. Alternatively or additionally, the filtering may include identifying germline mutations in peripheral blood mononuclear cells of a normal cell sample and removing said germline mutations from the first genetic marker profile.

[0142] like Figure 1B The workflow provided in step 130 of method 100 may further include detecting readings in a second subject-specific whole genome profile of genetic markers in a second biological sample from the subject to generate a tumor-associated whole genome representative of the genetic markers in the second sample.

[0143] like Figure 1BThe workflow provided in step 140 of method 100 may include using at least one error suppression scheme to filter noise from first and second read genome profiles to produce a first filtered read set for the first read genome profile and a second filtered read set for the second read genome profile. At least one error suppression scheme may include calculating the probability that any single nucleotide variant in the first and second profiles is an artificial mutation and removing said mutation. The probability may be calculated as a function of features selected from: mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof. Alternatively or in combination, at least one error suppression scheme may include removing artificial mutations using a dissimilarity test between independent repeats of the same DNA fragment generated by polymerase chain reaction or sequencing processing. Alternatively, or in combination with a dissimilarity test, removing artificial mutations may include a repeat consistency test, wherein artificial mutations are identified and removed when consistency is lacking in the majority of a given repeat family.

[0144] like Figure 1B The workflow provided in step 150 of method 100 may include calculating the estimated tumor fraction (eTF) of the first and second biological samples by applying a background noise model to one or more integrated mathematical models and using first and second filtered reading sets.

[0145] As in Figure 1B The workflow provided in step 160 of method 100 may include detecting residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.

[0146] like Figure 1C The exemplary workflow 100 shown, further provided according to various embodiments, offers a method for detecting residual disease in subjects in need. (As in...) Figure 1C The workflow provided in step 110 of method 100 may include receiving a genome-wide summary of first subject-specific readings associated with genetic markers from a first biological sample from a subject. The first biological sample may include a baseline sample. The first reading summary may each include copy number variations (CNVs). The baseline sample may include a tumor sample or a plasma sample.

[0147] As in Figure 1C The workflow provided in step 120 of method 100 may include receiving a whole-genome summary of second subject-specific readings associated with genetic markers from a second biological sample of a subject. The second biological sample may include a peripheral blood mononuclear cell (PBMC) sample. The second summary of the genetic markers may each include copy number variations (CNVs).

[0148] As in Figure 1C The workflow provided in step 130 of method 100 may include filtering artificial sites from first and second reading profiles, wherein the filtering includes removing duplicate sites generated in a reference healthy sample cohort from the first and second reading profiles. Alternatively or in combination, the filtering may include identifying CNVs shared between the first and second profiles as germline mutations and removing said mutations from the first and second reading profiles.

[0149] like Figure 1C The workflow provided in step 140 of method 100 may further include readings in a third subject-specific whole genome profile of genetic markers from a third biological sample of the subject to generate a tumor-associated whole genome representative of the genetic markers in the third sample.

[0150] As in Figure 1C The workflow provided in step 150 of method 100 may include normalizing each of the first, second, and third read summaries to produce a first filtered read set for the first read genome summary, a second filtered read set for the second read genome summary, and a third filtered read set for the third read genome summary.

[0151] As in Figure 1C The workflow provided in step 160 of method 100 may include calculating an estimated tumor fraction (eTF) of a third biological sample by applying a background noise model to one or more integrated mathematical models, using a third filtered set of readings, wherein the one or more models generate a first eTF using a first filtered set of readings, and / or the one or more models generate a second eTF using a second filtered set of readings.

[0152] As in Figure 1C The workflow provided in step 170 of method 100 may include detecting residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold.

[0153] plan

[0154] Figure 1D and Figure 1E An illustrative workflow for practicing the methods of this disclosure is shown. Figure 1D This outlines the workflow typically used when the destination tag contains an SNV / indel; Figure 1EThe workflows commonly used when the target marker includes CNV / CV are outlined. It should be noted that although separate workflows are provided for illustrative purposes, they are not necessarily to be performed separately to achieve the methods of this disclosure. For example, certain features / elements of the workflows may be combined to generate an output (e.g., an estimated tumor score based on a combination of SNV / indel and CNV / SV) that is correlated with the target outcome (e.g., whether a subject with MRD responds to chemotherapy).

[0155] like Figure 1D As shown, SNV / indel-based MRD detection typically utilizes the following steps: receiving data; generating patient-specific signatures for SNV / indels; removing / filtering artificial sites; detecting reads / sites in subsequent samples; suppressing errors using specific algorithms (including machine learning); correcting reads; detecting sites that provide an estimated tumor score; and optionally, performing orthogonal integral analysis (e.g., analyzing fragment size shifts) on minor features in the genomic data to improve the sensitivity, specificity, and / or reliability of the detection.

[0156] exist Figure 1D In the first step, genetic data from baseline samples (typically tumor samples, but may also include pre-treatment plasma, alone or with tumor samples) and normal samples (typically PBMCs, but may also include adjacent normal tissue or buccal swabs) are received to generate patient-specific marker signatures (e.g., including SNVs / indels). Next, a reference list of somatic mutations is invoked from the baseline samples by filtering artificial sites. Germline mutations are removed from the samples here. Additionally, somatic mutation invocation is performed independently using multiple invocations (e.g., MUTECT, STRELKA), and the intersection of the invocations is used to generate a high-confidence mutation list. Recurring artificial sites are generated sequentially or in parallel from a cohort of healthy plasma samples (a blacklist or mask of normal plasma groups (PONs)) and removed from patient-detected mutations to eliminate common sequencing or alignment artifacts. The filtered high-confidence patient-specific mutation dataset is then used to detect mutations in subsequent plasma samples. Typically, subsequent plasma is obtained after surgery, during or after treatment (e.g., chemotherapy), or during follow-up (e.g., checking for recurrence or relapse).

[0157] Next, a highly sensitive method capable of detecting individual mutant fragments was employed. This step involved one or more error suppression steps. In the first error suppression step, a filtering scheme was used to analyze individual reads and quantify the probability that the reads represented artificial mutations. Representative methods included a multidimensional classification framework using a support vector machine (SVM) with a linear kernel. This classification engine was trained on germline SNPs and compared with low variant allele fraction (VAF) sequencing artifacts in normal PBMC samples. The classification decision boundary was defined in a multidimensional space including variant base quality (VBQ), mapping quality (MQ), read position (PIR), and mean read base quality (MRBQ). To evaluate the classification scheme, the SVM validation metric after 10-fold cross-validation was compared with random forest under the same protocol. SVM classification showed high classification performance, moderately outperforming the random forest model. In all patients, the mean sensitivity of SVM was 90.7%, and the specificity was 83.9% (N = 10 samples, F1 = 87.7%, PPV = 84.9%).

[0158] In the second error suppression step, artificial mutations generated by PCR or sequencing are corrected by comparison of independent repeats of the same original DNA fragment. In cfDNA samples, paired 150bp sequencing is typically used, which, given the short size (~165bp) of typical cfDNA fragments, results in overlapping paired reads (R1 and R2 sequence overlap). Therefore, any inconsistencies between R1 and R2 pairs are considered potential sequencing artifacts, which are corrected back to the corresponding reference genome. Furthermore, recognizing the potential for any DNA molecule that has been replicated multiple times in sequencing and PCR to establish independent replication, repeat families are identified by 5' and 3' similarity and alignment position. Each repeat family is then used to check for consistency of specific mutations in the independent repeats, thereby correcting for artificial mutations that do not show consistency in most repeat families.

[0159] Next, the fraction of patient-specific mutations appearing in the plasma is estimated. This parameter follows a binomial distribution across N independent Bernoulli trials, where N is the patient mutation burden. Each such trial comprises multiple rounds of random sampling, which depend on local coverage, where the probability of sampling a mutated fragment in each round is the tumor fraction. Therefore, a mathematical relationship exists between coverage, mutation burden, the number of detected mutations, and the tumor fraction, corresponding to the following equation: , where M represents the number of mutations detected in follow-up plasma samples, N represents the mutation load in patient-specific mutation patterns, TF represents the tumor score, cov represents the local coverage of the patient's mutation site, and μ represents the noise rate corresponding to a specific patient's mutation site. This relationship allows the calculation of a patient's tumor score from the mutation detection rate, even in extremely low allele scores where the mutation allele score itself does not provide information (primarily representing random sampling between 0 and 1 within the effective coverage range—only one supporting reading).

[0160] To address noise variability among patients with different mutation patterns, patient-specific mutation signatures are used to calculate the expected noise distribution in the healthy plasma sample (normal group, PON) cohort. The same procedures described above are primarily performed to detect patient-specific patterns in healthy samples (PON) or other patients (cross-patient analysis). These detections represent a background noise model for which we calculate the mean and standard deviation (μ, σ) of the artificial mutation detection rate. Confidence-based tumor detection and tumor score estimation are achieved if the detected tumor score in a patient is higher than the artificial tumor score, which corresponds to an error rate of 1.5 * σ above the mean.

[0161] Next, optionally, the workflow may include orthogonal integrals calculated based on fragment size offsets. Here, to make the prediction / diagnostic method more robust, accurate, and / or sensitive, read-based features (e.g., DNA fragment size offsets) can be orthogonally integrated into the model. The importance of the orthogonal features (in determining the MRD) can be determined using statistical methods or probabilistic mixture models (e.g., Gaussian models). See Example 3A for a detailed overview.

[0162] In the exemplary method, high-confidence tumor-specific detections in plasma samples are aggregated and converted into an estimate of tumor DNA fraction (TF) based on a probabilistic dilution model. The entire detection protocol (detection, false suppression, and tumor fraction estimation) is also performed on the healthy plasma sample group (PON) as follows: using a patient-specific mutation profile, the distribution of noisy TF values ​​in healthy samples is calculated using the same signature. Subsequently, using a statistical significance framework (z-score) that ensures a low false positive rate (high specificity), tumor detection and evaluation are performed only on samples showing tumor fractions significantly higher than the PON noisy TF values. The presence of tumor DNA in plasma mutation detection is orthogonally confirmed using statistical methods (significance tests or GMM), which quantifies the intra-patient fragment size shift between the tumor-specific detection list and other random mutation detection lists.

[0163] Alternatively, or in conjunction with the above workflow, this disclosure also relates to the use of CNV / SV markers to detect residual disease (or monitor treatment). Figure 1EAs shown, MRD detection based on CNV / SV markers typically utilizes the following steps: receiving data; generating baseline sample-specific and / or normal sample-specific signatures for CNV / SV; removing germline CNV events; filtering artificial windows; detecting window-based median depth coverage in subsequent samples; normalizing using, for example, guanine-cytosine (GC) normalization and / or z-score normalization; detecting tumor CNV signals, which provide an estimate of the tumor score; and optionally, performing orthogonal integral analysis (e.g., fragment size shift analysis) on minor features in the genomic data to improve the sensitivity, specificity, and / or reliability of the detection.

[0164] exist Figure 1E In the first step, genetic data from baseline samples (typically tumor samples, but may also include pre-treatment plasma, alone or with tumor samples) and normal samples (typically PBMCs, but may also include adjacent normal tissue or buccal swabs) are received to generate tumor-specific marker signatures and normal marker signatures (e.g., signatures containing CNV / SV). Next, tumor copy number variants (T_CNVs) are invoked using the baseline against the normal group (PON). PBMC copy number variants (P_CNVs) are invoked using PBMC samples against the normal group (PON). Shared copy number variant events are considered germline. Tumor somatic events (sT_CNVs, detected only in tumor tissue) and PBMC somatic events (sP_CNVs, detected only in PBMC tissue) are used for tumor score detection and assessment.

[0165] Next, germline variations (e.g., CNV / SV events) are removed from the CNV / SV reference list to generate baseline sCNV / SV and / or normal-sCNV / SV. Similarly, windows with low mappability and / or coverage are filtered. Recurring artificial sites are generated sequentially or in parallel from a cohort of healthy plasma samples (a blacklist or mask of normal plasma groups (PONs)) and removed from the windows to filter artificial windows. The filtered high-confidence reference CNV / SV segments are used to detect mutations in subsequent plasma samples. Typically, subsequent plasma is obtained after surgery, during or after treatment (e.g., chemotherapy), or during follow-up (e.g., to check for recurrence or relapse).

[0166] Reproducible artificial CNV sites were generated on a cohort of healthy plasma samples (the blacklist of the normal PON group) and removed from mutations detected in patients to remove common sequencing or alignment artifacts, such as centromeres and repetitive regions.

[0167] The target region of interest (ROI) containing all genomic segments of sT_CNV and sP_CNV is then binned into windows (500 bp or more). Depth coverage (reading counts) for each window is estimated from subsequent plasma samples (postoperative, during treatment, and during follow-up for relapse). The median depth coverage for each window is calculated and then divided by the mean sample coverage.

[0168] Next, the depth coverage value was normalized by performing two LOESS regression curve fittings on the bin-by-bin GC score and mappability score to correct for GC content and mappability bias.

[0169] Further batch effect correction can be performed using robust z-score normalization, which is applied to each sample individually. In short, the median and median absolute deviation (MAD) are calculated based on the neutral region for each sample, and then all CNV bins are normalized by (B(i)-Median) / MAD.

[0170] Depth coverage skewness and fragment size center of mass (COM) skewness were calculated for each bin compared to healthy plasma samples from the normal group (PON). In this study, samples with low tumor fractions showed sparse depth coverage skewness, biased by the directionality of CNV segments; amplified segments showed a bias toward positive depth coverage skewness, while deletions showed a bias toward negative depth coverage skewness. On the other hand, neutral regions showed random skewness with no preferred directionality; therefore, multiplying the differential (plasma-PON) depth coverage skewness by the directionality of the CNV segments (amplification multiplied by +1, deletion multiplied by -1) summed the CNV signal across the entire genome, while noise in neutral regions was eliminated due to random directionality.

[0171] This step is based on the following equation Complete, where M is the number of windows for coverage ROI. P(i) and N(i) are the depth coverage values ​​of plasma samples and PON in window I, respectively. The symbol (T(i)-N(i)) represents the direction of the tumor CNV segment (amplification multiplied by +1, deletion multiplied by -1).

[0172] The tumor score can then be calculated by examining the linear dilution ratio between the cumulative signal detected in the plasma sample and the cumulative signal detected in the tumor. This step is accomplished by the following equation:

[0173]

[0174] Where N(i), P(i), and T(i) represent the patient's PBMC, plasma, and tumor depth coverage in window I, respectively.

[0175] To address noise variability among patients with different CNV patterns, patient-specific CNV signatures were used to calculate the expected noise distribution in the healthy plasma sample (normal group, PON) cohort. The same procedures described above in the SNV labeling analysis can be performed to detect patient-specific patterns in healthy plasma samples (PON) or other patients (cross-patient analysis). These detections represent a background noise model, for which we calculated the mean and standard deviation (μ, σ) of the artificial mutation detection rate. Confidence-based tumor detection and tumor score estimation are achieved if the tumor score detected in the patient is higher than the artificial tumor score, which corresponds to an error rate of 1.5 * σ above the mean.

[0176] It is also possible to infer tumor scores from the directional whole-genome depth coverage skewness in sP_CNVs. In this paper, due to the increased tumor DNA score (since tumor DNA does not include this CNV event), PBMC-specific CNV events are expected to have reduced signals. Therefore, a negative correlation is expected between tumor scores and sP_CNV detection signals in plasma. Therefore, multiplying the differential (PBMC-plasma) depth coverage skewness by the directionality of the PBMC CNV segment (amplification multiplied by +1, deletion multiplied by -1) will sum the PBMC CNV signals across the entire genome. Figure 11A ).

[0177] The tumor score can then be calculated by examining the proportion of CNV signal loss on the PBMC, for example, using the following equation:

[0178]

[0179] In MRD estimation using SNV / indel markers, secondary features can be orthogonally integrated into the final calculation. Here, to improve the robustness, accuracy, and / or sensitivity / specificity of the detection method, read-based features (e.g., DNA fragment size offset) can be orthogonally integrated into the model. A generalized linear model (GLM) can be used to determine the importance of orthogonal features (in determining the MRD) to orthogonally determine tumor scores based on the relationship between CNV depth coverage and fragment size offset. See Example 3B for a detailed overview.

[0180] It should be understood that, with some modifications, the workflow disclosed herein can also be widely used to detect residual disease during or after chemotherapy, immunotherapy, targeted therapy, or combinations thereof, and / or in the process of monitoring the effects of such therapies.

[0181] The exemplary method is partly based on the understanding that whole-genome CNV signals accumulate in plasma samples only if the coverage skewness in plasma follows the same directionality (amplification and deletion) as copy number variations in baseline tissues (e.g., tumors). Therefore, the tumor DNA ratio can be calculated based on the signal gain in plasma samples for patient-specific CNV events, for example, by dividing the accumulated CNV signal in plasma by the linear dilution ratio between accumulated CNVs in the tumor. A similar mixture dilution model can be used to orthogonally estimate tumor fractions based on signal loss only for patient PBMC-specific CNV events (hematopoietic somatic CNV events). The entire CNV detection protocol is also performed on a healthy plasma sample pool (PON) by calculating the distribution of noise TF values ​​in healthy samples using the same CNV signature, using a patient-specific copy number variation profile. Subsequently, using a statistical significance framework (z-score) that ensures a low false positive rate (high specificity), tumor detection and evaluation are performed only on samples showing tumor fractions significantly higher than the PON noise TF values. The presence of tumor DNA in plasma can be orthogonally confirmed by examining the relationship (negative correlation) between CNV log2 values ​​and fragment size centroid (COM) values ​​in patient-specific CNV segments, and then converting this relationship into an orthogonal estimate of CNV-based TF estimation based on a generalized linear model (GLM).

[0182] Machine Learning

[0183] Not limited to a single implementation, and for illustrative purposes only, machine learning (ML) algorithms are integrated into existing methods in single steps or combinations thereof, according to various implementations herein. ML can be incorporated to optimize results from algorithms (e.g., neural networks, ML algorithms, etc.) by using the input training dataset, cross-referencing the output to a known answer, backpropagation, and adjusting weighting factors and parameters associated with a given ML algorithm in a loop of repetition to achieve a threshold quality of the data output. In subsequent steps, such as using a probabilistic model like logistic regression (e.g., in combination or alternatively optimized or trained), the model's predictive power on a test dataset can be validated. Optionally, resampling can be performed to obtain a fair assessment of the model's possible future performance. The consistency probability derived from characteristics of the ROC curve (e.g., the area under the curve) (also known as the c-index) or statistical tests (e.g., the Wilcoxon-Mann-Whitney test) can provide a good generalization measure for simple predictive discrimination.

[0184] Preferably, based on one or more quality filters or reading characteristics, the ML algorithm adaptively and / or systematically filters sequencing noise associated with each reading in the profile. In some embodiments, the ML algorithm performs a base quality (BQ) filter (more specifically, variable base quality (VBQ) or average reading base quality (MRBQ)) for filtering noise. In some embodiments, the ML algorithm performs a mapped quality (MQ) filter for filtering noise. In some embodiments, the ML algorithm performs a reading position (PIR) filter for filtering noise. In some embodiments, the ML algorithm performs a combination of filters.

[0185] In some embodiments, the machine learning (ML) methods used in the systems and / or methods of this disclosure include deep convolutional neural networks (CNNs), recurrent neural networks (RNNs), random forests (RFs), support vector machines (SVMs), discriminant analysis, nearest neighbor analysis (KNNs), ensemble classifiers, or combinations thereof, preferably support vector machines (SVMs). In some embodiments, the ML has been trained to distinguish between sequencing reads altered by cancer and reads altered due to sequencing or PCR errors. In some embodiments, the ML has been trained on a large whole-genome sequencing (WGS) cancer dataset containing billions of reads from tumor mutations and normal sequencing errors. In some embodiments, the ML is capable of (a) identifying sequencing or PCR artifacts with high accuracy, and (b) integrating sequence background and read-specific features.

[0186] This disclosure further relates to systems and procedures that utilize ML (e.g., engines) to adaptively and / or systematically filter sequencing noise. This disclosure also relates to a computer-readable storage medium containing a procedure for detecting tumor markers, including somatic mutations in genomic reads, the procedure utilizing ML, such as support vector machines (SVM).

[0187] As is known in the art, Convolutional Neural Networks (CNNs) typically accomplish a high-level form of processing and classification / detection through the following steps: first, they look for low-level features (such as, for example, repeating sequences in readings), and then advance through a series of convolutional layers to more abstract concepts (e.g., unique to the type of reading being classified). CNNs achieve this by passing the data through a series of convolutional, non-linear, pooling (or downsampling, as described below), and fully connected layers, and obtaining an output. Again, the output can be a single category or the probability of classifying an object in the optimal description or detection data.

[0188] Regarding the layers in a CNN, the first layer is typically a convolutional layer (conv). This first layer processes a representative array of readings using a set of parameters. Instead of treating the data as a whole, the CNN analyzes a collection of subsets of data using filters (or neurons or kernels). These subsets will include points in the array that are focal points and those surrounding them. For example, a filter might be able to examine a series of 5x5 regions (or areas) within a 32x32 representative. These regions can be called receptive fields. Since filters typically have the same depth as the input, a representative of size 32x32x3 will have filters of the same depth (e.g., 5x5x3). The actual steps for convolution using the exemplary dimensions above would involve sliding the filter along the input data, multiplying the filter values ​​by the original representative values ​​of the data to compute element-wise multiplication, and summing these values ​​to arrive at a single number for the region of the representative being examined.

[0189] After this convolutional step, using a 5 x 5 x 3 filter yields an activation map (or filter map) of size 28 x 28 x 1. For each additional layer used, it's better to preserve the spatial dimensions so that using two filters will result in an activation map of 28 x 28 x 2. Each filter typically has a unique feature it represents, and these features together represent the feature identifiers needed for the final data output. When these filters are combined, the CNN is allowed to process the data input to detect those features present on each representation. Thus, if the filters are used as curve detectors, the convolution of the filters along the data input will generate an array of numbers in the activation map that corresponds to the high probabilities of the curve (product of the high summation elements in sequence), the low probabilities of the curve (product of the low summation elements in sequence), or zero values ​​(where the input volume at some point does not provide anything for the activation curve detector filters). In this way, the more filters (also called channels) in the Conv, the more depth (or data) is provided on the activation map, and therefore more information about the input that will lead to a more accurate output.

[0190] The accuracy of a CNN is commensurate with the processing time and power required to produce the results. In other words, the more filters (or channels) used, the more time and processing power are required to perform Conv. Therefore, the selection and number of filters (or channels) should be carefully chosen to meet the requirements of the CNN method, producing the most accurate output possible while taking into account available time and power.

[0191] To further enable CNNs to detect more complex features, additional Convs can be added to analyze the output from previous Convs (e.g., activation maps). For example, if the first Conv looks for basic features (e.g., curves or edges), the second Conv looks for more complex features (e.g., shapes), which could be combinations of individual features detected in earlier Conv layers. By providing a series of Convs, CNNs are able to detect increasingly higher levels of features, eventually reaching a probability of detecting a specific desired object. Furthermore, since Convs are stacked on top of each other, analyzing the output of previous activation maps, each Conv in the stack naturally analyzes an increasingly larger receptive field through scaling down at each Conv level, allowing the CNN to respond to regions representing progressively larger representation spaces when detecting target objects.

[0192] A CNN architecture typically consists of a set of processing blocks, including at least one block for convolving the input volume (data) and at least one block for deconvolution (or transpose convolution). Additionally, each processing block may include at least one pooling block and an unpooling block. Pooling blocks can be used to scale down the data resolution to produce an output usable by Conv. This provides computational efficiency (effective time and capability), which in turn can improve the practical performance of the CNN. These pooling or subsampling blocks keep the filters small and computationally efficient, and they can coarsen the output (causing loss of spatial information in the receptive field), thereby reducing the size of the input by a specific factor.

[0193] Unpooling blocks can be used to reconstruct these coarse outputs to produce an output volume with the same size as the input volume. Unpooling blocks can be viewed as the inverse operation of convolutional blocks, returning the activation outputs to the original input volume size. However, the unpooling process typically just amplifies the coarse outputs into a sparse activation map. To avoid this, deconvolutional blocks densify this sparse activation map to generate an enlarged and denser activation map. Ultimately, after any further necessary processing, the final output volume will have a size and density much closer to the input volume. Instead of reducing multiple array points in the receptive field to a single digit, deconvolutional blocks associate a single activation output point with multiple outputs to amplify and densify the resulting activation output.

[0194] It should be noted that while pooling blocks can be used to scale down data and depooling blocks can be used to scale up these scaled activation maps, convolution and deconvolution blocks can be constructed to perform convolution / deconvolution and scale down / scale up without separate pooling and depooling blocks.

[0195] Depending on the target object being detected in the data input, pooling blocks and unpooling processes can have drawbacks. Since pooling typically scales down the data by looking at sub-data windows that do not overlap, spatial information is significantly lost during scaling.

[0196] The processing block may include other layers wrapped with convolutional or deconvolutional layers. These layers may include, for example, rectified linear unit (ReLU) layers or exponential linear unit (ELU) layers, which are activation functions that examine the Conv output in their processing block. The ReLU or ELU layer acts as a gating function, advancing only those values ​​of positive detections that correspond to the target features specific to Conv.

[0197] Given a basic architecture, the training process for a CNN is then prepared to hone its accuracy in data classification / detection (of the target object). This involves a process called backpropagation (backprop), which uses the training dataset or sample data used to train the CNN to update its parameters to achieve optimal (or threshold) accuracy. Backpropagation involves a series of repeated steps (training iterations) that, depending on the backpropagation parameters, will train the CNN slowly or quickly. Backpropagation steps typically include forward propagation, loss function, backpropagation, and parameter (weight) updates based on a given learning rate. Forward propagation involves passing the training data through the CNN. The loss function is a measure of error in the output. Backpropagation determines the factors influencing the loss function. Weight updates involve updating the parameters of the filters to move the CNN to an optimal state. The learning rate determines how much the weights are updated in each iteration to reach the optimal state. If the learning rate is too low, training may take a very long time and involve too much processing power. If the learning rate is too fast, each weight update may be too large to accurately achieve the given optimal value or threshold.

[0198] Backpropagation can complicate training, leading to a lower learning rate and more specific, carefully determined initial parameters at the start of training. One contributing factor is that changes to the Conv parameters amplify the network's depth as weight updates occur at the end of each iteration. For example, if a CNN has multiple Convs, as mentioned above, that allow for higher-level feature analysis, the parameter updates for the first Conv will be doubled at each subsequent Conv. The net result depends on the depth of the given CNN, where even a small change to the parameters can have a significant impact. This phenomenon is called internal covariate shift.

[0199] Typically, the CNNs disclosed herein are capable of adaptively and / or systematically filtering sequencing noise. In some implementations, the CNN architecture is designed based on the inventors' understanding that the trinucleotide background contains different features involved in mutagenesis. Therefore, the CNN uses a receptive field of size 3 to convolve over all features (columns) at a given location. After two consecutive convolutional layers, downsampling is performed using max pooling with a receptive field of 2 and a stride of 2, forcing the model in the engine to retain only the most important features in a smaller spatial region. The resulting architecture maintains spatial invariance when convolving over the trinucleotide window and captures a “quality map” by collapsing the readout fragments into 25 segments (each segment representing a region of approximately 8 nucleotides). Final classification is performed by directly applying the output of the last convolutional layer to a fully connected sigmoid layer. The CNN employs a simple logistic regression layer instead of a multilayer perceptron or global average pooling to preserve features relevant to location within the genome readout.

[0200] To train the engine, multiple lung cancer patients and their matched systematic error profiles are first sampled. The goal of the training exercises is to use a training protocol that allows for the detection of true somatic mutations with high sensitivity while also rejecting candidate mutations caused by systematic errors. Mixtures of samples can be used during training, such as a mixture of whole tumor samples and healthy tissue samples from subjects, for example, those with or suspected of having cancer.

[0201] Upstream steps:

[0202] Receive genetic data

[0203] In some implementations, genetic data are received in situ from the subject's biological samples (e.g., tumor samples or normal cell samples containing PBMCs). This is primarily accomplished via sequencing. In some implementations, samples can be purified using conventional methods to obtain cell subpopulations. For example, PBMCs can be purified from whole blood using various known Ficoll-based centrifugation methods (e.g., Ficoll-Hypaque density gradient centrifugation). Other cells, such as T cells, can also be purified by selecting appropriate phenotypes using techniques such as immunomagnetic cell sorting (e.g., DYNABEADS, Invitrogen, Carlsbad, CA, USA). For example, T cells can be purified using a two-step selection process that first removes CD8+ cells and then selects CD4+ cells. Cell population purity can be confirmed by evaluating appropriate markers (e.g., CD19-FITC, CD3-PE, CD8-PerCP, CD11c-PE Cy7, CD4-APC, and CD14-APC Cy7) using commercially available antibodies (e.g., BD Biosciences).

[0204] After sample preparation, DNA is extracted from the sample for labeling analysis. In one instance, the DNA is genomic DNA. Various methods for isolating DNA, particularly genomic DNA, are known to those skilled in the art. Typically, known methods involve disrupting and lysing the starting material, then removing proteins and other contaminants, and finally recovering the DNA. For example, techniques involving alcohol precipitation; extraction with organic phenols / chloroform and salting out have been used for many years to extract and isolate DNA. An example of DNA isolation is illustrated below (e.g., the Qiagen ALL-PREP kit). However, there are various other commercially available kits for genomic DNA extraction (Thermo-Fisher, Waltham, MA; Sigma-Aldrich, St. Louis, MO). The purity and concentration of DNA can be assessed using various methods, such as spectrophotometry.

[0205] In some implementations, genetic data includes a summary of genetic markers compiled in a Variance Call Format (VCF) file. As understood in the art, VCF files are used in bioinformatics to store gene sequence variations. The VCF format was developed with the advent of large-scale genotyping and DNA sequencing projects, such as the 1000 Genomes Project. Alternatively, a summary can be provided in a Generic Characteristic Format (GFF) containing all genetic data. Typically, GFFs offer redundant functionality because they are shared between genomes. Conversely, with VCF, variations only need to be stored along with a reference genome.

[0206] Microarray technology is widely used to detect markers disclosed herein, such as SNVs / indels and CNVs / SVs. For example, array comparative genomic hybridization (array CGH) and single nucleotide polymorphism (SNP) microarrays can be used. In conventional array CGH, reference DNA and test DNA are fluorescently labeled and hybridized with the array, and the signal ratio is used as an estimate of the copy number (CN) ratio. SNP microarrays are also based on hybridization, but a single sample is processed on each microarray, and an intensity ratio is formed by comparing the intensity of the sample under study with the intensity of a set of reference samples or all other samples under study. While microarrays / genotyping arrays are effective for detecting large CNVs, they are less sensitive for detecting CNVs of short genes or DNA sequences (e.g., less than about 50 kb in length).

[0207] In some implementations, the markers of this disclosure can be detected using next-generation sequencing (NGS). By providing a base-by-base view of the genome, NGS allows the detection of small or novel CNVs that might not be detected by an array. Examples of suitable NGS methods may include whole-genome sequencing (WGS), whole-exome sequencing (WES), or targeted exome sequencing (TES). Preferably, the sequencing method used is WGS.

[0208] In some implementations, samples from subjects are sequenced, for example, using whole-genome sequencing (WGS), and then invoked using standard methods (for SNV / indel and / or CNV / CV markers). For instance, SNV invocation from NGS data utilizes computational methods to determine the presence of single nucleotide variants (SNVs) from the results of next-generation sequencing (NGS) experiments. These techniques are becoming increasingly popular for SNP genotyping due to the ever-growing availability of NGS data, and a wide variety of algorithms have been designed for specific experimental designs and applications. Similarly, several bioinformatics methods exist for detecting CNVs from next-generation sequencing data (Pirooznia et al., Front Genet., 6: 138, 2015). In some implementations, samples are processed and sequenced to obtain sequence files, and these sequence files are processed, for example, using tools such as genomic VCF or exome VCF (eVCF).

[0209] In some embodiments, the methods of this disclosure may involve generating a summary of genetic markers. A typical summary includes genetic data from a tumor sample with whole-genome sequencing and a control (e.g., PMBC). The tumor sample preferably includes excised tumors or FNAs, such as lung adenocarcinoma or skin melanoma. The control sample preferably contains PMBCs isolated using Ficoll as described above. A mixture is then created and the markers therein are analyzed using the computational methods of this disclosure.

[0210] In some embodiments, the methods of this disclosure may include classifying genetic data into different components based on markers contained therein, such as SNVs, CNVs, indels, SVs, mutations, deletions, fusions, etc. In a preferred embodiment, the classification step may include binning the noise-filtered and analyzed somatic SNV (sSNV) and somatic CNV (sCNV) markers separately based on the computational methods of this disclosure. Here, the computational method used to analyze the noise and uniqueness of SNV markers may differ from the method used to analyze CNVs. In some embodiments, computational analysis of SNVs or indels may be performed sequentially with computational analysis of CNVs or SVs. In some embodiments, the analyses may be performed together.

[0211] This disclosure provides mathematical algorithms and computational methods for use in (a) filtering artificial noise; and (b) screening for real markers.

[0212] Regarding noise cancellation, where the marker is SNV or indel, artificial noise is eliminated based on multiple parameters including base quality and / or mapping quality. Typically, the BQ score is a measure of the identification quality of nucleobases generated by automated DNA sequencing. It can be determined using conventional methods, such as the Phred quality score, which is assigned to each nucleotide base in the automated sequencer trace. The Phred quality score (Q) is defined as a property related to the logarithm of the probability (P) of a base call error. For example, if Phred assigns a quality score of 30 to a base, the chance of incorrectly calling that base is one in a thousand. Typically, the BQ of sequencing reads is between 10 and 50, such as BQ scores of 10, 15, 20, 25, 30, 35, or 40.

[0213] Similarly, in the case of sSNV or indel markers, the Mapping Quality (MQ) score is the confidence that the reading actually comes from the position aligned by the mapping algorithm. It can be determined using conventional methods, such as mapping quality scores (see, Li et al., Genome Research 18:1851-8, 2008). Typically, the MQ of a reading is between 10 and 50, for example, MQ scores of approximately 10, 15, 20, 25, 30, 35, or 40.

[0214] In some implementations, noise is removed by performing an optimal receiver operating characteristic (ROC) curve based on the joint base quality (BQ) and mapping quality (MQ) scores, where the curve includes the probabilistic classification of the genetic marker in the summary. Typically, the joint BQMQ score is provided as a matrix (x, y), where x is the BQ score and y is the MQ score. In exemplary implementations, a joint BQMQ score between 10 and 50 (for each parameter) is typically used, for example, BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), etc.

[0215] While not bound by any particular theory, in some implementations, the deletion step filters out “noisy” markers with low base quality and / or mapping quality from a marker profile initially identified as strongly associated with the disease. In some implementations, the deletion step may include: obtaining a threshold probability (P) that satisfies the detection. D For each tag, the tag is classified as either signal or noise based on its ROC curve; and if a tag is classified as noise, it is removed from the summary. Alternatively, this may include, for example, a detection probability (P0). D) and noise probability (P) N A scoring system based on the ratio of 1 / 2 to 1 / 3 can be used to remove tags that do not meet the threshold score.

[0216] Besides BQ and MQ, readout position (RP) can also affect signal quality. In the case of sSNV or indel tags, RP can be mapped, for example, by mapping the position of the starting base of the sequencing read. Other factors affecting tag quality include, for example, the specific sequence context associated with a higher probability of sequencing errors (Chen et al., Science, 355(6326):752–756, 2017). In this respect, genuine mutations can usually be mapped to their own specific sequence context, while errors cannot. For example, tobacco-related mutations tend to occur in a CC context, and mutations associated with APOBEC enzyme activity tend to be used in a TpC context for inserting somatic mutations (see Greenman et al., Nature, 446(7132):153–158, 2007). Therefore, sequence context can be used to help identify variations more likely to be caused by sequencing artifacts as well as variations more likely to be caused by generalized mutational processes.

[0217] Regarding noise cancellation, where the marker is a CNV, artificial noise is eliminated based on multiple CNV-specific parameters. In some embodiments, CNV-specific noise parameters include the “location attribute” of the CNV. Typically, centromeres, telomeres, and / or heterochromatin regions of chromosomes exhibit extensive variability due to their involvement in rearrangements. CNVs located in or near these regions (detected by in situ methods and by computer software) may be disadvantageous. In some embodiments, the location attribute of a CNV can be measured based on whether it is at least 1000 kb, at least 400 kb, at least 100 kb, at least 20 kb, or less than 1 kb, such as 1 kb, from a telomere, centromere, or heterochromatin region of the chromosome. In some embodiments, CNVs located in subtelomere or centromere regions characterized by chromosomal rearrangement hotspots are disadvantageous. A further feature that can be employed in the methods of this disclosure includes the position of the read (PIR) or read location. Read location information can be obtained using various techniques with different location measurements (e.g., genomic coordinates of the read, position on a reference sequence, or chromosomal position). In further implementations, unique molecular indexes (UMIs) and reading locations can be combined to collapse the readings.

[0218] In some implementations, CNV-specific noise parameters include an assessment of the “representativeness” of CNVs with disease characteristics. For example, previous studies have found that CNV calls in immunoglobulin regions are not representative of gDNA and tend to depend primarily on DNA origin—e.g., saliva versus blood or lymphoblastic cell line versus blood (Need et al., 2009; Wang et al., 2007; Sebat et al., 2004). Such non-representative CNVs can be disadvantageous.

[0219] In some implementations, CNV-specific noise parameters include an assessment of the “depth coverage” of the CNV, which refers to the number of unique readings whose mapping overlaps with specific genomic coordinates within the CNV genomic segment.

[0220] Once the noise labels are filtered, the next step in the diagnostic approach involves integrating whole-genome profiling signals from a plasma sample into a mathematical inference model that outputs an estimated score of tumor DNA in the biological sample (e.g., plasma). Depending on the label, the mathematical model integrates multiple process quality metrics and patient-specific attributes to estimate the tumor score (TF). Recognizing the fundamental differences in frequency between SNVs (or indels) and CNVs (SVs) and their association with traits such as cancer, the systems and methods disclosed herein involve using label-specific mathematical algorithms to estimate the tumor score.

[0221] From a workflow perspective, CNV-based detection methods can perform variations of the previously described SNV-based detection methods. In some implementations, baseline samples (e.g., plasma and / or tumor samples) and normal cell samples (e.g., PBMCs) are processed and analyzed separately. In the final analysis step, tumor signals and PBMC signals are binned separately, for example, based on orientation coverage skewness and local fragment size skewness. If the signal is identified as originating from the tumor (tumor CNV / SV), the mathematical model used to estimate the tumor score has forward directionality; conversely, if the signal is identified as originating from the PBMC, the mathematical model used to estimate the tumor score has reverse directionality. Although tumor scores can be estimated using only tumor samples (i.e., without using PBMC samples), this method preferably integrates bidirectionality (i.e., both tumor-based and PBMC-based tumor score estimations are integrated).

[0222] Similar to SNV-based detection methods, CNV-based detection methods also allow for orthogonal integration of secondary features (e.g., fragment size offset). Here, the provisional application (specifically, tumor-based eTF estimation using CNV) covers the main approaches for determining the estimated tumor score (eTF) using mathematical equations incorporating directional features. However, to make the prediction / diagnostic methods more robust, accurate, and / or sensitive, read-based features (e.g., DNA fragment size offset) can be orthogonally integrated into the model. The importance of orthogonal features (in determining MRD) can be determined using a generalized linear model (GLM) to orthogonally determine the tumor score based on the relationship between CNV depth coverage and fragment size offset.

[0223] In some implementations, the CNV-based approach is performed as follows: Germ markers are removed from baseline samples (typically tumor samples, but may also include plasma samples optionally containing tumor samples) and normal samples (typically PBMCs). Next, artificial CNV sites are generated on a cohort of healthy plasma samples (a blacklist of normal PON groups) and removed from mutations detected in patients to remove common sequencing or alignment artifacts, such as centromeres and repetitive regions. The target region of interest (ROI) containing all genomic segments of the tumor (sT_CNV) and PMBC (sP_CNV) is then binned into discrete windows (500 bp or more), and the depth coverage (read counts) for each window is estimated from subsequent plasma samples (post-surgery, during treatment, follow-up for recurrence). The median depth coverage for each window is calculated and then divided by the average sample coverage.

[0224] Next, depth coverage values ​​were normalized to correct for GC content and mappability bias by performing two LOESS regression curve fittings on bin-wise GC scores and mappability scores. Further batch effect correction was performed using robust z-score normalization, applied individually to each sample. In short, the median and median absolute deviation (MAD) were calculated based on the neutral region for each sample, and then all CNV bins were normalized using (B(i)-Median) / MAD. Next, depth coverage skewness and fragment size center of mass (COM) skewness were calculated for each bin compared to healthy plasma samples from the normal group (PON). In this study, samples with low tumor scores showed sparse depth coverage skewness biased by the directionality of CNV segments; amplified segments showed a bias towards positive depth coverage skewness, while missing segments showed a bias towards negative depth coverage skewness. On the other hand, neutral regions exhibit random skewness and no preferred directionality. Therefore, multiplying the differential (plasma-PON) depth coverage skewness by the directionality of the CNV segment (amplification multiplied by +1, deletion multiplied by -1) will sum the CNV signal across the entire genome, while the noise in neutral regions will be eliminated due to random directionality.

[0225] This step is performed mathematically and estimates the tumor score by examining the linear dilution ratio between the cumulative signal detected at the plasma sample and the cumulative signal detected in the tumor. To address noise variations among patients with different CNV patterns, patient-specific CNV signatures are used to calculate the expected noise distribution in the healthy plasma sample (normal group, PON) cohort. The same procedure described above in SNV labeling analysis can be performed to detect patient-specific patterns in healthy plasma samples (PON) or other patients (cross-patient analysis). These detections represent background noise patterns, for which the mean and standard deviation (μ, σ) of the artificial mutation detection rate are calculated. Confidence-based tumor detection and tumor score estimation are achieved if the tumor score detected in a patient is above a threshold (e.g., an artificial tumor score corresponding to an error rate of 1.5 * σ above the mean).

[0226] It is also possible to infer the tumor score from the skewness of the directional whole-genome depth coverage in sP_CNV, for example, using the reverse approach in the workflow described above. Finally, orthogonal features can be integrated into this computational model to improve the robustness, accuracy, sensitivity, or specificity of the algorithm and method. In some embodiments, the method of this disclosure includes estimating TF based on the detection of multiple SNV markers. In this document, the estimated TF (eTF[SNV]) is calculated by integrating process quality metrics including estimated genome coverage and sequencing noise with patient-specific parameters including mutational burden (N). Preferably, the method includes calculating the estimated tumor score (eTF) of SNV markers, where eTF[SNV] = 1 - [1 - (ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific profile detections in the patient sample, σ is a measure of noise estimated empirically, R is the total number of unique reads in the target region (ROI), N is the tumor mutational burden, and cov is the average number of unique reads per site in the ROI.

[0227] In some embodiments, the method of this disclosure includes estimating TF based on the detection of multiple SNV markers. Herein, the estimated TF (eTF[CNV]) is calculated by the coverage direction depth based on the directional integral skewness of the tumor CNV, where copy number amplification is positively skewed and copy number deletion is negatively skewed. Preferably, the method includes calculating an estimated tumor score (eTF) of CNV markers, where eTF[CNV] = (sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth value in a genomic window indexed by {i} representing plasma depth coverage, T is the median depth value in a genomic window indexed by {i} representing tumor depth coverage, and N is the median depth value in a genomic window indexed by {i} representing normal depth coverage.

[0228] On one hand, determining the TF score may include establishing an optimized base / mapping quality filter, using the optimal receiver operating point to filter SNV noise, and analyzing the filtered SNV signal using the integrated mathematical model described above. Example 2 provides a representative method, and the results are shown in Figure 2. The error rate distribution can be evaluated among multiple replicates using control samples and tumor samples. A theoretical threshold for the cutoff can be established using a statistical model (e.g., a binomial model), empirical measurements can be plotted against this statistical model, and the mean / confidence interval for each measurement can be calculated. The noise level in the distribution is identified using the statistical model. A baseline tumor score (TF) that can diagnose tumors is established based on the statistical measurements. From Figure 3D to 3GThe data shows that it is higher than approximately 1×10 -5 The tumor score of the baseline TF value represents the minimum residual disease for most solid tumors, including melanoma, lung and breast tumors.

[0229] In one aspect, determining the TF score may include establishing an appropriate filter for filtering CNV noise and analyzing the filtered CNV signal using the integrated mathematical model described above. Example 3 provides a representative method, the results of which are shown in Figure 5. First, genetic data for the resected tumor, germline (e.g., PBMC), and preoperative biological samples (preferably cfDNA) are obtained. Profiles of tumor read depth, germline read depth, and preoperative plasma cfDNA read depth in representative amplified segments (e.g., 500kb; preferably 100kb) are generated. Depth coverage is normalized across all samples to minimize bias. An integrated mathematical model is used to assess the differences between the genomes of the three samples, wherein the model integrates the read depth skewness across the entire genome as described above. The results show that detection has high sensitivity when integrating the whole-genome CNV pattern using the aforementioned method. More specifically, the above method has a surprising and unexpected ability to detect tumors with TFs as low as approximately 1 / 100,000. This characteristic is evident from the signal-to-noise ratio (SNR) for each TF, where all 10 -5 All the above TF values ​​show positive (>0) signal detection.

[0230] Figures 7A-7C An exemplary system for using the methods of this disclosure is illustrated. In this document, a genetic marker profile is received from a subject (e.g., a cancer patient). The genetic marker profile includes, for example, tumor DNA (e.g., obtained from a resected tumor) and control DNA (e.g., PMBC). Genetic data are analyzed using mutation call factors, and somatic SNVs (sSNVs) are set as a reference for downstream analysis. In some aspects, this reference standard can be personalized, e.g., for a specific subject. In some aspects, this reference standard can be used in conjunction with an additional cohort of reference standards.

[0231] Preferably, to utilize a very clean and high-quality reference set, the outputs of three different mutation call packages, MUTECT, LOFREQ, and STRELKA, are intersected. MUTECT can reliably and accurately identify somatic point mutations in next-generation sequencing data of cancer genomes (Cibulskis et al., Nature Biotechnology, 31, 213–219, 2013); LOFREQ models the sequencing run-specific error rate to accurately call variants occurring in a population of <0.05% (Wilm et al., Nucleic Acids Res., 40(22): 11189–11201, 2012); STRELKA is an analysis package designed to detect somatic SNVs and small indels from aligned sequencing reads of matched tumor-normal samples (Saunders et al., Bioinformatics, 28(14):1811-7, 2012).

[0232] Typically, mutation caller intersection involves using multiple callers known in the art. In some implementations, three mutation callers (MUTECT, LOFREQ, and STRELKA) are used on patient tumor and normal sequencing reads, and the intersecting list of variants is defined as the detection of identical substitutions (the same genomic coordinates and nucleotide changes) in all callers.

[0233] Next, reads from patient-specific mutation sites are collected and filtered. In some implementations, the collection and / or filtering steps include removing low-map quality reads. For example, any reads with a mapping quality score less than 29 (ROC optimized) will be filtered. Additionally or alternatively, filtering may involve establishing repeat families. For example, repeats may include multiple PCR / sequencing copies of the same DNA fragment (i.e., repeats of non-unique marker and target regions). Finally, corrected reads based on a concordance test can be generated. The filtering step may include removing low-base-quality reads. For example, any reads with a base-quality score less than 21 (ROC optimized) can be filtered. Finally, the filtering step may include removing high fragment size reads. For example, any reads with a fragment size greater than 160 (ROC optimized) can be filtered. The principle is that tumor DNA tends to be shorter than normal DNA, so a low fragment size filter enriches tumor DNA. See Jiang et al., PNAS USA, 112.11 (2015):E1317-E1325; and Mouliere et al., bioRxiv, 134437, 2017.

[0234] The next step involves calculating the number of patient-specific mutation sites with at least one supporting read (in the filtered set) that has the exact same permutation as in the tumor. Among the aspects labeled as SNVs, the calculation step may include an integral probability model comprising: 1) an integral signal from plasma SNV detection; 2) a process quality metric including an estimated genome coverage and sequencing noise model; and 3) a patient-specific parameter incorporating the mutational burden (N). More specifically, the integrated mathematical model may involve calculating the estimated eTF[SNV] = 1 - [1 - (ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific SNV profile detections in the patient's plasma sample, σ is a measure of the error rate estimated empirically, R is the total number of unique reads in the SNV profile target region (ROI), N is the tumor mutational burden, and cov is the average number of unique reads per site in the SNV profile ROI. Next, the estimated TF is checked against a detection threshold defined by an empirically measured baseline noise TF estimate of healthy samples. In some respects, a TF is defined as detected if it is above a threshold, such as two standard deviations of the noise TF distribution (e.g., FPR < 2.5%).

[0235] In some implementations where CNV is denoted, the filtering step may include running CNV calls (e.g., amplification and / or deletion analysis) on tumor and normal (e.g., PBMC) samples from the patient and generating reference partitions for all CNV segments that meet threshold characteristics (e.g., length greater than 5 Mega base pairs) and the directionality of the variants (where amplification is assigned a positive factor, e.g., +1, and deletion is assigned a negative factor, e.g., -1). Next, single-base-pair depth coverage information for plasma, tumor, and PBMC samples covering the patient-specific CNV partition ROI is collected. Next, the patient-specific CNV partition ROI is normalized to 500 bp windows, and the median (artificial suppression) for each window is calculated for all samples and windows. Next, normalized depth coverage information for all 500 bp windows is generated.

[0236] In some implementations, normalization can be performed using (1) robust z-score normalization per sample and / or (2) robust principal component analysis (RPCA) methods. For example, the z-score method may include using the algebraic function preop_median = (preop_median - median(preop_median)). / (1.4826*mad(preop_median,1)). Alternatively, the robust principal component analysis (RPCA) method may involve solving an optimization problem of M = L + S to remove noise and high-frequency artifacts (S matrix). A combination of the above methods may also be used.

[0237] Next, readings / windows from patient-specific regions are filtered. In some implementations, the filtering step may include removing low-mapped-quality readings (e.g., <29, ROC optimized); removing readings adjacent to centromere regions, for example, removing windows with normalized normal values ​​above a threshold (e.g., 10). Regarding centromere-near filters, approximately 70%–80% of CNV noise hotspots co-located with centromere regions have been identified and can be detected by abnormally high depth coverage values ​​in PBMC samples. These centromere hotspots can be removed in the filtering step.

[0238] Next, unrepresented regions in the cfDNA are removed. For example, windows not included in a cfDNA representation mask composed of multiple cfDNA samples can be removed. The basic principle of this filtering step is that including these unrepresented regions in the calculations can lead to bias and errors, especially when the cfDNA is biased towards displaying only nucleosome-protected genomic regions and showing unrepresented gaps in accessible chromatin genomic regions. Therefore, a mask representing the regions (>0 readings) in the cfDNA queue is generated using the cfDNA sample queue.

[0239] Next, a computational method is used to integrate the coverage parameters of plasma and normal samples. Therefore, the equation sum can be used. i [(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma) is the integral of the coverage direction depth of the skewness between plasma and normal (PBMC) patient samples. Similarly, the sum can be used. i [abs(T(i)-N(i))]-E(σ)) is used to integrate the cumulative coverage depth of the skewed distribution between tumor and normal (PBMC) patient samples.

[0240] Next, the dilution ratio of the aforementioned signal relative to directional depth and cumulative coverage depth is calculated, corresponding to the estimated tumor fraction (eTF). In some aspects, the calculation steps may include calculating the eTF of CNV markers by utilizing a probabilistic dilution model, which includes: 1) integrating the skewed directional coverage depth between plasma and normal (PBMC) patient samples according to the directionality of the tumor CNV, where copy number amplification is positively skewed and copy number loss is negatively skewed; 2) integrating the skewed cumulative coverage depth between tumor and normal (PBMC) patient samples; and 3) determining the dilution ratio between the aforementioned signals. More specifically, the integrated mathematical model can involve calculating the estimated eTF[CNV]=(sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth coverage value in the genomic window indexed by {i} representing plasma depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; and N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort. Next, the estimated TF (CNV) is checked against a detection threshold defined by an empirically measured baseline noise TF estimate of a healthy sample. In some respects, if the eTF (CNV) is above a threshold, such as two standard deviations of the noise TF distribution (e.g., FPR < 2.5%), it is defined as detected.

[0241] In some implementations, a probabilistic model is used to calculate the effective coverage of each genomic locus based on the mathematical operation A*PBMC_cov + B*tumor_cov, where PBMC coverage and tumor coverage are not identical if a particular locus is associated with amplification or deletion, and A+B=1. In some implementations, A and B are as follows for various samples: controls (e.g., PBMC samples) A=1 and B=0; tumor samples B=purity and A=1-purity; plasma samples B=TF and A=1-TF. In some implementations, the relationship between signal in plasma and tumor is linearly correlated with dilution (or change in the proportion of mixtures) between purity and TF. As is known in the art, the model is also susceptible to noise, which may be included in the probabilistic model.

[0242] Application of the method in the treatment of postoperative patients

[0243] For cancer patients who have undergone surgical removal of tumors (e.g., removal of breast tumors via mastectomy; removal of lung tumors via pulmonary resection or lobectomy; or removal of the prostate via prostatectomy), prognosis is crucial. For example, among breast cancer patients, the vast majority of women considering adjuvant therapy expressed a desire to know their prognosis without adjuvant therapy (Ravdin et al., J Clin Oncol., 16(2):515-521, 1998). Adjuvant therapy is undesirable because it is unpleasant and inconvenient (Duric et al., Lancer Oncol., 2(11):691-697, 2001). In some cases, it may only provide modest benefits (Simes et al., J Natl CancerInst Monogr., 30, 146-152, 2001). Whether or not to have it is a legitimate decision (Duric et al., ibid.). It may involve weighing the options of Wouters et al. (Ann Oncol., 24(9):2324-9, 2013). There are calls to improve the determination of the risks posed by cancer (Kratz et al., Transl Lung Cancer Res., 2(3): 222–225, 2013).

[0244] Many studies have indicated that tumor size is an important prognostic variable. However, in the context of MRD (Morphological Research and Development), tumor size is not relevant because tumors are often undetectable using conventional diagnostic tools such as CT scans. Therefore, the threshold for tumor size is problematic.

[0245] Therefore, a computerized version of the predictive model will provide an important step in this direction and may be the most accurate predictive method currently available. Figure 7 illustrates model predictions in postoperative patients based on estimated tumor scores. For example, estimated tumor scores above a threshold (e.g., approximately 10 for SNV markers) -4 And / or approximately 10 for SNV markers -5 This will indicate to the subject that they require adjunctive therapy.

[0246] Beyond its simple use in patient consultations, this model can also be used by physicians to make decisions regarding adjuvant therapy. Therefore, the disclosed approach provides physicians and clinicians with a tool to predict outcomes (e.g., metastasis or even death) in the absence of adjuvant therapy. It is presumably that patients with very low baseline risk (as a function of estimated tumor fraction (eTF)) may wish to avoid toxicities associated with adjuvant therapy. Thus, the predictive tool can be an effective decision-making aid. This predictive tool could also serve as a benchmark for judging the predictive ability of any new therapy (e.g., chemotherapy, immunotherapy, or targeted therapy, such as the use of investigational drugs).

[0247] system

[0248] This disclosure further relates to systems for performing the methods of this disclosure. Figure 7A The schematic diagram provides a representative system illustrating an exemplary system for performing the diagnostic methods of this disclosure. As depicted herein, a system 500 is provided, which may include an analysis unit 510, a classification unit 520, a calculation unit 530, and a display 540 for outputting data and receiving user input (via an associated input device (not shown)). The analysis unit 510 typically includes input of genetic data, such as a VCF file containing readings of tumor samples from a subject, optionally normal (e.g., PBMC) samples, and a second biological sample, such as a plasma sample from the same subject (note: the first and second sample collections may be performed together or sequentially, i.e., separated in time). The classification unit 520 may include one or more engines for classifying various types of markers, such as CNV / SV versus SNP / indel. It should be noted that... Figure 7A One configuration of the system is shown. The orientation and configuration of these components can be changed as needed. Furthermore, additional components can be added to the system. These various components, their various operations, their various orientations, and their various relationships with each other will be discussed in detail below.

[0249] In some embodiments, this disclosure relates to a system for detecting residual disease in subjects in whom it is required. The system may include an analysis unit 510 configured and arranged to filter artificial noise markers from a labeled whole-genome profile, wherein the labeled whole-genome profile is generated from multiple genetic markers in biological samples from a subject, including tumor samples and normal cell samples, wherein the genetic marker profile is selected from the group consisting of single nucleotide variants (SNVs), indels, copy number variants (CNVs), structural variants (SVs), and combinations thereof. The analysis unit further includes detecting a subject-specific whole-genome profile of genetic markers in a second biological sample to generate a representative of tumor whole-genome genetic markers in the second sample. The analysis unit further includes a classification engine 520. In some embodiments, the classification engine 520 statistically classifies each marker in the profile as either a signal or noise. For example, where the markers are SNVs or indels (grouped together due to similar structural features but not necessarily using the same classification scheme), the classification engine classifies noise based on the probability (P) of detection. NThe classification engine categorizes SNVs or indels as signals or noise based on 1) the mapping quality (MQ) of the read group containing the SNV or indel, 2) the fragment size length of the read group containing the SNV or indel, 3) the consistency test within the read repeat family containing a specific SNV, and 4) the base quality (BQ) of the SNV or indel. Similarly, where the label is an SNV or indel (grouped together due to similar structural features but not necessarily using the same classification scheme), the classification engine categorizes SNVs or indels as signals or noise based on 1) their position relative to the centromere, 2) the mapping quality (MQ) of the read group containing the CNV or SV window, and 3) the representativeness of the CNV or SV window in the cfDNA data.

[0250] In some implementations, the SNV / indel classification unit 520 is based on the probability of detecting noise (P) N The CNV / SV classification unit 520 classifies each SNV / SV statistically in the profile as a signal or noise, as a function of the base quality (BQ) and the mapping quality (MQ) of the SNV / SV. In some embodiments, the CNV / SV classification unit 520 classifies each CNV / SV statistically in the profile as a signal or noise based on its position relative to the centromere, its non-representativeness at a given coverage depth, and its reading ability. In some embodiments, the classification unit 520 classifies SNV / SV markers and CNV / SV markers based on one or more of the aforementioned parameters.

[0251] In some embodiments, the system of this disclosure includes a computational unit 530 configured and arranged to compute an estimated tumor fraction (eTF) of a sample based on one or more integrated mathematical models. For example, the computational unit may be configured and arranged to compute the eTF of a sample based on one or more integrated mathematical models specific to SNV / indel markers or specific to CNV / SV markers. In such an embodiment, where the marker is an SNV / indel, the computational unit may integrate process quality metrics including estimated genomic coverage and sequencing noise with patient-specific parameters including mutational burden (N). Similarly, where the marker is a CNV or SV, the computational unit may compute the eTF of a CNV marker by integrating a skewed directional coverage depth consistent with the CNV orientation of the tumor, where copy number amplification is positively skewed and copy number deletion is negatively skewed.

[0252] The system of this disclosure further includes a display unit 540 that outputs a residual disease profile of the subject based on an estimated tumor score, wherein if the estimated tumor score exceeds an empirical threshold calculated by a background noise model, the subject's residual disease is output in the residual disease profile. In some embodiments, in the system of this disclosure, a classification engine unit and / or a calculation unit may be coupled individually or jointly to the display unit, which outputs the subject's residual disease profile based on the estimated tumor score.

[0253] In some embodiments, the system 500 of this disclosure includes an analysis unit 510, which includes a classification unit 520. The classification unit 520 includes at least one engine selected from the group consisting of an SNV classification engine 520-1, a CNV classification engine 520-2, an indel classification unit 520-3, a structural variant (SV) classification unit 520-4, or combinations thereof 520-5, wherein: the SNV / indel classification engine is based on the probability (P) of detecting noise. N The system classifies each SNV in the summary as signal or noise as a function of the base quality (BQ) or mapping quality (MQ) of the SNV; and / or the CNV / SV classification engine classifies each CNV / SV in the summary as signal or noise based on its position relative to the centromere, its non-representativeness at a given coverage depth, and its readability. System 500 may further include a computation unit 530 configured to compute an estimated tumor fraction (eTF) of the sample based on one or more integrated mathematical models specific to the marker type. For example, where the marker is an SNV, computation unit 530 may be configured to compute an estimated tumor fraction (eTF) of the sample based on the mathematical model eTF[SNV] = 1 - [1 - (ME(σ)] / (σ)] / (σ) R The eTF is calculated using [1 / cov]^(M / N), where M is the number of tumor-specific profile detections in the patient sample, σ is a measure of noise estimated empirically, R is the total number of unique readings in the target region (ROI), N is the tumor mutation burden, and cov is the average number of unique readings at each site in the ROI. Similarly, where the label is CNV, the calculation unit 530 can be configured to calculate the eTF based on the mathematical model eTF[CNV] = (sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth value in a genomic window indexed by {i} representing plasma depth coverage, T is the median depth value in a genomic window indexed by {i} representing tumor depth coverage, and N is the median depth value in a genomic window indexed by {i} representing normal depth coverage.

[0254] In some embodiments, calculation unit 530 may be configured to calculate the eTF based on an indel-specific mathematical model (typically similar to or the same as the mathematical model used to calculate the eTF of the SNP). In some embodiments, calculation unit 530 may be configured to calculate the eTF based on an SV-specific mathematical model (typically similar to or the same as the mathematical model used to calculate the eTF of the CNV). In some embodiments, calculation unit 530 may be configured to calculate the eTF based on both an SNP-specific mathematical model and a CNV-specific mathematical model, wherein the SNP-specific mathematical model includes the equation eTF[SNV] = 1 - [1 - (ME(σ)]] R ) / N]^(1 / cov), where M is the number of tumor-specific profile detections in the patient sample, σ is the empirically estimated noise measurement, R is the total number of unique readings in the target region (ROI), N is the tumor mutation burden, and cov is the average number of unique readings at each site in the ROI. The CNV-specific mathematical model includes the equation eTF[CNV] =(sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth value in the genomic window indexed by {i} representing plasma depth coverage, T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, and N is the median depth value in the genomic window indexed by {i} representing normal depth coverage.

[0255] In some implementations, the computation unit 530 is configured to calculate the eTF of SNV or Indel markers using an integral probability model, wherein the probability model includes 1) an integral signal of plasma SNV or Indel detection, 2) a process quality metric including an estimated genome coverage and sequencing noise model, and / or 3) a patient-specific parameter including mutational burden (N); and / or to calculate the eTF of CNV or SV markers using a probabilistic mixture model, wherein the probabilistic dilution model includes 1) integrating the skewed directional coverage depth between plasma and normal patient samples based on the orientation of tumor CNV or SV, where copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative coverage depth of the skewed skew between tumor and normal patient samples; and / or 3) finding the dilution ratio between the above signals.

[0256] According to various embodiments of this document, a computer-readable medium is provided, comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for filtering noise in a genetic marker profile received from a subject sample, wherein the genetic markers include SNVs (preferably sSNVs), CNVs (preferably sCNVs), indels, and / or SVs (preferably translocations, gene fusions, or combinations thereof) in genome readings. Preferably, the filter removes artificial noise markers from the whole-genome profile of the markers by: based on the probability of detecting noise (P... N The summary statistically classifies each SNV or Indel in the summary as signal or noise, as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and 4) the base quality (BQ) of SNVs or Indels; and / or statistically classifies each CNV or SV window in the summary as signal or noise based on 1) its position relative to the centromere, 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and 3) the representativeness of CNV windows in cfDNA data. The computer-readable medium may further include computer-executable instructions that, when executed by a processor, cause the processor to perform a set of methods or steps: calculating an estimated tumor fraction (eTF) of a biological sample based on one or more integrated mathematical models; and diagnosing residual disease in a subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model.

[0257] In some embodiments, the system includes a computing unit 530 comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a set of methods or steps: estimating a tumor fraction (eTF) based on one or more of the aforementioned mathematical models used to calculate the eTF; and a diagnostic unit for making a qualifying diagnosis based on the calculated eTF (e.g., making a positive diagnosis if the eTF ≥ 2 std is above a noise threshold). The system may further include a display 540 for outputting data and receiving user input via an associated input device (e.g., a mouse). In some embodiments, results may be displayed on the display 540 as binary output (i.e., “MRD +ve” or “MRD -ve”) or ordinal scores (e.g., in a scale of 1 to 5), where a score of 1 indicates that the subject is unlikely to have MRD, and a score of 5 indicates that the subject is likely to have MRD.

[0258] like Figure 7B As shown, an example system 100 is provided, which is configured and arranged to detect residual disease in subjects for whom it is required. (Reference) Figure 7B System 100 may include an analysis unit 110 and a calculation unit 150. The analysis unit 110 may include a pre-filter engine 120 and a correction engine 130. These system components and associated engines will be discussed in detail below.

[0259] Refer again Figure 7B The pre-filter engine 120 of the analysis unit 110 can be configured and arranged to receive a whole-genome summary of first subject-specific readings associated with genetic markers from a first biological sample of the subject. As discussed in the workflow herein, and according to various embodiments, the first biological sample may include a baseline sample; the first reading summary may each contain readings of single base pair lengths; the baseline sample may include a tumor sample or a plasma sample.

[0260] Figure 7B The pre-filter engine 120 can also be configured and arranged to filter artificial sites from the first readout summary. As discussed in the workflow herein, and according to various embodiments, filtering may include removing duplicate sites generated on a reference healthy sample cohort from the first readout summary of genetic markers, and / or identifying germline mutations in peripheral blood mononuclear cells of normal cell samples and removing said germline mutations from the first readout summary of genetic markers.

[0261] exist Figure 7B In this configuration, the correction engine 130 of the analysis unit 110 can be configured and arranged to receive output from the engine 120. The correction engine 130 can also be configured and arranged to receive readings from a second subject-specific whole-genome summary of genetic markers in a second biological sample from the subject, to generate a tumor-associated whole-genome representation of the genetic markers in the second sample. Figure 7B As shown, the detection unit 140 can be used to detect readings of the second biological sample. The detection unit 140 may or may not be part of the system 100; in the latter case, the readings may be simply received by the calibration engine 130 from the external system 100. Furthermore, these readings can be received at any point in the system before noise filtering, as will be discussed below. Moreover, if the readings are provided to the system 110 after noise filtering, they can even be received after noise filtering. Furthermore, the detection unit 140 may be integrated into the analysis unit 110 or separated from it, such as... Figure 7B As shown.

[0262] The correction engine 130 can also be configured and arranged to use at least one error suppression scheme to filter noise from the first and second read genome summaries to produce a first filtered read set for the first read genome summaries and a second filtered read set for the second read genome summaries.

[0263] As discussed in the workflow of this document, and according to various embodiments, the at least one error suppression scheme may include calculating the probability that any single nucleotide variant in the first and second summaries is an artificial mutation, and removing the mutation.

[0264] As discussed in the workflow of this paper, and according to various embodiments, the probability can be calculated as a function of features selected from: mapping quality (MQ), variant base quality (MBQ), reading position (PIR), average reading base quality (MRBQ), and combinations thereof.

[0265] As discussed in the workflow of this document, and according to various embodiments, the at least one error suppression scheme may include removing artificial mutations by using inconsistency tests and / or repeat consistency tests between independent repeats of the same DNA fragment generated by polymerase chain reaction or sequencing processing (where artificial mutations can be identified and removed when consistency is lacking in most of a given repeat family).

[0266] The computation unit 150 of system 100 can be configured and arranged to receive the output from the correction engine 130 and to compute the estimated tumor fraction (eTF) of the first and second biological samples using first and second filtered reading sets by applying a background noise model to one or more integrated mathematical models. The computation unit 150 can be further configured and arranged to detect residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold. The background noise model, integrated mathematical models, and empirical thresholds will be discussed in detail herein.

[0267] System 100 may also include a display 160, such as Figure 7B As shown. The display can be configured and arranged to receive output from computing unit 150. The output may include data related to the detection of residual disease in the subject / user. Alternatively, system 100 may not include a display and the data output from computing unit 150 may be sent to any form of storage or display device or location outside system 100. Also, as discussed herein, components of system 100 may be integrated into a single unit or may be broken down into smaller units. Figure 7BThe diagram shows more separate physical units. Furthermore, system 100 can be part of a distributed network of systems, each performing essentially similar tasks and transmitting data from each system to a hub.

[0268] like Figure 7C As shown, an example system 100 is provided, which is configured and arranged to detect residual disease in subjects for whom it is required. Figure 7C As shown in the example system, system 100 may include an analysis unit 110 and a calculation unit 150. Figure 7B The system is the opposite of, Figure 7C The analysis unit 110 may include a pre-filter engine 120 and a normalization engine 130. These system components and associated engines will be discussed in detail below.

[0269] Refer again Figure 7C The pre-filter engine 120 of the analysis unit 110 can be configured and arranged to receive a whole-genome summary of first subject-specific readings associated with genetic markers from a first biological sample of the subject. As discussed in the workflow herein, and according to various embodiments, the first biological sample may include a baseline sample; the first reading summary may each contain readings of single base pair lengths; the baseline sample may include a tumor sample or a plasma sample.

[0270] The prefilter engine 120 can also be configured and arranged to receive a whole-genome summary of second subject-specific readings associated with genetic markers from a second biological sample of the subject. As discussed in the workflow described herein, and according to various embodiments, the second biological sample may include a peripheral blood mononuclear cell sample (PBMC); and the second summary of the genetic markers may each include copy number variations (CNVs).

[0271] The pre-filter engine 120 can also be configured and arranged to filter artificial sites from the first and second reading profiles. As discussed in the workflow of this document, and according to various embodiments, filtering may include removing duplicate sites generated on a reference healthy sample cohort from the first and second reading profiles; identifying CNVs shared between the first and second profiles as germline mutations, and removing said mutations from the first and second reading profiles.

[0272] The normalization engine 130 of the analysis unit 110 can be configured and arranged to receive output from the engine 120. The normalization engine 130 can also be configured and arranged to receive readings from a third subject-specific whole genome summary of genetic markers in a third biological sample from the subject to generate a tumor-associated whole genome representative of the genetic markers in the second sample.

[0273] like Figure 7CAs shown, detection unit 140 can be used to detect readings of a third biological sample. Detection unit 140 may be part of system 100 or not, in which case the readings may be received by normalization engine 130 from external system 100. Furthermore, these readings can be received at any point in the system before noise filtering in analysis unit 110, as will be discussed below. Moreover, if readings are provided to system 110 after noise filtering, these readings can even be received after noise filtering. Furthermore, detection unit 140 may be integrated into analysis unit 110 or separated from analysis unit 110, such as... Figure 7C As shown.

[0274] The normalization engine 130 can also be configured and arranged to normalize each of the first, second, and third read summaries to produce a first set of filtered reads for the first read genome summary, a second set of filtered reads for the second read genome summary, and a third set of filtered reads for the third read genome summary. Normalization methods are discussed in detail herein and can be used in any intended combination to normalize the reads in question.

[0275] Figure 7C The computation unit 150 of system 100 can be configured and arranged to receive output from normalization engine X30 and, for example, to calculate an estimated tumor fraction (eTF) of a third biological sample using a third set of filtered readings by applying a background noise model to one or more integrated mathematical models, wherein the one or more models generate a first eTF using a first set of filtered readings, and / or the one or more models generate a second eTF using a second set of filtered readings. The computation unit 150 can be further configured and arranged to detect residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold. The background noise model, integrated mathematical model, and empirical threshold will be discussed in detail herein.

[0276] System 100 may also include a display 160, such as Figure 7C As shown. The display can be configured and arranged to receive output from computing unit 150. The output may include data related to the detection of residual disease in the subject / user. Alternatively, system 100 may not include a display and the data output from computing unit 150 may be sent to any form of storage or display device or location outside system 100. Also, as discussed herein, components of system 100 may be integrated into a single unit or may be broken down into smaller units. Figure 7C The diagram shows more separate physical units. Furthermore, system 100 can be part of a distributed network of systems, each performing essentially similar tasks and transmitting data from each system to a hub.

[0277] Other relevant implementation plans

[0278] estimation of transplant rejection

[0279] This disclosure further relates to using the aforementioned systems, methods, and algorithms to estimate transplant rejection. Preferably, it can be used... Figure 1B and Figure 1D The workflow based on SNV / indel, as outlined in the document, is used to estimate transplant rejection.

[0280] In some implementations, the assessment of transplant rejection is based on a protocol that uses a reference only for donor-specific SNPs (and not present in the recipient). Based on the detection rate of these donor-specific SNPs in the recipient's blood (e.g., post-transplant), the donor DNA score can be calculated using the methods and systems of this disclosure. The expected donor DNA score is correlated with the apoptosis rate or rejection rate of the transplanted tissue. For example, a high donor DNA score is associated with a high rejection phenotype, and a low donor DNA score is associated with a low rejection phenotype.

[0281] In some implementations, the differential SNP between the donor and recipient, as measured using the methods of this disclosure, can be used to estimate the fraction of donor DNA (eDF) in the recipient's blood sample. Based on the eDF, the probability / likelihood of a transplant that will be rejected is calculated. For example, if the eDF is greater than a certain threshold, it indicates that the transplanted tissue will be rejected by the host or is incompatible with the host. Conversely, if the eDF is equal to or below the threshold level, it indicates that the transplanted tissue will be accepted by the host or is compatible with the host.

[0282] Non-invasive prenatal testing (NIPT) for chromosomal abnormalities

[0283] This disclosure further relates to non-invasive prenatal testing (NIPT) for chromosomal aberrations using the aforementioned systems, methods, and algorithms. Preferably, it can be used... Figure 1C and Figure 1E This paper outlines a CNV / SV-based workflow for performing NIPT. Known amplifications and deletions are used as a CNV reference set for measurement of subject samples (e.g., amniotic fluid or blood obtained from pregnant women carrying fetuses suspected of having chromosomal abnormalities). Figure 1C and Figure 1E The workflow is designed to detect copy number variations even with low and sparse signal, assuming the target region and directionality (amplification, deletion) are known. In the context of NIPT, assuming the need to test for trisomy 21 in maternal blood, the target region (chromosome 21) and the direction of the variation (amplification) are both known. Example

[0284] The structures, materials, compositions, and methods described herein are intended as representative examples of this disclosure, and it should be understood that the scope of this disclosure is not limited to the scope of the embodiments. Those skilled in the art will recognize that this disclosure can be implemented through variations of the disclosed structures, materials, compositions, and methods, and such variations are considered to be within the scope of this disclosure.

[0285] Example 1: Methods and systems for detecting and validating tumor-specific low-abundance tumor markers and their use in cancer diagnosis.

[0286] The systems and methods disclosed herein can be used to detect minimal residual disease. As is known in the art, the abundance of ctDNA limits the use of targeted sequencing technologies in the case of residual disease detection, compared to metastatic cancer (characterized by high disease burden and significantly elevated ctDNA). Considering the known limited amount of cfDNA in cases of low tumor burden, the potential for optimizing cfDNA extraction was first investigated. First, to reduce variability arising from sample collection and inter-individual differences, commercially available extraction kits and methods were compared using homogeneous cfDNA material generated from large-volume plasma collections (approximately 300 cc) produced by plasma removal from healthy subjects and cancer patients undergoing hematopoietic stem cell collection. Large volumes of plasma allow for testing of multiple method and protocol parameters on the same cfDNA input, enabling precise measurement of subtle differences in yield and quality.

[0287] In this comparative study, kits and / or extraction methods from Capital Biosciences (Gaithersburg, MD, USA; Catalog # CFDNA-0050), Qiagen (Germantown, MD, USA), Zymo (Irvine, CA, USA; Catalog # D4076), Omega BIO-TEK (Norcross, GA, USA; Catalog # M3298), and NEOGENESTAR (Somerset, NJ, USA; Catalog # NGS-cfDNA-WPR) were used. These kits and reagents were used uniformly according to the manufacturer's instructions for extraction from large 1 mL plasma samples. Multiple plasma aliquots were processed in parallel to assess inter- and intra-method variability. The yield and purity of each recovered cfDNA sample were determined using quantitative PCR (total mass), UV absorbance (for detecting salt and protein contaminants), and on-chip electrophoresis (for size distribution and gDNA contamination).

[0288] The results showed that the MAG-BIND cfDNA extraction kit from Omega BIO-TEK outperformed all other assay methods. Further systematic optimization of each step of the manufacturer's protocol was performed to reduce contaminant residues and improve cfDNA recovery. Even so, cfDNA yields in early NSCLC (n = 21) remained low and highly variable (median 5 ng / ml (<1000 genomic equivalents); range 3–30 ng / ml).

[0289] The data above support the view that the detection of single-point mutations in patient plasma samples is generated by two consecutive statistical sampling processes: (i) the probability of sampling the mutated fragment with a limited number of genomic equivalents present in a typical plasma sample; and (ii) the probability of detecting the mutated fragment in the sample given its abundance, sequencing depth, and sequencing errors (signal-to-noise ratio). While the latter process has been the focus of in-depth research and technological development in the scientific community (e.g., ultra-deep error-free sequencing protocols), the stochastic process of the former has been rarely addressed. However, both processes play a crucial role in the detection of low disease burden ctDNA, as shown in Figure 2. Even ideal ultra-deep targeted sequencing will fail to detect cancer signals if there is no physical fragment containing the target mutation. In practice, the fact that a single observation (mutation sequencing read) is rarely sufficient for reliable detection further complicates the issue.

[0290] Therefore, the genomic equivalent present in plasma samples constitutes a random sampling of the entire pool of cfDNA fragments in the patient's circulation, which can be formulated using a Bernoulli trial random sampling model. This model predicts that the probability of detection of TFs associated with early cancer treatment regimens (TF < 1%) will decrease rapidly for low TFs. Even at a frequency of 0.1% (1 / 1000), the predicted detection probability will be less than 0.65 ( Figure 2A However, by repeating Bernoulli trials at a large number of sites, the breadth of sequencing introduced can compensate for the limited coverage at each site (a function of the finite genome equivalent). Using this model, it was found that even with a TF of 1:100,000, with modest sequencing work (e.g., 20X coverage), Figure 2B With an integration of over 20,000 point mutations (approximately 10 mutations / mb found in 17% of human cancers), it can provide a high detection probability (up to 0.98), making it easily achievable with standard whole-genome sequencing (WGS).

[0291] The optimized extraction protocol was then applied to patient samples. This cohort included six postoperative (~14 days) plasma samples from the same patients for estimating minimal residual disease (MRD), and four plasma samples from benign patients (controls). Despite optimized extraction, cfDNA yields in low disease burden samples remained low and showed high variability among patients, ranging from 0.13 ng / mL to 1.6 ng / mL. These data confirm the small and variable number of DNA molecules available for cfDNA sequencing.

[0292] In summary, these results indicate that, in the case of MRD detection, the limited input material constitutes a major obstacle to the effective application of ultra-deep targeted sequencing (minimum ctDNA frequency of 0.1-1%), given that the number of genomic equivalents is far lower than the sequencing depth applied.

[0293] Example 2: Whole genome integration allows for sensitive WGS-based ctDNA detection of residual disease in NSCLC after surgery, aiding in therapy stratification and therapy optimization.

[0294] Hypersensitive identification of MRD using cfDNA may have fundamental prognostic implications and allow for patient stratification for subsequent adjuvant chemotherapy. Current approaches primarily expand the paradigm for driver hotspot mutation detection by increasing sequencing depth to offset the low fraction of ctDNA in cfDNA. However, these approaches are inherently limited by the upper limit of genomic equivalents. To overcome this limitation, whole-genome information is integrated, the rationale being that information incorporated across the entire genome will allow for the utilization of the high mutation rate in lung cancer. Thus, instead of relying on deeper sequencing at a few sites, the breadth of mutation detection is extended to the entire genome to improve sensitivity. Consequently, WGS has been used for base sensitivity detection of the cumulative signal provided by the 10,000–30,000 somatic mutations observed in a significant proportion of NSCLC. Notably, most of these mutations are thought to occur prior to transformation, and therefore, they may even be present in early-stage NSCLC. To evaluate methods for detecting residual disease in postoperative NSCLC patients for therapeutic purposes, samples from five early-stage lung cancer patients were analyzed (full clinical details are provided in Table 1).

[0295] Table 1: Clinical information of patients currently undergoing sequencing.

[0296]

[0297] First, matched tumor DNA and germline DNA from peripheral blood mononuclear cells (PBMCs) were subjected to whole-genome sequencing (WGS) to generate patient-specific whole-genome sSNV profiles. Additionally, plasma samples were collected before surgery and approximately 14 days after surgical resection. cfDNA was extracted using the optimized MAG-BIND cfDNA extraction kit, and libraries were prepared from only 1 ng of patient cfDNA according to the kit.

[0298] Next, point mutation pattern matching was used to detect MRD. To this end, a robust mathematical model was established to estimate tumor fractions for SNV and CNV markers. The mathematical model showed that increasing the number of sites would lead to a significant increase in the detection probability. To validate this prediction, a computer-generated mixture of tumor and normal WGS data from multiple lung adenocarcinoma patients was used to simulate cfDNA detection by mixing tumor and normal WGS readings at different proportions to obtain different TF(10) values. -2 Up to 10 -6 Virtual plasma samples (n = 5 replicates) were generated. To simulate noise and potential false detection, a supplementary dataset of sequencing reads was generated from matched normal germline WGS, while a mixture without tumor reads (TF = 0, n = 20 replicates) was used. To simulate detection in the context of residual disease, somatic mutation calls were performed on the original tumor and germline WGS data, and patient-specific somatic SNV profiles were obtained. The number of tumor-associated mutation sites in the computer-simulated plasma mixture was then measured by detecting at least one supporting read for the patient-specific SNV profile. By analyzing simulated plasma with and without ctDNA, sequencing noise was identified as a major obstacle to sensitive detection. To reduce the impact of sequencing artifacts, errors associated with low base quality (BQ) and mapping quality (MQ) markers were filtered out. Optimal receiver point analysis (ROC) was used to further reduce the impact of sequencing artifacts. Figure 3A A combined BQ and MQ optimization filter was developed, reducing the measurement error rate by -10x. Figure 3B The number of SNVs detected is reduced to approximately 2 / 10,000. In summary, this optimized SNV detection method is effective in our proposed mathematical approach (red line). Figure 3C ) and empirical data of measurements (mean + / - confidence interval, Figure 3C The results showed high consistency with the mathematical model, and a high sensitivity close to TF = 1 / 100,000. Furthermore, the high consistency between the experimental results and the mathematical model allowed us to accurately convert empirical SNV detections into TF estimates. Figure 3D This allows for quantitative MRD monitoring. Furthermore, computer validation of TF estimation shows that for 5 × 10⁻⁶... -5 All of the above TFs have been accurately and specifically estimated. Figure 3E , Figure 3F and Figure 3G Here, in three different samples (e.g., melanoma), Figure 3E ),lung( Figure 3F ) and breast ( Figure 3G In tumor samples, a high correlation (R0) was observed between the input mixture TF (x-axis) and the TF (y-axis) estimated by the mutation pattern. 2 = 0.999).

[0299] Data shows that filters reduce noise in samples. For example, for lung cancer and melanoma cancer types, pre-filter noise appears at ~2 x 10⁻⁶. -3 For both cancer types, the post-filter noise rate decreased to ~2 x 10⁻⁶. -4 ( Figure 3C The application of a combined base quality (BQ) and mapping quality (MQ) optimized filter with reduced 35X coverage allows for the detection of markers in samples with TFs as low as 1 / 20,000. Here, the red line represents the theoretical (binomial model) expected value, and the empirical measurements (mean and confidence interval of 5 independent replicates) are shown in black. Figure 3D The noise level is represented by a gray area based on the detection distribution with TF = 0. Further, in the computer validation of TF estimation in melanoma samples, for 5 × 10⁻⁶... -5 All of the above TFs have been accurately and specifically estimated. Figure 3E ).

[0300] Analytical validation using labeled synthetic plasma mixtures further confirmed somatic SNVs and somatic cCNVs in all TF > 5 x 10⁻⁶. -5 Especially TF> 5 x 10 -4 The effectiveness in tumor score estimation. Data are shown in 3H and Figure 3I middle.

[0301] Further analysis using synthetic samples demonstrated a strong correlation between the SNV and CNV detection methods (R0). 2 = 83.5%. See also Figure 3J .

[0302] A comparative evaluation of the method disclosed herein, compared with ICHOR, shows that only when TF > 5 x 10 -3 Only then does the ICHOR method provide the correlation between the input tumor score and the output tumor score. Figure 3K ).

[0303] Figure 4A graph showing the SNV detection rate in ctDNA samples obtained from a computer or from control subjects (BB601) or cancer patients (BB1122 or BB1125) using the methods and systems of this disclosure.

[0304] To evaluate a method for detecting residual disease in postoperative NSCLC patients for therapeutic purposes, five early-stage lung cancer samples were collected (Table 1). First, matched tumor and germline DNA (PBMCs) were subjected to whole-genome sequencing (WGS) to generate patient-specific whole-genome SNV profiles. Additionally, plasma samples were collected from subjects before surgery and approximately 14 days after surgical resection. CfDNA was extracted and sequenced using an optimized WGS protocol, and SNV detection in all plasma samples was then analyzed based on the patient-specific whole-genome SNV profile.

[0305] The results show Figure 5A Data shows that in all five preoperative plasma samples from early-stage NSCLC adenocarcinoma cases, whole-genome SNV detection above the noise threshold (…) Figure 5A Furthermore, postoperative plasma testing was recorded in two-fifths of patients, and it was associated with these patients' clinical outcomes (recurrence or death). Figure 5A Specifically, only two patients showed a postoperative TF higher than 5x10. -5 The noise threshold was determined. However, all healthy control samples showed TF levels below the detection threshold. ND indicates undetected. The data showed results consistent with the SNV method in terms of plasma detection and TF correlation.

[0306] To validate this innovative method clinically and facilitate its implementation in clinical practice, the method was applied to 30 patients with early-stage lung cancer (stages I and II). First, matched previously collected tumor and PBMC DNA samples, as well as preoperative and postoperative plasma samples, were subjected to whole-cell sequencing (WGS). A SNV-based detection algorithm was used to quantify preoperative and postoperative progression-free survival (TF). Clinical variables associated with high preoperative or postoperative plasma TF (e.g., disease stage, lymph node involvement, pathological features, and patient demographics) were identified. The impact of positive postoperative plasma samples on progression-free survival in these patients was examined. Data from a representative cohort of 11 patients showed… Figure 5B (Adenocarcinoma against healthy plasma controls) and Figure 5C In adenocarcinoma (with negative controls across patients), it demonstrated >60% sensitivity and >85% specificity. The concordance between sSNV and sCNV detection was shown in… Figure 5D middle.

[0307] Postoperative tumor DNA testing can serve as a prognostic marker for aggressive disease requiring adjuvant therapy. For example, in a postoperative analysis of results from 11 patients (plasma collected 2 weeks post-surgery), a recurrence-free time was found to be inversely correlated with sSNV-based z-score testing. Figure 11H ).

[0308] Example 3A: Orthogonal integration of fragment size features in SNV-based methods

[0309] Due to DNA degradation during blood circulation, the distribution of cfDNA fragments has unique characteristics. Healthy, normal cfDNA samples show... Figure 10A The fragment size distribution is shown. Circulating DNA fragments derived from tumors show shorter fragment sizes compared to "normal" DNA fragments, which are primarily derived from apoptosis of hematopoietic cells (immune cells). Breast tumor cfDNA (red and purple) shows a fragment size shift compared to normal cfDNA samples. Figure 10B Calculations of the centroid (COM) of the first nucleosome (peak around 170 bp) showed a shift towards a lower COM that corresponds linearly to TF. In a mouse model of human tumor xenograft (PDX), circulating DNA from tumor origin (red, compared to humans) was shown to be significantly shorter than circulating DNA from normal origin (black, compared to mice). See also Figure 10C .

[0310] To generate robust models that quantify the probability of a single DNA fragment originating from either a tumor or normal origin, we used a joint Gaussian mixture model (GMM) to characterize the fragment size distribution of circulating DNA. By applying GMM analysis to circulating tumor DNA extracted from our PDX samples, using only circulating DNA aligned to the human genome, the circulating tumor DNA model (red dashed line) was estimated. By applying GMM analysis to circulating DNA in plasma samples from healthy human volunteers, the circulating normal DNA model (grey dashed line) was estimated. The joint log odds ratio (yellow line) was then used to estimate the probability that a particular circulating DNA fragment size originated from either a tumor or normal origin. Data are shown in... Figure 10D middle.

[0311] Based on DNA fragment size distribution and the combined log-ratio of the GMM, patient-specific mutation detection can be used to examine whether these DNA fragments correspond to tumor origin. To increase confidence and reduce batch effect bias, in-patient controls were developed using cross-patient detections. For example, in the specific patients shown below, detected tumor mutations (grey, matching detection results) are present and show a trend of fragment size shift towards lower fragment sizes. Mutations associated with other patients were detected on the same patient samples (red, cross-patient detections). These artifact detections have the same tobacco-feature contextual information pattern but are not true detections. Interestingly, these cross-patient detection results do not show a trend towards lower fragment size shifts, and their fragment size distribution is significantly different from that of true tumor detection results (Wilcoxonrank-sum, P-value 3*10). -9 The GMM combined with the log-odds ratio was used to confirm that patient-specific mutations were of tumor origin (combined log-odds ratio = 0.3), while artificial mutations from the same patient sample were of normal origin (combined log-odds ratio = -0.35). Representative data from three patients are shown in... Figure 10E middle.

[0312] Example 3B: Orthogonal integral of fragment size with CNV marking

[0313] Due to DNA degradation during blood circulation, cfDNA fragment distribution exhibits unique characteristics. Healthy, normal cfDNA samples show variations in fragment size distribution (see above). Figure 10A and Figure 10B Here, in analyzing the centroid (COM) distribution, the calculated COM of the first nucleosome (peak around 170 bp) shows a shift towards a lower COM that corresponds linearly to the TF.

[0314] Comparative analyses of fragment size centroids (COM) between patients may be limited in sensitivity and susceptible to batch effects. The local fragment size COM within a patient can vary due to epigenetic traits or copy number events. Indeed, in amplified regions, there is a local increase in tumor fraction (due to an increased proportion of tumor DNA), thus decreasing the local fragment size COM. Conversely, in deleted regions, there is a local decrease in tumor fraction (due to a decreased proportion of tumor DNA), thus increasing the local fragment size COM.

[0315] Validating this concept in plasma samples from cancer patients identified a significant negative correlation between log2 (log2 > 0.5 = amplification, log2 < -0.5 = deletion) of depth coverage in this segment and the local fragment size centroid (COM). See also Figure 11B Further validation from plasma samples from 12 different cancer patients showed a clear relationship between depth-coverage-based CNV detection and fragment-size centroid (COM)-based CNV detection. Figure 11C This relationship is not obvious in normal (healthy) plasma samples. Figure 11D ).

[0316] Several quantitative features can be extracted from the relationship between depth coverage (Log2) and fragment size (COM) for each sample. More specifically, these include the centroid of the neutral region (Log2=0), the slope of the Log2 / COM relationship, and the R-squared value of the Log2 / COM relationship. 2 These features demonstrate the dynamic response to changes in a patient's tumor score after surgery or during treatment. For example, below is a patient with cancer that progressed during treatment, showing a decrease in COM, absolute slope value, and R... 2 Elevation (Fig. 11E and Fig. 11F). Changes can be detected even in trace amounts of tumor DNA (e.g., in a second patient during treatment).

[0317] Using multiple linear regression or GLM allows the log2 / COM features to be converted into tumor scores for monitoring patients post-surgery and during treatment. Figure 11G For example, monitor the outcomes of patients who have received treatment over a period of 6 weeks (42 days). Estimated tumor score ( Figure 11I ) and normalized CNV score ( Figure 11J The data were tabulated and presented in a comparative bar chart for monitoring residual disease. The data showed that, over time, patient 4 (rather than patients 1–3) responded to treatment, as evidenced by the fact that this patient had a significantly lower eTF at 42 days after treatment compared to the eTF at the time of treatment. Figure 11I Analysis of normalized CNV scores yielded similar conclusions, with positive responses in 4 patients receiving a combination of immunotherapy and chemotherapy, contrasting with 1–3 patients receiving monotherapy (chemotherapy or immunotherapy alone). Treatment response outcomes were confirmed by imaging and long-term clinical follow-up and showed consistency with eTF predictions.

[0318] Example 4: Sensitivity of ctDNA detection using genome-wide integrals of large somatic copy number variations (sCNVs)

[0319] Besides somatic point mutations, cancer genomes are also characterized by fundamental aneuploidy. Through this process, large swath regions of the genome undergo amplification and deletion, generating potentially robust signals for ctDNA detection. This is primarily because WGS coverage depth is a function of DNA content at each site. Other notable examples include shorter ctDNA fragment lengths compared to normal cfDNA, and nucleosome localization information.

[0320] Therefore, WGS offers more advantages than targeted sequencing due to the abundance of orthogonal information sources to enhance detection. To leverage the orthogonal whole-genome signal provided by WGS, a similar approach was developed to utilize differential read depth coverage in large amplified and deleted genomic segments. This read depth detection method aims to integrate millions of small genomic windows to sensitively detect minute depth variations in patient-specific sCNV regions, thereby allowing for sensitive differentiation between low-TF plasma and healthy (TF=0) controls.

[0321] Therefore, this disclosure provides an analytical method to integrate a large number of directional depth coverage skewnesses across large genome CNV regions. Figure 6A Testing on our NSCLC virtual plasma samples, using an integrated whole-genome CNV mode, achieves high detection sensitivity down to TF 1 / 100,000. Figure 6B Furthermore, the comparison between the detection signal and TF shows linearity (R). 2 =1, P-value = 2 * 10 -24 The relationship demonstrates that appropriate modeling can be achieved through a simple dilution model, where differences in local tumor depth coverage (amplification, deletion) are diluted by mixing with the conventional readings. This explicit relationship allows for the calculation of TF from empirical patient measurements. This method, along with the SNV method, will be validated side-by-side in the same patient cohort described above and will be used to construct a joint classification model to synergistically improve sensitivity by integrating these orthogonal signals.

[0322] It is important to note that the method of this invention provides a supplementary and sensitive detection for patients with low SNV mutational burden but high CNV burden. Alternatively, the method described herein can be integrated with SNV-based methods to further improve detection independent of cfDNA abundance. Integration of the two methods on exemplary samples demonstrates the potential for detection of minimal residual disease. Data indicate that even in the absence of a matching tumor sample, genome-wide sSNV scores can provide sensitive MRD detection by applying mutation inference features.

[0323] The methods disclosed herein are not limited to the types of markers exemplified herein. For example, residual disease detection / diagnosis can be performed by analyzing insertions or deletions (indels) in the genomic profile of readings in a manner similar to SNV analysis (exemplified above in Example 2). Similarly, residual disease detection / diagnosis can be performed by analyzing structural variants (SVs) in the genomic profile of readings in a manner similar to CNV analysis (exemplified above in Example 3).

[0324] While many exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain modifications, arrangements, additions, and sub-combinations thereof. Therefore, it is intended that the appended claims and the claims subsequently introduced be construed as including all such modifications, arrangements, additions, and sub-combinations within their true spirit and scope.

[0325] Example 5: Comparative Evaluation

[0326] The systems and methods of this disclosure are compared with those known in the art.

[0327] Current mutational calls do not work in low TF protocols. More specifically, MUTECT does not work at TF below 1%. Suitable alternatives for identifying ctDNA markers include high-coverage targeted sequencing with error suppression (e.g., double-strand sequencing). Phallen et al. provide examples of methods in this field in “Direct Detection of Early Stage Cancers Using Circulating Tumor DNA” (Science Translational Medicine, 9, 203, 2017). The methods described in Phallen and other publications have limited sensitivity at low TF (i.e., almost no detection below 1 / 1000 TF). A second method in this field from the Broad Institute (called ICHOR) has similar limitations. ICHOR (see, Adalsteinsson et al., “Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors,” Nature communications 8.1, 1324, 2017) attempts to integrate CNV information across WGS; however, the ICHOR method is entirely different from this method. Figure 9The comparison results shown demonstrate that the BroadICHOR method exhibits significantly lower sensitivity compared to the present method. In particular, the 100-fold increase in sensitivity obtained using the method and system of this disclosure compared to the ICHOR method is significantly superior and unexpectedly advantageous.

[0328] Therefore, this disclosure relates to the following non-restrictive embodiments:

[0329] Implementation Scheme 1. A method for detecting residual disease in a subject in need, comprising: (A) receiving a subject-specific whole-genome profile of genetic markers from a plurality of genetic markers in a first biological sample of the subject, said biological sample including a tumor sample and optionally a normal cell sample, wherein said genetic marker profile is selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome profile of genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; (C) filtering artificial noise markers from the whole-genome profiles marked in the first and second biological samples, wherein said filtering comprises: (a) based on the probability of detecting noise (P) N (a) classify each SNV or Indel in the profile as signal or noise as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and / or 4) the base quality (BQ) of SNVs or Indels; and / or (b) classify each CNV or SV window in the profile as signal or noise based on 1) the position relative to the centromere, 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and / or 3) the overlap with the cfDNA mask (blacklist); (d) calculate the estimated tumor fraction (eTF) of the first and second biological samples based on one or more integrated mathematical models; and (e) detect residual disease in the subject if the estimated tumor fraction exceeds an empirical threshold calculated using a background noise model.

[0330] Implementation Scheme 2. According to the method of Implementation Scheme 1, wherein step (A) includes receiving a subject-specific whole genome summary of multiple genetic markers, wherein the multiple genetic markers are derived from a biological sample containing a tumor sample and a normal cell sample of the subject.

[0331] Implementation Scheme 3. The method according to any one of Implementation Schemes 1 and 2, wherein the read array comprises a read set covering a specific SNV or indel site, or a read set contained within a specific CNV or SV genomic window.

[0332] Implementation Scheme 4. The method according to any one of Implementation Schemes 1 to 3, wherein the tumor sample comprises a resected tumor or FNA, which includes frozen tissue, OCT-embedded tissue, or FFPE.

[0333] Implementation Scheme 5. The method according to any one of Implementation Schemes 1 to 4, wherein the normal sample includes peripheral blood mononuclear cells (PMBC) or saliva or skin samples.

[0334] Implementation Scheme 6. The method of any one of Implementation Schemes 1-5, wherein multiple genetic markers are received via whole-genome sequencing of the subject's biological sample.

[0335] Implementation Scheme 7. The method according to any one of Implementation Schemes 1 to 6, wherein the genetic marker profile of multiple genetic markers from the first biological sample of the subject includes a high mutation rate and / or a high number of CNVs or SVs.

[0336] Implementation Scheme 8. The method according to Implementation Scheme 7, wherein the high mutation rate includes a mutation rate of at least one somatic single nucleotide polymorphism or indel per megabase pair, and wherein the high copy number variation includes somatic CNVs or SVs with a cumulative size of at least 5 megabase pairs.

[0337] Implementation Scheme 9. The method according to any one of Implementation Schemes 1 to 8, wherein the background noise model includes measuring the detection error rate in a normal healthy sample and converting the error rate into a basic noise eTF estimation model.

[0338] Implementation Scheme 10. According to the method of Implementation Scheme 9, wherein the threshold calculated by the eTF estimation model is in the range of 10. -4 Up to 10 -6 between.

[0339] Implementation Scheme 11. The method according to any one of Implementation Schemes 1 to 11, wherein step (A) includes receiving a subject-specific whole-genome summary of somatic genetic markers from a plurality of genetic markers from a subject biological sample, said biological sample including a tumor sample and a normal cell sample; and step (B) includes subsequently detecting the subject-specific whole-genome summary of the genetic markers in a second biological sample containing a sample of the subject's plasma to generate a tumor-associated whole-genome representative of the genetic markers in the patient's plasma that is updated over time.

[0340] Implementation Scheme 12. The method according to any one of Implementation Schemes 1-11, wherein the normal cell sample includes PMBC, saliva sample, hair sample or skin sample.

[0341] Implementation Scheme 13. The method according to any one of Implementation Schemes 1-12, wherein the subject is a human being, and the second biological sample of the subject is a biological material selected from the group consisting of: blood, cerebrospinal fluid, pleural fluid, ocular fluid, feces, urine, or a combination thereof.

[0342] Implementation Scheme 14. A method for quantitatively estimating the minimum residual disease burden of a patient during treatment, observation, or follow-up, comprising: (A) receiving a subject-specific whole-genome profile of genetic markers from a plurality of genetic markers in a first biological sample of a subject, said biological sample including a tumor sample and optionally a normal cell sample, wherein said genetic marker profile is selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variants, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome profile of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; (C) filtering artificial noise markers from the whole-genome profiles of the markers in the first and second biological samples, wherein said filtering comprises: (a) based on the probability of detecting noise (P N (a) classify each SNV or Indel in the profile as signal or noise as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and / or 4) the base quality (BQ) of SNVs or Indels; and / or (b) classify each CNV or SV window in the profile as signal or noise based on 1) the position relative to the centromere, 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and / or 3) the overlap with the cfDNA mask (blacklist); (d) calculate the estimated tumor fraction (eTF) of the first and second biological samples based on one or more integrated mathematical models; and (e) detect residual disease in the subject if the estimated tumor fraction exceeds an empirical threshold calculated using a background noise model.

[0343] Implementation Scheme 15. The method according to Implementation Scheme 14, wherein (E) further includes detecting residual disease in the subject after resection; and detecting residual disease during or after treatment; detecting residual disease to monitor treatment effectiveness; detecting residual disease to monitor cancer recurrence or recurrence; or a combination thereof.

[0344] Implementation Scheme 16. According to the method of Implementation Scheme 15, the resection includes lymph node biopsy; head or neck surgery; uterine or endometrial biopsy; bladder biopsy; mastectomy; prostatectomy; removal of skin lesions; small bowel resection; gastrectomy; thoracotomy; adrenalectomy; colectomy; oophorectomy; thyroidectomy; hysterectomy; tongue resection; or colon polyp removal.

[0345] Implementation Scheme 17. The method according to Implementation Scheme 15, wherein the therapy includes chemotherapy, immunotherapy, targeted therapy, radiotherapy, or a combination thereof.

[0346] Implementation Scheme 18. The method of any one of Implementation Schemes 14 to 17, wherein the BQ, MQ and fragment size parameters of the label are optimized using ROC curves.

[0347] Implementation Scheme 19. The method according to any one of Implementation Schemes 14 to 18, including the use of a combination of base quality mapping quality (BQ MQ) parameters.

[0348] Implementation Scheme 20. The method according to any one of Implementation Schemes 14-19 further includes receiving multiple genetic markers from a biological sample of a subject, said biological sample including tumor samples and normal cell samples, and generating a subject-specific whole-genome summary of the genetic markers from the received multiple genetic markers.

[0349] Implementation Scheme 21. The method according to any one of Implementation Schemes 14 to 20 further includes detecting a subject-specific whole-genome profile of genetic markers in a third biological sample of the subject for comparison with a subject-specific whole-genome profile of genetic markers generated in a first biological sample of the subject.

[0350] Implementation Scheme 22. The method according to Implementation Scheme 21, wherein the third biological sample is a plasma sample of a subject, said subject's plasma sample being obtained to generate a time-updated representative of tumor whole-genome genetic markers in the patient's plasma.

[0351] Implementation Scheme 23. The method according to any one of Implementation Schemes 14 to 22 further includes empirically determining a background noise threshold, wherein a tumor score above the background noise threshold provides a quantitative estimate of the tumor burden.

[0352] Implementation Scheme 24. The method according to any one of Implementation Schemes 14 to 23, wherein tumor fractions below a noise threshold are considered undetected (ND).

[0353] Implementation Scheme 25. The method according to any one of Implementation Schemes 14 to 24, wherein the detection includes quantitative monitoring over time.

[0354] Implementation Scheme 26. The method according to any one of Implementation Schemes 14 to 25, wherein the tumor is brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, melanoma, osteosarcoma, or a solid tumor that is either heterogeneous or homogeneous.

[0355] Implementation Scheme 27. The method according to any one of Implementation Schemes 14 to 26, wherein the tumor is lung adenocarcinoma, ductal adenocarcinoma, non-small cell lung adenocarcinoma (NSCLC LUAD), cutaneous melanoma, urothelial carcinoma, or osteosarcoma.

[0356] Implementation Scheme 28. The method according to any one of Implementation Schemes 14 to 27, wherein the calculation step further comprises: calculating the eTF of SNV or indel markers by means of an integral probability model, wherein the probability model comprises 1) an integral signal of plasma SNV or indel detection, 2) a process quality metric including an estimated genome coverage and sequencing noise model, and / or 3) a patient-specific parameter including mutational burden (N); and / or calculating the eTF of CNV or SV markers by means of a probabilistic dilution model, wherein the probabilistic dilution model comprises 1) integrating the skewed directional coverage depth between plasma and normal patient samples based on the directionality of tumor CNV or SV, wherein copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative coverage depth of the skewed skew between tumor and normal (PBMC) patient samples; and / or 3) finding the dilution ratio between the above signals.

[0357] Implementation Scheme 29. A system for detecting residual disease in subjects in need, comprising, (A) an analysis unit configured and arranged to filter artificial noise markers from a labeled whole-genome profile, wherein the labeled whole-genome profile is generated from multiple genetic markers from biological samples from the subject, including tumor samples and normal cell samples, wherein the genetic marker profile is selected from the group consisting of single nucleotide variants (SNVs), indels, copy number variants, SVs, and combinations thereof, the analysis unit further comprising detecting a subject-specific whole-genome profile of genetic markers in a second biological sample to generate a representative of tumor whole-genome genetic markers in the second sample, the analysis unit further comprising a classification engine, wherein the classification engine: (a) is based on the probability (P) of detecting noise. N(a) As a function of 1) the mapping quality (MQ) of the read group containing SNVs or Indels, 2) the fragment size length of the read group containing SNVs or Indels, 3) the consistency test within the read repeat family containing specific SNVs, and / or 4) the base quality (BQ) of SNVs or Indels, each SNV in the summary is statistically classified as signal or noise; and / or (b) based on 1) the position relative to the centromere; 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, 3) the representativeness of CNVs or SV windows in the cfDNA data, each CNV or SV window in the summary is statistically classified as signal or noise; (B) a calculation unit configured and arranged to calculate the estimated tumor score (eTF) of the sample based on one or more integrated mathematical models; (C) a display unit that outputs a residual disease profile of the subject based on the estimated tumor score, wherein if the estimated tumor score exceeds an empirical threshold calculated by the background noise model, the subject's residual disease is output in the residual disease profile.

[0358] Implementation Scheme 30. A system or method according to any of the foregoing embodiments, wherein the computing unit is further configured and arranged to: calculate the eTF of SNV or Indel markers by means of an integral probability model, wherein the probability model includes 1) an integral signal of plasma SNV or Indel detection, 2) a process quality metric including an estimated genome coverage and sequencing noise model, and / or 3) a patient-specific parameter including mutational burden (N); and / or calculate the eTF of CNV or SV markers by means of a probabilistic mixture model, wherein the probabilistic dilution model includes 1) integrating the skewed directional coverage depth between plasma and normal patient samples based on the directionality of tumor CNV or SV, wherein copy number amplification is positively skewed and copy number deletion is negatively skewed; 2) integrating the cumulative coverage depth of the skewed skew between tumor and normal patient samples; and / or 3) finding the dilution ratio between the above signals.

[0359] Implementation Scheme 31. The system or method according to Implementation Scheme 30, wherein the computing unit (B) includes a processor configured to execute computer-readable instructions that, when executed, estimate the sample tumor fraction (eTF) based on one or more of the following integrated mathematical models: (1) eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov), where M is the number of tumor-specific SNV profile detections in the patient's plasma sample, σ is a measurement of the error rate estimated empirically, R is the total number of unique readings in the SNV profile target region (ROI), N is the tumor mutation burden, and cov is the average number of unique readings at each site in the SNV profile ROI; and / or (2) eTF[CNV]=(sum_{i}[(P(i)-N(i))*sign[T(i)-N(i)]) ]-E(sigma)) / (sum_{i}[abs(T(i)-N(i))]-E(σ)), where P is the median depth coverage value in the genomic window indexed by {i} representing plasma depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; and {i} is a discrete index used to count all genomic windows covering patient tumor-specific amplified and deleted genomic regions.

[0360] Implementation Scheme 32. A computer-readable medium including computer-executable instructions, which, when executed by a processor, cause the processor to perform a method or set of steps for detecting residual disease, the method or steps comprising: (A) receiving a subject-specific whole-genome summary of genetic markers from a plurality of genetic markers in a biological sample of a subject, the biological sample including a tumor sample and optionally a normal cell sample, wherein the genetic marker summary is selected from the group consisting of: single nucleotide variants (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; (B) detecting the subject-specific whole-genome summary of the genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the second sample; (C) filtering artificial noise markers from the whole-genome summary of the markers by: based on the probability of detecting noise (P) N(a) as a function of 1) the mapping quality (MQ) of the read group containing SNVs, 2) the fragment size length of the read group containing SNVs, 3) the consistency test within the read repeat family containing SNVs or Indels, and 4) the base quality (BQ) of SNVs or Indels, by statistically classifying each SNV or Indel in the profile as signal or noise; or based on 1) the position relative to the centromere; 2) the mapping quality (MQ) of the read group containing CNVs or SV windows, and 3) the overlap with the cfDNA mask (blacklist), by statistically classifying each CNV or SV window in the profile as signal or noise; (d) calculating the estimated tumor fraction (eTF) of the biological sample based on one or more integrated mathematical models; and (e) diagnosing residual disease in the subject based on the estimated tumor fraction and an empirical threshold calculated by a background noise model.

[0361] Implementation Scheme 33. A method for detecting minimal residual disease in a subject, comprising (A) receiving a genome-wide summary of readings from genetic data sequenced from multiple biological samples received from the subject, said multiple biological samples including tumor samples, normal samples, and plasma samples; (B) performing mutation calls on tumor and peripheral blood mononuclear cell (PBMC) samples from the subject, said mutation calls comprising MUTECT, LOFREQ, and / or STRELKA mutation calls to generate subject-specific readings of somatic SNVs (sSNVs) or indels as a personalized reference set; (C) collecting and filtering readings from subject-specific somatic SNVs (sSNVs) or indels, said collection and filtering comprising (1) removing low-mapped quality readings (e.g., <29, ROC optimized); (2) establishing replication families (multiple PCRs representing the same DNA fragment). (3) Remove low-base-quality reads (e.g., <21, ROC optimized); and (4) Remove high fragment size reads (e.g., >160, ROC optimized); (D) Calculate the number of subject-specific mutation sites with at least one supporting read (in the filtered set), said subject-specific mutation sites having identical substitutions to those in the tumor; (F) Estimate the tumor fraction of SNV based on the mathematical model eTF[SNV]=1-[1-(ME(σ)*R) / N]^(1 / cov)…(Equation 1), where M is the number of tumor-specific profiles detected in the patient sample, σ is a measure of noise estimated empirically, R is the total number of unique reads within the target region (ROI), N is the tumor mutational burden, and cov is the average number of unique reads per site in the ROI; (G) Transform eTF [SNV] is compared to a detection threshold, which includes an empirically measured baseline noise TF estimate from healthy samples, where an eTF [SNV] above the threshold level (e.g., 2 standard deviations of the noise TF distribution (FPR < 2.5%)) indicates a positive detection; and (K) detects residual disease in subjects based on eTF estimates exceeding the detection threshold level.

[0362] Implementation Scheme 34. A method for detecting minimal residual disease in a subject, comprising (A) receiving a whole-genome summary of readings from genetic data sequenced from multiple biological samples received from the subject, said multiple biological samples including tumor samples, normal samples, and plasma samples; (B) performing CNV or SV recall on tumor and peripheral blood mononuclear cell (PBMC) samples from the subject to generate multiple CNV or SV fragments or reference segmentations of SVs exceeding a threshold length (e.g., >2 Mbp, preferably >5 Mbp), and annotating the directionality of the segments, wherein amplification is positively annotated and deletion is negatively annotated; (C) collecting single-bp depth coverage information of plasma, tumor, and PBMC samples covering the target region (ROI) of the patient-specific CNV or SV segmentation; (D) converting the patient-specific CNV or SV segmentation into a single-bp depth coverage information. (E) Divide the NV or SV partitioned ROI into 500bp windows and calculate the median (artificially suppressed) for each window across all samples; (G) Generate normalized depth coverage information for all 500bp windows using: (a) robust z-score normalization for each sample; and / or (2) robust principal component analysis (RPCA); (F) Filter windows from patient-specific partitions, wherein the filtering includes: (1) removing low-map quality readings (e.g., <29, ROC optimized); and / or (2) removing centromere regions (e.g., removing windows with normalized normal values ​​greater than 10); and / or (3) removing unrepresented regions in cfDNA (e.g., removing windows not included in a cfDNA representation mask consisting of multiple cfDNA samples); (G) Using a mathematical model sum i [(P(i)-N(i))*sign[T(i)-N(i)]]-E(σ)…(Equation 2) integrates the skewed directional coverage depth between plasma and normal (PBMC) patient samples, where P is the median depth coverage value in the genomic window indexed by {i} representing plasma depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; E(sigma) is a measure of the error rate estimated empirically; T is the median depth value in the genomic window indexed by {i} representing tumor depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; and N is the median depth value in the genomic window indexed by {i} representing normal depth coverage, normalized by a robust z-score method or robust PCA compared to the normal sample cohort; (H) uses the mathematical model sum i[abs(T(i)-N(i))]-E(σ)) …(Equation 3), integrates the skewed cumulative coverage depth between tumor and normal (PBMC) patient samples, where T, N, and E(σ) are provided above; (I) calculates the dilution rate between the directional depth coverage (G) and the cumulative depth coverage (H), which corresponds to the estimated tumor fraction (eTF[CNV]) of CNV or SV = (sum i [(P(i)-N(i))*sign[T(i)-N(i)]]-E(σ)) / (sum i [abs(T(i)-N(i))]-E(σ))…(Equation 4); (J) compares the eTF [CNV] with a detection threshold that includes an empirically measured baseline noise TF estimate from healthy samples, where an eTF [CNV] above the threshold level (e.g., 2 standard deviations of the noise TF distribution (FPR < 2.5%)) indicates a positive detection; and (K) detects residual disease in subjects based on eTF estimates exceeding the detection threshold level.

[0363] Implementation Scheme 35. A method for detecting residual disease in a subject in need, comprising: (A) receiving a first subject-specific whole-genome profile of genetic marker-related readings from a first biological sample of the subject, the first biological sample comprising a baseline sample and a normal cell sample, wherein each of the first reading profiles contains a reading of a single base pair length, and wherein the baseline sample comprises a tumor sample or a plasma sample; (B) filtering artificial sites from the first reading profile, wherein the filtering comprises removing duplicate sites generated on a reference healthy sample cohort from the first reading profile of the genetic marker, and / or identifying germline mutations in peripheral blood mononuclear cells of the normal cell sample and removing the germline mutations from the first reading profile of the genetic marker; (C) detecting readings of a second subject-specific whole-genome profile of the genetic marker in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic marker in the second sample; (D) filtering noise from the first and second reading whole-genome profiles using at least one error suppression scheme to generate a tumor-associated whole-genome representative for the first reading whole-genome profile. A first filtered read set and a second filtered read set for a second read genome-wide profile, wherein at least one error suppression protocol comprises: (a) calculating the probability that any single nucleotide variant in the first and second profiles is an artificial mutation and removing said mutation, wherein said probability is calculated as a function of a feature selected from a group consisting of map quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof; and / or (b) removing artificial mutations using a dissimilarity test between independent repeats of the same DNA fragment generated by polymerase chain reaction or sequencing processing, and / or repeat consistency (wherein artificial mutations are identified and removed when consistency is lacking in most of a given repeat family); (e) calculating an estimated tumor score (eTF) for the first and second biological samples using the first and second filtered read sets by applying a background noise model to one or more integrated mathematical models; and (f) detecting residual disease in the subject if the estimated tumor score in the second biological sample exceeds an empirical threshold.

[0364] Implementation Scheme 36. A method for detecting residual disease in subjects in need, comprising: (A) receiving a first subject-specific whole-genome summary of readings associated with genetic markers from a first biological sample of the subject, the first biological sample comprising a baseline sample, wherein each of the first reading summaries contains copy number variations (CNVs) or structural variations (SVs), and wherein the baseline sample comprises a tumor sample or a plasma sample; (B) receiving a second subject-specific whole-genome summary of readings associated with genetic markers from a second biological sample of the subject, the second biological sample comprising a peripheral blood mononuclear cell (PBMC) sample, wherein each of the second summaries of genetic markers comprises a CNV or an SV; (C) filtering artificial sites from the first and second reading summaries, wherein the filtering comprises removing duplicate sites generated on a reference healthy sample cohort from the first and second reading summaries; identifying CNVs / SVs shared between the first and second summaries as germline mutations, and filtering from the first and second reading summaries (D) Remove the mutations; detect readings of a third subject-specific whole-genome profile of genetic markers in a third biological sample from the subject to generate a tumor-associated whole-genome representative of the genetic markers in the third sample; (E) normalize each of the first, second, and third reading profiles to generate a first filtered set of readings for the first reading whole-genome profile, a second filtered set of readings for the second reading whole-genome profile, and a third filtered set of readings for the third reading whole-genome profile; (F) compute an estimated tumor score (eTF) of the third biological sample using the third filtered set of readings by applying a background noise model to one or more integrated mathematical models, the one or more models using the first filtered set of readings to generate a first eTF, and / or the one or more models using the second filtered set of readings to generate a second eTF; and (G) detect residual disease in the subject if the estimated tumor score in the third biological sample exceeds an empirical threshold.

[0365] Implementation Scheme 37. A system for detecting residual disease in subjects in need, comprising: an analysis unit including a pre-filter engine configured and arranged to receive a first subject-specific whole-genome profile of genetic marker-related readings from a first biological sample of a subject, the first biological sample comprising a baseline sample and a normal sample, wherein each of the first reading profiles contains a reading of a single base pair length, and wherein the baseline sample comprises a tumor sample or a plasma sample; and filtering artificial sites from the first reading profile, wherein the filtering includes removing duplicate sites generated on a reference healthy sample cohort from the first reading profile of the genetic marker, and / or identifying germline mutations in peripheral blood mononuclear cells of a normal cell sample and removing the germline mutations from the first reading profile of the genetic marker; and a correction engine configured and arranged to receive readings from a second subject-specific whole-genome profile of the genetic marker in a second biological sample of the subject to generate a tumor-associated whole-genome representation of the genetic marker in the second sample; and filtering noise from the first and second reading whole-genome profiles using at least one error suppression scheme. The method generates a first filtered read set for a first read genome profile and a second filtered read set for a second read genome profile, wherein at least one error suppression scheme comprises: (a) calculating the probability that any single nucleotide variant in the first and second profiles is an artificial mutation and removing said mutation, wherein said probability is calculated as a function of a feature selected from a group consisting of mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof; and / or (b) removing artificial mutations using a dissimilarity test between independent repeats of the same DNA fragment generated by polymerase chain reaction or sequencing processing, and / or repeat consistency (wherein artificial mutations are identified and removed when consistency is lacking in most of a given repeat family); and a computational unit configured and arranged to compute an estimated tumor fraction (eTF) for the first and second biological samples using the first and second filtered read sets by applying a background noise model to one or more integrated mathematical models; and detecting residual disease in the subject if the estimated tumor fraction in the second biological sample exceeds an empirical threshold.

[0366] Implementation Scheme 38. A system for detecting residual disease in subjects in need, comprising: a pre-filter engine configured and arranged to receive a first subject-specific whole-genome summary of readings associated with genetic markers from a first biological sample of the subject, the first biological sample comprising a baseline sample, wherein each of the first reading summaries contains readings of single base pair lengths, and wherein the baseline sample comprises a tumor sample or a plasma sample; receiving a second subject-specific whole-genome summary of readings associated with genetic markers from a second biological sample of the subject, the second biological sample comprising a peripheral blood mononuclear cell (PBMC) sample, wherein each of the second summaries of genetic markers comprises copy number variations (CNVs); and filtering artificial sites from the first and second reading summaries, wherein the filtering comprises removing duplicate sites generated on a reference healthy sample cohort from the first and second reading summaries; identifying CNVs shared between the first and second summaries as germline mutations, and removing said mutations from the first and second reading summaries; The system includes a correction engine configured and arranged to receive readings from a third subject-specific whole-genome profile of genetic markers in a second biological sample of the subject to generate a tumor-associated whole-genome representative of the genetic markers in the third sample; and to normalize each of the first, second, and third reading profiles to generate a first filtered set of readings for the first reading whole-genome profile, a second filtered set of readings for the second reading whole-genome profile, and a third filtered set of readings for the third reading whole-genome profile; and a computation unit configured and arranged to compute an estimated tumor fraction (eTF) of the third biological sample using the third filtered set of readings by applying a background noise model to one or more integrated mathematical models, the one or more models using the first filtered set of readings to generate a first eTF, and / or the one or more models using the second filtered set of readings to generate a second eTF; and to detect residual disease in the subject if the estimated tumor fraction in the third biological sample exceeds an empirical threshold.

[0367] Implementation Scheme 39. The method of Implementation Scheme 35, wherein the marker comprises a single nucleotide variant (SNV) or an insertion / deletion (indel); preferably an SNV.

[0368] Implementation Scheme 40. The methods of Implementation Schemes 35 and 39, wherein filtering duplicate sites generated on a reference healthy sample cohort includes generating a normal group (PON) blacklist or mask.

[0369] Implementation Scheme 41. The method of any one of Implementation Schemes 35 and 39 to 40, wherein the normal sample contains peripheral blood mononuclear cells (PBMCs), and germline mutations in the PBMCs are removed in the artificial site filtering step (B).

[0370] Implementation Scheme 42. The method of any one of Implementation Schemes 35 and 39 to 41, wherein in step (A), the first biological sample comprises a plasma sample, said plasma sample being obtained from the subject before surgery or treatment.

[0371] Implementation Scheme 43. The method of any one of Implementation Schemes 35 and 39 to 42, wherein in step (C), the second biological sample comprises a plasma sample obtained from the same subject after treatment or surgery.

[0372] Implementation Scheme 44. The method of any one of Implementation Schemes 35 and 39 to 43, wherein step (D) includes employing a machine learning (ML) algorithm, such as a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), discriminant analysis, nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof; preferably a support vector machine (SVM), to filter artificial noise.

[0373] Implementation Scheme 45. The method of any one of Implementation Schemes 35 and 39 to 44, wherein in step (D), the second error suppression step includes correcting artificial mutations generated by PCR or sequencing by comparing independent repeats of the same original nucleic acid fragment.

[0374] Implementation Scheme 46. The method of Implementation Scheme 45, wherein in step (D), the second error suppression step includes correcting artificial mutations generated by sequencing the paired end 150 bp, resulting in overlapping paired reads (R1 and R2), and the inconsistency between the R1 and R2 pairs is corrected back to the corresponding reference genome.

[0375] Implementation Scheme 47. The method of any one of Implementation Schemes 35 and 39 to 46, wherein in step (D), the second error suppression step includes correcting repeat families generated during sequencing and / or PCR amplification, wherein the repeat families are identified by 5' and 3' similarity and alignment positions, and wherein each repeat family is used to check the consistency of specific mutations across independent repeats, thereby correcting artificial mutations that do not show consistency in most repeat families.

[0376] Implementation Scheme 48. The method of any one of Implementation Schemes 35 and 39 to 47, wherein in step (E), a mathematical model integrates the relationship between coverage, mutation burden, number of detected mutations, and tumor score (TF).

[0377] Implementation Scheme 49. The method according to any one of Implementation Schemes 35 and 39 to 48, wherein in step (E), the background noise calculation includes using patient-specific mutation features to calculate (1) the expected noise distribution in a healthy plasma sample cohort (normal group or PON) or (2) the expected noise distribution in other patients (cross-patient analysis).

[0378] Implementation scheme 50. The method of implementation scheme 49, wherein the background noise model provides the mean and standard deviation (μ, σ) of the estimated artificial mutation detection rate.

[0379] Implementation scheme 51. The method of any one of implementation schemes 35 to 50, further comprising orthogonal integration of a second feature including a fragment size offset.

[0380] Implementation scheme 52. The method of implementation scheme 51, wherein statistical methods such as significance tests or Gaussian mixture models (GMM) are used to analyze intra-patient fragment size shifts in a list of tumor-specific markers and random markers.

[0381] Implementation scheme 53. The method of implementation scheme 36, wherein the markers include copy number variation (CNV).

[0382] Implementation scheme 54. The method of any one of implementation schemes 36 and 37, wherein filtering duplicate sites generated on a reference healthy sample cohort includes generating a normal group (PON) blacklist or mask.

[0383] Implementation scheme 55. The method of any one of implementation schemes 36 and 53 to 54, wherein phylogenetic events in the PBMC are removed in the artificial site filtering step (C).

[0384] Implementation Scheme 56. The method of any one of Implementation Schemes 36 and 53 to 55, wherein in step (A), the first biological sample includes a plasma sample obtained from the subject before surgery or treatment, and the second biological sample includes PBMCs obtained from the same subject before surgery or treatment.

[0385] Implementation Scheme 57. The method of any one of Implementation Schemes 36 and 53 to 56, wherein in step (C), the third biological sample comprises a plasma sample obtained from the same subject after treatment or surgery.

[0386] Implementation Scheme 58. The method of any one of Implementation Schemes 36 and 53 to 57, wherein step (C) includes binning the target region (ROI) containing all genomic segments of somatic tumor CNV (sT_CNV) and somatic PBMC CNV (sP_CNV) into windows of ≥500 bp; estimating the depth coverage (read counts) in each window from subsequent plasma samples; and calculating the median depth coverage in each window.

[0387] Implementation Scheme 59. The method of any one of Implementation Schemes 36 and 53 to 58, wherein the subsequent plasma sample is obtained after surgery, during treatment, or at follow-up.

[0388] Implementation Scheme 60. The method of any one of Implementation Schemes 36 and 53 to 59, wherein the normalization step includes normalizing the depth coverage value to correct for GC content and mappability bias by performing two LOESS regression curve fittings on the bin-by-bin GC score and the mappability score.

[0389] Implementation Scheme 61. The method of any one of Implementation Schemes 36 and 53 to 60, wherein the normalization step includes batch effect correction using robust z-score normalization, the robust z-score normalization being applied to each sample respectively.

[0390] Implementation Scheme 62. The method described in Implementation Scheme 62, wherein zscore normalization includes calculating the median and median absolute deviation (MAD) based on the neutral region of each sample, and normalizing all CNV bins by subtracting the median and dividing the differential by the MAD.

[0391] Implementation Scheme 63. The method of any one of Implementation Schemes 36 and 53 to 62, wherein step (E) includes calculating the depth coverage skewness and / or fragment size centroid (COM) skewness of the third sample compared to the normal group (PON) healthy plasma sample.

[0392] Implementation Scheme 64. The method of any one of Implementation Schemes 36 and 53 to 63, wherein step (E) includes calculating a tumor score by examining a linear dilution ratio between a cumulative signal detected in a subsequent plasma sample and a cumulative signal detected in a tumor sample.

[0393] Implementation Scheme 65. The method of any one of Implementation Schemes 36 and 53 to 64, wherein in step (F), the background noise calculation includes using patient-specific CNV / SV features to calculate (1) the expected noise distribution in a healthy plasma sample cohort (normal group or PON) or (2) the expected noise distribution in other patients (cross-patient analysis).

[0394] Implementation Scheme 66. The method of Implementation Scheme 65, wherein the background noise model provides the mean and standard deviation (μ, σ) of the estimated artificial SNV / SV detection rate.

[0395] Implementation Scheme 67. The method of any one of Implementation Schemes 36 and 53 to 66, further comprising orthogonal integration of a second feature including a fragment size offset.

[0396] Implementation Scheme 68. The method of Implementation Scheme 67, wherein the correlation between depth coverage skewness and fragment size skewness in CNV segments is analyzed in order to infer tumor scores, for example, using a generalized linear model (GLM).

[0397] For convenience, certain terms used in the specification, embodiments, and claims are collected herein. Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0398] Throughout this disclosure, various patents, patent applications, and publications are cited. The entire contents of these patents, patent applications, included information (e.g., information identified by PubMed, PUBCHEM, NCBI, UNIPROT, or EBI Registry Numbers), and publications are incorporated herein by reference to provide a more comprehensive description of the level of technology known to those skilled in the art as of the date of this disclosure. In the event of any inconsistency between the cited patents, patent applications, and publications and this disclosure, this disclosure shall prevail.

Claims

1. A method comprising: receiving a genetic marker subject-specific whole genome profile from a first biological sample of a subject, wherein the genetic marker subject-specific whole genome profile is selected from the group consisting of: single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variants (SVs), and combinations thereof; detecting a genetic marker subject-specific whole genome profile in a second biological sample of the subject to generate a tumor-related whole genome representation of genetic markers in the second biological sample; filtering out artifact markers from the genetic marker subject-specific whole genome profiles in the first and second biological samples, wherein filtering comprises classifying each SNV, Indel, CNV, and / or SV in the subject-specific whole genome profile as signal or noise; and detecting whether residual disease is present in the subject based on one or more of model integration coverage, mutational burden, detected mutations, and tumor fraction between the genetic marker subject-specific whole genome profiles in the first and second biological samples.

2. The method of claim 1, wherein the biological samples from which genetic marker subject-specific whole genome profiles are received comprise a tumor sample and a normal cell sample of the subject.

3. The method of claim 1, wherein the genetic marker subject-specific whole genome profile comprises SNVs and / or Indels, wherein filtering comprises classifying each SNV and / or Indel in the subject-specific whole genome profile as signal or noise.

4. The method of claim 3, wherein the genetic marker subject-specific whole genome profile has a mutation rate of at least 1 somatic SNV per megabase pair and / or at least 1 somatic indel per base pair, respectively.

5. The method of claim 1, wherein the genetic marker subject-specific whole genome profile comprises CNVs and / or SVs, wherein filtering comprises classifying each CNV and / or SV in the subject-specific whole genome profile as signal or noise.

6. The method of claim 5, wherein the genetic marker subject-specific whole genome profile has somatic CNVs and / or SVs with a cumulative size of at least 5 megabase pairs.

7. The method of claim 1, wherein the genetic marker subject-specific whole genome profile is received from a biological sample by whole genome sequencing.

8. The method of claim 1, wherein detecting whether residual disease is present in the subject comprises: calculating an estimated tumor fraction (eTF) for the first and second biological samples based on one or more of model integration coverage, mutational burden, detected mutations, and tumor fraction between the genetic marker subject-specific whole genome profiles in the first and second biological samples; and detecting residual disease in the subject if the eTF exceeds an empirical threshold.

9. The method of claim 8, wherein the empirical threshold is calculated using a background noise model based on a rate of detection errors in normal healthy samples.

10. The method of claim 8, wherein the empirical threshold is between 10 -4 and 10 -6 .

11. The method of claim 1, wherein the subject-specific whole genome profile of genetic markers received from the subject comprises a subject-specific whole genome profile of somatic genetic markers received from a biological sample of the subject, wherein the biological sample comprises a tumor sample and a normal cell sample.

12. The method of claim 11, wherein the second biological sample from which the subject-specific whole genome profile of genetic markers is detected comprises a plasma sample of the subject; wherein detecting the subject-specific whole genome profile of genetic markers in the second biological sample comprises generating a time-updated tumor-related whole genome representation of genetic markers in a plasma sample of the subject, wherein the subject-specific whole genome profile of genetic markers in the second biological sample comprises a time-updated tumor-related whole genome representation of genetic markers in a plasma sample of the subject.

13. The method of claim 1, wherein the biological sample from which the subject-specific whole genome profile of genetic markers is received comprises a tumor sample and a normal cell sample of the subject, wherein the normal cell sample comprises PMBCs, a saliva sample, a hair sample, or a skin sample.

14. The method of claim 1, wherein the subject is a human, wherein the second biological sample is a biological material selected from the group consisting of blood, cerebrospinal fluid, pleural fluid, ocular fluid, stool, urine, and combinations thereof.

15. The method of claim 1, further comprising: determining whether to administer an adjuvant therapy to the subject based on the detected residual disease.

16. The method of claim 15, wherein determining whether to administer an adjuvant therapy comprises predicting an outcome of the adjuvant therapy based on the detected residual disease.

17. The method of claim 1, wherein filtering out artificial noise markers comprises: based on the probability (P N ) of detecting noise as a function of 1) mapping quality (MQ) of the read set containing the SNV, 2) fragment size length of the read set containing the SNV, 3) concordance testing within the read duplication family containing the SNV and / or Indel, and / or 4) base quality (BQ) of the SNV and / or Indel, classifying each SNV and / or Indel in the subject-specific whole genome summary as signal or noise.

18. The method of claim 1, wherein filtering out artificial noise markers comprises: classifying each CNV and / or SV in the subject-specific whole genome profile as signal or noise based on 1) position relative to centromere, 2) mapping quality (MQ) of read sets comprising CNV or SV windows, and / or 3) overlap with cfDNA mask.

19. The method of claim 1, wherein detecting whether residual disease is present in the subject comprises: determining a z score based on one or more of model integration coverage, mutational burden, detected mutations, and tumor fraction between the subject-specific whole genome profiles of genetic markers in the first and second biological samples; and detecting whether residual disease is present in the subject based on the z score.

20. A system comprising: a memory storing computer-readable instructions; and a processor connected to the memory and configured to perform the method of any one of claims 1-19.