Ultrasound-sensitive detection of circulating tumor DNA by genome-wide integration
A machine learning-based method using a CNN filters sequencing noise to accurately detect low-abundance somatic mutations in ctDNA, addressing the limitations of current methods and enhancing early cancer detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-03-10
AI Technical Summary
Current methods for detecting somatic mutations in low-abundance circulating tumor DNA (ctDNA) are limited by low sensitivity and accuracy, especially in early-stage cancer, due to the challenges of distinguishing true mutations from sequencing errors and the limited amount of input material.
A machine learning-based approach using a convolutional neural network (CNN) to filter sequencing noise and distinguish between cancer-associated mutational signatures and sequencing errors by analyzing features such as base quality, mapping quality, and allele fraction, enabling accurate detection of low-abundance markers.
The method significantly improves the sensitivity and accuracy of detecting early-stage cancers by filtering sequencing noise and identifying true mutations, even with limited input material, achieving superior performance compared to existing techniques.
Smart Images

Figure 2026041712000005 
Figure 2026041712000006 
Figure 2026041712000007
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 62 / 636,135, filed February 27, 2018, the entire contents of which are incorporated herein by reference. FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate generally to the field of medical diagnostics. In particular, aspects of the present disclosure relate to compositions, methods, and systems for tumor detection and diagnosis. [Background technology]
[0002] The significant burden that cancers, such as solid tumors of the lung, breast, prostate, liver, and brain, pose on human health is well documented in the medical literature. Many subjects are diagnosed with advanced neoplastic disease, which is associated with poor outcomes. Recently, computed tomography (CT) has been found to improve early detection and has been used by a US task force to screen high-risk populations. Nevertheless, this approach is limited by a high false-positive rate, high costs, and risky follow-up evaluations.
[0003] One approach used in cancer diagnosis is the analysis of tumor samples for genetic indexes or markers. Cancer genomes acquire somatic mutations that drive their proliferative potential (Non-Patent Document 1). Cancer genome mutations also provide important information about the active evolutionary history and mutational processes of each cancer (Non-Patent Documents 2 and 3). Cancer mutation calling in patient biopsies is a crucial step in assessing patient outcomes and treatment recommendations. Identification of cancer driver mutations in liquid biopsy specimens, such as cell-free circulating DNA (cfDNA), is suggested as a transformative platform for early cancer screening.
[0004] Statistical methods for analyzing genomic markers, such as DNA somatic mutations (e.g., single-nucleotide variants (SNVs)), require multiple independent observations (supportive reads) of a somatic mutation at any genomic location to distinguish true mutations from sequencing errors. One technique used to distinguish true mutations from sequencing errors is to increase sequencing depth, which is useful when the proportion of tumor cells in a tumor sample is high. For example, if the tumor cell content in a sample is reduced due to the presence of normal cells such as immune cells, each somatic variant will no longer be supported by multiple reads, and the variant call will no longer be valid. For example, MUTECT is the current state-of-the-art low-allele frequency somatic variant calling method. At its core, MUTECT separates SNVs into two Bayesian classifiers: one assumes that SNVs are due to random noise, and the other assumes that the site contains a true variant. SNVs are then filtered based on the log-likelihood ratio derived from the two models. This is fundamentally different from the cfDNA setting. In benchmarking settings where the frequency of mutant alleles is reduced to 0.05 and the sequencing depth of tumor samples is reduced 10-fold, the sensitivity of MUTECT drops to less than 0.1 (Non-Patent Document 4). Although MUTECT is currently the state-of-the-art for somatic mutation calling in low-frequency settings, it remains unable to identify somatic mutations in the tumor fraction observed in cfDNA.
[0005] A fundamental limitation of MUTECT and other variant calling methods is that clinical sensitivity falls below acceptable levels when input material is limited (e.g., in the early-stage cancer setting). The small amount of cfDNA involved is only a few hundred to a few thousand genome equivalents. Therefore, ultra-deep sequencing (e.g., 100,000X) can be rendered ineffective by the limited number of physical fragments covering each site present in the sample (e.g., 1,000 genome equivalents in 6 ng of cfDNA). Even with ultra-deep sequencing and advanced molecular error suppression, the limit of detection with limited input material is a tumor fraction (TF) frequency of less than 0.1–1%.
[0006] This limitation is exemplified by [5], who applied advanced sequencing methods, including technically challenging patient-specific targeted deep sequencing for lung adenocarcinoma, and identified approximately 18 mutations at a median sequencing depth of 42,000x. However, the rarity of cfDNA likely meant that cancer was detected in only 19% of early-stage subjects, even with the inclusion of more advanced stage III tumors in the study group. Furthermore, all of these positively identified patients had lesions detectable by CT scan. These data suggest that even ultra-deep sequencing underperforms current imaging technologies in terms of comprehensiveness and / or accuracy in the early stages of disease. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] Lawrence et al.,Nature,505(7484):495-501,2014 [Non-patent document 2] Martincorena et al.,Cell,171(5):1029-1041.e21,2017 [Non-patent document 3] Alexandrov et al.,Nature,500(7463):415-421,2013 [Non-patent document 4] Cibulskis et al.,Nature Biotechnology,31(3),213,2013 [Non-patent document 5] Abbosh et al.,Nature,545(7655):446-451, 2017 Summary of the Invention [Problem to be solved by the invention]
[0008] Improved methods and systems for identifying low-abundance disease markers, such as somatic mutations in cfDNA (including subject-specific features) that are indicative of tumor disease are desirable. Furthermore, systems and methods are desirable that utilize high-quality markers that can be used for early tumor diagnosis, providing clinicians with better options for disease management and / or therapeutic intervention, and significantly improving disease treatment outcomes (e.g., improved survival and / or quality of life). [Means for solving the problem]
[0009] Provided herein are programs, systems, and methods for screening subjects for cancer and using information obtained from the screening for early detection and disease stratification. In certain embodiments, the programs, systems, and methods of the present disclosure enable a user, e.g., a clinician, to diagnose cancer early.
[0010] In one aspect, the present disclosure provides a classifier trained to distinguish between systematic errors and somatic mutations induced by cancer (e.g., tobacco-induced lung cancer). Taking advantage of the fact that both cancer mutations and sequencing errors are systematic and controlled by distinct features that can be learned and used for efficient signal-to-noise discrimination, the classifier integrates this knowledge to improve the accuracy of cancer diagnosis and / or detection. For example, in a genomic context, cancer patterns can include base substitutions that induce cancer-associated mutagenesis. Such genomic patterns are particularly unique in cancers induced by tobacco and UV exposure, including cancers associated with deregulated DNA checkpoint and / or repair enzyme activity, such as BRCA (BRCA1 or BRCA2), p53, and APOBEC1.
[0011] The present disclosure also relates to several indicators that may suggest that a variant detected by sequencing is not a true somatic mutation, but rather an artifact of the sequencing or mapping technique. In this context, previous studies have shown that sequencing errors are not random, but are likely related to both the context of the DNA sequence resulting from the sequencing technique and technical factors. The fidelity of sequencing is also limited by the length of each sequencing read, with the probability of errors increasing as the read length increases. Errors may occur when a read is mapped to a reference genome. The process of mapping is computationally intensive and complex due to the fact that genomes contain variable regions, motifs, and repeatable elements. Short nucleotide reads may map to more than one location, or may not map at all. These limitations of existing methodologies for sequencing / mapping genomic data may be remedied using the systems and methods of the present disclosure. The disclosed metrics may analyze multiple factors, such as (i) low base quality, (ii) low mapping quality, (iii) read estimated fragment size (RP), (iv) read estimated allele fraction (VAF), (v) sequence context, (vi) abundance, (vii) sequencing depth, and / or (viii) sequencing errors, to derive true mutations from errors.
[0012] The systems and methods of the present disclosure are particularly applicable to the detection of low-abundance markers that predict cancer. The inventors of the present disclosure recognized that breadth of sequencing, which is not limited by the abundance of input material, can be used to replace methods that rely on depth sequencing. Breadth sequencing is less dependent on the abundance of input material and can be used to improve both the accuracy and sensitivity of detection. From a statistical perspective, the inventors first demonstrated that breadth of sequencing (e.g., 10x sequencing of 10,000 mutations) is equivalent to depth (100,000x sequencing of a single mutation) and can be performed with as little as 1 ng of cfDNA. Thus, the analytical approach of the present disclosure easily and non-invasively integrates genome-wide mutation information for sensitive analysis of samples containing cfDNA for the detection and / or accurate diagnosis of tumors (e.g., tobacco-induced cancer).
[0013] In this context, simulated plasma somatic mutation calling tests using a synthetic mixture of tumor and normal whole-genome sequence data from lung patients, in which the read rates for various tumor fractions ranged from 1% to 0.001% (1 / 10,000), demonstrate the strength and accuracy of our method over existing techniques. The performance of our technique was further benchmarked by first characterizing patient-specific somatic cancer SNVs using standard mutation calling in pure tumor and normal samples from the patient; then detecting cancer mutations in plasma samples using a method involving the disclosed convolutional network. The sensitivity and accuracy of each method using pure tumor mutation calling as a reference demonstrate high signal and / or low noise for our analytical method. Finally, validation studies conducted using actual cfDNA samples obtained from early-stage lung cancer patients demonstrate significantly superior sensitivity and accuracy compared to current state-of-the-art methods.
[0014] The present disclosure relates to the following non-limiting embodiments: In various aspects, a method for genetic screening of cancer in a subject is provided. The method includes receiving a subject-specific genome-wide compendium of reads associated with a plurality of genetic markers from a biological sample of the subject. The biological sample may include a tumor sample. The read compendium may each include single base pair reads. The method may further filter artifactual sites from the read compendium. The filtering may remove repetitive sites generated across a cohort of reference healthy samples from the read compendium. Alternatively, or in combination, the filtering may include identifying germline mutations in the biological sample and / or identifying mutations shared between a tumor sample and a peripheral blood mononuclear cell of a normal cell sample as germline mutations, and removing the germline mutations from the read compendium. The method may further filter noise from the read compendium using at least one error suppression protocol to generate the filtered read set of the genome-wide compendium of reads. The at least one error suppression protocol may calculate the probability that any single nucleotide variation in the list is an artifact and remove the variation. The probability may be calculated as a function of features selected from the group including mapping quality (MQ), variant base quality (MBQ), position in read (PIR), mean read base quality (MRBQ), and combinations thereof. Alternatively or in combination, artifact removal may include testing for discrepancies between independent copies of the same DNA fragment generated by polymerase chain reaction or sequencing processing, and / or overlap consensus, which identifies and removes artifacts when there are no matches in a given overlap family. The method may include compiling subject-specific patterns using the filtered read set based on comparison of specific mutation patterns associated with a given mutagenesis process. The method may further include statistically quantifying a confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern based on comparison of the subject-specific pattern with a cohort of background mutation patterns of cancer-associated mutation pattern exposure values.The method may include screening the subject for cancer if the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern exceeds a predetermined threshold.
[0015] In various aspects, a method for genetically screening a subject for cancer is provided. The method includes receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject. The biological sample may include a tumor sample. Each read list may include copy number variations (CNVs). The method may include dividing the read list into a plurality of windows. The method may include calculating a set of features per window. The features may include a median depth coverage per window and a representative fragment size per window. The method may include filtering artifacts from the read list. The filtering may include removing repetitive sites generated on a cohort of reference healthy samples from the read list. The method may include normalizing the read list, creating a filtered set of reads of the genome-wide read list. The method may include calculating a linear relationship between the feature sets per window, converting the calculated relationship to an estimated tumor fraction using a regression model, and calculating an estimate of tumor fraction using the filtered read set. Alternatively, or in combination, the method may include calculating an estimate of tumor fraction using the filtered read set based on one or more integrative mathematical models as a function of the calculated feature sets per window across the subject-specific genome-wide list of reads. The method may include screening the subject for cancer if the estimated tumor fraction exceeds an empirical threshold.
[0016]
[0003] A system for genetically screening a subject for cancer is provided. In various embodiments, the system includes an analysis unit, the analysis unit including a pre-filter engine configured and arranged to receive a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, the biological sample including a tumor sample, and the list of reads each including a single base pair in length. The pre-filter engine can be configured and arranged to filter artifactual sites from the list of reads, the filtering including removing repetitive sites generated across a cohort of reference healthy samples from the list of reads. The pre-filter engine can be configured and arranged to identify germline mutations in the biological sample and / or identify mutations shared between the tumor sample and peripheral blood mononuclear cells of a normal cell sample as germline mutations and remove the germline mutations from the list of reads, or in combination. The analysis unit may include a correction engine constructed and arranged to filter noise in the read inventory using at least one error suppression protocol to create a filtered read set for the subject-specific genome-wide inventory of reads, wherein the at least one error suppression protocol comprises: (a) calculating the probability that any single nucleotide variation in the inventory is an artifact and removing the variation, wherein the probability is calculated as a function of features selected from the group including mapping quality (MQ), variant base quality (MBQ), position in read (PIR), mean read base quality (MRBQ), and combinations thereof.
[0017] The at least one error suppression protocol may include removing mutations using discrepancy testing between independent copies of the same DNA fragment generated by polymerase chain reaction or sequencing processing, and / or combining overlap consensus to identify and remove artifactual mutations when there are no matches in a majority of predetermined overlap families. The system may include a computing unit configured and arranged to compile a subject-specific pattern using the filtered read set based on a comparison of specific mutation patterns associated with a predetermined mutagenesis process. The computing unit may be configured and arranged to statistically quantify a confidence estimate via the subject-specific pattern that the subject's biological sample contains a cancer-associated mutation pattern based on a comparison of the cancer-associated mutation pattern exposure value with a cohort of background mutation patterns. The computing unit may be configured and arranged to screen the subject for cancer if the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern exceeds a predetermined threshold.
[0018] In various aspects, a system for genetically screening a subject for cancer is provided. The system includes an analysis unit, the analysis unit including a binning engine configured and arranged to receive a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, the biological sample including a tumor sample, and the read list each including copy number variations (CNVs). The binning engine may be configured and arranged to divide the read list into a plurality of windows and calculate a set of features per window, the features including a median depth coverage per window and a representative fragment size per window. The system may include a pre-filter engine configured and arranged to filter artifactual sites from the read list, the filtering including removing repetitive sites generated on a cohort of reference healthy samples from the read list. The system may include a normalization engine configured and arranged to normalize the read list to generate a filtered set of reads for the genome-wide list of reads. The system may include a computing unit configured and arranged to calculate an estimated tumor fraction (eTF) using the filtered read set. The computing unit may calculate a linear relationship between feature sets per window and convert the calculated relationship into an eTF using a regression model. Alternatively, or in combination, the computing unit may calculate eTF as a function of the calculated feature sets per window across the subject-specific genome-wide read list based on one or more integrative mathematical models. The computing unit may be configured and arranged to screen the subject for cancer if the estimated tumor fraction exceeds an empirical threshold.
[0019] In an embodiment, the present disclosure relates to a method for genetic screening of cancer in a subject, comprising: (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) statistically classifying each read in the list as signal or noise based on (1) the base quality (BQ) of the read, (2) the mapping quality (MQ) of the read, (3) the estimated fragment size of the read, and / or (4) the probability of detection of noise (PN) as a function of the estimated allele fraction of the read, to identify a cancer from the list. (c) adaptively and / or systematically filtering sequencing noise associated with each read in the inventory using a machine learning (ML) approach that distinguishes between cancer-associated mutational signatures and PCR or sequencing error-associated signatures; (d) compiling subject-specific signatures comprising multiple true reads in the inventory based on the noise removal step (c) and the filtering step (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signatures and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern.
[0020] In some aspects of the methods, the subject's biological sample comprises plasma, cerebrospinal fluid, pleural effusion, ocular fluid, stool, urine, or a combination thereof.
[0021] In certain embodiments of the methods, the cancer pattern includes a COSMIC tobacco pattern, a UV pattern, a breast cancer (BRCA) pattern, a microsatellite instability (MSI) pattern, an apolipoprotein B mRNA editing enzyme, poly(ADP-ribose) polymerase (PARP) multiactivation pattern, or an APOBEC pattern. In particular, in certain embodiments, the cancer pattern includes a pattern associated with a tissue-specific epigenetic pattern, such as a tissue-specific chromatin accessibility pattern.
[0022] In some aspects of the method, the sequencing noise associated with each read in the list is filtered using a machine learning (ML) approach to distinguish between mutational signatures associated with cancer (true positives) and signatures associated with PCR or sequencing errors (false positives). In some embodiments, the machine learning method includes a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), discriminant analysis, a nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof. In some embodiments, the ML is trained to distinguish between cancer-altered sequencing reads and reads altered by sequencing or PCR errors. In some embodiments, the ML is trained on a large whole-genome sequencing (WGS) cancer dataset containing billions of reads across tumor mutations and normal sequencing errors. In some embodiments, the ML can (a) identify sequencing or PCR artifacts with high accuracy and (b) integrate sequence context and read specific signatures.
[0023] In an embodiment, the present disclosure relates to a method for genetic screening of cancer in a subject, comprising the steps of: (a) receiving a subject-specific genome-wide inventory of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the inventory of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; (b) statistically classifying each read in the inventory as signal or noise based on the read's (1) base quality (BQ), (2) read mapping quality (MQ), (3) read estimated fragment size, and / or (4) probability of detection of noise (PN) as a function of the read's estimated allele fraction; and (c) removing artifactual reads from the inventory. (d) compiling a subject-specific pattern comprising a plurality of true reads in the list based on the noise removal step (c) and the filtering step (b); (e) statistically quantifying whether the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific pattern and the cancer pattern; and (f) screening the subject for cancer based on a confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern.
[0024] In some embodiments of the method, the tumor is a heterogeneous or allogeneic brain tumor, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, renal cancer, oral cancer, gastric cancer, solid tumor, non-small cell lung cancer (NSCLC), tobacco-induced cancer (TIC), UV light-induced cancer, cancer mediated by apolipoprotein B mRNA editing enzyme catalytic protein (APOBEC) activity, cancer containing breast cancer protein (BRCA) mutations, cancer containing poly(ADP-ribose) polymerase (PARP) activity, and tumor containing microsatellite instability (MSI). In some embodiments of the method, the screening method enables the diagnosis of early-stage cancer disease in undiagnosed and / or asymptomatic patients. In particular, the subject is a patient with early-stage cancer at stages I to III.
[0025] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide inventory of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the inventory of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; (b) statistically classifying each read in the inventory as signal or noise based on the probability of detection of noise (PN) as a function of (1) the base quality (BQ) of the read, (2) the mapping quality (MQ) of the read, (3) the estimated fragment size of the read, and / or (4) the estimated allele fraction of the read, to remove artifactual reads from the inventory; and (c) using a machine learning (ML) approach to distinguish between cancer-associated mutational signatures and PCR or sequencing error-associated signatures. (d) compiling a subject-specific pattern comprising a plurality of true reads in the list based on the noise removal step (c) and the filtering step (b); (e) statistically quantifying whether the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific pattern and the cancer pattern; (f) screening the subject for cancer based on a confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern; and (g) prescribing a pattern-based treatment based on the patient-specific pattern used for diagnosis. In one embodiment, the treatment prescribing includes a PARP inhibitor for a BRCA pattern and an immunotherapy for an MSI pattern. In one embodiment, the PARP inhibitor is niraparib, olaparib, veliparib, rucaparib, and / or talazoparib. In some embodiments, the immunotherapy for the MSI pattern includes an anti-PD-1 antibody (e.g., nivolumab or pembrolizumab) or an anti-CTLA4 antibody (e.g., nivolumab or pembrolizumab).In some embodiments, the tumor comprises a heterogeneous or homogeneous brain tumor, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, mouth cancer, stomach cancer, solid tumor, lung adenocarcinoma, breast ductal adenocarcinoma (breast tumor), non-small cell lung cancer (NSCLC LUAD), cutaneous melanoma, urothelial carcinoma (bladder tumor), colorectal cancer (Lynch), or osteosarcoma.
[0026] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N(c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising multiple true reads in the inventory based on the noise removal (c) and filtering (b); (e) determining whether the subject's biological sample is matched between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern, wherein step (f) comprises solving a linear optimization function minllAx-bll,x≧0, where A is the mutation pattern sequence context matrix, x is the contribution of each cosmic mutation pattern (variable), and b is the patient-specific sequence context list. In some embodiments, the optimization problem is solved by non-negative least squares (NNLS), cross-entropy global optimization, golden section search, or a combination thereof. In some aspects, the method further comprises verifying the reliability using comparison of the cancer mutation pattern with multiple random background patterns, e.g., using comparison of the cancer mutation pattern with multiple random background patterns. In some aspects, the comparing step comprises evaluating a z-score, wherein a z-score above a threshold indicates that the subject-specific feature is specific to the cancer feature and not associated with random features.
[0027] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N (c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising multiple true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) quantifying whether the subject's biological sample contains a cancer-associated mutational pattern. and / or (4) removing high fragment size reads (e.g., >160, ROC optimized), and wherein step (f) comprises calculating the sequence context similarity between a particular cosmic sequence context list and a patient sequence context list.
[0028] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N (c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising a plurality of true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains the cancer-associated mutational signature, wherein step (f) comprises estimating a similarity between the subject-specific signature and the cancer pattern based on cosine-similarity, correlation, mutual information, or a combination thereof. In some embodiments, the method further includes verifying the reliability of the cancer mutation pattern using a comparison of the cancer mutation pattern to a plurality of random background patterns, e.g., using a comparison of the cancer mutation pattern to a plurality of random background patterns. In some embodiments, the comparing step includes assessing a z-score, where a z-score above a threshold indicates that the subject-specific feature is specific to the cancer feature and is not associated with random background features.
[0029] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N (c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising a plurality of true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains the cancer-associated mutational pattern, wherein step (f) comprises comparing the cancer-specific pattern confidence (z-score) to an empirical threshold calculated by a background noise model. In one embodiment, an empirical noise model is defined by measuring cancer-specific feature confidence (z-score) in normal healthy samples and converting it into a base noisy z-score estimate, where the z-score estimate noise threshold is between 1 and 5.
[0030] In certain embodiments of the above cancer screening / diagnostic methods, the subject-specific signature matches a cancer-specific mutation signature comprising markers that are differentially expressed in tumors but not in normal samples. In certain embodiments, the tumor sample comprises a lung tumor, a breast tumor, a melanoma, a bladder tumor, a colorectal tumor, or a bone tumor.
[0031] In certain embodiments of the cancer screening / diagnostic methods, the methods allow for early detection in at least 50% of subjects.
[0032] In some embodiments of the cancer screening / diagnosis methods, the method further comprises performing computed tomography (CT) screening, where the CT screening step occurs before, simultaneously with, or after the genetic screening. In some embodiments, the cancer is a solid tumor, and the CT screening comprises, for example, detecting suspicious nodules in patients with benign lesions. In some embodiments, the benign lesions are identified via advanced CT screening, histopathology, and / or biopsy.
[0033] In certain embodiments of the cancer screening / diagnostic methods, the methods include distinguishing between malignant and benign nodules to increase the positive predictive value (PPV) of the CT screening, e.g., by at least 30%, at least 40%, at least 50%, at least 60%, at least 80%, or at least 90%.
[0034] In certain embodiments of the cancer screening / diagnostic methods, the methods involve early detection (ED) of malignancies.
[0035] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N(c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising a plurality of true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains a cancer-associated mutational pattern, wherein step (a) comprises aggregating genome-wide mutation data by whole genome sequencing, and step (c) comprises detecting the mutational pattern using a mathematical optimization process. In some embodiments, the mathematical optimization step includes using non-negative least squares.
[0036] In an embodiment, the disclosure provides a method for (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, wherein the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allelic fraction (VAF) of the reads. N(c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising a plurality of true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern. In certain embodiments, the pre-malignant tumor comprises heterogeneous or homogeneous brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, mouth cancer, stomach cancer, solid tumor, lung adenocarcinoma, breast ductal adenocarcinoma (breast tumor), non-small cell lung cancer (NSCLC LUAD), cutaneous melanoma, urothelial carcinoma (bladder tumor), colorectal cancer (Lynch), or osteosarcoma, particularly Lynch syndrome or BRCA gene deletion.
[0037] In some embodiments of the above methods, the machine learning (ML) comprises a deep convolutional neural network (CNN) that adaptively and / or systematically filters ordered noise. In some aspects, the CNN comprises: using a deep learning algorithm on a pan-tumor cohort to identify features that distinguish between true tumor mutations and artifactual errors; assigning a confidence estimate to each individual mutation detected in samples from tumor patients; integrating the confidence estimates across the entire genome; and using non-negative least squares for specific cosmic mutations in the samples.
[0038] In an embodiment, the present disclosure relates to a computer-readable medium comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for early detection of tumors or detection of precancerous tumor lesions, the method or set of steps including: (a) receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject, where the list of genetic markers is selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (indels), copy number variations, structural variations (SVs), and combinations thereof; and (b) calculating noise (PQ) as a function of (1) the base quality (BQ) of the reads, (2) the mapping quality (MQ) of the reads, (3) the estimated fragment size of the reads, and / or (4) the estimated allele fraction (VAF) of the reads. N (c) utilizing a machine learning (ML) approach to adaptively and / or systematically filter the sequencing noise associated with each read in the inventory to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-associated signatures; (d) compiling a subject-specific signature comprising a plurality of true reads in the inventory based on the noise removal (c) and filtering (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on a match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer based on the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern. In one embodiment, the ML comprises a layered convolutional neural network (CNN) with a single fully connected layer at one end, where the CNN maintains spatial invariance when convolving over 3-nucleotide windows and maintains mapping quality by collapsing read fragments into multiple segments, each representing an approximately 8-nucleotide region.
[0039] In some embodiments of the computer-readable medium or method, the CNN includes eight layers, including a single fully connected layer at one end and two consecutive convolutional layers with a two-step receptive field and a two-step receptive field, where the output is downsampled by max-pooling. The eight-layer CNN maintains mapping quality by collapsing read fragments into approximately 25 individual segments using a receptive field of size 3 and convolving them onto columns corresponding to their position in the genomic read. The output of the final convolutional layer is applied directly to a sigmoid fully connected layer, where final classification of markers is performed. In some embodiments, the CNN includes a read representation that simultaneously captures the genomic context of the alignment, the complete read sequence, and an integrated per-base quality score. In some aspects of the computer-readable medium or method, the CNN provides enrichment of tumor-specific markers, including somatic mutations, in genomic reads by about 1.12-fold to about 30-fold compared to MUTECT.
[0040] In one embodiment, the present disclosure relates to a computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method or series of steps for diagnosing cancer in a subject in need of diagnosis, the medium comprising: (A) receiving a list of genetic markers from a sample from the subject, the genetic markers including somatic single nucleotide variations (sSNVs), somatic copy number variations (sCNVs), insertions / deletions (indels), or structural variations (SVs) in genomic reads; (B) processing the list of genetic markers for each subject in a pan-tumor cohort to identify features that distinguish between true cancer markers and artifactual errors; (C) assigning a confidence estimate to each feature in the collection based on processing step (B); (D) for each feature in the genomic reads, integrating the confidence estimates across step (C) to construct a tumor signature; and (E) mathematically optimizing the tumor signature. In some embodiments, assigning a confidence estimate involves (1) calculating a confidence measure for the contribution of a cosmic mutation pattern using linear mixture optimization, or (2) calculating the similarity of a patient sequence-context list to a particular cosmic pattern. In some embodiments, linear mixture optimization involves solving the algebraic function minllAx-bll,x≧0, where A is the mutation pattern sequence-context matrix, x is the contribution of each cosmic mutation pattern (variable), and b is the patient-specific sequence context list. In some aspects, A in the algebraic function minllAx-bll,x≧0 includes at least 5, at least 10, at least 15, at least 20, at least 25, or at least 30 COSMIC patterns along with 100 random mutation patterns.In one embodiment, the linear mixed optimization includes calculating the distribution of random pattern contributions including extraction E_random (mean contribution score) and std_random (std contribution score), checking the reliability of contribution detection for each COSMIC pattern by z-score, and calculating the metric (cosmic_sig_contribution-E_random) / std_random, where the metric represents the significance of a particular pattern compared to the random set. In one embodiment, the mathematical optimization process includes using non-negative least squares (NNLS).
[0041] In an embodiment, the present disclosure provides a system for detecting residual tumor in a subject in need thereof, comprising: a data collection unit configured and arranged to receive a plurality of amplified and sequenced read lists from a plasma sample of the subject and a normal biological sample, including a normal cell sample; and a marker identification unit configured and arranged to identify a plurality of subject-specific markers in a subject-specific list of genetic markers, the marker identification unit being communicatively coupled to the data collection unit and configured to identify a plurality of subject-specific markers in a subject-specific list of genetic markers based on a base quality (BQ) of the reads, a mapping quality (MQ) of the reads, a fragment size of the reads, and / or a variable allele frequency (VAF) of the reads. a noise reduction unit that removes actual noise based on a confidence interval score indicating a statistical level of statistical association between the read and the tumor; and a classification engine configured and arranged to statistically classify each denoised read in the list based on a confidence interval score that indicates a statistical level of statistical association between the read and the tumor, wherein the classification engine utilizes machine learning (ML) to adaptively and systematically filter noise introduced during the amplification or sequencing step; further comprising: a step of matching the denoised ML filtered reads with one or more known cancer patterns in the list; and a diagnostic unit configured and arranged to diagnose the tumor based on the match.
[0042] In certain embodiments of the disclosed system, the classification engine is further configured to calculate a reliability metric using a linear mixed optimization problem to match the denoised ML filter readings with one or more known cancer features.
[0043] In certain aspects of the disclosed systems, optimizing the linear mixture includes calculating a z-score confidence estimate for the association between tumor incidence and a tumor mediator selected from tobacco exposure, UV exposure, deregulated DNA repair, defective DNA editing, microsatellite instability, or a combination thereof.
[0044] In some embodiments of the disclosed system, the artificial noise removal engine is configured to perform a best fit receiver operating characteristic curve, including a probabilistic classification of the reads in the list based on the read's base quality (BQ) score, the read's mapping quality (MQ) score, the read's fragment size, or the read's variable allele frequency (VAF). In some embodiments of the disclosed system, the artificial actual noise removal engine is further configured to filter noise based on (iii) intra-read position (RP), (iv) read sequence context (SC), (v) read abundance, (vi) sequencing depth, and / or (vii) sequencing error.
[0045] In some embodiments of the disclosed system, the reliability metric calculation involves solving the algebraic function minllAx-bll,x≧0, where A is the mutation pattern sequence context matrix, x is the contribution of each cosmic mutation pattern (variable), and b is the patient-specific sequence context list. In some embodiments, the z-score reliability estimate involves solving the algebraic function A including 30 cosmic patterns and 100 random mutation patterns, calculating the distribution of contributions of the cosmic patterns (CSC) and random patterns (E_random) including the average contribution score (ACS) and standard contribution score (std_random), and calculating a z-score metric using the function (CSC-E_random) / std_random to check the reliability of the contribution for each cosmic pattern, where the z-score represents the importance of the particular pattern contribution compared to a random set. In some embodiments, the z-score reliability estimate involves calculating the similarity of the patient sequence-context list to a particular cosmic feature. In some embodiments, the reliability estimate of the z-score includes normalizing the list of patient-sequence contexts to obtain a density function, calculating the cosine similarity between the density function of the patient-sequence context and the cosmic signature density function, and normalizing the cosine similarity by dividing by the cosine similarity between the density function of the patient-sequence context and a non-informative uniform density function. In some embodiments, the reliability estimate of the z-score includes checking whether the z-score exceeds a detection threshold, the threshold including an empirically estimated baseline noise for healthy samples. In some aspects, the cancer feature includes a tobacco feature, and a positive confidence interval includes a z-score greater than 2, 3, 4, and preferably greater than 5 standard deviations.
[0046] In some embodiments of the methods and systems of the present disclosure, the genetic markers include SNVs, CNVs, indels, and / or SVs in DNA, and the receiving unit receives genetic data from whole genome sequencing (WGS), e.g., from a biological sample including a plasma sample. The normal cell sample includes cell-free DNA (cfDNA). The normal cell sample includes peripheral mononuclear cells (PMBCs), and the genetic data includes multiple markers including somatic single nucleotide variations (sNVs) or somatic copy number variations (sCNVs), or a combination thereof. In some embodiments, the amount of cfDNA in the sample is about 0.1 ng / ml to about 20.0 ng / ml. In some embodiments, the sample has a low tumor fraction (TF), as measured by the ratio of tumor DNA molecules to normal DNA molecules, e.g., between about 0.0001% (1 to 1 million molecules) and about 20%.
[0047] The details of one or more embodiments of this disclosure are set forth in the accompanying drawings / tables and the description below. Other features, objects, and advantages of this disclosure will be apparent from the drawings / tables, detailed description, and claims. [Brief explanation of the drawings]
[0048] [Figure 1]A representative flowchart of the diagnostic method of the present disclosure is shown. In the first step 110 of Figure 1A, a subject-specific genome-wide list of reads associated with multiple genetic markers (e.g., somatic SNVs) is received from the subject's sample (e.g., generated via whole-genome sequencing). In step 120, each read is statistically classified as signal or noise (N) based on the probability of detecting noise as a function of (1) base quality (BQ), (2) mapping quality (MQ), (3) estimated fragment size, and / or (4) estimated allele fraction (VAF) to remove artifactual reads. Other secondary parameters, such as (v) read position, (vi) sequence context position size (SC), (vii) abundance, (viii) sequencing depth, and / or (ix) sequencing error, may also be used. The noise-reduced reads may be fed into an in silico dataset and / or a convolutional neural network trained using a dataset from a pan-cancer cohort. The neural network adaptively and systematically filters sequencing noise in step 130. Next, based on the noise removal step 120 and the filtering step 130, a subject-specific feature including multiple true reads in the list is compiled in step 140. Next, in step 150, a confidence estimate that the subject's biological sample contains circulating tumor DNA is made by matching the subject-specific pattern and the cancer pattern. The subject is screened for cancer in step 160 based on the confidence estimate. Figure 1B shows a representative workflow for cancer screening of a subject according to various embodiments. Figure 1C shows a representative workflow for cancer screening of a subject according to various embodiments. Figure 1D shows a representative workflow for cancer screening of a subject based on measurement of single nucleotide polymorphisms (SNVs) or indels. Figure 1E shows a representative workflow for cancer screening of a subject based on measurement of copy number variations (CNVs) or structural variations (SVs). Figure 1F shows the scheme for the generation of an in silico database for synthetic plasma generated in seven cancer patients—two melanoma, three lung adenocarcinoma, and two breast (SCHEME A).
[0049] [Figure 2]An exemplary flowchart is provided outlining the use of the disclosed systems and methods to aid in early cancer detection, which reduces, if not eliminates, the need for surgical and / or therapeutic intervention. The many economic and health benefits that result from early cancer detection include risks of surgery (e.g., pneumonia, bleeding, infection, blood clots (hematomas), and reactions to anesthesia), side effects of chemotherapy or immunotherapy (e.g., fatigue, hair loss, easy bruising and bleeding, infections, anemia, nausea and vomiting, loss of appetite, constipation, diarrhea, mouth, tongue, and throat problems, nerve and muscle problems (e.g., numbness, tingling, pain), skin and nail changes (e.g., dry skin and changes in color), urinary and bladder changes, kidney problems, weight fluctuations, etc.
[0050] [Figure 3] Figures A–C show the probability of detection of a parameter as a function of various parameters. In Figure 3A, the chart shows that the probability of detection decreases rapidly in samples containing low tumor fractions (TF). Figure 3B shows the predicted average number of detected sites and the probability of at least one detection as a function of the number of unique DNA fragments (genome equivalents or coverage), mutation burden (N), and tumor fraction (TF). Figure 3C shows that by incorporating over 20,000 sSNVs (approximately 10 mutations / megabase pair found in 17% of human cancers), even a moderate sequencing effort (20X coverage), easily achievable with standard whole genome sequencing (WGS), can provide a high probability of detection (up to 0.98).
[0051] [Figure 4]Figures A–E show the optimization of SNV markers. Figure 4A shows the linear relationship between the number of artifactual SNV detections (errors) and the total number of unique reads checked. This represents an error probability equivalent to 1 error per 1,000 reads, indicating that this error is primarily due to the sequencing error rate (1 / 1,000). Each point is a control sample (TF=0), and these points were generated from PBMC data from six patients with multiple coverage (2X–25X) and multiple independent replicates of three different cancer types (lung cancer, melanoma, and breast cancer). This is invariant to cancer type, as all appear to fall on the same regression line. Figure 4B shows the receiver operating characteristic curve for base quality filtering. Figure 4C shows a line graph of the number of checked reads (x-axis) against the number of errors detected (y-axis) for the filtered multiple cancer error model, demonstrating the linear relationship between the number of artifactual SNV detections (errors) and the total number of unique reads checked. SNV detection (errors) is performed after applying the optimized BQ and MQ filters. Figure 4D shows the effect of applying the joint BQ and MQ optimization filters, resulting in a reduction of approximately 7-fold changes in sequencing error. Evaluation of the error probability distribution across multiple replicates using control samples showed that the noise before filtering was ~2 × 10-3 for both lung cancer and melanoma types, while the noise after filtering was reduced to ~2 × 10-4 for both cancer types. Figure 4E shows a heat map of the error rate (red indicating more errors and blue indicating fewer errors) as a function of plasma coverage (x-axis) and tumor burden (y-axis). The estimated error probability (e.g., the number of detected SNVs divided by the total number of unique reads checked) and tumor mutation burden (tumor mutation burden was corrected by subsampling the original patient-specific tumor mutation list) at various coverage levels are shown. Each entry in the matrix is the average of multiple independent replicates. This shows a relatively invariant error probability (approximately 2–3 × 10-4) with respect to coverage and mutation burden for all mutation burdens over 2000. This indicates that the results are robust to all tumors with one or more mutations per megabase pair (>1 / Mbp).
[0052] [Figure 5] A chart of deep learning-based de novo mutation detection and noise suppression is provided.
[0053] [Figure 6] A typical padding matrix for genomic reads (e.g., 16 x 200 base pairs for a 150 base pair read) is shown. The top panel shows the reads and their alignment as displayed by the engine. The bottom panel shows how genomic context is added to the end of the read. Zeros are padded for features other than the context.
[0054] [Figure 7]
[0023] Figure 1 shows a schematic diagram of an exemplary method of the disclosure applied in a clinical setting. As shown, a biopsy sample obtained from a subject (e.g., a cancer patient or a subject suspected of having a tumor) containing cell-free DNA (cfDNA) (e.g., a plasma sample) is processed (e.g., sequenced) to obtain the patient's genetic data (e.g., a VCF file) that is cataloged using PILEUP (or a similar program). A VAF filter is applied to filter out germline markers (e.g., SNVs, CNVs, indels, or SVs), and mapping quality (MQ), positional filter (PIR), and / or base quality (BQ) filters are further applied to filter out artifactual noise. In a next step, deep learning is applied to the filtered genetic data. The deep learning method involves training a machine using genetic data including a list of markers obtained from mixed tumor biopsy specimens and peripheral blood mononuclear cells (PMBCs; control), which are subjected to the above filters (e.g., artifactual read catalog via PILEUP, VAF filter to filter out germline mutations, BQ filter to remove markers with poor base quality, and MQ to remove markers with poor mapping). The device may also be trained using the dataset. The output of the above system and method is the identification of multiple markers in cfDNA that are clinically relevant in the context of cancer diagnosis, useful for early cancer diagnosis and prognosis.
[0055] [Figure 8] "Dataset characteristics and engine signature analysis results are shown. Figure 8A shows the COSMIC patterns associated with tobacco (top) and melanoma (bottom) from Alexandrov et al. (supra, 2013). Figure 8B shows trinucleotide frequencies from sample-specific tumor and PBMC readouts. Specific trinucleotides bound to tobacco (purple) and ultraviolet light (green). Figure 8C shows the correlation between the relative difference in trinucleotide frequency and the average activity of the engine (engine)."
[0056] [Figure 9] 9A and 9B show line graphs of various performance-related characteristics of the engine of the present disclosure compared to known variant callers. FIG. 9A shows sensitivity using patient CA0044 synthetic plasma. It can be seen that the engine of the present disclosure (KitTYHAWK) outperforms variant callers known in the art, such as MUTECT, SNOOPER, and / or STRELIKA, in terms of sensitivity. FIG. 9B shows a comparative line graph of accuracy (measured in terms of positive predictive value or PPV) obtained using the engine on patient CA0044 synthetic plasma. MUTECT was excluded because it only detected two variants. It can be seen that the engine outperforms variant callers known to those skilled in the art in terms of accuracy. FIG. 9C shows enrichment achieved using the engine on patient CA0044 synthetic plasma. MUTECT was excluded because it only detected two variants. It can be seen that the engine outperforms known variant callers in terms of enrichment.
[0057] [Figure 10] 1 shows SNV detection rates in ctDNA samples obtained in silico or from control subjects (BB600; BB601) or cancer patients (BB1122 or BB1125) using the methods and systems of the present disclosure.
[0058] [Figure 11] 1 is a table showing the clinical characteristics of subjects diagnosed with adenocarcinoma or benign nodules.
[0059] [Figure 12]Figures 12A-C show tumor-specific features differentially expressed in various tumors. Figure 12A shows that application of tumor-specific features (UV, tobacco) provides high specificity in lung cancer and melanoma samples. Figure 12B shows differential expression of gene features in normal (PBMC) versus tumor samples in lung patients (left panel) and / or melanoma patients (right panel). Figure 12C shows the expression of various COSMIC patterns (and their associated z-scores) in patients with breast cancer, melanoma, or lung adenocarcinoma.
[0060] [Figure 13] Figures 13A-C show that cancer patterns can be detected in synthetic plasma down to about 1 / 1000 of the tumor fraction (TF). Figures 13A and 13B, which represent data from two seeds, seed 3 and seed 4, show that tobacco patterns can be detected in synthetic plasma down to about 1 / 1000 of the tumor fraction (TF). Figure 13C, which represents data from a single seed, shows that lung features can be detected in synthetic plasma down to about 1 / 1000 of the tumor fraction (TF).
[0061] [Figure 14] A-B show the z-scores of various patient samples. Figure 14A shows the z-scores of tobacco-related mutation pattern detection versus background random patterns for lung cancer patients (blue) and patients with benign nodules (red, detected by CT). This demonstrates the ability to distinguish benign from malignant nodules based on a noninvasive blood test. The tobacco signature (signature 4 / 8) is detected in the plasma of early-stage cancers from tobacco-exposed patients but not in benign nodules or patients with no history of smoking. "ND" indicates nondetectable samples. PY indicates the number of pack-years each patient has smoked. ED means early detection. Figure 14B shows the expansion of the z-scores of mutation pattern detection in a cohort of samples obtained from subjects with various stages of lung cancer (e.g., stages IA, IB, IIA, IIb, and IIIa) compared to benign controls. For most cancer samples, baseline sensitivity reached at least 67%, which increased to approximately 100% for all high-stage (e.g., stage IIIa or higher) cases.
[0062] [Figure 15] FIG. 1 is a schematic diagram of a computer system of the present disclosure.
[0063] [Figure 16] 1A-C provide schematic diagrams of various systems of the present disclosure, showing the various units included in a representative system.
[0064] [Figure 17] Figures 17A and 17B show the use of orthogonal features, such as fragment size, in the diagnostic methods of the present disclosure and the concomitant effects of applying such orthogonal features in SNV-based methods. Figure 17A shows the fragment size distributions exhibited in healthy normal cfDNA samples. Figure 17B shows the fragment size shift in breast tumor cfDNA (red and purple) compared to normal cfDNA samples. Figure 17C shows that in a mouse xenograft (PDX) model, tumor-derived circulating DNA is significantly shorter than normal-derived circulating DNA. Figure 17D shows a line graph of fragment DNA size (x-axis; number of bases) plotted against the frequency of observing fragments of that length across tumor and normal samples. Figure 17E shows patient-specific mutation detection using orthogonal features, such as the correspondence between DNA fragments and tumor origin, based on fragment size distribution (x-axis) and GMM-bound log-odds ratios (y-axis).
[0065] [Figure 18]Figures 18A and 18B show the use of orthogonal features such as fragment size in the diagnostic methods of the present disclosure and the concomitant effects of applying such orthogonal features in CNV-based methods. Figure 18A shows line graphs of genomic region (bp) versus cumulative plasma depth coverage skew (bottom panel), plasma versus vertical depth coverage skew (middle panel), and coverage (top panel). Figure 18B shows the relationship between the log2 of depth coverage (log2 > 0.5 = amplification, log2 < -0.5 = deletion) and the local fragment size center of mass (COM) for that segment. Figure 18C shows a dot plot of depth coverage Log2 versus fragment size center of mass (COM). Using the estimated Log2 and COM values for all windows across the genome, and the median sample center of mass (COM), the slope and R^2 of the Log2 / COM linear model are calculated at various time points (e.g., baseline, day 0, day 21, and day 42). Figure 18D shows the correlation between Log2 / FS estimates and tumor DNA fraction. Figure 18E shows the relationship between depth coverage-based CNV detection and fragment size center of mass-based CNV detection in patient samples. Figure 18F shows the lack of relationship between depth coverage-based CNV detection and fragment size center of mass (COM)-based CNV detection in normal (healthy) plasma samples. DETAILED DESCRIPTION OF THE INVENTION
[0066] The present disclosure will now be described in more detail with reference to the accompanying drawings, in which preferred embodiments of the present disclosure are shown. However, the present disclosure may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, the embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0067] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those of ordinary skill in the art. The terms used in describing the disclosure herein are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. Furthermore, unless the context otherwise requires, singular terms include plural terms and plural terms include singular terms. In general, the nomenclature utilized in connection with molecular biology and the protein and oligo- or polynucleotide chemistry and hybridization described herein is well known and commonly used in the art. Standard techniques are used, for example, for nucleic acid purification and preparation, chemical analysis, recombinant nucleic acid, and oligonucleotide synthesis. Enzymatic reactions and purification techniques are performed according to manufacturer's specifications or as commonly accomplished in the art or as described herein. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and described in the various general and more specific references cited and discussed throughout this specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY 2000). The nomenclature used in connection with the experimental procedures and techniques described herein is well known and commonly used in the art.
[0068] Various embodiments of the present disclosure are described in further detail in the following paragraphs.
[0069] [Definition] As used in the description of this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, "and / or" encompasses any and all possible combinations of one or more of the associated listed items, as well as the alternative ("or"), which indicates the lack of combinations in the interpretation.
[0070] The term "about" refers to a range of plus or minus 10% of the value; for example, "about 5" refers to 4.5 to 5.5, "about 100" refers to 90 to 100, etc., unless the context of the disclosure indicates otherwise; for example, in a list of values such as "about 49, about 50, about 55," "about 50" refers to less than half the interval between the preceding and succeeding values, e.g., a range of more than 49.5 or less than 52.5. Furthermore, the terms "less than about" or "greater than about" should be understood in light of the definition of the term "about" provided herein.
[0071] When a range of values is provided in this disclosure, each intervening value between the upper and lower limits of that range, and any other stated or intervening value within that stated range, is intended to be included within the scope of the disclosure. For example, if a range of 1 μM to 8 μM is recited, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, and 7 μM are also intended to be expressly disclosed.
[0072] As used herein, the term "plurality" can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.
[0073] The term "screening" or "screening" as used herein has a broad meaning. It includes processes intended for diagnosis or for determining the susceptibility, propensity, risk, or risk assessment of an asymptomatic subject to develop a disease later in life. Screening also includes determining the prognosis of a subject, i.e., if the subject has been diagnosed with a disorder, predicting the progression of the disorder, as well as assessing the effectiveness of therapeutic options to treat the disorder.
[0074] As used herein, the term "detecting" refers to the process of determining a value or set of values associated with a sample by measuring one or more parameters in the sample, and may further include comparing the test sample to a reference sample. According to the present disclosure, detecting a tumor includes identifying, assaying, measuring, and / or quantifying one or more markers.
[0075] As used herein, the term "diagnosis" refers to a method that can determine whether a subject is likely to suffer from a given disease or condition, including, but not limited to, diseases or conditions characterized by genetic mutations. Those skilled in the art often base their diagnosis on one or more diagnostic indicators, e.g., markers, the presence, absence, amount, or change in amount, the amount of which indicates the presence, severity, or absence of a disease or condition. Other diagnostic indicators include a patient's medical history, physical symptoms (e.g., unexplained weight loss, fever, fatigue, pain, or skin abnormalities), phenotype, genotype, or environmental or genetic factors. Those skilled in the art will understand that the term "diagnosis" refers to an increased likelihood of a particular course or outcome, i.e., an increased likelihood of a course or outcome in patients who exhibit a given characteristic, e.g., the presence or level of a diagnostic indicator, compared to individuals who do not exhibit that characteristic. The diagnostic methods of the present disclosure can be used independently or in combination with other diagnostic methods to determine whether a course or outcome is more likely to occur in patients who exhibit a given characteristic.
[0076] As used herein, the term "early detection" of a disease, e.g., cancer, refers to discovering the potential manifestation of the disease, e.g., before metastasis to a cancerous state. Preferably, early detection refers to identifying the disease before the observation of morphological changes in tissues or cells. Furthermore, the term "early detection" of cellular transformation refers to the likelihood that a cell will undergo transformation at its early stage before the cell is designated as transformed.
[0077] As used herein, the term "cellular transformation" refers to a change in a cell from one characteristic morphology to another, e.g., from normal to abnormal, from non-tumor to tumor, from undifferentiated to differentiated, or from homogeneous to heterogeneous. Furthermore, transformation can be recognized by changes in cell morphology, phenotype, and biochemical characteristics, such as growth characteristics, apoptotic characteristics, segregation, and invasive characteristics.
[0078] The term "tumor" as used herein includes any cell or tissue that may have undergone transformation at a genetic, cellular, or physiological level compared to normal or wild-type cells. The term generally refers to a neoplastic growth that may be benign (e.g., a tumor that does not form metastases and destroys adjacent normal tissue) or malignant / cancerous (e.g., a tumor that invades surrounding tissue and can usually give rise to metastases), and may be fatal to the host unless properly treated. Steadman's Medical Dictionary, 28 th See Ed Williams & Wilkins, Baltimore, MD (2005).
[0079] The term "cancer" (used interchangeably with "tumor") refers to human cancers and carcinomas, sarcomas, adenocarcinomas, lymphomas, leukemias, solid and lymphatic cancers, etc. Examples of various types of cancer include, but are not limited to, lung cancer, pancreatic cancer, breast cancer, stomach cancer, bladder cancer, oral cancer, ovarian cancer, thyroid cancer, prostate cancer, uterine cancer, testicular cancer, neuroblastoma, squamous cell carcinoma of the head, cervix, cervix and vagina, multiple myeloma, soft tissue and osteogenic sarcoma, colon cancer, colorectal cancer, renal cancer (e.g., RCC), pleural cancer, cervical cancer, anal cancer, bile duct cancer, gastrointestinal carcinoid tumor, esophageal cancer, gallbladder cancer, small intestine cancer, central nervous system cancer, skin cancer, choriocarcinoma; osteogenic sarcoma, fibrosarcoma, glioma, melanoma, etc. In certain aspects, "liquid" cancers, e.g., blood cancers, e.g., lymphoma and / or leukemia, are excluded.
[0080] Examples of cancer include adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, anorectal cancer, anal canal cancer, appendix cancer, childhood cerebellar astrocytoma, childhood cerebral astrocytoma, basal cell carcinoma, skin cancer (non-melanoma), biliary tract cancer, extrahepatic bile duct cancer, intrahepatic bile duct cancer, bladder cancer, bone and joint cancer, osteosarcoma and malignant fibrous histiocytoma, brain cancer, brain tumor, cerebral glioma, cerebral astrocytoma / malignant fibrous histiocytoma Glioma, ependymoma, medulloblastoma, supratentorial primitive extraneural tumor, visual pathway and hypothalamic glioma, breast cancer, bronchial adenoma / carcinoid, carcinoid, gastrointestinal cancer, nervous system cancer, nervous system lymphoma, central nervous system cancer, cervical cancer, chronic lymphocytic leukemia, chronic myeloproliferative disorder, colon cancer, colorectal cancer, cutaneous T-cell lymphoma, lymphoma, mycosis fungoides, Thesia syndrome, endometrial cancer of the esophagus, extracranial germ cell tumorCell tumors, extragonadal germ cell tumors, extrahepatic bile duct cancer, eye cancer, intraocular melanoma, retinoblastoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid, gastrointestinal stromal tumor (GIST), germ cell tumors, ovarian germ cell tumors, gestational trophoblastic tumor glioma, head and neck cancer, hepatocellular (liver) cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, eye cancer, pancreatic islet cancer (endocrine pancreas), Kaposi's sarcoma, renal cancer, kidney cancer, laryngeal cancer, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia , chronic myeloid leukemia, hairy cell leukemia, cancer of the lip and oral cavity, liver cancer, lung cancer, non-small cell lung cancer, AIDS-related lymphoma, non-Hodgkin's lymphoma, primary central nervous system lymphoma, Waldenstrom's macroglobulinemia, medulloblastoma, melanoma, intraocular melanoma, Merkel cell carcinoma, malignant mesothelioma, mesothelioma, metastatic squamous cell carcinoma, oral cancer, tongue cancer, multiple endocrine neoplasia, mycosis fungoides, myelodysplastic syndrome, myelodysplastic / myeloproliferative disorder, chronic myeloid leukemia , acute myeloid leukemia, multiple myeloma, chronic myeloproliferative disorders, nasopharyngeal carcinoma, neuroblastoma, oral cavity cancer, oral cancer, oropharyngeal cancer, ovarian cancer, ovarian epithelial cancer, ovarian low malignant potential tumor, pancreatic cancer, pancreatic islet cell carcinoma, cancer of the paranasal sinuses and nasal cavity, parathyroid cancer, pharyngeal cancer, pheochromocytoma, pineoblastoma and supratentorial primitive neuroectodermal tumor, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma, prostate cancer, rectal cancer, renal pelvis and ureter cancer, transitional cell carcinoma, retinoblastoma, salivary gland cancer These include, but are not limited to, liquid adenocarcinoma, Ewing's sarcoma, Kaposi's sarcoma, uterine cancer, uterine sarcoma, skin cancer (non-melanoma), skin cancer, Merkel cell carcinoma, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, gastric cancer, supratentorial primitive neuroectodermal tumor, testicular cancer, thymoma, thymic carcinoma, thyroid cancer, transitional cell carcinoma, renal pelvis and ureter and other urinary tract cancers, gestational trophoblastic neoplasia, urethral cancer, endometrial cancer, uterine sarcoma, uterine cancer, vaginal cancer, vulvar cancer, and Wilms' tumor.
[0081] As used herein, a "high rate of somatic mutations" refers to a tumor that has about 1, about 2, about 3, about 5, about 7, about 10, about 12, about 15, about 20, about 25, about 30, about 40, about 50, about 60, about 75, about 80, about 100, about 125, about 150, or more mutations per megabase pair of the genome (mutations / MBP). See Collisson et al., Nature, 511(7511):543-50, 2014.
[0082] As used herein, the term "non-small cell lung cancer" or NSCLC refers to all lung cancers other than small cell lung cancer, including certain subtypes, including, but not limited to, large cell carcinoma, squamous cell carcinoma, and adenocarcinoma, and includes all stages and metastases. Squamous cell carcinoma, which accounts for 25% of lung cancers, usually begins near the central bronchi. Tumors usually contain cavities and associated necrosis in the center. Well-differentiated squamous cell carcinomas often grow more slowly than other types of cancer. Adenocarcinoma accounts for 40% of non-small cell lung cancers and usually arises in peripheral lung tissue. While most cases of adenocarcinoma are associated with smoking, it is the most common type of lung cancer among never-smokers. See Rosell et al., Lung Cancer, 46(2), 135-48, 2004; Coate et al., Lancet Oncol, 10, 1001-10, 2009.
[0083] As used herein, the term "cell" is used interchangeably with "biological cell." Non-limiting examples of biological cells include animal cells such as eukaryotic cells, plant cells, mammalian cells, reptilian cells, avian cells, and fish cells; prokaryotic cells, bacterial cells, fungal cells, and protozoan cells; cells dissociated from tissues such as muscle, cartilage, fat, skin, liver, lung, and nervous tissue; immunological cells such as T cells, B cells, natural killer cells, and macrophages; embryos (e.g., zygotes), oocytes, eggs, sperm cells, hybrid cells, cultured cells, cells derived from cell lines, cancer cells, infected cells, transfected and / or transformed cells, and reporter cells. Mammalian cells can be obtained, for example, from humans, mice, rats, horses, goats, sheep, cattle, primates, and the like.
[0084] As used herein, the term "subject" refers to a mammal, including humans, veterinary or farm animals, livestock or pets, and animals commonly used in clinical research. In particular, the subject is a human subject, e.g., a human patient diagnosed with or suspected of having a tumor.
[0085] As used herein, the term "subject-specific dataset" refers to a variety of information unique to each individual, such as, for example, genomic information, phenotypic information, biochemical information, metabolic information, microbiome sequence information, electronic medical record data, electronic health record data, drug prescriptions, biometric data, nutritional information, exercise information, family medical history information (e.g., as obtained from a family health history survey), in-application chat logs, the subject's personal healthcare provider records and notes, the subject's insurance provider, patient submitted network information, social network information, etc. In some embodiments, one or more of the subject-specific datasets are routinely updated and / or supplemented. In some embodiments, one or more datasets are added to the plurality of subject-specific datasets.
[0086] "Subject-specific genomic information" refers to the genetic makeup of an individual, including mutations (e.g., SNPs, Del / Dups, VUS) and mutation frequencies, familial genomic sequence information, structural genomic information (including mutations (e.g., sequences, deletions, insertions)), single nucleotide polymorphisms, personal immunology information (study of immune system regulation and response to pathogens using genome-wide approaches), functional genomic information (functional genomic information focusing on dynamic aspects such as gene transcription, translation, and protein-protein interactions), computational genomics information (the use of computational statistical analysis to decipher, discover, or predict biology from genome sequence and associated data), epigenomics information (reversible modifications of DNA or histones that affect gene expression without changing the DNA sequence, such as DNA methylation and histone modifications), etiology information including personal genomes, pathogenesis, regenomics information, behavioral genomics information, and metagenomics (i.e., personal genetic material recovered directly from environmental samples) involving microbial interactions.
[0087] "Subject-specific phenotypic information" refers to sex, race, height, weight, hair color, eye color, heart rate, preferences, blood pressure, self-reported medical symptoms, medically diagnosed symptoms, medically diagnosed symptoms, test results and / or medically provided diagnoses, proteomic profile, etc., and "subject-specific biochemical information" refers to clinical tests (e.g., sodium, magnesium, potassium, iron, blood urea nitrogen (BUN), uric acid, etc.), drug / drug concentrations in tissues, blood, etc.
[0088] The terms "subject-specific electronic medical record data" (EMR), "electronic health record" (EHR), and "personal health record" (PHR) refer to data from individual healthcare providers, clinics, hospitals, healthcare facilities, subject health history, subject disease predispositions, subject medical history, diagnoses, medications / prescriptions, treatment plans, immunization dates, allergies, radiological images, laboratory tests and lab results, advance directives, biopsies, home and portable monitoring devices (e.g., FITBIT, iWatch, With Scale, wireless blood pressure monitors, etc.).
[0089] As used herein, the term "sample" refers to a composition obtained or derived from a subject containing cells and / or other molecular entities to be characterized and / or identified, for example, based on physical, biochemical, chemical, and / or physiological characteristics. Sources of tissue samples can be blood or any blood component; body fluids; fresh, frozen, and / or preserved organ or tissue samples, or solid tissue from biopsies or aspirates; and cells from any time during pregnancy or development of a subject or plasma. Samples include, but are not limited to, primary cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous humor, ocular fluid, lymphatic fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, urine, cerebrospinal fluid (CSF), saliva, sputum, tears, sweat, mucus, tumor lysates, and tissue culture media, as well as tissue extracts such as homogenized tissues, tumor tissues, and cell extracts. Samples also include biological samples that have been manipulated in some way after their procurement, such as by adding reagents, solubilizing, or enriching for certain components such as proteins or nucleic acids, or by embedding in a semi-solid or solid matrix for sectioning, such as thin tissue sections or cells in a histological sample. Preferably, the sample is obtained from blood or blood components, including, for example, whole blood, plasma, serum, lymph, etc.
[0090] As used herein, the term "marker" refers to a characteristic that can be objectively measured as an indicator of normal biological processes, pathogenic processes, or therapeutic intervention, such as the pharmacological response to anti-cancer drug treatment. Representative types of markers include molecular changes in the structure (e.g., sequence) or number of markers, including multiple differences such as gene mutation, gene duplication, or somatic mutation of cfDNA, copy number variation, tandem repeat, or combinations thereof.
[0091] The term "genetic marker" as used herein refers to a sequence of DNA having a specific location on a chromosome that can be measured in a laboratory. The term "genetic marker" can also be used to refer, for example, to cDNA and / or mRNA encoded by a genomic sequence, as well as the genomic sequence itself. A genetic marker can include two or more alleles or variants. A genetic marker can be a direct marker (e.g., a marker located within a gene of interest or a locus of interest (e.g., a candidate gene)) or an indirect marker (e.g., a marker closely associated with a gene of interest or a locus of interest because it is adjacent to the gene of interest or a locus of interest but not within the gene of interest or a locus of interest). Furthermore, a genetic marker can also be unrelated to a gene or locus present in a non-coding region of the genome, such as an SNV, CNV, or tandem repeat. A genetic marker includes nucleic acid sequences that may or may not encode a gene product (e.g., a protein). In particular, a genetic marker includes a single nucleotide polymorphism / mutation (SNP / SNV) or a copy number variation (CNV), or a combination thereof. Preferably, the genetic markers comprise somatic mutations in DNA, such as sSNVs or sCNVs, or combinations thereof compared to a reference sample.
[0092] As used herein, the term "cell-free DNA" or "cfDNA" refers to cell-free deoxyribose nucleic acid (DNA) strands, such as those extracted or isolated from circulating blood plasma / serum, lymph, cerebrospinal fluid (CSF), urine, or other bodily fluids. The term "cfDNA" is in contrast to "circulating tumor DNA" or "ctDNA." Cell-free DNA (cfDNA) is a broader term that describes DNA that circulates freely in the bloodstream, but is not necessarily derived from a tumor.
[0093] As used herein, the term "single nucleotide polymorphism" or "single nucleotide variation" ("SNP" or "SNV"), with respect to a mutation, refers to a difference of at least one nucleotide in a sequence compared to another sequence, and the term "copy number variation" or "CNV" refers to a comparative numerical change in the presence, absence, or deletion of a gene fragment having an identical nucleotide sequence.
[0094] The term "indels," as used herein and generally in the art, refers to a genomic location where one or more bases are present in one allele and no bases are present in the other allele. Although insertions and deletions are distinct from an evolutionary perspective, analyses such as those described herein often do not distinguish an insertion in one allele as equivalent to a deletion in the other allele. Thus, the term "indel" refers to the location of an insertion / deletion between two alleles.
[0095] "Structural mutation" refers to a change in the portion of a chromosome, rather than a change in the number or set of chromosomes in the genome. There are four general types of mutations that result in structural mutations: deletions and insertions, e.g., duplications (changes in the amount of DNA in a chromosome, deletions and gains of genetic material), inversions (changes in the arrangement of chromosomal segments), and translocations (changes in the position of chromosomal segments that can result in gene fusions). In the present invention, the term "structural variant" includes losses of genetic material, gains of genetic material, translocations, gene fusions, and combinations thereof.
[0096] As used herein, the term "germline DNA" or "gDNA" refers to DNA isolated or extracted from a patient's peripheral mononuclear cells, including lymphocytes, which in turn are obtained from circulating blood.
[0097] The term "mutation" as used herein refers to a change or deviation. With respect to nucleic acids, mutation refers to a difference(s) or change between DNA nucleotide sequences, including copy number differences (CNVs). This actual difference in nucleotides between DNA sequences can be SNPs and / or changes in the DNA sequence, such as fusions, deletions, additions, repeats, etc., observed when comparing the sequence with a reference, such as germline DNA (gDNA) or the reference human genome HG38 sequence. Preferably, mutation refers to a difference between a cfDNA sequence and a control DNA sequence not derived from tumor cells, such as when cfDNA is compared with the reference HG38 sequence; when cfDNA is compared with gDNA. Differences identified in both gDNA and cfDNA may be considered "constitutional" and disregarded.
[0098] The term "locus" (plural "locuses") corresponds to an identified position in a genome and can span a single base or a contiguous series of bases. A locus is typically identified using a discrimination value or range of discrimination values for a reference genome and / or its chromosome. For example, a discrimination value range of "5100001" to "5800000" refers to a specific location on chromosome 1 in a reference human genome. A "heterozygote locus" (also called a "het") is a locus in a genome where the two copies of the chromosome do not have identical sequences. The different sequences at a locus are called "alleles." A het can be a single nucleotide polymorphism (SNP) if a location in a reference genome has two alleles that differ by only one base. A "het" can also be a location in a reference genome where there is an insertion or deletion (collectively known as an "indel"). The term "homozygous locus" refers to a locus in a reference or baseline genome where two copies of a chromosome have the same allele, and the "haplotype" of a chromosome refers to whether the chromosome is present once or twice in the genome; in the genomes of cancer cells and other tumor cells, the haplotype of a chromosome may be non-integer or may be greater than two. A "region" in a genome may contain one or more loci.
[0099] "Fragment" refers to a nucleic acid molecule (e.g., DNA) contained in or derived (e.g., via amplification) from a biological sample extracted from a target organism, e.g., a human. A fragment can include an entire chromosomal arm, an entire chromosome, or a portion thereof.
[0100] "Fragment size" refers to the length of the fragment and can be expressed in any acceptable units, such as base pairs or daltons. Exemplary fragments may be less than 200 bps, 200-500 bps, 500-1 Kb, where 1 Kb = 1000 bps, 1 Kb-10 Kb, 10 Kb-50 Kb, 50 Kb-100 Kb, and longer than 100 Kb, e.g., 1 megabase pair. Sequencing is used to determine information identifying one or more sequences (reads) of nucleotides in a fragment. Partial and complete sequence information for a fragment can be generated. Sequence information can be determined with varying degrees of statistical confidence or reliability.
[0101] As used herein, the term "variant allele frequency" (VAF) or "variant allele fraction" refers to the fraction of one allele relative to the total amount of alleles in a DNA sample after genotyping. Traditionally, for biallelic polymorphic variants (PV), VAF refers to the B allele frequency (BAF), which is the proportion of the B allele in PV typing data, which can be obtained from a DNA sample by high-throughput genotyping methods, such as SNP arrays or NGS. In some embodiments, VAF is the B-allele frequency. Alternatively, A allele frequency (AAF) can be used as well. The B allele frequency contains the information of the A allele frequency, and vice versa.
[0102] Generally, VAF values are expressed using a value between 0 and 1 because they refer to frequencies or fractions. In principle, VAF values can be expressed using the multiplicity of the value, e.g., using a value between 0 and 100. For example, a VAF value of 0.5, indicating that half of the total amount of alleles have a polymorphic variant allele, can be expressed as, e.g., 50. In this case, a VAF value of 1 (i.e., all alleles have a particular genotype) is expressed as 100. Typically, VAFmax indicates the maximum VAF value (i.e., all alleles have a particular genotype), and VAFminindidis indicates the minimum VAF value (i.e., none of the alleles have a particular genotype). Throughout this application, VAF (especially BAF) values are expressed using a value between 0 and 1, so VAFmin is 0 and VAFmax is 1. Nevertheless, embodiments of the present invention are not limited to VAF values expressed using this particular range. Detailed guidance regarding VAF, including "flip" VAF, is described in US 2016 / 0210402.
[0103] As used herein, a "read" refers to a set of one or more data values representing one or more nucleotide bases. A read may be generated by a sequencing device and / or associated logic that has sequenced all or a portion of a nucleic acid fragment. A "mate pair" (also referred to as a "mated read" or "paired end read") refers to at least two reads (also referred to as "arm reads") determined from opposite ends of the same fragment. Two arm reads (also referred to as "arm reads") can be collectively referred to as a mate pair. Two arm reads have a gap between them with respect to the fragment from which they were sequenced. Two arm reads can be individually referred to as a "left" arm read and a "right" arm read. However, it is understood that any "left" (or right) designation does not strictly limit the position of the arm to the left (or right). Fragments may be reported with respect to various reference points, such as the orientation of the observer, the orientation of the DNA strand (e.g., from the 5' end to the 3' end or vice versa), or a genomic coordinate system selected for a reference genome. Reads may be stored with various information, such as a unique read identifier, a fragment identifier, or a mate pair identifier for reads that are part of a mate pair.
[0104] As used herein, "artifact" refers to an observation in scientific research or experimentation that does not occur in nature but arises as a result of preparatory or exploratory procedures. Artifacts in sequencing include, for example, artifacts (shadow bands) and template-related artifacts (false terminations). Artifacts refer to peaks seen in separations that do not correspond to correctly sized fragments terminated by the respective dideoxynucleotide triphosphates (ddNTPs), which are used in the Sanger dideoxy method to generate DNA strands of different lengths for DNA sequencing. Artifacts can be subdivided into primer-induced artifacts and template-induced artifacts. Primer-related artifacts arise when the primers used have affinity for binding to other regions of the template, resulting in the formation of DNA fragments unrelated to the intended sequence. In contrast, termination artifacts occur as a result of DNA polymerase being disengaged from the template before the ddNTPs are included. It is believed that secondary structures of the template DNA are involved in this false transcription termination. DNA polymerases also have a finite periodicity in associating with the template, called processivity, and the frequency of short processive reactions is thought to increase the number of artifacts. For example, Taq DNA polymerase has a processive reaction of approximately 40 base pairs and is thought to contain no primer-related artifacts. When DNA polymerase encounters a ddNTP, it prevents the extending chain from extending, and if DNA chain extension is terminated without a ddNTP, false termination can occur during Sanger chain termination.
[0105] The term "allele" refers to one of two or more alternative nucleotide sequences present at a particular genetic locus.
[0106] "Allelic fraction" refers to the proportion of one or more alleles at a particular locus in a genome sequenced from nucleic acid fragments contained in a biological sample. Exceptionally, diploid organisms, such as the human Y chromosome, typically have two copies of each chromosome. Thus, loci in a genome can typically be either homozygous (e.g., have the same allele on both chromosome copies) or heterozygous (e.g., have different alleles on the two chromosome copies). Thus, "equal allelic fraction" refers to a data value of 1.0 (e.g., 100% allelic fraction of alleles at a homozygous locus) or 0.5 (e.g., 50% allelic fraction of alleles at a heterozygous locus).
[0107] "Variable allele fraction" or "VAF" refers to a data value greater than zero but different than 0.5 and 1.0. Variable allele fraction values can be used to address situations in which alleles for a given locus may be represented in nucleic acid fragments of a biological sample at fractions greater than 0%, 50%, and 100%. Such situations include, but are not limited to, heterogeneity, contamination, and aneuploidy. For example, a tumor sample (e.g., a cancer sample) may be heterogeneous due to normal / stromal tissue contamination within the sample or multiple different tumor populations within the same tumor sample. In another example, a tumor sample may be aneuploid, such that a chromosome (or region thereof) has a copy number different from two, causing the allele fraction to deviate from 50% for a given locus to 33% or 66% if three copies are present. Examples of variable allele fraction values include, but are not limited to, the following ranges of values and / or combinations of ranges: 0.005 to 0.10, 0.10 to 0.20, 0.20 to 0.30, 0.30 to 0.40, 0.40 to 0.49, 0.51 to 0.60, 0.60 to 0.70, 0.70 to 0.80, 0.80 to 0.90, 0.90 to 0.99, and more commonly, values of 0.005 to 0.49 and 0.51 to 0.99.
[0108] The term "control," as used herein, refers to a reference for a test sample, such as control DNA isolated from peripheral blood mononuclear cells and lymphocytes (where the cells are not cancer cells), and a "reference sample" refers to a sample of tissue or cells that may or may not have cancer, used for comparison. Thus, a "reference" sample provides a basis against which another sample, such as a plasma sample containing cfDNA, can be compared. In contrast, a "test sample" refers to a sample that is compared to a reference or control sample. The reference sample does not need to be cancer-free, as would be the case if the reference and test samples were obtained from the same patient separated in time.
[0109] In some embodiments, the reference sample or control may comprise a reference assembly. The term "reference assembly" refers to a digital nucleic acid sequence database, such as the Human Genome (HG38) database (assembled December 2013), which includes the HG38 assembly sequence. The gateway may be accessed via the Human (Homo sapiens) University of California, Santa Cruz (UCSC) Genome Browser Gateway at the world-wide web URL GENOME(dot)UCSC(dot)EDU at GENOME(dot)UCSC(dot)EDU. Alternatively, the reference assembly may refer to the Genome Reference Consortium's Human Genome Assembly (Build #38; assembled June 2017), accessible on the internet via the National Center for Biotechnology Information (NCBI) website.
[0110] As used herein, the term "sequencing" or "sequencing" as a verb refers to the process by which the nucleotide sequence of DNA, or the order of nucleotides, is determined, such as the nucleotide order AGTCC. The term "sequence" as a noun refers to the actual nucleotide sequence resulting from sequencing, e.g., DNA having the sequence AGTCC. While "sequencing" may be provided and / or received in digital form, e.g., on a disk or remotely via a server, "sequencing" refers to a collection of DNA that is grown, manipulated, and / or analyzed using the methods and / or systems of the present disclosure.
[0111] As used herein, "substantially" means sufficient to function for its intended purpose. Thus, the term "substantially" allows for small, insignificant variations from an absolute or perfect state, dimension, measurement, result, etc., as would be expected by one of ordinary skill in the art, but which do not affect overall performance. When used in reference to a number or a parameter or characteristic that can be expressed as a number, "substantially" means within 10%.
[0112] As used herein, the term "substantially purified" refers to cfDNA molecules that have been removed from their natural environment, isolated or separated or extracted, and are at least 60% free, preferably 75% free, more preferably 90% free, and most preferably 99% free from other components that are naturally associated with them.
[0113] The term "whole genome sequencing" refers to the laboratory process of determining the DNA sequence of each DNA strand in a sample, and the resulting sequence can be referred to as "raw sequencing data" or "read". As used herein, a read is "mappable" if its sequence is similar to that of a region of a reference chromosomal DNA sequence. The term "mappable" refers to a region that shows similarity to a reference sequence and is therefore "mapped", for example, a segment of cfDNA that shows similarity to a reference sequence in a database, for example, a cfDNA that has a high percentage of similarity with the human chromosomal region 8q248q24.3 in the Human Genome (HG38) database is a "mappable read".
[0114] In addition to "WGS," genome catalogs can be obtained using targeted sequencing. In contrast to WGS, the term "targeted sequencing," as used herein, refers to an experimental process of determining the DNA sequence of one or more selected DNA loci in a sample, for example, determining the sequence of a selected group (e.g., target) of cancer-related genes or markers. In this context, the term "target sequence" refers to a selected target polynucleotide, for example, a sequence present in a cfDNA molecule whose presence, quantity, and / or nucleotide sequence, or alteration, is desired to be determined. The target sequence is examined for the presence or absence of somatic mutations. The target polynucleotide can be a region of a gene associated with a disease, for example, cancer. In some embodiments, the region is an exon.
[0115] As used herein, the term "low abundance" with respect to cfDNA refers to less than about 20 ng / mL, e.g., about 15 ng / mL, about 10 ng / mL, or less, e.g., about 9 ng / mL, 8 ng / mL, 7 ng / mL, 6 ng / mL, 5 ng / mL, 4 ng / mL, 3 ng / mL, 2 ng / mL, 1 ng / mL, 0.7 ng / mL, 0.5 ng / mL, 0.3 ng / mL, or less, e.g., 0.1 ng / mL or 0.05 ng / mL. In some embodiments, the term "low abundance" can be understood in the context of the uniqueness of a marker, e.g., length or base composition. For example, a subject's sample may contain an abundant amount of cfDNA (e.g., >20 ng / mL), but the actual number of unique genetic markers (e.g., sSNVs) contained in the cfDNA may be very low. Typically, this parameter is expressed as genome equivalent (GE) or coverage, as described below. In some embodiments, the term "low abundance" can be understood in the context of the tumor specificity of the marker. For example, a subject's sample may contain an abundant amount of cfDNA (e.g., >20 ng / mL), but most of the genetic markers (e.g., sSNVs) contained in the cfDNA may be redundant and / or may be related to a reference (e.g., PBMC gDNA). Typically, this parameter is expressed as a tumor fraction, as described below.
[0116] As used herein, the terms "tumor-specific" or "tumor-associated" with respect to cfDNA refer to differences in the DNA sequence of cfDNA in a subject with cancer, such as a lung cancer patient, when the cfDNA is compared to reference DNA, such as when the cfDNA is compared to control DNA (gDNA) from cells that are not tumorous, as described herein.
[0117] The term "genomic equivalent" or "GE" as used herein refers to the number of unique DNA fragments. In one embodiment, a sample contains 5 to about 10,000 GE, preferably 100 to about 5,000 GE, particularly about 200 to about 2,000 GE, e.g., about 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,200, 1,400, 1,600, 2,000, or 5,000 GE. As understood in the art, a typical sample containing about 6 ng of cfDNA contains about 1,000 or less GE. Preferably, the GE is greater than 1 (e.g., greater than 2, 5, 10, 15, 20, 25, 50, 100, 200, 500, or 1,000). It is believed that 10 to 20 ml of blood contains about 10,000 GE. Thus, in some embodiments, a suitable sample may comprise about 20 ml, 15 ml, 10 ml, 5 ml, 4 ml, 3 ml, 2 ml, 1 ml, 0.5 ml, 0.1 ml, 0.01 ml, or 0.001 ml of plasma.
[0118] The terms "coverage" or "read depth" refer to sequencing effort. For example, 20X coverage refers to a moderate sequencing effort, 35X or more coverage refers to a high sequencing effort, and 5X coverage refers to a low sequencing effort. In embodiments of the present disclosure, coverage is typically about 5X to about 100X, particularly 15X to about 40X, e.g., 20X, 30X, 35X, 40X, 50X, 70X, or more.
[0119] As used herein, the term "mutation burden" or "N" refers to the level, e.g., number, of alterations (e.g., one or more genetic alterations, particularly one or more somatic alterations) per preselected unit (e.g., per megabase pair) in a given genomic window. Mutational burden may be measured, for example, on a whole-genome or exome basis, or based on a subset of the genome or exome. In certain embodiments, the mutational burden measured on a subset of the genome or exome may be extrapolated to determine the whole-genome or exome mutational burden. In certain embodiments, mutational burden is measured in a sample from a subject, e.g., a subject described herein, such as a tumor sample (e.g., a lung tumor sample, or an obtained or derived sample). Preferably, mutational burden is a measure of the number of mutations per megabase pair (1,000,000 bp or MBP) of cfDNA. As known in the art, mutational burden may vary depending on tumor type, genetic lineage, and other subject-specific characteristics, such as age, sex, and tobacco consumption. For tumor diagnosis, the mutational burden can be about 1,000 to about 10,000 mutations per MBP, e.g., about 1,000, 2,000, 4,000, 6,000, 8,000, 10,000, 12,000, 15,000, 20,000, 25,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 10,000, or more, e.g., about 200,000 mutations per MBP. Typically, the mutational burden is about 8,000 / MBP in non-smokers and greater than 40,000 / MBP in subjects with melanoma.
[0120] As used herein, the term "genomic window" refers to a region of DNA within selected nucleotide sequence boundaries. Windows are separated from each other and overlap each other.
[0121] As used herein, the term "tumor fraction" or "TF" refers to the level, e.g., amount, of tumor DNA molecules relative to normal DNA molecules. In one embodiment, "tumor fraction" refers to the ratio of circulating cell-free tumor DNA (cfDNA) to the total amount of cell-free DNA. The tumor fraction is considered to indicate the size of the tumor. Typically, the tumor fraction (TF) is about 0.001% to about 1%, for example, about 0.001%, 0.05%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1% or more, for example, 2%.
[0122] The term "abundance" can refer to binary (e.g., absent / present), qualitative (e.g., absent / low / medium / high), or quantitative information (e.g., a value proportional to number, frequency, or concentration) indicating the presence of a particular molecular species. In this context, mutations present at higher relative concentrations are associated with more malignant cells, e.g., cells transformed early in the tumorigenesis process, compared to other malignant cells in the body (Welch et al., Cell, 150:264-278, 2012). Due to their high relative abundance, such mutations are expected to have higher diagnostic sensitivity for detecting cancer DNA than mutations with lower relative abundance.
[0123] As used herein, "sequencing error probability" refers to the inaccurate proportion of sequenced nucleotides. For example, in the context of whole genome sequencing, a sequencing error probability of approximately 1 / 1000 bases is reported in the literature (range: error probability is on the order of 0.1-1% per base call; see Wu et al., Bioinformatics, 33(15):2322-2329, 2017).
[0124] The term " sequencing depth " used herein refers to the number of times that sequencing area is covered by sequence reading.For example, the average sequencing depth of 10 times means that each nucleotide in the sequencing area is covered by 10 sequence readings on average.It is expected that the greater the sequencing depth, the greater the probability of detecting cancer-related mutations.However, in reality, the probability of detection does not increase linearly with sequencing depth, as evidenced by the fact that even at a median depth of 42,000X, the fundamental limitations of cfDNA abundance only lead to the positive detection of early-stage lung adenocarcinoma at only 19% (Abbosh et al., Nature, 545(7655):446-451, 2017).
[0125] As used herein, the term "base quality" score refers to the probability that a given base in a sequencing read is miscalled by the sequencer. Each base in a read is assigned a quality score using a Phred-like algorithm (Ewing et al., Genome Res. 8(3):175-185, 1998; Ewing et al., Genome Res. 8(3):186-194, 1998, representative methods are described). This was similar to the method originally developed for Sanger sequencing experiments. In some embodiments, base quality (BQ) includes variable base quality (VBQ) or mean read base quality (MRBQ), both of which are variations of the base quality metric.
[0126] As used herein, the term "PCR error" refers to errors introduced through the polymerase chain reaction (PCR) amplification step in sequencing. A typical PCR error rate is approximately 1 error per 10 base pairs (Barnes et al., PNAS USA, 91:2216, 1994).
[0127] As used herein, the term "mapping quality" score indicates the confidence that a particular sequence read is accurately positioned relative to a reference sequence. A method for determining a mapping quality score is provided by Li et al. Genome Research, 18:1851-1858, 2008. A mapping quality score can be provided by a mapping algorithm after mapping a read sequence to a reference sequence.
[0128] The term "read position" or "read position (PIR)" refers to the position of a read (e.g., a marker) in a nucleotide sequence. As understood in genomics, many sequencing protocols are prone to various types of amplification-induced bias and errors, which can be reduced by implementing filters such as the "read direction" and "read position" filters. The read direction filter removes variants that are present almost exclusively in either the forward or reverse read. In many sequencing protocols, such variants are most likely the result of amplification-induced errors. The read position filter is implemented in a similar manner to the "read direction filter" to remove systematic errors, but is also suitable for hybridization-based data. It removes variants located within a read that differ from those expected based on the general position of the read covering the mutation site. This is done by sorting each sequenced nucleotide (or gap) by the mapping direction of the read and where the nucleotide is found in the read; each read is divided into portions (e.g., 5 portions) along its length, and the portion number of the nucleotide is recorded. This results in a total of 10 categories for each nucleotide sequenced, and a given site will be distributed among these 10 categories for the reads that cover that site. If a variant is present at this site, the variant nucleotides are expected to follow the same distribution. The read position filter performs a test to measure the significance of the read position, for example, whether the read position distribution of a variant differs from that of the entire set of reads that cover the site.
[0129] As used herein, the term "bin" refers to a group of DNA sequences grouped together, such as a "genomic bin." In certain cases, a bin may include a group of DNA sequences binned based on a "genomic bin window," which involves grouping DNA sequences using a genomic window.
[0130] By way of example only, and reviewing the detailed description below, various embodiments herein relate to algorithms and software involved in implementing the diagnostic engine (engine) of the present disclosure. The engine utilizes a read representation that simultaneously captures the genomic context of the alignment, the complete read sequence, and an integrated base-by-base quality score. In contrast, representations used in sequence analysis software known in the art consider the pile of reads as a single feature, losing valuable information about the sequence alignment itself and the base-by-base quality associated with the read (Poplin et al., bioRxiv, pp. 092890, 2016; Torracinta & Campagne, bioRxiv, pp. 097469, 2016).
[0131] 〔method〕 The disclosed systems and methods are useful in the diagnosis, prognosis, and monitoring of various human diseases. For example, numerous cancers can be detected using the methods and systems described herein. Most cancer cells are characterized by a high rate of turnover, where old cells die and are replaced by new cells. Generally, dying cells can release DNA or fragments of DNA into the bloodstream upon contact with the vasculature of a given subject. This is true even when cancer cells are at various stages of the disease. Depending on the stage of the disease, cancer cells can also be characterized by various genetic abnormalities, such as copy number variations and mutations. This phenomenon can be used to detect the presence or absence of cancer in individuals using the methods and systems described herein.
[0132] According to the present disclosure, blood may be collected from a subject at risk for cancer and prepared as described herein to generate a population of cell-free polynucleotides. In one example, the population may include cell-free DNA. The systems and methods of the present disclosure may be used to detect markers (e.g., SNVs, CNVs, indels, and / or SVs) present in certain cancers. The methods may be useful for detecting the presence of cancer cells in the body despite the absence of symptoms or other features of disease. The methods of the present disclosure may be applied to the diagnosis or prognosis of any type of cancer or tumor. Accordingly, types of cancer that may be detected include, but are not limited to, blood cancer, brain cancer, lung cancer, skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, and solid tumors. Both heterogeneous and homogeneous tumors may be diagnosed or prognosed according to the disclosure.
[0133] The present systems and methods can be used to detect any number of genetic abnormalities that may lead to or result in cancer, including, but not limited to, mutations, alterations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural variations, gene fusions, chromosomal fusions, gene truncations, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, nucleic acid methylation, infections, and cancers. Furthermore, the systems and methods described herein can also be used to aid in the characterization of specific cancers. Genetic data obtained from the disclosed systems and methods can enable practitioners to better characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can enable characterization of specific subtypes of cancer, which may be important in diagnosing or treating that particular subtype. This information may also provide clues to subjects or practitioners regarding the prognosis of a particular type of cancer. The systems and methods provided herein can be used to monitor a known cancer or other disease in a particular subject. This allows either the subject or a medical practitioner to adapt treatment options as the disease progresses. In this example, the systems and methods described herein can be used to build a genetic profile of a particular subject's disease progression. In some cases, the cancer may progress and become more aggressive and genetically unstable. In other instances, the cancer may remain benign, inactive, or dormant. The systems and methods of the present disclosure can be useful for determining disease progression.
[0134] Furthermore, the systems and methods described herein may be useful in determining the effectiveness of a particular treatment option. In one example, a successful treatment may actually increase the amount of copy number mutations or mutations detected in the patient's blood, as more cancer cells die and release DNA. In another example, this does not occur. In another example, perhaps a particular treatment option may be correlated with the cancer's genetic profile over time. This correlation may be useful in selecting a treatment. Furthermore, if a cancer is observed to be in remission after treatment, the systems and methods described herein may be useful in monitoring residual disease or disease recurrence.
[0135] The methods and systems described herein are not limited to detecting mutations and copy number variations associated with cancer alone. Preferably, the methods and systems of the present disclosure are useful for early diagnosis or detection of cancer.
[0136] Furthermore, the disclosed methods can be used to characterize heterogeneity of an abnormal condition in a subject, the method comprising generating a genetic profile of extracellular polynucleotides in the subject, the genetic profile comprising multiple data obtained from copy number variation and mutation analysis. Diseases, including but not limited to cancer, can be heterogeneous. Disease cells may not be identical. In the example of cancer, some tumors contain different types of tumor cells, and some cells are known to be at different stages of cancer. In other examples, heterogeneity can include multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps the result of one or more foci metastasizing from the primary site.
[0137] The disclosed methods can be used to generate or profile data that is the sum of genetic information from different cells in heterogeneous diseases. The datasets can include copy number variation and mutation analysis, alone or in combination. Furthermore, the disclosed systems and methods can be used to diagnose, prognose, monitor, or monitor cancer or other diseases of fetal origin. That is, the methods can be used in pregnant subjects to diagnose, prognose, monitor, or monitor cancer or other diseases in fetal subjects that have DNA and other polynucleotides that may co-circulate with maternal molecules.
[0138] The above diagnostic methods may be used in combination with other common diagnostic procedures, such as health checkups, physical examinations, laboratory tests (blood, urine, etc.), biopsy imaging tests (e.g., X-ray, PET / CT, MRI, ultrasound, etc.), nuclear medicine scans (e.g., bone scans), endoscopies, family history, etc.
[0139] Preferably, the diagnostic methods of the present disclosure improve the predictive prognostic value (PPV) of a common diagnostic procedure (e.g., a CT scan) by at least 20%, at least 30%, at least 40%, or more (e.g., at least 50%).
[0140] Representative, non-limiting schematic diagrams of diagnostic methods are shown in Figures 1, 2 and 7 of the drawings. [Work procedure] 1A is a flowchart illustrating a method 100 for diagnosing a neoplastic disease, e.g., an early-stage neoplastic disease, according to various embodiments of the present disclosure. Method 100 is merely exemplary, and embodiments may employ variations of method 100. Method 100 may include receiving a collection of markers, filtering noise associated with the markers based on multiple features, applying a convolutional neural network trained on an in silico dataset and / or a patient dataset to adaptively and systematically filter the noise, removing artifactual noise markers from the collection to generate subject-specific markers, which are statistically matched to the dataset for generation of confidence intervals, and diagnosing the disease based on the confidence intervals.
[0141] In step 110 of method 100 of FIG. 1A, a list of genetic markers is received from a subject. In one embodiment, the list of genetic markers is received in a Variant Call Format (VCF) file. As understood in the art, VCF files are used in bioinformatics to store genetic sequence variations. The VCF format was developed with the advent of large-scale genotyping and DNA sequencing projects, such as the 1000 Genomes Project. Alternatively, the list can be provided in a Generic Feature Format (GFF), which contains all of the genetic data. GFFs generally provide redundant features because they are shared genome-wide. In contrast, VCF only stores variations along with the reference genome. In one embodiment, the subject's sample is sequenced, for example, using whole genome sequencing (WGS), and the sequence file is processed using a tool such as genomic VCF (gVCF).
[0142] In step 120 of method 100 of FIG. 1A, artifactual reads are filtered by statistically classifying each read as signal or noise based on the probability of detecting noise as a function of (1) base quality (BQ), (2) mapping quality (MQ), (3) estimated fragment size, and / or (4) estimated allele fraction (VAF). Other parameters, such as (5) read position (RP); (6) sequence context (SC); (7) abundance; (8) sequencing depth, and / or (9) sequencing error, can also be used. The noise removal step 120 can include implementing a best-fit receiver operating characteristic curve (ROC) that includes a probabilistic classification of the genetic markers in the list based on the combined base quality score and mapping quality score. Typically, the combined BQMQ score is provided as a matrix (x, y), where x is the BQ score and y is the MQ score. In exemplary embodiments, a combined BQMQ score of 10-50 (for each parameter) is typically used, e.g., BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), and (30, 40). In certain aspects, marker classification involves measuring the area under the ROC curve (AUC), which typically represents the probability that a randomly selected candidate marker from among potential markers will exhibit a higher value than a randomly selected control marker. For completely uninformative markers, the ROC curve approaches the ascending diagonal (referred to as the "chance diagonal" or "chance line"), and the AUC is 0.5 (i.e., the expected probability of classification by chance alone). Conversely, for perfect classification, the ROC curve peaks at theoretical accuracy (both sensitivity and specificity are 100%), and the AUC tends to be 1, i.e., the highest probability value. A representative ROC is shown in Figure 3B. The pre-filtration error model is shown in Figure 3A, and the post-filtration error model is shown in Figure 3C.
[0143] Optionally, genetic markers are weighted. In some aspects where the markers are SNVs or CNVs, a weighting step is performed to distinguish between true markers (e.g., mutations that are likely to be associated with the disorder) and common mutations (e.g., random somatic SNPs that are not associated with the disorder). In some embodiments, the weighting step weights the markers based on a probability score (PD). Preferably, the weighting step 120 is performed using the Bernoulli equation PD=1-[(1-TF)] GE The weighting step may further include binning markers based on increasing weight or weight range, where PD is the probability of detection, TF is the tumor fraction, and GE is the number of genome equivalents present in the patient DNA. Performing a weighting step is advantageous because sequencing depth is overcome as the number of detection sites (SNVs) increases through spread resulting from repeated Bernoulli trials for each SNV (binomial distribution for Bernoulli trial probability). In some embodiments, the weighting step may further include binning markers based on, for example, increasing weight or weight range. For example, genetic markers may be binned based on PD, where markers with high PD may be binned separately from markers with low PD. For example, genetic markers may be binned based on a PD threshold of at least about 0.60, e.g., at least about 0.65, 0.70, 0.75, 0.80, 0.90, 0.95, or higher, e.g., at least about 0.98. Thus, if a marker's PD is below the threshold, it is classified as a false positive and is not included in the analysis.
[0144] In step 130 of method 100 of FIG. 1A, a machine learning (ML) approach is used to filter sequencing noise in each read in the pamphlet in order to distinguish between cancer-associated mutational signatures and PCR- or sequencing error-related signatures. In certain embodiments, the diagnostic methods of the present disclosure may utilize a neural network to systematically remove or reduce noise. A neural network may be applied at any step of the method, although it is advantageous to implement the neural network after artifactual markers have been removed according to step 120 above. In this regard, in the purely exemplary method 100 of FIG. 1A, a deep convolutional neural network is optionally applied in step 130 to adaptively and / or systematically filter sequencing noise present in the filtered dataset. Preferably, the CNN involves employing deep learning algorithms on a pan-tumor cohort to identify features that distinguish between true tumor mutations and artifactual errors, assigning a confidence estimate to each individual mutation detected in samples from tumor patients, integrating the confidence estimates across the entire genome, and rigorously analyzing the characteristics of specific cosmic mutations in the samples, e.g., using non-negative least squares for each marker.
[0145] In some embodiments, the CNN is trained on an in silico dataset. For example, the in silico dataset may include synthetic plasma samples obtained from a cohort of actual cancer patients, such as breast cancer or lung cancer patients. The accuracy, sensitivity, and / or precision of the CNN may be evaluated according to the methods described below. For example, sensitivity may be determined as the ratio [TP / (TP+FN)], where TP is true positive and FN is false negative; accuracy may be determined as the ratio [TP / (TP+FP)], and specificity may be determined as the ratio [TN / (TN+FN)], where TN is true negative and FN is false negative. Under a typical validation method, the accuracy of the CNN may be evaluated based on the average F1 score. For example, the F1 score may be calculated as 2×[(accuracy×recovery) / (accuracy+recovery)]. In certain embodiments, the CNN may achieve an F1 score of at least about 0.5, about 0.6, about 0.7, about 0.8, or about 0.9 or greater in tumor controls, e.g., 0.95.
[0146] In one embodiment, the CNN can be trained on an in silico patient-specific dataset containing tumor and normal WGS reads mixed in various proportions across different tumor fractions (0.00001, 0.00005, 0.0001, 0.0005, 0.005, 0.01) and coverages (5, 10, 15, 20, 35). Repetition and / or random seeds can be further used to increase the variability of the training dataset. The structure of CNN will be described later.
[0147] In step 140 of method 100 of FIG. 1A, a subject-specific signature comprising multiple true reads in the list is compiled by removing artifactual noise (see step 120) and / or sequencing noise (see step 130). Without being bound by any particular theory, in one aspect, the removal step filters "noise" markers with low base quality and / or mapping quality from a list of markers initially identified as strongly associated with disease. In one embodiment, the removal step may include taking each marker that meets a threshold probability of detection (PN) based on step 120, classifying the marker as signal or noise based on an ROC curve, and removing the marker from the list if classified as noise. Alternatively, a scoring system including, for example, the ratio of probability of detection (PD) to probability of noise (PN) may be used to remove markers that do not meet a preset threshold score.
[0148] In step 150 of method 100 of Figure 1A, a match is made between the subject-specific pattern and the cancer pattern to quantify a confidence estimate that the subject's biological sample contains circulating tumor DNA. This can be accomplished, for example, using probability density function (PDF) estimates and / or z-score estimates, both of which are described in more detail below.
[0149] A weighting step can optionally be used in estimating the confidence interval. For example, all markers classified as true positives based on the noise removal step 120 and noise filtering step 130 can be weighed identically. For example, a pan-tumor network can use a modified weighing system based on the scores assigned to the markers. Diagnosis can further include using a threshold score, e.g., a score obtained based on performing the same noise removal step 120 and noise filtering step 130, on test markers, e.g., markers known to be associated with tumors. For example, the test markers can include unique SNVs and / or CNVs in cancer patient samples that are not present in control (non-tumor) subjects.
[0150] Further, various embodiments provide a method for genetic screening of a subject for cancer, as provided by the exemplary workflow 100 illustrated in FIG. 1B. As provided in step 110, the method can include receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject. The biological sample can include a tumor sample. The list of reads can each include reads of a single base pair in length.
[0151] As provided in step 120 of method 100 of FIG. 1B, the method can include filtering actual sites from the read list. Filtering can include removing repetitive sites generated on a cohort of reference healthy samples from the read list. Alternatively, or in combination, filtering can include identifying germline mutations in the biological sample and / or identifying shared mutations between the tumor sample and peripheral blood mononuclear cells of the normal cell sample as germline mutations and removing the germline mutations from the read list.
[0152] As provided in step 130 of method 100 of FIG. 1B, the method may include using at least one error suppression protocol to filter noise from a genome-wide inventory of reads to generate a filtered read set for the genome-wide inventory of reads. The at least one error suppression protocol may include calculating the probability that any single nucleotide variation in the inventory is an artifact and removing the variation. The probability may be calculated as a function of features selected from the group including mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof. Alternatively, or in combination, the at least one error suppression protocol may include removing artifacts using mismatch testing between independent copies of the same DNA fragment generated from polymerase chain reaction or sequencing processing, and / or overlap consensus, in which artifacts are identified and removed in the event of mismatches across a large portion of a given overlap family.
[0153] As provided in step 140 of method 100 of FIG. 1B, the method may include compiling subject-specific patterns using the filtered read set based on a comparison of specific mutation patterns associated with a given mutagenesis process.
[0154] As provided in step 150 of method 100 of FIG. 1B, the method can include statistically quantifying a confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern based on a comparison of the cancer-associated mutation pattern exposure value to a cohort of background mutation patterns via a feature specific to the subject.
[0155] As provided in step 160 of method 100 of FIG. 1B, the method can include screening the subject for cancer if a confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern exceeds a given threshold.
[0156] Further, as provided by the exemplary workflow 100 shown in FIG. 1C, a method provides for genetic screening of a subject for cancer. As provided in step 110, the method can include receiving a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample of the subject. The biological sample can include a tumor sample. The read list can each include copy number variations (CNVs).
[0157] As provided in step 120 of method 100 of FIG. 1C, the method may include dividing the reading list into multiple windows.
[0158] As provided in step 130 of method 100 of FIG. 1C, the method may include calculating a set of features per window. The features may include a median depth coverage per window and a representative fragment size per window.
[0159] As provided in step 140 of method 100 of FIG. 1C, the method can include filtering actual sites from the read list. Filtering can include removing repetitive sites generated on a cohort of reference healthy samples from the read list.
[0160] As provided in step 150 of method 100 of FIG. 1C, the method may include normalizing the read list to generate a filtered set of reads for the genome-wide read list.
[0161] As provided in step 160 of method 100 of Figure 1C, the method may include calculating the estimated tumor fraction using the filtered read set by calculating a linear relationship between the set of features per window and converting the calculated relationship to an estimated tumor fraction using a regression model. Alternatively, or in combination, the method may include calculating the estimated tumor fraction based on one or more integrative mathematical models as a function of the calculated set of features per window across the subject-specific genome-wide list of reads.
[0162] As provided in step 170 of method 100 of FIG. 1C, the method can include screening the subject for cancer if the estimated tumor fraction exceeds an empirical threshold.
[0163] Exemplary Workflow for Performing Marker Type-Based Screening Methods 1D and 1E show schematic workflows for implementing the methods of the present disclosure. FIG. 1D outlines a workflow typically used when the marker of interest includes SNVs / indels, and FIG. 1E outlines a workflow typically used when the marker of interest includes CNVs / CVs. Note that while the separate workflows are provided for illustrative purposes, they are not required to be separate for implementing the methods of the present disclosure. For example, certain features / elements of the workflows may be used in combination to generate an output (e.g., a combined estimated tumor fraction based on SNVs / indels and CNVs / SVs) related to an outcome of interest (e.g., whether a subject develops cancer).
[0164] [SNV-based cancer screening] The present disclosure provides systems, methods, and algorithms for cancer screening based on the detection of SNV / indel markers in a subject's biological sample. As shown in Figure 1B, cancer diagnosis based on SNV / indel markers typically utilizes the following steps to improve the sensitivity, specificity, and / or reliability of detection: receiving genetic data; detecting mutations (e.g., single mismatches); removing / filtering artificial and real sites; suppressing errors using algorithms including machine learning; correcting reads; detecting cancer based on one or more mathematical models; and, optionally, orthogonally integrating analysis of secondary features in the genomic data.
[0165] In the first step of Figure 1D, genetic data from a biological sample (typically a plasma sample) is received. Next, sensitive variant calling is performed on the plasma sample using PILEUP (or other single-supported caller). Germline SNPs are detected using the GATK germline calling device on the plasma sample or variant calling on matched peripheral blood mononuclear cells (PBMCs). Buccal swabs may be used instead of PBMCs. Serially or in parallel, recurrent artifacts are generated across a cohort of normal plasma samples (blacklist or masked normal panel (PON)) and removed from the detected variants to eliminate common sequencing or alignment artifacts.
[0166] Next, a highly sensitive method capable of detecting a single mutant fragment is used. This process includes one or more error suppression steps. In the first error suppression step, a filtering scheme is used to analyze single read bases and quantify the probability that the read represents an artifact. In one embodiment, a multidimensional classification framework using a support vector machine (SVM) classifier with a linear kernel may be implemented in this process. The classification framework is trained on germline SNPs compared to low variant allele fraction (VAF) sequencing artifacts in normal PBMC samples. Here, classification decision boundaries are defined in a multidimensional space, including variant base quality (VBQ), mapping quality (MQ), read position (PIR), and / or mean read base quality (MRBQ). To evaluate the classification scheme, the validation metrics of the SVM classification scheme were compared with random forests under the same protocol after 10-fold cross-validation. The SVM classifier demonstrated high classification performance, slightly outperforming the random forest model. SVM achieved an average sensitivity of 90.7% and specificity of 83.9% across all patients (N = 10 samples, F1 = 87.7%, PPV = 84.9%).
[0167] In the second error suppression step, artifacts introduced by PCR or sequencing were corrected using comparison of independent replicates of the same original DNA fragment. In cfDNA samples, 150 bp of the paired ends were typically sequenced, and given the short size of typical cfDNA fragments (approximately 165 bp), duplicated paired reads (overlapping R1 and R2 sequences) were obtained. Therefore, discrepancies between R1 and R2 pairs were considered potential sequencing artifacts that were matched back to the corresponding reference genome. Furthermore, to recognize the potential generation of independent duplicates due to any DNA molecules copied multiple times during sequencing and PCR, duplicate families were identified by 5' and 3' similarity as well as alignment position. Each duplicate family was then used to check for consensus of specific mutations across the independent replicates, correcting for artifacts that showed no match across the majority of the duplicate family.
[0168] The resulting set of reliable de novo plasma mutations is used to identify tumor patterns using one or more identification steps. A first method involves using a mutation pattern inference method, such as non-negative least squares (NNLS), to identify tumor patterns in the resulting set. The method outputs a confidence score (e.g., z-score) that can be used to determine whether a subject has cancer. In this regard, a threshold confidence score (e.g., a z-score of about 2) can be used to make a reliable determination that the subject has cancer. A second method can be used that utilizes deep learning methods for detecting mutation patterns. The method outputs a tumor fraction score (e.g., eTF) that can be used to determine whether a subject has cancer. This method is described in more detail below.
[0169] [Cancer-specific mutation induction pattern] Cancer mutagenesis is dominated by sequence-specific signatures associated with different mutagenesis processes, such as tobacco smoking and UV radiation. These mutational signatures are specific to cancer tissue and are not expressed in normal PBMC samples. Here, we demonstrate that genetic patterns are differentially expressed in lung cancer patients (exposed to tobacco) and melanoma patients (exposed to UV radiation) compared with normal samples (PBMCs). Recognizing this feature, we developed a novel, highly sensitive detection analytical method. The method is based on a non-negative least squares (NNLS) model of specific mutation patterns in a single plasma sample. Signature detection was further validated for reliability using a comparison of the cancer-specific mutation signature exposure values to exposure values estimated for 100 random background signatures, with a confidence threshold of z-score >2 std.
[0170] [Detection of mutation patterns using deep learning] To further suppress sequencing artifacts and increase ctDNA sensitivity, a machine learning method was developed to distinguish between cancer-altered sequencing reads and those altered by sequencing errors, enabling adaptive and specific filtering of whole-body sequencing noise. We applied a deep convolutional neuron network (CNN) based on artificial intelligence technology. CNNs allow for the learning and integration of multiple features in a supervised manner for classification problems. This rapid approach relies on rethinking the mutation calling challenge to distinguish between reads containing true variants and those containing sequencing artifacts. This allows for training the CNN on millions of true variant reads and errors using a large collection of tumor and normal WGS data, achieving extremely high sensitivity and specificity across a variety of patient and tumor types.
[0171] The implementation of these features in deep CNN training results in the independent capture of sequence context patterns known to occur in lung cancer and melanoma. First, to apply CNN to an early detection (ED) framework, we trained a CNN algorithm on a pan-lung cancer cohort (five patients with deep tumors and PBMC WGS) and utilized supervised learning to identify features that distinguish true tumor mutations from artifacts. The resulting model was used to infer and assign confidence estimates to individual mutations detected in ED plasma samples from early-stage lung adenocarcinoma patients. These estimates can be integrated into estimates of tumor read rates in a given sample. The model was able to identify specific tobacco and / or UV patterns, which, when applied to patient samples, allowed for highly accurate detection of early-stage cases of each cancer.
[0172] Furthermore, the ability of this method to improve the low positive predictive value (PPV) of current lung cancer CT screening in at-risk tobacco-exposed populations was evaluated by applying the method to plasma samples from 21 patients with early-stage lung cancer and 12 patients with CT-detected benign nodules. Results showed 14 positive detections in early-stage lung cancer samples and 3 positive detections in benign nodules, representing an 80% improvement in PPV compared with 40%-50% for current CT-based screening schemes. These data demonstrate a significant improvement over existing methods for the early detection of lung cancer and melanoma patients.
[0173] [Integration of orthogonal characteristics] In some cases, the basic workflow described above may orthogonally integrate secondary features contained in the genetic data into the final analysis model. For example, to improve the robustness, accuracy, and / or sensitivity / specificity of the detection method, read-based features, such as shifts in DNA fragment size, may be orthogonally integrated into the mathematical model. The significance of the orthogonal feature integration (in cancer detection) may be calculated using a probabilistic mixture model (e.g., a Gaussian mixture model). See the Examples section and corresponding data in Figures 17 and 18.
[0174] [CNV cancer screening] Alternatively or additionally, the present disclosure provides systems, methods, and algorithms for cancer screening based on the detection of CNV / SV markers in a subject's biological sample. As shown in Figure 1E, cancer detection based on CNV / SV markers typically utilizes steps of receiving genetic data, extracting window-based feature vectors in the genetic data, filtering artificial real CNV windows, normalizing the filtered genetic data using one or more normalization steps, detecting tumors after feature vector segmentation, and, in some cases, orthogonally integrating analysis of secondary features in the genomic data (e.g., fragment size shift analysis) to improve detection sensitivity, specificity, and / or reliability.
[0175] In the first step of Figure 1E, genetic data from a biological sample (typically a plasma sample) is received. Next, windowed feature vectors are extracted from the genetic data. For example, depth coverage characteristics (represented by Log2) and / or fragment size characteristics (represented by COM) are extracted. The estimated Log2 and COM values of all windows across the genome are used to determine the median center of mass of the sample (median COM relative to the neutral region), and the slope and R^2 of the Log2 / COM linear model are calculated. Additionally, split reads can also be extracted. Split reads typically occur when a portion of an NGS read maps to one location in the genome and another portion of the same read maps to a different location in the genome, resulting in discrepancies.
[0176] Next, the windows with low mappability and / or coverage are filtered. Subsequently, the artificial sites are generated and the artificial windows are filtered across a cohort of healthy plasma samples (blacklist or masked normal panel (PON)) that have been removed from the windows, either serially or in parallel. The filtered high-confidence CNV / SV segments are normalized. Typically, the normalization step involves guanine-cytosine (GC) normalization and / or z-score normalization.
[0177] The feature vectors are then segmented using one or more mathematical models. In one embodiment, a Hidden Markov Model (HMM) is used. In one embodiment, a Self-Organizing Neural Network (SONN) based on a mathematical model such as Adaptive Resonance Theory (ART) or Self-Organizing Maps (SOM) is used. The segmented data is analyzed using one or more of the mathematical models for copy number variation (CNV) detection and cancer diagnosis.
[0178] Again, it is possible to orthogonally integrate secondary features of the genetic data into the final analysis model, for example, to improve the robustness, accuracy, and / or sensitivity / specificity of the detection method. Log2 / COM correlation (R^2), Log2 / COM slope, and center of mass (COM) of sample median fragment size may be taken together to define a classification model between tumor and healthy samples, and estimated TFs may be calculated using, for example, a generalized linear model (GLM). The significance of the orthogonal feature integral (in cancer detection) can be calculated using a probabilistic mixture model (eg, a Gaussian mixture model).
[0179] It should be understood that the workflow disclosed herein, with certain modifications, may also be used generally in the detection of residual disease during or after chemotherapy, immunotherapy, targeted therapy, or a combination thereof, and / or in the process of monitoring the effectiveness of such treatment.
[0180] [Use of the above method for early tumor diagnosis] The disclosed methods are particularly useful for early diagnosis of tumors. Preferably, the disclosed diagnostic methods are performed non-invasively. The diagnostic methods can be performed before surgery or treatment of the tumor.
[0181] The disclosed method can be performed even with a low tumor fraction (TF). Generally, samples with lower TF have a lower probability of detection, and prior art methods cannot accurately and reliably diagnose tumor diseases. In contrast, the disclosed method allows for the detection of markers and accurate diagnosis of tumor diseases even with low tumor fractions, e.g., 1 / 1000, 1 / 10,000, or 1 / 20,000. The sensitivity of the disclosed method and system is particularly demonstrated by the fact that even with very low tumor fractions (e.g., 1 / 10,000 or less), the disclosed method detects approximately 10 to 15 sSNVs contained in a single support read. This detection can distinguish between normal and tumor samples with a high level of fidelity and accuracy not achievable by prior art means. It should be understood that diagnosis is not limited to sSNV detection. For example, diagnosis can be based on detection of about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more, e.g., 150, 200, or 250 copies of the altered segment (genome-wide) frequently observed in human cancers.
[0182] The present disclosure relates in particular to a method for early diagnosis of tumors characterized by a high rate of somatic mutations. Tumor types that can be diagnosed or detected by the present disclosure preferably include, for example, non-small cell lung cancer (NSCLC), tobacco-induced cancer (TIC), UV light-induced cancer, cancers mediated by apolipoprotein B mRNA-editing enzyme catalytic protein (APOBEC) activity, cancers containing breast cancer protein (BRCA) mutations and / or cancers containing increased poly(ADP-ribose) polymerase (PARP) activity, and tumors containing microsatellite instability (MSI). The method may be applied to diagnose liquid tumors, solid tumors, or a mixture thereof, including heterogeneous tumors, such as lymphomas that have metastasized to extra-lymphatic organs, such as the liver, lung, or brain.
[0183] The following tumors can be diagnosed particularly early by the present invention: lung adenocarcinoma, ductal adenocarcinoma (breast tumor), skin melanoma, urothelial carcinoma (bladder tumor) or osteosarcoma. In particular, tumors include non-small cell lung cancer and lung adenocarcinoma (NSCLC LUAD).
[0184] The present disclosure particularly relates to the early diagnosis or detection of non-small cell lung cancer, preferably tobacco-induced cancer of the lung, which is characterized by a high rate of somatic mutations. Smoking (e.g., smoking or chewing) is a well-established risk factor or causative agent for epithelial cancers of the oral cavity, pharynx, larynx, esophagus, lung, stomach, cervix, and colon / rectum. See Sasco et al., Lung Cancer 45, Suppl 2, S3-9, 2004.
[0185] The present disclosure also relates to the diagnosis or detection of ultraviolet light-induced cancers, such as skin cancer. Exposure to ultraviolet (UV) light is associated with approximately 65% of melanoma cases and 90% of non-melanoma skin cancers (NMSCs), including basal cell carcinoma (BCC) and squamous cell carcinoma (SCC). See Kim et al., Genes & Disease, 1(2):188-198, 2014. Preferably, the UV-induced cancer is selected from melanoma and SCC, both of which are characterized by a high rate of somatic mutations. See Alexandrov et al., Curr Opin Genet Dev. 24, 52-60, 2014.
[0186] The present disclosure also relates to early diagnosis of cancers with high somatic mutation rates due to perturbations in gene editing / DNA checkpoint-related enzymes. In certain embodiments, the present disclosure relates to diagnosis of cancers mediated by gene editing enzymes, such as apolipoprotein B mRNA editing enzyme catalytic proteins (APOBEC). APOBEC-mediated mutation patterns are common in bladder, cervical, breast, head and neck, and lung cancers. See Roberts et al., Nat Genet., 45(9):970-6, 2013.
[0187] In one embodiment, the present disclosure relates to early diagnosis of cancers mediated by breast cancer protein (BRCA) mutations, such as BRCA1 or BRCA2 mutations, or a combination thereof. Reports estimate that more than 50% of women with BRCA1 mutations will develop breast cancer by age 70, and more than one-third of them will develop ovarian cancer by that age. In addition to breast and ovarian cancer, BRCA2 mutations are associated with male breast and pancreatic cancer, as well as melanoma. Both BRCA1 / 2 mutations are associated with male prostate cancer risk. See Ngeow et al., npj Genomic Medicine 1, 15006, 2016.
[0188] In one embodiment, the present disclosure relates to the early diagnosis of cancers induced by microsatellite instability (MSI). MSI-induced cancers generally result from mutations in DNA mismatch repair genes (e.g., MLH1, MSH2, or MSH6) and are characterized by errors in repetitive sequences. MSI can occur in tumors of many organs but is primarily a hallmark of colorectal cancer. See Kurzawski et al., Annals of Oncology, 15(Supp. 4), 283-284, 2004. MSI has also been observed in endometrial cancer, ovarian cancer, gastric cancer, sebaceous gland carcinoma, glioblastoma, lymphoma / leukemia, and Lynch syndrome (hereditary nonpolyposis colorectal cancer (HNPCC)) tumors. See Vilar et al., Nat Rev Clin Oncol., 7(3):153-62, 2010.
[0189] In certain embodiments, the present disclosure relates to the early diagnosis of cancers induced by PPAR activity, for example, through the compensatory homologous recombination activity of PARP. For example, tumors with defects in the homologous recombination mechanism depend on PARP-mediated DNA repair for survival and are sensitive to its inhibitory effect on PARP. Therefore, PARP inhibition is a potential synthetic lethal therapeutic strategy for treating cancers with specific DNA repair deficiencies, such as those occurring in carriers of BRCA1 or BRCA2 mutations (Morales et al., Crit Rev Eukaryot Gene Expr., 24(1):15-28, 2014; Fong et al., N Engl J Med., 361(2):123-34, 2009).
[0190] The diagnostic method of the present disclosure first includes receiving a subject sample containing a plurality of genetic markers. In some embodiments, the subject sample containing DNA / RNA is sequenced, and the genetic markers therein are received for analysis. In other embodiments, the genetic markers can be received from a dataset, e.g., genome sequencing information compiled and / or stored on a computer or stored remotely (e.g., on a server). The genetic markers can be received by sequencing various samples. Preferably, the sample includes a biological sample, e.g., an organ containing cells, tissue, or biological fluid, e.g., blood, plasma, lymph, etc. Alternatively, the sample includes a primary or metastatic tumor.
[0191] Samples can be obtained using a variety of methods. Tissue biopsies are often used to obtain a representative piece of tumor tissue. Tumor cells can also be obtained indirectly in the form of tissues or bodily fluids known or suspected to contain tumor cells from a subject. For example, biological samples of lung cancer lesions can be obtained by resection, bronchoscopy, fine needle aspiration, bronchial brushing, or from sputum, pleural effusion, or blood. Metastatic tumors can be sampled from surrounding tissues or lymph nodes (primary metastasis) or from distant sites (distant metastasis).
[0192] The sample preferably comprises a plasma sample containing circulating DNA and peripheral blood mononuclear cells (PMBC). In this context, the sample can be obtained from the subject using conventional techniques, such as blood sampling (phlebotomy), biopsy (including liquid biopsy), surgical resection, tracheal swab, sputum, etc. The sample thus obtained can optionally be processed to, for example, purify and / or isolate markers useful for diagnosis. The presence of cfDNA in the sample can be determined using conventional methods, such as PCR using universal primers, followed by electrophoresis. The cfDNA in the sample of the subject can be purified using conventional techniques, such as the DNA isolation kit described in the Examples section of this disclosure.
[0193] In certain embodiments, the sample comprises a biological fluid selected from blood, cerebrospinal fluid, pleural fluid, ocular fluid, urine, or a combination thereof.
[0194] In one embodiment, a sample containing somatic mutations in cfDNA is obtained using liquid biopsy technology (LBT), a transforming, non-invasive technique that allows for the detection of tumor DNA in a patient's plasma cfDNA sample and characterization of the somatic malignant genome.
[0195] In a specific embodiment, the biological sample is a plasma sample containing cell-free DNA (cfDNA). Typically, the amount of cfDNA in the sample is about 0.1 ng / ml to about 20.0 ng / ml, preferably about 1 ng / ml to about 10 ng / ml. A normal cell sample containing peripheral mononuclear blood cells (PMBCs) can be used as a control. Both samples can be analyzed for genetic markers, including single nucleotide variations (SNVs) (preferably somatic SNVs), copy number variations (preferably somatic CNVs), short insertions and deletions (indels), structural variants (SVs), or combinations thereof.
[0196] In some embodiments, the genetic markers include a combination of SNVs and CNVs. Such a combination is typically used in samples with low SNV mutation burden and high CNV burden. In an exemplary embodiment, samples with an SNV mutation burden of less than 8,000 mutations per megabase pair (MBP) can be analyzed to additionally detect CNVs. Typically, in such cases, a CNV burden of at least 60, at least 70, at least 80, at least 90, at least 100 or more, e.g., 200, CNVs / MBP of DNA is desirable, as it is likely to be diagnostically significant.
[0197] As is well known in the field of genetics, the significance of mutations, e.g., SNVs or CNVs, is greatly influenced by the distinction between germline and somatic cells. Mutations in somatic cells (body cells) are not passed on to offspring. Mutations occurring in somatic cells, such as those in the lungs, can damage, cancerize, or even kill cells. However, mutated DNA can only be passed on to future generations if it is present in the germline of gametes. Therefore, germline sequences can be compared (e.g., used as a control) to identify alterations in somatic or cancerous cells specific to a subject that are not present in non-cancerous cells of the same subject. While comparison of germline and gamete sequences indicates mutations, comparison of cancerous and non-cancerous cells is also useful. For example, a subject's peripheral leukocytes or lymphocytes can be used as a control, representing non-cancerous somatic sequences. In this way, mutations found in both cancerous and non-cancerous cells can be ignored.
[0198] Preferably, the genetic markers of the present disclosure, such as sSNVs, sCNVs, indels or SVs in cfDNA, can be detected by comparing the cfDNA sequence to a reference sequence, such as a germline DNA sequence.
[0199] In some embodiments, the method of the present disclosure may include detecting variations between genetic markers and a control (e.g., control) sequence. In some embodiments, variations may be uniform, semi-uniform, or dynamic between samples. Temporally dynamic variations include, for example, differences between cfDNA collected during or after treatment and a sample before treatment.
[0200] Mutations in cfDNA can also be detected by creating a genome-wide list of genetic markers and then subtracting the genetic markers present in a control (e.g., germline) sample. In this context, the term "genome-wide" refers to and includes the genetic material of an organism's germline and somatic cells. The list of markers can include, for example, multiple sSNVs, sCNVs, indels, SVs, and other variations such as fusions in DNA.
[0201] Typically, the specimen is characterized by a low tumor fraction (TF), in one embodiment, the TF is between about 0.0001% and about 1%, preferably between about 0.001% and about 0.1%, particularly less than 0.1%, e.g., 0.005%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%.
[0202] Additionally, samples containing cfDNA are characterized by a genomic equivalent (e.g., the number of unique DNA fragments measured via random sampling of the entire pool of cfDNA fragments in a subject's sample) of between about 100 and about 20,000, preferably between about 1,000 and about 10,000. In certain embodiments, the cfDNA sample is characterized by a mutational burden (N) of about 3,000 to about 100,000, preferably about 5,000 to about 40,000.
[0203] Exemplary methods for generating a genome-wide inventory can include sequencing. Typically, sequencing is performed using purified nucleic acid samples. In particular, genome-wide inventory used in the diagnostic methods and / or systems of the present disclosure is achieved using whole-genome sequencing. For example, WGS can be performed using conventional techniques and can include amplification (e.g., PCR amplification). Amplification-free sequencing can also be used using methods and reagents known in the art. See Karlsson et al., Genomics, 105(3):150-8, 2015. Purely by way of example, in one embodiment, genetic markers in cfDNA can be detected by whole-genome sequencing (WGS) of a subject's tumor, whole-genome sequencing (WGS) of the subject's normal cells, mixing the tumor and normal WGS in various ratios to generate subject-specific sample datasets with different tumor fractions and coverages, downsampling the datasets, not mixing in reads from the tumor, and generating a complementary dataset of downsampled normal reads. The complementary data set can be filtered to remove noise-associated markers as described below.
[0204] A genome-wide inventory of genetic markers can also be generated by targeted sequencing (TS) or by combining WGS with TS. The following U.S. Patent Nos. 7,115,400, 7,718,403, 7,741,463, 8,932,812, 7,572,584, and 9,218,450 relating to whole genome sequencing and / or targeted sequencing are incorporated herein by reference in their entirety.
[0205] Once a DNA sample is received, a diagnostic method can be performed. Genetic markers contained in the sample are preferably analyzed for mutations, e.g., somatic mutations. The most common type of somatic mutation in DNA is a single nucleotide variant (SNV), occurring at a frequency of 1-100 / Mbp (megabase pair). Such mutations are typically identified in shotgun sequencing data through careful comparison of DNA sequencing reads that map to specific loci in cancer samples and germline normal DNA samples (controls). This complex process is being developed using ever-more sophisticated technologies / tools that refine statistical comparisons between the number of supportive reads and mutations in cancer and germline samples. For reference, Cibulskis et al., Nature Biotechnology, 31(3):213-219, 2013; Saunders et al., Bioinformatics, 28(14): 1811-1817, 2012; Wilm et al., Nucleic acids research, 40(22): 11189-11201, 2012.
[0206] Mutation analysis can be performed using a variety of technologies, including, but not limited to, array-based methods (e.g., DNA microarrays, etc.), real-time / digital / quantitative PCR instrument methods, and whole nucleic acid sequencing systems (e.g., whole genome sequencing (WGS) services provided by Illumina, Helicos Biosciences, Pacific Biosciences, Complete Genomics, Sequenom, ION Torrent Systems, and Halcyon Molecular).
[0207] Preferably, genetic markers are analyzed for somatic mutations and / or copy number variations using whole genome sequencing (WGS). Whole genome sequencing can decipher genetic readouts at single-base resolution. In the context of DNA (deoxyribonucleic acid), this method deciphers at the level of the basic components of DNA, such as A (adenine), T (thymine), C (cytosine), and G (guanine). In the context of RNA (ribonucleic acid), this method deciphers at the level of the basic components of DNA, such as A, U (uracil), G, and C.
[0208] The products of the sequencing methods include "sequencing data," "sequencing information," or "sequencing reads," which contain information about the order of one or more of the bases in a polynucleotide molecule (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, a fragment, etc.). The order of reads of DNA in a sample (e.g., cfDNA contained in a patient plasma sample) can be compared to a control (e.g., a whole genome sequence of PMBC) to identify genetic markers of interest (e.g., somatic SNVs or somatic CNVs). It should be understood that the identification methods of the present disclosure are applicable to all types of sequencing technologies, platforms, or techniques, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion-based or pH-based detection systems, electronic pattern-based systems, etc.
[0209] The next step in the disclosed early diagnostic method involves identifying low abundance tumor-specific markers.
[0210] The present disclosure relates to determining the probability of error in a read based on multiple factors selected from (1) the base quality (BQ) of the read, (2) the mapping quality (MQ) of the read, and / or (3) the fragment size of the read, (4) the variable allele frequency (VAF) of the read, which alone or together affect the quality of the signal. Other secondary parameters may also be used, such as (5) the position within the read (RP), (6) the read sequence context (SC), (7) the read abundance, (8) sequencing depth, and / or (9) sequencing error.
[0211] Generally, base quality (BQ) relates to the reliability of the sequencing quality of each base, while mapping quality (MQ) scores relate to a confidence estimate of the accuracy of the marker's mapping to the genome. In the context of sSNV markers, the base quality (BQ) score is a measure of the quality of the nucleobase identification generated by automated DNA sequencing. It can be determined using conventional methods, such as the Phred quality score assigned to each nucleotide base call in an automated sequencer trace. The Phred quality score (Q) is defined as a property logarithmically related to the base call error probability (P). For example, if Phred assigns a quality score of 30 to a base, the probability that this base will be miscalled is 1 in 1000. Typically, the BQ of a sequencing read is between 10 and 50, e.g., a BQ score of 10, 15, 20, 25, 30, 35, or 40.
[0212] Also, in the context of sSNV markers, a mapping quality (MQ) score is a measure of the confidence that a read is actually derived from a position aligned by a mapping algorithm. This can be determined using conventional methods, such as the mapping quality score (see Li et al., Genome Research 18:1851-8, 2008). Typically, the MQ of a read is between 10 and 50, e.g., an MQ score of about 10, 15, 20, 25, 30, 35, or 40.
[0213] In one embodiment, the denoising step involves performing an optimal receiver operating characteristic (ROC) curve involving a probabilistic classification of the genetic markers in the list based on the combined base quality (BQ) and mapping quality (MQ) scores. Typically, the combined BQMQ scores are provided as a matrix (x, y), where x is the BQ score and y is the MQ score. In an exemplary embodiment, combined BQMQ scores of 10-50 (for each parameter) are typically used, such as BQMQ scores of (10, 40), (15, 30), (20, 20), (20, 30), and (30, 40).
[0214] The noise removal process may involve the implementation of additional filters. For example, extra information contained in read pairs derived from DNA fragments can be used to determine replication origins (Watson or Crick) and estimate DNA fragment size. ctDNA has been observed to have a different fragment size distribution than normal circulating healthy DNA (Underhill et al., PLoS Genetics, 12(7):e1006162, 2016). More specifically, fragment lengths obtained from cell-free DNA of tumor patients and healthy controls were found to be shorter for mutant alleles than for wild-type alleles. Similarly, size selection for shorter cell-free DNA fragment lengths substantially increased mutant allele frequencies in human lung cancer (Jiang et al., PNAS USA, 112.11, E1317-E1325, 2015; Mouliere et al., bioRxiv, 134437, 2017; Underhill, supra). Therefore, a specific subset of fragment lengths from cell-free DNA detection can be used to improve ctDNA detection. In one embodiment, the fragment size of the readout is preferably less than 160 bp, e.g., 160 bp, 140 bp, 120 bp, 100 bp, 75 bp, 50 bp, or less, e.g., 20 bp.
[0215] Furthermore, artifactual noise can be removed based on variable allele frequency (VAF). In some embodiments, low allele fraction mutation sites are removed from the sample, for example, VAF is about 1% or less. In some embodiments, only markers (e.g., SNVs) with a threshold VAF are retained for downstream analysis. For example, mutation sites with a VAF of at least 1%, the last 2%, at least 3%, at least 4%, or at least 5% (as determined by amplicon sequencing on a PGM instrument) can be retained. As known in the art, the VAF value of a particular allele (e.g., BRAF V600R) is not static and may change over time (due to cancer development and / or progression) and with treatment, for example, immunotherapy, chemotherapy, or targeted therapy. However, a threshold VAF of less than 1%, for example, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, or 0.8, can be used to reliably infer that a particular allele is not associated with tumors.
[0216] In one specific embodiment, artifactual noise is removed by performing one or more, preferably all, of the following steps: (a) removing reads with low mapping quality (e.g., <29, ROC optimized); (b) constructing duplicate families (e.g., representing multiple PCR / sequencing copies of the same DNA fragment) and generating corrected reads based on consensus testing; (c) removing reads with low base quality (e.g., <21, ROC optimized); and / or (d) removing reads with high fragment size (e.g., >160, ROC optimized).
[0217] In addition to using the BQ / MQ, VAF, and fragment size filters described above, other factors, such as the read position (RP or PIR), can be used to filter artifacts, since RP affects signal quality. In the context of sSNV markers, RP can be mapped, for example, by mapping the first base position of a sequencing read. Other factors that affect marker quality include specific sequence contexts, which are associated with a higher probability of sequencing errors (Chen et al., Science, 355(6326):752-756, 2017). In this regard, true mutations are often mappable to their own specific sequence contexts, while errors are not. For example, tobacco-related mutations tend to occur in CC contexts, while mutations associated with the activity of APOBEC enzymes prefer TpC contexts to insert somatic mutations (see Greenman et al., Nature, 446(7132):153-158, 2007). Thus, sequence context helps identify changes that are likely to be due to sequencing artifacts and those that are likely to be due to dominant mutational processes.
[0218] In one embodiment, the markers are determined by the Bernoulli equation PD=1−[(1−TF)] GE The probability of detection can be further measured by determining the probability of detection based on PD, where PD is the probability of detection, TF is tumor fraction, and GE is the number of genome equivalents present in patient DNA.Then, genetic markers are weighted based on PD, and markers with high PD are bound.For example, genetic markers can be bound based on a PD threshold of at least about 0.60, for example, at least about 0.65, 0.70, 0.75, 0.80, 0.90, 0.95 or more, for example, at least about 0.98.Therefore, if the PD of a marker is below the threshold, it will be classified as a false positive and will not be included in the analysis.
[0219] Once the artifactually noisy reads are removed from the read list, the remaining markers are fed into a deep learning inference model trained to separate tumor-related features from PCR / sequencing error features. In this stage, a read-based method is used to classify reads that support cancer mutations from artifactually mutated (error) reads. In one embodiment, the sequence-context distribution of cancer mutation-supporting reads is calculated, and the contribution of known mutation patterns is classified using machine learning.
[0220] Reads that have been noise filtered for artifacts and / or classified as supported by cancer mutations are matched to cancer patterns. In one embodiment, a dataset containing the cancer patterns (e.g., Catalogue of Somatic Mutations in Cancer (COSMIC)) can be used. As of February 2018, 30 different cancer patterns have been registered in the database, and their details are as follows:
[0221] Pattern 1 (found in all types of cancer) is the result of an endogenous mutational process initiated by spontaneous deamination of 5-methylcytosine; Pattern 2 (seen in 22 cancer types) is due to activity of the AID / APOBEC family. Based on similarities in the sequence context of cytosine mutations induced by APOBEC enzymes in experimental systems, a role for APOBEC1, APOBEC3A, and / or APOBEC3B in human cancer seems more likely than for other members of this family; Pattern 3 (breast, ovarian, and pancreatic cancers) is associated with failure of DNA double-strand break repair by homologous recombination; Pattern 4 (head and neck cancer, liver cancer, lung adenocarcinoma, lung squamous cell carcinoma, small cell lung cancer, and esophageal cancer) is associated with smoking, and its profile resembles the mutation patterns observed in experimental systems exposed to tobacco carcinogens (e.g., benzo[a]pyrene). Pattern 4 is likely due to tobacco mutagens; Pattern 5 (etiology unknown) is found in all cancers and most cancer samples; Pattern 6 (seen in 17 cancer types, most commonly in colorectal and uterine cancers) is associated with defects in DNA mismatch repair and is seen in microsatellite-unstable tumors; Pattern 7 (skin and lip cancer; head and neck cancer or oral squamous cell carcinoma) is associated with UV exposure; Pattern 8 (seen in breast cancer and medulloblastoma) is of unknown etiology; Pattern 9 (seen in CLL and aggressive B-cell lymphomas) is caused by polymerase η, which is associated with AID activity during somatic hypermutation; Pattern 10 (found in six cancer types, particularly colorectal and uterine cancers) is due to altered activity of the error-prone polymerase POLE. Recurrent somatic POLE mutations, Pro286Arg and Val411Leu, are primarily associated with pattern 10 mutations; Pattern 11 (seen in melanoma and glioblastoma) shows a mutation pattern similar to alkylating agents; Pattern 12 (seen in liver cancer) is of unknown etiology; cytidine 13, resulting from the activity of the AID / APOBEC family of cytidine deaminases that convert cytidine to uracil (found in 22 cancer types, and appears to be most frequent in cervical and bladder cancers); Pattern 14 (etiology unknown) was found in four uterine cancer and one adult low-grade glioma specimen; Pattern 15 (seen in some gastric cancers and a single small cell lung cancer) is associated with a defect in DNA mismatch repair; Pattern 16 (seen in liver cancer) is of unknown etiology; The etiology of pattern 17 (seen in esophageal cancer, breast cancer, liver cancer, lung adenocarcinoma, B-cell lymphoma, gastric cancer, and melanoma) is unknown; Pattern 18 (found in neuroblastoma and also observed in breast and gastric cancer) is of unknown etiology; Pattern 19 (seen in pilocytic astrocytoma) is of unknown etiology; Pattern 20 (seen in gastric and breast cancers) is associated with defects in DNA mismatch repair; The etiology of pattern 21 (seen in gastric cancer) is unknown; Pattern 22 (seen in urothelial (renal pelvis) cancer and liver cancer) is associated with exposure to aristolochic acid; The etiology of pattern 23 (seen in liver cancer) is unknown; Pattern 24 (seen in a subset of liver cancers) is associated with aflatoxin exposure; Pattern 25 (seen in Hodgkin lymphoma) is of unknown etiology; Pattern 26 (seen in breast, cervical, gastric, and uterine cancers) is associated with DNA mismatch repair; Pattern 27 (seen in a subset of renal clear cell carcinoma) is of unknown etiology; Pattern 28 (seen in gastric cancer) is of unknown etiology; Pattern 29 (seen in gingivobuccal squamous cell carcinoma) is associated with smoking; Pattern 30 (seen in areas of breast cancer) is of unknown etiology.
[0222] In one embodiment, the matching process involves a linear mixed optimization (e.g., z-score confidence estimates of the contributions from tobacco exposure or BRCA mutations or APBEC1 activity) and is used to calculate a confidence measure for the contribution of the COSMIC mutation patterns. As a purely representative, non-limiting example, the linear optimization problem may be solved using the algebraic function minllAx-bll,x≧0, where A is the mutation pattern sequence context matrix, x is the contribution of each COSMIC mutation pattern (variable), and b is the patient-specific sequence context list.
[0223] In one embodiment, in the linear optimization method used above, A can include any number of COSMIC patterns, including randomly mutated patterns. For example, A can include about 20, 30, 40, 50, or more, e.g., 70 COSMIC features and about 50, 60, 80, 100, or more, e.g., 150, randomly mutated features. The distribution of random pattern contributions can be calculated using a sampling method. For example, E_random calculates an average contribution score; and std_random calculates a standard contribution score. The reliability associated with the contribution of each COSMIC feature can be calculated statistically, e.g., using a z-score. For example, the Z-score can be calculated as (cosmic_sig_contribution-E_random) / std_random. Thus, as with the permutation score, the Z-score represents the significance of the pattern contribution compared to a random set.
[0224] In one embodiment, the similarity of a patient sequencing context list to a particular COSMIC label is calculated using a statistical method, such as a probability density function (PDF). As a purely representative example, the patient sequencing context list is normalized to generate a density function and the PDF is calculated. The cosine similarity between the patient sequencing context density function and the COSMIC pattern density function is calculated. The cosine similarity is then normalized by dividing it by the cosine similarity between the patient sequencing context density function and a non-informative uniform density function.
[0225] In step 160 of method 100 of FIG. 1A, the confidence estimate calculated in step 150 is used to screen subjects for early detection of cancer, e.g., tumors. A confidence interval, as known in the art, consists of a set of values (intervals) that serve as a good estimate of an unknown population parameter (e.g., the likelihood that an asymptomatic subject has cancer). The desired confidence level is set by the researcher (not determined by the data). Most commonly, a 95% confidence level is used, although other confidence levels can be used, e.g., any value between 80% and 99%, e.g., 80%, 90%, 98%, or 99%.
[0226] In some embodiments, the confidence interval may be simple (e.g., based on a single reading) or composite (e.g., based on multiple readings). Confidence bands or confidence intervals may also be used. Confidence intervals generalize the concept of confidence intervals to handle multiple quantities and are useful for revealing the degree of possible sampling error and / or lack of reliability of quantities used in statistical analysis. Confidence bands are used to represent the uncertainty of an estimate of a curve or function based on limited or noisy data, while prediction bands may be used to represent the uncertainty regarding the value of a new data point on the curve (subject to noise).
[0227] In some cases, the calculated reliability metric for the contribution of COSMIC mutation patterns can be checked against a detection threshold. In one embodiment, the threshold is defined by an empirically measured baseline noise detection estimate from healthy samples, for example, a z-score of at least 2 standard deviations (STD) above the threshold, particularly at least 3 standard deviations above the threshold, preferably at least 4 standard deviations above the threshold, particularly at least 5 standard deviations above the threshold.
[0228] For illustrative purposes, and purely by way of example, the disclosed method may include first receiving a plurality of genetic markers sequenced from a subject's biological sample (e.g., a sample including a plasma sample and a normal cell sample) to generate a subject-specific genome-wide genetic read inventory containing the markers (e.g., sSNVs, CNVs, indels, and / or SVs), then filtering artifactual noise from the read inventory using one or more parameters selected from BQ, MQ, position in read (PIR), fragment size, and / or VAF; inputting the denoised reads into a neural network capable of distinguishing from noise generated by PCR and / or sequencing errors; generating a filtered, denoised subject-specific signature that is matched to a cancer signature (e.g., a COSMIC signature), where matching includes calculating z-scores or evaluating probabilities for all markers or a subset thereof, and a density function between the subject's pattern and a reference cancer pattern; and outputting a confidence interval indicating that the subject's pattern contains the tumor pattern to diagnose the subject's tumor. An exemplary method is shown in the flowchart of FIG. 1A. See the example below for details of this method.
[0229] In certain embodiments, the cancer pattern may include a pattern associated with a tissue-specific epigenetic pattern, such as a tissue-specific chromatin accessibility pattern (e.g., methylation status).
[0230] In some embodiments, the diagnostic method may further involve karyotyping. For example, a dataset containing tumor-specific, low-abundance markers may be further karyotyped, for example, by excluding markers proximal to the centrosome. This step may be performed using the mapping techniques described above. Furthermore, a dataset containing low-abundance markers may be orthogonally integrated with aneuploidy markers, for example, markers indicating gene amplification or gene deletion.
[0231] Systems and Devices for Performing Diagnostic / Screening Methods For example, methods described herein, such as method 100, may be implemented using computer system 400 as a standalone device or on a distributed network of shared computer processing resources, such as a cloud computing network. Accordingly, a non-transitory computer-readable medium may be provided having stored thereon a first program for causing a computer to execute the disclosed method of removing artificial noise (e.g., associated with low BQ / MQ markers, markers greater than a threshold fragment size of approximately 160 bp, and markers with a VAF less than a threshold of approximately 4%). A non-transitory computer-readable medium may be provided having stored thereon a second program for adaptively and systematically filtering noise (e.g., associated with PCR / sequencing errors). For example, a non-transitory computer-readable medium may be provided having stored thereon a second program for determining z-scores or analyzing probability density functions to match noise-filtered CNN-processed subject-specific patterns with cancer patterns and outputting a confidence interval (CI) of the match, where a CI equal to or greater than a threshold (e.g., 80%, 90%, 95%, or 99%) indicates that the subject is affected by a tumor. In some embodiments, the first, second, and third programs may each be provided or used separately (e.g., standalone), and in some embodiments, the first, second, and third programs may each be provided or used together (e.g., as a package).
[0232] It should also be understood that the above embodiments may be provided, in whole or in part, as a system of components integrated to carry out the described method. For example, the workflow of Figure 1A may be provided as a system of components or stations to identify high-quality, low-abundance tumor-specific markers present in the cfDNA of cancer patients, further enabling early diagnosis in a sensitive, accurate, and precise manner.
[0233] From the detailed description above, one of the distinguishing features of the systems and methods of the present disclosure is the use of an engine that can adaptively and systematically filter noise. An exemplary engine is described in detail below. The engine may be implemented in the diagnostic method of the present disclosure (discussed in detail below), for example, according to the flowchart of FIG. 1A (note: the positioning of the engine in the flowchart is merely exemplary, so as to fit the exemplary methodology). The engine may include a convolutional neural network (CNN) that can capture invariances within markers (e.g., somatic mutations, including sSNVs). CNNs and their corresponding structures are described in more detail below with reference to the section "Convolutional Neural Networks (CNNs)."
[0234] The engine's ability to remove low-quality markers can be evaluated with synthetic plasma samples and real plasma DNA samples. Synthetic plasma samples can be generated by randomly sampling the patient's healthy DNA and the patient's tumor DNA from a test sample (e.g., a lung sample). For real plasma DNA analysis, plasma samples from smoking lung cancer patients can be used. Patient PMBCs can be used as a control. Alternatively, plasma samples from non-cancer or healthy subjects can be used as controls.
[0235] An exemplary schematic of how machine learning (ML) can be used to detect mutations in a subject's sample while suppressing noise de novo (e.g., errors during amplification (PCR), errors during sequencing, errors during mapping, and other false-positive markers (e.g., mutations found in control samples)) is provided in FIG. 5 . As shown, genetic data is received from a subject in a suitable format (e.g., a variant of VCF format), which may represent true or false positives. The data is input into a machine learning tool, such as an n-dimensional convolutional neural network (CNN). The CNN may have K filters per position, resulting in a total of 32D trainable filters, where D is the number of dimensions in the CNN. The genetic data is maximally pooled, for example, using a size of 2 and a step count of 2. Sequencing reads are captured in a discrete feature representation using any method. For example, a spatially oriented representation containing up to 1, 2, 3, 4, ... n feature lengths may be used.
[0236] Exemplary features are provided in Figure 8. As shown, the first five rows represent the reference context (e.g., a sequence in the human genome), the next five rows represent the read sequence (base pairs in the read), rows 11-15 represent the alignment string (CIGAR), and the last row represents the quality score at each position in the read. Each column of the feature represents an indicator vector that indicates the presence or absence of a particular base. The read, genome context, and CIGAR rows are mutually exclusive, as in one-hot encoding. Details regarding the construction and implementation of the feature are described in the representative examples below.
[0237] The engine may be implemented as a standalone tool or using other known prior art callers, such as PILEUP (Li et al., Bioinformatics, 25(16):2078-2079, 2009), STRELKA (Saunders et al., Bioinformatics, 28(14):1811-1817, 2012), or LOFREQ (Wilm et al., Nucleic Acids Research, 40(22):11189-11201, 2012). An exemplary schematic of the engine's location and inputs / outputs is shown in Figure 7. Note: While the engine is shown at the distal end of the pipeline in this illustration, in practice, the engine may be located at any level or stage of the process. To train the engine, genetic data comprising a collection of markers from mixed tumor biopsy samples and peripheral blood mononuclear cells (PMBCs; control) are optionally subjected to the filters described above (e.g., artificial read cataloging via PILEUP; removing germline mutations using VAF; removing markers with low base quality using an appropriate BQ filter; and removing poorly mapped markers using an appropriate MQ filter). The device may also be trained using the dataset.
[0238] When implemented using independent samples from lung cancer patients, the engine was found to be able to distinguish true somatic mutations from noise with high sensitivity and accuracy. The results are shown in Figures 8 and 9. Experiments performed with synthetic plasma revealed that the engine was particularly accurate and sensitive at low tumor fractions (TFs), outperforming state-of-the-art callers such as MUTECT (Cibulskis et al., Nature Biotechnology, 31(3):213-219, 2013) and / or PILEUP. In particular, the engine demonstrated superior performance in both in silico analysis and clinical settings. The engine performed particularly well in balanced tumor fraction settings compared to programs such as MUTECT. For example, sensitivity metrics showed it outperformed MUTECT, SNOOPER (Spinella et al., BMC Genomics, 17(1):912, 2016), and STRELKA. See Figure 9A. In terms of accuracy metrics, it outperformed PILEUP across all tumor fractions, approximately 25-fold at low TFs (TF = 0.0001). Furthermore, much of its performance was maintained in simulated plasma. The engine also achieved approximately 30-fold enrichment at TFs of 0.0001 (better than PILEUP), suggesting that it can capture relevant somatic mutations at frequencies less than 10-fold lower than the sequencing noise itself (see Figure 9C). In contrast, MUTECT achieved only a modest improvement of approximately 2-fold (compared to PILEUP) across all tumor fractions. Furthermore, the engine allows users to minimize false negatives, and for applications where specificity is a priority, the engine can be configured to minimize false positives. The engine variant identification system can simultaneously minimize false positives and false negatives, detecting mutations with discordance accuracy and precision (see Table 4 for a list).
[0239] In particular, the engine can be applied, in some cases, in conjunction with noise cancellation filters such as a mutation frequency filter and / or a base quality mapping quality filter, to significantly improve the accuracy of variant callers known in the art. The following example describes a representative pipeline using the variant caller PILEUP in conjunction with a downstream noise cancellation filter and engine. In the context of actual plasma samples, the pipeline includes PILEUP, a noise reduction filter (based on mutation frequency (MF) and quality (BQMQ)), and an engine, which significantly enriches samples for tumor DNA analysis while significantly suppressing false positives. Thus, the results demonstrate that the engine can be used to significantly improve variant calling performance, with little, if any, loss of sensitivity.
[0240] The performance of the engine demonstrates that the integration of features across reads and their alignments provides deep coverage and generates a new set of somatic mutation calls using the complete mutation profile of a sample, capturing mutations at a sensitive level using simple measurement tools, enabling new and improved diagnostic platforms that can be used to treat and / or manage cancer patients.
[0241] The present disclosure relates to at least three potential applications of the engine: improved detection of somatic SNV mutations, particularly in cancer diagnosis, prognosis and care, and other clinical settings; improved detection of structural variants for genetic disease diagnosis and disease risk estimation; and / or improved detection of germline genomic SNVs in biomedical research, disease diagnosis, and / or treatment. See Figure 10.
[0242] Based on the current prior art, the engine is the first somatic variant caller designed to work in settings with low allele frequencies, such as liquid biopsies for early cancer detection. To achieve early detection goals, a new list of reads was implemented using a custom structure to best capture the expected features associated with the reads and their alignment. Thus, the present disclosure provides a new family of somatic variant callers that can contribute to detection in liquid biopsies, a critically important, non-invasive method of cancer diagnosis, particularly in the context of early tumor detection and residual tumor detection.
[0243] [Computer Systems] In one embodiment, the diagnostic method of the present disclosure is implemented on a computer system. By way of purely representative example, a schematic diagram of such a computer system is shown in FIG. 15. FIG. 15 is a block diagram illustrating a computer system 400, which may implement portions or multiple embodiments of the present disclosure. In various embodiments of the present disclosure, the computer system 400 may include a bus 402 or other communication mechanism for communicating information, and a processor 404 coupled with the bus 402 for processing information. In various embodiments, the computer system 400 may also include memory, which may be a random access memory 406 or other dynamic storage device coupled to the bus 402, for determining instructions to be executed by the processor 404. The memory may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 404. In various embodiments, the computer system 400 may further include a read-only memory 408 or other static storage device coupled to the bus 402 for storing static information and instructions for the processor 404. A storage device 410, such as a magnetic disk or optical disk, may be provided and coupled to bus 402 for storing information and instructions. In various embodiments, computer system 400 may be coupled via bus 402 to a display 412, such as a cathode ray tube or liquid crystal display, for displaying information to a computer user. An input device 414, including alphanumeric and other keys, may be coupled to bus 402 for communicating information and command selections to processor 404. Another type of user input device is a cursor control device 416, such as a mouse, trackball, or cursor direction keys, which communicates directional information and command selections to processor 404 and controls cursor movement on display 412. This input device 414 typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane. However, it should be understood that input devices 414 that allow for three-dimensional (x, y, and z) cursor movement are also contemplated herein.
[0244] Consistent with certain embodiments of the present disclosure, results may be provided by computer system 400 in response to processor 404 executing one or more sequences of one or more instructions contained in memory 406. The instructions may be read into memory 406 from another computer-readable medium or computer-readable storage medium, such as storage device 410. Execution of the sequences of instructions contained in memory 406 may cause processor 404 to perform the processes described herein. Alternatively, hardwired circuitry may be used in place of or in combination with software instructions to implement the present teachings. Thus, embodiments of the present teachings are not limited to any specific combination of hardware circuitry and software.
[0245] The terms "computer-readable medium" (e.g., data storage device, data storage device, etc.) or "computer-readable storage medium" as used herein refer to any medium that participates in providing instructions to processor 404 for execution. Such media may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical, solid state, and magnetic disks, such as storage device 410. Examples of volatile media include, but are not limited to, dynamic memory, such as memory 406. Examples of transmission media include, but are not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 402.
[0246] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or other magnetic media, CD-ROMs, any other optical media, punch cards, paper tape, any other physical media with patterns of holes, RAM, PROMs, and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, or any other tangible medium from which a computer can read.
[0247] In addition to computer-readable media, data may be provided as signals on a transmission medium included in a communication device or system to provide a sequence of one or more instructions to the processor 404 of the computer system 400 for execution. For example, a communication device may include a transceiver that provides signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to perform the functions outlined in this disclosure. Representative examples of data communication transmission connections include, for example, a telephone modem connection, a wide area network, a local area network, an infrared data connection, an NFC connection, etc.
[0248] It should be understood that the methods described herein, including the flowcharts, diagrams, and accompanying disclosure, may be practiced using computer system 400 as a stand-alone device or on a distributed network of shared computer processing resources, such as a cloud computing network.
[0249] 〔system〕 The present disclosure further relates to systems for implementing the methods of the present disclosure. Representative systems are shown in the schematic diagrams of FIGS. 16A-16C. FIG. 16A illustrates an exemplary system for implementing the diagnostic methods of the present disclosure. As shown herein, a system 500 is provided that may include a data collection unit 510, a marker identification unit 520, a diagnostic unit 550, and a display 412 that outputs data and receives user input via associated input devices (not shown). The marker identification unit 520 may include a noise removal unit 530 and a classification engine 540. Note that FIG. 16A illustrates one configuration of the system. The orientation and configuration of the components may be changed as needed. Furthermore, additional components (e.g., convolutional neural networks) may be added to the system. The various components, their various operations, their various orientations, and various relationships between each other are discussed in detail below.
[0250] The data collection 510 unit of FIG. 16A may be configured and arranged to receive a genetic inventory from a subject, e.g., a plurality of genetic markers sequenced from biological samples, including a plasma sample and a normal cell sample, to generate a subject-specific genome-wide genetic marker inventory. In an embodiment, the genetic marker inventory is received in a Variant Call Format (VCF) file on a physical disk (e.g., a compact disk, DVD) or via the internet (e.g., as provided by a server or cloud). In an embodiment, the subject's sample is sequenced, e.g., using whole genome sequencing (WGS), and the sequence file is sent directly to the data collection unit 510. In an embodiment, the data collection unit 510 may reformat, organize, sort, or otherwise restructure the received data for further analysis within the system 500. In an embodiment, the unit 510 may receive data, for example, via the display 412, data or user input associated therewith, a memory associated therewith, or another memory component associated with the computer system 400.
[0251] The data acquired by the data acquisition unit may be transferred to the marker identification unit 520. The marker identification unit 520 may include one or more engines for analyzing markers in a list of genetic markers specific to the subject. The noise removal unit 530, as one of the components of the unit 520, may include one or more programs for weighing markers based on BQ, MQ, fragment size, and / or VAF, and filtering out artificial noise, including, for example, one or more of intra-read position (RP), sequence context, abundance, sequencing depth, and / or sequencing error. Preferably, the noise removal unit includes a program for calculating an optimal receiver operating characteristic curve, including a probabilistic classification of compendial genetic markers, based on a score integrated with the fragment size score and / or VAF score, such as a joint base quality score and a mapping quality score. The noise removal unit may typically include a program for measuring the area under the ROC curve, which represents the probability that a candidate marker randomly selected from potential markers exhibits a higher value than a randomly selected control marker. A classifier may include a program that evaluates whether a particular binned marker is a "chance" marker or a "true" marker based on a ROC curve.
[0252] In one embodiment, the noise removal unit may weigh the markers based on a probability score (PD). Preferably, the program uses the Bernoulli equation PD=1-[(1-TF)] GE The probability of detection (PD) is measured based on the following formula: where PD is the probability of detection, TF is the tumor fraction, and GE is the number of genome equivalents present in patient DNA.Each genetic marker can be weighted based on PD, and the marker with the highest PD is included.For example, the PD threshold of a genetic marker can be bound based on at least about 0.60, for example, at least about 0.65, 0.70, 0.75, 0.80, 0.90, 0.95 or more, for example, at least about 0.98.Therefore, if the PD of a marker is below the threshold, it can be classified as a false positive and not included in the analysis.
[0253] The marker identification unit 520 may include a classification engine 540 that may, for example, examine the likelihood that a marker is associated with noise. The classifier may include a classification scheme that includes an algorithm or neural network that may adaptively recognize erroneous markers (e.g., errors due to PCR or sequencing). In one specific embodiment, the classification unit 540 includes a deep convolutional neural network that adaptively and / or systematically filters sequencing noise that may affect the accurate detection of tumor-specific low-abundance markers. The CNN may be provided as a separate engine within the marker identification unit 520, or may be provided as a separate unit, for example, between the marker identification unit 520 and the diagnosis unit 550. Features of the CNN (not shown in FIG. 16A ) are described in detail below.
[0254] Finally, the noise-filtered, CNN-processed subject-specific features, including the markers, can be provided as a file to a diagnostic unit 550, configured and arranged to diagnose a disease (e.g., a tumor disease) based on a statistical score indicating a match between the subject-specific features and the cancer features. The diagnostic unit can include a repository containing cancer patterns, such as the Catalog of Somatic Mutations in Cancer (COSMIC) database or the Latin American Consortium for Lung Cancer Research (CLICaP) database. The diagnostic unit 550 can include one or more software or algorithms for comparing known cancer mutation patterns (e.g., any of COSMIC patterns 1-30) with the subject-specific mutation patterns. Representative examples of such comparison software include, for example, measurement of reliability estimates at the level of individual markers, as well as pools containing 2, 5, 10, 20, 50, 100, 200, 500, 1000 or more, e.g., 5000, unique markers. Exemplary methods include estimating Z-score confidence levels using linear optimization (see above), or checking the similarity of normalized probability density functions (PDFs) using the cosine similarity function (see above).
[0255] The output of the diagnostic engine may be output, for example, to a display 412 for user review. In some embodiments, the output may include a raw confidence interval (CI) score or an ordinal score (e.g., a score on a 1 to 10 scale, with 10 being a high probability and 1 being a low probability that the subject has a neoplastic disease).
[0256] In terms of orientation, the marker identification unit 520 of the system 500 of FIG. 16A may be communicatively coupled to the data collection unit 510. Additionally, each component (e.g., engine, module, etc.) shown as part of the marker identification unit 520 (and described herein) may be implemented as hardware, firmware, software, or any combination thereof. In various embodiments, the marker identification unit 520 may be implemented as an integrated metrology system assembly with the data collection unit 510. That is, the units 520 and 510 may be housed within the same housing assembly and communicate via conventional device / component connection means (e.g., serial bus, optical cable, electrical cable, etc.). In various embodiments, the marker identification unit 520 may be implemented as a stand-alone computing device (shown in FIG. 16) communicatively coupled to the data collection unit 510 via an optical, serial port, network, or modem connection, for example, via a LAN or WAN connection that enables image data acquired by the data collection unit 510 to be transmitted to the marker identification unit 520 for analysis. In various embodiments, the functionality of the marker identification unit 520 may be realized on a distributed network of shared computer processing resources (e.g., a cloud computing network) that is communicatively connected to the data acquisition unit 510 via a WAN (or equivalent) connection. For example, the functionality of the marker identification unit 520 may be divided and implemented on one or more computing nodes on a cloud processing service such as Amazon Web Services™.
[0257] FIG. 16B illustrates a second exemplary system for implementing the diagnostic method of the present disclosure. As shown in FIG. 16B, exemplary system 100 is configured and arranged for genetic screening of a subject in need thereof. Referring to FIG. 16B, system 100 may include an analysis unit 110 and a computing unit 140. Analysis unit 110 may include a pre-filter engine 120 and a correction engine 130. The system components and associated engines are described in further detail below.
[0258] 16B, the pre-filter engine 120 of the analysis unit 110 may be configured and arranged to receive a subject-specific genome-wide read list associated with a plurality of genetic markers from a biological sample of the subject. As described in the workflow herein, according to various embodiments, the biological sample may include a tumor sample, and the read list may each include reads of a single base pair in length.
[0259] The pre-filter engine 120 may also be constructed and arranged to filter artifactual sites from the read list. As described in the workflows herein, according to various embodiments, filtering may include removing repetitive sites generated across a cohort of reference healthy samples from the read list, and / or identifying germline mutations in the biological sample, and / or identifying shared mutations between tumor samples and peripheral blood mononuclear cells of normal cell samples as germline mutations and removing said germline mutations from the read list.
[0260] Correction engine 130 of analysis unit 110 may be configured and arranged to receive output from engine 120. Correction engine 130 may be configured and arranged to filter noise from the genome-wide list of reads using at least one error suppression protocol to generate a filtered read set for the genome-wide list of reads.
[0261] As described in the workflow herein, according to various embodiments, the at least one error suppression protocol can include calculating the probability that any single nucleotide variation in the list is an artifact and removing the variation.
[0262] As described in the workflows herein, in various embodiments, the probability may be calculated as a function of features selected from the group consisting of mapping quality (MQ), variant base quality (MBQ), read position (PIR), mean read base quality (MRBQ), and combinations thereof.
[0263] As described in the workflow herein, and according to various embodiments, at least one error suppression protocol may include removing artifacts by testing for mismatches between independent copies of the same DNA fragment generated from polymerase chain reaction or sequencing processing, and / or by using overlap consensus, whereby artifacts are identified and removed upon mismatches across the majority of a given overlap family.
[0264] The computing unit 140 of the system 100 may be configured and arranged to receive the output from the correction engine 130 and compile a subject-specific pattern using the filtered read set based on comparison with a specific mutation pattern associated with a predetermined mutagenesis process.
[0265] Computing unit 140 may also be configured and arranged to statistically quantify a confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern based on a cohort comparison of the cancer-associated mutation pattern exposure value with background mutation patterns via the subject-specific features. Computing unit 150 may further be configured and arranged to screen the subject for cancer if the confidence estimate that the subject's biological sample contains the cancer-associated mutation pattern exceeds a given threshold.
[0266] System 100 may also include a display 150, as shown in FIG. 16B. The display may be constructed and arranged to receive output from computing unit 140. The output may include data related to the cancer screening of the subject / user. Alternatively, system 100 may omit a display and instead transmit data output from computing unit 140 to any type of storage or display device or location external to system 100. Also, as described herein, the components of system 100 may be integrated into one single unit or may be separated into separate physical units rather than as shown in FIG. 16B. Furthermore, system 100 may be part of a distributed network of systems, each performing substantially similar tasks and transmitting data from each system to a hub.
[0267] 16C illustrates a third exemplary system for implementing the diagnostic method of the present disclosure. As shown in FIG. 16C, exemplary system 100 is configured and arranged to perform genetic screening for cancer in a subject in need thereof. System 100 may include an analysis unit 110 and a computing unit 150. Analysis unit 110 may include a binning engine 120, a pre-filter engine 130, and a normalization engine 140. The system components and associated engines are described in further detail below.
[0268] 16C, the binning engine 120 may be configured and arranged to receive a genome-wide list of subject-specific reads associated with a plurality of genetic markers from a biological sample of the subject. As described in the workflow herein, according to various embodiments, the first biological sample may include a tumor sample, and the first list of reads may include copy number variations (CNVs).
[0269] The binning engine 120 may be constructed and arranged to divide the read list into windows and compute a set of features for each window, which may include median depth coverage per window and typical fragment size per window.
[0270] The pre-filter engine 130 may be constructed and arranged to filter artifactual sites from the read list. Filtering may include removing repetitive sites from the read list that were generated on a cohort of reference healthy samples.
[0271] Normalization engine 140 of analysis unit 110 may be configured and arranged to receive output from engine 130. Normalization engine 140 may be configured and arranged to normalize the read list to generate a filtered read set for the genome-wide list of reads. Normalization methods are discussed in detail herein and may be used in any combination contemplated to normalize reads as discussed.
[0272] Computing unit 150 of system 100 may be configured and arranged to receive output from normalization engine 140, calculate a linear relationship between the set of features per window, convert the calculated relationship using a regression model to an estimated tumor fraction (eTF), and calculate the estimated tumor fraction using the filtered read set. Computing unit 150 may also be configured and arranged to calculate the estimated tumor fraction based on one or more integrative mathematical models as a function of the calculated set of features per window across the genome-wide list of reads specific to the subject. Computing unit 150 may further be configured and arranged to screen the subject for cancer if the estimated tumor fraction exceeds an empirical threshold. Regression models, integrative mathematical models, and empirical thresholds are discussed in detail herein.
[0273] System 100 may also include a display 160, as shown in FIG. 16C. The display may be constructed and arranged to receive output from computing unit 150. The output may include data related to the detection of residual disease in the subject / user. Alternatively, system 100 may omit a display and instead transmit data output from computing unit 150 to any form of storage or display device or location external to system 100. Also, as described herein, the components of system 100 may be integrated into one single unit or separated into separate physical units rather than as shown in FIG. 16C. Furthermore, system 100 may be part of a distributed network of systems, each performing substantially similar tasks and transmitting data from each system to a hub.
[0274] [Convolutional Neural Network (CNN)] The present disclosure further relates to systems and programs that utilize convolutional neural networks (CNNs), e.g., engines, to adaptively and / or systematically filter ordered noise.
[0275] The present disclosure further relates to a computer-readable storage medium comprising a program for detecting tumor markers comprising somatic mutations in a genomic read, the program comprising a layered convolutional neural network.
[0276] Convolutional neural networks, as known in the art, generally achieve advanced forms of processing and classification / detection by first looking for low-level features, such as repetitive sequences in the reads, and then progressing to more abstract concepts through a series of convolutional layers. CNNs may do this by passing the data through a series of convolutional, nonlinear, pooling (or downsampling, as described below), and fully connected layers to obtain an output. Again, the output may be a single class or class probabilities that best describe the data, or detect objects in the data.
[0277] In a CNN, the first layer is typically a convolutional layer (conv). This first layer processes a representative array of reads using a set of parameters. Rather than processing the entire data, a CNN uses filters (or neurons or kernels) to analyze a set of data subsets. The subsets include the focal point and surrounding points in the array. For example, a filter might examine a series of 5x5 regions (or areas) in a 32x32 representation. These regions are called the receptive field. The filter is generally the same depth as the input; a filter for a representation with dimensions of 32x32x3 would have the same depth (e.g., 5x5x3). Using the above example dimensions, the actual convolution process involves sliding the filter along the input data, multiplying the filter value by the original representation value of the data, calculating the element-wise multiplication, and adding the values together to arrive at a single number for the examined region of the representation.
[0278] Using a 5x5x3 filter, an activation map (or filter map) with dimensions of 28x28x1 is obtained after this convolution process is completed. For each additional layer used, the spatial dimensions are better preserved, as using two filters results in a 28x28x2 activation map. Each filter typically has a unique signature that together indicate the feature identifier required for the final data output. Using such filters in combination, the CNN can process the data input and detect the features present in each representation. Thus, if a filter functions as a curve detector, convolution of the filter along the data input generates an array of numbers in the activation map corresponding to a high probability of a curve (high sum element-wise multiplication), a low probability of a curve (low sum element-wise multiplication), or a zero value if the input volume at a particular point does not provide anything to activate the curve detector filter. Thus, the more filters (also called channels) in a Conv, the more depth (or data) is provided on the activation map, and therefore, more information about the input, leading to a more accurate output.
[0279] The balance between accuracy of CNN is the processing time and power required to produce the results. In other words, the more filters (or channels) there are, the more time and processing power is required to perform the convolution. Therefore, the selection and number of filters (or channels) that satisfy the requirements of a CNN method should be specifically chosen to produce the most accurate output possible while taking into account the available time and power.
[0280] Furthermore, to enable the CNN to detect more complex features, additional convs can be added to analyze the output (e.g., activation maps) from previous convs. For example, if the first conv looks for basic features such as curves and edges, the second conv can search for more complex features, which can be combinations of individual features detected in previous conv layers. By providing a series of convs, the CNN can detect increasingly higher-level features, eventually reaching a probability of detecting a particular desired object. Furthermore, because conv stacks overlap each other and each conv level in the stack shrinks due to analysis of previous activation map outputs, each conv naturally analyzes a wider receptive field, thereby allowing the CNN to accommodate an expanded representation space when detecting the desired object.
[0281] A CNN structure generally consists of a set of processing blocks, including at least one processing block for convolution of an input volume (data) and at least one processing block for deconvolution (or deconvolution). Furthermore, the processing blocks may include at least one pooling block and a non-pooling block. Pooling blocks can be used to reduce the resolution of data to generate an output usable by Conv. This provides computational efficiency (time and power efficient) and may improve the actual performance of the CNN. The pooling, or subsampling, blocks make filters smaller and keep computational requirements reasonable. They may coarsen the output (which may lose spatial information within the acceptable field) and reduce the size of the input by only a certain factor.
[0282] A deconvolution block may be used to reconstruct this coarse output, generating an output volume of the same dimensions as the input volume. The deconvolution block may be viewed as the inverse of the convolution block, which returns the activation output to the original input volume dimensions. However, the deconvolution process generally simply spreads the coarse output into a sparse activation map. To avoid this result, the deconvolution block densifies this sparse activation map, ultimately generating an enlarged and dense activation map that, after any necessary processing, produces a final output volume whose size and density more closely resembles the input volume. Rather than reducing multiple array points in the receiving region to a single number as the inverse of the convolution block, the deconvolution block associates a single activation output point with multiple outputs, enlarging and densifying the resulting activation output.
[0283] It should be noted that although pooling blocks can be used to reduce data and non-pooling blocks can be used to expand the reduced activation map, convolution and deconvolution blocks can structure both convolution / deconvolution and reduction / expansion without separate pooling and non-pooling blocks.
[0284] Pooling and non-pooling processes can suffer from object-dependent drawbacks found in the data input. Pooling generally reduces the data by looking at sub-windows of data without overlapping windows, resulting in a loss of spatial information as the data is reduced.
[0285] A processing block may include other layers packaged with the convolutional or deconvolutional layers. These may include, for example, rectified linear unit layers or exponential linear unit layers, which are activation functions that examine the output from the Conv in that processing block. The ReLU or ELU layer acts as a gating function that advances only values that correspond to positive detection of the feature of interest specific to the Conv.
[0286] After the CNN is given a basic structure, it is prepared for a training process to improve its accuracy in data classification / detection (of the subject of interest). This involves a process called backpropagation. In this process, a training dataset, or sample data for training the CNN, is used to update parameters to reach an optimal, or threshold, accuracy. Backpropagation involves a series of iterative steps (training iterations), which train the CNN slowly or quickly depending on the backpropagation parameters. The backpropagation process generally includes a forward pass, a loss function, a backward pass, and parameter (weight) updates, depending on the learning rate. The forward pass involves passing training data through the CNN. The loss function is a measure of the error in the output. The backward pass determines the contribution factor of the loss function. The weight updates involve updating the filter parameters, which moves the CNN toward the optimal. The learning rate determines the extent of the weight updates at each iteration to reach the optimal. If the learning rate is too low, training may take too long and require high processing power. If the learning rate is too fast, each weight update may be too large and may not accurately achieve a given optimum value or threshold.
[0287] The backpropagation process can complicate training, resulting in a slower learning rate and requiring more specific and carefully determined initial parameters at the beginning of training. One such complication is the deep amplification of the network due to changes in the parameters of convs, with weight updates at the end of each iteration. For example, as mentioned above, if a CNN has multiple convs capable of higher-level feature analysis, parameter updates to the first conv are multiplied by each subsequent conv. The net effect is that, depending on the depth of a given CNN, even small changes to the parameters have a large impact. This phenomenon is called internal covariate shift.
[0288] In general, the CNN disclosed herein can adaptively and / or systematically filter ordering noise. In one embodiment, the CNN structure was designed based on the inventor's recognition that trinucleotide contexts contain distinct features involved in mutagenesis. Therefore, the CNN uses a receptive field of size 3 to cover all features (columns) at a given position. After two successive convolutional layers, downsampling is applied using max pooling with a receptive field of 2 and a step count of 2, forcing the engine model to retain only the most important features in a narrow spatial region. The resulting structure maintains spatial invariance when convolved across a 3-nucleotide window and captures "mapping quality" by collapsing read fragments into 25 segments corresponding to regions of approximately 8 nucleotides. Final classification is performed by directly applying the output of the last convolutional layer to a sigmoidally fully connected layer. The CNN employs a simple logistic regression layer, rather than a multilayer perceptron or global average pooling, to preserve position-related features in genomic reads.
[0289] To train the engine, first, various lung cancer patients and their corresponding systemic error profiles are sampled. The goal of training is to use a training scheme that enables highly sensitive detection of true somatic mutations and rejects candidate mutations caused by systemic errors. For this training, four separate samples, each containing a complete tumor sample and a healthy tissue sample from the same patient, are selected from various tobacco-smoking lung cancer patients (see, e.g., Table 3). For example, the consensus of three art-known calls (STRELKA, LOFREQ, and MUTECT) can be adopted for the final call of the somatic mutation. Next, the reads supporting the mutation are used as tumor reads to train the engine.
[0290] To ensure that the model engine learns to distinguish between sequencing artifacts, reads are taken from healthy samples containing mutations that occur exactly once. Because the variations are not supported by multiple reads, they can be considered with a high degree of certainty to be the product of systematic errors. Low-quality variants are then filtered. For example, variants with a base quality score of less than 20 or a mapping quality of less than 40 (e.g., BQ20, MQ40) may be filtered. These thresholds are purely exemplary and may be identified by inspection of the reads. If necessary, low-quality samples may be included in the training engine. A subset of the training set may be used as a validation dataset, which can be used to monitor training progress and verify model performance on independent reads.
[0291] In various embodiments herein, a computer-readable medium is provided, the computer-readable medium comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for identifying low-abundance tumor-specific markers in a list of genetic markers received from a sample from a subject, wherein the genetic markers in a genomic read include SNVs (preferably SNVs), CNVs (preferably SCNVs), indels, and / or SVs (preferably translocations, gene fusions, or combinations thereof). Preferably, the medium comprises a layered convolutional neural network (CNN) with a single fully connected layer at one end, the CNN being spatially invariant when convolving over a 3-nucleotide window and maintaining mapping quality by collapsing the read fragments into multiple segments, each representing a region of approximately 8 nucleotides, and the CNN measures each genetic marker in the list. For example, a CNN of the present disclosure can include eight layers, including a single fully connected layer at one end and successive convolutional layers whose output is downsampled by max pooling with receptive fields of 2 and 2 increments; the eight-layer CNN maintains spatial invariance when using a receptive field of size 3 to collapse read fragments into approximately 25 individual segments and convolve onto columns of position in the genome read; and the output of the final convolutional layer is applied directly to a sigmoidal fully connected layer where the final classification of the markers takes place.
[0292] The CNN can include a read representation that simultaneously captures the genomic context of the alignment, the complete read sequence, and an integrated per-base quality score. In part, due to this arrangement and structure, the CNN of the present disclosure provides about 1.12-fold to about 12-fold enrichment of tumor-specific markers, including somatic mutations, in the genomic read compared to the state-of-the-art variant caller MUTECT.
[0293] The present disclosure also relates to a computer-readable medium comprising computer-executable instructions that, when executed by a processor, cause the processor to perform a method or set of steps for diagnosing cancer in a subject in need of diagnosis, the medium comprising a convolutional neural network. In one embodiment, the CNN is developed using a training dataset comprising tumor-associated patterns and PCR / sequencing error patterns, and the CNN is trained to distinguish between cancer mutation-supportive reads and artifactual mutation (error) reads; and optionally, is validated using actual samples from cancer patients or synthetic plasma obtained from the dataset. In the list of genetic markers received from the subject's sample, the genetic markers include SNVs (preferably sSNVs), CNVs (preferably sCNVs), indels, and / or SVs (preferably translocations, gene fusions, or a combination thereof) in the genomic reads.
[0294] In one embodiment, the mathematical optimization process used in developing the CNN includes using non-negative least squares (NNLS). Other exemplary methods include cross-entropy global optimization, golden section search, or a combination thereof.
[0295] Preferably, the CNN of the present disclosure includes a single fully connected layer at one end, where the program maintains spatial invariance when convolving over a 3-nucleotide window; and maintains mapping quality by collapsing the read fragments into multiple segments, each representing an approximately 8-nucleotide region.
[0296] In one embodiment, the system of the present disclosure includes an 8-layer CNN that maintains mapping quality by collapsing read fragments into approximately 25 individual segments, which further collapses into every feature (column) at its position in the genomic read using a receptive field of size 3. In the context of analyzing genetic markers (e.g., sSNVs) in cfDNA, the CNN may include two successive convolutional layers, the output of which is sampled by max-pooling with a receptive field of 2 and a step count of 2, and the output of the final convolutional layer is applied directly to a sigmoid fully connected layer, where the final classification of the marker occurs.
[0297] The CNN constructed in this way accounts for true somatic mutations and spatial invariance in errors due to mapping, while preserving base quality across the read, providing a read representation that simultaneously captures the genomic context of the alignment, the full read sequence, and the integral of the quality score per base.
[0298] The embodiments disclosed herein offer several advantages over known CNNs, including, for example, providing CNNs that significantly improve accuracy and sensitivity. In particular, the disclosed systems and networks provide enrichment (measured as the ratio of output precision to input precision) of tumor-specific markers, including identified somatic mutations, of about 1.12-fold to about 12-fold, e.g., about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, about 10-fold, or more, compared to programs known in the art, such as MUTECT.
[0299] In one embodiment, the CNN involves identifying features that distinguish true tumor mutations from artifacts using a deep learning algorithm across a generic cancer cohort. The algorithm performs this function by assigning a confidence estimate to each individual mutation detected in a sample from a tumor patient, integrating the confidence estimates across the entire genome, and using an algorithm to analyze the characteristics of the mutations in the sample. For example, in the context of diagnosing lung cancer, the algorithm may analyze lung tumor patterns in the sample. Similarly, in the context of diagnosing UV-induced melanoma, the algorithm may analyze UV patterns in the sample. Similarly, in the context of diagnosing breast cancer, the algorithm may analyze breast tumor (BRCA) patterns in patient samples.
[0300] In some embodiments, the CNNs of the present disclosure include algorithms that can perform NNLS analysis using art-recognized / registered mutation patterns, such as mutation patterns deposited across samples in the Catalog of Somatic Mutations in Cancer (COSMIC) database. The present disclosure further relates to the CNNs of the present disclosure integrated with specific genome atlases, such as the TCGA Pan-Cancer dataset.
[0301] According to various embodiments, the CNN of the present disclosure can be trained with a deep learning algorithm developed on a pan-lung cancer cohort. In this case, the cohort can include WGS data on deep-stage tumor patients and PBMCs (controls). Using supervised learning, the CNN can be trained to identify features that distinguish true tumor mutations from human error. The resulting model can be used to infer and assign confidence estimates to individual mutations detected in plasma samples from cancer patients (e.g., patients with early-stage lung adenocarcinoma). A tumor detection signal can then be derived using a novel analytical method for sensitive detection by integrating the confidence estimates across the entire genome, followed by non-negative least squares (NNLS) of specific COSMIC mutation patterns in a single plasma sample. The detection signal can be further validated for reliability by comparing the COSMIC mutation exposure value with exposure values estimated for 100 random background patterns.
[0302] In some embodiments, the machine learning (ML) methods used in the systems and / or methods of the present disclosure include deep convolutional neural networks (CNNs), recurrent neural networks (RNNs), random forests (RFs), support vector machines (SVMs), discriminant analysis, nearest neighbor analysis (KNNs), ensemble classifiers, or combinations thereof.
[0303] The systems and / or methods of the present disclosure may allow for early detection in, for example, at least 50%, at least 60%, at least 70%, at least 80%, or even a greater percentage, such as 90% or 95%, of subjects.
[0304] [Other applications] Patient reports compiled using the above method can be transmitted and accessed electronically via the Internet. For example, analysis of sequence data can be performed at a location other than the subject's location. Reports can be generated, optionally annotated, and transmitted to the subject's location, for example, via an Internet-enabled computer. The annotated information can be used by healthcare providers to select other drug treatment options or to provide information about drug treatment options to insurance companies. This method includes annotating drug treatment options for a disease, such as the NCCN Clinical Practice Guidelines in Cancer (TM) or the American Society of Clinical Oncology (ASCO) practice guidelines. The drug treatment options stratified within the report can be annotated within the report, listing additional drug treatment options. Additional drug treatments are FDA-approved drugs for off-label use. A provision of the Omnibus Budget Reconciliation Act of 1993 (OBRA) requires health care coverage to cover off-label uses of anticancer drugs included in standard medical prescriptions. Drugs used in the annotated list include those listed in the CMS-approved National Comprehensive Cancer Network (NCCN), Drug and Biologics Directory (trademark), Thomson Micromedics Drug Dex (registered trademark), Elsevier Gold Standard Clinical Drug Directory, and the American Hospital Prescription Services Directory (registered trademark).
[0305] In some embodiments, drug treatment options may be annotated by listing experimental drugs that may be useful in treating cancer at one or more molecular markers of a particular condition. The investigational drugs may be drugs for which in vitro data, in vivo data, animal model data, preclinical study data, or clinical trial data are available. The data may be disclosed in peer-reviewed medical literature published in journals listed in the CMS Medicare Beneficial Policy Manual, such as American Journal of Medicine, Annals of Internal Medicine, Annals of Oncology, Annals of Surgical Oncology, Biology of Blood and Marrow Transplantation, Blood, Bone Marrow Transplantation, British Journal of Cancer, British Journal of Hematology, British Medical Journal, Cancer, Clinical Cancer Research, Drugs, European Journal of Cancer, Gynecologic Oncology, International Journal of Radiation, Oncology, Biology, and Physics, The Journal of the American Medical Association, Journal of Clinical Oncology, Journal of the National Cancer Institute, Journal of the National Comprehensive Cancer Network (NCCN), Journal of Urology, Lancet, Lancet Oncology, Leukemia, The New England Journal of Medicine, or Radiation Oncology.
[0306] Medications options can be annotated by providing links in the electronic report that associate the listed drug with scientific information about the drug. For example, a link to information about clinical trials of a drug can be provided. If the report is provided via a computer or computer website, the link can be a footnote, a hyperlink to a website, a pop-up box, or a flyover box with information. The report and annotated information can be provided in print format, and the annotation can be, for example, a footnote to a reference. The information annotating one or more medication options in the report can be provided by a commercial organization that preserves scientific information. A healthcare provider can treat a subject, such as a cancer patient, with an investigational drug listed in the annotated information, and the healthcare provider can access the annotated medication options, retrieve the scientific information (e.g., print a medical journal article), submit it to an insurance company (e.g., print a medical journal article), and file a claim for reimbursement for the medication. Physicians can use any of a variety of diagnosis-related group (DRG) codes to enable reimbursement.
[0307] Drug treatment options can also be annotated with information about other molecular components of the pathway in which the drug acts (e.g., information about drugs that target kinases downstream of the cell surface receptor that is the drug target). Drug treatment options can be annotated with information about drugs that target one or more molecular pathway components. The identification and / or annotation of pathway-related information can be outsourced or subcontracted to other companies.
[0308] The annotated information can be, for example, drug names (e.g., FDA-approved off-label drugs, CMS-approved listed drugs, and / or drugs described in scientific journal articles), scientific information about one or more drug treatment options, one or more links to scientific information about one or more drugs, clinical trial information about one or more drugs, one or more links to citations for scientific information about drugs, etc. Annotated information can be inserted anywhere within a report. Annotated information can be inserted in multiple places on a report. Annotated information can be inserted in a section near the stratified drug treatment options. Annotated information can be inserted in a report on a separate page from the stratified drug treatment options. Information can also be annotated in reports that do not include stratified drug treatment options.
[0309] The system may also include reporting on the effect of drugs on samples (e.g., tumor cells) isolated from a subject (e.g., a cancer patient). In vitro cultures using tumors from cancer patients may be established using techniques known to those skilled in the art. The system may also include high-throughput screening of FDA-approved, off-label, or experimental drugs using the in vitro cultures and / or xenograft models described above. The system may also include monitoring tumor antigens for recurrence detection. In a preferred embodiment, the annotated information may include treatment recommendations, including annotation of the effect of PARP inhibitors on BRCA patterns, and immunotherapy for MSI patterns. The above embodiments of the present disclosure are further described in view of the following non-limiting examples.
[0310] [Example] It will be understood that the structures, materials, compositions, and methods described herein are intended to be representative examples of the present disclosure, and that the scope of the present disclosure is not limited by the scope of the examples. Those skilled in the art will understand that the present disclosure can be practiced with variations on the disclosed structures, materials, compositions, and methods, and that such variations are considered to be within the scope of the present disclosure.
[0311] 〔background〕 [ To overcome the limitations of cfDNA abundance in highly sensitive cancer detection, extensive sequencing depth may be an alternative to sequencing. ] The above data demonstrate that the detection of a single sSNV in a patient's plasma sample is the result of two sequential statistical sampling processes. The first process provides the probability that a variant fragment will be sampled in the limited number of genome equivalents typically present in a blood sample. The second process evaluates the probability of detecting a variant fragment in a sample given its abundance, sequencing depth, and sequencing error (signal-to-noise). While the latter process has been the focus of intensive research and technological development by the scientific community (e.g., ultra-deep error-free sequencing protocols), the former probabilistic process has received less attention. However, as noted above, both processes play an important role in low-burden disease ctDNA detection. Even ideal ultra-deep targeted sequencing fails to discover a cancer signal if the physical fragment representing the target sSNV is absent. This is thought to be one of the reasons for the low sensitivity of this approach (approximately 40%, Rosenfeld et al.). In practice, this issue is further complicated by the fact that a single observation (mutation read) is rarely sufficient for confident detection.
[0312] We modeled cfDNA sampling as a Bernoulli test, mixing cfDNA fragments from two populations, normal cells and malignant cells, in a ratio defined by the tumor fraction (TF), to formulate the probability of sampling a variant fragment in a given cfDNA sample. Thus, the genome equivalents present in a plasma sample constitute a random sampling of the entire pool of cfDNA fragments in the patient's circulation. Therefore, the probability of sampling at least one variant fragment in a plasma sample that supports a particular substitution is given by the following: P = 1 - (1 - TF) GEwhere P is probability, TF is the tumor fraction, and GE corresponds to the number of genome equivalents present in the patient's cfDNA. In this model, the detection probability in TF associated with early cancer regimens (TF < 1%) rapidly declines for low TFs, and even at a frequency of 0.1% (1 / 1000), the detection probability is predicted to be less than 0.65 (Figure 3A). This limit was observed even under ideal conditions for full sequencing, efficiently utilizing 1000 genome equivalents (approximately 6 ng of cfDNA), and was noted to be based on detection based on a single supporting DNA fragment with ideal signal-to-noise. These results demonstrate that plasma sampling probability imposes a severe upper limit on mutation detection in low TF regimens, such as MRD and early cancer stage detection.
[0313] On the other hand, this model also shows that this restriction on sequencing depth can be effectively overcome by increasing the number of detected sites (SNVs) with increasing width, resulting from repeating the Bernoulli test for each SNV (Bernoulli distribution over the Bernoulli trial probability). This model can be represented by a binomial distribution of Bin(N,P), where N represents the number of test sites (mutations) and P = 1-(1-TF). GE is the probability of detection of a single site. Importantly, the mathematical model predicts the average number of detected sites as well as the probability of at least one detection as a function of the number of unique DNA fragments (genome equivalents or coverage), mutational burden (N, which can also be used as panel size), and TF (Figure 3B). Using the model, we integrated 20,000 sSNVs (approximately 10 mutations / mb found in 17% of human cancers) and found that even with a TF of 1:100,000, a high detection probability (up to 0.98) was obtained, easily achievable with standard whole genome sequencing (WGS) (Figure 3C).
[0314] Related Applications Non-invasive prenatal testing (NIPT) for chromosomal abnormalities The present disclosure further relates to noninvasive prenatal testing for chromosomal abnormalities using the above-described systems, methods, and algorithms. Preferably, NIPT can be performed using the CNV / SV-based workflow outlined in Figures 1C and 1E. Herein, non-de novo amplifications and deletions can be evoked and used for diagnosis in subject samples (e.g., amniotic fluid or blood from a pregnant woman with a suspected chromosomal abnormality). The method enhances sensitivity and specificity by utilizing the unique log2 / fragment size relationship (the same phenomenon that appears in fetal versus normal DNA), as shown in Figures 18E and 18F. Thus, the workflow of Figures 1C and 1E allows researchers or clinicians to combine two sources of information: CNVs generated from fetal DNA only, uncorrelated by noise corresponding to sequencing, alignment, and GC artifacts. Therefore, the disclosed methods and systems can enable clinicians to achieve higher sensitivity and specificity for NIPT using de novo CNV detection, even when prior information about the CNV segments is not readily available. [Example]
[0315] [Design of a somatic mutation classifier] When designing a model for somatic mutation classification, it is important to recognize sources of error that can lead to false-positive somatic mutations. True mutations are likely to exhibit high base quality regardless of read position. Similarly, the read base, reference base, and alignment string (CIGAR) at the position of a true mutation are likely to be independent of read alignment. More specifically, true somatic mutations are expected to be spatially invariant. Systemic errors in sequencing experiments are well known to depend on read position. Thus, while the mutation itself may be spatially invariant, the read position is usually not. Errors resulting from mismapping may involve repetitive sequences or highly specific sequence motifs (e.g., TTAGGG at telomeres). Therefore, it is desirable for the model to accurately represent both the spatial invariance of true somatic mutations and mapping errors while maintaining a model of base quality across reads. Therefore, shallow convolutional networks, whose classification relies on fully connected layers above the target read, may not be able to capture the invariance of mutations.
[0316] The inventors designed a somatic mutation classification engine recognizing these constraints and / or requirements. The engine utilizes a convolutional neural network (CN) to compensate for spatial dependencies. It uses an eight-layer CN with a single fully connected layer at the end, inspired by a VGG architecture (Simonyan & Zisserman, arXiv:1409.1556, revised April 10, 2015; Alexandrov et al., Nature, 500(7463):415-421, 2013). All features (columns) at a given location are convolved with a perceptual field of size 3. After two consecutive convolutional layers, downsampling is applied by max pooling with a receptive field of 2 steps and a step count of 2 steps, forcing the model to retain only the most important features in a narrow spatial region. Two main advantages are expected from this architecture. We found that: 1) spatial invariance is maintained when convolution is performed on 3-nucleotide windows, and 2) high mapping quality is achieved by collapsing the read fragments into 25 segments corresponding to approximately 8-nucleotide regions. The output of the final convolutional layer was applied directly to a sigmoidally fully connected layer used for final classification. To preserve features related to read position, we used a simple logistic regression layer instead of a multilayer perceptron or global average pooling.
[0317] The rapidly developed model and training scheme is called the Engine. The Engine uses an initial read representation that simultaneously captures the genomic context of the alignment, the complete read sequence, and the integration of base quality scores. The Engine's capabilities include high-depth coverage of feature integration across reads and their alignments, as well as the creation of a new set of somatic variant calls using the complete mutation profile of the sample.
[0318] The performance of the rapid model was evaluated by examining its predictive performance on an independently selected lung cancer dataset. The dataset was combined with healthy WGS data from the same patients. The model was evaluated using the following metrics: F1 score, accuracy, sensitivity, and specificity: Sensitivity =TP / (TP+FN) (Equation 1) Accuracy =TP / (TP+FP)(Equation 2) Specificity =TN / (TN+FN) (Equation 3) F1 score = 2 x (precision x recall) / (precision + recall) (Equation 4) [Table 1]
[0319] The model was found to maintain an average F1 score of 0.961 for the validation set. The model achieved an F1 score of 0.71 for tumor controls. While the model remains highly sensitive for tumor control, specificity decreased slightly compared to the validation dataset. However, for independent lung samples, specificity was high, with an F1 of 0.92 (Table 1). The low accuracy and specificity for cancer control indicated that while the engine learned specific mutation patterns associated with tobacco-related lung cancer, it also learned general error patterns.
[0320] We further investigated the engine's learning capabilities using an additional sample from a melanoma patient (CA0040; Table 1). Melanoma samples typically exhibit a significantly different mutation profile upon exposure to UV light compared to the mutation profile associated with tobacco exposure (Figure 8A). The engine model achieves an F1 score of 0.71 for melanoma samples. Thus, while the model remains sensitive, its accuracy and specificity for melanoma samples are lower, demonstrating that the engine learned specific mutation patterns associated with tobacco-exposed lung cancer while also learning more general sequencing artifact patterns applicable to both tumor types.
[0321] To further explore this issue, we investigated the differences in trinucleotide context frequencies between true cancer mutation variant reads and sequencing artifacts using the following datasets: (i) a lung cancer patient sample included in training (CA0046, validation dataset), (ii) a lung cancer patient not included in training (CA0044), and (iii) a melanoma patient (CA0040). The results are shown in Figure 8B.
[0322] As expected, tobacco-associated lung adenocarcinoma samples were noted to exhibit a high enrichment of C>A transversions, consistent with tobacco-associated mutation patterns (Figure 8B). Therefore, we hypothesized that the engine could learn specific sequence contexts that are prevalent in tumor mutation data (i.e., tumor-specific mutational signatures). To test this hypothesis, we measured the difference in frequency between true cancer mutations versus sequencing artifacts in each trinucleotide context and correlated this with the average model prediction for these same reads. We reasoned that if the model had learned a (lung) cancer-specific sequence context, we would expect a high correlation between the frequency of the trinucleotide sequence and the model output. Consistent with this hypothesis, we observed a high correlation between model predictions and trinucleotide enrichment for both CA0046 (Pearson's r = 1, included in training) and CA0044 (Pearson's r = 0.95, not included in training). The results are shown in Figure 8C.
[0323] To directly examine whether high correlation results from accurate classification independent of sequence context (an alternative scenario), we performed a similar analysis using melanoma samples (CA0040). Results showed a persistent positive correlation (Pearson's r = 0.64) between trinucleotide context and model predictions, indicating accurate classification derived from features other than mutation pattern alone. This was significantly lower than in the tobacco-exposed lung cancer data. This finding is consistent with model training on lung cancer-specific mutational features. This finding led to the training of a separate model specifically for detecting melanoma-associated somatic mutations. Using the NSCLC procedure described above, we examined an additional dataset of three melanoma patients. High F1 scores for the melanoma validation dataset and the independent melanoma samples revealed similar performance, although the F1 scores were lower when the model was applied to NSCLC data (controls).
[0324] Engine Sensitivity and Precision at Low Tumor Fractions in Synthetic Plasma To evaluate the performance of the present system and / or method in a low tumor fraction setting, the accuracy and sensitivity of the engine were compared to state-of-the-art callers, MUTECT, SNOOPER, and STRELIKA. Figure 9A shows the results, demonstrating the excellent sensitivity achieved by the engine, especially at low tumor fractions. In contrast, MUTECT failed to detect more than two mutations in the synthetic samples at any tumor fraction, and every successful mutation prediction had the same call in the tumor fraction. Thus, the engine increased sensitivity over MUTECT by more than 200-fold, while improving accuracy over a simple filter with a tumor fraction of 0.01. Based on these surprisingly good results, the disclosed system and method were applied in the context of real plasma samples.
[0325] A comparison was also made between the engine and the simplified calling method PILEUP. The results are shown in Figures 9B and 9C. A comparative evaluation was performed across the filters implemented using the engine. The comparative evaluation was also performed using another metric called enrichment, which provides information on the increase in the ratio of tumor to normal mutations when referring to the filter. The enrichment factor is calculated as follows: [Concentration] = [Precision out] / [Precision in]... (Formula 5) It can be calculated using Equation 5.
[0326] PILEUP is sensitive enough to detect somatic mutations in pseudo-traits, but it includes all mutations. This is not reflected in the enrichment and precision metrics. In the next stage of the pipeline, we applied a mutation frequency filter. While the MF and BQ+MQ filters actually deplete the sample of tumor reads, an increase in enrichment was observed when TF = 0.01. This is a good indicator that the filters are useful for the evaluation pipeline as well as for removing a large portion of the noise before presentation to the CNN. Applying the CNN filter, we observed an additional (third) order reduction in noise. Most importantly, this was accompanied by only a loss of sensitivity of about 25%. With the full pipeline, we observed a 30-fold enrichment (over PILEUP; green line) at both tumor fraction 0.01 and tumor fraction 0.0001. The data are shown in Figure 9C.
[0327] [Analysis of somatic mutations in actual cfDNA samples using the engine] To verify that the disclosed method and system are robust in actual clinical settings, actual evaluations were performed on two different types of samples: one cfDNA sample from a healthy individual (identifiers: BB600; BB601), and two cfDNA samples from early-stage lung cancer patients (identifiers: BB1122; BB1125) collected before surgery. In the actual clinic, the clinician performing the test did not have access to mutation information for the patient. However, because BB1125 underwent surgery, the clinician was able to measure true somatic mutations using a standard mutation calling pipeline. These calls, combined with the reads obtained from cfDNA, can provide a second qualitative estimate of the engine's sensitivity, precision, and enrichment.
[0328] After applying the filtering pipeline, we found that 27 of the 413 mutations present in the samples were successfully captured. Most notably, false positives were reduced from 266 to 3 in the control group (see Table 2). The results showed that the overall pipeline actually reduced tumor signal by approximately 50%, while in contrast, the engine enriched the samples by approximately 1.7-fold. [Table 2]
[0329] The results indicate that differences in the pretreatment process may have led to poor settings of the BQMQ filter. In this sample, a base quality score of 20 was estimated to be too lenient.
[0330] Recognizing the advantage of using a training scheme that can detect true somatic mutations with high sensitivity while rejecting candidate mutations caused by systemic errors, we sampled various lung cancer patients and matched their systemic error profiles. Four representative samples from various smoking lung cancer patients were selected for training to implement this scheme (Table 3). [Table 3]
[0331] An additional patient with lung cancer who smoked was tested. Samples were processed and provided by the Cancer Alliance at the New York Genome Center. The samples included a complete tumor specimen and a healthy tissue specimen from the same patient. A three-call consensus of STRELKA, LOFREQ, and MUTECT was selected to perform the final somatic mutation call. The reads supporting the mutation were then used as training tumor reads.
[0332] Because we wanted a learning model that could discriminate against sequencing artifacts, we collected accurate reads from a single healthy sample in which a mutation occurred. Because the variant was not supported by multiple reads, it was likely due to a systemic error. We then filtered out low-quality variants, filtering out variants with a base quality score below 20 or a mapping quality below 40 (BQ 20, MQ 40). The threshold BQMQ values were determined by inspection, but a window was created to allow for the inclusion of lower-quality samples in training. A small subset of the training set was additionally prepared for use as a validation dataset. This dataset was used to monitor training progress and validate model performance on independent reads (not independent mutations). Next, model performance was evaluated on a test lung dataset.
[0333] [Synthetic plasma] To test the model's ability to detect somatic mutations at low frequencies, four simulated plasma samples from a test lung sample (CA0044, Table 3) were generated by randomly sampling the patient's healthy DNA and the patient's tumor DNA. Sampling was performed with 35% coverage and tumor mixtures of 0%, 0.01%, 0.001%, and 0.0001%. Mixing was performed with three random seeds for stability. A threshold rate of approximately 0.1 was selected as the somatic mutation rate in cfDNA. Therefore, when preparing a synthetic plasma readout, 1 / 10 of the readouts covered in the mixture were used. th Only mutations supported by <1 were selected.
[0334] To evaluate the performance of the disclosed method and / or system in a low tumor fraction setting, parameters such as accuracy, sensitivity, and enrichment were compared between the engine and a state-of-the-art low-frequency caller, MUTECT. Additionally, a simplified calling method called PILEUP, which tolerates observed mismatches, was included in the comparison. After PILEUP, the same filters used in the engine were repeatedly applied to measure the performance of each step. The filters implemented in the method are MF (mutation frequency), which filters PILEUP reads where PILEUP occurs more frequently than expected in plasma (mutations occur 10% of the time), and BQMQ, which filters reads where the mutation base quality is less than 20 or the mapping quality is less than 40. Finally, an instant filtering method using the engine is used.
[0335] [Evaluation of cfDNA samples] After evaluating the engine on synthetic samples, we tested its performance on real plasma DNA samples. We analyzed control samples (BB600; BB601) and samples from smoking lung cancer patients (BB1125 or BB1122). Because these patients also had tumor biopsies, we measured true positives by assuming all mutations from the biopsy, called MUTECTs, that were also present in the cfDNA. Using these calls, we performed the same analysis as with synthetic plasma (see above).
[0336] [Evaluation of sensitivity, precision, and concentration] For controls, all measurements were performed for the mutations in subject BB1125.
[0337] [Characteristics] To fully capture the sequencing reads, alignments, and genomic context, we created a spatially oriented representation of the reads (Figure 5). For reference insertions, we maintained spatial alignment by marking the deletion in the reference as an "N." For deletions in the reference, we placed the location of the deletion in the sequencing read as an "N." The soft-masked region is segmented so that the reads are adjacent to the read-mapping portion and the reference context is broken with consecutive "Ns" until the end of the soft-masked region. This is done for two reasons: first, to ensure that the signal in the soft-masked region is strong, and second, to maintain the concept that the read is independent of its alignment.
[0338] Segments (e.g., + / - 25 bases) were inserted on either side of the read from the genomic context (Figure 6). This resulted in a 16 x 200 base matrix for 150-base reads, with extra context bases added if the read was not 150 bases. The maximum base quality score was set to 40 (p = 99:99%), with scores in the interval [0,1]. Bases not covered by the read (genome-related) received a base quality score of zero. Deletions in a read received a quality score that was the average of the two flanking positions in the read.
[0339] [Hyperparameters and implementation details] The model was trained using mini-batch stochastic gradient decent with an initial learning rate of 1 and momentum of 9. Once the validation loss reached a plateau as outlined in He et al. (In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016), the learning rate was reduced by a factor of 10. A mini-batch size of 256 was used, as this appeared to offer the best tradeoff between validation loss and training speed. 64 filter bases were used, doubling after each downsampling layer to maintain a consistent number of parameters at each stage. This was chosen empirically after observing that a 32-base filter model did not perform well. After each convolutional layer, batch normalization was applied, followed by a rectified linear unit. Before each pooling layer, dropout with a descent probability of 0.5 was applied.
[0340] The engine demonstrated robust performance in a balanced tumor fraction setting. Furthermore, much of its performance was maintained in simulated plasma. The engine also achieved a 2-fold enrichment of tumor fractions of 0.0001, suggesting its ability to capture relevant somatic mutations even at frequencies 10-fold lower than the sequencing noise itself. In contrast, MUTECT, a tool not intentionally designed to function in a cfDNA setting, only predicts less than 2 across all tumor fractions. See Figures 9A-9C.
[0341] Detailed engine results are shown in Table 4. [Table 4]
[0342] Other Embodiments Therefore, the system and method can be developed into a complete early detection engine. The engine uses a fully connected sigmoidal layer to capture the position of the reads, but there are other structures that are more suitable for capturing the relative position of the reads. Furthermore, additional information contained in the read pairs from DNA fragments excluded in preliminary testing can be used to determine the origin strand (Watson or Crick) and estimate the size of the DNA fragments. The fragment size distribution of ctDNA has been observed to differ from that of normal circulating healthy DNA (Underhill et al., PLoS Genetics, 12(7):e1006162, 2016).
[0343] The above systems and methods can be integrated with recurrent neural networks (RNNs). RNNs have been shown to be powerful tools that can be used to characterize length in bioinformatics, even at distances up to 1 kb (Hill et al., bioRxiv, pp. 2007-58, 2017). Integrating a recurrent neural network instead of a logistic regression layer may further improve the performance of the disclosed methods and systems. [Example]
[0344] Methods and systems for detecting and validating tumor-specific low-abundance tumor markers, and their use in cancer diagnosis The disclosed systems and methods are useful for early cancer diagnosis. In contrast to metastatic cancer (where disease burden is high and ctDNA is significantly elevated), as known in the art, the abundance of ctDNA limits the use of targeted sequencing techniques in the detection of early-stage cancer or residual disease. Given the known limited amount of cfDNA in settings with low tumor burden, we first investigated the possibility of optimizing cfDNA extraction. First, to reduce variability arising from sample acquisition and interindividual variability, commercially available extraction kits and methods were compared using homogenous cfDNA material generated through large-volume plasma collection (approximately 300 cc) via plasmapheresis from healthy subjects and cancer patients undergoing hematopoietic stem cell collection. The large volume of plasma allowed multiple method and protocol parameters to be tested on the same cfDNA input, allowing subtle differences in yield and quality to be accurately measured.
[0345] Kits and reagents from Capital Biosciences (Gaithersburg, MD, USA; Catalog #CFDNA-0050), Qiagen (Germantown, MD, USA), Zymo (Irvine, CA, USA; Catalog #D4076), OmegaBIO-TEK (Norcross, GA, USA; Catalog #M3298), and NEOGENESTAR (Somerset, NJ, USA; Catalog #NGS-cfDNA-WPR) were used in this comparative study. Extractions were performed on large 1 ml plasma samples using the kits and reagents according to the manufacturer's instructions. Multiple plasma aliquots were processed in parallel to assess inter- and intra-method variability. The yield and purity of each recovered cfDNA sample were measured using fluorimetry (total mass), UV absorbance (detection of salt and protein contaminants), and on-chip electrophoresis (size distribution and gDNA contamination).
[0346] Results demonstrated that OmegaBIO-TEK's MAG-BIND cfDNA extraction kit outperformed all other tested methods. Systematic optimization of each step of the manufacturer's protocol was further performed to reduce carryover of contaminants and improve cfDNA recovery.
[0347] The optimized extraction protocol was then applied to samples from early-stage lung cancer. This cohort included 11 preoperative early-stage lung cancer plasma samples and four plasma samples from benign patients (controls). Exemplary patient characteristics are shown in Figure 11. Despite optimal extraction, cfDNA yields from low-disease burden samples were low and highly variable between patients, ranging from 0.13 ng / mL to 1.6 ng / mL. The data confirm the low and variable number of DNA molecules available for cfDNA sequencing.
[0348] [Extensive sequencing depth can be an alternative to sequencing to overcome the limitations of cfDNA abundance in highly sensitive cancer detection] The above data demonstrate that the detection of a single sSNV in a patient's plasma sample is the result of two sequential statistical sampling processes. The first process provides the probability that a variant fragment will be sampled in the limited number of genome equivalents present in a typical blood sample. The second process evaluates the probability of detecting a variant fragment in a sample given its abundance, sequencing depth, and sequencing error (signal-to-noise). While the latter process has been the focus of intensive research and technological development by the scientific community (e.g., ultra-deep error-free sequencing protocols), the former probabilistic process has received less attention. However, as noted above, both processes play an important role in the detection of low-burden disease ctDNA. Even ideal ultra-deep targeted sequencing cannot detect a cancer signal if the physical fragment representing the target sSNV is not present. This is thought to be one of the reasons for the low sensitivity of this approach (approximately 40%, Rosenfeld et al.). In reality, this problem is further complicated by the fact that a single observation (mutation read) is rarely sufficient for confident detection.
[0349] To formulate the probability of sampling a variant fragment in a given cfDNA sample, cfDNA sampling was modeled as a Bernoulli test in which cfDNA fragments from two populations, those from normal cells and those from malignant cells, were mixed in a ratio defined by the tumor fraction (TF). Thus, the genome equivalents present in a plasma sample constitute a random sampling of the entire pool of cfDNA fragments in the patient's circulation. Therefore, the probability of sampling at least one variant fragment in a plasma sample that supports a particular substitution is given by: P = 1 - (1 - TF). GE It can be defined as: where P is probability, TF is the tumor fraction, and GE corresponds to the number of genome equivalents present in the patient's cfDNA. The rapid model predicts that the detection probability in TF-associated early cancer regimens (TF < 1%) rapidly decreases for low TFs, with even a frequency of 0.1% (1 / 1000) predicting a detection probability below 0.65 (Figure 3A). This limit is observed even under ideal conditions of exhaustive sequencing, efficiently utilizing 1000 genome equivalents (approximately 6 ng of cfDNA), and is based on detection based on a single supported DNA fragment with ideal signal-to-noise. These results indicate that plasma sampling probability imposes a severe upper limit on mutation detection in low-TF regimens, such as MRD and early cancer stage detection.
[0350] On the other hand, this model also shows that this restriction on sequencing depth can be effectively overcome by increasing the number of detected sites (SNVs) with increasing width, resulting from repeating the Bernoulli test for each SNV (Bernoulli distribution over the Bernoulli trial probability). This model can be represented by a binomial distribution of Bin(N,P), where N represents the number of test sites (mutations) and P = 1-(1-TF). GEis the detection probability of a single site. Importantly, the mathematical model predicts the average number of detected sites as well as the probability of at least one detection as a function of the number of unique DNA fragments (genome equivalents or coverage), mutational burden (N, which can also be used as panel size), and TFs (Figure 3B). Using this model, we found that by integrating 20,000 sSNVs (approximately 10 mutations / mb found in 17% of human cancers), a high detection probability (up to 0.98) was obtained even with a TF of 1:100,000, which is easily achievable with standard whole genome sequencing (WGS) (Figure 3C).
[0351] [In silico validation of genome-wide integrated sSNV detection] The rapid model indicates that increasing the number of sites results in a significant increase in the probability of detection. To test this prediction, we simulated cfDNA detection using in silico mixtures of tumor and normal WGS data from 11 cancer patients with a variety of cancers, including high-grade tumors from lung adenocarcinoma, ductal adenocarcinoma (breast), cutaneous melanoma, urothelial carcinoma (bladder), and osteosarcoma (full clinical details in Schematic A, Figure 1F).
[0352] All samples were deep sequenced with ~80x tumor WGS and ~40x PBMC WGS. To create in silico mixtures, tumor and normal WGS reads were mixed in various ratios to generate datasets of patient-specific virtual plasma samples with different TFs (0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01) and coverage (0.01). This dataset consisted of five independent replicates for each condition, obtained through various random seeds used during the downsampling process. To simulate detection in the setting of residual disease, somatic mutation calling was performed on the original tumor and germline WGS data to obtain a patient-specific list of sSNVs. The number of tumor-associated mutation sites in the in silico plasma simulation mixture was then measured through the detection of at least one support for the patient-specific sSNV list. We also found that integrating many sites accumulates noise due to sequencing errors, which can limit the detection of the signal. To estimate the degree of noise in WGS-based cfDNA detection, complementary data bases of downsampled normal reads without intermixing of reads from tumor WGS were generated (TF = 0, 20x, and 35x coverage, 20 repeats). The data allow for signal-to-noise measurements and demonstrate that integrated genome-wide SNV detection can reliably detect TFs > 1:2000 in high mutation burden tumors with 20x coverage for various tumor types.
[0353] The data also show how noise from sequencing errors shapes the relationship between the number of detected sites and TFs, as the relative contribution of noise increases with decreasing TFs. Comparison of estimated sequencing noise with integrated mathematical model predictions showed high agreement across different TFs and coverage values for all patients and cancer types. The analysis also shows how detection signal increases with increasing mutation burden (N) and coverage, with detection counts in 1% TFs ranging from 40K mutation burden (melanoma) to 8K mutation burden (non-tobacco lung).
[0354] By characterizing the variables underlying the estimated noise and developing optimized filters, we can significantly improve the signal-to-noise ratio and detection sensitivity. We model the noise distribution with other independent variables, such as mutation load, coverage, and cancer type. Results show cancer-type-independent error rates that reflect previously published sequencing error rates (~1 / 1000 bases). Furthermore, the detected signal exhibits patient-specific correlation with negligible germline-associated noise.
[0355] The data showed that the sequencing error was related to parameters such as base quality (BQ), mapping quality (MQ), fragment length, and variable allele frequency (VAF). Therefore, to reduce the sequencing error probability, a combined base quality (BQ) and mapping quality (MQ) optimization filter was developed through optimal receiver operating characteristic (ROC) analysis, which reduced the measured error probability by 3FC (approximately 3 × 10 -4 ) was reduced. Even though tumors with a 35X WGS depth had a 20,000-fold reduction in TFs, applying this filter at a reduced coverage of 35X enabled marker detection. These data support the use of integrated genome-wide sSNV profiling with patient matching, as cancer detection can be achieved even with very low TFs, regardless of the abundance of cfDNA (e.g., 100X WGS with 1 ng input). Furthermore, the high agreement between experimental results and mathematical models indicates that measurements of the number of detected sites (patient-specific sSNVs) can be translated into estimates of plasma TFs, enabling quantitative TF monitoring in early detection settings.
[0356] Additional parameters beyond quality metrics can be further utilized to filter remaining noise, including utilizing information about specific motifs, patterns, etc. Representative methods include, for example, performing fragment size filtering (e.g., only fragments of approximately 200 bp or less are considered) and variable allele frequency (VAF) filtering (e.g., only alleles with a VAF above a threshold of 2%, 5%, 10%, etc. are considered). The various mutation patterns of tobacco exposure and UV exposure are shown in the upper and lower panels of Figure 12A, respectively. Differentially expressed COSMIC patterns in lung tumor, breast tumor, and melanoma samples are shown in Figures 12B and 12C.
[0357] [Application] We then applied this highly sensitive de novo mutation detection to preoperatively sequenced plasma from five early-stage patients to generate genome-wide cfDNA mutation detection. We aggregated genome-wide mutation data and calculated a mutation list for each patient. We then used novel analytical methods for highly sensitive mutation pattern detection using novel machine learning algorithms and tools, such as convolutional neural networks (CNNs).
[0358] Based on the application of a two-pronged strategy, a deep learning algorithm was first trained on a generalized lung cancer cohort (five patients with deep-seated tumors and PBMC WGS) using supervised learning to identify features that distinguish true tumor mutations from artifacts. The resulting model was then used to infer and assign confidence estimates to individual mutations detected in early-stage lung adenocarcinoma plasma samples. Second, detection signals were derived through integration of genome-wide confidence estimates, followed by a novel analytical method for sensitive detection of specific COSMIC mutation patterns in single plasma samples using non-negative least squares (NNLS). Pattern detection was further validated for reliability using a comparison of COSMIC mutation exposure values with exposure values inferred for 100 random background patterns (z-score >2 STDs).
[0359] The results shown in Figure 13 demonstrate that the disclosed CNN is particularly useful for early tumor detection. The method detected tobacco-specific patterns in lung cancer patients, UV-specific patterns in melanoma patients, and BRCA-specific patterns in breast cancer patients, even at TF levels below 1 / 1000. To improve the low positive predictive value (PPV) of current lung cancer CT screening in subjects at risk for tobacco exposure, the method was evaluated by applying it to plasma samples from five early-stage lung patients and four benign nodules detected as positive by CT screening. The data showed positive detection in early-stage lung cancer specimens and fewer (false) positive detections in benign nodules, indicating improved PPV.
[0360] The patient-specific feature scores (z-scores) were then mapped to patient characteristics, such as smoker or non-smoker status and smoking history (e.g., number of pack-years smoked by each patient (smoker)), including histopathological features such as positive or negative (ND) for nodule detection. The results, shown in Figure 14A, clearly demonstrate that the tobacco pattern is detectable in plasma from early-stage cancer patients exposed to tobacco but not in benign nodules or patients with no smoking history. Our method successfully detected the tobacco pattern in three of four early-stage lung patients with a history of tobacco exposure, but no pattern was detected in the plasma samples of three non-smoker lung patients who underwent benign lung nodule resection. The specificity of the tobacco pattern for detecting lung cancer patients was at least 67% in all but one stage of disease, and the specificity approached 100% in patients with higher-stage disease (e.g., stage IIIa or later).
[0361] [Combined CT screening and diagnostic methods to improve PPV] To improve the low positive predictive value (PPV) of CT screening methods, we applied the above screening method to the diagnosis / prognosis of at-risk subjects exposed to tobacco, with or without CT screening. First, markers and patterns (including SNVs, CNVs, indels, and / or SVs) were de novo detected via whole-genome sequencing (WGS), and the markers were analyzed for noise / error using the above method. A total of 30 preoperative specimens from patients with early-stage NSCLC (stages I and II) were analyzed in this manner. Additionally, WGS was performed on patients matched for age 30 and tobacco exposure who had benign lesions identified by the institution's CT-based screening program. The detection signal derived from cfDNA data was integrated with CT-based readings in a blinded manner to determine whether cfDNA information could improve the positive predictive value of CT screening. This cohort demonstrated an effect size of approximately 40% PPV with current methods, increasing to approximately 60% with integrated cfDNA and CT screening, demonstrating detectability. Depending on the results of the study, larger prospective, local clinical trials may be conducted.
[0362] [Consideration] The data demonstrate that the disclosed methods and / or systems are superior to existing methods, particularly in the context of detecting low-abundance markers for early detection of tumors (ED). Early cancer detection, where matching tumor DNA is unavailable, presents the challenge of detecting de novo cancer mutations. The disclosed genome-wide integrated method utilizes sSNV sequence context information to detect mutation patterns associated with specific mutagenic processes, such as exposure to tobacco, UV light, APOBEC hyperactivity, BRCA mutations, PARP activity, or MSI. The patterns were specifically manifested in tumor somatic mutations and completely absent in PBMC somatic mutations in all samples tested.
[0363] Sensitive and specific de novo mutation detection in cfDNA from low-TF samples is fundamentally challenging with existing mutation detection algorithms. All known methods rely on comparing tumor and normal DNA at specific genomic sites. Detection performance for detecting mutation sites in the genome is derived from the observation of multiple supporting reads covering the site, a statistical framework that distinguishes these multiple observations from sources of sequencing noise (sequencing errors, mapping errors, etc.). However, in early detection, the amount of mutated ctDNA is significantly smaller than the sequencing depth (or the number of fragments available for sequencing at a specific site), and therefore, at most, only one supporting read is observed at each site. For example, application of MUTECT to virtual plasma data shows a rapid decrease in true tumor-associated somatic mutations with decreasing TF, even when considering all detections (before variant filtering) contained in the callstat file, but when considering detections with a single supporting read, the mutation site is called more frequently.
[0364] To enable error-free de novo single ctDNA detection at low TFs, a new framework is required that can distinguish between alternating reads resulting from cancer mutations and reads resulting from sequencing artifacts. While mutation patterns typically utilize trinucleotide context, recent data suggest that sequence context may extend far beyond this range, making it difficult to capture with supervised feature selection.
[0365] The present disclosure provides a novel method and pipeline for filtering sequencing errors. For example, tumors resulting from specific mutation processes produce distinct mutation patterns, which can be utilized to effectively filter artifacts, providing enriched markers with improved subject specificity, sensitivity, and accuracy. The disclosed neural network utilizes machine learning, thereby overcoming the limitations of known calling methods in the art. The machine learning architecture distinguishes between cancer-altered sequencing reads and reads altered by sequencing errors, specifically and adaptively filtering whole-body sequencing noise. In this context, the disclosed deep convolutional neural network provides an artificial intelligence platform that integrates multiple features in a supervised manner, which is unique to solving classification problems in the context of genomic sequence readings. The approach used to design the CNN is based on rethinking difficult mutations. Unlike known mutation calling methods in the art, such as MUTECT, the disclosed CNN can distinguish between reads containing true variants and reads containing sequencing artifacts. The disclosed CNNs are dynamic rather than static, as they can be trained with millions of true mutation reads and errors using a large collection of tumor and normal WGS data. The above characteristics of CNN provide advantages over mutation calling methods known in the art, as evidenced by the high sensitivity and specificity associated with detecting a wide variety of tumor types in many patients.
[0366] Application of the disclosed methods and systems to lung cancer detection These results demonstrate that genome-wide information integration can overcome the major barriers associated with detecting low-abundance markers that indicate disease states. Applying the disclosed methods and systems to analytical methods breaks the detection limit, enabling detection of tumor fractions as low as 1 / 10,000, with improvements depending on sequencing depth. This advantage is particularly useful in the fields of lung cancer detection and residual disease detection in patients after surgery and / or treatment.
[0367] In the context of pre-malignant lung lesions, detecting low-invasive lesions may be more difficult than detecting early-stage NSCLC.Notably, most cancer mutations are thought to occur before malignant transformation, so pre-malignant growths are likely to also exist.Therefore, the systems and methods described herein can also be used to detect pre-malignant lesions, particularly in the context of lung tumors.
[0368] Orthogonal integration of fragment size features in SNV-based methods cfDNA fragment distribution exhibits a unique profile of DNA degradation in the blood circulation. Figure 17A shows the fragment size distribution of a healthy cfDNA sample. Tumor-derived circulating DNA fragments are shorter compared to "normal" DNA fragments, primarily derived from apoptosis of hematopoietic (immune) cells. Breast tumor cfDNA (red and purple) exhibits a fragment size shift compared to a normal cfDNA sample (Figure 17B). Calculation of the center of mass (COM) of the first nucleosome (peak at approximately 170 bp) shows a shift to a lower COM, which corresponds linearly to TFs. Using a human tumor xenograft model (PDX) in mice, tumor-derived circulating DNA (red, aligned to human) was significantly shorter than normal-derived circulating DNA (black, aligned to mouse). See Figure 17C.
[0369] To create a stable model that can quantify the probability that a single DNA fragment is derived from tumor or normal origin, we used a joint Gaussian mixture model (GMM) to characterize the fragment size distribution of circulating DNA. The circulating tumor DNA model (red dashed line) was estimated by applying GMM analysis to circulating tumor DNA extracted from our PDX samples using only circulating DNA aligned to the human genome. The circulating normal DNA model (gray dashed line) was estimated by applying GMM analysis to circulating DNA from plasma samples of healthy human volunteers. The joint log-odds ratio (yellow line) was then used to estimate the probability that a particular circulating DNA fragment size is of tumor or normal origin. The data are shown in Figure 17D.
[0370] Patient-specific mutation detection can be used to confirm whether the DNA fragment is tumor-derived based on its fragment size distribution and GMM joint log odds ratio. To increase reliability and reduce batch effect bias, we developed intra-patient controls using inter-patient cross-detection. For example, specific patients, shown below the detected tumor mutation (gray, concordant detection), show a tendency for fragment size to shift to smaller sizes. Mutations associated with other patients in the same patient sample (red inter-patient detection) share the same tobacco pattern context information, but are not true detections. Interestingly, the inter-patient detections did not show a tendency for a lower fragment size shift, and their fragment size distributions were significantly different from the true tumor detections (Wilcoxon rank sum, P value 3 × 10-9). Using the GMM joint log odds ratio, the patient-specific mutation detection was confirmed to be tumor-derived (joint log odds ratio = 0.3), while the artifactual mutation from the same patient sample was confirmed to be normal-derived (joint log odds ratio = -0.35). Representative data from three patients are shown in Figure 17E.
[0371] Orthogonal integration of fragment sizes in relation to CNV markers cfDNA fragment distribution has a unique profile due to DNA degradation in the blood circulation. Normal cfDNA samples show a shift in fragment size distribution (see Figures 17A and 17B above). Here, in the context of analyzing the center of mass distribution (COM), calculation of the COM of the first nucleosome (peak at approximately 170 bp) shows a shift to lower COM, which corresponds linearly to TFs.
[0372] Comparative analysis of fragment size center of mass (COM) between patients may be limited in terms of sensitivity and prone to batch effects. The local fragment size COM within a patient may vary due to epigenetic patterns and copy number events. Indeed, at amplification segments, the local increase in tumor fraction (due to an increase in the proportion of tumor DNA) results in a decrease in the local fragment size center of mass (COM). On the other hand, at deletion sites, the local decrease in tumor fraction (due to a decrease in the proportion of tumor DNA) results in an increase in the local fragment size center of mass (COM). The data are shown in Figure 18B.
[0373] Using the estimated Log2 and COM values for all windows across the genome, we calculated the median center of mass (COM), slope, and R^2 of the Log2 / COM linear model, which itself corresponds to the fraction of tumor DNA (Figure 18C). More specifically, the data show that the Log2 / FS correlation (R2) is strongly related to the fraction of tumor DNA (Figure 18D).
[0374] Each dot in this Figure 18D corresponds to a patient sample. The x-axis represents the correlation (R^2) between all Log2 and COM values for all 1 Mbp bins for this patient. This value shows a strong correlation with the orthogonal estimate of the specimen TF (y-axis). When checking the correlation between Log2 and COM in healthy plasma samples, the correlation (R^2 = 0.008) (see Figure 18F) is extremely low compared to the correlation value seen in cancer patients (Figure 18E).
[0375] The present disclosure relates to the following non-limiting embodiments. Embodiment 1: A method of genetic screening for cancer in a subject, comprising: (A) receiving a subject-specific genome-wide compendium of reads associated with a plurality of genetic markers from a biological sample of the subject, the genetic marker compendium being selected from the group consisting of single nucleotide variations (SNVs), short insertions and deletions (Indels), copy number variations, structural variations (SVs), and combinations thereof; (b) statistically classifying each read in the compendium as signal or noise based on (1) the read base quality (BQ), (2) the read mapping quality (MQ), (3) the read's estimated fragment size, and / or (4) the read's estimated allele fraction (VAF), and removing artifactual reads from the compendium; (c) (d) compiling a subject-specific signature comprising a plurality of true reads in the signature based on the noise removal step (c) and the filtering step (b); (e) statistically quantifying a confidence estimate that the subject's biological sample contains circulating tumor DNA (ctDNA) based on the degree of match between the subject-specific signature and the cancer pattern; and (f) screening the subject for cancer if the confidence estimate that the subject's biological sample contains a cancer-associated mutation pattern exceeds a given threshold.
[0376] Embodiment 2: The method of embodiment 1, wherein the subject's biological sample comprises plasma, cerebrospinal fluid, pleural effusion, ocular fluid, stool, urine, or a combination thereof.
[0377] Embodiment 3: The method of any one of embodiments 1 and 2, wherein the cancer pattern comprises a COSMIC tobacco pattern, a UV pattern, a breast cancer (BRCA) pattern, a microsatellite instability (MSI) pattern, an apolipoprotein B mRNA editing enzyme, a poly(ADP-ribose) polymerase (PARP) multiactivation pattern, or a catalytic polypeptide-like pattern.
[0378] Embodiment 4: The method of any of embodiments 1 to 3, wherein the cancer pattern comprises a pattern associated with a tissue-specific epigenetic pattern, such as a tissue-specific chromatin accessibility pattern.
[0379] Embodiment 5: The method of any of embodiments 1-4, further comprising filtering sequencing noise associated with each read in said list using a machine learning (ML) model that distinguishes between mutational signatures associated with cancer (true positives) and signatures associated with PCR or sequencing errors (false positives).
[0380] Embodiment 6: The method of any of embodiments 1 to 5, wherein the machine learning model comprises a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), a discriminant analysis, a nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof. The method of any of the embodiments, wherein the ML is trained...
Claims
1. 1. An in vitro method for screening a subject for cancer genes, comprising: (A) receiving a subject-specific genome-wide compendium of read(s) associated with a plurality of genetic markers from a biological sample isolated from the subject, the biological sample comprising a plasma sample, wherein the compendium of reads each comprises a single base pair in length, and the biological sample has a tumor fraction (TF), a ratio of tumor DNA molecules to normal DNA molecules, of less than 0.1%; (B) filtering recurring sites and / or germline mutations from the subject-specific genome-wide read list, wherein the recurring sites are based on a cohort of reference healthy plasma samples; (C) filtering noise from the subject-specific genome-wide read list using at least one error suppression protocol to generate a set of filtered subject-specific genome-wide reads; (D) using the filtered set of subject-specific genome-wide reads to compile a subject-specific pattern based on a comparison with a specific mutation pattern associated with a given mutagenesis process; and (E) statistically quantifying a confidence estimate that the biological sample isolated from the subject contains the cancer-associated mutation pattern based on a comparison of the exposure value of the cancer-associated mutation pattern to a cohort of background mutation patterns via the subject-specific pattern; An in vitro method comprising:
2. 1. An in vitro method for genetic screening of cancer in a subject, comprising: (A) receiving a subject-specific genome-wide read list associated with a plurality of genetic markers from a biological sample isolated from the subject, the biological sample comprising a plasma sample, the read list each comprising a copy number variation (CNV) or a structural variation (SV), and the biological sample having a tumor fraction (TF), which is a ratio of tumor DNA molecules to normal DNA molecules, of less than 0.1%; (B) dividing the list of readings into a plurality of windows; (C) calculating a set of features per said window, said features including a median depth coverage per said window and a representative fragment size per window, and optionally including split reads; (D) filtering artifactual sites from said list of reads, said filtering comprising removing repetitive sites generated on a cohort of reference healthy plasma samples from said list of reads; (E) normalizing the read list to create a filtered read set of the subject-specific genome-wide read list; and (F) using the filtered read set, (i) calculating a linear relationship between the set of features per window and converting the calculated relationship to an estimated tumor fraction using a regression model; and / or (ii) calculating an estimate of tumor fraction based on one or more integrated mathematical models as a function of the feature set per window calculated across the subject-specific genome-wide list of reads; An in vitro method comprising:
3. 1. A system for genetic screening of a subject for cancer, comprising: An analytical unit comprising: A pre-filter engine configured and arranged to receive a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample isolated from a subject, wherein the biological sample comprises a plasma sample, and wherein the list of reads each comprises a single base pair in length read, and wherein the biological sample has a low tumor fraction (TF), which is a ratio of tumor DNA molecules to normal DNA molecules, of less than 0.1%; and filtering repetitive sites and / or germline mutations from the list of reads, wherein the repetitive sites are based on a cohort of reference healthy plasma samples; and an analysis unit comprising: a correction engine configured and arranged to filter noise in the list of reads using at least one error suppression protocol to create the filtered read set of the subject-specific genome-wide list of reads; and a computing unit configured and arranged to use the filtered read set to compile a subject-specific pattern based on comparison with a specific mutation pattern associated with a predetermined mutagenesis process, and via the subject-specific pattern, to statistically quantify a confidence estimate that a biological sample isolated from the subject contains a cancer-associated mutation pattern based on comparison of the cancer-associated mutation pattern with a cohort of background mutation patterns of an exposure value.
4. 1. A system for detecting residual tumor in a subject in need thereof, comprising: An analytical unit comprising: a binning engine configured and arranged to receive a subject-specific genome-wide list of reads associated with a plurality of genetic markers from a biological sample isolated from a subject, wherein the biological sample comprises a plasma sample, the list of reads each comprising copy number variations (CNVs), and the biological sample has a tumor fraction (TF), a ratio of tumor DNA molecules to normal DNA molecules, of less than 0.1%; Dividing the read list into windows and computing a set of features per window, wherein the features include a median depth coverage per window and a representative fragment size per window; a pre-filter engine configured and arranged to filter artifactual sites from the list of reads, wherein the filtering comprises removing repetitive sites generated on a cohort of reference healthy plasma samples from the list of reads; and an analysis unit comprising: a normalization engine configured and arranged to normalize the list of reads to create a filtered read set of the subject-specific genome-wide list of reads; A computing unit comprising: Using the filtered read set, (i) calculating a linear relationship between the set of features per window and converting the calculated relationship to an estimated tumor fraction using a regression model; and / or (ii) a computing unit configured and arranged to calculate an estimate of tumor fraction based on one or more integrated mathematical models as a function of the computed feature set per window across the subject-specific genome-wide list of reads.
5. 2. The in vitro method of claim 1, wherein the genetic marker comprises a single nucleotide variation (SNV) or an insertion / deletion (indel).
6. 2. The in vitro method of claim 1, wherein filtering the repetitive sites based on a cohort of reference healthy plasma samples comprises generating a panel of normal (PON) blacklist or mask.
7. 2. The in vitro method of claim 1, wherein the reference healthy plasma sample comprises peripheral blood mononuclear cells (PBMCs).
8. 2. The in vitro method of claim 1, comprising using a machine learning model in step (C), wherein the machine learning model comprises a deep convolutional neural network (CNN), a recurrent neural network (RNN), a random forest (RF), a support vector machine (SVM), a discriminant analysis, a nearest neighbor analysis (KNN), an ensemble classifier, or a combination thereof.
9. 2. The in vitro method of claim 1, wherein in step (C), at least one error suppression protocol comprises removing artifacts using mismatch testing and / or overlap consensus between independent copies of the same DNA fragment generated by PCR or sequencing, wherein artifacts are identified and removed when there is no match in a majority of predetermined overlap families, and wherein correction of said artifacts comprises correction of artifacts generated by PCR or sequencing using comparison of independent copies of the original nucleic acid fragment.
10. 10. The in vitro method of claim 9, wherein in step (C), artifacts generated by paired-end 150 bp sequencing that result in overlapping paired reads (R1 and R2) are removed by correcting mismatches between R1 and R2 pairs to the corresponding reference genome.
11. In step (C), the at least one error suppression protocol comprises: (a) calculating the probability that any single nucleotide variation in said list is an artifact, and removing said artifacts, wherein said probability is calculated as a function of features selected from the group including mapping quality (MQ), variant base quality (MBQ), position in read (PIR), mean read base quality (MRBQ), and combinations thereof; and / or 2. The in vitro method of claim 1, wherein (b) removing artifacts using mismatch testing and / or overlap consensus between independent copies of the same DNA fragment generated by polymerase chain reaction or sequencing processing, wherein artifacts are identified and removed when there is a majority non-match in a given overlap family, correcting for artifacts generated by overlaps during sequencing and / or PCR amplification, wherein overlap families are recognized by 5' and 3' similarity and alignment position, and each overlap family is used to identify a consensus of specific mutations across the independent copies, thereby correcting for artifacts where there is a majority non-match in the majority of the overlap families.
12. 2. The in vitro method of claim 1, wherein in step (D), specific mutation patterns in a single plasma sample are identified using the non-negative least squares (NNLS) method.
13. 2. The in vitro method of claim 1, wherein step (E) further verifies the reliability of the specific mutation pattern by comparing the exposure value of the cancer-associated mutation pattern with exposure values estimated for a plurality of random background patterns.
14. 14. The in vitro method of claim 13, wherein in step (F), the subject is identified as having cancer if the confidence estimate that the biological sample isolated from the subject contains the cancer-associated mutation pattern exceeds a predetermined threshold of z-score > 2std.
15. 2. The in vitro method of claim 1, wherein step (D) comprises using a machine learning model, wherein the machine learning model is trained to distinguish between sequencing reads altered by cancer and reads altered by sequencing errors.
16. 16. The in vitro method of claim 15, wherein the machine learning model is trained on a plurality of true mutation reads and errors, and the trained machine learning model is capable of distinguishing between reads containing true mutations and reads containing artifactual sequences.
17. 10. The in vitro method of claim 1, further comprising orthogonal integration of secondary features comprising fragment size shift.
18. 18. The in vitro method of claim 17, wherein intra-patient fragment size shifts in the list of tumor-specific markers and random markers are analyzed using statistical methods, such as tests of significance or Gaussian mixture models (GMM).
19. 3. The in vitro method of claim 2, wherein the genetic marker comprises a copy number variation (CNV).
20. 3. The in vitro method of claim 2, wherein in step (B), each window is at least ≧150 bp.
21. 3. The in vitro method of claim 2, wherein step (C) comprises extracting depth coverage (Log2) and fragment size (COM) relationships (gradient, R^2) from the subject-specific genome-wide feature vectors.
22. 3. The in vitro method of claim 2, wherein step (D) comprises generating a blacklist or mask of normal panels (PONs) to filter repetitive sites generated on a cohort of reference healthy plasma samples and / or to filter windows with low mappability or coverage.
23. 3. The in vitro method of claim 2, wherein the normalization step comprises normalizing depth coverage values to correct for GC content and mappability bias by performing two LOESS regression curve fits to the bin-wise GC fraction and mappability scores.
24. 3. The in vitro method of claim 2, wherein the normalization step comprises batch effect correction using robust z-score normalization applied to each sample separately.
25. 25. The in vitro method of claim 24, wherein normalizing the z-scores comprises calculating the median and median absolute deviation (MAD) based on the neutral region for each sample to normalize all CNV bins, where normalization is achieved by subtracting the median and dividing by the MAD.
26. 3. The in vitro method of claim 2, wherein step (E) comprises calculating depth coverage skew and / or fragment size center of gravity (COM) skew in the plasma sample compared to a normal panel of healthy plasma samples (PON).
27. 3. The in vitro method of claim 2, wherein step (F) comprises calling copy number variations (CNVs) and calculating tumor fraction of the filtered read set using a hidden Markov model or a self-organizing neural network, such as a neural network based on adaptive resonance theory or a self-organizing map.
28. 3. The in vitro method of claim 2, further comprising orthogonal integration of secondary features comprising fragment size shifts.
29. 29. The in vitro method of claim 28, wherein intra-patient fragment size shifts in the list of tumor-specific markers and random markers are analyzed using statistical methods, such as tests of significance or Gaussian mixture models (GMM).