Machine learning implementation for multi-analyte assays of biological samples
By using machine learning to analyze multiple classes of molecules in biological samples, particularly focusing on non-cellular components in plasma, the challenges of current cancer screening methods are addressed, leading to improved early detection and treatment guidance.
Patent Information
- Application Number
- JP2024038608
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-27
- Filing Date
- 2024-03-13
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2039-04-15
AI Technical Summary
Current cancer screening methods face challenges such as low patient compliance due to the need for non-serum specimens, and existing blood-based tests are limited in analyzing a single class of molecules, leading to biological noise and inefficiencies in early detection and treatment guidance.
The development of methods and systems that incorporate machine learning approaches to analyze multiple classes of molecules in biological samples, focusing on the non-cellular portion of circulation, such as plasma, to stratify individuals at risk for cancer and guide treatment decisions.
This approach enables the characterization of early cancer and improves treatment decision-making by reducing the need for significant blood volumes, enhancing the sensitivity and specificity of cancer detection, and providing a surrogate indicator of cancer status.
Smart Images

Figure 0007681145000013 
Figure 0007681145000014 
Figure 0007681145000015
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is U.S. Provisional Patent Application No. US62 / 657,602, filed April 13, 2018; No. US62 / 749,955, filed October 24, 2018; No. US62 / 679,641, filed June 18, 2018; No. US62 / 767,435, filed November 14, 2018; No. US62 / 679,587, filed June 1, 2018; No. US62 / 731,557, filed September 14, 2018; No. US62 / 742,799, filed October 8, 2018; No. US62 / 804,614, filed February 2, 2019; No. US62 / 767,369, filed November 14, 2018; and This application claims the benefit of US 62 / 824,709, filed March 29, 2019, the entire contents of which are incorporated by reference.
[0002] Cancer screening is complex, and various cancer types require different approaches for screening and early detection. Patient compliance remains an issue. Screening methods that require non-serum specimens often have low participation rates. Screening rates for breast, cervical, and colorectal cancer using mammograms, Pap tests, and sigmoidoscopy / fecal occult blood tests, respectively, are far from the 100% compliance recommended by the United States Preventive Services Task Force (USPSTF) (Sabatino et al, Cancer Screening Test Use-United States, 2013, MMWR, 2015 64(17):464-468; Adler et al. BMC Gastroenterology 2014,14:183). A recent report found that the percentage of adults eligible for colorectal cancer screening by state ranged from 58.5% (New Mexico) to 75.9% (Maine) in 2016, with a mean of 67.3% (Joseph DA, et al. Use of Colorectal Cancer Screening Tests by State. Prev Chronic Dis 2018;15:170535).
[0003] Blood-based tests hold great promise for cancer diagnostics and precision medicine. However, most current tests are limited to the analysis of a single class of molecules (e.g., circulating tumor DNA, platelet mRNA, circulating proteins). A wide range of biological analytes are present in blood for potential analysis, and the generation of relevant data is critical. However, analysis of the entire analyte is laborious, uneconomical, and injects a very large amount of biological noise compared to the useful signal, which can confound useful analyses for diagnostic or precision medicine applications.
[0004] Even with early detection and genomic characterization, there remain a significant number of cases where genomic analysis fails to designate an effective drug or applicable clinical trial. Even when targetable genomic alterations are found, patients do not always respond to treatment. (Pauli et al.,Cancer Discov.2017,7(5):462-477). Furthermore, with regard to detection methods, sensitivity barriers exist for the use of circulating tumor DNA (ctDNA). ctDNA has recently been evaluated as a prospective specimen for detecting early cancer and has been found to require significant amounts of blood to detect ctDNA with the required specificity and sensitivity. (Aravanis,A.et al.,Next-Generation Sequencing of Circulating Tumor DNA for Early Cancer Detection,Cell,168:571-574). Thus, a simple and readily available single specimen test remains unclear.
[0005] In the field of cancer diagnostics, machine learning can enable large-scale statistical approaches and automated characterization of signal strength. However, machine learning applied to biology in the context of molecular diagnostics remains a largely unexplored field and has not been previously applied in aspects of diagnostics and precision medicine such as specimen selection, assay selection, and overall optimization.
[0006] Thus, there is a need for methods to analyze readily obtained biological specimens to stratify individuals at risk for or with cancer and provide effective characterization of early cancer to guide treatment decisions. There is also a need for methods to incorporate machine learning approaches with specimen datasets to develop and refine classifiers for use in stratifying populations of individuals and detecting diseases such as cancer. Summary of the Invention
[0007] Described herein are methods and systems that incorporate machine learning approaches using one or more biological analytes in a biological sample for various applications to stratify individual populations. In certain embodiments, the methods and systems are useful for predicting disease, treatment efficacy, and guiding treatment decisions for affected individuals.
[0008] This approach differs from other methods and systems in that it focuses on an approach that characterizes the non-cellular portion of the circulation, including specimens derived from tumor cells, healthy non-tumor cells induced or educated by the microenvironment, as well as circulating immune cells that may be educated by tumor cells present in an individual.
[0009] While other approaches are directed to characterizing the cellular parts of the immune system, the present method and system investigates the cancer-inducing non-cellular parts of the circulation to provide biological information that can then be combined with machine learning tools for useful applications. Studying non-cellular specimens in liquid biological samples (e.g., plasma) can allow for deconvolution of the sample to recreate the molecular state of an individual's tissues and immune cells in a living cellular state. Studying the non-cellular parts of the immune system provides a surrogate indicator of cancer status and avoids the requirement of significant blood volume to detect cancer cells and associated biological markers when screening for ctDNA alone.
[0010] In a first aspect, the present disclosure provides a method of using a classifier capable of discriminating between populations of individuals, comprising: a) assaying a plurality of classes of molecules in a biological sample, the assay providing a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules that are input into a machine learning or statistical model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of the computer system a machine learning model including a classifier, the machine learning model trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified as having a specified property, and a second subset of the training biological samples identified as not having the specified property; e) inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample has the specified property or not, thereby identifying a population of individuals having the specified property.
[0011] By way of example, the class of molecules can be selected from nucleic acids, polyamino acids, carbohydrates, or metabolites. As further examples, the class of molecules can include nucleic acids including deoxyribonucleic acid (DNA), genomic DNA, plasmid DNA, complementary DNA (cDNA), cell-free (e.g., non-encapsulated) DNA (cfDNA), circulating tumor DNA (ctDNA), nucleosomal DNA, chromatosomal DNA, mitochondrial DNA (miDNA), artificial nucleic acid analogs, recombinant nucleic acids, plasmids, viral vectors, and chromatin. In one example, the sample includes cfDNA. In one example, the sample includes peripheral blood mononuclear cell-derived (PBMC-derived) genomic DNA.
[0012] As a further example, the class of molecules can include nucleic acids, including ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (mitoRNA), ribosomal RNA (rRNA), circular RNA (cRNA), alternatively spliced mRNA, small nuclear RNA (snRNA), antisense RNA, short hairpin RNA (shRNA), or small interfering RNA (siRNA).
[0013] As a further example, the class of molecules can include polyamino acids, including polyamino acids, peptides, proteins, autoantibodies, or fragments thereof.
[0014] As further examples, the classes of molecules can include sugars, lipids, amino acids, fatty acids, phenolic compounds, or alkaloids.
[0015] In various embodiments, the multiple classes of molecules include at least two of cfDNA molecules, cfRNA molecules, circulating proteins, antibodies, and metabolites.
[0016] As with the aspects of the disclosure and various examples of the systems and methods herein, the classes of molecules can be selected from 1) cfDNA, cfRNA, polyamino acids and small chemical molecules, or 2) cfDNA and cfRNA and polyamino acids, 3) cfDNA and cfRNA and small chemical molecules, or 4) cfDNA, polyamino acids and small chemical molecules, or 5) cfRNA, polyamino acids and small chemical molecules, or 6) cfDNA and cfRNA, or 7) cfDNA and polyamino acids, or 8) cfDNA and small chemical molecules, or 9) cfRNA and polyamino acids, or 10) cfRNA and small chemical molecules, or 11) polyamino acids and small chemical molecules.
[0017] In one example, the multiple classes of molecules are cfDNA, proteins, and autoantibodies.
[0018] In various examples, the plurality of assays can include at least two of whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), small RNA sequencing, quantitative immunoassays, enzyme-linked immunosorbent assays (ELISA), proximity extension assays (PEA), protein microarrays, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing (WO2019 / 051484), CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56 gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing, as well as mixture proportions of cell types or cell phenotypes derived from any of the above assays.
[0019] In one embodiment, the whole genome bisulfite sequencing includes methylation analysis.
[0020] In various embodiments, the classification of the biological sample is performed by a classifier trained and constructed according to one or more of Linear Discriminant Analysis (LDA), Partial Least Squares (PLS), Random Forest, k-Nearest Neighbor (KNN), Support Vector Machine (SVM) with a radial basis function kernel (SVMRadial), SVM with a linear basis function kernel (SVMLinear), SVM with a polynomial basis function kernel (SVMPoly), decision trees, multi-layer perceptrons, mixture of experts, sparse factor analysis, hierarchical decomposition, and a combination of linear algebraic routines and statistics.
[0021] In various examples, the designated property may be a clinically diagnosed disorder. The clinically diagnosed disorder may be cancer. By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer. In some examples, the designated property is responsiveness to treatment. In one example, the designated property may be a continuous measure of a patient's trait or phenotype.
[0022] In a second aspect, the present disclosure provides a system for classifying a biological sample, comprising: a) a receiver for receiving a plurality of training samples, each of the plurality of training samples having a plurality of classes of molecules, each of the plurality of training samples including one or more known labels; b) a feature module that identifies, for each of a plurality of training samples, a set of features corresponding to an assay operable to be input to the machine learning model, the set of features corresponding to properties of molecules in the plurality of training samples; For each of a plurality of training samples, the system is operable to subject a plurality of classes of molecules in the training sample to a plurality of different assays to obtain a set of measurements, each set of measurements being from one assay applied to a class of molecules in the training sample, the plurality of sets of measurements being obtained for the plurality of training samples; and c) an analysis module for analyzing the set of measurements to obtain a training vector of the training sample, the training vector including N sets of feature quantities of features of the corresponding assays, each feature quantity corresponding to a feature and including one or more measurements, the training vector being formed using at least one feature from at least two of the N sets of features corresponding to a first subset of the plurality of different assays; d) a label module that uses parameters of the machine learning model to inform the system of the training vectors to obtain output labels for a plurality of training samples; e) a comparator module for comparing the output labels with known labels of the training samples; f) a training module that iteratively searches for optimal values of parameters as part of training the machine learning model based on comparing the output labels with known labels of the training samples; g) an output module that provides machine learning model parameters and a set of machine learning model features.
[0023] In a third aspect, the present disclosure provides a system for classifying a subject based on a multi-analyte analysis of a biological sample composition, comprising: (a) a computer-readable medium comprising a classifier operable to classify the subject based on the multi-analyte analysis; and (b) one or more processors for executing instructions stored on the computer-readable medium.
[0024] In one embodiment, the system includes a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and an NMF prediction algorithm classifier.
[0025] In one embodiment, the system includes means for performing any of the aforementioned methods. In one embodiment, the system includes one or more processors configured to perform any of the aforementioned methods. In one embodiment, the system includes modules for performing each of the steps of any of the aforementioned methods.
[0026] Another aspect of the present disclosure provides a non-transitory computer-readable medium that includes machine-executable code that, when executed by one or more computer processors, implements the above or any other method herein.
[0027] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or otherwise herein.
[0028] In a fourth aspect, the present disclosure provides a method for detecting the presence of cancer in an individual, comprising: a) assaying a plurality of classes of molecules in a biological sample obtained from an individual, the assay providing a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules to be input to the machine learning model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of the computer system a machine learning model that is trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals with cancer, and a second subset of the training biological samples identified from individuals without cancer; e) detecting the presence of cancer in the individual by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with cancer or not.
[0029] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value being indicative of the presence of cancer in the individual.
[0030] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value being indicative of a stage of cancer in the individual.
[0031] By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer or breast cancer. In one embodiment, the cancer is colon cancer.
[0032] In a fifth aspect, the present disclosure provides a method of determining a prognosis for an individual with cancer, comprising: a) assaying a plurality of classes of molecules in a biological sample, the assay providing a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to properties of multiple classes of molecules to be input into a machine learning model; creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; c) loading into the memory of the computer system a machine learning model that is trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals having a favorable cancer prognosis, and a second subset of the training biological samples identified from individuals not having a favorable cancer prognosis; d) determining a prognosis for an individual with cancer by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with a good cancer prognosis.
[0033] By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer or breast cancer.
[0034] In a sixth aspect, the present disclosure provides a method for determining responsiveness to cancer treatment, comprising: a) assaying a plurality of classes of molecules in a biological sample, the assay providing a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules to be input to the machine learning model; creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; c) loading into the memory of the computer system a machine learning model that is trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals that are responsive to the treatment, and a second subset of the training biological samples identified from individuals that are not responsive to the treatment; d) determining responsiveness to cancer treatment by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with treatment response.
[0035] In one embodiment, the cancer therapy is selected from an alkylating agent, a plant alkaloid, an antitumor antibiotic, an antimetabolite, a topoisomerase inhibitor, a retinoid, a checkpoint inhibitor therapy, or a VEGF inhibitor.
[0036] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value being indicative of a response to treatment in the individual.
[0037] These and other embodiments are described in more detail below. For example, other embodiments are directed to systems, devices, and computer-readable media associated with the methods described herein.
[0038] The nature and advantages of the embodiments of the present disclosure may be better understood by reference to the following detailed description and accompanying drawings. [Brief description of the drawings]
[0039] [Figure 1] FIG. 1 illustrates an exemplary system that is programmed or otherwise configured to implement the methods provided herein. [Diagram 2] FIG. 2 is a flow chart illustrating a method for analyzing a biological sample. [Diagram 3] FIG. 3 illustrates an overall framework in accordance with various embodiments. [Figure 4] FIG. 4 shows an overview of the multi-analyte approach. [Diagram 5] FIG. 5 illustrates an iterative process for designing an assay and a corresponding machine learning model, according to various embodiments. [Figure 6] FIG. 6 is a flow chart illustrating a method for classifying a biological sample, according to one embodiment. [Figure 7] 7A and 7B show the classification performance for different analytes. [Figure 8A] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8B] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8C] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8D] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8E] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8F] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8G] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8H] Figures 8A-H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 9] FIG. 9 shows CpG methylation analysis at LINE-1 sites. [Figure 10] FIG. 10 shows cf-miRNA sequence analysis. [Figure 11A] FIG. 11A shows circulating protein biomarker distribution. [Figure 11B] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 11C] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 11D] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 11E] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 11F] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 11G] Figures 11B-11G show that protein levels differed significantly across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. [Figure 12A] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12B] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12C] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12D]Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12E] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12F] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12G] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12H] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 13] FIG. 13 shows a heatmap of chromosome structure scores determined from nuance structure of correlation matrices generated using Pearson / Spearman / Kendall correlation of regions of the genome using cfDNA samples. [Figure 14] FIG. 14 shows a heat map of chromosome structure scores determined from Hi-C sequencing of the same genomic region as in FIG. [Figure 15A-C] Figures 15A-C show correlation maps generated from Hi-C, spatially correlated fragment length distributions from multiple cfDNA samples, and spatially correlated fragment length distributions from a single cfDNA sample. [Figure 15D] Figure 15D shows genome browser tracks of partition A / B from Hi-C, multi-sample cfDNA, and single-sample cfDNA. [Figure 15E] Figures 15E-F show scatter plots of concordance at the compartment level between Hi-C, multi-sample cfDNA (Figure 15E), and single-sample cfDNA (Figure 15F). [Figure 15F] Figures 15E-F show scatter plots of concordance at the compartment level between Hi-C, multi-sample cfDNA (Figure 15E), and single-sample cfDNA (Figure 15F). [Figure 16A] FIG. 16A shows the correlation between Hi-C and cfHi-C at the pixel level (500 kb bins). [Figure 16B] FIG. 16B shows the correlation between Hi-C and cfHi-C at the compartment level (500 kb bins). [Figure 17A-D] Figure 17A shows a heat map of cfHi-C before G+C% is regressed from the fragment length of each bin on chr1 by LOWESS. Figure 17B shows a heat map of cfHi-C after G+C% is regressed from the fragment length of each bin on chr1 by LOWESS. Figure 17C shows a heat map of gDNA before G+C% is regressed from the fragment length of each bin on chr1 by LOWESS. Figure 17D shows a heat map of gDNA after G+C% is regressed from the fragment length of each bin on chr1 by LOWESS. [Figure 17E] Figure 17E shows boxplots of pixel-level correlations (Pearson and Spearman) with Hi-C (WBC, replicate 2) across all chromosomes represented in Figures 17A-17D. [Figure 18A] FIG. 18A shows G+C % and mappability bias analysis in two-dimensional space from multi-sample cfHi-C. [Figure 18B] FIG. 18B shows G+C % and mappability bias analysis in two-dimensional space from a single sample cfHi-C. [Figure 18C] FIG. 18C shows G+C % and mappability bias analysis in two-dimensional space from multi-sample genomic DNA. [Figure 18D] FIG. 18D shows G+C % and mappability bias analysis in two-dimensional space from a single sample genomic DNA. [Figure 18E] FIG. 18E shows G+C % and mappability bias analysis in two-dimensional space from multi-sample cfHi-C. [Figure 18F] FIG. 18F shows G+C % and mappability bias analysis in two-dimensional space from Hi-C (WBC). [Figure 19A] FIG. 19A shows a heatmap of multi-sample cfHi-C where one pair of bins was randomly shuffled from any other individual (chr14). [Figure 19B] FIG. 19B shows a multi-sample cfHi-C heatmap for samples from the same batch as in FIG. 19A (11 samples; chr14). [Figure 19C] FIG. 19C shows a heatmap of multi-sample cfHi-C for samples with the same sample size as in FIG. 19B (11 samples; chr14). [Figure 19D] FIG. 19D shows a boxplot of pixel-level correlation with Hi-C (WBC, repeat 2) across all chromosomes represented in FIGS. 19A-19C. [Figure 20A] FIG. 20A shows the Pearson correlation between Hi-C (WBC, replicate 1) and multi-sample cfHi-C at different sample sizes. [Figure 20B] FIG. 20B shows the Spearman correlation between Hi-C (WBC, replicate 1) and multi-sample cfHi-C at different sample sizes. [Figure 20C] FIG. 20C shows the Pearson correlation between Hi-C (WBC, replicate 2) and multi-sample cfHi-C at different sample sizes. [Figure 20D] FIG. 20D shows the Spearman correlation between Hi-C (WBC, replicate 2) and multi-sample cfHi-C at different sample sizes. [Figure 21A] FIG. 21A shows pixel-level Pearson correlation between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21B] FIG. 21B shows pixel-level Spearman correlation between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21C] FIG. 21C shows pixel-level Pearson correlation between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 21D]FIG. 21D shows pixel-level Spearman correlation between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 21E] FIG. 21E shows Pearson correlation at the compartment level between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21F] FIG. 21F shows Spearman correlation at the compartment level between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21G] FIG. 21G shows Pearson correlations at the compartment level between Hi-C and single-sample cfHi-C at different bin sizes. [Fig. 21H] FIG. 21H shows the Spearman correlation at the compartment level between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 22A] FIG. 22A shows pixel-level Pearson and Spearman correlations between Hi-C and single-sample cfHi-C at different read numbers after downsampling. [Figure 22B] FIG. 22B shows Pearson and Spearman correlations at the section level between Hi-C and single-sample cfHi-C at different read numbers after downsampling. [Figure 23A] FIG. 23A shows kernel PCA (RBF kernel) of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23B] 23B-F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23C] 23B-F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23D] 23B-F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23E] 23B-F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23F]23B-F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 24] FIG. 24 shows a correlation map between DNA accessibility from Hi-C from the same cell type (GM12878) and compartment-level eigenvalues. [Figure 25-1] Figure 25A shows a heat map of the cellular composition inferred from single-sample cfDNA of healthy, colon cancer, lung cancer, and melanoma samples. Figure 25B shows a pie chart of the cellular composition inferred from single-sample cfDNA of healthy, colon cancer, lung cancer, and melanoma samples. [Figure 25-2] Figure 25C shows box plots of leukocyte and tumor fractions inferred from single sample cfDNA from 100 healthy individuals. [Figure 26] FIG. 26 shows a comparison of ichorCNA-derived tumor fractions with cfHi-C-derived tumor fractions by using only genomic regions without CNV alterations for lung, melanoma, and colon cancer. [Figure 27] Figure 27A shows the k-fold, k-batch, balanced k-batch, and ordered k-batch training schemes, and Figure 27B shows the k-batch with systematic downsampling scheme. [Figure 28A] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28B] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28C] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28D]28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28E] FIG. 28E shows sensitivity by CRC stage across all validation approaches evaluated. [Figure 28F] FIG. 28F shows the AUC by IchorCNA estimated tumor fraction across all validation approaches evaluated. [Figure 28G] FIG. 28G shows the AUC by age bin across all validation approaches evaluated. [Fig. 28H] FIG. 28H shows the AUC by gender bin across all validation approaches evaluated. [Figure 29A] 29A-B show classification performance in cross-validation (ROC curves) for breast cancer. [Figure 29B] 29A-B show classification performance in cross-validation (ROC curves) for breast cancer. [Figure 29C] 29C to 29D show classification performance in cross-validation (ROC curve) for liver cancer. [Figure 29D] 29C to 29D show classification performance in cross-validation (ROC curve) for liver cancer. [Figure 29E] Figures 29E-F show classification performance in cross-validation (ROC curves) for pancreatic cancer. [Figure 29F] Figures 29E-F show classification performance in cross-validation (ROC curves) for pancreatic cancer. [Diagram 30] FIG. 30 shows the distribution of estimated tumor fractions (TF) by class. [Figure 31A] FIG. 31A shows the AUC performance of CRC classification when the training set for each fold is downsampled as either a proportion of the samples. [Figure 31B] FIG. 31B shows the AUC performance of CRC classification when the training set for each fold is downsampled as either a proportion of the samples or a proportion of the batch. [Figure 32A] 32A-C show examples of healthy samples with high tumor fractions. [Figure 32B] 32A-C show examples of healthy samples with high tumor fractions. [Figure 32C] 32A-C show examples of healthy samples with high tumor fractions. [Diagram 33] Figure 33A shows the training method and cross-validation procedure for the k-fold model. Figure 33B shows the k-fold, k-batch, and balanced k-batch training schemes. [Figure 34A] FIG. 34A shows sensitivity by CRC stage in patients aged 50-84 years. [Figure 34B] FIG. 34B shows sensitivity by tumor fraction in patients aged 50-84 years. [Figure 34C] FIG. 34C shows the AUC performance of CRC classification among the total number of samples. [Diagram 35] Figure 35 shows a schematic of V-plots derived from cfDNA captured protein-DNA associations showing chromatin structure and transcriptional state. TF = transcription factor (small footprint region protected), NS = nucleosome (large region protected, complete wrap of DNA). [Diagram 36] FIG. 36 shows V-plots from cfDNA surrounding the TSS regions used to predict gene expression. [Figure 37] FIG. 37 shows a classifier that uses fragment length and position representations to accurately classify on and off genes using different cutoffs. [Figure 38A] Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. Tumor fraction estimates (ITFs) based on IchorCNA increase with stage, but most stage I-III CRCs have low estimated ITFs (<1%) (Figure 38A). Performance increases with stage, with the most pronounced increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 38B]Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. Tumor fraction estimates (ITFs) based on IchorCNA increase with stage, but most stage I-III CRCs have low estimated ITFs (<1%) (Figure 38A). Performance increases with stage, with the most pronounced increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 38C] Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. Tumor fraction estimates (ITFs) based on IchorCNA increase with stage, but most stage I-III CRCs have low estimated ITFs (<1%) (Figure 38A). Performance increases with stage, with the most pronounced increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 39A] FIG. 39A shows tumor fraction estimates versus 44 colon gene mean P(on). [Figure 39B] FIG. 39B shows the fold change from the mean coverage for healthy samples containing strong evidence of copy number alterations in chr8 and chr9.
[0040] term The terms "a," "an," and "the" are intended to mean "one or more," unless specifically indicated to the contrary. Unless specifically indicated to the contrary, the use of "or" is intended to mean "inclusive or" rather than "exclusive or." A reference to a "first" element does not necessarily require that a second element be provided. Further, unless specifically stated, a reference to a "first" or "second" element does not limit the referenced element to a particular location. The term "based on" is intended to mean "based at least in part on."
[0041] The term "area under the curve" or "AUC" refers to the area under the curve of a receiver operating characteristic (ROC) curve. The AUC measure is useful for comparing the accuracy of classifiers across the complete data range. A classifier with a larger AUC has a greater ability to correctly classify unknowns between two groups of interest (e.g., cancer samples and normal or control samples). ROC curves are useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in discriminating between two populations (e.g., individuals who respond to a therapeutic agent and those who do not). Typically, the feature data across populations (e.g., cases and controls) are sorted in ascending order based on the value of a single feature. Then, for each value of that feature, the true positive rate and false positive rate of the data are calculated. The true positive rate is determined by counting the number of cases above the value of that feature and dividing by the total number of cases. The false positive rate is determined by counting the number of controls above the value of that feature and dividing by the total number of controls. While this definition refers to scenarios where a feature is elevated in cases compared to controls, this definition also applies to scenarios where a feature is lower compared to controls (in such scenarios, samples below the value of the feature may be counted). ROC curves can be generated for single features as well as other single outputs, for example, combinations of two or more features can be mathematically combined (e.g., added, subtracted, multiplied, etc.) to provide a single total value, which can be plotted on a ROC curve. Additionally, any combination of multiple features whose combination derives a single output value can be plotted on a ROC curve. These combinations of features may include tests. A ROC curve is a plot of the true positive rate (sensitivity) of a test against the false positive rate (1 specificity) of the test.
[0042] The term "biological sample" (or simply "sample") refers to any material obtained from a subject. A sample may contain or be presumed to contain an analyte from a subject, such as those described herein (nucleic acid, polyamino acid, carbohydrate, or metabolite). In some embodiments, a sample may include cells and / or acellular material obtained in vivo, cultured in vitro, or processed in situ, as well as lineages, including lineages and genealogies. In various embodiments, a biological sample may be tissue, such as a normal or healthy tissue from a subject (e.g., solid or liquid tissue). Examples of solid tissue include primary tumors, metastatic tumors, polyps, or adenomas. Examples of liquid samples (e.g., bodily fluids) include whole blood, buffy coat from blood (which may contain lymphocytes), urine, saliva, cerebrospinal fluid, plasma, serum, ascites, sputum, sweat, tears, buccal samples, cavity rinses, or organ rinses. In some cases, the liquid is an acellular liquid, is an essentially acellular liquid sample, or contains acellular nucleic acid, e.g., acellular DNA. In some cases, cells, including circulating tumor cells, can be enriched or isolated from the liquid.
[0043] The terms "cancer" and "cancerous" refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. Neoplasia, malignancy, cancer, and tumor are often used interchangeably and refer to the abnormal growth of tissue or cells resulting from excessive cell division.
[0044] The term "cancer-free" refers to a subject whose organs have not been diagnosed with cancer or have no detectable cancer.
[0045] The term "genetic variant" (or "variant") refers to a deviation from one or more expected values. Examples include sequence variants or structural polymorphisms. In various examples, a variant can refer to a known variant, such as a variant that has been scientifically confirmed and reported in the literature, a putative variant associated with a biological change, a putative variant that has been reported in the literature but has not yet been biologically confirmed, or a putative variant that has not been reported in the literature but is inferred based on computational analysis.
[0046] The term "germline variant" refers to a nucleic acid that induces a natural or normal polymorphism (e.g., skin color, hair color, and normal weight). Somatic mutations can refer to nucleic acids that induce acquired or abnormal polymorphisms (e.g., cancer, obesity, conditions, diseases, disorders, etc.). Germline variants are inherited and therefore correspond to the genetic differences of an individual born relative to the reference human genome. Somatic variants are variants that arise at zygote or later, at any time during cell division, development, and aging. In some examples, the analysis can distinguish between germline variants, e.g., private variants, and somatic variants.
[0047] The term "input feature" (or "feature") refers to a variable used by a model to predict the output classification (label) of a sample, e.g., a status, sequence content (e.g., mutations), a recommended data collection action, or a recommended treatment. A value of the variable can be determined for a sample and used to determine the classification. Examples of input features of genetic data include aligned variables related to the alignment of sequence data (e.g., sequence reads) to a genome, and unaligned variables related, for example, to the sequence content of sequence reads, protein or autoantibody measurements, or average methylation levels in a genomic region.
[0048] The term "machine learning model" (or "model") refers to a collection of parameters and functions where the parameters are trained on a set of training samples. The parameters and functions can be a collection of linear algebraic operations, nonlinear algebraic operations, and tensor algebraic operations. The parameters and functions can include statistical functions, tests, and probability models. The training samples can correspond to samples with measured characteristics of the samples (e.g., genomic data and other subject data such as images or health records), as well as known classifications / labels of the subjects (e.g., phenotypes or treatments). The model can learn from the training samples in a training process that optimizes the parameters (and potentially the features) to provide an optimal quality metric (e.g., accuracy) for classifying new samples. The training functions can include Bayesian parameter estimation methods such as expectation maximization, maximum likelihood, Markov chain Monte Carlo, Gibbs sampling, Hamiltonian Monte Carlo, and variational inference, or gradient-based methods such as stochastic gradient descent and the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm. Exemplary parameters include weights (e.g., vector or matrix transformations) that multiply values, for example, in regression or neural networks, probability distribution families, or loss, cost, or objective functions that assign scores and guide model training. Exemplary parameters include weights that multiply values, for example, in regression or neural networks. A model can include multiple sub-models, may be models of different layers or independent models, and may have different structural forms, for example, a combination of a neural network and a support vector machine (SVM). Examples of machine learning models include deep learning models, neural networks (e.g., deep learning neural networks), kernel-based regression, adaptation-based regression or classification, Bayesian methods, ensemble methods, logistic regression and augmentation, Gaussian processes, support vector machines (SVM), probabilistic models, and probabilistic graphical models.Machine learning models can further include feature engineering (e.g., collection of features into data structures such as one-dimensional, two-dimensional, or larger dimensional vectors) and feature representation (e.g., processing of the data structure of features into transformed features used for training for classification inference).
[0049] A "marker" or "marker protein" is a diagnostic indicator found in a patient and is detected directly or indirectly by the method of the present invention. Indirect detection is preferred. In particular, all of the markers of the present invention have been shown to cause the production of (auto)antigens in cancer patients or in patients at risk of developing cancer. Therefore, a simple way to detect these markers is to detect these (auto)antibodies in blood or serum samples from the patient. Such antibodies can be detected by binding to their respective antigens in an assay. Such antigens are in particular the marker proteins themselves or antigen fragments thereof. Suitable methods can be used to specifically detect such antibody-antigen reactions and can be used according to the system and method of the present disclosure. Preferably, the entire antibody content of the sample is normalized (e.g., diluted to a pre-set concentration) and applied to the antigen. Preferably, IgG, IgM, IgD, IgA or IgE antibody fractions are used exclusively. The preferred antibody is IgG.
[0050] The term "non-cancerous tissue" refers to tissue from the same organ in which a malignant neoplasm has formed but which does not have the pathology characteristic of the neoplasm. Generally, non-cancerous tissue appears histologically normal. As used herein, "normal tissue" or "healthy tissue" refers to tissue from an organ in which the organ is not cancerous.
[0051] The terms "polynucleotide", "nucleotide", "nucleic acid", and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, minimally linked together in length. In some examples, polynucleotides can have any three-dimensional structure and perform any function, known or unknown. Nucleic acids can include RNA, DNA, such as genomic DNA, mitochondrial DNA, viral DNA, synthetic DNA, cDNA reverse transcribed from RNA, bacterial DNA, viral DNA, and chromatin. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers, and can also be single-base nucleotides. In some examples, polynucleotides comprise modified nucleotides, such as methylated or glycosylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure, if present, may be imparted before or after assembly of the polymer. In some examples, the sequence of nucleotides is interrupted by non-nucleotide components. In certain examples, a polynucleotide is further modified after polymerization, such as by conjugation with a labeling component.
[0052] The term "polypeptide" or "protein" or "peptide" is specifically intended to encompass naturally occurring proteins, as well as recombinantly or synthetically produced proteins. It should be noted that the term "polypeptide" or "protein" may include naturally occurring modified forms of proteins, such as glycosylated forms. As used herein, the term "polypeptide" or "protein" or "peptide" is intended to encompass any amino acid sequence, including modified sequences such as glycoproteins.
[0053] The term "prediction" is used herein to refer to the likelihood, probability or score that a patient will respond favorably or unfavorably to a drug or set of drugs, and the extent of their response, as well as the detection of disease. The exemplary predictive methods of the present disclosure can be used clinically to make treatment decisions by selecting the most appropriate treatment modality for any particular patient. The predictive methods of the present disclosure are valuable tools for predicting whether a patient is likely to respond well to a treatment regimen, such as surgical intervention, chemotherapy with a given drug or drug combination, and / or radiation therapy.
[0054] The term "prognosis" as used herein refers to the likelihood of the clinical outcome of a subject suffering from a particular disease or disorder.With respect to cancer, prognosis is the expression of the likelihood (probability) that a subject will survive (e.g., 1, 2, 3, 4 or 5 years) and / or the likelihood (probability) that tumor will metastasize.
[0055] The term "specificity" (also called true negative rate) refers to a measure of the proportion of actual negatives that are correctly identified as such (e.g., the proportion of healthy individuals that are correctly identified as not having a pathology). Specificity is a function of the number of true negative calls (TN) and false positive calls (FP). Specificity is measured as (TN) / (TN+FP).
[0056] The term "sensitivity" (also called true positive rate, or probability of detection) refers to a measure of the proportion of actual positives that are correctly identified as such (e.g., the proportion of sick people who are correctly identified as having a disease condition). Sensitivity is a function of the number of true positive calls (TP) and false negative calls (FN). Sensitivity is measured as (TP) / (TP+FN).
[0057] The term "structural polymorphism (SV)" refers to a DNA region that is about 50 bp and is different in size from the reference genome. Examples of SV include inversions, translocations, and copy number variants (CNVs), such as insertions, deletions, and amplifications.
[0058] The term "subject" refers to a biological entity that contains genetic material. Examples of biological entities include plants, animals, or microorganisms, including, for example, bacteria, viruses, fungi, and protozoa. In some examples, the subject is a mammal, e.g., a human, and may be male or female. Such humans may be of various ages, e.g., from 1 to about 1 year old, from about 1 to about 3 years old, from about 3 to about 12 years old, from about 13 to about 19 years old, from about 20 to about 40 years old, from about 40 to about 65 years old, or over 65 years old. In various examples, the subject may be healthy or normal, abnormal, or diagnosed with a disease or suspected to be at risk for a disease. In various examples, the disease includes cancer, a disorder, a condition, a syndrome, or any combination thereof.
[0059] The term "training sample" refers to a sample whose classification may be known. The training sample may be used to train a model. The feature values of the sample may form an input vector, e.g., a training vector of the training sample. Each element of the training vector (or other input vector) may correspond to a feature that includes one or more variables. For example, the elements of the training vector may correspond to a matrix. The label values of the sample may form a vector that includes strings, numbers, byte codes, or any collection of the aforementioned data types of any size, dimension, or combination.
[0060] The terms "tumor," "neoplasia," "malignancy," or "cancer," as used herein, generally refer to the growth and proliferation of neoplastic cells, and all pre-cancerous and cancerous cells and tissues, and the results of abnormal and unregulated growth of cells, whether malignant or benign.
[0061] The term "tumor burden" refers to the amount of tumors in an individual and can be measured as the number, volume, or weight of tumors. Tumors that do not metastasize are termed "benign." Tumors that can invade surrounding tissues and / or metastasize are termed "malignant."
[0062] The term "nucleic acid sample" as used herein encompasses "nucleic acid library" or "library" including nucleic acid libraries prepared by any suitable method. The adaptor may anneal to a PCR primer to facilitate amplification by PCR, or may be a universal primer region, such as, for example, a sequencing tail adaptor. The adaptor may be a universal sequencing adaptor. As used herein, the term "efficiency" may refer to a measurable metric calculated as the ratio of the number of unique molecules for which sequences may be available after sequencing over the number of unique molecules originally present in the primary sample. In addition, the term "efficiency" may also refer to reducing the initial nucleic acid sample material required, reducing sample preparation time, reducing the amplification process, and / or reducing the overall cost of nucleic acid library preparation.
[0063] As used herein, the term "barcode" may be a known sequence used to associate a polynucleotide fragment with the input or target polynucleotide from which it is generated. The barcode sequence may be a sequence of synthetic or natural nucleotides. The barcode sequence may be contained within an adapter sequence such that the barcode sequence is contained in the sequencing read. Each barcode sequence may comprise at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more nucleotides in length. In some cases, the barcode sequences may be of sufficient length and may be sufficiently different from each other to allow for identification of samples based on the barcode sequence with which they are associated. In some cases, the barcode sequences are used to tag and subsequently identify "original" nucleic acid molecules (nucleic acid molecules present in a sample from a subject). In some cases, the barcode sequence or a combination of barcode sequences is used in combination with endogenous sequence information to identify the original nucleic acid molecule. For example, a barcode sequence (or a combination of barcode sequences) can be used in conjunction with endogenous sequence flanking the barcode (e.g., the beginning and end of the endogenous sequence), and / or the length of the endogenous sequence.
[0064] In some examples, the nucleic acid molecules used herein may be subjected to a "tagmentation" or "ligation" reaction. "Tagmentation" combines a fragmentation reaction and a ligation reaction into a single step of the library preparation process. The tagged polynucleotide fragments are "tagged" with transposon end sequences during tagmentation and may further include additional sequences added during extension during several cycles of amplification. Alternatively, the biological fragments may be "tagged" directly, and may include performing nucleic acid amplification, to process the nucleic acid molecule or a fragment thereof. For example, any type of nucleic acid amplification reaction may be used to amplify the target nucleic acid molecule or a fragment thereof to generate an amplification product. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0065] Methods and systems are provided for detecting analytes in biological samples, measuring various metrics of the analytes, and inputting the metrics as features into a machine learning model to train a classifier for medical diagnosis. The trained classifiers generated using the methods described herein are useful for multiple approaches, including disease detection and staging, identifying treatment responders, and stratifying patient populations in need thereof.
[0066] Provided herein are methods and systems that incorporate machine learning approaches to one or more biological analytes in a biological sample for various applications such as stratifying individual populations. Provided are methods and systems that detect analytes in a biological sample, measure various metrics of the analytes, and input the metrics as features into a machine learning model to train a classifier for medical diagnosis. The trained classifiers generated using the methods described herein are useful for multiple approaches, including disease detection and staging, identifying treatment responders, and stratifying patient populations in need thereof. In certain examples, the methods and systems are useful for predicting disease, treatment efficacy, and guiding treatment decisions for diseased individuals.
[0067] The present approach differs from other methods and systems in that the present method focuses on approaches to characterize the non-cellular parts of the circulating immune system, although the cellular parts may also be used. The process of hematopoietic turnover is the natural death and lysis of circulating immune cells. The plasma fraction of blood contains a fragment-enriched sample of the immune system at the time the cells die and release their intracellular contents into the circulation. Specifically, plasma provides an information-rich biological specimen sample that reflects the population of immune cells that have been educated by the presence of cancer cells before clinical symptoms appear. While other approaches are directed to characterizing the cellular parts of the immune system, the present method investigates the cancer-educated non-cellular parts of the immune system to provide biological information that is then combined with machine learning tools for useful applications. Studying non-cellular specimens in liquids such as plasma allows for the deconvolution of liquid samples to recapitulate the molecular state of the immune cells when they were alive. Studying the non-cellular parts of the immune system provides a surrogate indicator of cancer status and avoids the need for significant blood volumes to detect cancer cells and associated biological markers.
[0068] I. Circulating specimens and cell deconstruction by biological assays For health-related or biological predictions (e.g., drug resistance / susceptibility prediction) based entirely or partially on the diagnostics of body fluids, it is important to develop cost-effective and high-quality assays for each question. It is essential to be able to quickly and efficiently generate data representing the different analytes that may carry the strongest signals required to successfully train high-performance (precision) predictive models.
[0069] A. Specimen In various embodiments, the biological sample includes different specimens that provide a source of features for the models, methods, and systems described herein. The specimens can be derived from apoptosis, necrosis, and secretion from tumor, non-tumor, or immune cells. Four classes of highly informative molecular biomarkers include: 1) genomic biomarkers based on analysis of DNA profiles, sequences, or modifications; 2) transcriptomic biomarkers based on analysis of RNA expression profiles, sequences, or modifications; 3) proteomic or protein biomarkers based on analysis of protein profiles, sequences, or modifications; and 4) metabolomic biomarkers based on analysis of metabolite abundance.
[0070] 1.DNA Examples of nucleic acids include, but are not limited to, deoxyribonucleic acid (DNA), genomic DNA, plasmid DNA, complementary DNA (cDNA), cell-free (e.g., non-encapsulated) DNA (cfDNA), circulating tumor DNA (ctDNA), nucleosomal DNA, chromatosomal DNA, mitochondrial DNA (miDNA), artificial nucleic acid analogs, recombinant nucleic acids, plasmids, viral vectors, and chromatin. In one embodiment, the sample comprises cfDNA. In one embodiment, the sample comprises genomic DNA from PBMCs.
[0071] 2.RNA In various embodiments, the biological sample contains coding and non-coding transcripts, including ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (miRNA), ribosomal RNA (rRNA), circular RNA (cRNA), alternatively spliced mRNA, small nuclear RNA (snRNA), antisense RNA, short hairpin RNA (shRNA), and small interfering RNA (siRNA).
[0072] The nucleic acid molecule or fragment thereof may be single stranded or may be double stranded. A sample may contain one or more types of nucleic acid molecule or fragment thereof.
[0073] A nucleic acid molecule or a fragment thereof can contain any number of nucleotides.For example, a single-stranded nucleic acid molecule or a fragment thereof can contain at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 350, at least 400, or more nucleotides. In the case of a double-stranded nucleic acid molecule or a fragment thereof, the nucleic acid molecule or a fragment thereof may comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 350, at least 400, or more base pairs (bp), e.g., nucleotide pairs. In some cases, the double-stranded nucleic acid molecule or a fragment thereof may comprise 100-200 bp, e.g., 120-180 bp. For example, the sample may comprise cfDNA molecules comprising 120-180 bp.
[0074] 3. Polyamino acids, peptides, and proteins In various embodiments, the analyte is a polyamino acid, a peptide, a protein, or a fragment thereof. As used herein, the term polyamino acid refers to a polymer in which the monomers are amino acid residues linked together through amide bonds. When the amino acids are α-amino acids, either the L-optical isomer or the D-optical isomer can be used, with the L-isomer being preferred. In one embodiment, the analyte is an autoantibody.
[0075] In cancer patients, serum-antibody profiles change and autoantibodies against cancerous tissues are generated. These profile changes provide many possibilities for tumor-associated antigens as markers for early cancer diagnosis. The immunogenicity of tumor-associated antigens is conferred by mutated amino acid sequences that expose altered non-self epitopes. Other explanations related to this immunogenicity include alternative splicing, expression of embryonic proteins in adulthood (e.g., ectopic expression), deregulation of apoptotic or necrotic processes (e.g., overexpression), and abnormal cellular localization (e.g., secreted nuclear proteins). Examples of epitopes of tumour-restricted antigens that are encoded by intronic sequences (e.g., partially unspliced translated RNA) have been shown to make tumor-associated antigens highly immunogenic.
[0076] Exemplary markers of the present invention are suitable protein antigens that are overexpressed in tumors.Markers usually cause antibody responses in patients.Therefore, the most convenient method for detecting the presence of these markers in patients is to detect (auto)antibodies against these marker proteins in samples from patients, particularly body fluid samples such as blood, plasma or serum.
[0077] 4. Other specimens In various embodiments, the biological sample includes small chemical molecules such as, but not limited to, sugars, lipids, amino acids, fatty acids, phenolic compounds, and alkaloids.
[0078] In one embodiment, the analyte is a metabolite. In one embodiment, the analyte is a carbohydrate. In one embodiment, the analyte is a carbohydrate antigen. In one embodiment, the carbohydrate antigen is attached to an O-glycan. In one embodiment, the analyte is a monosaccharide, a disaccharide, a trisaccharide, or a tetrasaccharide. In one embodiment, the analyte is a tetrasaccharide. In one embodiment, the tetrasaccharide is CA19-9. In one embodiment, the analyte is a nucleosome. In one embodiment, the analyte is platelet rich plasma (PRP). In one embodiment, the analyte is a cellular element, such as lymphocytes (neutrophils, eosinophils, basophils, lymphocytes, PBMCs, and monocytes), or platelets.
[0079] In one embodiment, the analyte is a cellular element such as lymphocytes (neutrophils, eosinophils, basophils, lymphocytes, PBMCs and monocytes) or platelets.
[0080] In various embodiments, a combination of analytes is assayed to obtain information useful for the methods described herein. In various embodiments, the combination of analytes assayed varies depending on the type or classification of the cancer.
[0081] In various embodiments, the analyte combination is selected from: 1) cfDNA, cfRNA, polyamino acids and small chemical molecules, or 2) cfDNA and cfRNA and polyamino acids, 3) cfDNA and cfRNA and small chemical molecules, or 4) cfDNA, polyamino acids and small chemical molecules, or 5) cfRNA, polyamino acids and small chemical molecules, or 6) cfDNA and cfRNA, or 7) cfDNA and polyamino acids, or 8) cfDNA and small chemical molecules, or 9) cfRNA and polyamino acids, or 10) cfRNA and small chemical molecules, or 11) polyamino acids and small chemical molecules.
[0082] II. Sample Preparation In some embodiments, the sample is obtained, for example, from tissue or bodily fluid from a subject, or both. In various embodiments, the biological sample is a liquid sample, such as plasma, or serum, buffy coat, mucus, urine, saliva, or cerebrospinal fluid. In one embodiment, the liquid sample is a cell-free liquid. In various embodiments, the sample comprises cell-free nucleic acid (e.g., cfDNA or cfRNA).
[0083] A sample containing one or more analytes can be treated to provide or purify a specific nucleic acid molecule or fragment thereof or a collection thereof. For example, a sample containing one or more analytes can be treated to separate one type of analyte (e.g., cfDNA) from another type of analyte. In another example, a sample is divided into aliquots for analysis of a different analyte in each aliquot from the sample. In one example, a sample containing one or more nucleic acid molecules or fragments thereof of different sizes (e.g., lengths) can be treated to remove higher molecular weight and / or longer nucleic acid molecules or fragments thereof, or lower molecular weight and / or shorter nucleic acid molecules or fragments thereof.
[0084] The methods described herein may include treating or modifying the nucleic acid molecule or a fragment thereof. For example, the nucleotides of the nucleic acid molecule or a fragment thereof can be modified to include modified nucleobases, sugars, and / or linkers. The modification of the nucleic acid molecule or a fragment thereof can include oxidation, reduction, hydrolysis, tagging, barcoding, methylation, demethylation, halogenation, deamination, or any other process. The modification of the nucleic acid molecule or a fragment thereof can be achieved using enzymes, chemical reactions, physical processes, and / or exposure to energy. For example, for methylation analysis, deamination of unmethylated cytosine can be achieved through the use of bisulfite.
[0085] Sample processing may include one or more processes such as, for example, centrifugation, filtration, selective precipitation, tagging, barcoding, and partitioning. For example, cellular DNA can be separated from cfDNA by selective polyethylene glycol and bead-based precipitation processes such as centrifugation or filtration processes. The cells contained in the sample may or may not be lysed prior to separation of different types of nucleic acid molecules or fragments thereof. In one example, the sample is substantially cell-free. In one example, cellular components are assayed for measurements that may be input as features into a machine learning method or model. In various examples, cellular components such as PBMC lymphocytes may be detected (e.g., by flow cytometry, mass spectrometry, or immunopanning). A processed sample can contain, for example, at least 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 nanogram (ng), 10 ng, 50 ng, 100 ng, 500 ng, 1 microgram (μg), or more of a particular size or type of nucleic acid molecule or fragment thereof.
[0086] In some examples, blood samples are obtained from healthy individuals and individuals with cancer, e.g., individuals with stage I, II, III, or IV cancer. In one example, blood samples are obtained from healthy individuals and individuals with benign polyps, advanced adenomas (AA), and stages I-IV colorectal cancer (CRC). The systems and methods described herein are useful for detecting the presence and differentiating between stages and sizes of AA and CRC. Such differentiation is useful for stratifying individuals in a population for changes in behavior and / or treatment decisions.
[0087] A. Library Preparation and Sequencing The purified nucleic acid (e.g., cfDNA) can be used to prepare a library for sequencing. The library can be prepared using a platform-specific library preparation method or kit. The method or kit can be commercially available and can generate a sequencer-ready library. The platform-specific library preparation method can add a known sequence to the end of the nucleic acid molecule, which known sequence can be referred to as an adapter sequence. Optionally, the library preparation method can incorporate one or more molecular barcodes.
[0088] To sequence a population of double-stranded DNA fragments using a massively parallel sequencing system, the DNA fragments must be flanked by known adapter sequences. Such a collection of DNA fragments with adapters at both ends is called a sequencing library. Two examples of suitable methods for generating a sequencing library from purified DNA are (1) ligation-based attachment of known adapters to either end of the fragmented DNA, and (2) transposase-mediated insertion of adapter sequences. Any suitable massively parallel sequencing technology can be used for sequencing.
[0089] For methylation analysis, nucleic acid molecules are treated prior to sequencing. Treating nucleic acid molecules (e.g., DNA molecules) with bisulfite, enzymatic methyl-seq or hydroxymethyl-seq deaminates unmethylated cytosine bases and converts them to uracil bases. This bisulfite conversion process does not deaminate cytosines that are methylated or hydroxymethylated at the 5' position (5mC or 5hmC). When used in conjunction with sequencing analysis, the process involving bisulfite conversion of nucleic acid molecules or fragments thereof may be referred to as bisulfite sequencing (BS-seq). In some cases, nucleic acid molecules may be oxidized prior to undergoing bisulfite conversion. Oxidation of nucleic acid molecules may convert 5hmC to 5-formylcytosine and 5-carboxylcytosine, both of which are susceptible to bisulfite conversion to uracil. When used in conjunction with sequence analysis, the oxidation of a nucleic acid molecule or a fragment thereof prior to subjecting the nucleic acid molecule or a fragment thereof to bisulfite sequencing can be referred to as oxidized bisulfite sequencing (oxBS-seq).
[0090] 1. Sequencing Nucleic acids may be sequenced using sequencing methods such as next-generation sequencing, high-throughput sequencing, massively parallel sequencing, sequencing-by-synthesis, paired-end sequencing, single-molecule sequencing, nanopore sequencing, pyrosequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-seq, Digital Gene Expression, Single Molecule Sequencing by Synthesis (SMSS), Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and Sanger sequencing.
[0091] The sequencing method may include targeted sequencing, whole genome sequencing (WGS), low pass sequencing, bisulfite sequencing, whole genome bisulfite sequencing (WGBS), or a combination thereof. The sequencing method may include preparation of a suitable library. The sequencing method may include amplification of nucleic acid (e.g., by targeted or universal amplification such as PCR). The sequencing method may be performed at a desired depth, for example, at least about 5 times, at least about 10 times, at least about 15 times, at least about 20 times, at least about 25 times, at least about 30 times, at least about 35 times, at least about 40 times, at least about 45 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, at least about 100 times. Targeted sequencing methods may be performed to a desired depth, for example, at least about 500-fold, at least about 1000-fold, at least about 1500-fold, at least about 2000-fold, at least about 2500-fold, at least about 3000-fold, at least about 3500-fold, at least about 4000-fold, at least about 4500-fold, at least about 5000-fold, at least about 6000-fold, at least about 7000-fold, at least about 8000-fold, at least about 9000-fold, at least about 10000-fold.
[0092] Biological information can be generated using any useful method. Biological information can include sequencing information. Sequencing information can be generated, for example, using the assay for transposase accessible chromatin using sequencing (ATAC-seq), micrococcal nuclease sequencing (MNase-seq), deoxyribonuclease hypersensitive site sequencing (DNase-seq), or chromatin immunoprecipitation sequencing (ChIP-seq).
[0093] Sequencing reads can be obtained from a variety of sources, including, for example, whole genome sequencing, whole exome sequencing, targeted sequencing, next generation sequencing, pi sequencing, sequencing by synthesis, ion semiconductor sequencing, tag-based next generation sequencing, semiconductor sequencing, single molecule sequencing, nanopore sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (DGE), massively parallel sequencing, clonal single molecule arrays (Solexa / Illumina), sequencing using PacBio, and sequencing by oligonucleotide ligation and detection (SOLiD).
[0094] In some examples, sequencing includes modification of a nucleic acid molecule or fragment thereof, for example, by ligating a barcode, a unique molecular identifier (UMI), or another tag to the nucleic acid molecule or fragment thereof. Ligating a barcode, UMI, or tag to one end of the nucleic acid molecule or fragment thereof can facilitate analysis of the nucleic acid molecule or fragment thereof after sequencing. In some examples, the barcode is a unique barcode (i.e., a UMI). In some examples, the barcode is non-unique, and the barcode sequence can be used in conjunction with endogenous sequence information, such as the start and stop sequence of the target nucleic acid (e.g., the target nucleic acid is adjacent to the barcode, and the barcode sequence is associated with sequences at the start and end of the target nucleic acid to create a uniquely tagged molecule).
[0095] Sequencing reads may be processed using methods such as demultiplexing, de-redundancy (e.g., using unique molecular identifiers, UMIs), adapter trimming, quality filtering, GC correction, amplification bias correction, correction for batch effects, depth normalization, removal of sex chromosomes, and removal of low quality genomic bins.
[0096] In various embodiments, the sequencing reads can be aligned to a reference nucleic acid sequence. In one embodiment, the reference nucleic acid sequence is a human reference genome. For example, the human reference genome can be hg19, hg38, GrCH38, GrCH37, NA12878, or GM12878.
[0097] 2. Assay The selection of which assay to use is integrated based on the results of training the machine learning model, taking into account the clinical goal of the system. As used herein, the term "assay" includes known biological assays and may also include computational biology approaches to convert biological information into useful features as input for machine learning analysis and modeling. The assays described herein may include various pre-processing computational tools, and the term "assay" is not intended to be limiting. Various classes of samples, fractions of samples, fractions / samples of those fractions with different classes of molecules, and multiple types of assays can be used to generate feature data for use in computational methods and models to inform classifiers useful in the methods described herein. In one example, the sample is divided into aliquots for performing bioassays.
[0098] In various embodiments, biological assays are performed on different portions of the biological sample to provide a data set corresponding to the biological assays for the specimens of the portions. Various assays are known to those skilled in the art and are useful for investigating biological samples. Examples of such assays include, but are not limited to, whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), small RNA sequencing, quantitative immunoassays, enzyme-linked immunosorbent assays (ELISA), proximity extension assays (PEA), protein microarrays, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing (WO2019 / 051484), CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56 gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing, as well as cell type or cell phenotype mixture ratios derived from any of the above assays. The ability to simultaneously analyze multiple analytes (such as, but not limited to, DNA, RNA, proteins, autoantibodies, metabolites, or combinations thereof) from the same biological sample or fractions thereof can increase the sensitivity and specificity of diagnostic testing of such bodily fluids by taking advantage of the independent information between the signals.
[0099] In one example, cell-free DNA (cfDNA) content is assessed by low-coverage whole genome sequencing (lcWGS) or targeted sequencing, or whole genome bisulfite sequencing (WGBS) or whole genome enzymatic methyl sequencing, cell-free microRNA (cf-miRNA) is assessed by small RNA sequencing or PCR (digital droplet or quantitative), and circulating protein levels are measured by quantitative immunoassays. In one example, cell-free DNA (cfDNA) content is assessed by whole genome bisulfite sequencing (WGBS), proteins are measured by quantitative immunoassays (including ELISA or proximity extension assays), and autoantibodies are measured by protein microarrays.
[0100] B. cf-DNA assay using WGS In various examples, assays that profile cfDNA features are used to generate features useful in computational applications. In one example, cf-DNA features are used in machine learning models to generate classifiers that stratify individuals or detect disease, as described herein. Exemplary features include, but are not limited to, those that provide biological information about gene expression, 3D chromatin, chromatin state, copy number variants, tissue of origin, and cellular composition in cfDNA samples. Metrics of cfDNA concentration that can be used as input features for machine learning methods and models can be obtained by methods including, but not limited to, methods that quantify dsDNA within a specified size range (e.g., Agilent TapeStation, Bioanalyzer, Fragment Analyzer), methods that quantify all dsDNA using dsDNA binding dyes (e.g., QuantiFluor, PicoGreen, SYBR Green), and methods that quantify DNA fragments (either dsDNA or ssDNA) below a certain size (e.g., short fragment qPCR, long fragment qPCR, and long / short qPCR ratio).
[0101] The biological information may also include information regarding transcription start sites, transcription factor binding sites, assay for transposase accessible chromatin using sequencing (ATAC-seq) data, histone marker data, DNAse hypersensitive sites (DHS), or combinations thereof.
[0102] In one example, the sequencing information includes information regarding multiple genetic features such as, but not limited to, transcription start sites, transcription factor binding sites, chromatin gating state, and nucleosome positioning or occupancy.
[0103] 1.cfDNA plasma concentration The plasma concentration of cfDNA may be assayed as a feature indicative of the presence of cancer in various examples. In various examples, both the total amount of cfDNA in circulation and an estimate of the tumor-derived contribution to cfDNA (also referred to as the "tumor fraction") are used as prognostic biomarkers and indicators of response and resistance to treatment. Sequencing fragments aligned within annotated genomic regions are counted and normalized for sequencing depth to generate a 30,000-dimensional vector per sample, with each element corresponding to a gene count (e.g., the number of reads aligned to that gene in the reference genome). In one example, for a list of known genes with annotated regions, a sequence read count is determined for each of those annotated regions by counting the number of fragments aligned to that region. The gene read counts are normalized in various ways, for example, using a global expectation where the genome is located, within-sample normalization, and cross-feature normalization. Cross-feature normalization refers to all of those features being averaged to a specified value, for example, 0, a different negative value, 1, or a range from 0 to 2. With respect to cross-feature normalization, the total reads from a sample can be variable and therefore dependent on the preparation process and sequencing load process. Normalization can be done to a fixed number of reads as part of a global normalization.
[0104] For intra-sample normalization, it is possible to normalize by some features or screening features of some regions, especially GC bias. Thus, the composition of base pairs in each region is different and can be used for normalization. In some cases, the number of GC is significantly higher or lower than 50%, which has a thermodynamic impact since its bases are more energetic, biasing the process. Some regions give more reads than expected due to biological artifacts of sample preparation in the laboratory. Therefore, during modeling, it may be necessary to correct such biases by applying another kind of feature / feature transformation / normalization method.
[0105] In one example, the software tool ichorCNA is used to identify tumor fraction components of cfDNA via copy number changes detected by sparse (approximately 0.1× coverage) to deep (approximately 30× coverage) whole genome sequencing (WGS). In another example, measurements of tumor content by quantifying the presence of individual alleles are used to assess response or resistance to treatment in cancers where those alleles are known clonal drivers.
[0106] Copy number variation (CNV) is recognized as a major source of average human genome survival rate, and can be amplified or deleted in the region of genome that significantly contributes to phenotypic polymorphism. Tumor-derived cfDNA has genomic changes corresponding to copy number changes. Copy number changes play a role in carcinogenesis in many cancers, including CRC. Detection of copy number changes across the genome can be characterized in cfDNA to act as tumor biomarker. In one embodiment, detection uses deep WGS. In another embodiment, chromosomal instability analysis in cell-free DNA by low-coverage whole genome sequencing can be used as an assay for cfDNA. Other examples of cfDNA assays useful for detecting tumor DNA fragments include Length Mixture Model (LMM) and Fragment Endpoint Analysis.
[0107] In one example, high tumor fraction samples (>20%) are identified through manual inspection of large-scale CNVs.
[0108] In one embodiment, the change in gene expression is also reflected in plasma cfDNA concentration level, and methods such as microarray analysis can be used to assay the change in gene expression level in cfDNA samples.The cfDNA concentration metrics that can be used as input features for machine learning methods and models include, but are not limited to, Tape Station, short qPCR, long qPCR, and long / short qPCR ratio.
[0109] 2. Somatic Mutation Analysis In one example, low-coverage whole genome sequencing (lcWGS) can be used to sequence the cf-DNA in a sample and then interrogated for somatic mutations associated with a particular cancer type. Somatic mutations from lcWGS, deep WGS, or targeted sequencing (by NGS or other techniques) can be used to generate features that can be input into the machine learning methods and models described herein. Somatic mutation analysis has matured to include highly complex techniques such as microarrays and next-generation sequencing (NGS) or massively parallel sequencing. This approach may allow for extensive multiplexing capabilities in a single test. These types of hotspot panels may range in number from a few to hundreds of genes in a single assay. Other types of gene panels include whole-exon or whole-gene sequencing, offering the advantage of identifying novel mutations in specific sets of genes.
[0110] 3. Transcription Factor Profiling Inference of transcription factor binding from cfDNA has tremendous diagnostic potential in cancer. Components involved in nucleosomal signatures in transcription factor binding sites (TFBS) are assayed to evaluate and compare transcription factor binding site accessibility in different plasma samples. In one example, deep whole genome sequencing (WGS) data is used, obtained from blood samples taken from healthy donors and plasma samples from cancer patients with metastatic prostate, colon or breast cancer, where cfDNA also contains circulating tumor DNA (ctDNA). Instead of using a mixture of cfDNA signals originating from multiple cell types to establish a general tissue-specific pattern, shallow WGS data profiles individual transcription factors and performs analysis by Fourier transformation and statistical summarization. Thus, the approach provided herein provides a more nuanced view of both tissue contributions and biological processes, which allows for the identification of lineage-specific transcription factors suitable for both tissue of origin and tumor of origin analysis. In one example, the plasticity of transcription factor binding sites in cfDNA from patients with cancer is used to classify cancer subtypes, stages, and response to treatment.
[0111] In one embodiment, cfDNA fragmentation patterns are used to detect non-hematopoietic signatures. To identify transcription factor-nucleosome interactions mapped from cfDNA, hematopoietic transcription factor-nucleosome footprints are first identified in plasma samples from healthy controls. A curated list of transcription factor binding sites from publicly accessible databases (e.g., the Gene Transcription Regulation Database (GTRD)) may be used to generate a comprehensive transcription factor binding site-nucleosome occupancy map from cfDNA. Different stringency criteria are used to measure the nucleosome signatures at transcription factor binding sites, and a metric called "accessibility score" and z-score statistics are established to objectively compare significant changes in transcription factor binding site accessibility in different plasma samples. For clinical purposes, a set of lineage-specific transcription factors suitable for identifying the tissue of origin of cfDNA or the tumor of origin in patients with cancer can be identified. The accessibility score and z-score statistics are used to elucidate altered transcription factor binding site accessibility from cfDNA of patients with cancer.
[0112] In one aspect, the disclosure provides a method for diagnosing a disease in a subject, comprising: (a) providing sequence reads from deoxyribonucleic acid (DNA) extracted from the subject; (b) generating a coverage pattern of a transcription factor; (c) processing the coverage pattern to provide a signal; (d) comparing the signal to a reference signal, wherein the signal and the reference signal have different frequencies; and (e) diagnosing the disease in the subject based on the signal.
[0113] In some embodiments, (b) comprises aligning the sequence reads to a reference sequence to provide an aligned sequence pattern, selecting a region of the aligned sequence pattern that corresponds to a binding site of the transcription factor, and normalizing the aligned sequence pattern within the region.
[0114] In some examples, the transcription factor is selected from the group consisting of GRH-L2, ASH-2, HOX-B13, EVX2, PU.1, Lyl-1, Spi-B, and FOXA1.
[0115] In some embodiments, (e) comprises identifying an indication of higher accessibility of a transcription factor. In some embodiments, the transcription factor is an epithelial transcription factor. In some embodiments, the transcription factor is GRHH-L2.
[0116] 4. Estimated chromosome structure / chromatin state In other examples, the assay is used to infer the three-dimensional structure of the genome using cell-free DNA (cfDNA). In particular, the present disclosure provides methods and systems for detecting chromatin abnormalities associated with diseases or conditions, such as cancer. Without being bound to any particular mechanism, it is believed that DNA fragments are released from cells, for example, into the bloodstream. Once released from cells, the half-life of the released DNA fragments, known as cell-free DNA (cfDNA), may depend on the chromatin remodeling state. Thus, the abundance of cfDNA fragments in a biological sample can indicate the chromatin state of the gene from which the cfDNA fragment is derived (known as the "location" of the cfDNA). The chromatin state of a gene may be altered in disease. Identifying changes in the chromatin state of a gene can serve as a method for identifying the presence of a disease in a subject. The chromatin state of a gene can be predicted from the abundance and location of cfDNA fragments in a biological sample using computer-aided techniques. Chromatin state can also be useful in inferring gene expression in a sample. A non-limiting example of a computer-assisted technique that can be used to predict chromatin states is a probabilistic graphical model (PGM). PGMs can be estimated using statistical techniques such as expectation maximization or gradient methods to identify cfDNA profiles of open and closed TSSs (or between states), by fitting those parameters to a training set and statistical techniques to estimate the parameters of the PGM. The training set can be a cfDNA profile of known open and closed transcription start sites. Once trained, the PGM can predict the chromatin state of one or more genes in a naive (previously unexamined) sample. The predictions can be analyzed and quantified. By comparing predictions in the chromatin state of one or more genes from healthy and diseased samples, biomarkers or diagnostic tests can be developed. PGMs can include a variety of information, measurements, and mathematical objects that contribute to make the model more accurate.These objects can include other measured covariates, such as the biological context of the data and the laboratory process conditions of the samples.
[0117] In one example where the genetic characteristic is a chromatin state, a first array provides a measure of the constitutive openness of a plurality of cell types as a reference, a second array provides the relative proportions of the cell types in the sample, and a third array provides a measure of the chromatin state in the sample.
[0118] The expression of genes can be controlled by the access of cellular machinery to the transcription start site. The access of the transcription start site can be determined by the state of chromatin where the transcription start site is located. The state of chromatin can be controlled by chromatin remodeling, which can condense (close) or relax (open) the transcription start site. A closed transcription start site leads to a decrease in gene expression, and an open transcription start site leads to an increase in gene expression. The length of cfDNA fragments can also depend on the chromatin state. Chromatin remodeling can occur through the modification of histones and other associated proteins. Non-limiting examples of histone modifications that can control the state of chromatin and transcription start site include, for example, methylation, acetylation, phosphorylation, and ubiquitination.
[0119] Gene expression is also controlled by more distal elements such as enhancers that interact with the transcription machinery in the physical 3D space of the genome. ATAC-seq and DNAse-seq provide measurements of open chromatin that may not be clearly associated with specific genes, but correlates with the binding of these more distal elements. For example, ATAC-seq data can be obtained for a number of cell types and conditions and can be used to identify regions of the genome that have open chromatin for various basal regions such as active transcription start sites or bound enhancers or repressors.
[0120] The half-life of cfDNA, once released from a cell, may depend on the chromatin remodeling state. Thus, the abundance of cfDNA fragments in a biological sample can be indicative of the chromatin state of the gene from which the cfDNA fragment originates (referred to herein as the "location" of the cfDNA). The chromatin state of a gene may be altered in disease. Identifying changes in the chromatin state of a gene may serve as a method to identify the presence of disease in a subject. When comparing expressed and unexpressed genes, there is a quantitative shift in both the number and location distribution of cell-free DNA (cfDNA) fragments. More specifically, there is a strong depletion of reads within the approximately 1000-3000 bp region surrounding the transcription start site (TSS), and nucleosomes downstream of the TSS are strongly positioned (their location becomes much more predictable). The present disclosure provides a method to resolve the inverse correlation, and starting from cfDNA, one can infer the expression or chromatin openness of a gene. In one example, the assay is used in the multi-analyte method described herein.
[0121] The present disclosure also provides methods to generate predictions for other chromatin states as well, e.g., in repressive regions, active or poised promoters, etc. These predictions can quantify differences between different individuals (or samples), e.g., healthy individuals, colorectal cancer (CRC) patients, or samples diagnosed with other diseases or cancers.
[0122] Because the presence of open chromatin is broadly captured by the absence of nucleosomes, or also by the presence of tightly positioned nucleosomes adjacent to internal regions of open chromatin, the methods described herein can also be used for enhancers, repressors, or simply regions of open chromatin identified by other means in a reference sample.
[0123] The position of the cfDNA sequence read in the genome can be determined by "mapping" the sequence to the reference genome. Mapping can be performed using computer algorithms, including, for example, the Needleman-Wunsch algorithm, the BLAST algorithm, the Smith-Waterman algorithm, the Burrows-Wheeler alignment, the suffix tree, or custom-developed algorithms.
[0124] The three-dimensional structure of chromosomes is responsible for compartmentalizing the nucleus and linking spatially distant functional elements in close proximity. Analysis of the spatial arrangement of chromosomes and understanding how chromosomes fold provides insight into the relationship between chromatin structure, gene activity, and the biological state of the cell.
[0125] Detection of DNA interaction and modeling of three-dimensional chromatin structure can be achieved by using chromosome conformation technology.Such technology includes, for example, 3C (chromosome conformation capture), 4C (circularized chromosome conformation capture), 5C (chromosome conformation capture carbon copy), Hi-C (3C with high-throughput sequencing), ChIP-loop (3C with ChIP-seq) and ChIA-PET (Hi-C with ChIP-seq).
[0126] Hi-C sequencing is used to explore the three-dimensional structure of entire genomes by coupling proximity-based ligation with massively parallel sequencing. Hi-C sequencing utilizes high-throughput next-generation sequencing to unbiasedly quantify interactions across the genome. In Hi-C sequencing, DNA is cross-linked with formaldehyde, the cross-linked DNA is digested with a restriction enzyme to obtain 5'-overhangs, which are then filled with biotinylated residues, and the resulting blunt-ended fragments are ligated under conditions that favor ligation between the cross-linked DNA fragments. The resulting DNA sample contains the ligation products, whose junctions are labeled with biotin and consist of fragments that were in spatial proximity in the nucleus. Hi-C libraries can be generated by shearing the DNA and selecting the biotinylated products with streptavidin beads. The libraries can be analyzed by using massively parallel paired-end DNA sequencing. Using this technique, all pairwise interactions in the genome can be calculated to infer potential chromosomal structures.
[0127] In one embodiment, nucleosome occupancy of cfDNA provides an index of DNA openness and the ability to infer transcription factor binding. In certain embodiments, nucleosome occupancy correlates with tumor cell phenotype.
[0128] cfDNA represents a unique specimen generated by endogenous physiological processes to generate an in vivo map of nucleosome occupancy by whole genome sequencing. Nucleosome occupancy at transcription start sites has been exploited to infer expressed genes from cells that release their DNA into circulation. cfDNA nucleosome occupancy can reflect the footprint of transcription factors.
[0129] In various examples, cfDNA includes non-encapsulated DNA, e.g., in a blood or plasma sample, and can include ctDNA and / or cfDNA, and cfDNA can be less than 200 base pairs (bp) in length, e.g., 120-180 bp in length. The cfDNA fragmentation pattern generated by mapping cfDNA fragment ends to a reference genome can include regions of increased read depth (e.g., fragment pile-ups). These regions of increased read depth can be about 120-180 bp in size, reflecting the size of nucleosomal DNA. A nucleosome is a core of eight histone proteins wrapped by about 147 bp of DNA. A chromatosome includes nucleosomes plus histones (e.g., histone H1), and about 20 bp of associated DNA anchored outside the nucleosome. Regions of increased read depth of cfDNA can correlate with nucleosome placement. Thus, the methods of analyzing cfDNA disclosed herein can facilitate mapping of nucleosomes. The fragment pile-up seen when cfDNA reads are mapped to a reference genome may reflect nucleosome binding that protects certain regions from nuclease digestion during the process of cell death (apoptosis) or systemic clearance of circulating cfDNA by the liver and kidney.The cfDNA analysis method disclosed herein can be complemented by, for example, digestion of DNA or chromatin with MNase and subsequent sequencing (MNase sequencing).This method can reveal regions of DNA that are protected from MNase digestion, since nucleosomal histones are bound at regular intervals and the intervening regions are preferentially degraded, thus reflecting the footprint of nucleosome arrangement.
[0130] 5. Tissue of Origin Assay The plurality of nucleic acid molecules in a cfDNA sample originates from one or more cell types. In various embodiments, the assay is used to identify the tissue of origin of the nucleic acid sequences in the sample. Inferring the contribution of the cells from which the analytes in the sample originate is useful for deconstructing the analyte information in a biological sample. In various embodiments, methods such as learning regulatory regions (LRR) and immune DHS signatures are useful as methods for determining the cell type of origin and contributing cell types of the analytes in a biological sample. In various embodiments, genetic features such as V-plot measurements, FREE-C, cfDNA measurements on transcription start sites, and DNA methylation levels on cfDNA fragments are used as input features to machine learning methods and models.
[0131] In one embodiment, a first array of values corresponding to the state of the plurality of genetic features for the plurality of cell types may be created. In one embodiment, the values corresponding to the state of the plurality of genetic features are obtained for a reference population. The reference population provides values used to provide an indication of the constituent states of the plurality of genetic features.
[0132] In one example, a second array of values corresponding to a plurality of genetic features for a plurality of nucleic acid molecules of the nucleic acid sample can also be generated. The first and second arrays can then be used to generate a third array of values.
[0133] In one embodiment, the first and second arrays are matrices, and are used to generate values for the third array by matrix multiplication and parameter optimization. In one embodiment, the values for the third array correspond to the estimated proportions of multiple cell types for the multiple nucleic acid molecules of the sample. The nucleic acid data from the sample is used in combination with a reference population of information to estimate the mixture of reference populations that best matches the multiple nucleic acids of the sample. This mixture is normalized to 1 and can be used to represent the proportion or score of those reference populations in the sample.
[0134] Thus, the type and proportion of one or more cell types from which a plurality of nucleic acid molecules is derived may be determined.
[0135] In a first aspect, the present disclosure provides a method of processing a sample comprising a plurality of nucleic acid molecules, comprising: (a) providing sequencing information for a sample comprising a plurality of nucleic acid molecules, the sequencing information comprising information regarding a plurality of genetic characteristics, the plurality of nucleic acid molecules being derived from one or more cell types; (b) generating a first array of values corresponding to aspects of the plurality of genetic characteristics for a plurality of cell types, the plurality of cell types including one or more cell types; (c) generating a second array of values corresponding to aspects of the plurality of genetic characteristics for a plurality of nucleic acid molecules of the sample; (d) using the values of the first array and the values of the second array to generate a third array of values for the plurality of nucleic acid molecules of the sample corresponding to the plurality of cell types, thereby determining the type and proportion of one or more cell types from which the plurality of nucleic acid molecules are derived.
[0136] C. cfDNA Assay for Methylation Using WGBS 1. Methylation Sequencing The assay is used to sequence the entire genome (e.g., via WGBS). Enzymatic methyl sequencing ("EMseq") can provide the ultimate resolution by characterizing DNA methylation at nearly every nucleotide in the genome. Other targeted methods, such as high-throughput sequencing, pyrosequencing, Sanger sequencing, qPCR, or ddPCR, can be useful for methylation analysis. DNA methylation refers to the addition of methyl groups to DNA and is one of the most extensively characterized epigenetic modifications with important functional consequences. Typically, DNA methylation occurs at cytosine bases in nucleic acid sequences. Enzymatic methyl sequencing is particularly useful because it uses a three-step conversion and requires a smaller volume of sample for analysis.
[0137] Some embodiments of any of the foregoing aspects include subjecting the DNA or barcoded DNA to conditions sufficient to convert cytosine nucleobases of the DNA or barcoded DNA to uracil nucleobases to perform bisulfite conversion. In some embodiments, performing bisulfite conversion includes oxidizing the DNA or barcoded DNA. In some embodiments, oxidizing the DNA or barcoded DNA includes oxidizing 5-hydroxymethylcytosine to 5-formylcytosine or 5-carboxylcytosine. In some embodiments, bisulfite conversion includes reduced representation bisulfite sequencing.
[0138] In other examples, the assay used for methylation analysis is selected from mass spectrometry, methylation-specific PCR (MSP), reduced representation bisulfite sequencing (RRBS), HELP assay, GLAD-PCR assay, ChIP-on-chip assay, restriction landmark genomic scanning, methylated DNA immunoprecipitation (MeDIP), pyrosequencing of bisulfite treated DNA, molecular cleavage photoassay, methyl-sensitive Southern blotting, high-resolution melting analysis (HRM or HRMA, ancient DNA methylation reconstruction, or methylation-sensitive single nucleotide primer extension assay (msSNuPE).
[0139] In one example, the assay used for methylation analysis is whole genome bisulfite sequencing (WGBS). Modification of nucleic acid molecules or fragments thereof can be achieved using enzymes or other reactions. For example, deamination of cytosine can be achieved through the use of bisulfite. Treatment of nucleic acid molecules (e.g., DNA molecules) with bisulfite deaminates unmethylated cytosine bases and converts them to uracil bases. This bisulfite conversion process does not deaminate cytosines that are methylated or hydroxymethylated at the 5-position (5mC or 5hmC). When used in conjunction with sequencing analysis, a process involving bisulfite conversion of nucleic acid molecules or fragments thereof can be referred to as bisulfite sequencing (BS-seq). In some cases, nucleic acid molecules can be oxidized before undergoing bisulfite conversion. Oxidation of nucleic acid molecules can convert 5hmC to 5-formylcytosine and 5-carboxylcytosine, both of which are susceptible to bisulfite conversion to uracil. When used in conjunction with sequence analysis, the oxidation of a nucleic acid molecule or a fragment thereof prior to subjecting the nucleic acid molecule or a fragment thereof to bisulfite sequencing can be referred to as oxidized bisulfite sequencing (oxBS-seq).
[0140] Methylation of cytosines at CpG sites can be significantly enriched in nucleosome-spanning DNA compared to adjacent DNA. Thus, CpG methylation patterns can also be used to infer nucleosome placement using machine learning approaches. Matched nucleosome placement and 5mC datasets from the same cfDNA samples generated by micrococcal nuclease-seq (MNase-seq) and WGBS, respectively, can be used to train machine learning models. BS-seq or EM-seq datasets can be analyzed according to the same methods used in WGS to generate features for input into machine learning methods and models, regardless of methylation conversion. 5mC patterns can then be used to predict nucleosome placement, which can help infer gene expression and / or classification of disease and cancer. In another example, features may be derived from a combination of methylation status and nucleosome placement information.
[0141] Metrics used in methylation analysis include, but are not limited to, M-bias (per base methylation for CpG, CHG, CHH), conversion efficiency (100% average methylation for CHH), hypomethylated blocks, methylation levels (global average methylation for CPG, CHH, CHG, chrM, LINE1, ALU), dinucleotide coverage (normalized coverage of dinucleotides), coverage uniformity (unique CpG sites at 1x and 10x average genome coverage for (S4 run)), global average CpG coverage (depth), and average coverage at CpG islands, CGI shelves, CGI shores. These metrics can be used as feature inputs for machine learning methods and models.
[0142] In one aspect, the disclosure provides a method, comprising: (a) providing a biological sample comprising deoxyribonucleic acid (DNA) from a subject; (b) subjecting the DNA to conditions sufficient to convert unmethylated cytosine nucleobases of the DNA to uracil nucleobases, where the conditions at least partially degrade the DNA; (c) sequencing the DNA to generate sequence reads; (d) computationally processing the sequence reads to (i) determine a degree of methylation of the DNA based on the presence of uracil nucleobases, and (ii) model the at least partial DNA degradation to generate degradation parameters; and (e) determining a genetic sequence characteristic using the degradation parameters and the degree of methylation.
[0143] In another aspect, the disclosure provides a method, comprising: (a) providing a biological sample comprising deoxyribonucleic acid (DNA) from a subject; (b) subjecting the DNA to conditions sufficient to optionally enrich for methylated DNA in the sample; (c) converting unmethylated cytosine nucleobases of the DNA to uracil nucleobases; (d) sequencing the DNA to generate sequence reads; (e) computationally processing the sequence reads to (i) determine a degree of methylation of the DNA based on the presence of uracil nucleobases, and (ii) model at least partial DNA degradation to generate degradation parameters; and (f) determining a genetic sequence characteristic using the degradation parameters and the degree of methylation.
[0144] In some embodiments, (d) comprises determining the degree of DNA methylation based on a ratio of unconverted cytosine nucleobases to converted cytosine nucleobases. In some embodiments, converted cytosine nucleobases are detected as uracil nucleobases. In some embodiments, uracil nucleobases are observed as thymine nucleobases in sequence reads.
[0145] In some embodiments, generating the decomposition parameters includes using a Bayesian model.
[0146] In some embodiments, the Bayesian model is based on strand bias or bisulfite transformation or over-transformation. In some embodiments, (e) includes using the decomposition parameters under the framework of a paired HMM or a naive Bayesian model.
[0147] In certain embodiments, the methylation of specific genetic markers is assayed for use in informing the classifiers described herein. In various embodiments, the methylation of promoters such as APC, IGF2, MGMT, RASSF1A, SEPT9, NDRG4, and BMP3, or combinations thereof, is assayed. In various embodiments, the methylation of 2, 3, 4, or 5 of these markers is assayed.
[0148] 2. DMR In one example, the methylation analysis is a variably methylated region (DMR) analysis. DMRs are used to quantify CpG methylation across regions of the genome. Regions are dynamically assigned by discovery. Several samples from different classes can be analyzed and the most variably methylated regions can be identified across different classifications. A subset can be selected to be variably methylated and used for classification. The number of CpGs captured in a region can be used for analysis. Regions can tend to be of variable size. In one example, a pre-discovery process is performed that bundles several CpG sites together as regions. In one example, DMRs are used as input features for machine learning methods and models.
[0149] 3. Haplotype blocks In one embodiment, a haplotype block assay is applied to the sample. Identification of methylation haplotype blocks aids in deconvolution of heterogeneous tissue samples and mapping of tumor tissue of origin from plasma DNA. In WGBS data, tightly coupled CpG sites, known as methylation haplotype blocks (MHB), can be identified. A metric called methylation haplotype load (MHL) is used to perform tissue-specific methylation analysis at the block level. This method provides information blocks useful for deconvolution of heterogeneous samples. This method is useful for quantitative estimation of tumor burden and tissue of origin mapping in circulating cf DNA. In one embodiment, the haplotype blocks are used as input features for machine learning methods and models.
[0150] D. cfRNA Assay In various embodiments, assaying for cfRNA can be accomplished using methods such as RNA sequencing, whole transcriptome shotgun sequencing, Northern blot, in situ hybridization, hybridization arrays, linkage analysis of gene expression (SAGE), reverse transcription PCR, real-time PCR, real-time reverse transcription PCR, quantitative PCR, digital droplet PCR, or microarrays, Nanostring, FISH assays, or combinations thereof.
[0151] When using small cfRNA (including onc-RNA and miRNA) as specimens, the measurements relate to the abundance of these cfRNAs. Their transcripts are of a certain size, and each transcript is stored, and the number of cfRNAs found for each can be counted. The RNA sequences can be aligned to a reference cfRNA database, such as a set of sequences corresponding to known cfRNAs in the human transcriptome. Each cfRNA found can be used as its own feature, and multiple cfRNAs found across all samples can become a feature set. In one example, RNA fragments aligned to annotated cfRNA genomic regions are counted and normalized for sequencing depth to generate a multidimensional vector of the biological sample.
[0152] In various embodiments, all measurable cfRNA (cfRNA) is used as a feature. Some samples have a feature value of 0 for that cfRNA, where no expression is detected.
[0153] In one embodiment, all samples are collected and the read values are aggregated together. For each microRNA found in the sample, multiple aggregated reads can be found. Note that microRNAs with high expression ranks may provide better markers since they provide a more reliable signal with larger absolute changes.
[0154] In one example, cfRNA can be detected in a sample using direct detection methods such as molecular "barcoding" and microscopic imaging from the nCounter Analysis System® (nanoString, South Lake Union, WA) to detect and count up to hundreds of unique transcripts in a single hybridization reaction.
[0155] In various examples, assaying mRNA levels involves contacting a biological sample with a polynucleotide probe capable of specifically hybridizing to one or more sequences of mRNA, thereby forming a probe-target hybridization complex. Hybridization-based RNA assays include, but are not limited to, traditional "direct probe" methods such as Northern blots or in situ hybridization. These methods can be used in a wide variety of formats, including, but not limited to, substrate (e.g., membrane or glass) bound methods, or array-based approaches. In a typical in situ hybridization assay, cells are fixed to a solid support, typically a glass slide. If nucleic acid is being probed, the cells are typically denatured with heat or alkali. The cells are then contacted with a hybridization solution at a moderate temperature to allow annealing of labeled probes specific to protein-encoding nucleic acid sequences. The targets (e.g., cells) are then typically washed at a given stringency or with increasing stringency until an appropriate signal-to-noise ratio is obtained. The probes are typically labeled, for example, with radioisotopes or fluorescent reporters. Preferred probes are sufficiently long to specifically hybridize with the target nucleic acid(s) under stringent conditions. In one embodiment, the size range is from about 200 bases to about 1000 bases. In another embodiment, for small RNA molecules, shorter probes are used in the size range of about 20 bases to about 200 bases.For hybridization protocols suitable for use with the methods of the present invention, see, for example, Albertson (1984) EMBO J. 3:1227-1234, Pinkel (1988) Proc. Natl. Acad. Sci. USA 85:9138-9142; EPO Pub. No. 430, 402; Methods in Molecular Biology, Vol. 33: In situ Hybridization Protocols, Choo, ed., Humana Press, Totowa, NJ (1994), Pinkel, et al. (1998) Nature Genetics 20:207-211, and / or Kallioniemi (1992) Proc. Natl Acad Sci USA 89:5321-5325 (1992). In some applications, it is necessary to block the hybridization capacity of repetitive sequences. Thus, in some embodiments, tRNA, human genomic DNA, or Cot, I DNA is used to block non-specific hybridization.
[0156] In various examples, assaying the mRNA level comprises contacting the biological sample with a polynucleotide primer capable of specifically hybridizing to the mRNA of a single exon gene (SEG), forming a primer-template hybridization complex, and performing a PCR reaction. In some examples, the polynucleotide primer comprises about a 15-45, 20-40, or 25-35 bp sequence that is identical (in the case of a forward primer) or complementary (in the case of a reverse primer) to a sequence of a SEG listed in Table 1. As a non-limiting example, a polynucleotide primer for STMN1 (e.g., NM_203401, Homo sapiens stathmin 1 (STMN1), transcript variant 1, mRNA, 1730 bp) can include sequences identical (in the case of a forward primer) to or complementary (in the case of a reverse primer) to 1-20, 5-25, 10-30, 15-35, 20-40, 25-45, 30-50 (bp) of STMN1, as well as 1690-1710, 1695-1715, 1700-1720, 1705-1725, 1710-1730 (bp) to the end of STMN1, etc. Although not exhaustively listed herein for space reasons, all of these polynucleotide primers for STMN1 and other SEGs listed in Table 1 can be used in the systems and methods of the present disclosure. In various embodiments, the polynucleotide primers are labeled with radioisotopes or fluorescent molecules, and because the labeled primers emit a radioactive or fluorescent signal, PCR products containing the labeled primers can be detected and analyzed with various imaging instruments.
[0157] The method of "quantitative" amplification is a variety of suitable methods. For example, quantitative PCR involves using the same primers to simultaneously coamplify a known amount of a control sequence. This can be used to provide an internal standard and calibrate the PCR reaction. Detailed protocols for quantitative PCR are provided in Innis, et al. (1990) PCR Protocols, A Guide to Methods and Applications, Academic Press, Inc. NY). Measurement of DNA copy number at microsatellite loci using quantitative PCR anlysis is described in Ginzonger, et al. (2000) Cancer Research 60:5405-5409. The known nucleic acid sequence of a gene makes it sufficient to routinely select primers to amplify any part of the gene. Fluorogenic quantitative PCR can also be used in the method of the present invention. In fluorogenic quantitative PCR, quantification is based on the amount of fluorescent signal, e.g., TaqMan and SYBR green. Other suitable amplification methods include, but are not limited to, ligase chain reaction (LCR) (see Wu and Wallace (1989) Genomics 4:560, Landegren, et al. (1988) Science 241:1077, and Barringer et al. (1990) Gene 89:117), transcriptional amplification (Kwoh, et al. (1989) Proc. Natl. Acad. Sci. USA 86:1173), self-sustained sequence replication (Guatelli, et al. (1990) Proc. Nat. Acad. Sci. USA 87:1874), dot PCR, and linker-adapter PCR.
[0158] In various embodiments, the cancer-associated RNA markers are selected from miR-125b-5p, miR-155, miR-200, miR21-5pm, miR-210, miR-221, miR-222, or combinations thereof.
[0159] E. Polyamino Acid and Autoantibody Assays 1. Proteins and peptides In various embodiments, proteins are assayed using immunoassays or mass spectrometry. For example, proteins can be measured by liquid chromatography-tandem mass spectrometry (LC-MS / MS).
[0160] In various embodiments, proteins are measured by affinity reagents or immunoassays such as protein arrays, SIMOA (antibodies; Quanterix), ELISA (Abcam), O-link (DNA-conjugated antibodies; O-link Proteomics), or SOMASCAN (aptamers; SomaLogic), Luminex, and Meso Scale Discovery.
[0161] In one embodiment, the protein data is normalized by a standard curve. In various embodiments, each protein is essentially treated as its own immunoassay, each with its own standard curve, which can be calculated in various ways. The concentration relationship is typically non-linear. Samples may then be run and calculated based on the expected fluorescence concentration in the primary sample.
[0162] Several cancer-associated peptide and protein sequences are known and, in various embodiments, are useful in the systems and methods described herein.
[0163] In one embodiment, the assay comprises a combination that detects at least 2, 3, 4, 5, 6 or more markers.
[0164] In various embodiments, the cancer-associated peptide or protein marker is selected from carcinoembryonic antigens (e.g., CEA, AFP), glycoprotein or carbohydrate antigens (e.g., CA125, CA19.9, CA15-3), enzymes (e.g., PSA, ALP, NSE), hormone receptors (ER, PR), hormones (b-hCG, calcitonin), or other known biomolecules (VMA, 5HIAA).
[0165] In various embodiments, the cancer associated peptide or protein marker is 1p / 19q deletion, HIAA, ACTH, AE1, 3, ALK(D5F3), AFP, APC, ATRX, BOB-1, BCL-6, BCR-ABL1, β-hCG, BF-1, BTAA, BRAF, GCDFP-15, BRCA1, BRCA2, b72.3, c-MET, calcitonin, CALR, calretinin, CA125, CA27.29, CA19-9, CEA M, CEA P, CEA, CBFB-MYH11, CALA, c-Kit, syndical-1, CD14, CD15, CD19, CD2, CD20, CD200, CD23, CD3, CD30, CD33, CD4, CD45, CD5, CD56, CD57, CD68, CD7, CD79A, CD8, CDK4, CDK2, chromogranin A, creatine kinase isoenzyme, Cox-2, CXCL13, cyclinD, CK19, CYFRA21-1, CK20, CK5, CK6, CK7, CAM5.2, DCC, death-gamma-carboxyprothrombin, E-cadherin, EGFR T790M, EML4-ALK, ERBB2, ER, ESR1, FAP, gastrin, glucagon, HER-2 / neu, SDHB, SDHC, SDHD, HMB45, HNPCC, HVA, beta-hCG, HE4, FBXW7, IDH1 R132H, IGH-CCND1, IGHV, IMP3, LOH, MUM1 / IRF4, JAK exon 12, JAK2 V617F, Ki-67, KRAS, MCC, MDM2, MGMT, MelanA, MET, metanephrine, MSI, MPL codon 515, Muc-1, Muckiest-4, MEN2, MYC, MYCN, MPO, myf4, myoglobin, myosin, napsin A, neurofilament, NSE P, NMP22, NPM1, NRAS, Oct2, p16, p21, p53, pancreatic polypeptide, PTH, Pax-5, PAX8, PCA3, PD-L1 28-8, PIK3CA, PTEN, ERCC-1, ezrin, STK11, PLAP, PML / RARa translocation, PR, proinsulin, prolactin, PSA, PAP, PGP, RAS, ROS1, S-100, S100A2, S100B, SDHB, serotonin, SAMD4, MESOMARK, squamous cell carcinoma antigen, SS18 SYT 18q11, synaptophysin, TIA-1, TdT, thyroglobulin, TNIK, TP53, TTF-1, TNF-α, TRAFF2, urovysion, VEGF, or a combination thereof.
[0166] In one embodiment, the cancer is colorectal cancer and the CRC-associated markers are selected from APC, BRAF, DPYD, ERBB2, KRAS, NRAS, RET, TP53, UGT1A1, and combinations thereof.
[0167] In one embodiment, the cancer is lung cancer and the lung cancer associated markers are selected from ALK, BRAF, EGFR, ERBB2, KRAS, MET, NRAS, RET, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is breast cancer and the breast cancer associated markers are selected from BRCA1, BRCA2, ERBB2, TP53, and combinations thereof. In one embodiment, the cancer is gastric cancer and the gastric cancer associated markers are selected from APC, ERBB2, KRAS, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is glioma and the glioma associated markers are selected from APCAPC, BRAF, BRCA2, EGFR, ERBB2, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is melanoma and the melanoma associated markers are selected from BRAF, KIT, NRAS, and combinations thereof. In one embodiment, the cancer is ovarian cancer and the ovarian cancer associated markers are selected from BRAF, BRCA1, BRCA2, ERBB2, KRAS, TP53, and combinations thereof. In one embodiment, the cancer is thyroid cancer and the thyroid cancer associated markers are selected from BRAF, KRAS, NRAS, RET, and combinations thereof. In one embodiment, the cancer is pancreatic cancer and the pancreatic cancer associated markers are selected from APC, BRCA1, BRCA2, KRAS, TP53, and combinations thereof.
[0168] 2. Autoantibodies In another embodiment, antibodies (e.g., autoantibodies) are detected in the sample and are markers of early tumor formation. It has been demonstrated that autoantibodies are generated early in tumor formation and can be detected months or years before clinical symptoms develop. In one embodiment, plasma samples are screened using a mini-APS array (ITSI-Biosciences, Johnstown, PA, USA) using the protocol described in Somiari RI et al. (Somiari RI, et al., A low-density antigen array for detection of disease-associated autoantibodies in human plasma. Cancer Genom Proteom 13:13-19, 2016). Autoantibody markers can be used as input features in machine learning methods or models.
[0169] Assays for detecting autoantibodies include immunosorbent assays such as ELISA or PEA. When detecting an epitope containing autoantibody, preferably a marker protein, or at least a fragment thereof, it is bound to a solid support, for example a microtiter well. The autoantibody of the sample binds to this antigen or fragment. The bound autoantibody can be detected by a secondary antibody with a detectable label, for example a fluorescent label. The label is then used to generate a signal that depends on the binding to the autoantibody. The secondary antibody may be an anti-human antibody if the patient is human, or may be directed to any other organism, depending on the patient sample to be analyzed. The kit may include means for such an assay, such as a solid support, and preferably also a secondary antibody. Preferably, the secondary antibody binds to the Fc portion of the patient's (auto)antibody. Also, buffers and washing or rinsing solutions are possible. The solid support may be coated with a blocking compound to avoid non-specific binding.
[0170] In one example, autoantibodies are assayed with a protein microarray or other immunoassay.
[0171] Metrics for autoantibody assays that may be used as input features include, but are not limited to, adjusted quantile-normalized z-scores for all autoantibodies, binary 0 / 1, or absence / presence for each autoantibody based on a particular z-score cutoff.
[0172] In various embodiments, the autoantibody markers are associated with different subtypes or stages of cancer. In various embodiments, the autoantibody markers are directed or capable of binding with high affinity to tumor associated antigens. In various embodiments, the tumor associated antigen is selected from carcinoembryonic antigen / immature laminin receptor protein (OFA / iLRP), alpha fetoprotein (AFP), carcinoembryonic antigen (CEA), CA-125, MUC-1, epithelial tumor antigen (ETA), tyrosinase, melanoma associated antigen (MAGE), aberrant products of ras, aberrant products of p53, wild type of ras, wild type of p53, or fragments thereof.
[0173] In one example, ZNF700 was shown to be a capture antigen for the detection of autoantibodies in colon cancer. In a panel with other zinc finger proteins, the detection of ZNF-specific autoantibodies allowed the detection of colon cancer (O'Reilly et al., 2015). In one example, anti-p53 antibodies are assayed, as such antibodies can develop months to years before the clinical diagnosis of cancer.
[0174] F. Carbohydrates Assays exist for measuring carbohydrates in biological samples. Thin layer chromatography (TLC), gas chromatography (GC) and high performance liquid chromatography (HPLC) can be used to separate and identify carbohydrates. Carbohydrate concentrations can be determined gravimetrically (Manson and Walker), spectrophotometrically, or titrimetrically (e.g., Lane-Eynon). There are also calorimetric methods (Anthrone, Phenol-Sulfuric Acid) that analyze carbohydrates. Other physical methods to characterize carbohydrates include polarimetry, refractive index, IR, and density. In one example, metrics from carbohydrate assays are used as input features for machine learning methods and models.
[0175] III. Exemplary Systems In some examples, the present disclosure provides a system, method, or kit that may include software code executing on a data analysis, computing hardware implemented in a measurement device (e.g., laboratory equipment such as a sequencing machine). The software may be stored in memory and executed on one or more hardware processors. The software may be organized into routines or packages that can communicate with each other. Modules may include one or more devices / computers and one or more software routines / packages executing on the one or more devices / computers. For example, an analysis application or system may include at least a data reception module, a data pre-processing module, a data analysis module (which may operate on one or more types of genomic data), a data interpretation module, or a data visualization module.
[0176] The data receiving module can interface laboratory hardware or equipment with a computer system that processes laboratory data. The data preprocessing module can perform operations on the data in preparation for analysis. Examples of operations that can be applied to the data in the preprocessing module include affine transformation, noise removal operations, data cleaning, reformatting, or subsampling. The data analysis module, which can be specialized to analyze genomic data from one or more genomic materials, can, for example, use constructed genomic sequences and perform probabilistic and statistical analysis to identify abnormal patterns associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analytical methods derived from, for example, statistics, mathematics, or biology to assist in understanding the relationship between the identified abnormal patterns and a health state, functional state, prognosis, or risk. The data analysis module and / or data interpretation module can include one or more machine learning models, which can be implemented, for example, in hardware that executes software that embodies the machine learning model. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to generate visual representations of the data that can facilitate understanding or interpretation of the results. The present disclosure provides a computer system programmed to implement the methods of the present disclosure.
[0177] In some examples, the methods disclosed herein can include computational analysis of nucleic acid sequencing data of samples from an individual or multiple individuals. The analysis can identify variants inferred from sequence data and identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analysis methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix decomposition, and clustering. Non-limiting examples of variants include germline polymorphisms or somatic mutations. In some examples, the variants can refer to known variants. Known variants can be scientifically confirmed or reported in the literature. In some examples, the variants can refer to putative variants associated with biological changes. The biological changes can be known or unknown. In some examples, the putative variants can be reported in the literature but not yet biologically confirmed. Alternatively, the putative variants can be not yet reported in the literature but can be inferred based on the computational analysis disclosed herein. In some examples, a germline variant can refer to a nucleic acid that gives rise to a naturally occurring or normal polymorphism.
[0178] Natural or normal polymorphisms can include, for example, skin color, hair color, and normal body weight. In some examples, somatic mutations can refer to nucleic acids that induce acquired or abnormal polymorphisms. Acquired or abnormal polymorphisms can include, for example, cancer, obesity, conditions, symptoms, diseases, and disorders. In some examples, the analysis can include identifying germline variants. Germline variants can include, for example, private variants and somatic mutations. In some examples, the identified variants can be used by clinicians or other medical professionals to improve health care methodologies, diagnostic accuracy, and cost reduction.
[0179] FIG. 1 illustrates a system 100 that is programmed or otherwise configured to perform the methods described herein. As various examples, the system 100 can process and / or assay samples, perform sequencing analysis, measure a set of values representative of classes of molecules, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train the machine learning model (e.g., iteratively search for optimal values of parameters of the machine learning model). The system 100 includes a computer system 101 and one or more measurement devices 151, 152, or 153 that can measure various analytes. As shown, the measurement devices 151-153 measure respective analytes 1-3.
[0180] The computer system 101 can regulate various aspects of the sample processing and assays of the present disclosure, such as activating valves or pumps to move reagents or samples from one chamber to another or apply heat to the sample (e.g., during an amplification reaction), other aspects of sample processing and / or assays, performing sequencing analysis, measuring a set of values representative of a class of molecules, identifying a set of features and feature vectors from the assay data, processing the feature vectors using a machine learning model to obtain an output classification, and training the machine learning model (e.g., iteratively searching for optimal values of parameters of the machine learning model). The computer system 101 can be a user's electronic device or a computer system located remotely relative to the electronic device.
[0181] The computer system 101 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 105, which may be a single-core or multi-core processor, or multiple processors for parallel processing, a memory 110 (e.g., cache, random access memory, read-only memory, flash memory, or other memory), an electronic storage unit 115 (e.g., hard disk), a communication interface 120 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 125, such as cache, other memory, adapters for data storage and / or electronic displays. The memory 110, storage unit 115, interface 120, and peripheral devices 125 may communicate with the CPU 105 via a communication bus (solid lines), such as a motherboard. The storage unit 115 may be a data storage unit (or data repository) for storing data. Input of one or more analyte characteristics may be input from one or more measurement devices 151, 152, or 153. Exemplary analytes and measurement devices are described herein.
[0182] The computer system 101 can be operatively coupled to a computer network ("network") 130 using the communication interface 120. The network 130 can be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The network 130 is, in some cases, a telecommunications and / or data network. The network 130 can include one or more computer servers, enabling distributed computing, such as cloud computing via the network 130 ("cloud"), to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as activating valves or pumps to move reagents or samples from one chamber to another or applying heat to the sample (e.g., during an amplification reaction), other aspects of sample processing and / or assays, performing sequencing analysis, measuring a set of values representative of classes of molecules, identifying a set of features and feature vectors from the assay data, processing the feature vectors using a machine learning model to obtain an output classification, and training the machine learning model (e.g., iteratively searching for optimal values of parameters of the machine learning model). Such cloud computing may be provided, for example, by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. Network 130 may in some cases implement a peer-to-peer network with computer system 101, allowing devices coupled to computer system 101 to act as clients or servers.
[0183] The CPU 105 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 110. The instructions may be directed to the CPU 105, which may then program or otherwise configure the CPU 105 to implement the methods of the present disclosure. The CPU 105 may be part of a circuit, such as an integrated circuit. One or more other components of the system 101 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0184] The storage unit 115 can store files, such as drivers, libraries, and saved programs. The storage unit 115 can store user data, such as user preferences and user programs. The computer system 101 can optionally include one or more additional data storage units external to the computer system 101, such as those located on remote servers that communicate with the computer system 101 via an intranet or the Internet.
[0185] Computer system 101 can communicate with one or more remote computer systems through network 130. For example, computer system 101 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad®, Samsung® Galaxy Tab), a phone, a smartphone (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or a personal digital assistant. A user can access computer system 101 via network 130.
[0186] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored in electronic storage locations of computer system 101, such as, for example, on memory 110 or electronic storage unit 115. Machine executable or machine readable code may be provided in the form of software. In use, the code may be executed by CPU 105. In some cases, the code may be retrieved from storage unit 115 and stored in memory 110 for immediate access by CPU 105. In some circumstances, electronic storage unit 115 may be eliminated and machine executable instructions are stored on memory 110.
[0187] The code may be precompiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that may be selected to allow the code to be executed in a precompiled or in-place compiled manner.
[0188] Aspects of the systems and methods provided herein, such as the computer system 101, may be embodied in programming. Various aspects of the technology may be considered "products" or "articles of manufacture" typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine executable code may be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium may include any or all of the tangible memory of a computer, a processor, etc., or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that may provide non-transitory storage at any time for software programming. All or a portion of the software may be communicated from time to time over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another computer, such as a management server or host computer to the computer platform of an application server. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and through various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry the software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0189] Thus, a machine-readable medium such as a computer-executable code may take many forms, including but not limited to a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer(s), such as may be used to implement, for example, the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus in a computer system.
[0190] Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer readable media thus include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards paper tapes, any other physical storage media with patterns of holes, RAMs, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transmitting data or instructions, cables or links transmitting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0191] The computer system 101 may include or communicate with an electronic display 135 including a user interface (UI) 140 for providing, for example, the current stage of sample processing or assay (e.g., a particular step such as a lysis step, or a sequencing step being performed). Input is received by the computer system from one or more measurement devices 151, 152, or 153. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. The algorithm may, for example, process and / or assay the sample, perform a sequencing analysis, measure a set of values representative of a class of molecules, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train the machine learning model (e.g., iteratively search for optimal values of parameters of the machine learning model).
[0192] IV. Machine Learning Tools To determine the set of assays to be used in experimental testing, machine learning systems can be leveraged to evaluate the validity of a given assay or a given dataset generated from multiple assays, run on a given analyte, adding to the overall predictive accuracy of the classification. In this way, new biological / health / diagnostic questions can be addressed to design new assays.
[0193] Machine learning can be used to reduce a set of data generated from all (primary samples / specimens / tests) combinations to a set of optimal predictive features that meet, for example, a specified criterion. In various embodiments, statistical learning and / or regression analysis can be applied. Simple to complex small to large models making various modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes consideration of linear to nonlinear and non-hierarchical to hierarchical representation of features. Small to large models include consideration of the size of the basis vector space for projecting the data as well as the number of interactions between features included in the modeling process.
[0194] Machine learning techniques can be used to evaluate commercial test modalities that best fit the cost / performance / commercial range as defined by the initial question. Threshold checks can be performed. If the method applied to the holdout data set not used in cross-validation outperforms the initialized constraints, the assay is locked and put into operation. For example, assay performance thresholds can include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or combinations thereof. For example, the desired minimum precision, PPV, NPV, clinical sensitivity, clinical specificity, or a combination thereof can be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, the desired minimum AUC can be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.A subset of assays may be selected from the set of assays performed on a given sample based on the total cost of running the subset of assays according to assay performance thresholds such as desired minimum precision, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and combinations thereof. If the thresholds are not met, the assay engineering procedure can loop back to either constraint setting for possible relaxation or wet lab to change the parameters under which the data was acquired. Considering a clinical question, biological constraints, budget, lab machines, etc. may constrain the problem.
[0195] In various embodiments, the computational processing of the machine learning techniques may include method(s) of statistics, mathematics, biology, or any combination thereof. In various embodiments, any one of the computational processing methods may include dimensionality reduction methods, logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosted trees, logistic regression, matrix decomposition, network clustering, statistical testing, and neural networks.
[0196] In various embodiments, the computer processing of machine learning techniques may include logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix decomposition, multidimensional scaling (MDS), dimensionality reduction methods, t-distributed stochastic neighbor embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson product moment correlation coefficient, Spearman rank correlation coefficient, Kendall rank correlation coefficient, or any combination thereof.
[0197] In some embodiments, the computational methods are supervised machine learning methods, including, for example, regression, support vector machines, tree-based methods, and neural networks. In some embodiments, the computational methods are unsupervised machine learning methods, including, for example, clustering, networks, principal component analysis, and matrix decomposition.
[0198] For supervised learning, the training samples (e.g., thousands) may include measured data (e.g., of various analytes) and known labels, which may be determined through other time-consuming processes such as imaging of the subject and analysis by a trained practitioner. Exemplary labels may include a classification of the subject, e.g., a discrete classification of whether the subject has cancer or not, or a continuous classification that provides a probability of a discrete value (e.g., risk or score). The learning module may optimize the parameters of the model to obtain a quality metric (e.g., accuracy of prediction for known labels) on one or more specified criteria. The determination of the quality metric may be implemented for any function, including the set of all risk, loss, utility, and decision functions. The gradient may be used in conjunction with the learning step (e.g., a measure of how much the parameters of the model should be updated for a given time step of the optimization process).
[0199] As described above, the embodiments can be used for a variety of purposes. For example, plasma (or other samples) can be collected from symptomatic subjects (e.g., known to have a condition) and healthy subjects. Genetic data (e.g., cfDNA) can be obtained and analyzed to obtain a variety of different features, including features based on genome-wide analysis. These features can form a feature space, which can be searched, stretched, rotated, translated, and linearly or non-linearly transformed to generate an accurate machine learning model to distinguish between healthy subjects and subjects with a pathology (e.g., to identify a diseased or non-disease state of a subject). This data and the output derived from the model (which can include a probability of a condition, a stage (level) of a condition, or other values) can be used to generate another model that can be used to recommend further procedures (e.g., whether to recommend a biopsy or continue to monitor the subject's condition).
[0200] V. Input Feature Selection As described above, a large set of features can be generated to provide a feature space from which feature vectors can be determined. This feature vector from each of a set of training samples can then be used to train the current version of the machine learning model. The type of features used can depend on the type of specimen used.
[0201] Examples of features can include variables related to structural polymorphisms (SVs), such as copy number polymorphisms and translocations, fusions, mutations (e.g., SNPs or other single nucleotide polymorphisms (SNVs), or polymorphisms of slightly larger sequences), telomere shortening, and nucleosome occupancy and distribution. These features can be calculated genome-wide. Exemplary classes (types) of features are provided below. If genetic sequence data is obtained from at least one of the specimens, exemplary features can include aligned features (e.g., compared to one or more reference genomes) and non-aligned features. Exemplary aligned features can include sequence polymorphisms and sequence counts within a genomic window. Examples of non-aligned features include kmers from sequence reads, and organism-derived information from reads.
[0202] In some embodiments, at least one of the features is a gene sequence feature. For example, the gene sequence feature can be selected from DNA methylation status, single nucleotide polymorphism, copy number polymorphism, insertion deletion, and structural variant. In various embodiments, the methylation status can be used to determine nucleosome occupancy and / or to determine the methylation density in the CpG islands of DNA or barcoded DNA.
[0203] Ideally, feature selection can select features that are invariant or have low variability within samples with the same classification (e.g., having the same probability or associated risk of a particular phenotype), but that vary between groups of samples with different classifications. Procedures can be implemented to identify features that are most likely to be invariant within a particular population (e.g., if the classifications are real numbers, those that share a classification or lease have similar classifications). Procedures can also identify features that differ between populations. For example, read counts of sequence reads that overlap partially or completely with various genomic regions of the genome can be analyzed to determine how they vary within a population, and such read counts can be compared to read counts of a separate population (e.g., subjects known to have a disease or disorder, or subjects who are asymptomatic for a disease or disorder).
[0204] Various statistical metrics can be used to analyze the variation of features across a population with the goal of selecting features that may predict classification and therefore be advantageous for training. A further example can be the selection of a particular type of model based on the analysis of the feature space and the selected features used in the feature vector.
[0205] A. Creating a feature vector The feature vector can be created as any data structure that can be reproduced for each training sample such that corresponding data appears in the same place in the data structure across training samples. For example, the feature vector can be associated with an index where a particular value is present at each index. As explained above, a matrix can be stored at a particular index of the feature vector, and the elements of the matrix can have further sub-indexes. Other elements of the feature vector can be generated from summary statistics of such a matrix.
[0206] As another example, a single element of a feature vector may correspond to a set of sequence reads across a set of windows of a genome. Thus, an element or feature vector may itself be a vector. Such a count of reads may be of all reads or a particular group (class) of reads, e.g., reads with a particular sequence complexity or entropy. The set of sequence reads may be filtered or normalized, such as for GC bias and / or mappability bias.
[0207] In some examples, an element of a feature vector may be the result of a concatenation of multiple features. This may differ from other examples where the element is itself an array (e.g., a vector or matrix) in that the concatenated value can be treated as a single value, as opposed to a collection of values. Thus, features can be concatenated, combined, and combined to be used as manipulated features or feature representations for machine learning models.
[0208] Multiple combinations and approaches to merge features can be performed. For example, when different measures are counted over the same window (bin), the ratio between those bins (e.g., inversions divided by deletions) can be a useful feature. In addition, the ratio of bins that are spatially close and whose merger can convey biological information, such as the number of transcription start sites divided by the number of gene bodies, can also serve as a useful feature.
[0209] Features can also be manipulated, for example, by setting up a multitask unsupervised learning problem in which the joint probability of all feature vectors given a set of parameters and latent vectors is maximized. The latent vectors in this probabilistic procedure often serve as good features when trying to predict phenotypes (or other classifications) from biological sequence data.
[0210] B. Weights used for training Weights may be applied to features when they are added to a feature vector. Such weights may be based on elements in the feature vector, or on specific values within elements of the feature vector. For example, every region (window) in the genome may have a different weight. Some windows may have a weight of zero, meaning that the window does not contribute to the classification. Other windows may have a larger weight, e.g., between 0 and 1. Thus, weighting masks may be applied to the values of the features used to create the feature vector, e.g., different values of the mask applied to features for counts in a population, sequence complexity, frequency, sequence similarity, etc.
[0211] In some examples, the training process can learn the weights to be applied. In this way, any prior knowledge or biological insight into the data does not need to be known prior to the training process. The weights initially applied to the features can be considered part of the first layer of the model. Once the model is trained and one or more specified criteria are met (e.g., a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof), the model can be used in a production run to classify new samples. In such a production run, any features with initial weights of zero do not need to be calculated. Thus, the size of the feature vector may shrink from training to production. In some examples, principal component analysis (PCA) can be used to train the machine learning model. For the machine learning model, in various examples, each principal component may be a feature, or all principal components concatenated together may be a feature. Based on the output of the PCA for each of these for the analyte, a model can be created. The model can be updated based on the raw features prior to the PCA (not necessarily the PCA output). In various approaches, the raw features can use every bit of the data, can be done by taking a random selection of each batch of data, can do a random forest, or create other trees or random data sets. The features can be the measurements themselves, but can also be both, as opposed to the results of any dimensionality reduction.
[0212] C. Feature Selection Between Training Iterations As mentioned above, the training process may not produce a model that meets the desired criteria. At such times, feature selection may be performed again. Since the feature space may be quite large (e.g., 35 or 100,000), the number of possible permutations with different differential features to use in the feature vector may be huge. Certain features (potentially many) may belong to the same class (type), e.g., read counts in a window, ratios of counts from different regions, variants at different sites, etc. Furthermore, features may be concatenated into a single element to further increase the number of permutations.
[0213] The set of new features can be selected based on information from previous iterations of the training process. For example, weights associated with the features can be analyzed. These weights can be used to determine whether to keep or discard features. Features associated with weights or average weights above a threshold can be retained. Features associated with weights or average weights below a threshold (either the same or different) can be removed.
[0214] The selection of features and the creation of feature vectors for training the models can be repeated until one or more desired criteria, such as a quality metric suitable for the model (e.g., a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof) is met. Another criterion can be to select the model with the optimal quality metric among a set of models generated with different feature vectors. Thus, the model with the best statistical performance and generalizability in its ability to detect phenotypes from the data can be selected. Furthermore, the set of training samples can be used to train various models for different purposes, such as classification of condition (e.g., individuals with or without cancer), treatment (e.g., individuals with or without treatment response), prognosis (e.g., individuals with good prognosis or not with good prognosis), etc. A good cancer prognosis may correspond to when an individual has the potential for resolution or improvement of symptoms or is expected to recover after treatment (e.g., the tumor shrinks or the cancer is not expected to recur), and as used herein, a good cancer prognosis refers to a prognosis associated with a less aggressive and / or more treatable form of the disease. For example, less aggressive and more treatable forms of cancer have a higher expected survival than more aggressive and / or less treatable forms. In various examples, a good prognosis refers to a tumor that remains the same or decreases in size in response to treatment, remission, or overall survival in remission.
[0215] Similarly, poor prognosis (or individuals without a good prognosis), as used herein, refers to a prognosis associated with a more aggressive and / or less treatable form of the disease. For example, an aggressive and less treatable form has a lower survival than a less aggressive and / or less treatable form. In various examples, poor prognosis refers to the tumor size remaining the same or increasing, or the cancer recurring or not decreasing.
[0216] VI. Using Machine Learning Models for Multi-analyte Assays 2 illustrates an exemplary method 200 for analyzing a biological sample, according to one embodiment. Method 200 may be implemented by any of the systems described herein. In one embodiment, the method uses a machine learning model capable of class discrimination in a population of individuals. In various embodiments, this model (e.g., a classifier) capable of class discrimination is used to discriminate between healthy and diseased populations, between treatment responders and non-responders, and between stages of disease, providing information useful for guiding treatment decisions.
[0217] At block 210, the system receives a biological sample that includes multiple classes of molecules. Exemplary biological samples, such as blood, plasma, or urine, are described herein. Individual samples can also be received. A single sample (e.g., blood) may be collected in a set of multiple containers, such as vials.
[0218] In block 220, the system separates the biological sample into a plurality of portions, with each of the plurality of classes of molecules being in one of the plurality of portions. The sample may already be a fraction of plasma obtained from a larger sample, for example, a blood sample. The portions may then be obtained from such fractions. In some embodiments, the portions may contain a plurality of classes of molecules. An assay in a portion may test only one class of molecules, and thus a class of molecules in one portion may not be measured, but may be measured in a different portion. As an example, the measurement devices 151, 152, and 153 may perform respective assays on different portions of the sample. The computer system 101 may analyze the measured data from the various assays.
[0219] In block 230, for each of the plurality of assays, the system identifies a set of features that are input into the machine learning model. The set of features can correspond to a characteristic of one of the plurality of classes of molecules in the biological sample. The definition of the set of features to be used can be stored in the memory of the computer system. The set of features can be identified in advance, for example, using the machine learning techniques described herein. When using a particular assay, the corresponding set of features can be retrieved from memory. Each assay can have an identifier used to obtain the corresponding set of features, along with any specific software code for creating the features. Such code can be modularized so that sections can be updated independently, and the final set of features is defined based on the assays used and the stored definitions of the various sets of features.
[0220] In block 240, for each of the plurality of parts, the system performs an assay on the class of molecules in that part and obtains a set of measurements of the classes of molecules in the biological sample. The system can obtain multiple sets of measurements for the biological sample from the plurality of assays. Depending on which assay is specified (e.g., via an input file or a measurement configuration specified by the user), a particular set of measurement devices can be used to provide specific measurements to the computer system.
[0221] In block 250, the system forms a feature vector of feature quantities from the multiple sets of measurements. Each feature quantity corresponds to a feature and can include one or more measurements. The feature vector can include at least one feature quantity formed using each of the multiple sets of measurements. Thus, the feature vector can be determined using the values measured from each of the assays for different classes of molecules. The formation of the feature vector, and other details for the extraction of the feature vector, are described in other sections, but apply to all examples for the formation of the feature vector.
[0222] The characteristics of a given specimen can be determined using principal component analysis. For machine learning models, in various embodiments, each principal component may be a feature, or all principal components concatenated together may be a feature. Based on the output of PCA for each of these specimens, a model can be created. In other examples, the model can also be updated based on the raw features before any PCA, and thus the features may not necessarily include any PCA output. In various approaches, the raw features can include all bits of the data, random extraction of each batch of data for the specimen can be used, random forests can be performed, or other trees or random datasets can be created. The features may also be the measurements themselves, as opposed to the result of any dimensionality reduction (e.g., PCA), and both can also be used.
[0223] In block 260, the system loads into the computer system's memory a machine learning model trained using training vectors obtained from training biological specimens. The training specimens can have the same measurements and thus can generate the same feature vectors. The training specimens can be selected based on a desired classification, as indicated by, for example, a clinical suspicion. Different subsets can have different characteristics, as determined by, for example, the labels assigned to them. A first subset of the training biological specimens can be identified as having a specified property, and a second subset of the training biological specimens can be identified as not having the specified property. Examples of properties are various diseases or disorders, but can also be intermediate classifications or measurements. Examples of such properties include, for example, the presence or stage of cancer, or the prognosis of cancer, for the treatment of cancer. By way of example, the cancer can be colorectal cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer.
[0224] At block 270, the system inputs the feature vector into a machine learning model to obtain an output classification of whether the biological sample has a specified property. The classification may be provided in various ways, for example, as a probability of one or more classifications, respectively. For example, the presence of cancer may be assigned a probability and an output. Similarly, the absence of cancer may be assigned a probability and an output. The classification with the highest probability may be used and may be subject to one or more criteria, for example, such that one classification has a sufficiently higher probability than the second highest classification. The difference may need to exceed a threshold. If one or more criteria are not met, the output classification may be undetermined. Thus, the output classification may include a detection value (e.g., a probability) indicating the presence of cancer in the individual. And the machine learning model may further output another classification providing a probability that the biological sample does not have cancer.
[0225] Following such classification, the subject may be provided with a treatment. Examples of treatment regimens include surgical intervention, chemotherapy with a given drug or combination of drugs, and / or radiation therapy.
[0226] VII.Classifier generation The disclosed methods and systems relate to identifying a set of informative features (e.g., genetic loci) that correlate with class distinctions between samples, including sorting features (e.g., genes) by the degree to which their presence in a sample correlates with the class distinction, and determining whether the correlation is stronger than expected by chance. Machine learning techniques can implicitly use such informative features from input feature vectors. In one embodiment, the class distinction is a known class distinction, and in one embodiment, the class distinction is a disease class distinction. In particular, the disease class distinction can be a cancer class distinction. In various embodiments, the cancer is colon cancer, lung cancer, liver cancer, or pancreatic cancer.
[0227] Some embodiments of the present disclosure may also be directed to ascertaining at least one previously unknown class (e.g., disease class, proliferative disease class, cancer stage, or treatment response) into which at least one sample tested is classified, the sample being obtained from an individual. In one aspect, the present disclosure provides a classifier capable of identifying an individual within a population of individuals. The classifier may be part of a machine learning model. The machine learning model may receive as input a set of features corresponding to properties of each of a plurality of classes of molecules of a biological sample. The plurality of classes of molecules in the biological sample may be assayed to obtain a plurality of sets of measurements representative of the plurality of classes of molecules. A set of features corresponding to properties of each of the plurality of classes of molecules may be identified and input to the machine learning model. A feature vector of features from each of the plurality of sets of measurements may be generated, such that each feature corresponds to a feature of the set of features and includes one or more measurements. The feature vector may include at least one feature obtained using each set of the plurality of sets of measurements. The machine learning model including the classifier may be loaded into a computer memory. The machine learning model may be trained using training vectors obtained from the training biological samples, such that a first subset of the training biological samples are identified as having the specified property and a second subset of the training biological samples are identified as not having the specified property. The feature vectors may be input into the machine learning model to obtain an output classification of whether the biological samples have the specified property, thereby identifying a population of individuals having the specified property. In one example, the specified property is whether the individual has cancer or not.
[0228] In one aspect, the present disclosure provides a system for classifying a subject based on multi-analyte analysis of a biological sample, comprising: (a) a computer-readable medium including a classifier operable to classify the subject based on the multi-analyte analysis; and (b) one or more processors for executing instructions stored on the computer-readable medium.
[0229] In one embodiment, the system includes a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and an NMF prediction algorithm classifier.
[0230] In one example, informative features (e.g., genomic loci) of biomarkers are assayed in a cancer sample (e.g., tissue) to form a profile. The threshold of the scalar output of the linear classifier is optimized to maximize accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, e.g., the sum of sensitivity and specificity under cross-validation, as observed in the training dataset.
[0231] The overall multi-analyte assay data (e.g., expression data or sequence data) for a given sample may be normalized using methods known to those skilled in the art to correct for different amounts of starting material, various efficiencies of extraction and amplification reactions, etc. Using a linear classifier on normalized data to effectively make a diagnostic or prognostic call (e.g., responsiveness or resistance to a therapeutic agent) means dividing the data space, e.g., all possible combinations of expression values of all features (e.g., genes) in the classifier, into two disjoint parts by a separating hyperplane. This division is derived experimentally from a large set of training examples, e.g., from patients showing responsiveness or resistance to a therapeutic agent. Without loss of generality, a certain fixed set of values can be assumed for all biomarkers except one, and a threshold for this remaining biomarker can be automatically defined, with the decision varying with, e.g., responsiveness or resistance to a therapeutic agent. Expression values above this dynamic threshold can then indicate either resistance (for biomarkers with negative weights) or responsiveness (for biomarkers with positive weights) to the therapeutic agent. The exact value of this threshold depends on the actual measured expression profile of all other biomarkers in the classifier, but the general indicator of a particular biomarker remains fixed, e.g., high values or "relative overexpression" always contribute to either responsiveness (genes with positive weights) or resistance (genes with negative weights). Thus, in the context of a global gene expression classifier, relative expression can indicate whether either upregulation or downregulation of a particular biomarker is indicative of responsiveness or resistance to a therapeutic agent.
[0232] In one example, the biomarker profile (e.g., expression profile) of a patient's biological (e.g., tissue) sample is evaluated by a linear classifier. As used herein, a linear classifier refers to the weighted sum of individual biomarker features into a compound's decision score ("decision function"). The decision score is then compared to a predefined cutoff score threshold that corresponds to a particular set value in terms of accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, that indicates whether the sample is above (decision function is positive) or below (decision function is negative) the score threshold. Effectively, this means that the data space, e.g., the set of all possible combinations of biomarker features, is divided into two mutually exclusive halves that correspond to different clinical classifications or predictions, e.g., one half that corresponds to responsiveness to a therapeutic agent and the other half that corresponds to resistance.
[0233] This quantity, i.e., the interpretation of the cut-off threshold responsiveness or resistance to the therapeutic agent, is derived in the development phase ("training") from a series of patients with known outcomes. The corresponding weights for the decision scores and the responsiveness / resistance cut-off thresholds are pre-fixed from the training data by methods known to those skilled in the art. In one example, partial least squares discriminant analysis (PLS-DA) is used to determine the weights. (L. Stale, S. Wold, J. Chemom. 1 (1987) 185-196; DV Nguyen, D M Rocke, Bioinformatics 18 (2002) 39-50). Other methods for classification known to those skilled in the art may be used with the methods described herein when applied to assay data (e.g., transcripts) of a cancer classifier.
[0234] Using different methods, the quantitative assay data measured with these biomarkers can be translated into prognostic or other predictive uses. These methods include, but are not limited to, pattern recognition (Duda et al. Pattern Classification, 2nd ed., John Wiley, New York 2001), machine learning (Scholkopf et al. Learning with Kernels, MIT Press, Cambridge 2002; Bishop, Neural Networks for Pattern Recognition, Clarendon Press, Oxford 1995), statistics (Hastie et al. The Elements of Statistical Learning, Springer, New York 2001), bioinformatics (Dudoit et al., 2002, J. Am. Statist. Assoc. 97:77-87; Tibshirani et al., 2002, Proc. Natl. Acad. Sci. USA 99:6567-6572), or chemometrics (Vandeginste, et al., Handbook of Chemometrics and Examples of methods include those in the field of quantitative characterization (Qualimetrics, Part B, Elsevier, Amsterdam 1998).
[0235] In the training step, a set of patient samples for both responsive and resistant cases (e.g., including patients responsive to treatment, patients not responsive to treatment, patients resistant to treatment, and / or patients not resistant to treatment) are measured, and the prediction method is optimized using unique information from this training data to best predict the training set or future sample sets. In this training step, the method is trained or parameterized to predict from a specific assay data profile to a specific predictive call. Appropriate transformation or pre-processing steps may be performed with the measurement data before the measurement data is subjected to a classification (e.g., diagnostic or prognostic) method or algorithm.
[0236] For each of the assay data (e.g., transcripts), a weighted sum of the preprocessed feature (e.g., intensity) values is formed and compared to a threshold optimized on the training set (Duda et al. Pattern Classification, 2013). nd ed., John Wiley, New York 2001). The weights can be derived by a number of linear classification methods, including but not limited to partial least squares (PLS, (Nguyen et al., 2002, Bioinformatics 18(2002)39-50)) or support vector machines (SVM, (Scholkopf et al. Learning with Kernels, MIT Press, Cambridge 2002)).
[0237] The data may be nonlinearly transformed before applying the weighted sum as described above. This nonlinear transformation may include increasing the dimensionality of the data. The nonlinear transformation and weighted sum may also be performed implicitly, for example, through the use of a kernel function. (Scholkopf et al. Learning with Kernels, MIT Press, Cambridge 2002).
[0238] In another example, decision trees (Hastie et al., The Elements of Statistical Learning, Springer, New York 2001) or random forests (Breiman, Random Forests, Machine Learning 45:5 2001) are used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).
[0239] In another example, neural networks (Bishop, Neural Networks for Pattern Recognition, Clarendon Press, Oxford 1995) are used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).
[0240] In another example, discriminant analysis (Duda et al., Pattern Classification, 2nd ed., John Wiley, New York 2001) uses methods including linear, diagonal, quadratic, and logistic discriminant analysis to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).
[0241] In another example, predictive analytics for microarrays (PAM, (Tibshirani et al., 2002, Proc. Natl. Acad. Sci. USA 99:6567-6572)) is used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).
[0242] In another embodiment, Soft Independent Modeling of Class Analogy (SIMCA, (Wold, 1976, Pattern Recogn. 8:127-139)) is used to make predictive calls from measured intensity data of a set of transcripts or their products.
[0243] Machine learning models can be used to process various types of signals and infer classifications (e.g., phenotypes or phenotype probabilities). One type of classification corresponds to a subject's condition (e.g., disease and / or stage or severity of disease). Thus, in some embodiments, the model can classify subjects based on the type of condition for which the model was trained. Such conditions may correspond to the labels of training samples, or a collection of categorical variables. As mentioned above, these labels can be determined through more intensive measurements, or through measurements of patients with a later stage of the condition (where the condition is more easily identified).
[0244] Such models created using training samples with predefined conditions can provide certain advantages. The advantages of this technology include (a) pre-screening of diseases or disorders (e.g., age-related diseases before the onset of symptoms, or reliable detection by alternative methods), applications including, but not limited to, cancer, diabetes, Alzheimer's disease, and other diseases that may have genetic signatures, e.g., somatic genetic signatures, (b) diagnostic confirmation or supplementary evidence to existing diagnostic methods (e.g., cancer biopsy / medical imaging scan), and (c) treatment and post-treatment monitoring for prognostic reporting, treatment response, treatment resistance, and recurrence detection.
[0245] In various examples, the biological state can include a disease or disorder (e.g., an age-related disease, an aging condition, a treatment effect, a drug effect, a surgery effect, a measurable trait, or a biological state) following a lifestyle modification (e.g., a change in diet, a change in smoking, a change in sleep patterns, etc.). In some examples, the biological state may be unknown and the classification may be determined as the absence of another state. Thus, the machine learning model can infer an unknown biological state or interpret an unknown biological state.
[0246] In some examples, there may be gradations in the classification, and thus there may be many levels of classification, corresponding to conditions, e.g., real numbers. Thus, the classification may be a probability, risk, or measure for a subject having a condition or other biological state. Each such value may correspond to a different classification.
[0247] In some examples, the classification may include recommendations that may be based on previous classifications of the condition. The previous classifications may be made by separate models using the same training data (albeit potentially different input features), or by earlier sub-models that are part of a larger model that includes various classifications, and the output classification of one model may be used as input to another model. For example, if a subject is classified as being at high risk for myocardial infarction, the model may recommend lifestyle changes, such as regular exercise, eating a healthy diet, maintaining a healthy weight, quitting smoking, and lowering LDL cholesterol. As another example, the model may recommend a clinical trial for the subject to confirm the classification (e.g., a diagnostic or prognostic call). The clinical trial may include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest x-rays, positron emission tomography (PET) scans, PET-CT scans, or any combination thereof. Such recommended actions may be made as part of the methods and systems described herein.
[0248] Thus, an embodiment may provide many different models, each directed to a different type of classification. As another example, an initial model may determine whether a subject has cancer. A further model may determine whether a subject has a particular stage of a particular cancer. A further model may determine whether a subject has a particular cancer. A further model may classify a subject's predicted response to a particular surgery, chemotherapy (e.g., drug), radiation therapy, immunotherapy, or other type of treatment. As another example, an early model in a chain of sub-models may determine whether a particular genetic polymorphism is accurate or relevant, and then use that information to generate input features to a subsequent sub-model (e.g., later in the pipeline).
[0249] In some examples, phenotypic classification is derived from physiological processes, such as changes in cell turnover due to infection or physiological stress that induce changes in the types and distribution of molecules that an experimenter can observe in a patient's blood, plasma, urine, etc.
[0250] Thus, some embodiments may include active learning, where a machine learning procedure may suggest future experiments or data to obtain based on the probability of that data reducing the uncertainty in the classification. Such issues may be related to insufficient coverage of the subject's genome, lack of time point resolution, insufficient patient background sequences, or other reasons. In various embodiments, the model may suggest one of many follow-up steps based on the missing variables, including one or more of the following: (i) whole genome sequencing (WGS) resequencing, (ii) whole chromosome sequencing (WES) resequencing, (iii) targeted sequencing of specific regions of the subject's genome, (iv) specific primers or other approaches, and (v) other wet-lab approaches. Recommendations may vary between patients (e.g., due to genetic or non-genetic data of the subject). In some examples, the analysis aims to minimize some function, such as cost, risk, or morbidity to the patient, or maximize classification performance, such as accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, and suggests the best next steps to obtain the most accurate classification.
[0251] VIII. Cancer Diagnosis and Detection The machine learning methods, models, and discriminatory classifiers trained herein are useful for a variety of medical applications, including cancer detection, diagnosis, and treatment response. Once models are trained with individual metadata and specimen-derived features, applications may be tailored to stratify individuals in a population and guide treatment decisions accordingly.
[0252] A. Diagnosis The methods and systems provided herein may perform predictive analysis using an artificial intelligence based approach to analyze data acquired from a subject (patient) to generate an output of a diagnosis of the subject having cancer (e.g., colon cancer, CRC). For example, the application may apply a predictive algorithm to the acquired data to generate a diagnosis of the subject having cancer. The predictive algorithm may include an artificial intelligence based predictor, such as a machine learning based predictor configured to process the acquired data to generate a diagnosis of the subject having cancer.
[0253] The machine learning predictor can be trained using a dataset from a set of one or more cohorts of patients with cancer (e.g., a dataset generated by performing multi-analyte assays of individual biological samples) as input and the subject's known diagnostic (e.g., staging and / or tumor fraction) outcomes as output to the machine learning predictor.
[0254] A training dataset (e.g., a dataset generated by performing a multi-analyte assay on an individual's biological samples) may be generated, for example, from one or more sets of subjects with common characteristics (features) and results (labels). The training dataset may include a set of features and labels corresponding to features related to diagnosis. The features may include, for example, features such as a particular range or category of cfDNA assay measurements (e.g., counts of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap or fall into each of a set of bins (genomic windows) of a reference genome). For example, a set of features collected from a given subject at a given time point may collectively function as a diagnostic signature that may indicate a specific cancer of the subject at the given time point. The features may also include a label that indicates a diagnostic outcome of the subject, such as one or more cancers.
[0255] The label may include, for example, a result, such as a known diagnostic (e.g., stage classification and / or tumor fraction) result for the subject. The result may include a characteristic associated with the subject's cancer. For example, the characteristic may indicate that the subject has one or more cancers.
[0256] The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). The training set may be balanced across data sets corresponding to one or more sets of subjects (e.g., patients from different clinical sites or clinical trials). The machine learning predictor may be trained until certain predetermined conditions for accuracy or performance are met, e.g., until it has a minimum desired value corresponding to a diagnostic accuracy measure. For example, the diagnostic accuracy measure may correspond to a prediction of one or more cancer diagnoses, staging, or tumor fraction in a subject.
[0257] Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC) of a receiver operating characteristic (ROC) curve, which correspond to the diagnostic accuracy of detecting or predicting cancer (e.g., colon cancer).
[0258] In another aspect, the disclosure provides a method of identifying cancer in a subject, comprising: (a) providing a biological sample comprising cell-free nucleic acid (cfNA) molecules from the subject; (b) sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 distinct regions, each of the at least about 10 distinct regions comprising at least a portion of a gene selected from the group consisting of genes in Table 1; and (e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has the cancer.
[0259] In some embodiments, the at least about 10 different regions include at least about 20 different regions, and each of the at least about 20 different regions includes at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 30 different regions, and each of the at least about 30 different regions includes at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 40 different regions, and each of the at least about 40 different regions includes at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 50 different regions, and each of the at least about 50 different regions includes at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 60 different regions, and each of the at least about 60 different regions includes at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 70 different regions, and each of the at least about 70 different regions includes at least a portion of a gene selected from the group of Table 1. [Table 1]
[0260] For example, such a predetermined state can be that the sensitivity for predicting cancer (e.g., colorectal cancer, breast cancer, pancreatic cancer, or liver cancer) is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0261] As another example, such a predetermined condition may be a specificity for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0262] As another example, such a predetermined condition may be one in which the positive predictive value (PPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) includes a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0263] As another example, such a predetermined condition may be one in which the negative predictive value (NPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) includes a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0264] As another example, such a predetermined condition may be that the area under the curve (AUC) of a receiver operating characteristic (ROC) curve predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0265] In some embodiments of any of the aforementioned aspects, the method further comprises monitoring disease progression in the subject, the monitoring being based, at least in part, on the gene sequence characteristics. In some embodiments, the disease is cancer.
[0266] In some embodiments of any of the aforementioned aspects, the method further includes determining a tissue of origin of the cancer in the subject, where the determination is based, at least in part, on the gene sequence characteristics.
[0267] In some embodiments of any of the aforementioned aspects, the method further includes estimating tumor burden in the subject, the estimation being based, at least in part, on the gene sequence signature.
[0268] B. Treatment Response The predictive classifiers, systems, and methods described herein are useful for classifying populations of individuals (e.g., based on performing multi-analyte assays on biological samples of the individuals) for a number of clinical applications. Examples of such clinical applications include detecting early cancer, diagnosing cancer, classifying cancer into specific disease stages, and determining responsiveness or resistance to therapeutic agents for treating cancer.
[0269] The methods and systems described herein are applicable to various cancer types, as well as grades and stages, and are therefore not limited to a single cancer disease type. Thus, a combination of specimens and assays may be used in the systems and methods to predict the responsiveness of cancer therapeutics across different cancer types in different tissues and classify individuals based on treatment responsiveness. In one embodiment, the classifiers described herein can stratify groups of individuals into treatment responders and treatment non-responders.
[0270] The present disclosure also provides a method for determining drug targets (e.g., genes associated / significant to a particular class) for a condition or disease of interest, comprising evaluating a sample obtained from an individual for gene expression levels of at least one gene and using adjacent analysis routines to determine genes associated with the classification of the sample, thereby validating one or more drug targets associated with the classification.
[0271] The disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, comprising obtaining a sample from an individual having the disease class, subjecting the sample to the drug, evaluating the drug-exposed sample for gene expression levels of at least one gene, and classifying the drug-exposed sample into a disease class as a function of the relative gene expression levels of the sample to the gene expression levels of the model using a computer model constructed with a weighted voting scheme.
[0272] The disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, comprising subjecting an individual to the drug, obtaining a sample from the individual subjected to the drug, evaluating the sample for gene expression levels of at least one gene, and classifying the sample into a disease class using a model constructed with a weighted voting scheme, and evaluating the gene expression levels of the sample compared to the gene expression levels of the model.
[0273] Yet another application is a method for determining whether an individual belongs to a phenotypic class (e.g., intelligence, responsiveness to treatment, life span, likelihood of viral infection or obesity) comprising obtaining a sample from the individual, assessing the sample for gene expression levels of at least one gene, and classifying the sample into a disease class using a model constructed with a weighted voting scheme, comprising assessing the gene expression levels of the sample compared to the gene expression levels of the model.
[0274] There is a need to identify biomarkers that are useful for predicting the prognosis of patients with colon cancer. The ability to classify patients as high risk (poor prognosis) or low risk (good prognosis) may allow for the selection of appropriate therapy for these patients. For example, high risk patients are more likely to benefit from aggressive treatment, whereas treatment may not have significant benefits for low risk patients. However, despite this need, there has been no solution to this problem.
[0275] Predictive biomarkers that can guide treatment decisions have been sought to identify subsets of patients who may be "exceptional responders" to particular cancer therapies, or individuals who may benefit from alternative treatment modalities.
[0276] In one aspect, the system and method described herein related to classifying populations based on therapeutic response refers to cancers treated with, but not limited to, the following classes of chemotherapeutic agents: DNA damage agents, DNA repair targeting therapy, inhibitors of DNA damage signal transduction, inhibitors of DNA damage-induced cell cycle arrest, and inhibitors of processes that indirectly lead to DNA damage.Each of these chemotherapeutic agents, as the term is used herein, is considered to be "DNA damage therapeutic agents".
[0277] Patient sample data can be classified into high-risk and low-risk patient groups, such as patients with high or low risk of clinical recurrence, and the results can be used to determine the course of treatment. For example, patients determined to be high-risk patients can be treated with adjuvant chemotherapy after surgery. For patients considered to be low-risk patients, adjuvant chemotherapy can be withheld after surgery. Thus, in certain aspects, the present disclosure provides a method for generating gene expression profiles of colon cancer tumors that indicate risk of recurrence.
[0278] In various examples, the classifiers described herein can stratify populations of individuals between responders and non-responders to a treatment.
[0279] In various embodiments, the treatment is selected from an alkylating agent, a plant alkaloid, an antitumor antibiotic, an antimetabolite, a topoisomerase inhibitor, a retinoid, a checkpoint inhibitor therapy, or a VEGF inhibitor.
[0280] Examples of treatments that may stratify populations into responders and non-responders include sorafenib, regorafenib, imatinib, eribulin, gemcitabine, capecitabine, pazopanib, lapatinib, dabrafenib, sutinib malate, crizotinib, everolimus, trisirolimus, sirolimus, axitinib, gefitinib, anastrol, bicalutamide, fulvestrant, ralitrexed, pemetrexed, goserin acetate, erlotinib, vemurafenib, Biciophenib, tamoxifen citrate, paclitaxel, docetaxel, cabazitaxel, oxaliplatin, ziv-aflibercept, bevacizumab, trastuzumab, pertuzumab, panziumumab, taxanes, bleomycin, melphalen, plumbagin, camptosar, mitomycin-C, mitoxantrone, sumanx, doxorubicin, pegylated doxorubicin, forfoli, 5-fluorouracil, temozolomide, pasireotide , tegafur, gimeracil, oterasin, itraconazole, bortezomib, lenalidomide, irintotecan, epirubicin, and romidepsin, resminostat, tasquinimod, refametinib, lapatinib, tyverb, arenezil, pasireotide, signifor, ticilimumab, tremelimumab, lansoprazole, PrevOnco, ABT-869, linifanib, borolanib, tivantinib, tarceva, erlotinib, stivarga, lego Chemotherapeutic agents including rafenib, fluorosorafenib, brivanib, liposomal doxorubicin, lenvatinib, ramucirumab, peretinoin, Ruchiko, mupafostat, TS-1, tegafur, gimeracil, oteracil, and orantinib; and antibody therapies including alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, or rituximab.
[0281] In other examples, populations can be stratified into responders and non-responders to checkpoint inhibitor therapy, such as compounds that bind to PD-1 or CTLA4.
[0282] In other examples, populations can be stratified into responders and non-responders to anti-VEGF therapies that bind to targets in the VEGF pathway.
[0283] IX. Efficacy In some examples, the biological state may include a disease. In some examples, the biological state may be a stage of a disease. In some examples, the biological state may be a gradual change in a biological state. In some examples, the biological state may be a treatment effect. In some examples, the biological state may be a drug effect. In some examples, the biological state may be a surgical effect. In some examples, the biological state may be a biological state following a lifestyle modification. Non-limiting examples of lifestyle modifications include a change in diet, a change in smoking, and a change in sleep patterns.
[0284] In some examples, the biological state is unknown. The analyses described herein can include machine learning to infer the unknown biological state or to interpret the unknown biological state.
[0285] In one embodiment, the system and method are particularly useful for applications related to colon cancer. Cancer forms in the tissues of the colon (the longest part of the large intestine). Most colon cancers are adenocarcinomas (cancers that begin in the cells that line the internal organs and have gland-like characteristics). The progression of cancer is characterized by the stage or extent of the cancer in the body. Staging is usually based on the size of the tumor, whether lymph nodes contain cancer, and whether the cancer has spread from the original site to other parts of the body. Stages of colon cancer include stage I, stage II, stage III, and stage IV. Unless otherwise stated, the term colon cancer refers to colon cancer at stage 0, stage I, stage II (including stage IIA or IIB), stage III (including stage IIIA, IIIB, or IIIC), or stage IV. In some embodiments herein, colon cancer is from any stage. In one embodiment, the colon cancer is stage I colon cancer. In one embodiment, the colon cancer is stage II colon cancer. In one embodiment, the colon cancer is stage III colon cancer.In one embodiment, the colon cancer is stage IV colon cancer.
[0286] Conditions that may be predicted by the methods of the present disclosure include, for example, cancer, gut-related diseases, immune-mediated inflammatory diseases, neurological diseases, renal diseases, prenatal diseases, and metabolic diseases.
[0287] In some examples, the methods of the present disclosure can be used to diagnose cancer.
[0288] Non-limiting examples of cancer include adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.
[0289] Non-limiting examples of cancers that may be predicted by the disclosed methods and systems include acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brain stem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pine fruit tumor, breast cancer, bronchial tumor, Burkitt's lymphoma, non-Hodgkin's lymphoma, carcinoid tumor, and cervical cancer. , chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing's sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrotic Histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cavity cancer, lung cancer, non-small cell cancer, small cell cancer, melanoma, oral cancer cancer), myelodysplastic syndromes, multiple myeloma, medulloblastoma, nasal cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumors, plasma cell neoplasms, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sezary syndrome, skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, and Wilms' tumor.
[0290] Non-limiting examples of bowel-related diseases that may be predicted by the disclosed methods and systems include Crohn's disease, colitis, ulcerative colitis (UC), inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), and celiac disease. In some examples, the disease is inflammatory bowel disease, colitis, ulcerative colitis, Crohn's disease, microscopic colitis, collagenous colitis, lymphocytic colitis, diversion colitis, Behcet's disease, and undifferentiated colitis.
[0291] Non-limiting examples of immune-mediated inflammatory diseases that can be inferred by the disclosed methods and systems include psoriasis, sarcoidosis, rheumatoid arthritis, asthma, rhinitis (hay fever), food allergies, eczema, lupus, multiple sclerosis, fibromyalgia, type 1 diabetes, and Lyme disease. Non-limiting examples of neurological diseases that can be inferred by the disclosed methods and systems include Parkinson's disease, Huntington's disease, multiple sclerosis, Alzheimer's disease, stroke, epilepsy, neurodegeneration, and neuropathy. Non-limiting examples of renal diseases that can be inferred by the disclosed methods and systems include interstitial nephritis, acute renal failure, and nephropathy. Non-limiting examples of prenatal diseases that can be inferred by the disclosed methods and systems include Down's syndrome, aneuploidy, spina bifida, trisomy, Edwards syndrome, teratoma, sacrococcygeal teratoma (SCT), ventriculomegaly, renal agenesis, cystic fibrosis, and fetal hydrops. Non-limiting examples of metabolic diseases that can be inferred by the disclosed methods and systems include cystinosis, Fabry disease, Gaucher disease, Lesch-Nyhan syndrome, Niemann-Pick disease, phenylketonuria, Pompe disease, Tay-Sachs disease, von Gierke disease, obesity, diabetes, and heart disease.
[0292] The specific details of the particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the disclosed embodiments of the invention. However, other embodiments of the invention may be directed to the particular embodiments in relation to each individual aspect, or to particular combinations of these individual aspects. All patents, patent applications, publications, and descriptions referred to herein are incorporated by reference in their entirety for all purposes.
[0293] X. Example The above description and the following examples of the present invention have been presented for purposes of illustration and description, and are not intended to be exhaustive or to limit the invention to the precise form described, many modifications and variations are possible in light of the above teaching.
[0294] A. Example 1: Preparation of Multi-Analyte Assays of Biological Samples This example provides a multi-analyte approach to exploit the independent information between signals. A process diagram is described below for different components of the system for assaying with corresponding machine learning models to perform accurate classification. The selection of which assay to use can be integrated based on the results of training the machine learning models, taking into account the clinical goals of the system. Various classes of samples, fractions of samples, parts of those fractions / samples with different classes of molecules, and types of assays can be used.
[0295] 1. System diagram 3 shows an overall framework 300 of the disclosed systems and methods. The framework 300 can use measurements of samples (wet lab 320) and other data about a subject, in combination with machine learning, to identify a set of assays and features for classifying the subject, e.g., diagnostic or prognostic. In this example, the steps of the process can be as follows:
[0296] In block 311 of stage 310, a question having clinical, scientific, and / or commercial relevance is asked, e.g., early colorectal cancer detection for viable follow-up. In block 312, subjects (new or previously tested) are identified. The subjects may have known classifications (labels) for later use in machine learning. Thus, different cohorts may be identified. In block 313, the analysis may select the type of sample to be mined (i.e., the sample may not remain in the final assay) and determine the collection of biomolecules in each of the samples (e.g., blood) that can generate a sufficient signal to assess the presence or absence of a pathology / disorder (e.g., early colorectal cancer malignancy). Constraints may be imposed on the assay / model, e.g., related to accuracy. Exemplary constraints include the minimum sensitivity of the assay, the minimum specificity of the assay, the maximum cost of the assay, the time available to develop the assay, the available biological materials and expected accumulation rates, the available set of previously developed processes that determine the maximum set of experiments that can be performed on those biological materials, and the available hardware that limits the number of processes that can be performed on those biological materials to obtain data.
[0297] Patient cohorts can be designed and sampled to accurately represent the different classifications needed to adequately achieve the clinical goal (healthy, colorectal and other, advanced adenoma, colorectal cancer (CRC)). The patient cohort can be selected and the selected cohort can be viewed as a constraint on the system. An example cohort is 100 CRC, 200 advanced adenoma, 200 non-advanced adenoma, and 200 healthy subjects. The selected cohort can correspond to the population intended for use in the final assay and the cohort can specify the number of samples for which the performance of the assay is calculated.
[0298] Once a cohort is selected, samples may be collected to meet the cohort design. Various samples may be collected, such as blood, cerebrospinal fluid (CSF), and other samples as mentioned herein. Such analysis may be performed in block 313 of FIG. 3.
[0299] At stage 320, wet lab experiments can be performed for the first set of assays. For example, an unconstrained test set can be selected (primary sample / analyte / test combinations). Protocols and modalities for isolating analytes from primary samples can be performed. Protocols and modalities for test execution can be generated. Performance of wet lab activities can be performed using hardware devices including sequencers, fluorescence detectors, and centrifuges.
[0300] In block 321, the sample is divided into smaller components (also called fractions or portions) by, for example, centrifugation. As an example, blood is divided into fractions of plasma, buffy coat (white blood cells and platelets), serum, red blood cells, and extracellular vesicles such as exosomes. The fractions (e.g., plasma) can be divided into aliquots to assay different analytes. For example, different aliquots are used to extract cfDNA and cfRNA. Thus, analytes can be isolated from the fractions or aliquots to enable multi-analyte assays. A fraction (e.g., some plasma) can be retained to measure protein concentration.
[0301] In block 323, experimental procedures are performed to measure the characteristics and amounts of the above molecules in each fraction, such as (1) the sequence and assigned location along the genome of the cell-free DNA fragments found in the plasma, (2) the methylation pattern of the cfDNA fragments found in the plasma, (3) the amount and type of microRNAs found in the plasma, and (4) the concentrations of proteins known to be associated with CRC from the literature (CRP, CEA, FAP, FRIL, etc.).
[0302] The QC of each of the samples processed on any given pipeline can be verified. cfDNA QCs include insert size distribution, relative representation of GC bias, barcode sequence of spike-ins (introduced for sample traceability), etc. Examples of methylation QCs include bisulfite conversion efficiency of control DNA, insert size distribution, average sequencing depth, % overlap, etc. Examples of miRNA QCs include insert size distribution, relative representation of normalized spike-ins, etc. Examples of protein QCs include linearity of standard curve, control sample concentration, etc.
[0303] The samples are then processed to obtain data for all patients in the cohort. The raw data is indexed by patient metadata. Data can be obtained from other sources and stored in the database. Data can be curated from relevant open databases such as GTEX, TCGA, and ENCODE. This includes ChIP-seq, RNA-seq, and eQTL.
[0304] At stage 340, data can be obtained from other sources, e.g., wearables, images, etc. Such other data corresponds to data determined external to the biological specimen. Such measurements can be heart rate, activity measurements, or other such data available from a wearable device. Imaging data can provide information such as organ size and location and identify unknown masses.
[0305] The database 330 can store data. The data can be curated from relevant open databases such as GTEX, TCGA, and ENCODE. This includes ChIP-seq, RNA-seq, and eQTL. Each subject record can include fields with the measured data and the subject's label, such as whether a condition is present, the severity (stage) of the condition, etc. A subject may have multiple labels.
[0306] At block 350, dry lab operations can occur. The "dry lab" operations can begin with a query to a database to generate a matrix of relevant data and metadata values to perform a predictive task. Features are generated by processing the incoming data and possibly selecting a subset of relevant inputs.
[0307] In block 351, machine learning can be used to reduce the entire set of data generated from all (primary sample / specimen / test) combinations to the most predictive set of features in block 352. Accuracy metrics of different feature sets can be compared to each other to determine the most predictive set of features. In some embodiments, a set of features / models that meet an accuracy threshold can be identified and then other constraints (e.g., cost and number of tests) can be used to select the optimal model / feature grouping.
[0308] A variety of different functions and models can be tested. Simple to complex, small to large models making various modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes consideration of linear to nonlinear and non-hierarchical to hierarchical representations of features. Small to large models include consideration of the size of the basis vector space onto which the data is projected, as well as the number of interactions between features included in the modeling process.
[0309] Machine learning techniques can be used to evaluate the commercial test modality that is optimal for cost / performance / commercial scope as defined in the initial question. A threshold check can be performed. If the method applied to the holdout data set not used in cross-validation exceeds the initialized constraints, the assay is locked and put into operation. Thus, the assay can be output at block 360.
[0310] If the thresholds are not met, the assay engineering procedure loops back to either the constraint settings for possible relaxation, or to the wet lab to modify the parameters under which the data were acquired.
[0311] Considering the clinical question, biological constraints, budget, lab machines, etc. may constrain the problem. Cohort design can then be based on clinical samples, statistical and informative actionable items, and sample accrual rates, which are in fact based on performance or prior knowledge.
[0312] 2. Hierarchy of samples and their parts In one example, multiple samples are taken from patients in a cohort and analyzed for multiple molecular types via multiple assays. The results of the assays are then analyzed by an ML model, and after selection of significant features and samples, assay results related to clinically, scientifically, or commercially important questions are output.
[0313] FIG. 4 shows a hierarchical overview of the multi-analyte approach used for an exemplary "liquid biopsy". In stage 401, the different samples are collected. As shown, blood, CSF, and saliva are collected. In stage 402, the sample can be divided into fractions, e.g., blood is shown divided into plasma, platelets, and exosomes. In stage 403, each of the fractions can be analyzed to measure one or more classes of molecules, e.g., DNA, RNA, and / or proteins. In stage 404, each class of molecules can be subjected to one or more assays. For example, methylation and whole genome assays can be applied to DNA. For RNA, assays detecting mRNA or single-stranded RNA can be applied. For proteins, enzyme-linked immunosorbent assays (ELISAs) can be used.
[0314] In this example, collected plasma was analyzed using multi-analyte assays including: low-coverage whole genome sequencing, CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56-gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing. Whole blood can be collected in K3-EDTA tubes and centrifuged twice to separate plasma. Plasma can be divided into aliquots for cfDNA, lcWGS, WGS, WGBS, cf-miRNA sequencing, and quantitative immunoassays (either enzyme-linked immunosorbent assay [ELISA] or single molecule array [SIMOA]).
[0315] In stage 405, a learning module running on computer hardware can receive measured data from various assays of various fraction(s) of various sample(s). The learning module can provide metrics for various groupings of models / features. For example, various sets of features can be identified for each of the multiple models. Different models can use different techniques such as neural networks or decision trees. Stage 406 can select the model / feature group to use, or potentially provide instructions (commands) to make further measurements. Stage 407 can specify the samples, fractions, and individual assays to be used as part of the total assay used to measure and classify new samples.
[0316] 3. Iterative flow between modules FIG. 5 illustrates an iterative process for designing an assay and corresponding machine learning model according to an embodiment of the present invention. Wet lab components are shown on the left and computational components on the right. Omitted modules include external data, prior structures, clinical metadata, etc. These meta components may flow into both wet lab and dry lab (computational) components. In general, the iterative process may include various phases including an initialization phase, an exploration phase, a refinement phase, and a validation / confirmation phase. The initialization phase may include blocks 502-508. The exploration phase may include a first flow path through blocks 512-528. The refinement phase may include an additional flow path through blocks 512-528, as well as blocks 530 and 532. The validation / confirmation phase may occur using blocks 524 and 529. Various blocks may be optional or hard-coded to provide a specified outcome, for example, in a particular model, module 518 may always be selected.
[0317] At block 502, a clinical question is received, for example to screen for the presence of colorectal cancer (CRC). Such clinical question may also include the number of classifications required. For example, the number of classifications may correspond to different stages of the cancer.
[0318] In block 504, a cohort(s) is designed. For example, the number of cohorts is equal to the number of classifications, and subjects within a cohort can have the same label. Additional cohorts can be added at later stages or stages of the process.
[0319] In one embodiment, before any biochemical tests are performed, there is an initial selection of sample and / or test. For example, whole genome sequencing can be selected to obtain information about an initial sample, such as blood. Such initial sample and initial assay can be selected based on the clinical question, for example, based on the relevant organ.
[0320] At block 506, an initial sample is obtained. The sample can be of various types, e.g., blood, urine, saliva, cerebrospinal fluid, etc. As part of obtaining the initial sample, the sample can be divided into fractions as described herein (e.g., dividing blood into plasma, buffy coat, exosomes, etc.), and those fractions can be further divided into portions having particular classes of molecules.
[0321] In block 508, one or more initial assays are performed. The initial assays can be operated for individual classes of molecules. Some or all of the initial set of assays can be used as defaults across various clinical questions. The initial data 510 can be sent to a computer 511 to evaluate the data, determine machine learning models, and suggest further assays to potentially be performed. The computer 511 can perform the operations described in this section and other sections of this disclosure.
[0322] The data filter module 512 can filter the initial data 510 to provide one or more sets of filtered data. Such filtering can simply obtain data from different assays, but can be more complex, such as performing statistical analysis to provide measurements from the raw data, where the initial data 510 is considered raw data. The filtering can include dimensionality reduction, such as principal component analysis (PCA), non-negative matrix factorization (NMF), kernel PCA, graph-based kernel PCA, linear discriminant analysis (LDA), generalized discriminant analysis (GDA), or autoencoder. Multiple sets of filtered data can be determined from raw data of a single assay. Different sets of filtered data can be used to determine different sets of features. In some embodiments, the data filter module 512 can take into account processing performed by downstream modules. For example, the type of machine learning model can affect the type of dimensionality reduction used.
[0323] The feature extraction module 514 can extract features, for example, using genetic data, non-genetic data, filtered data, and reference sequences. Feature extraction may also be referred to as feature engineering. The features of the data obtained from an assay will correspond to the characteristics of the classes of molecules obtained in that assay. By way of example, the features (and their corresponding features) may be measurements output from filtering, only a portion of such measurements, further statistical results of such measurements, or measurements that are added together. Certain features are extracted with the goal that some features have different values between different groups of subjects (e.g., different values between subjects with and without a pathology), thereby allowing discrimination between different groups, or inference of the degree of a property, condition, or trait. Examples of features are provided in Section V.
[0324] The cost / loss selection module 516 can select a particular cost function (also referred to as a loss function) to optimize in training the machine learning model. The cost function can have various terms to define the accuracy of the current model. At this point, algorithmically, other constraints can be injected. For example, the cost function can measure the number of misclassifications (e.g., false positives and false negatives) and have a scaling factor for each of the different types of misclassifications, thereby providing a score that can be compared to a threshold to determine whether the current model is satisfactory. Such accuracy tests can also implicitly determine whether a set of features and a set of assays can provide a satisfactory model, and if not, a different set of features can be selected.
[0325] In some instances, the distribution of the data may influence the selection of a loss function for an unmonitored task, for example to have technical control of the system, in which case the loss function may correspond to a distribution that is consistent with the input data.
[0326] The model selection module 518 can select the model(s) to use. Examples of such models include logistic regression, support vector machines with different kernels (e.g., linear or nonlinear kernels), neural networks (e.g., multi-layer perceptrons), and various types of decision trees (e.g., random forests, gradient trees, or gradient boosting techniques). Multiple models can be used, e.g., models can be used sequentially (e.g., the output of one model goes into the input of another model) or in parallel (e.g., using voting to determine the final classification). When multiple models are selected, these can be referred to as sub-models.
[0327] Cost functions are distinct from models, and distinct from features. These different parts of the architecture can have significant effects on each other, but they are also defined by other components of the test design and their corresponding constraints. For example, a cost function can be defined by its components, including feature distributions, feature numerical values, diversity of label distributions, label types, label complexity, risk associated with different error types, etc. Changes to certain features can result in changes to the model and / or cost function, and vice versa.
[0328] The feature selection module 520 can select a set of features to be used for the current iteration in training the machine learning model. In various embodiments, all features extracted by the feature extraction module 514 can be used, or only a portion of the features can be used. Feature weights for the selected features can be determined and used as inputs for training. As part of the selection, some or all of the extracted features can undergo transformation. For example, weights can be applied to certain features based, for example, on the expected importance (probability) of the particular feature(s) relative to the other feature(s). Other examples include dimensionality reduction (e.g., of a matrix), distribution analysis, normalization or regularization, matrix decomposition (e.g., kernel-based discriminant analysis and non-negative matrix factorization), which can provide a low-dimensional manifold corresponding to a matrix. Another example is transforming raw data or features from one type of instrument to that of another type of instrument, for example, when different samples are measured using different instruments.
[0329] The training module 522 can perform parameter optimization of the machine learning model, which may include sub-models. Various optimization techniques can be used, such as gradient descent or the use of second derivatives (Hessian). In other embodiments, training can be implemented in a manner that does not require Hessian or gradient calculations, such as dynamic programming or evolutionary algorithms.
[0330] The evaluation module 524 can determine whether the current model (e.g., as defined by a set of parameters) satisfies one or more criteria included in the output constraint(s). For example, the quality metric can measure the predictive accuracy of the model for a training set and / or a validation set of samples for which the labels are known. Such accuracy metrics can include sensitivity and specificity. The quality metric can be determined using values other than accuracy, e.g., the expected cost of a number of assays, and the time to perform the assay measurements. If the constraints are satisfied, a final assay 529 can be provided. The final assay 529 can include a specific order for performing the assays on the test samples, e.g., when an assay not in the default list is selected.
[0331] If the output constraints are not met, various items can be updated. For example, the set of selected features can be updated, or the set of selected models can be updated. Some or all upstream modules can be evaluated, checked, and alternatives suggested. Thus, feedback can be provided anywhere in the upstream pipeline. When the evaluation module 524 determines that the space of features and models has been sufficiently searched (e.g., exhausted) without satisfying the constraints, the process can decide to flow to further modules to obtain new assays and / or sample types. Such decisions can be defined by constraints. For example, a user can only do so many assays (and associated time and cost), have so many samples, or do an iterative loop (or several loops) many times. These constraints can contribute to stopping the test design of the current set of features, models, and assays in exchange for exceeding a minimum metric.
[0332] The assay identification module 526 can identify new assays to perform. If a particular assay is determined to be unimportant, the data may be discarded. The assay identification module 526 can receive certain input constraints that can be used to determine which assay or assays to select, for example, based on the cost or timing of performing the assay.
[0333] The sample identification module 528 can determine the new sample type (or portion thereof) to use. The selection can depend on which new assay(s) are to be performed. Input constraints can also be provided to the sample identification module 528.
[0334] The assay identification module 526 and sample identification module 528 can be used if the assessment is that the assay and model do not meet the output constraints (e.g., accuracy). Discarding of the assay can be performed in the next round of assay design where that assay or sample type will not be used. The new assay or sample can be one that was previously measured but where no data was used.
[0335] In block 530, new sample types, or potentially more samples of the same type, are obtained, for example, to increase the number of samples in the cohort.
[0336] In block 532 , a new assay can be performed, for example, based on a suggested assay from the assay identification module 526 .
[0337] The final assay 529 can specify, for example, the order of the assays in the set, the amount of data, the data quality, and the data throughput. The order of the assays can be optimized for cost and timing. The order and timing of the assays can be parameters that are optimized.
[0338] In some embodiments, the computer module can inform other parts of the wet lab steps. For example, some computer module(s) may precede the wet lab steps for some assay development procedures, such as when external data can be used to inform the starting point of the wet lab experiments. Additionally, the output of the wet lab experiment components may be fed to the computer components, such as cohort design and clinical questions. Meanwhile, the computer results may be fed back to the wet lab, such as the impact of the choice of cost function on the cohort design.
[0339] 4. How to design a multi-analyte assay Figure 6 shows the overall process flow of the disclosed method. In this example, the process steps are as follows:
[0340] In operation, at block 610, the system receives a plurality of training samples, each containing a plurality of classes of molecules, for which one or more labels are known. Examples of analytes are provided herein, such as cell-free DNA, cell-free RNA (e.g., miRNA or mRNA), proteins, carbohydrates, autoantibodies, or metabolites. The labels may be for a particular condition (e.g., cancer or different classifications of a particular cancer), or for treatment response. Block 610 may be performed by a receiver that includes a measurement device, such as one or more receiving devices, such as measurement devices 151-153 of FIG. 1. The measurement devices may perform different assays. The measurement devices may convert the samples into usable features (e.g., a large library of information about each analyte from the samples) so that the computer can select the combination of input features required by a particular ML model to classify a particular biological sample.
[0341] In block 620, for each of the plurality of different assays, the system identifies a set of features operable to be input into the machine learning model for each of the plurality of training samples. The set of features may correspond to properties of molecules in the training samples. For example, the features may be read counts in different regions, percentages of methylation in regions, counts of different miRNAs, or concentrations of a set of proteins. Different assays may have different features. Block 620 may be performed by feature selection module 520 of FIG. 5. In FIG. 5, feature selection may occur before or after feature extraction, for example, if features are already known based on the type of assay being performed. As part of an iterative procedure, a new set of features may be identified, for example, based on results from evaluation module 524.
[0342] At block 630, for each of a plurality of training samples, the system subjects a group of classes of molecules in the training sample to a plurality of different assays to obtain a set of measurements. Each set of measurements may be from one assay applied to a class of molecules in the training sample. A plurality of sets of measurements may be obtained for a plurality of training samples. As examples, the different assays may be lcWGS, WGBS, cf-miRNA sequencing, and protein concentration measurements. In one example, a portion contains more than one class of molecules, but only one type of assay is applied to the portion. The measurements may correspond to values obtained from an analysis of raw data (e.g., sequence reads). Examples of measurements are read counts of sequences that partially or completely overlap different genomic regions of the genome, percentages of methylation in a region, counts of different miRNAs, or concentrations of a set of proteins. Features may be determined from a plurality of measurements, e.g., statistics of the distribution of measurements, or concatenation of measurements appended to each other.
[0343] In block 640, the system analyzes the set of measurements to obtain a training vector for the training samples. The training vector may include feature values of a set of features of the corresponding assay. Each feature value may correspond to a feature and may include one or more measurements. The training vector may be formed using at least one feature from at least two of the N sets of features corresponding to a first subset of the plurality of different assays, where N corresponds to the number of different assays. A training vector may be determined for each sample, where the training vector potentially includes features from some or all of the assays, and therefore all classes of molecules. Block 640 may be performed by feature extraction module 514 of FIG. 5.
[0344] In block 650, the system operates on the training vectors using parameters of the machine learning model to obtain output labels for a plurality of training samples. Block 650 may be performed by a machine learning module that implements the machine learning model.
[0345] In block 660, the system compares the output labels to the known labels of the training samples. A comparator module can perform such a comparison of the labels to form a measure of error of the current state of the machine learning model. The comparator module can be part of the training module 522 of FIG. 5.
[0346] A first subset of the plurality of training samples may be identified as having a designated label, and a second subset of the plurality of training samples may be identified as not having the designated label, in one example, the designated label is a clinically diagnosed disorder, e.g., colon cancer.
[0347] In block 670, the system iteratively searches for optimal values of parameters as part of training the machine learning model based on comparing the output labels with known labels of the training samples. Various techniques for performing the iterative search are described herein, for example, gradient techniques. Block 670 may be implemented by the training module 522 of FIG. 5.
[0348] Training of the machine learning model can provide a first version of the machine learning model after a refinement stage, which may include, for example, one or more additional passes through modules 512-528. A quality metric can be determined for the first version, and the quality metric can be compared to one or more criteria, for example, a threshold. The quality metric may consist of various metrics, for example, an accuracy metric, a cost metric, a time metric, etc., as described with respect to FIG. 4. Each of these metrics can be individually compared to a threshold or otherwise determined to determine whether the metric meets one or more criteria. Based on the comparison(s), for example, in the case of FIG. 5, at blocks 526 and 532, it can be determined whether to select a new subset of assays for determining a set of features.
[0349] The new subset of assays may include at least one of the multiple different assays that were not in the first subset and / or may potentially remove assays. The new subset of assays may include at least one assay from the first subset, and a new set of features may be determined for one assay from the first subset. If the quality metrics of the new subset of assays meet one or more criteria, the new subset of assays may be output, for example, as final assays 529 of FIG. 5.
[0350] If the new subset includes a new assay not previously performed, the molecules in the training sample can be subjected to the new assay not included in the plurality of different assays to obtain a new set of measurements based on the quality metric for the subset of new assays that do not meet one or more criteria. The new assay can be performed on a new class of molecules not in the group of classes of molecules.
[0351] In block 680, the system provides a set of machine learning model parameters and machine learning model features. The machine learning model parameters may be stored in a predefined format or may be stored with tags that identify the number and identity of each of the parameters. The feature definitions may be obtained from settings used in feature extraction and selection, for example, as specified by the current iteration through the feature extraction module 514 and feature selection module 520. Block 680 may be performed by an output module.
[0352] 5. How to identify cancer In one aspect, the disclosure provides a method of identifying cancer in a subject, comprising: (a) providing a biological sample comprising a cell-free nucleic acid (cfNA) molecule from the subject; (b) sequencing the cfNA molecule from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 15,000 distinct hypomethylated regions; and (e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has the cancer.
[0353] In some embodiments, the trained algorithm includes performing dimensionality reduction by singular value decomposition. In some embodiments, the method further includes generating a quantitative measure of the plurality of cfNA sequencing reads at each of a second plurality of genomic regions of the reference genome to generate a second cfNA feature set, the second plurality of genomic regions of the reference genome comprising at least about 20,000 distinct protein-coding gene regions, and applying the trained algorithm to the second cfNA feature set to generate the likelihood that the subject has the cancer. In some embodiments, the method further includes generating a quantitative measure of the plurality of cfNA sequencing reads at each of a third plurality of genomic regions of the reference genome to generate a third cfNA feature set, the third plurality of genomic regions of the reference genome comprising contiguous non-overlapping genomic regions of equal size, and applying the trained algorithm to the third cfNA feature set to generate the likelihood that the subject has the cancer. In some embodiments, the third plurality of non-overlapping genomic regions of the reference genome comprises at least about 60,000 distinct genomic regions. In some embodiments, the method further comprises generating a report comprising information indicative of the likelihood that the subject has the cancer. In some embodiments, the method further comprises generating one or more recommended steps for treating the cancer for the subject based at least in part on the generated likelihood that the subject has the cancer. In some embodiments, the method further comprises diagnosing the subject with the cancer when the likelihood that the subject has the cancer meets a predetermined criterion. In some embodiments, the predetermined criterion is that the likelihood is greater than a predetermined threshold. In some embodiments, the predetermined criterion is determined based on an accuracy metric of the diagnosis. In some embodiments, the accuracy metric is selected from the group consisting of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC).
[0354] In some examples, the computer modules can inform other parts of the wet lab steps. For example, some computer modules may precede the wet lab steps for some assay development procedures, such as when using external data to inform the starting point of the wet lab experiments. Additionally, the output of the wet lab experiment components can be fed into computer components such as cohort design and clinical questions. In turn, computer results may be fed back to the wet lab, such as the impact of the choice of cost function on the cohort design.
[0355] 6.Results Table 2 shows the results for different specimens and the corresponding best performing models according to embodiments of the present disclosure. [Table 2]
[0356] Samples that were similar across specimens were used.
[0357] In Table 2, SD refers to significant difference and is determined by comparing the read counts of different genes between different classification labels. This is part of dimensionality reduction. It filters out features that are significantly different between two classifications and then transfers them to the classification. PCA looks at groups of features that are put together but correlated in a certain way, while SD looks at individual features unilaterally. The feature with the highest SD (e.g., gene read counts) can be used for the feature vector of interest. PCA is about the projection of measurements through the first few components. This is, for example, a condensed representation of many features in a smaller dimensional space.
[0358] The table was created by analyzing the results of different models with different dimensionality reductions (including no reduction) for different combinations of analytes. The table includes the best performing models. As an example, for multi-analyte assay datasets involving proteins, there may be no need for PCA due to the small dimensionality (14), and therefore only logistic regression (LR) is used.
[0359] Of these models, LR was tried along with PCA (top 5 components) and showed significant differences in feature selection (retaining 10% of features). PCA can be performed across samples or within a sample.
[0360] The feature columns correspond to different combinations of analytes, e.g., genes (cell-free DNA analysis) + methylation. If more than one analyte is used, there are two options: combine the features into one feature set, or run two models and output two classifications (e.g., probabilities of the classifications) and use those as a vote, e.g., majority vote or some weighted average or probability, to determine which classification has the highest score. As another example, one could take the average or mode of the predictions as opposed to looking at the scores.
[0361] Cross-validation was performed five times to obtain AUC information for the receiver operating characteristic curves in Figures 7A and 7B. Samples can be divided into five different datasets (four datasets for training and the fifth dataset for validation). Sensitivity and specificity can be determined for the four sets. Furthermore, the assignment to the sets can be updated with a random number seed to provide more data. To determine sensitivity and specificity, the four classifications were reduced to four, with healthy and benign polyps as one classification and AA and CRC as the other classifications.
[0362] 7A and 7B show the classification performance of different analytes.
[0363] B. Example 2: Analysis of Individual Assays for Classification of Biological Samples This example describes the analysis of multiple specimens and multiple assays to differentiate between healthy individuals, AA, and stages of CRC.
[0364] Blood samples were separated into different fractions and three classes of molecules were investigated in four assays: cell-free DNA, cell-free miRNA, and circulating proteins. Two assays were performed on cfDNA.
[0365] De-identified blood samples were obtained from healthy individuals and individuals with benign polyps, advanced adenomas (AA), and stage I-IV colorectal cancer (CRC). After plasma isolation, multiple specimens were assayed as follows: First, cell-free DNA (cfDNA) content was assessed by low-coverage whole genome sequencing (lcWGS) and whole genome bisulfite sequencing (WGBS). Second, cell-free microRNAs (cf-miRNAs) were assessed by small RNA sequencing. Finally, levels of circulating proteins and RNA were measured by quantitative immunoassays.
[0366] Sequenced cfDNA, WGBS, and cf-miRNA reads were aligned to the human reference genome (hg38) and analyzed as follows, as provided in further detail in the Materials and Methods section. cfDNA (lcWGS): Fragments aligned within annotated genomic regions were counted and normalized for sequencing depth to generate a 30,000-dimensional vector per sample, where each element corresponds to the count of a gene (e.g., the number of reads that align to that gene in the reference genome). Samples with high tumor fraction (>20%) were identified via manual inspection of large-scale CNVs.
[0367] WGBS: Percent methylation was calculated per sample across LINE-1 CpGs and CpG sites in the target genes (56 genes).
[0368] cf-miRNA: Fragments that aligned to annotated miRNA genomic regions were counted and normalized for sequencing depth to generate a 1700-dimensional vector per sample.
[0369] Each of these sets of data can be filtered to identify measurements (e.g., reads aligned to a reference genome to obtain counts of reads for distinct genes). Measurements can be normalized. Normalization for each sample is described in more detail in a separate subsection for each sample. PCA analysis was performed for each sample and the results were obtained. The application of machine learning models is provided in a separate section.
[0370] 1.cf-DNA low-coverage whole-genome sequencing For a list of known genes with annotated regions, sequence read counts were determined for each of those annotated regions by counting the number of fragments that aligned to that region. The read counts of genes can be normalized in a variety of ways, for example, using a global expectation where the genome is spread, within-sample normalization, and cross-feature normalization. Cross-feature normalization can refer to all of those features being averaged to a specified value, for example, 0, a different negative value, 1, or a range of 0-2. With regard to cross-feature normalization, the total reads from a sample are variable and therefore may depend on the preparation process and the sequencing load process. Normalization can be made to a constant number of reads as part of the global normalization.
[0371] For intra-sample normalization, it is possible to normalize by some features or screening features of some regions, especially GC bias. Thus, the composition of base pairs in each region is different and can be used for normalization. In some cases, the number of GC is significantly higher or lower than 50%, which has a thermodynamic impact since its bases are more energetic and the process is biased. Some regions give more reads than expected due to biological artifacts of sample preparation in the laboratory. Therefore, it may be necessary to correct such biases by applying another kind of feature / feature transformation / normalization method.
[0372] Figures 8A and 8B show the distribution of high tumor fraction samples (i.e., greater than 20%) inferred by CNV across clinical stages, showing the difference between healthy and normal. In this example, lcWGS of plasma cfDNA allows identification of CRC samples with high tumor fraction (>20%) based on genome-wide CNV. Furthermore, high tumor fraction is more frequently observed in late-stage CRC samples, but was also observed in some samples of stage I and II. High tumor fraction was not observed in samples from healthy individuals or individuals with benign polyps or AA.
[0373] Figures 8A-H show CNV plots of individuals with high (>20%) tumor fraction based on cfDNA-seq data. Note that each plot in Figures 8A-H corresponds to a unique sample's histogram of self-read DNA copy number. Also note that tumor fraction can be calculated by estimation from CNV or by using open source software such as ichor DNA. Table 3 shows the distribution of high tumor fraction cfDNA samples across clinical stages. [Table 3]
[0374] High tumor fraction samples do not necessarily correspond to samples clinically classified as late stage. In the figure, the total number of healthy subjects is 26. "BP" refers to benign polyp, "AA" refers to advanced adenoma, and "Chr" refers to chromosome.
[0375] 2. Methylation Variable methylation regions (DMRs) are used for CpG sites. Regions can be dynamically assigned by discovery. It is possible to take several samples from different classes and discover which regions are most variably methylated between different classifications. Then, select the variably methylated subsets and use these for classification. The number of CpGs captured in the region is used. Regions can tend to be of variable size. Therefore, it is possible to do a pre-discovery process that bundles several CPG sites together as a region. In this example, 56 genes and LINE1 elements (regions repeated across the genome) were studied. To perform classification, the percentage of methylation in these regions was investigated and used as features to train the machine learning model. In this example, the classification utilizes essentially 57 features used for PCA. Specific regions can be selected based on regions that had sufficient coverage throughout the samples.
[0376] Figure 9 shows the CpG methylation analysis at LINE-1 sites, showing the difference between healthy and normal samples. The figure shows the methylation of all 57 regions used in the PCA. Each data point shown for the normal sample is for a different gene region and methylation.
[0377] In this example, genome-wide hypomethylation at LINE-1 CpG loci was observed only in individuals with CRC. No hypomethylation was observed in samples without CRC, e.g., from healthy individuals or individuals with benign polyps or AA. Note that each normal data point is for a different gene region and methylation. In one example, all reads that map to a region may be calculated. The system may determine whether a read is a methylated position, then sum the number of methylated CpGs (e.g., consecutively adjacent C and G bases) and methylated CpGs, and calculate the ratio of the number of methylated CpGs to the number of methylated CpGs.
[0378] In this example, significance was assessed by one-way analysis of variance (ANOVA) followed by Sidak's multiple comparison test. Only significant adjusted P values are shown. CpG hypomethylation of LINE-1 was observed only in CRC cases. Polyp (benign polyp), AA, CRC (stages I-IV). 5mC, 5-methylcytosine.
[0379] The proportion of DNA fragments that align to the site and have methylation can be studied in the entire region of interest.For example, a gene region may have two CpG sites (e.g., consecutive C and G bases adjacent to each other) for every 190 total reads (e.g., 100 reads that align to the first CpG site and 90 reads that align to the second CpG site).All reads that map to the region are found and observed whether the read is methylated or not.The number of methylated CpGs is then summed up and the ratio of the number of methylated CpGs to the number of unmethylated CpGs is calculated.
[0380] 3. MicroRNA In this example, virtually all microRNAs (miRNAs) that were measurable (approximately 1700 in this example) were used as features. Measurements are related to the expression data of these miRNAs. Their transcripts are of a certain size, and each transcript is stored, and the number of miRNAs found for each can be counted. For example, the RNA sequences can be aligned to a reference miRNA sequence, e.g., a set of 1700 sequences corresponding to known miRNAs in the human transcriptome. Each miRNA found can be used as its own feature, and across all samples, the whole can be a set of features. Some samples have a feature value that is 0 if no expression is detected for that miRNA.
[0381] Figure 10 shows cf-miRNA sequencing analysis to characterize microRNAs. After pooling reads from all samples, ranked by expression, the number of reads mapping to each miRNA is shown. miRNAs shown in red have been suggested in the literature as potential CRC biomarkers. Adapter-trimmed reads were mapped to adult human microRNA sequences (miRBase21) using bowtie2. Over 1800 miRNAs were detected in plasma samples with at least one read, while 375 miRNAs were present at higher abundance (detected with an average of ≥10 reads per sample).
[0382] In one example, all samples are taken and the read values are aggregated together. For each microRNA found in the sample, a large number of aggregated reads can be found. In this example, about 10 million aggregated reads are found to map to one single microRNA, with 300 microRNAs found in more than 1,000 reads, about 600 found in more than 100 reads, 1,200 found in 10 reads, and about 1,800 found in only a single read in the aggregate. It should be noted that microRNAs with high expression ranks may provide better markers because they can provide more reliable signals with larger absolute changes.
[0383] The cf-miRNA profiles in individuals with CRC were inconsistent with those in healthy controls. In this example, miRNAs suggested in the literature as potential CRC biomarkers tended to be present at higher abundance compared to other miRNAs.
[0384] 4. Protein Protein data was normalized by standard curves (14 proteins). Each of the 14 proteins has its own standard curve because it is essentially a unique immunoassay, typically a recombinant protein, and very stable in an optimized buffer. Thus, a standard curve is generated that can be calculated in many ways. The concentration relationship is typically nonlinear. Samples are then run and calculated based on the expected fluorescence concentration in the primary sample. Measurements may be measured in triplicate, but can be reduced to 14 individual values, for example, by averaging or more complex statistical analysis.
[0385] Figure 11A and Figures 11B-11G show the distribution of circulating protein biomarkers. Figure 11A is a box plot showing the levels of all circulating proteins analyzed, with outliers shown as diamonds. Figures 11B-11G show that protein levels are significantly different across tissue types by one-way ANOVA followed by Sidak's multiple comparison test. Only significant adjusted P values are shown. Proteins measured using SIMOA (Quanterix): ATP-binding cassette transporter A1 / G1 (A1G1), acylation stimulatory protein (C3a des Arg), cancer antigen 72-4 (CA72-4), carcinoembryonic antigen (CEA), cytokeratin fragment 21-1 (CYFRA21-1), FRIL u-PA. Proteins measured by ELISA (Abcam): AACT, cathepsin D (CATD), CRP, cutaneous T cell-induced chemokine (CTACK), FAP, matrix metalloproteinase-9 (MMP9), SAA1.
[0386] In this example, circulating levels of alpha-1-antichymotrypsin (AACT), C-reactive protein (CRP), and serum amyloid A (SAA) proteins were elevated in CRC samples, whereas urokinase-type plasminogen activator (u-PA) levels were lower compared to healthy controls. Circulating levels of fibroblast activation protein (FAP) and Flt3 receptor-interacting lectin precursor (FRIL) proteins were elevated in AA samples, whereas CRP levels were lower compared to CRC samples.
[0387] In this example, differences can be observed between some of the ANOVA plots. For example, CRP appears predictable. FAP varies among different. Thus, multi-analyte tests may show aggregation trends, but each one may be difficult to assess individually.
[0388] 5. Dimensionality reduction (e.g., PCA or significant differences) Principal component analysis (PCA) was performed on a per-sample basis. In one embodiment, PCA is performed on protein, cell-free DNA, methylation, and microRNA data. Thus, four PCAs can be performed in context.
[0389] In one example, 14 proteins can be considered as a single analyte. There are 14 measurements for a protein, and therefore 14 concentrations based on individual fluorescence. These are vectorized by 14. The output of the PCA can be: component 1 explaining 31% of the variance, component 2 explaining 17% of the variance, etc. This allows one to identify which proteins contribute the most variance.
[0390] For lcWGS on cell-free DNA, use difference in gene count statistics (e.g., mean, median, etc.) to identify genes with the highest variance.
[0391] Figures 12A-D show the output of PCA analysis of cf-DNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. Figures 12E-H show PCA of cf-DNA, CpG methylation, cf-miRNA and protein counts as a function of specimen. High tumor fraction samples have consistently abnormal behavior across all four specimens examined.
[0392] In the example of Figures 12A-12D, PCA is used to separate the distance between high and low tumor fractions. In Figures 12E-12H, it is a sample classification of different specimens (normal, healthy, benign polyp, and colon cancer). The disclosed system and method can be used to maximize the differentiation between such classes. In this example, the abnormal profile across specimens indicated high TF (inferred from cfDNA CNVs) but not the stage of cancer. Each dot shown corresponds to a separate sample and PCA is the value of the highest component.
[0393] Various implementations can be used for dimensionality reduction. For dimensionality reduction, several different hypothesis tests can be used to calculate, for example, there are several different criteria used to set the threshold for significant differences and how many to include. PCA or SVD (singular value decomposition) can be done on the correlation or covariance matrix rather than the data itself. Autoencoding or variational autoencoding can be used. Such filtering can filter measurements with low variance (e.g., region counts).
[0394] 6. Conclusion lcWGS of plasma cfDNA was able to identify CRC samples with high tumor fraction (>20%) based on genome-wide copy number variation (CNV). High tumor fraction was observed more frequently in late-stage cancer samples, but also in some stage I and II patients. Abnormal signals in each of three other samples (cf-miRNA profiles) that were mismatched in healthy controls, genome-wide hypomethylation at the LINE1 (long interspersed nuclear factor 1) CpG locus, and elevated levels of circulating carcinoembryonic antigen (CEA) and cytokeratin fragment 21-1 (CYFRA 21-1) proteins were also observed in cancer patients. Surprisingly, the multi-sample abnormal profiles indicated high tumor fraction (as estimated from cfDNA CNV) but not cancer stage.
[0395] These data suggest that tumor fraction correlates with cancer stage but has a large potential range even in early samples.Previous literature on blood-based screening for cancer detection has shown discrepancies in terms of the claimed ability of different single analytes to detect early cancer. The historical discrepancy may be explained by tumor fraction, as we found that abnormal profiles between cfDNA CpG methylation, cf-miRNA, and circulating protein levels were more strongly associated with high tumor fraction than with late stage. These findings suggest that some positive "early" findings may in fact be "high tumor fraction" findings. The results further demonstrate that assaying multiple analytes from a single sample allows the development of classifiers to reliably detect pre-malignant or early stage disease at low tumor fractions. Such a multi-analyte classifier is described below.
[0396] C. Example 3: Identification of Hi-C-like structures using covariance of sequencing depth in two different genomic regions from cfDNA across multiple samples. This example describes how to identify Hi-C-like structures in two distinct genomic regions from cfDNA in a single sample to identify cell type of origin as features for the generation of multi-analyte models.
[0397] Genomic sequences of multiple cfDNA samples were segmented into non-overlapping bins of various lengths (e.g., 10-kb, 50-kb, and 1-Mb non-overlapping bins). The number of high-quality mapped fragments in each bin was then quantified. High-quality mapped fragments met a quality threshold. Pearson / Kendall / Spearman correlation was then used to calculate correlations between pairs of bins in the same chromosome or between different chromosomes. The structure scores calculated from the nuance structures of the correlation matrix were used to generate the heat map shown in Figure 13. A similar heat map was generated using the structure scores determined using Hi-C sequencing, as shown in Figure 14. The similarity of the two heat maps suggests that the nuance structures determined using covariance were similar to those determined by Hi-C sequencing. Potential technical biases due to GC bias, genomic DNA, and correlated structures in MNase digestion were excluded.
[0398] Genomic regions (larger bin size) were divided into smaller bins and the correlation between the two larger bins was calculated using the Kolmogorov-Smirnov (KS) test. The KS test score provided information about the Hi-C-like structure that could be used to discriminate between cancer and control groups.
[0399] Two-dimensional segmentation (HiCseg) was used to segment and call domains in the correlation structures in cfDNA and Hi-C. The two approaches yielded similar numbers of domains and highly overlapping domains.
[0400] Identification of cfDNA-specific co-emission patterns. Covariance structure analysis in cfDNA showed mixed input signal patterns from multiple sources, including chromatin structure, genomic DNA, MNase digestion, and possible co-emission patterns of cfDNA. Deep learning was used to remove signals from other sources and retain only potential co-emission patterns of cfDNA.
[0401] The three-dimensional proximity of chromatin in cancer and non-cancer samples can be inferred from long-range spatially correlated fragmentation patterns. The fragmentation patterns of cfDNA from different genomic regions are not uniform and reflect the local epigenetic signatures of the genome. There is a high similarity between the long-range epigenetic correlation structure and higher-order chromatin organization. Thus, the long-range spatially correlated fragmentation patterns may reflect the three-dimensional proximity of chromatin. Using only fragment lengths in cfDNA, we generated a genome-wide map of in vivo higher-order chromatin organization inferred from co-fragmentation patterns. Fragments generated from endogenous physiological processes can reduce the possibility of technical variations associated with random ligation, restriction enzyme digestion, and biotin ligation during Hi-C library preparation. Sample collection and pre-processing: Retrospective human plasma samples (>0.27 mL) were obtained from 45 patients diagnosed with colon cancer (colorectal cancer), 49 patients diagnosed with lung cancer, and 19 patients diagnosed with melanoma. One hundred samples were also obtained from patients without a current cancer diagnosis. In total, samples were collected from commercial biobanks from Southern and Northern Europe, and the United States. All samples were de-identified. Plasma samples were stored at -80°C and thawed before use. Cell-free DNA was extracted from 250 μL of plasma (spiked with unique synthetic dsDNA fragments for sample tracking) using the MagMAX cell-free DNA isolation kit (Applied Biosystems) according to the manufacturer's instructions. Paired-end sequencing libraries were prepared using the NEBNext Ultra II DNA Library Preparation Kit (New England Biolabs) and sequenced on an Illumina NovaSeq 6000 sequencing system with dual indexing across multiple S2 or S4 flow cells at 2x51 base pairs.
[0402] Whole genome sequencing data processing: Reads were demultiplexed and aligned to the human genome (GRCh38 with decoys, alt contigs, and HLA contigs) using BWA-MEM 0.7.15. Unique molecular identifiers (UMIs) were used to remove PCR overlapping fragments. Contamination was assessed using a marginalized contamination model across all possible genotypes and contamination fractions for common SNPs identified by 1000 Genomes (IGSR).
[0403] Sequencing data were checked for quality and omitted from analysis if they met any of the following conditions: AT dropout >10 or GC dropout >2 (both calculated via Picard 2.10.5). Any samples suspected of contamination due to expected allele fraction <0.99, unexpected genotype calls, or failed negative controls were manually inspected before inclusion in the dataset. Adapters were trimmed by Atropos using default parameters. Only properly paired, high-quality reads with unique mapping at both ends (mapping quality score >60) and no PCR duplicates were used for all downstream analyses. Only autosomes were used for all downstream analyses.
[0404] Hi-C library preparation: In situ Hi-C library preparation of whole blood cells and neutrophils was performed by using Arima Genomics Services.
[0405] Hi-C data processing: Raw Fastq files were uniformly processed through Juicerbox command line tool v1.5.6. After filtering the reads, results with a mapping quality score >30 were used to generate a Pearson correlation matrix and parcel A / B. Principal component analysis (PCA) was calculated by the PCA function in scikit-learn 0.19.1 in Python 3.5. The first principal component was used to segment the parcels. For each chromosome, parcels were grouped into two groups based on sign. The group of parcels with a low average value of gene density was defined as parcel B. The other group was defined as parcel A. Gene density was determined by gene number annotated by Ensemble v84. Sequencing summary statistics and associated metadata information are shown in Table 4. [Table 4]
[0406] Multi-sample cfHi-C: 500 kb bins with mappability less than 0.75 were removed for downstream analysis. First, each 500 kb bin was divided into 50 kb sub-bins. The median fragment length of each sub-bin was first collapsed into 500 kb bins, then normalized by the z-score method using the mean and standard deviation of each chromosome and each sample. Pearson correlation between each pair of bins across all individuals was calculated.
[0407] Single-sample cfHi-C: 500 kb bins with mappability less than 0.75 were removed from downstream analysis. The fragment lengths of all high-quality fragments in each 500 kb bin were then determined. The distribution similarity of fragment lengths within each pair of 500 kb bins was calculated by a two-sample KS test (ks_2samp function implemented in SciPy 1.1.0 in Python 3.6). P values were then transformed to the log10 scale. Pearson correlations were then calculated for specific paired bins.
[0408] Sequence composition and mappability bias analysis: Mappability scores were generated by GEM17 for a read length of 51 bp. G+C% was calculated by gc5 base track from UCSC genome browser. For each pair of 500 kb bins, G+C% and mappability were obtained from bin1 and bin2. Gradient boosting machine (GBM) regression trees (GradientBoostingRegressor function implemented in scikit-learn 0.19.1 in Python 3.6) were then applied to regress the G+C% and mappability of each pixel of the correlation coefficient score from the matrix of cfHi-C, gDNA, and Hi-C data. N_estimators were varied at depth=5 for different model complexities. Residual values after regression were then used to calculate correlation with whole blood cell (WBC) Hi-C data at pixel level. r2 values were calculated to measure the goodness of fit of the model.
[0409] Tissue of origin analysis in cfHi-C: To infer tissue of origin from cfHi-C data, the parcels of cfHi-C data (first PC on correlation matrix in cfHi-C) were modeled as linear combinations of the parcels in each reference Hi-C data (first PC on correlation matrix in cfHi-C). Eigenvalues were re-evaluated to ensure that parcel A was a positive number. Genomic regions with mappability less than 0.75 were filtered out. Eigenvalues across cfHi-C and reference Hi-C panels were first transformed by quantile normalization. For each reference Hi-C dataset, only the genomic bins that showed the highest eigenvalues relative to the remainder of the reference Hi-C dataset (or the lowest if the eigenvalues were negative) were used for deconvolution analysis. Weights were constrained to sum to 1, allowing the weights to be interpreted as tissue contributions to cfDNA. Quadratic programming was used to solve the constrained optimization problem. To define the tumor fraction, tissue contribution fractions from cancer were summed.
[0410] ichorCNA analysis: ichorCNA v0.1.0 with default parameters was used to calculate the tumor fraction in each cfDNA WGS sample after normalization to the group of internal healthy samples. Code and Data Availability: All analysis code was implemented in Python 3.6 and R 3.3.3. Publicly available data used in the study are shown in Table 5. Detailed summary statistics of fragment lengths at the genomic bin level for each cfDNA sample. [Table 5]
[0411] Paired-end whole genome sequencing (WGS) was performed on cfDNA from 568 different healthy individuals. For each sample, an average of 395 million paired-end read values were obtained (approximately 12.8-fold coverage). After quality control and read filtering, an average of 310 million high-quality paired-end read values (approximately 10-fold coverage) were obtained for each sample. Autosomes were divided into non-overlapping bins of 500 kb, and normalized fragmentation scores were calculated for each individual sample from the fragment lengths in each bin alone. Pearson correlation coefficients were then calculated between each pair of bins with normalized fragmentation scores across all individuals. Similar patterns were seen between the fragmentation correlation maps of cfDNA and the partitions of Hi-C experiments from whole blood cells (WBCs) from two healthy individuals (Figures 15A-C and 15D). Figure 15A-C shows the correlation map generated from Hi-C, spatially correlated fragment length from multiple cfDNA samples, and spatially correlated fragment length distribution from a single cfDNA sample. Figure 15D shows genome browser tracks of section A / B from Hi-C (WBC), multiple sample cfDNA, and single sample cfDNA. All comparisons were made from chromosome 14 (chr14).
[0412] To quantify the degree of similarity, Pearson correlations between chromatin organization inferred from Hi-C and cfDNA were calculated at the pixel level (genome-wide average Pearson r=0.76, p<2.2e-16). Pixel-level correlation coefficients shown in Hi-C were calculated from replicates of two different healthy individuals. Pixel-level correlation coefficients shown in cfDNA (multiple samples in Figure 15E and single sample in Figure 15F) were calculated by correlation with WBC individual 2.
[0413] We further called the chromatin organization inferred from the Hi-C data and cfDNA. At the compartment level, we found a higher concordance between the chromatin organization inferred from Hi-C and cfDNA (Pearson r=0.89, p<2.2e-16). The compartments A / B called from Hi-C were highly overlapped with the results from cfDNA (hypergeometric test p<2.2e-16). This approach is referred to as cfHi-C.
[0414] To extend the application of cfHi-C to the single-sample level, we divided each 500 kb bin in each sample into smaller 5 kb subbins and used the Kolmogorov-Smirnov (KS) test to measure the similarity of fragmentation score distribution between each pair of 500 kb bins. The KS test further confirmed the high correlation between Hi-C and cfHi-C at both the pixel and compartment levels (Figure 16A and Figure 16B). To eliminate possible internal library preparation and sequencing biases due to the patterned flow cell technology in NovaSeq, we replicated the algorithm using a publicly available external cfDNA dataset generated by the HiSeq 2000 platform (BH01). Using this dataset, similar patterns were observed in healthy cfDNA samples (Figure 15D).
[0415] To eliminate possible technical biases caused by sequence composition, the locally weighted scatterplot smoothing (LOWESS) method was applied to normalize the fragment lengths of each bin by the average G+C% value. After regressing G+C%, a high similarity was observed between Hi-C and multi-sample cfHi-C in WBCs (Pearson correlation r=0.57, p<2.2e-16, Figure 17A in Figure 17A-D and Figure 17B in Figure 17A-D).
[0416] As a negative control, the same steps were repeated using genomic DNA (gDNA) from primary white blood cells from 120 individuals. Before regressing G+C%, again, a relatively high similarity was observed between Hi-C and gDNA (Pearson correlation r=0.40, p<2.2e-16, Fig. 17C in Fig. 17A-D and Fig. 17D in Fig. 17A-D). However, after normalizing with G+C% in gDNA, a low residual similarity was observed between Hi-C and gDNA (Pearson correlation r=0.15, p<2.2e-16, Fig. 17D in Fig. 17A-D) and Hi-C-like block structures were no longer observed. Figure 17E shows a boxplot of pixel-level correlations (Pearson and Spearman) with Hi-C (WBC, repeat 2) across all chromosomes represented in Figs. 17A-17D in Fig. 17A-D.
[0417] To clarify the effect of G+C% and mappability in two-dimensional space, GBM regression trees were applied to cfHi-C. For each pixel on the cfHi-C matrix, we obtained the two G+C% and mappability values in the interaction pair bins, and then regressed the G+C% and mappability from the signal at each pixel of the cfHi-C matrix. After regressing the bias of G+C% and mappability, we observed significant residual similarity between Hi-C in WBC and both multi-sample (Pearson correlation r=0.28, p<2.2e-16, n_estimator=500, FIG. 18A) and single-sample cfHi-C (Pearson correlation r=0.36, p<2.2e-16, n_estimator=500, FIG. 18B).
[0418] In the negative control using gDNA, no residual similarity was observed between Hi-C and both multi-sample (Pearson correlation r=0.009, p=0.0002, FIG. 18C) and single-sample gDNA (Pearson correlation r=-0.03, p<2.2e-16, FIG. 18D) in WBCs at the same range of model complexities. Furthermore, for each pair of bins in cfDNA, one of the bins was replaced with a random bin from another chromosome with the same G+C% and mappability, and the cofragmentation score was recalculated. By using the same GBM regression tree approach on the simulated cfHi-C matrix, significantly lower residual similarity was observed with Hi-C at the same range of model complexities (Pearson correlation r=0.13, p<2.2e-16, FIG. 18E).
[0419] To demonstrate that the model retained biological signal after regressing G+C% and mappability, the same regression tree approach was applied to WBC Hi-C from another individual (replicate 1). High similarity was still observed in the replicates (Pearson correlation r=0.53, p<2.2e-16, FIG. 18F).
[0420] To explore the model complexity effect on the analysis, the regression tree was repeated with different model complexities (n_estimator). Using multi-sample cfHi-C, single-sample cfHi-C, and Hi-C from different individuals, it was difficult to remove the correlation with Hi-C even at high model complexity. This phenomenon did not occur with negative control samples such as multi-sample gDNA, single-sample gDNA, and cfHi-C with permuted bins.
[0421] To exclude the possibility that the co-fragmentation patterns observed in multi-sample cfHi-C were due to batch defects during sequencing and library preparation, one bin was randomly shuffled between individuals for each pair of bins in cfHi-C. As expected, no correlation with Hi-C was observed (Pearson correlation r=-0.0002, p=0.74, Figures 19A and 19D). A multi-sample cfHi-C matrix was generated from samples in the same batch (18 samples). High correlation was observed between Hi-C at the pixel level (Pearson correlation r=0.60, p<2.2e-16, Figures 19B and 19D) and samples downsampled to the same size (Pearson correlation r=0.63, p<2.2e-16, Figures 19C and 19D).
[0422] To test the robustness of this approach, data from different sample sizes were randomly subsampled for multi-sample cfHi-C. With a sample size of 10, correlation coefficients of approximately 0.55 at the pixel level and 0.7 at the compartment level with WBC Hi-C were obtained. Saturation was obtained with a sample size of >80 (Figures 20A-20D).
[0423] To understand the effect of bin size, the same procedure was repeated with different bin sizes. High agreement with Hi-C experiments at different resolutions was consistently observed (Figures 21A-21H). To clarify the effect of sequencing depth in single-sample cfHi-C, fragment numbers were downsampled to different sizes. Even at about 0.7-fold coverage, correlation coefficients of about 0.45 at the pixel level and 0.7 at the compartment level with WBC Hi-C were still obtained (Figures 22A and 22B).
[0424] To determine whether the observed cfHi-C signals change in different pathological conditions, additional WGS was generated at similar sequencing depths for cfDNA obtained from 45 colon cancer, 48 lung cancer, and 19 melanoma cancer patients. After standardizing eigenvalues at the compartment level across all cfHi-C samples, principal component analysis (PCA) was applied to all healthy samples and selected cancer samples containing high tumor fractions (tumor fraction >= 0.2, estimated by ichorCNA). Even at 500 kb resolution, separation was observed between healthy and different types of cancer samples (Figure 23A). By further applying a semi-monitored dimensionality reduction method, canonical correlation analysis (CCA), a clear separation was observed between healthy and cancer samples (Figure 23B-23F).
[0425] To determine whether the in vivo chromatin organization measured through cfDNA could be used to infer the cell types contributing to cfDNA in healthy individuals and patients with cancer, the amplitude of eigenvalues observed in the Hi-C data was correlated with the amplitude of open / closed states in chromosomes. At 500 kb resolution from GM12878, a significantly high correlation was observed between DNase-seq signal intensity in Hi-C compartments and eigenvalues (Pearson correlation r=0.8, p<2.2e-16, FIG. 24). This observation suggested that eigenvalues at the compartment level could be further used to quantify the degree of chromosomal openness.
[0426] To generate a reference Hi-C panel for tissue of origin analysis, Hi-C data from 18 different cell types were uniformly processed from different pathological and healthy conditions. To determine whether correlation patterns were cell-specific, in situ Hi-C data were generated from neutrophil cells with 1.96 billion paired reads and 1.06 billion high-quality contacts (mapping quality score >30). Using quantile-normalized eigenvalues in cell type-specific compartments identified from the reference Hi-C panel, approximately 80% of cfDNA was detected from different types of white blood cells, and almost no cfDNA was detected from cancer cells in cfHi-C (Figures 25A-25C). In contrast to healthy samples, increased fractions of cancer components from related cell types were observed in colon cancer, lung cancer, and melanoma samples using cfHi-C (Figures 25A and 25B).
[0427] To exclude possible artifacts during library preparation and sequencing, the procedure was repeated using publicly available cfDNA WGS data from healthy individuals, colon cancer, lung squamous cell carcinoma, small cell lung adenocarcinoma, and breast cancer samples. Similar results were observed (Figures 25A and 25B).
[0428] To quantify the accuracy of the approach, tumor fractions estimated by cfHi-C were compared to those estimated by ichorCNA, an orthogonal method for estimating tumor fractions by coverage using copy number variation (CNV) in cfDNA. Similar low tumor fractions were observed in healthy individuals (median tumor fraction = 0.00, mean = 0.02, Figure 25C), and significantly higher concordance with ichorCNA was observed in different cancer patients (Figure 26). To avoid confounding CNVs from late-stage cancers, we excluded genomic regions with any significant CNV signals for tissue of origin analysis. The results were nearly identical to those before exclusion of late-stage cancer samples.
[0429] If the long-range, spatially correlated fragmentation patterns observed in cfDNA are primarily influenced by the epigenetic landscape, similar 2D Hi-C-like patterns may be observed in different epigenetic signals. To test this hypothesis at the single-sample level, a modified KS test was used to determine the similarity between paired bins in different epigenetic signals from GM12878. High concordance was observed with Hi-C experiments from the same cell type using DNase-seq, methylation levels from whole-genome bisulfite sequencing (WGBS), H3K4me1 ChIP-seq, and H3K4me2 ChIP-seq. This observation suggests that the "virtual compartments" inferred from these epigenetic marks are a comprehensive reference panel to perform nuanced tissue-of-origin analyses.
[0430] In conclusion, these analyses demonstrate the feasibility of using cfDNA as a biomarker to monitor changes in chromatin organization and cell type composition over time in vivo for different clinical conditions.
[0431] D. Example 4: Detection of colon, breast, pancreatic, or liver cancer In this example, an artificial intelligence based approach is used to perform predictive analytics to analyze cfDNA data obtained from a subject (to generate an output of a diagnosis of the subject having cancer, e.g., colon cancer, breast cancer, or liver cancer or pancreatic cancer).
[0432] Retrospective human plasma samples were obtained from 937 patients diagnosed with colorectal cancer (CRC), 116 patients diagnosed with breast cancer, 26 patients diagnosed with liver cancer, and 76 patients diagnosed with pancreatic cancer. In addition, a set of 605 control samples was obtained from patients with no current cancer diagnosis (but potentially with other comorbidities or undiagnosed cancer), 127 of which had a negative confirmed colonoscopy. In total, samples were collected from 11 institutions and commercial biobanks from Southern and Northern Europe and the United States. All samples were de-identified.
[0433] Control samples in the CRC model include all samples except liver control samples (n=524). Control samples in the breast cancer model (n=123) included samples from the same institution contributing the breast cancer samples. The liver cancer samples were derived from a case-control study with 25 matched control samples. The control samples are effectively HBV positive but negative for cancer. The pancreatic cancer samples and matched controls were also obtained from a single institution. Of the 66 controls, 45 control samples have some non-cancerous pathology including pancreatitis, CBD stones, benign strictures, pseudocysts, etc.
[0434] The age, sex, and stage of cancer (when available) of each patient was obtained for each sample. Plasma samples collected from each patient were stored at -80°C and thawed before use.
[0435] Cell-free DNA was extracted from 250 μL of plasma (spiked with a unique synthetic double-stranded DNA (dsDNA) fragment for sample tracking) using the MagMAX cell-free DNA isolation kit (Applied Biosystems) according to the manufacturer's instructions. Paired-end sequencing libraries were prepared using the NEBNext Ultra II DNA Library Prep Kit (New England Biolabs) with polymerase chain reaction (PCR) amplification and unique molecular identifiers (UMIs), and at least 400 million reads (median = 636 million reads) were sequenced using an Illumina NovaSeq6000 sequencing system at 2x51 base pairs across multiple S2 or S4 flow cells, except for liver cancer samples, where at least 4 million reads (median = 2.8 million reads) were sequenced.
[0436] The resulting sequencing reads were demultiplexed, adapter trimmed, and aligned to the human reference genome (GRCh38 with decoys, alt contigs, and HLA contigs) using Burrows Wheeler aligner (BWA-MEM 0.7.15). PCR overlapping fragments were removed using fragment endpoints or unique molecular identifiers (UMIs), if present.
[0437] For all samples except the liver cancer experiment, sequencing data were checked for quality and excluded from further analysis if they met any of the following conditions: AT dropouts greater than about 10 (calculated via Piccard 2.10.5), GC dropouts greater than about 2 (calculated via Piccard 2.10.5), or sequencing depth less than about 10-fold. Additionally, samples whose relative counts at sex chromosomes did not match the annotated sex were removed from further processing and discarded. Additionally, any samples suspected of contamination (e.g., due to expected allele fractions less than about 0.99, unexpected genotype calls, or batches with contaminated negative controls) were manually inspected prior to inclusion in the dataset.
[0438] A cfDNA "profile" was created for each sample by counting the number of fragments that aligned to each putative protein-coding region of the genome. This type of data representation can capture at least two types of signals: (1) somatic CNVs (where gene regions provide a sampling of the genome, allowing capture of any consistent large-scale amplifications or deletions), and (2) epigenetic changes in the immune system represented in cfDNA, with variable nucleosome protection causing the observed changes in coverage.
[0439] A set of functional regions of the human genome, including putative protein-coding gene regions (with genomic coordinate coverage including both intra- and exons), was annotated to the sequencing data. Annotations of protein-coding gene regions ("gene" regions) were obtained from the Comprehensive Human Expressed Sequences (CHESS) project (v1.0). A feature set was generated from the annotated human genome regions, including vectors of cfDNA fragment counts corresponding to a set of genomic regions. The feature set was obtained by counting several cfDNA fragments with a mapping quality of at least 60 that overlapped with each of the annotated gene regions by at least one base, thereby generating a set of "gene features" (D=24,152, covering 1352 Mb) for each sample.
[0440] The count feature vectors were preprocessed via the following transformations. First, counts of cfDNA fragments corresponding to sex chromosomes were removed (only autosomes were kept). Second, counts of cfDNA fragments corresponding to low-quality genomic bins were removed. Third, features were normalized for their length. Low-quality genomic bins were identified by having either an average mappability across bins of less than about 0.75, a GC percentage of less than about 30% or more than about 70%, or a reference genome N content of more than about 10%. Fourth, depth normalization was performed for the number of cfDNA fragments. For each sample depth normalization, a trimmed average was generated by removing the bottom and top 10 percent of the bins before calculating the average of counts across bins in the sample, and the trimmed average was used as a scaling factor. A GC correction was applied to the cIDNA fragment counts, and a Loess regression correction was used to address GC bias. After these filtering transformations, the resulting vector of gene features had a dimension of 17,582 features covering 1172 Mb.
[0441] Cross-validation procedures can be performed as part of machine learning techniques to obtain an approximation of a model's performance on new prospectively collected unseen data. Such an approximation can be obtained by sequentially training a model on a subset of the data and testing it on a retained dataset that is unseen to the model being trained. A k-fold cross-validation procedure can be applied, which requires randomly stratifying all data into k groups (or folds) and testing each group on models fitted to the other folds. This approach can be a general and traceable way to estimate generalization performance. However, if the class labels are confounded with known covariates, such a "k-fold" cross-validation scheme can result in the problem of inflated performance that may not generalize to new datasets. The machine may simply learn to identify batches of labels and associated distributions. This can lead to misleading results and poor generalization, as the classifier learns erroneous associations between class labels and confounding factors in the training set and applies them inaccurately to the test set. Cross-validation performance may overe...
Claims
1. 1. A method of screening an individual for colorectal advanced adenoma, comprising: a) receiving result measurements from analyzing a plurality of classes of molecules in a biological sample from an individual using a plurality of assays; said analyzing providing a plurality of sets of measurements representative of said plurality of classes of molecules, said biological sample being whole blood, plasma, or serum; the plurality of classes of molecules comprises a first class of nucleic acids endogenous to the individual, and a second class of polyamino acids endogenous to the individual; The first class is cell-free DNA (cfDNA), a first assay is applied to the cfDNA molecules to obtain a first set of measurements; a second assay is applied to the polyamino acid to obtain a second set of measurements; the first assay applied to the cfDNA molecules comprises methylation sequencing; Receiving and b) identifying a set of features corresponding to properties of each of said classes of molecules that are input into a machine learning model; c) generating a feature vector of features from the sets of measurements representative of the classes of molecules, each feature corresponds to a feature of the set of features and includes one or more measurements, and the feature vector includes at least one feature obtained using each set of measurements of the plurality of sets representing a class of the plurality of molecules; To create and d) loading the machine learning model into a memory of a computer system, The machine learning model is trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having colorectal advanced adenomas, and a second subset of the training biological samples identified as not having colorectal advanced adenomas. Loading and e) inputting the feature vector into the machine learning model to obtain an output classification of whether the individual has advanced colorectal adenoma; Including, method.
2. The second class of polyamino acids is a peptide, a protein, an autoantibody, or a fragment thereof. The method of claim 1.
3. the plurality of classes of molecules includes the second class being an autoantibody and a third class being a circulating protein; The method of claim 1.
4. The plurality of assays comprises:
2. The method of claim 1, comprising at least two of the following assays: whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), enzymatic methyl sequencing, quantitative immunoassay, enzyme-linked immunosorbent assay (ELISA), protein microarray, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing, CNV calling, tumor fraction (TF) estimation, LINE-1 CpG methylation, 56 gene CpG methylation, cf-protein immunoquantitative ELISA, single molecule array (SIMOA), and cf-miRNA sequencing, and a mixture ratio of cell types or cell phenotypes derived from any of the above assays.
5. The plurality of assays comprises: The whole genome bisulfite sequencing or enzymatic methyl sequencing, including the methylation sequencing, The method of claim 4 comprising:
6. The machine learning model is 2. The method of claim 1, wherein the support vector machine (SVM) is trained and constructed according to one or more of: Linear Discriminant Analysis (LDA), Partial Least Squares (PLS), Random Forest, k-Nearest Neighbor (KNN), Support Vector Machine (SVM) with Radial Basis Function Kernel (SVMRadial), SVM with Linear Basis Function Kernel (SVMLinear), SVM with Polynomial Basis Function Kernel (SVMPoly), Decision Tree, Multilayer Perceptron, Mixture of Experts, Sparse Factor Analysis, Hierarchical Decomposition, and a combination of linear algebra routines and statistics.
7. 2. The method of claim 1, wherein the biological sample is a plasma sample and the measurements comprise methylation patterns of cell-free DNA found in the plasma sample.
8. A system comprising one or more computer processors and a computer memory coupled thereto, the computer memory comprising machine executable code which, when executed by the one or more computer processors, performs a method according to any one of claims 1 to 7. system.
9. A non-transitory computer readable storage medium comprising machine executable code which, when executed by one or more computer processors, implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Biomarker compositions and methods
JP2014522993A
Methods and machine learning systems for predicting the likelihood or risk of having cancer
US20180068083A1
Population based treatment recommender using cell free DNA
WO2017062867A1
Methods for fragmentome profiling of cell-free nucleic acids
WO2018009723A1
Method and system for microbial pharmacogenomics
WO2018013865A1