Machine learning implementation for multi-analyte assay of biological sample

By analyzing multiple classes of molecules in biological samples using machine learning, the method enhances cancer screening efficiency and accuracy, addressing compliance and noise issues in existing tests.

JP2025118815APending Publication Date: 2025-08-13FREENOME HOLDINGS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025078549
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-03-27
Filing Date
2025-05-09
Publication Date
2025-08-13

Smart Images

  • Figure 2025118815000001_ABST
    Figure 2025118815000001_ABST
Patent Text Reader

Abstract

To provide systems and methods that analyze blood-based cancer diagnostic tests using multiple classes of molecules.SOLUTION: A method comprises: a step 210 of receiving a biological sample including multiple classes of molecules; a step 220 of separating the biological sample into multiple portions; a step 230 of, for each of multiple assays, identifying a set of features to be input to a machine learning model; a step 240 of, for each of the multiple portions, performing an assay on a class of molecules in the portion to obtain a set of measured values of the class of molecules in the biological sample; a step 250 of forming a feature vector of feature values from the multiple sets of measured values; and a step 270 of inputting the feature vector into the machine learning model to obtain an output classification of whether the biological sample has a specified property.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is U.S. Provisional Patent Application No. US62 / 657,602, filed April 13, 2018; No. US62 / 749,955, filed October 24, 2018; No. US62 / 679,641, filed June 18, 2018; No. US62 / 767,435, filed November 14, 2018; No. US62 / 679,587, filed June 1, 2018; No. US62 / 731,557, filed September 14, 2018; No. US62 / 742,799, filed October 8, 2018; No. US62 / 804,614, filed February 2, 2019; No. US62 / 767,369, filed November 14, 2018; and This application claims the benefit of US 62 / 824,709, filed March 29, 2019, the entire contents of which are incorporated by reference.

[0002] Cancer screening is complex, and various cancer types require different approaches for screening and early detection. Patient compliance remains a problem. Screening methods requiring non-serum specimens often result in low participation rates. Screening rates for breast, cervical, and colorectal cancer using mammograms, Pap tests, and sigmoidoscopy / fecal occult blood tests, respectively, are far from the 100% compliance recommended by the United States Preventive Services Task Force (USPSTF) (Sabatino et al., Cancer Screening Test Use—United States, 2013, MMWR, 2015 64(17):464-468; Adler et al. BMC Gastroenterology 2014, 14:183). A recent report found that the percentage of adults eligible for colorectal cancer screening by state ranged from 58.5% (New Mexico) to 75.9% (Maine) in 2016, with an average of 67.3% (Joseph DA, et al. Use of Colorectal Cancer Screening Tests by State. Prev Chronic Dis 2018;15:170535).

[0003] Blood-based tests hold great promise for cancer diagnostics and precision medicine. However, most current tests are limited to analyzing a single class of molecules (e.g., circulating tumor DNA, platelet mRNA, circulating proteins). A wide range of biological analytes exist in blood for potential analysis, making the generation of relevant data critical. However, analyzing the entire analyte is laborious, uneconomical, and introduces significant biological noise relative to the useful signal, potentially confounding useful analyses for diagnostic or precision medicine applications.

[0004] Even with early detection and genomic characterization, there remain a significant number of cases in which genomic analysis fails to identify effective drugs or applicable clinical trials. Even when targetable genomic alterations are identified, patients do not always respond to treatment (Pauli et al., Cancer Discov. 2017, 7(5):462-477). Furthermore, sensitivity barriers exist for the use of circulating tumor DNA (ctDNA) in detection methods. ctDNA has recently been evaluated as a prospective specimen for early cancer detection, but it has been found that significant amounts of blood are required to detect ctDNA with the required specificity and sensitivity (Aravanis, A. et al., Next-Generation Sequencing of Circulating Tumor DNA for Early Cancer Detection, Cell, 168:571-574). Therefore, a simple, readily available single-analyte test remains elusive.

[0005] In the field of cancer diagnostics, machine learning can enable large-scale statistical approaches and automated characterization of signal intensities. However, machine learning applied to biology in the context of molecular diagnostics remains a largely unexplored field and has not previously been applied to aspects of diagnostics and precision medicine, such as specimen selection, assay selection, and overall optimization.

[0006] Therefore, there is a need for methods to analyze easily obtained biological specimens to stratify individuals at risk for or with cancer and to provide effective characterization of early-stage cancer to guide treatment decisions. There is also a need for methods to incorporate machine learning approaches with specimen datasets to develop and refine classifiers for use in stratifying populations of individuals and detecting diseases such as cancer. Summary of the Invention

[0007] Described herein are methods and systems that incorporate machine learning approaches using one or more biological analytes in a biological sample for various applications to stratify populations of individuals. In certain embodiments, the methods and systems are useful for predicting disease, treatment efficacy, and guiding treatment decisions for affected individuals.

[0008] This approach differs from other methods and systems in that it focuses on characterizing the non-cellular portion of the circulation, including specimens derived from tumor cells, healthy non-tumor cells induced or educated by the microenvironment, and circulating immune cells that may be educated by tumor cells present in an individual.

[0009] While other approaches focus on characterizing the cellular portion of the immune system, the present method and system investigates the non-cellular portion of the circulating cancer cell to provide biological information that can then be combined with machine learning tools for useful applications. Studying non-cellular specimens in liquid biological samples (e.g., plasma) allows for sample deconvolution to recapitulate the molecular state of an individual's tissues and immune cells in a living cellular context. Studying the non-cellular portion of the immune system provides a surrogate indicator of cancer status and avoids the requirement for significant blood volume to detect cancer cells and associated biological markers when screening for ctDNA alone.

[0010] In a first aspect, the present disclosure provides a method of using a classifier capable of distinguishing between populations of individuals, comprising: a) assaying multiple classes of molecules in a biological sample, the assay providing multiple sets of measurements representative of the multiple classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules to be input into a machine learning or statistical model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of the computer system a machine learning model including a classifier, the machine learning model trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified as having a specified property, and a second subset of the training biological samples identified as not having the specified property; e) inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample has the specified property, thereby identifying a population of individuals having the specified property.

[0011] By way of example, the class of molecules can be selected from nucleic acids, polyamino acids, carbohydrates, or metabolites. As a further example, the class of molecules can include nucleic acids, including deoxyribonucleic acid (DNA), genomic DNA, plasmid DNA, complementary DNA (cDNA), cell-free (e.g., non-encapsulated) DNA (cfDNA), circulating tumor DNA (ctDNA), nucleosomal DNA, chromatosomal DNA, mitochondrial DNA (miDNA), artificial nucleic acid analogs, recombinant nucleic acids, plasmids, viral vectors, and chromatin. In one example, the sample includes cfDNA. In one example, the sample includes genomic DNA derived from peripheral blood mononuclear cells (PBMCs).

[0012] As a further example, classes of molecules can include nucleic acids, including ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (mitoRNA), ribosomal RNA (rRNA), circular RNA (cRNA), alternatively spliced mRNA, small nuclear RNA (snRNA), antisense RNA, short hairpin RNA (shRNA), or small interfering RNA (siRNA).

[0013] As a further example, the class of molecules can include polyamino acids, including polyamino acids, peptides, proteins, autoantibodies, or fragments thereof.

[0014] As further examples, classes of molecules can include sugars, lipids, amino acids, fatty acids, phenolic compounds, or alkaloids.

[0015] In various embodiments, the multiple classes of molecules include at least two of cfDNA molecules, cfRNA molecules, circulating proteins, antibodies, and metabolites.

[0016] As with aspects of the present disclosure, and various examples of the systems and methods herein, the classes of molecules can be selected from 1) cfDNA, cfRNA, polyamino acids and small chemical molecules, or 2) cfDNA and cfRNA and polyamino acids, 3) cfDNA and cfRNA and small chemical molecules, or 4) cfDNA, polyamino acids and small chemical molecules, or 5) cfRNA, polyamino acids and small chemical molecules, or 6) cfDNA and cfRNA, or 7) cfDNA and polyamino acids, or 8) cfDNA and small chemical molecules, or 9) cfRNA and polyamino acids, or 10) cfRNA and small chemical molecules, or 11) polyamino acids and small chemical molecules.

[0017] In one example, the multiple classes of molecules are cfDNA, proteins, and autoantibodies.

[0018] In various examples, the plurality of assays can include whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), small RNA sequencing, quantitative immunoassay, enzyme-linked immunosorbent assay (ELISA), proximity extension assay (PEA), protein microarray, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing (WO2019 / 051484), CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56-gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing, as well as at least two of the cell type or cell phenotype mixture ratios derived from any of the above assays.

[0019] In one embodiment, whole genome bisulfite sequencing includes methylation analysis.

[0020] In various embodiments, classification of the biological sample is performed by a classifier trained and constructed according to one or more of Linear Discriminant Analysis (LDA), Partial Least Squares (PLS), Random Forest, k-Nearest Neighbor (KNN), Support Vector Machine (SVM) with a Radial Basis Function Kernel (SVMRadial), SVM with a Linear Basis Function Kernel (SVMLinear), SVM with a Polynomial Basis Function Kernel (SVMPoly), Decision Trees, Multilayer Perceptron, Mixture of Experts, Sparse Factor Analysis, Hierarchical Decomposition, and a combination of linear algebraic routines and statistics.

[0021] In various examples, the specified property may be a clinically diagnosed disorder. The clinically diagnosed disorder may be cancer. By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer. In some examples, the specified property is responsiveness to treatment. In one example, the specified property may be a continuous measure of a patient's trait or phenotype.

[0022] In a second aspect, the present disclosure provides a system for classifying a biological sample, comprising: a) a receiver that receives a plurality of training samples, each of the plurality of training samples having a plurality of classes of molecules, each of the plurality of training samples including one or more known labels; b) a feature module that identifies, for each of a plurality of training samples, a set of features corresponding to an assay operable to be input to the machine learning model, the set of features corresponding to properties of molecules in the plurality of training samples; a feature module operable for each of a plurality of training samples to subject a plurality of classes of molecules in the training sample to a plurality of different assays to obtain sets of measurements, each set of measurements being from one assay applied to a class of molecules in the training sample, and the plurality of sets of measurements being obtained for the plurality of training samples; c) an analysis module for analyzing the set of measurements to obtain training vectors for the training samples, the training vectors including N sets of feature quantities of features of the corresponding assays, each feature quantity corresponding to a feature and including one or more measurements, the training vectors being formed using at least one feature from at least two of the N sets of features corresponding to a first subset of the plurality of different assays; d) a label module that uses parameters of the machine learning model to inform the system of the training vectors to obtain output labels for a plurality of training samples; e) a comparator module for comparing the output labels with known labels of the training samples; f) a training module that iteratively searches for optimal values for parameters as part of training the machine learning model based on comparing the output labels with known labels of training samples; g) an output module that provides machine learning model parameters and a set of machine learning model features.

[0023] In a third aspect, the present disclosure provides a system for classifying a subject based on a multi-analyte analysis of a biological sample composition, comprising: (a) a computer-readable medium comprising a classifier operable to classify the subject based on the multi-analyte analysis; and (b) one or more processors for executing instructions stored on the computer-readable medium.

[0024] In one embodiment, the system includes a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and an NMF prediction algorithm classifier.

[0025] In one embodiment, the system includes means for performing any of the aforementioned methods. In one embodiment, the system includes one or more processors configured to perform any of the aforementioned methods. In one embodiment, the system includes modules for performing each of the steps of any of the aforementioned methods.

[0026] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, implements the above or any other method herein.

[0027] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or other herein.

[0028] In a fourth aspect, the present disclosure provides a method of detecting the presence of cancer in an individual, comprising: a) assaying a plurality of classes of molecules in a biological sample obtained from an individual, the assay providing a plurality of sets of measurements representative of the plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules to be input into the machine learning model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into the memory of the computer system a machine learning model that is trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals with cancer, and a second subset of the training biological samples identified from individuals without cancer; e) detecting the presence of cancer in the individual by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with cancer.

[0029] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value being indicative of the presence of cancer in the individual.

[0030] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value indicating a stage of cancer in the individual.

[0031] By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer or breast cancer, hi one embodiment, the cancer is colon cancer.

[0032] In a fifth aspect, the present disclosure provides a method of determining a prognosis of an individual with cancer, comprising: a) assaying multiple classes of molecules in a biological sample, the assay providing multiple sets of measurements representative of the multiple classes of molecules; b) identifying a set of features corresponding to properties of multiple classes of molecules to be input into a machine learning model; creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; c) loading into the memory of the computer system a machine learning model trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals with a favorable cancer prognosis, and a second subset of the training biological samples identified from individuals without a favorable cancer prognosis; d) determining a prognosis for an individual with cancer by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with a good cancer prognosis.

[0033] By way of example, the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer or breast cancer.

[0034] In a sixth aspect, the present disclosure provides a method for determining responsiveness to cancer treatment, comprising: a) assaying multiple classes of molecules in a biological sample, the assay providing multiple sets of measurements representative of the multiple classes of molecules; b) identifying a set of features corresponding to properties of each of a plurality of classes of molecules to be input into the machine learning model; creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; c) loading into the memory of the computer system a machine learning model trained using training vectors obtained from the training biological samples, a first subset of the training biological samples identified from individuals who respond to the treatment, and a second subset of the training biological samples identified from individuals who do not respond to the treatment; d) determining responsiveness to cancer treatment by inputting the feature vector into a machine learning model to obtain an output classification of whether the biological sample is associated with treatment response.

[0035] In one embodiment, the cancer therapy is selected from an alkylating agent, a plant alkaloid, an antitumor antibiotic, an antimetabolite, a topoisomerase inhibitor, a retinoid, a checkpoint inhibitor therapy, or a VEGF inhibitor.

[0036] In one embodiment, the method includes combining the classification data from the classifier analysis to provide a detection value, the detection value being indicative of a response to treatment in the individual.

[0037] These and other embodiments are described in more detail below. For example, other embodiments are directed to systems, devices, and computer-readable media associated with the methods described herein.

[0038] The nature and advantages of embodiments of the present disclosure may be better understood by reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0039] [Figure 1] FIG. 1 illustrates an exemplary system that is programmed or otherwise configured to implement the methods provided herein. [Figure 2] FIG. 2 is a flow chart illustrating a method for analyzing a biological sample. [Figure 3] FIG. 3 illustrates an overall framework in accordance with various embodiments. [Figure 4] Figure 4 shows an overview of the multi-analyte approach. [Figure 5] FIG. 5 illustrates an iterative process for designing an assay and corresponding machine learning model, according to various embodiments. [Figure 6] FIG. 6 is a flow chart illustrating a method for classifying a biological sample, according to one embodiment. [Figure 7] 7A and 7B show the classification performance for different analytes. [Figure 8A] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8B] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8C] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8D] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8E] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8F] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8G] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 8H] Figures 8A-8H show the distribution of tumor fraction cfDNA samples for individuals with high (>20%) tumor fraction based on cfDNA sequencing data. [Figure 9] Figure 9 shows CpG methylation analysis at LINE-1 sites. [Figure 10] FIG. 10 shows cf-miRNA sequence analysis. [Figure 11A] FIG. 11A shows circulating protein biomarker distribution. [Figure 11B] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 11C] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 11D] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 11E] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 11F] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 11G] Figures 11B-11G show that protein levels differ significantly across tissue types by one-way analysis of variance followed by Sidak's multiple comparison test. [Figure 12A] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12B] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12C] Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12D]Figures 12A-D show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of tumor fraction. [Figure 12E] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12F] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12G] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 12H] Figures 12E-H show PCA of cfDNA, CpG methylation, cf-miRNA and protein counts as a function of patient diagnosis. [Figure 13] Figure 13 shows a heatmap of chromosome structure scores determined from nuanced structure of correlation matrices generated using Pearson / Spearman / Kendall correlation of regions of the genome using cfDNA samples. [Figure 14] FIG. 14 shows a heatmap of chromosome structure scores determined from Hi-C sequencing of the same genomic region as in FIG. [Figure 15A-C] Figures 15A-C show correlation maps generated from Hi-C, spatially correlated fragment lengths from multiple cfDNA samples, and spatially correlated fragment length distributions from a single cfDNA sample. [Figure 15D] Figure 15D shows genome browser tracks of partition A / B from Hi-C, multi-sample cfDNA, and single-sample cfDNA. [Figure 15E] Figures 15E-15F show scatter plots of compartment-level concordance between Hi-C, multi-sample cfDNA (Figure 15E), and single-sample cfDNA (Figure 15F). [Figure 15F] Figures 15E-15F show scatter plots of compartment-level concordance between Hi-C, multi-sample cfDNA (Figure 15E), and single-sample cfDNA (Figure 15F). [Figure 16A] Figure 16A shows the correlation between Hi-C and cfHi-C at the pixel level (500 kb bins). [Figure 16B] Figure 16B shows the correlation between Hi-C and cfHi-C at the compartment level (500 kb bins). [Figures 17A-D] Figure 17A shows a heatmap of cfHi-C before G+C% is regressed from the fragment length of each bin on chr1 using LOWESS. Figure 17B shows a heatmap of cfHi-C after G+C% is regressed from the fragment length of each bin on chr1 using LOWESS. Figure 17C shows a heatmap of gDNA before G+C% is regressed from the fragment length of each bin on chr1 using LOWESS. Figure 17D shows a heatmap of gDNA after G+C% is regressed from the fragment length of each bin on chr1 using LOWESS. [Figure 17E] Figure 17E shows boxplots of pixel-level correlations (Pearson and Spearman) with Hi-C (WBC, repeat 2) across all chromosomes represented in Figures 17A-17D. [Figure 18A] FIG. 18A shows G+C % and mappability bias analysis in two-dimensional space from multi-sample cfHi-C. [Figure 18B] Figure 18B shows G+C % and mappability bias analysis in two-dimensional space from a single sample cfHi-C. [Figure 18C] FIG. 18C shows G+C % and mappability bias analysis in two-dimensional space from multiple sample genomic DNA. [Figure 18D] FIG. 18D shows G+C % and mappability bias analysis in two-dimensional space from a single sample of genomic DNA. [Figure 18E] FIG. 18E shows G+C % and mappability bias analysis in two-dimensional space from multi-sample cfHi-C. [Figure 18F] FIG. 18F shows G+C % and mappability bias analysis in two-dimensional space from Hi-C (WBC). [Figure 19A] Figure 19A shows a heatmap of multi-sample cfHi-C where one pair of bins was randomly shuffled from any other individual (chr14). [Figure 19B] Figure 19B shows a multi-sample cfHi-C heatmap for samples from the same batch as Figure 19A (11 samples; chr14). [Figure 19C] Figure 19C shows a heatmap of multi-sample cfHi-C for samples with the same sample size as Figure 19B (11 samples; chr14). [Figure 19D] Figure 19D shows a boxplot of pixel-level correlation with Hi-C (WBC, repeat 2) across all chromosomes represented in Figures 19A-19C. [Figure 20A] FIG. 20A shows the Pearson correlation between Hi-C (WBC, replicate 1) and multi-sample cfHi-C at different sample sizes. [Figure 20B] FIG. 20B shows the Spearman correlation between Hi-C (WBC, replicate 1) and multi-sample cfHi-C at different sample sizes. [Figure 20C] FIG. 20C shows the Pearson correlation between Hi-C (WBC, replicate 2) and multi-sample cfHi-C at different sample sizes. [Figure 20D] FIG. 20D shows the Spearman correlation between Hi-C (WBC, replicate 2) and multi-sample cfHi-C at different sample sizes. [Figure 21A] FIG. 21A shows pixel-level Pearson correlation between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21B] FIG. 21B shows pixel-level Spearman correlation between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21C] FIG. 21C shows pixel-level Pearson correlation between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 21D]FIG. 21D shows the pixel-level Spearman correlation between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 21E] FIG. 21E shows Pearson correlations at the compartment level between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21F] FIG. 21F shows the Spearman correlation at the compartment level between Hi-C and multi-sample cfHi-C at different bin sizes. [Figure 21G] FIG. 21G shows Pearson correlations at the compartment level between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 21H] FIG. 21H shows the Spearman correlation at the compartment level between Hi-C and single-sample cfHi-C at different bin sizes. [Figure 22A] Figure 22A shows pixel-level Pearson and Spearman correlations between Hi-C and single-sample cfHi-C at different read numbers after downsampling. [Figure 22B] Figure 22B shows Pearson and Spearman correlations at the compartment level between Hi-C and single-sample cfHi-C at different read numbers after downsampling. [Figure 23A] FIG. 23A shows kernel PCA (RBF kernel) of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23B] Figures 23B-23F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23C] Figures 23B-23F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23D] Figures 23B-23F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23E] Figures 23B-23F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 23F]Figures 23B-23F show the CCA of healthy samples and high tumor fraction samples from colon cancer, lung cancer, and melanoma. [Figure 24] FIG. 24 shows the correlation map between DNA accessibility and compartment-level eigenvalues from Hi-C from the same cell type (GM12878). [Figure 25-1] Figure 25A shows a heat map of cellular composition inferred from single-sample cfDNA of healthy, colon cancer, lung cancer, and melanoma samples. Figure 25B shows a pie chart of cellular composition inferred from single-sample cfDNA of healthy, colon cancer, lung cancer, and melanoma samples. [Figure 25-2] Figure 25C shows box plots of leukocyte and tumor fractions inferred from single-sample cfDNA from 100 healthy individuals. [Figure 26] Figure 26 shows a comparison of ichorCNA-derived tumor fractions with cfHi-C-derived tumor fractions by using only genomic regions without CNV alterations for lung, melanoma, and colon cancers. [Figure 27] Figure 27A shows the k-fold, k-batch, balanced k-batch, and ordered k-batch training schemes, and Figure 27B shows the k-batch with systematic downsampling scheme. [Figure 28A] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28B] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28C] 28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28D]28A-28D show example receiver operating characteristic (ROC) curves for all validation approaches evaluated for cancer detection (e.g., k-fold, k-batch, balanced k-batch, and ordered k-batch). [Figure 28E] Figure 28E shows sensitivity by CRC stage across all validation approaches evaluated. [Figure 28F] Figure 28F shows the AUC by IchorCNA estimated tumor fraction across all validation approaches evaluated. [Figure 28G] Figure 28G shows the AUC by age bin across all validation approaches evaluated. [Figure 28H] Figure 28H shows the AUC by gender bin across all validation approaches evaluated. [Figure 29A] 29A-29B show classification performance in cross-validation (ROC curves) for breast cancer. [Figure 29B] 29A-29B show classification performance in cross-validation (ROC curves) for breast cancer. [Figure 29C] 29C to 29D show the classification performance in cross-validation (ROC curve) for liver cancer. [Figure 29D] 29C to 29D show the classification performance in cross-validation (ROC curve) for liver cancer. [Figure 29E] 29E-29F show classification performance in cross-validation (ROC curve) for pancreatic cancer. [Figure 29F] 29E-29F show classification performance in cross-validation (ROC curve) for pancreatic cancer. [Figure 30] FIG. 30 shows the distribution of estimated tumor fractions (TF) by class. [Figure 31A] Figure 31A shows the AUC performance of CRC classification when the training set for each fold is downsampled as either a proportion of the samples. [Figure 31B] Figure 31B shows the AUC performance of CRC classification when the training set for each fold is downsampled as either a proportion of the samples or a proportion of the batch. [Figure 32A] Figures 32A-C show examples of healthy samples with high tumor fractions. [Figure 32B] Figures 32A-C show examples of healthy samples with high tumor fractions. [Figure 32C] Figures 32A-C show examples of healthy samples with high tumor fractions. [Figure 33] Figure 33A shows the training method and cross-validation procedure for the k-fold model. Figure 33B shows the k-fold, k-batch, and balanced k-batch training schemes. [Figure 34A] FIG. 34A shows sensitivity by CRC stage in patients aged 50-84. [Figure 34B] Figure 34B shows sensitivity by tumor fraction in patients aged 50-84. [Figure 34C] Figure 34C shows the AUC performance of CRC classification among the total number of samples. [Figure 35] Figure 35 shows a schematic of V-plots derived from cfDNA-captured protein-DNA associations, indicating chromatin structure and transcriptional state. TF = transcription factor (small protected footprint region), NS = nucleosome (large protected region, complete wrap of DNA). [Figure 36] Figure 36 shows V-plots from cfDNA surrounding the TSS region used to predict gene expression. [Figure 37] Figure 37 shows a classifier using fragment length and position representations that accurately classifies on and off genes using different cutoffs. [Figure 38A] Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. The tumor fraction estimate (ITF) based on IchorCNA increases with stage, but most stage I-III CRCs have low estimated ITF (<1%) (Figure 38A). Performance increases with stage, with the most significant increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 38B]Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. The tumor fraction estimate (ITF) based on IchorCNA increases with stage, but most stage I-III CRCs have low estimated ITF (<1%) (Figure 38A). Performance increases with stage, with the most significant increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 38C] Figures 38A-38C show classification accuracy using tumor target genes set by stage and estimated tumor fraction. The tumor fraction estimate (ITF) based on IchorCNA increases with stage, but most stage I-III CRCs have low estimated ITF (<1%) (Figure 38A). Performance increases with stage, with the most significant increase in stage IV (Figure 38B). Performance increases most strongly with tumor fraction (Figure 38C). [Figure 39A] FIG. 39A shows tumor fraction estimates versus 44 colon gene mean P(on). [Figure 39B] Figure 39B shows the fold change from the mean coverage for healthy samples containing strong evidence of copy number alterations in chr8 and chr9.

[0040] term The terms "a," "an," and "the" are intended to mean "one or more" unless specifically indicated to the contrary. The use of "or" is intended to mean "inclusive or" rather than "exclusive or" unless specifically indicated to the contrary. A reference to a "first" element does not necessarily require that a second element be provided. Further, unless specifically stated, a reference to a "first" or "second" element does not limit the referenced element to a particular location. The term "based on" is intended to mean "based at least in part on."

[0041] The term "area under the curve" or "AUC" refers to the area under a receiver operating characteristic (ROC) curve. AUC measures are useful for comparing the accuracy of classifiers across the complete data range. A classifier with a larger AUC has a greater ability to correctly classify unknowns between two groups of interest (e.g., cancer samples and normal or control samples). ROC curves are useful for plotting the performance of a particular feature (e.g., any of the biomarkers described herein and / or any item of additional biomedical information) in distinguishing two populations (e.g., individuals who respond to a therapeutic agent and those who do not). Typically, feature data across populations (e.g., cases and controls) are sorted in ascending order based on the value of a single feature. Then, for each value of that feature, true positive and false positive rates are calculated for the data. The true positive rate is determined by counting the number of cases above the value of that feature and dividing by the total number of cases. The false positive rate is determined by counting the number of controls above the value of that feature and dividing by the total number of controls. While this definition refers to scenarios in which a feature is elevated in cases compared to controls, it also applies to scenarios in which a feature is lower compared to controls (in such scenarios, samples below the value of that feature may be counted). ROC curves can be generated for single features as well as other single outputs; for example, a combination of two or more features can be mathematically combined (e.g., by addition, subtraction, multiplication, etc.) to provide a single total value, which can be plotted on an ROC curve. Additionally, any combination of multiple features whose combination yields a single output value can be plotted on an ROC curve. These combinations of features may include tests. An ROC curve is a plot of the true positive rate (sensitivity) of a test against the false positive rate (specificity) of the test.

[0042] The term "biological sample" (or simply "sample") refers to any substance obtained from a subject. A sample may contain or be suspected to contain an analyte from a subject, such as those described herein (nucleic acids, polyamino acids, carbohydrates, or metabolites). In some embodiments, a sample may include cells and / or acellular material obtained in vivo, cultured in vitro, or processed in situ, as well as lineages, including lineages and genealogies. In various embodiments, a biological sample may be tissue (e.g., solid or liquid tissue), such as normal or healthy tissue from a subject. Examples of solid tissue include primary tumors, metastatic tumors, polyps, or adenomas. Examples of liquid samples (e.g., bodily fluids) include whole blood, blood-derived buffy coat (which may contain lymphocytes), urine, saliva, cerebrospinal fluid, plasma, serum, ascites, sputum, sweat, tears, buccal samples, cavity rinses, or organ rinses. In some cases, the fluid is acellular, is an essentially acellular fluid sample, or contains acellular nucleic acids, e.g., acellular DNA. In some cases, cells, including circulating tumor cells, can be enriched or isolated from the fluid.

[0043] The terms "cancer" and "cancerous" refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. Neoplasia, malignancy, cancer, and tumor are often used interchangeably and refer to the abnormal growth of tissue or cells that results from excessive cell division.

[0044] The term "cancer-free" refers to a subject whose organs have not been diagnosed with cancer or have no detectable cancer.

[0045] The term "genetic variant" (or "variant") refers to a deviation from one or more expected values. Examples include sequence variants or structural polymorphisms. In various embodiments, a variant can refer to a known variant, such as a variant that has been scientifically confirmed and reported in the literature, a putative variant associated with biological change, a putative variant that has been reported in the literature but not yet biologically confirmed, or a putative variant that has not been reported in the literature but is inferred based on computational analysis.

[0046] The term "germline variant" refers to a nucleic acid that induces a natural or normal polymorphism (e.g., skin color, hair color, and normal weight). Somatic mutations can refer to nucleic acids that induce acquired or abnormal polymorphisms (e.g., cancer, obesity, conditions, diseases, disorders, etc.). Germline variants are inherited and therefore correspond to the genetic differences of an individual born relative to the reference human genome. Somatic variants are variants that arise at any time during cell division, development, and aging, at zygotes or later. In some examples, analysis can distinguish between germline variants, e.g., private variants, and somatic mutations.

[0047] The term "input feature" (or "feature") refers to a variable used by a model to predict the output classification (label) of a sample, e.g., a status, sequence content (e.g., mutations), recommended data collection actions, or recommended treatment. Values of the variables can be determined for a sample and used to determine the classification. Examples of input features for genetic data include aligned variables related to the alignment of sequence data (e.g., sequence reads) to a genome, and unaligned variables related, for example, to the sequence content of sequence reads, protein or autoantibody measurements, or average methylation levels in a genomic region.

[0048] The term "machine learning model" (or "model") refers to a set of parameters and functions where the parameters are trained on a set of training samples. The parameters and functions can be a set of linear algebraic operations, nonlinear algebraic operations, and tensor algebraic operations. The parameters and functions can include statistical functions, tests, and probability models. The training samples can correspond to samples with measured characteristics of the samples (e.g., genomic data and other subject data such as images or health records) and known classifications / labels of the subjects (e.g., phenotypes or treatments). The model can learn from the training samples in a training process that optimizes the parameters (and potentially the functions) to provide optimal quality metrics (e.g., accuracy) for classifying new samples. Training functions can include Bayesian parameter estimation methods such as expectation maximization, maximum likelihood, Markov chain Monte Carlo, Gibbs sampling, Hamiltonian Monte Carlo, and variational inference, or gradient-based methods such as stochastic gradient descent and the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm. Exemplary parameters include, for example, weights (e.g., vector or matrix transforms) that multiply values in regression or neural networks, probability distribution families, or loss, cost, or objective functions that assign scores and guide model training. Exemplary parameters include, for example, weights that multiply values in regression or neural networks. A model can include multiple submodels, may be models with different layers or independent models, and may have different structural forms, for example, a combination of a neural network and a support vector machine (SVM). Examples of machine learning models include deep learning models, neural networks (e.g., deep learning neural networks), kernel-based regression, adaptation-based regression or classification, Bayesian methods, ensemble methods, logistic regression and augmentation, Gaussian processes, support vector machines (SVMs), probabilistic models, and probabilistic graphical models.Machine learning models can further include feature engineering (e.g., collecting features into data structures such as one-dimensional, two-dimensional, or larger dimensional vectors) and feature representation (e.g., processing the feature data structure into transformed features used for training for classification inference).

[0049] A "marker" or "marker protein" is a diagnostic indicator found in a patient and is detected directly or indirectly by the method of the present invention. Indirect detection is preferred. In particular, all of the markers of the present invention have been shown to cause the production of (auto)antigens in cancer patients or patients at risk of developing cancer. Therefore, a simple way to detect these markers is to detect these (auto)antibodies in blood or serum samples from the patient. Such antibodies can be detected by binding to their respective antigens in an assay. Such antigens are, in particular, the marker proteins themselves or antigenic fragments thereof. Suitable methods can be used to specifically detect such antibody-antigen reactions and can be used in accordance with the system and method of the present disclosure. Preferably, the entire antibody content of the sample is normalized (e.g., diluted to a predetermined concentration) and applied to the antigen. Preferably, IgG, IgM, IgD, IgA, or IgE antibody fractions are used exclusively. A preferred antibody is IgG.

[0050] The term "non-cancerous tissue" refers to tissue from the same organ in which a malignant neoplasm has formed but which does not have the pathology characteristic of the neoplasm. Generally, non-cancerous tissue appears histologically normal. As used herein, "normal tissue" or "healthy tissue" refers to tissue from an organ in which the organ is not cancerous.

[0051] The terms "polynucleotide," "nucleotide," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, minimally linked together in length. In some examples, polynucleotides can have any three-dimensional structure and perform any function, known or unknown. Nucleic acids can include RNA, DNA, such as genomic DNA, mitochondrial DNA, viral DNA, synthetic DNA, cDNA reverse transcribed from RNA, bacterial DNA, viral DNA, and chromatin. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers, and can also be single-base nucleotides. In some examples, a polynucleotide comprises modified nucleotides, such as methylated or glycosylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. In some examples, the sequence of nucleotides is interrupted by non-nucleotide components. In certain examples, a polynucleotide is further modified after polymerization, such as by conjugation with a labeling component.

[0052] The term "polypeptide" or "protein" or "peptide" is intended to encompass, among other things, naturally occurring proteins as well as recombinantly or synthetically produced proteins. Note that the term "polypeptide" or "protein" may include naturally occurring modified forms of proteins, such as glycosylated forms. As used herein, the term "polypeptide" or "protein" or "peptide" is intended to encompass any amino acid sequence and includes modified sequences such as glycoproteins.

[0053] The term "prediction" is used herein to refer to the likelihood, probability, or score that a patient will respond favorably or unfavorably to a drug or set of drugs, and the degree of their response, as well as the detection of disease. The exemplary predictive methods of the present disclosure can be used clinically to make treatment decisions by selecting the most appropriate treatment modality for any particular patient. The predictive methods of the present disclosure are valuable tools for predicting whether a patient is likely to respond well to a treatment regimen, such as surgical intervention, chemotherapy with a given drug or drug combination, and / or radiation therapy.

[0054] The term " prognosis " as used herein refers to the likelihood of the clinical outcome of the subject suffering from a particular disease or disorder.With respect to cancer, prognosis is the expression of the possibility (probability) that the subject will survive (for example, 1, 2, 3, 4 or 5 years) and / or the possibility (probability) that tumor will metastasize.

[0055] The term "specificity" (also called true negative rate) refers to a measure of the proportion of actual negatives that are correctly identified as such (e.g., the proportion of healthy individuals that are correctly identified as not having the condition). Specificity is a function of the number of true negative calls (TN) and false positive calls (FP). Specificity is measured as (TN) / (TN+FP).

[0056] The term "sensitivity" (also called true positive rate, or probability of detection) refers to a measure of the proportion of actual positives that are correctly identified as such (e.g., the proportion of sick people who are correctly identified as having a disease state). Sensitivity is a function of the number of true positive calls (TP) and false negative calls (FN). Sensitivity is measured as (TP) / (TP+FN).

[0057] The term "structural polymorphism (SV)" refers to a DNA region that is approximately 50 bp and differs from the reference genome in size. Examples of SV include inversions, translocations, and copy number variants (CNVs), such as insertions, deletions, and amplifications.

[0058] The term "subject" refers to a biological entity containing genetic material. Examples of biological entities include plants, animals, or microorganisms, including, for example, bacteria, viruses, fungi, and protozoa. In some examples, the subject is a mammal, e.g., a human, and can be male or female. Such humans can be of various ages, e.g., from 1 day to about 1 year old, from about 1 year to about 3 years old, from about 3 years to about 12 years old, from about 13 years to about 19 years old, from about 20 years to about 40 years old, from about 40 years to about 65 years old, or over 65 years old. In various examples, the subject can be healthy or normal, abnormal, or diagnosed with or suspected to be at risk for a disease. In various examples, the disease includes cancer, a disorder, a condition, a syndrome, or any combination thereof.

[0059] The term "training sample" refers to a sample whose classification may be known. The training sample can be used to train a model. The feature values of the sample can form an input vector, e.g., a training vector of the training sample. Each element of the training vector (or other input vector) can correspond to a feature that includes one or more variables. For example, the elements of the training vector can correspond to a matrix. The label values of the sample can form a vector that includes strings, numbers, byte codes, or any collection of the aforementioned data types of any size, dimension, or combination.

[0060] As used herein, the terms "tumor," "neoplasia," "malignancy," or "cancer" generally refer to the growth and proliferation of neoplastic cells, and all pre-cancerous and cancerous cells and tissues, and the results of abnormal and unregulated growth of cells, whether malignant or benign.

[0061] The term "tumor burden" refers to the amount of tumor in an individual and can be measured as the number, volume, or weight of tumors. Tumors that do not metastasize are called "benign." Tumors that can invade surrounding tissues and / or metastasize are called "malignant."

[0062] The term "nucleic acid sample," as used herein, encompasses a "nucleic acid library" or "library," including a nucleic acid library prepared by any suitable method. The adapter may anneal to a PCR primer to facilitate amplification by PCR, or may be a universal primer region, such as a sequencing tail adapter. The adapter may be a universal sequencing adapter. As used herein, the term "efficiency" may refer to a measurable metric calculated as the ratio of the number of unique molecules whose sequences may be available after sequencing over the number of unique molecules originally present in the primary sample. In addition, the term "efficiency" may also refer to reducing the initial nucleic acid sample material required, reducing sample preparation time, reducing the amplification process, and / or reducing the overall cost of nucleic acid library preparation.

[0063] As used herein, the term "barcode" refers to a known sequence used to associate a polynucleotide fragment with the input or target polynucleotide from which it is generated. The barcode sequence can be a sequence of synthetic or natural nucleotides. The barcode sequence can be contained within an adapter sequence such that the barcode sequence is contained in a sequencing read. Each barcode sequence can be at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more nucleotides in length. In some cases, barcode sequences can be sufficiently long and sufficiently different from each other to allow for sample identification based on the barcode sequence to which they are associated. In some cases, barcode sequences are used to tag and subsequently identify "original" nucleic acid molecules (nucleic acid molecules present in a sample derived from a subject). In some cases, a barcode sequence or a combination of barcode sequences is used in combination with endogenous sequence information to identify the original nucleic acid molecule. For example, a barcode sequence (or a combination of barcode sequences) can be used in conjunction with endogenous sequence flanking the barcode (e.g., the beginning and end of the endogenous sequence), and / or the length of the endogenous sequence.

[0064] In some examples, nucleic acid molecules used herein may be subjected to a "tagmentation" or "ligation" reaction. "Tagmentation" combines a fragmentation reaction and a ligation reaction into a single step in the library preparation process. Tagged polynucleotide fragments are "tagged" with transposon end sequences during tagmentation and may further contain additional sequences added during extension over several cycles of amplification. Alternatively, biological fragments can be "tagged" directly to process nucleic acid molecules or fragments thereof, which may include nucleic acid amplification. For example, any type of nucleic acid amplification reaction can be used to amplify target nucleic acid molecules or fragments thereof to generate amplification products. DETAILED DESCRIPTION OF THE INVENTION

[0065] Methods and systems are provided for detecting analytes in biological samples, measuring various metrics of the analytes, and inputting the metrics as features into machine learning models to train classifiers for medical diagnostics. The trained classifiers generated using the methods described herein are useful for multiple approaches, including disease detection and staging, identifying treatment responders, and stratifying patient populations in need thereof.

[0066] Provided herein are methods and systems that incorporate machine learning approaches into one or more biological analytes in a biological sample for various applications, such as stratifying individual populations. Methods and systems are provided for detecting analytes in a biological sample, measuring various metrics of the analytes, and inputting the metrics as features into a machine learning model to train a classifier for medical diagnosis. Trained classifiers generated using the methods described herein are useful for multiple approaches, including disease detection and staging, identifying treatment responders, and stratifying patient populations in need thereof. In certain embodiments, the methods and systems are useful for predicting disease, treatment efficacy, and guiding treatment decisions in affected individuals.

[0067] This approach differs from other methods and systems in that it focuses on characterizing the noncellular portion of the circulating immune system, although cellular portions can also be used. The process of hematopoietic turnover is the natural death and lysis of circulating immune cells. The plasma fraction of blood contains an enriched sample of immune system fragments at the time when cells die and release their intracellular contents into the circulation. Specifically, plasma provides an information-rich biological specimen that reflects the population of immune cells conditioned by the presence of cancer cells before clinical symptoms appear. While other approaches focus on characterizing the cellular portion of the immune system, this method investigates the noncellular portion of the immune system conditioned by cancer to provide biological information that can then be combined with machine learning tools for useful applications. Studying noncellular specimens in liquids, such as plasma, allows for deconvolution of the liquid sample to recapitulate the molecular state of immune cells when they were alive. Studying the noncellular portion of the immune system provides a surrogate indicator of cancer status and avoids the need for significant blood volumes to detect cancer cells and associated biological markers.

[0068] I. Circulating specimens and cell deconstruction by biological assays For health-related or biological predictions (e.g., drug resistance / susceptibility prediction) based entirely or partially on the diagnostics of body fluids, it is important to develop cost-effective and high-quality assays for each question. It is essential to be able to quickly and efficiently generate data representing the different analytes that may carry the strongest signals needed to successfully train high-performance (accuracy) predictive models.

[0069] A. Specimen In various embodiments, biological samples include different specimens that provide a source of features for the models, methods, and systems described herein. The specimens can be derived from apoptosis, necrosis, and secretions from tumor, non-tumor, or immune cells. Four classes of highly informative molecular biomarkers include: 1) genomic biomarkers, based on analysis of DNA profiles, sequences, or modifications; 2) transcriptomic biomarkers, based on analysis of RNA expression profiles, sequences, or modifications; 3) proteomic or protein biomarkers, based on analysis of protein profiles, sequences, or modifications; and 4) metabolomic biomarkers, based on analysis of metabolite abundance.

[0070] 1. DNA Examples of nucleic acids include, but are not limited to, deoxyribonucleic acid (DNA), genomic DNA, plasmid DNA, complementary DNA (cDNA), cell-free (e.g., non-encapsulated) DNA (cfDNA), circulating tumor DNA (ctDNA), nucleosomal DNA, chromatosomal DNA, mitochondrial DNA (miDNA), artificial nucleic acid analogs, recombinant nucleic acids, plasmids, viral vectors, and chromatin. In one example, the sample includes cfDNA. In one example, the sample includes genomic DNA from PBMCs.

[0071] 2.RNA In various embodiments, the biological sample contains coding and non-coding transcripts, including ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (miRNA), ribosomal RNA (rRNA), circular RNA (cRNA), alternatively spliced mRNA, small nuclear RNA (snRNA), antisense RNA, short hairpin RNA (shRNA), and small interfering RNA (siRNA).

[0072] The nucleic acid molecules or fragments thereof may be single-stranded or double-stranded. A sample may contain more than one type of nucleic acid molecule or fragment thereof.

[0073] A nucleic acid molecule or a fragment thereof can contain any number of nucleotides. For example, a single-stranded nucleic acid molecule or a fragment thereof can contain at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 350, at least 400, or more nucleotides. In the case of double-stranded nucleic acid molecules or fragments thereof, the nucleic acid molecules or fragments thereof may contain at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 220, at least 240, at least 260, at least 280, at least 300, at least 350, at least 400, or more base pairs (bp), e.g., nucleotide pairs. In some cases, double-stranded nucleic acid molecules or fragments thereof may contain 100 to 200 bp, e.g., 120 to 180 bp. For example, a sample may contain cfDNA molecules containing 120 to 180 bp.

[0074] 3. Polyamino acids, peptides, and proteins In various embodiments, the analyte is a polyamino acid, peptide, protein, or fragment thereof. As used herein, the term polyamino acid refers to a polymer in which the monomers are amino acid residues linked together through amide bonds. When the amino acids are α-amino acids, either the L-optical isomer or the D-optical isomer can be used, with the L-isomer being preferred. In one embodiment, the analyte is an autoantibody.

[0075] In cancer patients, serum antibody profiles change and autoantibodies against cancerous tissues are produced. These profile changes offer many possibilities for tumor-associated antigens as markers for early cancer diagnosis. The immunogenicity of tumor-associated antigens is conferred by mutated amino acid sequences that expose altered non-self epitopes. Other explanations for this immunogenicity include alternative splicing, expression of embryonic proteins in adulthood (e.g., ectopic expression), deregulation of apoptotic or necrotic processes (e.g., overexpression), and abnormal cellular localization (e.g., secreted nuclear proteins). Examples of epitopes of tumor-restricted antigens encoded by intronic sequences (e.g., translated partially unspliced RNA) have been shown to make tumor-associated antigens highly immunogenic.

[0076] Exemplary markers of the present invention are suitable protein antigens that are overexpressed in tumors. Markers usually cause antibody responses in patients. Therefore, the most convenient method for detecting the presence of these markers in patients is to detect (auto)antibodies against these marker proteins in samples from patients, particularly body fluid samples such as blood, plasma or serum.

[0077] 4. Other specimens In various embodiments, the biological sample includes small chemical molecules such as, but not limited to, sugars, lipids, amino acids, fatty acids, phenolic compounds, and alkaloids.

[0078] In one embodiment, the analyte is a metabolite. In one embodiment, the analyte is a carbohydrate. In one embodiment, the analyte is a carbohydrate antigen. In one embodiment, the carbohydrate antigen is attached to an O-glycan. In one embodiment, the analyte is a monosaccharide, disaccharide, trisaccharide, or tetrasaccharide. In one embodiment, the analyte is a tetrasaccharide. In one embodiment, the tetrasaccharide is CA19-9. In one embodiment, the analyte is a nucleosome. In one embodiment, the analyte is platelet-rich plasma (PRP). In one embodiment, the analyte is a cellular element such as lymphocytes (neutrophils, eosinophils, basophils, lymphocytes, PBMCs, and monocytes) or platelets.

[0079] In one embodiment, the analyte is a cellular element such as lymphocytes (neutrophils, eosinophils, basophils, lymphocytes, PBMCs and monocytes) or platelets.

[0080] In various embodiments, a combination of analytes is assayed to obtain information useful in the methods described herein. In various embodiments, different combinations of analytes are assayed depending on the type of cancer or classification needs.

[0081] In various embodiments, the analyte combination is selected from 1) cfDNA, cfRNA, polyamino acids and small chemical molecules, or 2) cfDNA and cfRNA and polyamino acids, 3) cfDNA and cfRNA and small chemical molecules, or 4) cfDNA, polyamino acids and small chemical molecules, or 5) cfRNA, polyamino acids and small chemical molecules, or 6) cfDNA and cfRNA, or 7) cfDNA and polyamino acids, or 8) cfDNA and small chemical molecules, or 9) cfRNA and polyamino acids, or 10) cfRNA and small chemical molecules, or 11) polyamino acids and small chemical molecules.

[0082] II. Sample Preparation In some embodiments, the sample is obtained from, for example, tissue or bodily fluid from a subject, or both. In various embodiments, the biological sample is a liquid sample, such as plasma, serum, buffy coat, mucus, urine, saliva, or cerebrospinal fluid. In one embodiment, the liquid sample is a cell-free liquid. In various embodiments, the sample contains cell-free nucleic acid (e.g., cfDNA or cfRNA).

[0083] A sample containing one or more analytes can be treated to provide or purify a specific nucleic acid molecule or fragment thereof, or a collection thereof. For example, a sample containing one or more analytes can be treated to separate one type of analyte (e.g., cfDNA) from another type of analyte. In another example, the sample is divided into aliquots for analysis of a different analyte in each aliquot from the sample. In one example, a sample containing one or more nucleic acid molecules or fragments thereof of different sizes (e.g., lengths) can be treated to remove higher molecular weight and / or longer nucleic acid molecules or fragments thereof, or lower molecular weight and / or shorter nucleic acid molecules or fragments thereof.

[0084] The methods described herein may include treating or modifying a nucleic acid molecule or a fragment thereof. For example, the nucleotides of a nucleic acid molecule or a fragment thereof can be modified to include modified nucleobases, sugars, and / or linkers. Modification of a nucleic acid molecule or a fragment thereof can include oxidation, reduction, hydrolysis, tagging, barcoding, methylation, demethylation, halogenation, deamination, or any other process. Modification of a nucleic acid molecule or a fragment thereof can be achieved using enzymes, chemical reactions, physical processes, and / or exposure to energy. For example, for methylation analysis, deamination of unmethylated cytosine can be achieved through the use of bisulfite.

[0085] Sample processing can include one or more processes, such as centrifugation, filtration, selective precipitation, tagging, barcoding, and partitioning. For example, cellular DNA can be separated from cfDNA by selective polyethylene glycol and bead-based precipitation processes, such as centrifugation or filtration processes. Cells contained in the sample may or may not be lysed prior to separation of different types of nucleic acid molecules or fragments thereof. In one example, the sample is substantially cell-free. In one example, cellular components are assayed for measurements that can be input as features into a machine learning method or model. In various examples, cellular components, such as PBMC lymphocytes, can be detected (e.g., by flow cytometry, mass spectrometry, or immunopanning). A processed sample can contain, for example, at least 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 nanogram (ng), 10 ng, 50 ng, 100 ng, 500 ng, 1 microgram (μg), or more of a particular size or type of nucleic acid molecule or fragment thereof.

[0086] In some examples, blood samples are obtained from healthy individuals and individuals with cancer, e.g., individuals with stage I, II, III, or IV cancer. In one example, blood samples are obtained from healthy individuals and individuals with benign polyps, advanced adenomas (AA), and stages I-IV colorectal cancer (CRC). The systems and methods described herein are useful for detecting the presence of AA and CRC and distinguishing between their stages and sizes. Such differentiation is useful for stratifying individuals in a population for changes in behavior and / or treatment decisions.

[0087] A. Library Preparation and Sequencing Purified nucleic acids (e.g., cfDNA) can be used to prepare a library for sequencing. The library can be prepared using platform-specific library preparation methods or kits. The methods or kits can be commercially available and can generate sequencer-ready libraries. The platform-specific library preparation methods can add known sequences to the ends of nucleic acid molecules, and the known sequences can be referred to as adapter sequences. Optionally, the library preparation methods can incorporate one or more molecular barcodes.

[0088] To sequence a collection of double-stranded DNA fragments using a massively parallel sequencing system, the DNA fragments must be flanked by known adapter sequences. A collection of such DNA fragments with adapters at both ends is called a sequencing library. Two examples of suitable methods for generating a sequencing library from purified DNA are (1) ligation-based attachment of known adapters to either end of fragmented DNA, and (2) transposase-mediated insertion of adapter sequences. Any suitable massively parallel sequencing technology can be used for sequencing.

[0089] For methylation analysis, nucleic acid molecules are treated before sequencing. Treating nucleic acid molecules (e.g., DNA molecules) with bisulfite, enzymatic methyl-seq, or hydroxymethyl-seq deaminates unmethylated cytosine bases and converts them to uracil bases. This bisulfite conversion process does not deaminate cytosines that are methylated or hydroxymethylated at the 5' position (5mC or 5hmC). When used in conjunction with sequencing analysis, a process involving bisulfite conversion of a nucleic acid molecule or a fragment thereof may be referred to as bisulfite sequencing (BS-seq). In some cases, nucleic acid molecules may be oxidized before undergoing bisulfite conversion. Oxidation of nucleic acid molecules can convert 5hmC to 5-formylcytosine and 5-carboxylcytosine, both of which are susceptible to bisulfite conversion to uracil. When used in conjunction with sequence analysis, the oxidation of a nucleic acid molecule or fragment thereof prior to subjecting the nucleic acid molecule or fragment thereof to bisulfite sequencing can be referred to as oxidized bisulfite sequencing (oxBS-seq).

[0090] 1. Sequencing Nucleic acids can be sequenced using sequencing methods such as next-generation sequencing, high-throughput sequencing, massively parallel sequencing, sequencing-by-synthesis, paired-end sequencing, single-molecule sequencing, nanopore sequencing, pyrosequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-seq, Digital Gene Expression, Single Molecule Sequencing by Synthesis (SMSS), Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and Sanger sequencing.

[0091] The sequencing method may include targeted sequencing, whole genome sequencing (WGS), low-pass sequencing, bisulfite sequencing, whole genome bisulfite sequencing (WGBS), or a combination thereof. The sequencing method may include the preparation of a suitable library. The sequencing method may include nucleic acid amplification (e.g., by targeted or universal amplification, such as PCR). The sequencing method may be performed at a desired depth, for example, at least about 5x, at least about 10x, at least about 15x, at least about 20x, at least about 25x, at least about 30x, at least about 35x, at least about 40x, at least about 45x, at least about 50x, at least about 60x, at least about 70x, at least about 80x, at least about 90x, or at least about 100x. Targeted sequencing methods may be performed to a desired depth, for example, at least about 500-fold, at least about 1000-fold, at least about 1500-fold, at least about 2000-fold, at least about 2500-fold, at least about 3000-fold, at least about 3500-fold, at least about 4000-fold, at least about 4500-fold, at least about 5000-fold, at least about 6000-fold, at least about 7000-fold, at least about 8000-fold, at least about 9000-fold, at least about 10000-fold.

[0092] Biological information can be generated using any useful method. Biological information can include sequencing information. For example, sequencing information can be generated using the assay for transposase-accessible chromatin using sequencing (ATAC-seq), micrococcal nuclease sequencing (MNase-seq), deoxyribonuclease hypersensitive site sequencing (DNase-seq), or chromatin immunoprecipitation sequencing (ChIP-seq).

[0093] Sequencing reads can be obtained from a variety of sources, including, for example, whole genome sequencing, whole exome sequencing, targeted sequencing, next generation sequencing, pi sequencing, sequencing by synthesis, ion semiconductor sequencing, tag-based next generation sequencing, semiconductor sequencing, single molecule sequencing, nanopore sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (DGE), massively parallel sequencing, clonal single molecule arrays (Solexa / Illumina), sequencing using PacBio, and sequencing by oligonucleotide ligation and detection (SOLiD).

[0094] In some examples, sequencing involves modifying a nucleic acid molecule or fragment thereof, for example, by ligating a barcode, unique molecular identifier (UMI), or another tag to the nucleic acid molecule or fragment thereof. Ligating a barcode, UMI, or tag to one end of a nucleic acid molecule or fragment thereof can facilitate analysis of the nucleic acid molecule or fragment thereof after sequencing. In some examples, the barcode is a unique barcode (i.e., a UMI). In some examples, the barcode is non-unique, and the barcode sequence can be used in conjunction with endogenous sequence information, such as the start and stop sequences of the target nucleic acid (e.g., the target nucleic acid is adjacent to the barcode, and the barcode sequence is associated with sequences at the start and end of the target nucleic acid to create a uniquely tagged molecule).

[0095] Sequencing reads may be processed using methods such as demultiplexing, de-redundancy (e.g., using unique molecular identifiers, UMIs), adapter trimming, quality filtering, GC correction, amplification bias correction, correction for batch effects, depth normalization, removal of sex chromosomes, and removal of low-quality genomic bins.

[0096] In various embodiments, the sequencing reads can be aligned to a reference nucleic acid sequence. In one embodiment, the reference nucleic acid sequence is a human reference genome. For example, the human reference genome can be hg19, hg38, GrCH38, GrCH37, NA12878, or GM12878.

[0097] 2. Assay The selection of which assay to use is based on the results of training a machine learning model, taking into account the clinical goals of the system. As used herein, the term "assay" includes known biological assays and may also include computational biology approaches for converting biological information into useful features as input for machine learning analysis and modeling. The assays described herein may include various preprocessing computational tools, and the term "assay" is not intended to be limiting. Various classes of samples, sample fractions, portions of those fractions / samples with different classes of molecules, and multiple types of assays can be used to generate feature data for use in computational methods and models to inform classifiers useful in the methods described herein. In one example, a sample is divided into aliquots for performing a bioassay.

[0098] In various embodiments, biological assays are performed on different portions of a biological sample to provide a dataset corresponding to the biological assay for the analytes in the portion. Various assays are known to those skilled in the art and are useful for investigating biological samples. Examples of such assays include, but are not limited to, whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), small RNA sequencing, quantitative immunoassays, enzyme-linked immunosorbent assays (ELISA), proximity extension assays (PEA), protein microarrays, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing (WO2019 / 051484), CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56-gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing, as well as cell type or cell phenotype mixture ratios derived from any of the above assays. The ability to simultaneously analyze multiple analytes (such as, but not limited to, DNA, RNA, protein, autoantibodies, metabolites, or combinations thereof) from the same biological sample or fraction thereof can increase the sensitivity and specificity of diagnostic testing of such bodily fluids by utilizing independent information between signals.

[0099] In one example, cell-free DNA (cfDNA) content is assessed by low-coverage whole genome sequencing (lcWGS) or targeted sequencing, or whole genome bisulfite sequencing (WGBS) or whole genome enzymatic methyl sequencing, cell-free microRNA (cf-miRNA) is assessed by small RNA sequencing or PCR (digital droplet or quantitative), and circulating protein levels are measured by quantitative immunoassay. In one example, cell-free DNA (cfDNA) content is assessed by whole genome bisulfite sequencing (WGBS), proteins are measured by quantitative immunoassays (including ELISA or proximity extension assays), and autoantibodies are measured by protein microarrays.

[0100] B. cf-DNA assay using WGS In various embodiments, assays that profile cfDNA characteristics are used to generate features useful in computational applications. In one embodiment, cfDNA features are used in machine learning models to generate classifiers that stratify individuals or detect disease, as described herein. Exemplary features include, but are not limited to, those that provide biological information about gene expression, 3D chromatin, chromatin state, copy number variants, tissue of origin, and cellular composition in cfDNA samples. Metrics of cfDNA concentration that can be used as input features for machine learning methods and models can be obtained by methods that include, but are not limited to, methods that quantify dsDNA within a specified size range (e.g., Agilent TapeStation, Bioanalyzer, Fragment Analyzer), methods that quantify all dsDNA using dsDNA-binding dyes (e.g., QuantiFluor, PicoGreen, SYBR Green), and methods that quantify DNA fragments (either dsDNA or ssDNA) below a certain size (e.g., short fragment qPCR, long fragment qPCR, and long / short fragment qPCR ratio).

[0101] The biological information may also include information regarding transcription start sites, transcription factor binding sites, assay for transposase accessible chromatin using sequencing (ATAC-seq) data, histone marker data, DNAse hypersensitive sites (DHS), or combinations thereof.

[0102] In one example, the sequencing information includes information regarding multiple genetic features such as, but not limited to, transcription start sites, transcription factor binding sites, chromatin gating state, and nucleosome positioning or occupancy.

[0103] 1.cfDNA plasma concentration In various embodiments, the plasma concentration of cfDNA can be assayed as a feature indicative of the presence of cancer. In various embodiments, both the total amount of circulating cfDNA and an estimate of the tumor-derived contribution to cfDNA (also referred to as the "tumor fraction") are used as prognostic biomarkers and indicators of response and resistance to treatment. Sequencing fragments aligned within annotated genomic regions are counted and normalized for sequencing depth to generate a 30,000-dimensional vector per sample, with each element corresponding to a gene count (e.g., the number of reads aligned to that gene in the reference genome). In one example, for a list of known genes with annotated regions, the sequence read count is determined for each of those annotated regions by counting the number of fragments aligned to that region. Gene read counts are normalized in various ways, including using a global expectation where the genome is aligned, within-sample normalization, and cross-feature normalization. Cross-feature normalization refers to all of those features being averaged to a specified value, such as 0, a distinct negative value, 1, or a range from 0 to 2. With respect to cross-feature normalization, the total reads from a sample are variable and may therefore depend on the preparation process and sequencing load process. Normalization can be done to a fixed number of reads as part of global normalization.

[0104] For intra-sample normalization, it is possible to normalize by some features or selective features of some regions, especially GC bias. Therefore, the base pair composition of each region is different and can be used for normalization. In some cases, the GC number is significantly higher or lower than 50%, which has a thermodynamic impact because the base is more energetic, biasing the process. Some regions give more reads than expected due to biological artifacts of sample preparation in the laboratory. Therefore, during modeling, it may be necessary to correct such biases by applying different types of features, feature transformations, or normalization methods.

[0105] In one example, the software tool ichorCNA is used to identify tumor fraction components of cfDNA via copy number alterations detected by sparse (approximately 0.1x coverage) to deep (approximately 30x coverage) whole genome sequencing (WGS). In another example, measuring tumor content by quantifying the presence of individual alleles is used to assess response or resistance to treatment in cancers where those alleles are known clonal drivers.

[0106] Copy number variation (CNV) is recognized as a major source of average human genome survival and can be amplified or deleted in regions of the genome that significantly contribute to phenotypic variation. Tumor-derived cfDNA has genomic alterations corresponding to copy number changes. Copy number changes play a role in carcinogenesis in many cancers, including CRC. The detection of genome-wide copy number changes can be characterized in cfDNA, which acts as a tumor biomarker. In one example, deep WGS is used for detection. In another example, chromosomal instability analysis in cell-free DNA by low-coverage whole genome sequencing can be used as an assay for cfDNA. Other examples of cfDNA assays useful for detecting tumor DNA fragments include the length mixture model (LMM) and fragment endpoint analysis.

[0107] In one example, high tumor fraction samples (>20%) are identified through manual inspection of large-scale CNVs.

[0108] In one embodiment, changes in gene expression are also reflected in plasma cfDNA concentration levels, and methods such as microarray analysis can be used to measure changes in gene expression levels in cfDNA samples.Methods of machine learning and modeling can be used as input features for cfDNA concentration, including but not limited to Tape Station, short qPCR, long qPCR, and long / short qPCR ratio.

[0109] 2. Somatic Mutation Analysis In one example, low-coverage whole genome sequencing (lcWGS) can be used to sequence the cf-DNA in a sample and then interrogate for somatic mutations associated with a particular cancer type. Somatic mutations from lcWGS, deep WGS, or targeted sequencing (by NGS or other techniques) can be used to generate features that can be input into the machine learning methods and models described herein. Somatic mutation analysis has matured to include highly complex technologies such as microarrays and next-generation sequencing (NGS), or massively parallel sequencing. This approach can enable extensive multiplexing capabilities in a single test. These types of hotspot panels can span from a few to hundreds of genes in a single assay. Other types of gene panels include whole-exon or whole-gene sequencing, offering the advantage of identifying novel mutations in specific sets of genes.

[0110] 3. Transcription Factor Profiling Inferring transcription factor binding from cfDNA has tremendous diagnostic potential in cancer. Components involved in nucleosome signatures at transcription factor binding sites (TFBSs) are assayed to assess and compare transcription factor binding site accessibility in different plasma samples. In one example, deep whole-genome sequencing (WGS) data obtained from blood samples collected from healthy donors and plasma samples from cancer patients with metastatic prostate, colon, or breast cancer is used, where the cfDNA also contains circulating tumor DNA (ctDNA). Instead of using a mixture of cfDNA signals from multiple cell types to establish general tissue-specific patterns, shallow WGS data profiles individual transcription factors and performs analysis by Fourier transform and statistical summarization. Thus, the approach provided herein provides a more nuanced view of both tissue contributions and biological processes, enabling the identification of lineage-specific transcription factors suitable for both tissue-of-origin and tumor-of-origin analysis. In one example, the plasticity of transcription factor binding sites in cfDNA from cancer patients is used to classify cancer subtypes, stages, and response to treatment.

[0111] In one example, cfDNA fragmentation patterns are used to detect non-hematopoietic signatures. To identify transcription factor-nucleosome interactions mapped from cfDNA, hematopoietic transcription factor-nucleosome footprints are first identified in plasma samples from healthy controls. A curated list of transcription factor binding sites from publicly accessible databases (e.g., the Gene Transcription Regulation Database (GTRD)) may be used to generate a comprehensive transcription factor binding site-nucleosome occupancy map from cfDNA. Different stringency criteria are used to measure nucleosome signatures at transcription factor binding sites, establishing a metric called an "accessibility score" and z-score statistics for objectively comparing significant changes in transcription factor binding site accessibility among different plasma samples. For clinical purposes, a set of lineage-specific transcription factors suitable for identifying the tissue of origin of cfDNA or the tumor of origin in patients with cancer can be identified. The accessibility score and z-score statistics are used to elucidate altered transcription factor binding site accessibility from cfDNA from patients with cancer.

[0112] In one aspect, the present disclosure provides a method for diagnosing disease in a subject, comprising: (a) providing sequence reads from deoxyribonucleic acid (DNA) extracted from the subject; (b) generating a coverage pattern of transcription factors; (c) processing the coverage pattern to provide a signal; (d) comparing the signal to a reference signal, wherein the signal and the reference signal have different frequencies; and (e) diagnosing the disease in the subject based on the signal.

[0113] In some embodiments, (b) includes aligning the sequence reads to a reference sequence to provide an aligned sequence pattern, selecting a region of the aligned sequence pattern that corresponds to a binding site of the transcription factor, and normalizing the aligned sequence pattern within the region.

[0114] In some examples, the transcription factor is selected from the group consisting of GRH-L2, ASH-2, HOX-B13, EVX2, PU.1, Lyl-1, Spi-B, and FOXA1.

[0115] In some embodiments, (e) comprises identifying an indication of increased accessibility of a transcription factor. In some embodiments, the transcription factor is an epithelial transcription factor. In some embodiments, the transcription factor is GRHH-L2.

[0116] 4. Estimated Chromosome Structure / Chromatin State In other examples, assays are used to infer the three-dimensional structure of the genome using cell-free DNA (cfDNA). In particular, the present disclosure provides methods and systems for detecting chromatin abnormalities associated with diseases or conditions, such as cancer. Without being bound by any particular mechanism, it is believed that DNA fragments are released from cells, for example, into the bloodstream. Once released from cells, the half-life of the released DNA fragments, known as cell-free DNA (cfDNA), can depend on the chromatin remodeling state. Therefore, the abundance of cfDNA fragments in a biological sample can indicate the chromatin state of the gene from which the cfDNA fragment is derived (known as the "location" of the cfDNA). The chromatin state of a gene can change in disease. Identifying changes in the chromatin state of a gene can serve as a method for identifying the presence of disease in a subject. The chromatin state of a gene can be predicted from the abundance and location of cfDNA fragments in a biological sample using computer-aided techniques. Chromatin state can also be useful in inferring gene expression in a sample. A non-limiting example of a computer-assisted technique that can be used to predict chromatin states is a probabilistic graphical model (PGM). PGMs can be estimated by using statistical techniques such as expectation maximization or gradient methods to identify cfDNA profiles of open and closed TSSs (or between states), and then fitting those parameters to a training set and statistical techniques to estimate the PGM's parameters. The training set can be a cfDNA profile of known open and closed transcription start sites. Once trained, the PGM can predict the chromatin state of one or more genes in naive (previously untested) samples. Predictions can be analyzed and quantified. By comparing predictions of the chromatin state of one or more genes from healthy and diseased samples, biomarkers or diagnostic tests can be developed. PGMs can include various information, measurements, and mathematical objects that contribute to making the model more accurate.These objects can include other measured covariates, such as the biological context of the data and the laboratory process conditions of the sample.

[0117] In one example where the genetic feature is a chromatin state, a first array provides a measure of the constitutive openness of multiple cell types as a reference, a second array provides the relative proportions of the cell types in the sample, and a third array provides a measure of the chromatin state in the sample.

[0118] Gene expression can be controlled by the cellular machinery's access to the transcription start site. Access to the transcription start site can be determined by the state of chromatin where the transcription start site is located. The state of chromatin can be controlled by chromatin remodeling, which can condense (close) or relax (open) the transcription start site. A closed transcription start site results in decreased gene expression, while an open transcription start site results in increased gene expression. The length of cfDNA fragments can also depend on the chromatin state. Chromatin remodeling can occur through modifications of histones and other associated proteins. Non-limiting examples of histone modifications that can control the state of chromatin and transcription start site include, for example, methylation, acetylation, phosphorylation, and ubiquitination.

[0119] Gene expression is also controlled by more distal elements, such as enhancers, that interact with the transcription machinery in the physical 3D space of the genome. ATAC-seq and DNAse-seq provide measurements of open chromatin, which may not be clearly associated with specific genes but correlates with the binding of these more distal elements. For example, ATAC-seq data can be obtained for numerous cell types and conditions and can be used to identify regions of the genome with open chromatin for various basal regions, such as active transcription start sites or bound enhancers or repressors.

[0120] The half-life of cfDNA, once released from cells, can depend on the chromatin remodeling state. Therefore, the abundance of cfDNA fragments in a biological sample can indicate the chromatin state of the gene from which the cfDNA fragment originates (referred to herein as the "location" of the cfDNA). The chromatin state of a gene can be altered in disease. Identifying changes in the chromatin state of a gene can serve as a method for identifying the presence of disease in a subject. When comparing expressed and unexpressed genes, quantitative shifts are observed in both the number and location distribution of cell-free DNA (cfDNA) fragments. More specifically, there is a strong depletion of reads within the approximately 1000-3000 bp region surrounding the transcription start site (TSS), and nucleosomes downstream of the TSS are strongly positioned (their location becomes much more predictable). The present disclosure provides a method to resolve this inverse correlation, allowing inference of gene expression or chromatin openness starting from cfDNA. In one example, this assay is used in the multi-analyte method described herein.

[0121] The present disclosure also provides methods to generate predictions for other chromatin states as well, e.g., in repressive regions, active or poised promoters, etc. These predictions can quantify differences between different individuals (or samples), e.g., healthy individuals, colorectal cancer (CRC) patients, or samples diagnosed with other diseases or cancers.

[0122] Because the presence of open chromatin is broadly captured by the absence of nucleosomes or by the presence of strongly positioned nucleosomes adjacent to internal regions of open chromatin, the methods described herein can also be used for enhancers, repressors, or simply regions of open chromatin identified by other means in a reference sample.

[0123] The position of cfDNA sequence reads in the genome can be determined by "mapping" the sequence to a reference genome. Mapping can be performed using computer algorithms, including, for example, the Needleman-Wunsch algorithm, the BLAST algorithm, the Smith-Waterman algorithm, the Burrows-Wheeler alignment, the suffix tree, or custom-developed algorithms.

[0124] The three-dimensional structure of chromosomes is responsible for compartmentalizing the nucleus and linking spatially separated functional elements in close proximity. Analysis of the spatial arrangement of chromosomes and understanding how chromosomes fold provides insight into the relationship between chromatin structure, gene activity, and the biological state of the cell.

[0125] The detection of DNA interaction and the modeling of three-dimensional chromatin structure can be achieved by using chromosome conformation techniques.Such techniques include, for example, 3C (chromosome conformation capture), 4C (circularized chromosome conformation capture), 5C (chromosome conformation capture carbon copy), Hi-C (3C with high-throughput sequencing), ChIP-loop (3C with ChIP-seq), and ChIA-PET (Hi-C with ChIP-seq).

[0126] Hi-C sequencing is used to explore the three-dimensional structure of entire genomes by coupling proximity-based ligation with massively parallel sequencing. Hi-C sequencing utilizes high-throughput next-generation sequencing to unbiasedly quantify interactions across the genome. In Hi-C sequencing, DNA is cross-linked with formaldehyde, the cross-linked DNA is digested with a restriction enzyme to generate 5'-overhangs, which are then filled with biotinylated residues. The resulting blunt-ended fragments are ligated under conditions that favor ligation between the cross-linked DNA fragments. The resulting DNA sample contains ligation products, consisting of fragments whose junctions are labeled with biotin and that were spatially close together in the nucleus. Hi-C libraries can be generated by shearing the DNA and selecting the biotinylated products with streptavidin beads. The libraries can then be analyzed using massively parallel paired-end DNA sequencing. This technique can be used to calculate all pairwise interactions within the genome and infer potential chromosomal structures.

[0127] In one example, nucleosome occupancy in cfDNA provides an index of DNA openness and the ability to infer transcription factor binding. In certain examples, nucleosome occupancy correlates with tumor cell phenotype.

[0128] cfDNA represents a unique specimen generated by endogenous physiological processes to generate an in vivo map of nucleosome occupancy through whole-genome sequencing. Nucleosome occupancy at transcription start sites has been exploited to infer expressed genes from cells that release their DNA into the circulation. cfDNA nucleosome occupancy can reflect the footprint of transcription factors.

[0129] In various examples, cfDNA includes unencapsulated DNA, e.g., in blood or plasma samples, and can include ctDNA and / or cfDNA. cfDNA can be less than 200 base pairs (bp) in length, e.g., 120-180 bp in length. The cfDNA fragmentation pattern generated by mapping cfDNA fragment ends to a reference genome can include regions of increased read depth (e.g., fragment pileups). These regions of increased read depth can be approximately 120-180 bp in size, reflecting the size of nucleosomal DNA. A nucleosome is a core of eight histone proteins wrapped by approximately 147 bp of DNA. A chromatosome contains the nucleosome plus histones (e.g., histone H1) and approximately 20 bp of associated DNA anchored outside the nucleosome. Regions of increased read depth in cfDNA can correlate with nucleosome configuration. Thus, the cfDNA analysis methods disclosed herein can facilitate nucleosome mapping. The fragment pileup seen when cfDNA reads are mapped to reference genome may reflect nucleosome binding, which protects certain regions from nuclease digestion during cell death (apoptosis) or the process of systemic clearance of circulating cfDNA by the liver and kidney.The cfDNA analysis method disclosed herein can be complemented by, for example, digesting DNA or chromatin with MNase and then sequencing (MNase sequencing).This method can reveal the regions of DNA that are protected from MNase digestion because nucleosomal histones bind at regular intervals, and the intervening regions are preferentially degraded, thus reflecting the footprint of nucleosome arrangement.

[0130] 5. Tissue of Origin Assay Multiple nucleic acid molecules in a cfDNA sample originate from one or more cell types. In various embodiments, assays are used to identify the tissue of origin of nucleic acid sequences in a sample. Inferring the contribution of the cells from which analytes in a sample originate is useful for deconstructing analyte information in biological samples. In various embodiments, methods such as learning regulatory regions (LRR) and immune DHS signatures are useful as methods for determining the cell type of origin and contributing cell types of analytes in biological samples. In various embodiments, genetic features such as V-plot measurements, FREE-C, cfDNA measurements on transcription start sites, and DNA methylation levels on cfDNA fragments are used as input features for machine learning methods and models.

[0131] In one example, a first array of values corresponding to the state of a plurality of genetic features for a plurality of cell types may be created. In one example, the values corresponding to the state of the plurality of genetic features are obtained for a reference population. The reference population provides values used to provide an indication of the composition state of the plurality of genetic features.

[0132] In one example, a second array of values corresponding to a plurality of genetic features for a plurality of nucleic acid molecules of the nucleic acid sample can also be generated. The first and second arrays can then be used to generate a third array of values.

[0133] In one embodiment, the first and second arrays are matrices, and are used to generate values for a third array through matrix multiplication and parameter optimization. In one embodiment, the values for the third array correspond to the estimated proportions of multiple cell types for multiple nucleic acid molecules of a sample. The nucleic acid data from the sample is used in combination with a reference population of information to estimate a mixture of reference populations that best matches the multiple nucleic acids of the sample. This mixture is normalized to 1 and can be used to represent the proportion or score of those reference populations in the sample.

[0134] Thus, the types and proportions of one or more cell types from which a plurality of nucleic acid molecules are derived may be determined.

[0135] In a first aspect, the present disclosure provides a method of processing a sample comprising a plurality of nucleic acid molecules, the method comprising: (a) providing sequencing information for a sample comprising a plurality of nucleic acid molecules, the sequencing information comprising information regarding a plurality of genetic characteristics, and the plurality of nucleic acid molecules being derived from one or more cell types; (b) generating a first array of values corresponding to aspects of the plurality of genetic characteristics for a plurality of cell types, the plurality of cell types including one or more cell types; (c) generating a second array of values corresponding to aspects of the plurality of genetic traits for a plurality of nucleic acid molecules of the sample; (d) using the values of the first array and the values of the second array to generate a third array of values corresponding to the plurality of cell types for the plurality of nucleic acid molecules of the sample, thereby determining the type and proportion of one or more cell types from which the plurality of nucleic acid molecules are derived.

[0136] C. cfDNA assay for methylation using WGBS 1. Methylation Sequencing Assays are used to sequence the entire genome (e.g., via WGBS). Enzymatic methyl sequencing ("EMseq") can provide ultimate resolution by characterizing DNA methylation at nearly every nucleotide in the genome. Other targeted methods, such as high-throughput sequencing, pyrosequencing, Sanger sequencing, qPCR, or ddPCR, can be useful for methylation analysis. DNA methylation refers to the addition of methyl groups to DNA and is one of the most extensively characterized epigenetic modifications with important functional consequences. Typically, DNA methylation occurs at cytosine bases in nucleic acid sequences. Enzymatic methyl sequencing is particularly useful because it uses a three-step conversion and requires a smaller sample volume for analysis.

[0137] In some embodiments of any of the foregoing aspects, the method includes subjecting the DNA or barcoded DNA to conditions sufficient to convert cytosine nucleobases of the DNA or barcoded DNA to uracil nucleobases to perform bisulfite conversion. In some embodiments, performing bisulfite conversion includes oxidizing the DNA or barcoded DNA. In some embodiments, oxidizing the DNA or barcoded DNA includes oxidizing 5-hydroxymethylcytosine to 5-formylcytosine or 5-carboxylcytosine. In some embodiments, the bisulfite conversion includes reduced representation bisulfite sequencing.

[0138] In other examples, the assay used for methylation analysis is selected from mass spectrometry, methylation-specific PCR (MSP), reduced representation bisulfite sequencing (RRBS), HELP assay, GLAD-PCR assay, ChIP-on-chip assay, restriction landmark genome scanning, methylated DNA immunoprecipitation (MeDIP), pyrosequencing of bisulfite-treated DNA, molecular cleavage photoassay, methyl-sensitive Southern blotting, high-resolution melting analysis (HRM or HRMA), ancient DNA methylation reconstruction, or methylation-sensitive single nucleotide primer extension assay (msSNuPE).

[0139] In one example, the assay used for methylation analysis is whole-genome bisulfite sequencing (WGBS). Modification of nucleic acid molecules or fragments thereof can be achieved using enzymes or other reactions. For example, deamination of cytosine can be achieved through the use of bisulfite. Treatment of nucleic acid molecules (e.g., DNA molecules) with bisulfite deaminates unmethylated cytosine bases and converts them to uracil bases. This bisulfite conversion process does not deaminate cytosines that are methylated or hydroxymethylated at position 5 (5mC or 5hmC). When used in conjunction with sequencing analysis, a process involving bisulfite conversion of nucleic acid molecules or fragments thereof can be referred to as bisulfite sequencing (BS-seq). In some cases, nucleic acid molecules can be oxidized before undergoing bisulfite conversion. Oxidation of nucleic acid molecules can convert 5hmC to 5-formylcytosine and 5-carboxylcytosine, both of which are susceptible to bisulfite conversion to uracil. When used in conjunction with sequence analysis, the oxidation of a nucleic acid molecule or fragment thereof prior to subjecting the nucleic acid molecule or fragment thereof to bisulfite sequencing can be referred to as oxidized bisulfite sequencing (oxBS-seq).

[0140] Cytosine methylation at CpG sites can be significantly enriched in nucleosome-spanning DNA compared to adjacent DNA. Therefore, CpG methylation patterns can also be used to infer nucleosome positioning using machine learning approaches. Matched nucleosome positioning and 5mC datasets from the same cfDNA sample, generated by micrococcal nuclease-seq (MNase-seq) and WGBS, respectively, can be used to train machine learning models. BS-seq or EM-seq datasets, regardless of methylation status, can be analyzed according to the same methods used in WGS to generate features for input into machine learning methods and models. 5mC patterns can then be used to predict nucleosome positioning, which can help infer gene expression and / or classification of diseases and cancers. In another example, features may be derived from a combination of methylation status and nucleosome positioning information.

[0141] Metrics used in methylation analysis include, but are not limited to, M-bias (percent base-by-base methylation for CpG, CHG, and CHH), conversion efficiency (percent methylation per base for CHH), hypomethylated blocks, methylation level (global average methylation for CPG, CHH, CHG, chrM, LINE1, and ALU), dinucleotide coverage (normalized coverage of dinucleotides), coverage uniformity (unique CpG sites at 1x and 10x average genome coverage for (S4 run)), overall average CpG coverage (depth), and average coverage at CpG islands, CGI shelves, and CGI shores. These metrics can be used as feature inputs for machine learning methods and models.

[0142] In one aspect, the disclosure provides a method, including: (a) providing a biological sample comprising deoxyribonucleic acid (DNA) from a subject; (b) subjecting the DNA to conditions sufficient to convert unmethylated cytosine nucleobases of the DNA to uracil nucleobases, where the conditions at least partially degrade the DNA; (c) sequencing the DNA to generate sequence reads; (d) computationally processing the sequence reads to (i) determine a degree of methylation of the DNA based on the presence of uracil nucleobases, and (ii) model the at least partial DNA degradation to generate degradation parameters; and (e) determining a genetic sequence characteristic using the degradation parameters and the degree of methylation.

[0143] In another aspect, the disclosure provides a method, comprising: (a) providing a biological sample comprising deoxyribonucleic acid (DNA) from a subject; (b) subjecting the DNA to conditions sufficient to optionally enrich for methylated DNA in the sample; (c) converting unmethylated cytosine nucleobases of the DNA to uracil nucleobases; (d) sequencing the DNA to generate sequence reads; (e) computationally processing the sequence reads to (i) determine a degree of methylation of the DNA based on the presence of uracil nucleobases, and (ii) model at least partial DNA degradation to generate degradation parameters; and (f) determining a genetic sequence characteristic using the degradation parameters and the degree of methylation.

[0144] In some embodiments, (d) comprises determining the degree of methylation of the DNA based on a ratio of unconverted cytosine nucleobases to converted cytosine nucleobases. In some embodiments, converted cytosine nucleobases are detected as uracil nucleobases. In some embodiments, uracil nucleobases are observed as thymine nucleobases in sequence reads.

[0145] In some embodiments, generating the decomposition parameters includes using a Bayesian model.

[0146] In some embodiments, the Bayesian model is based on strand bias or bisulfite transformation or over-transformation. In some embodiments, (e) includes using the decomposition parameters under the framework of a paired HMM or a naive Bayesian model.

[0147] In certain embodiments, the methylation of specific genetic markers is assayed for use in informing the classifiers described herein. In various embodiments, the methylation of promoters such as APC, IGF2, MGMT, RASSF1A, SEPT9, NDRG4, and BMP3, or combinations thereof, is assayed. In various embodiments, the methylation of two, three, four, or five of these markers is assayed.

[0148] 2. DMR (Dimethylated Maillard Receptor) In one example, the methylation analysis is a variably methylated region (DMR) analysis. DMRs are used to quantify CpG methylation across regions of the genome. Regions are dynamically assigned by discovery. Several samples from different classes can be analyzed, and the most variably methylated regions can be identified across different classifications. A subset can be selected as variably methylated and used for classification. The number of CpGs captured in a region can be used for analysis. Regions can tend to be of variable size. In one example, a pre-discovery process is performed that bundles several CpG sites together as regions. In one example, DMRs are used as input features for machine learning methods and models.

[0149] 3. Haplotype blocks In one example, a haplotype block assay is applied to a sample. Identification of methylation haplotype blocks aids in the deconvolution of heterogeneous tissue samples and the mapping of tumor tissue of origin from plasma DNA. In WGBS data, tightly coupled CpG sites known as methylation haplotype blocks (MHBs) can be identified. A metric called methylation haplotype burden (MHL) is used to perform tissue-specific methylation analysis at the block level. This method provides information blocks useful for the deconvolution of heterogeneous samples. This method is useful for quantitative estimation of tumor burden and tissue of origin mapping in circulating cf DNA. In one example, haplotype blocks are used as input features for machine learning methods and models.

[0150] D. cfRNA assay In various embodiments, assaying for cfRNA can be accomplished using methods such as RNA sequencing, whole transcriptome shotgun sequencing, Northern blot, in situ hybridization, hybridization array, linkage analysis of gene expression (SAGE), reverse transcription PCR, real-time PCR, real-time reverse transcription PCR, quantitative PCR, digital droplet PCR, or microarray, Nanostring, FISH assay, or combinations thereof.

[0151] When small cfRNAs (including oncRNAs and miRNAs) are used as samples, the measurement relates to the abundance of these cfRNAs. Their transcripts have a specific size, and each transcript is stored, and the number of cfRNAs found for each can be counted. RNA sequences can be aligned to a reference cfRNA database, such as a set of sequences corresponding to known cfRNAs in the human transcriptome. Each cfRNA found can be used as its own feature, and multiple cfRNAs found across all samples can become a feature set. In one example, RNA fragments aligned to annotated cfRNA genomic regions are counted and normalized for sequencing depth to generate a multidimensional vector of the biological sample.

[0152] In various embodiments, all measurable cfRNA (cfRNA) is used as a feature. Some samples have a feature value of 0, meaning no expression is detected for that cfRNA.

[0153] In one example, all samples are collected and the read values are aggregated together. For each microRNA found in the sample, multiple aggregate reads can be found. Note that microRNAs with higher expression ranks may provide better markers because they provide a more reliable signal with larger absolute changes.

[0154] In one example, cfRNA can be detected in a sample using direct detection methods such as molecular "barcoding" and microscopic imaging from the nCounter Analysis System® (nanoString, South Lake Union, WA) to detect and count up to hundreds of unique transcripts in a single hybridization reaction.

[0155] In various examples, assaying mRNA levels involves contacting a biological sample with a polynucleotide probe capable of specifically hybridizing to one or more sequences of mRNA, thereby forming a probe-target hybridization complex. Hybridization-based RNA assays include, but are not limited to, traditional "direct probe" methods such as Northern blots or in situ hybridization. These methods can be used in a variety of formats, including, but not limited to, substrate (e.g., membrane or glass)-bound methods or array-based approaches. In a typical in situ hybridization assay, cells are fixed to a solid support, typically a glass slide. If nucleic acids are being probed, the cells are typically denatured with heat or alkali. The cells are then contacted with a hybridization solution at a moderate temperature to allow annealing of labeled probes specific to protein-encoding nucleic acid sequences. The target (e.g., cells) is then typically washed at a predetermined stringency or with increasing stringency until an appropriate signal-to-noise ratio is obtained. The probe is typically labeled, for example, with a radioisotope or fluorescent reporter. Preferred probes are sufficiently long to specifically hybridize with the target nucleic acid(s) under stringent conditions. In one example, the size range is from about 200 bases to about 1000 bases. In another example, for small RNAs, shorter probes are used, in the size range of about 20 bases to about 200 bases.For hybridization protocols suitable for use with the methods of the present invention, see, for example, Albertson (1984) EMBO J. 3:1227-1234, Pinkel (1988) Proc. Natl. Acad. Sci. USA 85:9138-9142; EPO Pub. No. 430,402; Methods in Molecular Biology, Vol. 33: In situ Hybridization Protocols, Choo, ed., Humana Press, Totowa, NJ (1994), Pinkel, et al. (1998) Nature Genetics 20:207-211, and / or Kallioniemi (1992) Proc. Natl. Acad. Sci. USA 89:5321-5325 (1992). In some applications, it may be necessary to block the hybridization capacity of repetitive sequences. Thus, in some examples, tRNA, human genomic DNA, or Cot, I DNA is used to block nonspecific hybridization.

[0156] In various embodiments, assaying mRNA levels includes contacting the biological sample with a polynucleotide primer capable of specifically hybridizing to the mRNA of a single exon gene (SEG), forming a primer-template hybridization complex, and performing a PCR reaction. In some embodiments, the polynucleotide primer comprises an approximately 15-45, 20-40, or 25-35 bp sequence that is identical (in the case of a forward primer) or complementary (in the case of a reverse primer) to a sequence of a SEG listed in Table 1. As a non-limiting example, a polynucleotide primer for STMN1 (e.g., NM_203401, Homo sapiens stathmin 1 (STMN1), transcript variant 1, mRNA, 1730 bp) can contain sequences identical (in the case of a forward primer) or complementary (in the case of a reverse primer) to bp 1-20, 5-25, 10-30, 15-35, 20-40, 25-45, 30-50, etc., of STMN1, as well as bp 1690-1710, 1695-1715, 1700-1720, 1705-1725, 1710-1730, etc., up to the end of STMN1. While not exhaustively listed for space reasons, all of these polynucleotide primers for STMN1 and other SEGs listed in Table 1 can be used in the systems and methods of the present disclosure. In various embodiments, polynucleotide primers are labeled with radioisotopes or fluorescent molecules, and PCR products containing the labeled primers can be detected and analyzed with various imaging devices because the labeled primers emit a radioactive or fluorescent signal.

[0157] There are various suitable methods for "quantitative" amplification. For example, quantitative PCR involves simultaneously co-amplifying a known amount of a control sequence using the same primers. This can be used to provide an internal standard and calibrate the PCR reaction. Detailed protocols for quantitative PCR are provided in Innis, et al. (1990) PCR Protocols, A Guide to Methods and Applications, Academic Press, Inc. NY). Measurement of DNA copy number at microsatellite loci using quantitative PCR analysis is described in Ginzonger, et al. (2000) Cancer Research 60:5405-5409. The known nucleic acid sequence of a gene makes it possible to routinely select primers to amplify any portion of the gene. Fluorogenic quantitative PCR can also be used in the methods of the present invention. In fluorogenic quantitative PCR, quantification is based on the amount of fluorescent signal, for example, TaqMan and SYBR green. Other suitable amplification methods include, but are not limited to, ligase chain reaction (LCR) (see Wu and Wallace (1989) Genomics 4:560, Landegren, et al. (1988) Science 241:1077, and Barringer et al. (1990) Gene 89:117), transcription amplification (Kwoh, et al. (1989) Proc. Natl. Acad. Sci. USA 86:1173), self-sustained sequence replication (Guatelli, et al. (1990) Proc. Nat. Acad. Sci. USA 87:1874), dot PCR, and linker-adapter PCR.

[0158] In various embodiments, the cancer-associated RNA marker is selected from miR-125b-5p, miR-155, miR-200, miR21-5pm, miR-210, miR-221, miR-222, or a combination thereof.

[0159] E. Polyamino Acid and Autoantibody Assays 1. Proteins and peptides In various embodiments, proteins are assayed using immunoassays or mass spectrometry. For example, proteins can be measured by liquid chromatography-tandem mass spectrometry (LC-MS / MS).

[0160] In various embodiments, proteins are measured by affinity reagents or immunoassays such as protein arrays, SIMOA (antibodies; Quanterix), ELISA (Abcam), O-link (DNA-conjugated antibodies; O-link Proteomics), or SOMASCAN (aptamers; SomaLogic), Luminex, and Meso Scale Discovery.

[0161] In one example, protein data is normalized by a standard curve. In various examples, each protein is essentially treated as its own immunoassay, each with its own standard curve, which can be calculated in various ways. The concentration relationship is typically nonlinear. Samples may then be run and calculated based on the expected fluorescence concentration in the primary sample.

[0162] Several cancer-associated peptide and protein sequences are known and, in various embodiments, are useful in the systems and methods described herein.

[0163] In one example, the assay comprises a combination that detects at least 2, 3, 4, 5, 6 or more markers.

[0164] In various embodiments, the cancer-associated peptide or protein marker is selected from carcinoembryonic antigens (e.g., CEA, AFP), glycoprotein or carbohydrate antigens (e.g., CA125, CA19.9, CA15-3), enzymes (e.g., PSA, ALP, NSE), hormone receptors (ER, PR), hormones (b-hCG, calcitonin), or other known biomolecules (VMA, 5HIAA).

[0165] In various embodiments, the cancer-associated peptide or protein marker is 1p / 19q deletion, HIAA, ACTH, AE1, 3, ALK(D5F3), AFP, APC, ATRX, BOB-1, BCL-6, BCR-ABL1, β-hCG, BF-1, BTAA, BRAF, GCDFP-15, BRCA1, BRCA2, b72.3, c-MET, calcitonin, CALR, calretinin, CA125, CA27.29, CA19-9, CEA M, CEA P, CEA, CBFB-MYH11, CALA, c-Kit, syndical-1, CD14, CD15, CD19, CD2, CD20, CD200, CD23, CD3, CD30, CD33, CD4, CD45, CD5, CD56, CD57, CD68, CD7, CD79A, CD8, CDK4, CDK2, chromogranin A, creatine kinase isozyme, Cox-2, CXCL13, cyclin D, CK19, CYFRA21-1, CK20, CK5, CK6, CK7, and CAM5.2, DCC, des-γ-carboxyprothrombin, E-cadherin, EGFR T790M, EML4-ALK, ERBB2, ER, ESR1, FAP, gastrin, glucagon, HER-2 / neu, SDHB, SDHC, SDHD, HMB45, HNPCC, HVA, β-hCG, HE4, FBXW7, IDH1 R132H, IGH-CCND1, IGHV, IMP3, LOH, MUM1 / IRF4, JAK exon 12, JAK2 V617F, Ki-67, KRAS, MCC, MDM2, MGMT, MelanA, MET, metanephrine, MSI, MPL codon 515, Muc-1, Muckiest-4, MEN2, MYC, MYCN, MPO, myf4, myoglobin, myosin, napsin A, neurofilament, NSE The antibody is selected from P, NMP22, NPM1, NRAS, Oct2, p16, p21, p53, pancreatic polypeptide, PTH, Pax-5, PAX8, PCA3, PD-L1 28-8, PIK3CA, PTEN, ERCC-1, ezrin, STK11, PLAP, PML / RARa translocation, PR, proinsulin, prolactin, PSA, PAP, PGP, RAS, ROS1, S-100, S100A2, S100B, SDHB, serotonin, SAMD4, MESOMARK, squamous cell carcinoma antigen, SS18 SYT 18q11, synaptophysin, TIA-1, TdT, thyroglobulin, TNIK, TP53, TTF-1, TNF-α, TRAFF2, urovysion, VEGF, or a combination thereof.

[0166] In one embodiment, the cancer is colorectal cancer and the CRC-associated markers are selected from APC, BRAF, DPYD, ERBB2, KRAS, NRAS, RET, TP53, UGT1A1, and combinations thereof.

[0167] In one embodiment, the cancer is lung cancer and the lung cancer-associated markers are selected from ALK, BRAF, EGFR, ERBB2, KRAS, MET, NRAS, RET, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is breast cancer and the breast cancer-associated markers are selected from BRCA1, BRCA2, ERBB2, TP53, and combinations thereof. In one embodiment, the cancer is gastric cancer and the gastric cancer-associated markers are selected from APC, ERBB2, KRAS, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is glioma and the glioma-associated markers are selected from APCAPC, BRAF, BRCA2, EGFR, ERBB2, ROS1, TP53, and combinations thereof. In one embodiment, the cancer is melanoma and the melanoma-associated markers are selected from BRAF, KIT, NRAS, and combinations thereof. In one embodiment, the cancer is ovarian cancer and the ovarian cancer associated markers are selected from BRAF, BRCA1, BRCA2, ERBB2, KRAS, TP53, and combinations thereof. In one embodiment, the cancer is thyroid cancer and the thyroid cancer associated markers are selected from BRAF, KRAS, NRAS, RET, and combinations thereof. In one embodiment, the cancer is pancreatic cancer and the pancreatic cancer associated markers are selected from APC, BRCA1, BRCA2, KRAS, TP53, and combinations thereof.

[0168] 2. Autoantibodies In another example, antibodies (e.g., autoantibodies) are detected in the sample and are markers of early tumor formation. It has been demonstrated that autoantibodies are produced early in tumor formation and can be detected months or even years before clinical symptoms develop. In one example, plasma samples are screened using a mini-APS array (ITSI-Biosciences, Johnstown, PA, USA) using the protocol described by Somiari RI et al. (Somiari RI, et al., A low-density antigen array for detection of disease-associated autoantibodies in human plasma. Cancer Genom Proteom 13:13-19, 2016). Autoantibody markers can be used as input features in machine learning methods or models.

[0169] Assays for detecting autoantibodies include immunosorbent assays such as ELISA or PEA. When detecting an epitope containing an autoantibody, preferably a marker protein or at least a fragment thereof, it is attached to a solid support, such as a microtiter well. The autoantibody in the sample binds to this antigen or fragment. The bound autoantibody can be detected by a secondary antibody bearing a detectable label, such as a fluorescent label. The label is then used to generate a signal dependent on binding to the autoantibody. The secondary antibody may be an anti-human antibody if the patient is human, or may be directed against any other organism, depending on the patient sample being analyzed. The kit may include a means for such an assay, such as a solid support, and preferably also a secondary antibody. Preferably, the secondary antibody binds to the Fc portion of the patient's (auto)antibody. Buffers and washing or rinsing solutions can also be added. The solid support may be coated with a blocking compound to prevent nonspecific binding.

[0170] In one example, autoantibodies are assayed with a protein microarray or other immunoassay.

[0171] Metrics for autoantibody assays that can be used as input features include, but are not limited to, adjusted quantile-normalized z-scores for all autoantibodies, binary 0 / 1, or absence / presence for each autoantibody based on a specific z-score cutoff.

[0172] In various embodiments, the autoantibody markers are associated with different subtypes or stages of cancer. In various embodiments, the autoantibody markers are directed to or capable of binding with high affinity to a tumor-associated antigen. In various embodiments, the tumor-associated antigen is selected from carcinoembryonic antigen / immature laminin receptor protein (OFA / iLRP), alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), CA-125, MUC-1, epithelial tumor antigen (ETA), tyrosinase, melanoma-associated antigen (MAGE), aberrant products of ras, aberrant products of p53, wild-type ras, wild-type p53, or fragments thereof.

[0173] In one example, ZNF700 was shown to be a capture antigen for the detection of autoantibodies in colorectal cancer. In a panel with other zinc finger proteins, detection of ZNF-specific autoantibodies enabled the detection of colorectal cancer (O'Reilly et al., 2015). In one example, anti-p53 antibodies are assayed, as such antibodies can develop months to years before the clinical diagnosis of cancer.

[0174] F. Carbohydrates Assays exist for measuring carbohydrates in biological samples. Thin layer chromatography (TLC), gas chromatography (GC), and high performance liquid chromatography (HPLC) can be used to separate and identify carbohydrates. Carbohydrate concentrations can be determined gravimetrically (Manson and Walker), spectrophotometrically, or by titration (e.g., Lane-Eynon). There are also calorimetric methods for analyzing carbohydrates (anthrone, phenol-sulfuric acid). Other physical methods for characterizing carbohydrates include polarimetry, refractive index, IR, and density. In one example, metrics from carbohydrate assays are used as input features for machine learning methods and models.

[0175] III. Exemplary Systems In some examples, the present disclosure provides a system, method, or kit that may include software code executing on data analysis computing hardware implemented in a measurement device (e.g., laboratory equipment such as a sequencing machine). The software may be stored in memory and executed on one or more hardware processors. The software may be organized into routines or packages that can communicate with each other. Modules may include one or more devices / computers and one or more software routines / packages executing on the one or more devices / computers. For example, an analysis application or system may include at least a data reception module, a data preprocessing module, a data analysis module (which may operate on one or more types of genomic data), a data interpretation module, or a data visualization module.

[0176] The data receiving module can connect laboratory hardware or equipment with a computer system that processes laboratory data. The data preprocessing module can perform operations on the data in preparation for analysis. Examples of operations that can be applied to the data in the preprocessing module include affine transformations, noise removal operations, data cleaning, reformatting, or subsampling. The data analysis module, which can be specialized for analyzing genomic data from one or more genomic materials, can, for example, use assembled genomic sequences to perform probabilistic and statistical analysis to identify abnormal patterns associated with a disease, pathology, state, risk, condition, or phenotype. The data interpretation module can use analytical methods derived from, for example, statistics, mathematics, or biology to help understand the relationship between the identified abnormal patterns and a health state, functional status, prognosis, or risk. The data analysis module and / or data interpretation module can include one or more machine learning models, which can be implemented, for example, in hardware that executes software embodying the machine learning models. The data visualization module can use methods of mathematical modeling, computer graphics, or rendering to generate visual representations of the data that can facilitate understanding or interpretation of the results. The present disclosure provides a computer system programmed to implement the methods of the present disclosure.

[0177] In some examples, the methods disclosed herein can include computational analysis of nucleic acid sequencing data from a sample from an individual or multiple individuals. The analysis can identify inferred variants from the sequence data and identify sequence variants based on probabilistic modeling, statistical modeling, mechanistic modeling, network modeling, or statistical inference. Non-limiting examples of analytical methods include principal component analysis, autoencoders, singular value decomposition, Fourier basis, wavelets, discriminant analysis, regression, support vector machines, tree-based methods, networks, matrix decomposition, and clustering. Non-limiting examples of variants include germline polymorphisms or somatic mutations. In some examples, a variant can refer to a known variant. A known variant can be scientifically confirmed or reported in the literature. In some examples, a variant can refer to a putative variant associated with a biological change. The biological change can be known or unknown. In some examples, a putative variant can be reported in the literature but not yet biologically confirmed. Alternatively, a putative variant has not yet been reported in the literature but can be inferred based on the computational analysis disclosed herein. In some examples, a germline variant can refer to a nucleic acid that induces a natural or normal polymorphism.

[0178] Natural or normal polymorphisms can include, for example, skin color, hair color, and normal body weight. In some examples, somatic mutations can refer to nucleic acids that induce acquired or abnormal polymorphisms. Acquired or abnormal polymorphisms can include, for example, cancer, obesity, conditions, symptoms, diseases, and disorders. In some examples, the analysis can include identifying germline variants. Germline variants can include, for example, private variants and somatic mutations. In some examples, the identified variants can be used by clinicians or other medical professionals to improve healthcare methodologies, diagnostic accuracy, and cost reduction.

[0179] FIG. 1 illustrates a system 100 that is programmed or otherwise configured to perform the methods described herein. In various examples, the system 100 can process and / or assay samples, perform sequencing analysis, measure sets of values representing classes of molecules, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train the machine learning model (e.g., iteratively search for optimal values for the parameters of the machine learning model). The system 100 includes a computer system 101 and one or more measurement devices 151, 152, or 153 capable of measuring various analytes. As shown, the measurement devices 151-153 measure respective analytes 1-3.

[0180] The computer system 101 can regulate various aspects of the sample processing and assays of the present disclosure, such as activating valves or pumps to move reagents or samples from one chamber to another or applying heat to the sample (e.g., during an amplification reaction), other aspects of sample processing and / or assays, performing sequencing analysis, measuring sets of values representing classes of molecules, identifying sets of features and feature vectors from the assay data, processing the feature vectors using a machine learning model to obtain an output classification, and training the machine learning model (e.g., iteratively searching for optimal values for the parameters of the machine learning model). The computer system 101 can be a user's electronic device or a computer system located remotely relative to the electronic device.

[0181] The computer system 101 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 105, which may be a single-core or multi-core processor, or multiple processors for parallel processing; memory 110 (e.g., cache, random access memory, read-only memory, flash memory, or other memory); an electronic storage unit 115 (e.g., a hard disk); a communication interface 120 (e.g., a network adapter) for communicating with one or more other systems; and peripheral devices 125, such as adapters for cache, other memory, data storage, and / or electronic displays. The memory 110, storage unit 115, interface 120, and peripheral devices 125 may communicate with the CPU 105 via a communication bus (solid lines), such as a motherboard. The storage unit 115 may be a data storage unit (or data repository) for storing data. Input of one or more analyte characteristics may be input from one or more measurement devices 151, 152, or 153. Exemplary analyte and measurement devices are described herein.

[0182] The computer system 101 can be operatively coupled to a computer network ("network") 130 using the communications interface 120. The network 130 can be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. The network 130 is, in some cases, a telecommunications and / or data network. The network 130 can include one or more computer servers, enabling distributed computing, such as cloud computing via the network 130 (the "cloud"), to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as activating valves or pumps to move reagents or samples from one chamber to another or applying heat to a sample (e.g., during an amplification reaction), other aspects of sample processing and / or assays, performing sequencing analysis, measuring sets of values representing classes of molecules, identifying sets of features and feature vectors from assay data, processing the feature vectors using a machine learning model to obtain an output classification, and training the machine learning model (e.g., iteratively searching for optimal values for parameters of the machine learning model). Such cloud computing may be provided, for example, by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. Network 130 may, in some cases, implement a peer-to-peer network with computer system 101, allowing devices coupled to computer system 101 to act as clients or servers.

[0183] The CPU 105 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 110. The instructions may be directed to the CPU 105, which can then program or otherwise configure the CPU 105 to implement the methods of the present disclosure. The CPU 105 may be part of a circuit, such as an integrated circuit. One or more other components of the system 101 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0184] The storage unit 115 may store files, such as drivers, libraries, and saved programs. The storage unit 115 may store user data, such as user preferences and user programs. The computer system 101 may optionally include one or more additional data storage units external to the computer system 101, such as those located on remote servers that communicate with the computer system 101 via an intranet or the Internet.

[0185] Computer system 101 can communicate with one or more remote computer systems through network 130. For example, computer system 101 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 101 via network 130.

[0186] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored in electronic storage locations of computer system 101, such as, for example, on memory 110 or electronic storage unit 115. Machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by CPU 105. In some cases, the code may be retrieved from storage unit 115 and stored in memory 110 for immediate access by CPU 105. In some circumstances, electronic storage unit 115 may be eliminated, and machine-executable instructions are stored on memory 110.

[0187] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or in-place compiled manner.

[0188] Aspects of the systems and methods provided herein, such as computer system 101, may be embodied in programming. Various aspects of the technology can typically be thought of as a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of the tangible memory of a computer, processor, etc., or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, etc., that may provide non-transitory storage at any time for software programming. All or portions of the software may, from time to time, be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, e.g., from an administrative server or host computer to the computer platform of an application server. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0189] Thus, a machine-readable medium such as a computer-executable code may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer(s), such as those that may be used to implement, for example, the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system.

[0190] Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tape, any other physical storage media with patterns of holes, RAMs, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves transmitting data or instructions, cables or links transmitting such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0191] The computer system 101 can include or communicate with an electronic display 135 that includes a user interface (UI) 140 to provide, for example, the current stage of sample processing or assay (e.g., the lysis step, or a particular step, such as a sequencing step, being performed). Input is received by the computer system from one or more measurement devices 151, 152, or 153. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. The algorithm can, for example, process and / or assay samples, perform sequencing analysis, measure a set of values representing classes of molecules, identify a set of features and feature vectors from the assay data, process the feature vectors using a machine learning model to obtain an output classification, and train the machine learning model (e.g., iteratively search for optimal values for the parameters of the machine learning model).

[0192] IV. Machine Learning Tools To determine the set of assays to be used in experimental testing, machine learning systems can be leveraged to evaluate the validity of a given assay, or a given dataset generated from multiple assays, run on a given analyte and add to the overall predictive accuracy of the classification. In this way, new biological / health / diagnostic questions can be addressed to design new assays.

[0193] Machine learning can be used to reduce a set of data generated from all (primary sample / specimen / test) combinations to, for example, a set of optimal predictive features that meet specified criteria. In various embodiments, statistical learning and / or regression analysis can be applied. Simple to complex, small to large models making various modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes consideration of linear to nonlinear and non-hierarchical to hierarchical representations of features. Small to large models include consideration of the size of the basis vector space onto which the data is projected, as well as the number of interactions between features included in the modeling process.

[0194] Machine learning techniques can be used to evaluate commercial test modalities that best fit the cost / performance / commercial range, as defined by the initial question. Threshold checks can be performed. If the method applied to a holdout data set not used in cross-validation exceeds the initialized constraints, the assay is locked and put into operation. For example, assay performance thresholds can include a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof. For example, the desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or a combination thereof can be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. As another example, the desired minimum AUC can be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.A subset of assays can be selected from the set of assays performed on a given sample based on the total cost of running the subset of assays according to assay performance thresholds such as desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), and combinations thereof. If the thresholds are not met, the assay engineering procedure can loop back to either constraint setting for possible relaxation or to the wet lab to modify the parameters under which the data was acquired. When considering a clinical question, biological constraints, budget, lab machinery, etc. may constrain the problem.

[0195] In various embodiments, the computational processing of the machine learning techniques can include method(s) from statistics, mathematics, biology, or any combination thereof. In various embodiments, any one of the computational methods can include dimension reduction methods, logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosted trees, logistic regression, matrix decomposition, network clustering, statistical testing, and neural networks.

[0196] In various embodiments, the computer processing of machine learning techniques may include logistic regression, multiple linear regression (MLR), dimensionality reduction, partial least squares (PLS) regression, principal component regression, autoencoder, variational autoencoder, singular value decomposition, Fourier basis, wavelets, discriminant analysis, support vector machines, decision trees, classification and regression trees (CART), tree-based methods, random forests, gradient boosted trees, logistic regression, matrix decomposition, multidimensional scaling (MDS), dimensionality reduction methods, t-STEM Neighbor Embedding (t-SNE), multilayer perceptron (MLP), network clustering, neuro-fuzzy, neural networks (shallow and deep), artificial neural networks, Pearson product-moment correlation coefficient, Spearman's rank correlation coefficient, Kendall's rank correlation coefficient, or any combination thereof.

[0197] In some embodiments, the computational methods are supervised machine learning methods, including, for example, regression, support vector machines, tree-based methods, and neural networks. In some embodiments, the computational methods are unsupervised machine learning methods, including, for example, clustering, networks, principal component analysis, and matrix decomposition.

[0198] For supervised learning, training samples (e.g., thousands) may include measured data (e.g., of various analytes) and known labels, which may be determined through other time-consuming processes such as imaging the subject and analysis by a trained practitioner. Exemplary labels may include a classification of the subject, e.g., a discrete classification of whether the subject has cancer or not, or a continuous classification providing a probability of a discrete value (e.g., risk or score). The learning module can optimize the parameters of the model to obtain a quality metric (e.g., accuracy of prediction for known labels) for one or more specified criteria. The quality metric determination can be implemented for any function, including the set of all risk, loss, utility, and decision functions. Gradients can be used in conjunction with the learning step (e.g., a measure of how much the model's parameters should be updated for a given time step in the optimization process).

[0199] As described above, the embodiments can be used for a variety of purposes. For example, plasma (or other samples) can be collected from symptomatic subjects (e.g., known to have a condition) and healthy subjects. Genetic data (e.g., cfDNA) can be acquired and analyzed to obtain a variety of different features, including features based on genome-wide analysis. These features can form a feature space that can be searched, stretched, rotated, translated, and linearly or non-linearly transformed to generate an accurate machine learning model that can distinguish between healthy subjects and subjects with a pathological condition (e.g., identify a subject's diseased or non-disease state). Output from this data and the model (which may include a probability of the condition, a stage (level) of the condition, or other values) can be used to generate another model that can be used to recommend further procedures (e.g., whether to recommend a biopsy or continue monitoring the subject's condition).

[0200] V. Input Feature Selection As described above, a large set of features can be generated to provide a feature space from which feature vectors can be determined. This feature vector from each of a set of training samples can then be used to train the current version of the machine learning model. The type of features used can depend on the type of specimen used.

[0201] Exemplary features include variables related to structural polymorphisms (SVs), such as copy number variations and translocations, fusions, mutations (e.g., SNPs or other single nucleotide polymorphisms (SNVs), or small large sequence polymorphisms), telomere shortening, and nucleosome occupancy and distribution. These features can be calculated genome-wide. Exemplary classes (types) of functions are provided below. If genetic sequence data is obtained from at least one of the specimens, exemplary features may include aligned features (e.g., comparison with one or more reference genomes) and unaligned features. Exemplary aligned features may include sequence polymorphisms and sequence counts within a genome window. Examples of unaligned features include kmer from sequence reads and biological information from reads.

[0202] In some embodiments, at least one of the features is a gene sequence feature. For example, the gene sequence feature can be selected from DNA methylation status, single nucleotide polymorphism, copy number variation, insertion / deletion, and structural variant. In various embodiments, the methylation status can be used to determine nucleosome occupancy and / or determine the methylation density in CpG islands of DNA or barcoded DNA.

[0203] Ideally, feature selection can select features that are invariant or have low variability within samples with the same classification (e.g., having the same probability or associated risk of a particular phenotype), but that vary between groups of samples with different classifications. Procedures can be implemented to identify features that appear to be most invariant within a particular population (e.g., if classifications are real numbers, those that share a classification or lease have similar classifications). Procedures can also identify features that differ between populations. For example, read counts of sequence reads that partially or completely overlap various genomic regions of the genome can be analyzed to determine how they vary within a population, and such read counts can be compared to read counts in a separate population (e.g., subjects known to have a disease or disorder or subjects who are asymptomatic for the disease or disorder).

[0204] Various statistical metrics can be used to analyze the variation of features across a population with the goal of selecting features that may predict classification and therefore be advantageous for training. A further example can be the selection of a particular type of model based on the analysis of the feature space and the selected features used in the feature vector.

[0205] A. Creating a feature vector The feature vector can be created as any data structure that can be reproduced for each training sample so that corresponding data appears in the same place in the data structure across training samples. For example, the feature vector can be associated with an index, with a particular value present at each index. As explained above, a matrix can be stored at a particular index of the feature vector, and elements of the matrix can have further sub-indices. Other elements of the feature vector can be generated from summary statistics of such a matrix.

[0206] As another example, a single element of a feature vector may correspond to a set of sequence reads across a set of genome windows. Thus, an element or feature vector may itself be a vector. Such read counts may be for all reads or a specific group (class) of reads, for example, reads with a specific sequence complexity or entropy. The set of sequence reads may be filtered or normalized for GC bias and / or mappability bias, etc.

[0207] In some examples, an element of a feature vector may be the result of a concatenation of multiple features. This may differ from other examples where the element itself is an array (e.g., a vector or matrix) in that the concatenated value can be treated as a single value, as opposed to a collection of values. Thus, features can be concatenated, combined, and combined to be used as manipulated features or feature representations for machine learning models.

[0208] Multiple combinations and approaches for merging features can be performed. For example, when different measures are counted over the same window (bin), the ratio between those bins (e.g., inversions divided by deletions) can be a useful feature. In addition, the ratio of spatially adjacent bins whose combination can convey biological information, such as the number of transcription start sites divided by the number of gene bodies, can also serve as a useful feature.

[0209] Features can also be manipulated, for example, by setting up a multitask unsupervised learning problem in which the joint probability of all feature vectors is maximized given a set of parameters and latent vectors. The latent vectors in this stochastic procedure often serve as good features when trying to predict phenotypes (or other classifications) from biological sequence data.

[0210] B. Weights used for training Weights can be applied to features when they are added to a feature vector. Such weights can be based on elements in the feature vector or specific values within elements of the feature vector. For example, every region (window) in the genome can have a different weight. Some windows can have a weight of zero, meaning the window does not contribute to the classification. Other windows can have a larger weight, e.g., between 0 and 1. Thus, weighting masks can be applied to feature values used to create the feature vector, e.g., different values of the mask applied to features for population count, sequence complexity, frequency, sequence similarity, etc.

[0211] In some embodiments, the training process can learn the weights to apply. In this way, any prior knowledge or biological insight into the data need not be known prior to the training process. The weights initially applied to the features can be considered part of the first layer of the model. Once the model is trained and one or more specified criteria are met (e.g., a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof), the model can be used in a production run to classify new samples. In such a production run, any features with initial weights of zero do not need to be calculated. Thus, the size of the feature vector can be reduced from training to production. In some embodiments, principal component analysis (PCA) can be used to train a machine learning model. For machine learning models, in various embodiments, each principal component can be a feature, or all principal components concatenated together can be a feature. A model can be created based on the output of PCA for each of these analytes. The model can be updated based on the raw features prior to PCA (not necessarily the PCA output). In various approaches, the raw features can use every bit of the data, can be done by taking a random selection of each batch of data, doing a random forest, or creating other trees or random datasets. The features can be the measurements themselves, or both, as opposed to the results of any dimensionality reduction.

[0212] C. Feature Selection Between Training Iterations As mentioned above, the training process may not produce a model that meets the desired criteria. At such times, feature selection may be performed again. Because the feature space may be quite large (e.g., 35 or 100,000), the number of possible permutations in which the differential features used in the feature vector differ may be enormous. Particular features (potentially many) may belong to the same class (type), e.g., read counts within a window, ratios of counts from different regions, variants at different sites, etc. Furthermore, features may be concatenated into a single element to further increase the number of permutations.

[0213] The set of new features can be selected based on information from previous iterations of the training process. For example, the weights associated with the features can be analyzed. These weights can be used to determine whether to keep or discard features. Features associated with weights above a threshold or average weights can be retained. Features associated with weights below a threshold or average weights (either the same or different ones can be kept) can be removed.

[0214] The selection of features and the creation of feature vectors for training the models can be repeated until one or more desired criteria are met, such as a quality metric suitable for the model (e.g., a desired minimum accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof). Another criterion may be to select a model with the optimal quality metric from a set of models generated with different feature vectors. Thus, a model with the best statistical performance and generalizability in its ability to detect phenotypes from data can be selected. Furthermore, a set of training samples can be used to train various models for different purposes, such as classifying conditions (e.g., individuals with or without cancer), treatments (e.g., individuals with or without a treatment response), prognosis (e.g., individuals with a favorable prognosis), etc. A good cancer prognosis can correspond to when an individual has the potential for resolution or improvement of symptoms or is expected to recover after treatment (e.g., the tumor shrinks or the cancer is not expected to recur), and as used herein, a good cancer prognosis refers to a prognosis associated with less aggressive and / or more treatable forms of the disease. For example, less aggressive and more treatable forms of cancer have a higher expected survival than more aggressive and / or less treatable forms. In various examples, a good prognosis refers to a tumor that remains the same or decreases in size in response to treatment, remission, or improved overall survival.

[0215] Similarly, a poor prognosis (or an individual without a good prognosis), as used herein, refers to a prognosis associated with a more aggressive and / or less treatable form of the disease. For example, an aggressive, less treatable form has a lower survival rate than a less aggressive and / or less treatable form. In various examples, a poor prognosis refers to the tumor remaining the same or increasing in size, or the cancer recurring or not decreasing.

[0216] VI. Using Machine Learning Models for Multi-analyte Assays 2 illustrates an exemplary method 200 for analyzing a biological sample, according to one embodiment. Method 200 may be implemented by any of the systems described herein. In one embodiment, the method uses a machine learning model capable of class discrimination in a population of individuals. In various embodiments, this model (e.g., a classifier) capable of class discrimination is used to discriminate between healthy and diseased populations, between treatment responders and non-responders, and between stages of disease, providing information useful for guiding treatment decisions.

[0217] In block 210, the system receives a biological sample containing multiple classes of molecules. Exemplary biological samples, such as blood, plasma, or urine, are described herein. Individual samples can also be received. A single sample (e.g., blood) may be collected in multiple containers, such as a set of vials.

[0218] In block 220, the system separates the biological sample into multiple portions, with each of the multiple classes of molecules being in one of the multiple portions. The sample may already be a plasma fraction obtained from a larger sample, for example, a blood sample. The portions can then be obtained from such fractions. In some embodiments, the portions can include multiple classes of molecules. An assay in a portion may test for only one class of molecules; thus, molecules of a class may not be measured in one portion, but may be measured in a different portion. For example, measurement devices 151, 152, and 153 may perform respective assays on different portions of the sample. Computer system 101 may analyze the data measured from the various assays.

[0219] In block 230, for each of the plurality of assays, the system identifies a set of features to be input into the machine learning model. The set of features can correspond to the characteristics of one of the plurality of classes of molecules in the biological sample. Definitions of the set of features to be used can be stored in the memory of the computer system. The set of features can be pre-identified, for example, using the machine learning techniques described herein. When a particular assay is used, the corresponding set of features can be retrieved from memory. Each assay can have an identifier used to obtain the corresponding set of features, along with any specific software code for creating the features. Such code can be modularized so that sections can be updated independently, with the final set of features defined based on the assay used and the stored definitions of the various sets of features.

[0220] In block 240, for each portion of the plurality of portions, the system performs assays for classes of molecules in that portion to obtain a set of measurements of the classes of molecules in the biological sample. The system can obtain multiple sets of measurements for the biological sample from the multiple assays. Depending on which assays are specified (e.g., via an input file or a user-specified measurement configuration), specific measurements can be provided to the computer system using a specific set of measurement devices.

[0221] In block 250, the system forms a feature vector of features from the multiple sets of measurements. Each feature corresponds to a feature and can include one or more measurements. The feature vector can include at least one feature formed using each of the multiple sets of measurements. Thus, the feature vector can be determined using measured values from each of the assays for different classes of molecules. Other details for the formation of the feature vector and the extraction of the feature vector are described in other sections, but apply to all examples for the formation of the feature vector.

[0222] The features of a given analyte can be determined using principal component analysis. For machine learning models, in various embodiments, each principal component can be a feature, or all principal components concatenated together can be a feature. A model can be created based on the output of PCA for each of these analytes. In other examples, the model can be updated based on raw features prior to any PCA, and therefore the features may not necessarily include any PCA output. In various approaches, the raw features can include every bit of data, or a random sampling of each batch of data for the analyte can be used, a random forest can be performed, or other trees or random datasets can be created. Features can also be the measurements themselves, as opposed to the results of any dimensionality reduction (e.g., PCA), or both can be used.

[0223] In block 260, the system loads into the computer system's memory a machine learning model trained using training vectors obtained from training biological samples. The training samples may be identically measured and therefore generate the same feature vector. The training samples may be selected based on a desired classification, for example, as indicated by a clinical question. Different subsets may have different characteristics, for example, as determined by their assigned labels. A first subset of the training biological samples may be identified as having a specified property, and a second subset of the training biological samples may be identified as not having the specified property. Examples of characteristics are various diseases or disorders, but may also be intermediate classifications or measurements. Examples of such characteristics include, for example, the presence or stage of cancer, or the prognosis of cancer, for cancer treatment. By way of example, the cancer may be colon cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer.

[0224] In block 270, the system inputs the feature vector into a machine learning model to obtain an output classification of whether the biological sample has a specified property. The classification can be provided in various ways, such as a probability for each of one or more classifications. For example, the presence of cancer can be assigned a probability and an output. Similarly, the absence of cancer can be assigned a probability and an output. The classification with the highest probability can be used and, for example, can be subjected to one or more criteria to ensure that one classification has a sufficiently higher probability than the second-highest classification. The difference may need to exceed a threshold. If one or more criteria are not met, the output classification can be undetermined. Thus, the output classification can include a detection value (e.g., a probability) indicating the presence of cancer in the individual. The machine learning model can then further output another classification providing a probability that the biological sample does not have cancer.

[0225] After such classification, the subject can be provided with a treatment. Examples of treatment regimens include surgical intervention, chemotherapy with a given drug or drug combination, and / or radiation therapy.

[0226] VII.Classifier generation The disclosed methods and systems relate to identifying a set of informative features (e.g., genetic loci) that correlate with class distinctions between samples, including sorting features (e.g., genes) by the degree to which their presence in a sample correlates with the class distinction and determining whether the correlation is stronger than expected by chance. Machine learning techniques can implicitly use such informative features from input feature vectors. In one example, the class distinction is a known class distinction, and in one example, the class distinction is a disease class distinction. In particular, the disease class distinction can be a cancer class distinction. In various examples, the cancer is colon cancer, lung cancer, liver cancer, or pancreatic cancer.

[0227] Some embodiments of the present disclosure may also be directed to ascertaining at least one previously unknown class (e.g., disease class, proliferative disease class, cancer stage, or treatment response) into which at least one sample tested is classified, the sample being obtained from an individual. In one aspect, the present disclosure provides a classifier capable of identifying individuals within a population of individuals. The classifier may be part of a machine learning model. The machine learning model may receive as input a set of features corresponding to the properties of each of multiple classes of molecules in a biological sample. The multiple classes of molecules in the biological sample may be assayed to obtain multiple sets of measurements representing the multiple classes of molecules. A set of features corresponding to the properties of each of the multiple classes of molecules may be identified and input into the machine learning model. A feature vector of features may be generated from each of the multiple sets of measurements, such that each feature corresponds to a feature of the set of features and includes one or more measurements. The feature vector may include at least one feature obtained using each of the multiple sets of measurements. The machine learning model including the classifier may be loaded into computer memory. The machine learning model may be trained using training vectors obtained from training biological samples, such that a first subset of the training biological samples are identified as having a specified property and a second subset of the training biological samples are identified as not having the specified property. The feature vectors may be input into the machine learning model to obtain an output classification of whether the biological sample has the specified property, thereby identifying a population of individuals having the specified property. In one example, the specified property is whether the individual has cancer.

[0228] In one aspect, the present disclosure provides a system for classifying a subject based on multi-analyte analysis of a biological sample, comprising: (a) a computer-readable medium including a classifier operable to classify the subject based on the multi-analyte analysis; and (b) one or more processors for executing instructions stored on the computer-readable medium.

[0229] In one embodiment, the system includes a classification circuit configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and an NMF prediction algorithm classifier.

[0230] In one example, informative features (e.g., genomic loci) of biomarkers are assayed in a cancer sample (e.g., tissue) to form a profile. The threshold of the scalar output of the linear classifier is optimized to maximize accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, e.g., the sum of sensitivity and specificity under cross-validation, as observed in the training dataset.

[0231] The overall multi-analyte assay data (e.g., expression data or sequence data) for a given sample can be normalized using methods known to those skilled in the art to correct for different amounts of starting material, various efficiencies of extraction and amplification reactions, etc. Using a linear classifier on normalized data to effectively make a diagnostic or prognostic call (e.g., responsiveness or resistance to a therapeutic agent) means dividing the data space into two disjoint parts, e.g., by a separating hyperplane, all possible combinations of expression values for all features (e.g., genes) in the classifier. This division is derived experimentally from a large set of training examples, e.g., from patients exhibiting responsiveness or resistance to a therapeutic agent. Without loss of generality, a specific fixed set of values can be assumed for all biomarkers except one, and thresholds for this remaining biomarker can be automatically defined, with the determination varying depending, e.g., on responsiveness or resistance to a therapeutic agent. Expression values above this dynamic threshold can then indicate either resistance to the therapeutic agent (for biomarkers with negative weights) or responsiveness (for biomarkers with positive weights). While the exact value of this threshold depends on the actual measured expression profiles of all other biomarkers in the classifier, the general indicator of a particular biomarker remains fixed; for example, high values or "relative overexpression" always contribute to either responsiveness (genes with positive weights) or resistance (genes with negative weights). Thus, in the context of a global gene expression classifier, relative expression can indicate whether upregulation or downregulation of a particular biomarker is indicative of responsiveness or resistance to a therapeutic agent.

[0232] In one example, the biomarker profile (e.g., expression profile) of a patient's biological (e.g., tissue) sample is evaluated by a linear classifier. As used herein, a linear classifier refers to the weighted sum of individual biomarker features into a compound decision score ("decision function"). The decision score is then compared to a predetermined cutoff score threshold corresponding to a particular set value in terms of accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, indicating whether the sample is above (positive decision function) or below (negative decision function) the score threshold. Effectively, this means that the data space, e.g., the set of all possible combinations of biomarker features, is divided into two mutually exclusive halves corresponding to different clinical classifications or predictions, e.g., one half corresponding to responsiveness to a therapeutic agent and the other half corresponding to resistance.

[0233] This quantity, i.e., the interpretation of the cutoff threshold responsiveness or resistance to a therapeutic agent, is derived from a series of patients with known outcomes during the development phase ("training"). Corresponding weights for the decision scores and the responsiveness / resistance cutoff thresholds are pre-fixed from the training data by methods known to those skilled in the art. In one example, partial least squares discriminant analysis (PLS-DA) is used to determine the weights. (L. Stale, S. Wold, J. Chem. 1 (1987) 185-196; D.V. Nguyen, D.M. Rocke, Bioinformatics 18 (2002) 39-50). Other methods for classification known to those skilled in the art may be used in conjunction with the methods described herein when applied to assay data (e.g., transcripts) for cancer classifiers.

[0234] Different methods can be used to translate the quantitative assay data measured with these biomarkers into prognostic or other predictive uses. These methods include, but are not limited to, pattern recognition (Duda et al., Pattern Classification, 2nd ed., John Wiley, New York 2001), machine learning (Scholkopf et al., Learning with Kernels, MIT Press, Cambridge 2002; Bishop, Neural Networks for Pattern Recognition, Clarendon Press, Oxford 1995), statistics (Hastie et al., The Elements of Statistical Learning, Springer, New York 2001), bioinformatics (Dudoit et al., 2002, J. Am. Statist. Assoc. 97:77-87; Tibshirani et al., 2002, Proc. Natl. Acad. Sci. USA 99:6567-6572), or chemometrics (Vandeginste, et al., Handbook of Chemometrics and Examples of methods include those in the field of analytical methods (e.g., Qualimetrics, Part B, Elsevier, Amsterdam 1998).

[0235] In the training step, a set of patient samples for both responsive and resistant cases (e.g., including patients who respond to treatment, patients who do not respond to treatment, patients who are resistant to treatment, and / or patients who are not resistant to treatment) are measured, and the unique information from this training data is used to optimize the prediction method to best predict the training set or future sample sets. In this training step, the method is trained or parameterized to predict from a specific assay data profile to a specific predictive call. Appropriate transformations or preprocessing steps may be performed with the measurement data before it is subjected to a classification (e.g., diagnostic or prognostic) method or algorithm.

[0236] For each of the assay data (e.g., transcripts), a weighted sum of the preprocessed feature (e.g., intensity) values is formed and compared to a threshold optimized on the training set (Duda et al., Pattern Classification, 2014). nd ed., John Wiley, New York 2001). The weights can be derived by a number of linear classification methods, including but not limited to partial least squares (PLS, (Nguyen et al., 2002, Bioinformatics 18 (2002) 39-50)) or support vector machines (SVM, (Scholkopf et al. Learning with Kernels, MIT Press, Cambridge 2002)).

[0237] The data may be nonlinearly transformed before applying the weighted sum as described above. This nonlinear transformation may include increasing the dimensionality of the data. The nonlinear transformation and weighted sum may also be performed implicitly, for example, through the use of a kernel function (Scholkopf et al. Learning with Kernels, MIT Press, Cambridge 2002).

[0238] In another example, decision trees (Hastie et al., The Elements of Statistical Learning, Springer, New York 2001) or random forests (Breiman, Random Forests, Machine Learning 45:5 2001) are used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).

[0239] In another example, neural networks (Bishop, Neural Networks for Pattern Recognition, Clarendon Press, Oxford 1995) are used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).

[0240] In another example, discriminant analysis (Duda et al., Pattern Classification, 2nd ed., John Wiley, New York 2001) uses methods including linear, diagonal, quadratic, and logistic discriminant analysis to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).

[0241] In another example, predictive analysis of microarrays (PAM, (Tibshirani et al., 2002, Proc. Natl. Acad. Sci. USA 99:6567-6572)) is used to make classifications (e.g., diagnostic or prognostic calls) from assay data (e.g., transcript sets) or measurements of their products (e.g., intensity data).

[0242] In another embodiment, Soft Independent Modeling of Class Analogy (SIMCA, (Wold, 1976, Pattern Recogn. 8:127-139)) is used to make predictive calls from measured intensity data of a transcript set or its products.

[0243] Machine learning models can be used to process various types of signals and infer classifications (e.g., phenotypes or phenotype probabilities). One type of classification corresponds to a subject's condition (e.g., disease and / or disease stage or severity). Thus, in some embodiments, a model can classify subjects based on the type of condition for which the model was trained. Such conditions may correspond to the labels of training samples, or a set of categorical variables. As described above, these labels can be determined through more intensive measurements or through measurements of patients with later stages of the condition (where the condition is more easily identified).

[0244] Such models created using training samples with predetermined conditions can provide certain advantages, including (a) pre-screening of diseases or disorders (e.g., age-related diseases before the onset of symptoms or reliable detection by alternative methods), applications including, but not limited to, cancer, diabetes, Alzheimer's disease, and other diseases that may have genetic signatures, e.g., somatic genetic signatures, (b) diagnostic confirmation or supplemental evidence for existing diagnostic methods (e.g., cancer biopsies / medical imaging scans), and (c) treatment and post-treatment monitoring for prognostic reporting, treatment response, treatment resistance, and recurrence detection.

[0245] In various examples, the biological state can include a disease or disorder (e.g., an age-related disease, an aging condition, a treatment effect, a drug effect, a surgery effect, a measurable trait, or a biological state) following a lifestyle modification (e.g., a change in diet, a change in smoking, a change in sleep patterns, etc.). In some examples, the biological state may be unknown, and the classification can be determined as the absence of another state. Thus, the machine learning model can infer an unknown biological state or interpret an unknown biological state.

[0246] In some examples, there may be gradations in the classification, and thus there may be many levels of classification corresponding to conditions, e.g., real numbers. Thus, the classification may be a probability, risk, or measure for a subject having a condition or other biological state. Each such value may correspond to a different classification.

[0247] In some examples, the classification can include recommendations that may be based on previous classifications of the condition. The previous classifications can be made by separate models using the same training data (albeit potentially with different input features) or by sub-models that are part of a larger model that includes various classifications, and the output classification of one model can be used as input to another model. For example, if a subject is classified as being at high risk for myocardial infarction, the model can recommend lifestyle changes, such as regular exercise, eating a healthy diet, maintaining a healthy weight, quitting smoking, and lowering LDL cholesterol. As another example, the model can recommend clinical trials for the subject to confirm the classification (e.g., a diagnostic or prognostic call). The clinical trials can include imaging tests, blood tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest x-rays, positron emission tomography (PET) scans, PET-CT scans, or any combination thereof. Such recommended actions can be performed as part of the methods and systems described herein.

[0248] Thus, embodiments may provide many different models, each targeting a different type of classification. As another example, an initial model may determine whether a subject has cancer. A further model may determine whether a subject has a particular stage of a particular cancer. A further model may determine whether a subject has a particular cancer. A further model may classify a subject's predicted response to a particular surgery, chemotherapy (e.g., drug), radiation therapy, immunotherapy, or other type of treatment. As another example, an early model in a chain of sub-models may determine whether a particular genetic polymorphism is accurate or relevant, and then use that information to generate input features for a subsequent sub-model (e.g., later in the pipeline).

[0249] In some examples, phenotypic classification is derived from physiological processes, such as changes in cell turnover due to infection or physiological stress, which induce changes in the types and distribution of molecules that an experimenter can observe in a patient's blood, plasma, urine, etc.

[0250] Thus, some embodiments may involve active learning, where a machine learning procedure can suggest future experiments or data based on the probability of that data reducing the uncertainty in the classification. Such issues may be related to insufficient coverage of the subject's genome, lack of time point resolution, insufficient patient background sequences, or other reasons. In various embodiments, the model may suggest one of a number of follow-up steps based on missing variables, including one or more of the following: (i) whole genome sequencing (WGS) resequencing, (ii) whole chromosome sequencing (WES) resequencing, (iii) targeted sequencing of specific regions of the subject's genome, (iv) specific primers or other approaches, and (v) other wet-lab approaches. Recommendations may vary between patients (e.g., due to the subject's genetic or non-genetic data). In some examples, the analysis aims to minimize some feature, such as cost, risk, or morbidity to the patient, or maximize classification performance, such as accuracy, positive predictive value (PPV), negative predictive value (NPV), clinical sensitivity, clinical specificity, area under the curve (AUC), or a combination thereof, and suggests the best next steps to obtain the most accurate classification.

[0251] VIII. Cancer Diagnosis and Detection The machine learning methods, models, and discriminative classifiers trained herein are useful for a variety of medical applications, including cancer detection, diagnosis, and treatment response. Once models are trained with individual metadata and specimen-derived features, applications may be tailored to stratify individuals in a population and guide treatment decisions accordingly.

[0252] A. Diagnosis The methods and systems provided herein may perform predictive analysis using an artificial intelligence-based approach to analyze data acquired from a subject (patient) and generate an output of a diagnosis for the subject having cancer (e.g., colorectal cancer, CRC). For example, the application may apply a predictive algorithm to the acquired data to generate a diagnosis for the subject having cancer. The predictive algorithm may include an artificial intelligence-based predictor, such as a machine learning-based predictor, configured to process the acquired data to generate a diagnosis for the subject having cancer.

[0253] The machine learning predictor can be trained using a dataset from one or more cohorts of patients with cancer (e.g., a dataset generated by performing multi-analyte assays of individual biological samples) as input and the subject's known diagnostic (e.g., staging and / or tumor fraction) outcome as output to the machine learning predictor.

[0254] A training dataset (e.g., a dataset generated by performing a multi-analyte assay on an individual's biological samples) can be generated from one or more sets of subjects with common characteristics (features) and results (labels). The training dataset can include a set of features and labels corresponding to diagnostically relevant features. Features can include, for example, a specific range or category of cfDNA assay measurements (e.g., the count of cfDNA fragments in biological samples obtained from healthy and diseased samples that overlap or fall within a set of bins (genomic windows) of a reference genome). For example, a set of features collected from a given subject at a given time point can collectively function as a diagnostic signature that can indicate the subject's identified cancer at the given time point. Features can also include labels that indicate a subject's diagnostic outcome, such as one or more cancers.

[0255] The label may include, for example, a result, such as a known diagnosis (e.g., stage classification and / or tumor fraction) of the subject. The result may include a characteristic associated with the subject's cancer. For example, the characteristic may indicate that the subject has one or more cancers.

[0256] The training set (e.g., training dataset) may be selected by random sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). Alternatively, the training set (e.g., training dataset) may be selected by proportional sampling of a set of data corresponding to one or more sets of subjects (e.g., retrospective and / or prospective cohorts of patients with or without one or more cancers). The training set may be balanced across datasets corresponding to one or more sets of subjects (e.g., patients from different clinical sites or clinical trials). The machine learning predictor may be trained until certain predetermined conditions for accuracy or performance are met, for example, until it has a minimum desired value corresponding to a diagnostic accuracy measure. For example, the diagnostic accuracy measure may correspond to the diagnosis, staging, or prediction of tumor fraction of one or more cancers in a subject.

[0257] Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and area under the curve (AUC) of a receiver operating characteristic (ROC) curve, which correspond to the diagnostic accuracy of detecting or predicting cancer (e.g., colon cancer).

[0258] In another aspect, the disclosure provides a method of identifying cancer in a subject, comprising: (a) providing a biological sample comprising cell-free nucleic acid (cfNA) molecules from the subject; (b) sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 10 distinct regions, and each of the at least about 10 distinct regions comprises at least a portion of a gene selected from the group consisting of the genes in Table 1; and (e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has the cancer.

[0259] In some embodiments, the at least about 10 different regions include at least about 20 different regions, each of the at least about 20 different regions comprising at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 30 different regions, each of the at least about 30 different regions comprising at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 40 different regions, each of the at least about 40 different regions comprising at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 50 different regions, each of the at least about 50 different regions comprising at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 60 different regions, each of the at least about 60 different regions comprising at least a portion of a gene selected from the group of Table 1. In some embodiments, the at least about 10 different regions include at least about 70 different regions, and each of the at least about 70 different regions includes at least a portion of a gene selected from the group of Table 1. [Table 1]

[0260] For example, such a predetermined condition may be a sensitivity for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0261] As another example, such a predetermined condition may be a specificity for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0262] As another example, such a predetermined condition may be a positive predictive value (PPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0263] As another example, such a predetermined condition may be a negative predictive value (NPV) for predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) that includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0264] As another example, such a predetermined condition may be that the area under the curve (AUC) of a receiver operating characteristic (ROC) curve predicting cancer (e.g., colon cancer, breast cancer, pancreatic cancer, or liver cancer) comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0265] In some embodiments of any of the aforementioned aspects, the method further comprises monitoring disease progression in the subject, wherein the monitoring is based, at least in part, on the gene sequence signature. In some embodiments, the disease is cancer.

[0266] In some embodiments of any of the aforementioned aspects, the method further includes determining a tissue of origin of the cancer in the subject, the determination being based, at least in part, on the gene sequence signature.

[0267] In some embodiments of any of the aforementioned aspects, the method further includes estimating tumor burden in the subject, where the estimation is based, at least in part, on the gene sequence signature.

[0268] B. Treatment Response The predictive classifiers, systems, and methods described herein are useful for classifying populations of individuals (e.g., based on performing multi-analyte assays on biological samples from the individuals) for several clinical applications. Examples of such clinical applications include detecting early cancer, diagnosing cancer, classifying cancer into specific disease stages, and determining responsiveness or resistance to therapeutic agents for treating cancer.

[0269] The methods and systems described herein are applicable to various cancer types, as well as grades and stages, and are therefore not limited to a single cancer disease type. Thus, a combination of specimens and assays may be used in the systems and methods to predict response to cancer therapeutics across different cancer types in different tissues and to classify individuals based on treatment response. In one example, the classifiers described herein can stratify groups of individuals into treatment responders and treatment non-responders.

[0270] The present disclosure also provides methods for determining drug targets (e.g., genes associated / significant to a particular class) for a condition or disease of interest, comprising evaluating a sample obtained from an individual for gene expression levels of at least one gene and using adjacent analysis routines to determine genes associated with the sample classification, thereby confirming one or more drug targets associated with the classification.

[0271] The present disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, comprising obtaining a sample from an individual having the disease class, subjecting the sample to the drug, evaluating the drug-exposed sample for gene expression levels of at least one gene, and classifying the drug-exposed sample into a disease class as a function of the relative gene expression levels of the sample to the gene expression levels of the model using a computer model constructed with a weighted voting scheme.

[0272] The present disclosure also provides a method for determining the efficacy of a drug designed to treat a disease class, comprising: subjecting an individual to the drug; obtaining a sample from the individual subjected to the drug; evaluating the sample for a gene expression level of at least one gene; and classifying the sample into a disease class using a model constructed with a weighted voting scheme; and evaluating the gene expression level of the sample compared to the gene expression level of the model.

[0273] Yet another application is a method for determining whether an individual belongs to a phenotypic class (e.g., intelligence, response to treatment, lifespan, likelihood of viral infection, or obesity) comprising obtaining a sample from the individual, assessing the sample for gene expression levels of at least one gene, and classifying the sample into a disease class using a model constructed with a weighted voting scheme, and assessing the gene expression levels of the sample compared to the gene expression levels of the model.

[0274] There is a need to identify biomarkers useful for predicting the prognosis of patients with colon cancer. The ability to classify patients as high-risk (poor prognosis) or low-risk (good prognosis) may allow for the selection of appropriate therapy for these patients. For example, high-risk patients are more likely to benefit from aggressive treatment, while treatment may not have significant benefits for low-risk patients. However, despite this need, there has been no solution to this problem.

[0275] Predictive biomarkers that can guide treatment decisions have been sought to identify subsets of patients who may be "exceptional responders" to particular cancer therapies, or individuals who may benefit from alternative treatment modalities.

[0276] In one aspect, the systems and methods described herein that relate to classifying populations based on treatment response refer to cancers that are treated with, but are not limited to, the following classes of chemotherapeutic agents: DNA damaging agents, DNA repair target therapy, inhibitors of DNA damage signal transduction, inhibitors of DNA damage-induced cell cycle arrest, and inhibitors of processes that are indirectly linked to DNA damage.Each of these chemotherapeutic agents, as the term is used herein, is considered to be " DNA damaging therapeutic agents ".

[0277] Patient specimen data can be classified into high-risk and low-risk patient groups, such as patients at high or low risk of clinical recurrence, and the results can be used to determine the course of treatment. For example, patients determined to be high-risk can be treated with adjuvant chemotherapy after surgery. For patients considered low-risk, adjuvant chemotherapy can be withheld after surgery. Thus, in certain aspects, the present disclosure provides a method for generating gene expression profiles of colon cancer tumors that indicate the risk of recurrence.

[0278] In various examples, the classifiers described herein can stratify a population of individuals between responders and non-responders to a treatment.

[0279] In various embodiments, the treatment is selected from an alkylating agent, a plant alkaloid, an antitumor antibiotic, an antimetabolite, a topoisomerase inhibitor, a retinoid, a checkpoint inhibitor therapy, or a VEGF inhibitor.

[0280] Examples of treatments that may stratify populations into responders and non-responders include sorafenib, regorafenib, imatinib, eribulin, gemcitabine, capecitabine, pazopanib, lapatinib, dabrafenib, sutiginib malate, crizotinib, everolimus, trisirolimus, sirolimus, axitinib, gefitinib, anastrol, bicalutamide, fulvestrant, ralitrexed, pemetrexed, gosserin acetate, erlotinib, vemurafenib, Biciophenib, tamoxifen citrate, paclitaxel, docetaxel, cabazitaxel, oxaliplatin, ziv-aflibercept, bevacizumab, trastuzumab, pertuzumab, pancitumab, taxanes, bleomycin, melphalen, plumbagin, camptosar, mitomycin-C, mitoxantrone, smanthus, doxorubicin, pegylated doxorubicin, forfoli, 5-fluorouracil, temozolomide, pasireotide tegafur, gimeracil, otelas, itraconazole, bortezomib, lenalidomide, irinototecan, epirubicin, and romidepsin, resminostat, tasquinimod, refametinib, lapatinib, Tyverb, Arendil, pasireotide, Signifor, ticilimumab, tremelimumab, lansoprazole, PrevOnco, ABT-869, linifanib, borolanib, tivantinib, Tarceva, erlotinib, Stivarga, and Lego Chemotherapeutic agents including rafenib, fluorosorafenib, brivanib, liposomal doxorubicin, lenvatinib, ramucirumab, peretinoin, Ruchiko, mpafostum, TS-1, tegafur, gimeracil, oteracil, and orantinib; and antibody therapies including alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, or rituximab.

[0281] In other examples, populations can be stratified into responders and non-responders to checkpoint inhibitor therapy, such as compounds that bind to PD-1 or CTLA4.

[0282] In other examples, populations can be stratified into responders and non-responders to anti-VEGF therapies that bind to targets in the VEGF pathway.

[0283] IX. Efficacy In some examples, the biological state may include a disease. In some examples, the biological state may be a stage of a disease. In some examples, the biological state may be a gradual change in a biological state. In some examples, the biological state may be a treatment effect. In some examples, the biological state may be a drug effect. In some examples, the biological state may be a surgical effect. In some examples, the biological state may be a biological state after a lifestyle modification. Non-limiting examples of lifestyle modifications include a change in diet, a change in smoking, and a change in sleep patterns.

[0284] In some examples, the biological state is unknown. The analyses described herein can include machine learning to infer the unknown biological state or to interpret the unknown biological state.

[0285] In one embodiment, the present system and method are particularly useful for applications related to colon cancer. This cancer forms in colon tissue (the longest part of the large intestine). Most colon cancers are adenocarcinomas (cancers that begin in the cells that line the internal organs and have glandular characteristics). The progression of cancer is characterized by the stage or extent of the cancer in the body. Staging is usually based on the size of the tumor, whether lymph nodes contain cancer, and whether the cancer has spread from the original site to other parts of the body. Stages of colon cancer include stage I, stage II, stage III, and stage IV. Unless otherwise specified, the term colon cancer refers to colon cancer at stage 0, stage I, stage II (including stage IIA or IIB), stage III (including stage IIIA, IIIB, or IIIC), or stage IV. In some embodiments herein, colon cancer originates from any stage. In one embodiment, the colon cancer is stage I colon cancer. In one embodiment, the colon cancer is stage II colon cancer. In one embodiment, the colon cancer is stage III colon cancer. In one embodiment, the colon cancer is stage IV colon cancer.

[0286] Conditions that may be predicted by the methods of the present disclosure include, for example, cancer, gut-related diseases, immune-mediated inflammatory diseases, neurological diseases, renal diseases, prenatal diseases, and metabolic diseases.

[0287] In some examples, the methods of the present disclosure can be used to diagnose cancer.

[0288] Non-limiting examples of cancer include adenoma (adenomatous polyp), sessile serrated adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0289] Non-limiting examples of cancers that may be predicted by the disclosed methods and systems include acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adrenocortical carcinoma, Kaposi's sarcoma, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, osteosarcoma, malignant fibrous histiocytoma, brain stem glioma, brain cancer, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pine nut tumor, breast cancer, bronchial tumor, Burkitt's lymphoma, non-Hodgkin's lymphoma, carcinoid tumor, and cervical cancer. , chordoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), colon cancer, colorectal cancer, cutaneous T-cell lymphoma, ductal carcinoma in situ, endometrial cancer, esophageal cancer, Ewing's sarcoma, eye cancer, intraocular melanoma, retinoblastoma, fibrotic Histiocytoma, gallbladder cancer, gastric cancer, glioma, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular (liver) cancer, Hodgkin lymphoma, hypopharyngeal cancer, kidney cancer, laryngeal cancer, lip cancer, oral cancer, lung cancer, non-small cell cancer, small cell cancer, melanoma, oral cancer cancer), myelodysplastic syndrome, multiple myeloma, medulloblastoma, nasal cancer, paranasal sinus cancer, neuroblastoma, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, papillomatosis, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pituitary tumors, plasma cell neoplasms, prostate cancer, rectal cancer, renal cell carcinoma, rhabdomyosarcoma, salivary gland cancer, Sézary syndrome, skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, testicular cancer, pharyngeal cancer, thymoma, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Waldenstrom's macroglobulinemia, and Wilms' tumor.

[0290] Non-limiting examples of bowel-related diseases that may be predicted by the disclosed methods and systems include Crohn's disease, colitis, ulcerative colitis (UC), inflammatory bowel disease (IBD), irritable bowel syndrome (IBS), and celiac disease. In some examples, the disease is inflammatory bowel disease, colitis, ulcerative colitis, Crohn's disease, microscopic colitis, collagenous colitis, lymphocytic colitis, diversion colitis, Behcet's disease, and undifferentiated colitis.

[0291] Non-limiting examples of immune-mediated inflammatory diseases that can be inferred by the disclosed methods and systems include psoriasis, sarcoidosis, rheumatoid arthritis, asthma, rhinitis (hay fever), food allergies, eczema, lupus, multiple sclerosis, fibromyalgia, type 1 diabetes, and Lyme disease. Non-limiting examples of neurological diseases that can be inferred by the disclosed methods and systems include Parkinson's disease, Huntington's disease, multiple sclerosis, Alzheimer's disease, stroke, epilepsy, neurodegeneration, and neuropathy. Non-limiting examples of renal diseases that can be inferred by the disclosed methods and systems include interstitial nephritis, acute renal failure, and nephropathy. Non-limiting examples of prenatal diseases that can be inferred by the disclosed methods and systems include Down syndrome, aneuploidy, spina bifida, trisomy, Edwards syndrome, teratoma, sacrococcygeal teratoma (SCT), ventriculomegaly, renal agenesis, cystic fibrosis, and fetal hydrops. Non-limiting examples of metabolic diseases that can be inferred by the disclosed methods and systems include cystinosis, Fabry disease, Gaucher disease, Lesch-Nyhan syndrome, Niemann-Pick disease, phenylketonuria, Pompe disease, Tay-Sachs disease, von Gierke disease, obesity, diabetes, and heart disease.

[0292] Specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the disclosed embodiments of the invention. However, other embodiments of the invention may be directed to particular embodiments related to each individual aspect, or particular combinations of these individual aspects. All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes.

[0293] X. Working Example The foregoing description and the following examples of the present invention have been presented for purposes of illustration and description, and are not intended to be exhaustive or to limit the invention to the precise form described, many modifications and variations are possible in light of the above teaching.

[0294] A. Example 1: Preparation of Multi-Analyte Assays for Biological Samples This example provides a multi-analyte approach to exploit the independent information between signals. A process diagram is described below for the different components of the system for assaying with corresponding machine learning models to perform accurate classification. The selection of which assay to use can be integrated based on the results of training the machine learning model, taking into account the clinical goals of the system. Various classes of samples, fractions of samples, parts of those fractions / samples with different classes of molecules, and types of assays can be used.

[0295] 1. System diagram 3 shows an overall framework 300 of the disclosed systems and methods. The framework 300 can use measurements of a sample (wet lab 320) and other data about a subject, in combination with machine learning, to identify a set of assays and features for classifying the subject, e.g., diagnostic or prognostic. In this example, the process steps may be as follows:

[0296] In block 311 of stage 310, a question of clinical, scientific, and / or commercial relevance is asked, such as early colorectal cancer detection for viable follow-up. In block 312, subjects (new or previously tested) are identified. The subjects may have known classifications (labels) for later use in machine learning. Thus, different cohorts may be identified. In block 313, the analysis can select the type of sample to be mined (i.e., the sample may not remain in the final assay) and determine the collection of biomolecules in each sample (e.g., blood) that can generate a signal sufficient to assess the presence or absence of a pathology / disorder (e.g., early colorectal cancer malignancy). Constraints can be imposed on the assay / model, for example, related to accuracy. Exemplary constraints include the minimum sensitivity of the assay, the minimum specificity of the assay, the maximum cost of the assay, the time available to develop the assay, available biological materials and expected enrichment rates, the available set of previously developed processes that determine the maximum set of experiments that can be performed on those biological materials, and available hardware that limits the number of processes that can be performed on those biological materials to obtain data.

[0297] Patient cohorts can be designed and sampled to accurately represent the different classifications needed to adequately achieve clinical goals (healthy, colorectal and other, advanced adenoma, colorectal cancer (CRC)). Patient cohorts can be selected, and the selected cohort can be viewed as a system constraint. An example cohort is 100 CRC, 200 advanced adenoma, 200 non-advanced adenoma, and 200 healthy subjects. The selected cohort can correspond to the population intended for use in the final assay, and the cohort can specify the number of samples for which assay performance is calculated.

[0298] Once a cohort is selected, samples may be collected to meet the cohort design. Various samples may be collected, such as blood, cerebrospinal fluid (CSF), and other samples as described herein. Such analysis may be performed in block 313 of FIG. 3.

[0299] In stage 320, wet lab experiments can be performed for the first set of assays. For example, an unconstrained test set can be selected (primary sample / analyte / test combination). Protocols and modalities for isolating analytes from primary samples can be performed. Protocols and modalities for test execution can be generated. Performance of wet lab activities can be performed using hardware devices including sequencers, fluorescence detectors, and centrifuges.

[0300] In block 321, the sample is divided into smaller components (also called fractions or portions) by, for example, centrifugation. As an example, blood is divided into fractions of plasma, buffy coat (white blood cells and platelets), serum, red blood cells, and extracellular vesicles such as exosomes. The fractions (e.g., plasma) can be divided into aliquots to assay different analytes. For example, different aliquots are used to extract cfDNA and cfRNA. Thus, analytes can be isolated from the fractions or aliquots to enable multi-analyte assays. A fraction (e.g., some plasma) can be retained to measure protein concentration.

[0301] In block 323, experimental procedures are performed to measure the characteristics and amounts of the above molecules in each fraction, such as (1) the sequence and assigned location along the genome of the cell-free DNA fragments found in the plasma, (2) the methylation pattern of the cfDNA fragments found in the plasma, (3) the amount and type of microRNAs found in the plasma, and (4) the concentrations of proteins known to be associated with CRC from the literature (CRP, CEA, FAP, FRIL, etc.).

[0302] Each QC of a sample processed on any given pipeline can be verified. cfDNA QCs include insert size distribution, relative expression of GC bias, and spike-in barcode sequences (introduced for sample traceability). Methylation QCs include bisulfite conversion efficiency of control DNA, insert size distribution, average sequencing depth, and % overlap. miRNA QCs include insert size distribution and relative expression of normalized spike-ins. Protein QCs include standard curve linearity and control sample concentration.

[0303] The samples are then processed to obtain data for all patients in the cohort. The raw data is indexed by patient metadata. Data can be obtained from other sources and stored in the database. Data can be curated from relevant open databases such as GTEX, TCGA, and ENCODE. This includes ChIP-seq, RNA-seq, and eQTL.

[0304] In stage 340, data can be obtained from other sources, such as wearables, images, etc. Such other data corresponds to data determined external to the biological specimen. Such measurements can be heart rate, activity measurements, or other such data available from a wearable device. Imaging data can provide information such as organ size and location and identify unknown masses.

[0305] The database 330 can store data. The data can be curated from relevant open databases such as GTEX, TCGA, and ENCODE. This includes ChIP-seq, RNA-seq, and eQTL. Each subject record can include fields with measured data and subject labels, such as whether a condition is present, the severity (stage) of the condition, etc. A subject may have multiple labels.

[0306] At block 350, dry lab operations can occur. The "dry lab" work can begin with a query to a database to generate a matrix of relevant data and metadata values to perform a predictive task. Features are generated by processing the incoming data and possibly selecting a subset of relevant inputs.

[0307] In block 351, machine learning can be used to reduce the entire set of data generated from all (primary sample / analyte / test) combinations to the most predictive set of features in block 352. The accuracy metrics of different feature sets can be compared to each other to determine the most predictive set of features. In some embodiments, a set of features / models that meet an accuracy threshold can be identified, and then other constraints (e.g., cost and number of tests) can be used to select the optimal model / feature grouping.

[0308] A variety of different functions and models can be tested. Simple to complex, small to large models making different modeling assumptions can be applied to the data in a cross-validation paradigm. Simple to complex includes consideration of linear to nonlinear and non-hierarchical to hierarchical representations of features. Small to large models include consideration of the size of the basis vector space onto which the data is projected, as well as the number of interactions between features included in the modeling process.

[0309] Machine learning techniques can be used to evaluate commercial test modalities that best fit the cost / performance / commercial range, as defined by the initial question. Threshold checks can be performed. If the method applied to the holdout data set not used in cross-validation exceeds the initialized constraints, the assay is locked and put into operation. The assay can then be output at block 360.

[0310] If the threshold is not met, the assay engineering procedure loops back to either the constraint setting for possible relaxation or the wet lab to modify the parameters under which the data was acquired.

[0311] Considering the clinical question, biological constraints, budget, lab machines, etc. may constrain the problem. Cohort design can then be based on clinical samples, statistical and informative actionable items, and sample accrual rates, which are actually based on performance or prior knowledge.

[0312] 2. Sample and its sub-stratum In one example, multiple samples are collected from patients in a cohort and analyzed for multiple molecular types via multiple assays. The assay results are then analyzed by an ML model, and after selection of significant features and samples, assay results relevant to a clinically, scientifically, or commercially important question are output.

[0313] FIG. 4 shows a hierarchical overview of the multi-analyte approach used in an exemplary "liquid biopsy." In stage 401, different samples are collected. As shown, blood, CSF, and saliva are collected. In stage 402, the sample can be divided into fractions (portions), e.g., blood is shown divided into plasma, platelets, and exosomes. In stage 403, each fraction can be analyzed to measure one or more classes of molecules, e.g., DNA, RNA, and / or protein. In stage 404, each class of molecule can be subjected to one or more assays. For example, methylation and whole genome assays can be applied to DNA. For RNA, assays detecting mRNA or single-stranded RNA can be applied. For proteins, enzyme-linked immunosorbent assays (ELISAs) can be used.

[0314] In this example, collected plasma was analyzed using multi-analyte assays, including low-coverage whole-genome sequencing, CNV calling, tumor fraction (TF) estimation, whole-genome bisulfite sequencing, LINE-1 CpG methylation, 56-gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing. Whole blood was collected in K3-EDTA tubes and centrifuged twice to separate the plasma. Plasma was then divided into aliquots for cfDNA, lcWGS, WGS, WGBS, cf-miRNA sequencing, and quantitative immunoassays (either enzyme-linked immunosorbent assay [ELISA] or single molecule array [SIMOA]).

[0315] In stage 405, a learning module running on computer hardware can receive measured data from various assays of various fraction(s) of various sample(s). The learning module can provide metrics for various groupings of models / features. For example, various sets of features can be identified for each of multiple models. Different models can use different techniques, such as neural networks or decision trees. Stage 406 can select the model / feature group to use, or potentially provide instructions (commands) for performing further measurements. Stage 407 can specify the samples, fractions, and individual assays to be used as part of the overall assay used to measure and classify new samples.

[0316] 3. Iterative flow between modules Figure 5 illustrates an iterative process for designing an assay and corresponding machine learning model according to an embodiment of the present invention. Wet-lab components are shown on the left, and computational components are shown on the right. Omitted modules include external data, prior structures, clinical metadata, etc. These meta-components may flow into both the wet-lab and dry-lab (computational) components. Generally, the iterative process can include various phases, including an initialization phase, an exploration phase, a refinement phase, and a validation / confirmation phase. The initialization phase can include blocks 502-508. The exploration phase can include a first flow through blocks 512-528. The refinement phase can include an additional flow through blocks 512-528, as well as blocks 530 and 532. The validation / confirmation phase can occur using blocks 524 and 529. Various blocks may be optional or hard-coded to provide specified results; for example, in a particular model, module 518 may always be selected.

[0317] At block 502, a clinical question is received, for example, to screen for the presence of colorectal cancer (CRC). Such a clinical question may also include the number of classifications required. For example, the number of classifications may correspond to different stages of cancer.

[0318] In block 504, cohort(s) are designed. For example, the number of cohorts may equal the number of classifications, and subjects within a cohort may have the same label. Additional cohorts may be added at later stages in the process.

[0319] In one embodiment, before any biochemical testing is performed, there is an initial selection of sample and / or test. For example, whole genome sequencing can be selected to obtain information about an initial sample, such as blood. Such initial sample and initial assay can be selected based on the clinical question, for example, based on the relevant organ.

[0320] At block 506, an initial sample is obtained. The sample can be of various types, e.g., blood, urine, saliva, cerebrospinal fluid, etc. As part of obtaining the initial sample, the sample can be divided into fractions as described herein (e.g., dividing blood into plasma, buffy coat, exosomes, etc.), and those fractions can be further divided into portions having particular classes of molecules.

[0321] In block 508, one or more initial assays are performed. The initial assays can be tailored for individual classes of molecules. Some or all of the initial set of assays can be used as defaults across various clinical questions. The initial data 510 can be sent to a computer 511 to evaluate the data, determine machine learning models, and suggest additional assays to potentially perform. The computer 511 can perform the operations described in this section and other sections of this disclosure.

[0322] The data filter module 512 can filter the initial data 510 to provide one or more sets of filtered data. Such filtering can simply obtain data from different assays, but can be more complex, such as performing statistical analysis to provide measurements from the raw data, where the initial data 510 is considered raw data. The filtering can include dimensionality reduction, such as principal component analysis (PCA), non-negative matrix factorization (NMF), kernel PCA, graph-based kernel PCA, linear discriminant analysis (LDA), generalized discriminant analysis (GDA), or autoencoder. Multiple sets of filtered data can be determined from the raw data of a single assay. Different sets of filtered data can be used to determine different sets of features. In some embodiments, the data filter module 512 can take into account processing performed by downstream modules. For example, the type of machine learning model can affect the type of dimensionality reduction used.

[0323] The feature extraction module 514 can extract features using, for example, genetic data, non-genetic data, filtered data, and reference sequences. Feature extraction may also be referred to as feature engineering. Features of data obtained from an assay will correspond to characteristics of classes of molecules obtained in that assay. By way of example, features (and their corresponding features) may be measurements output from filtering, only a portion of such measurements, further statistical results of such measurements, or measurements added together. Certain features are extracted with the goal that some features have different values between different groups of subjects (e.g., different values between subjects with and without a pathological condition), thereby enabling discrimination between different groups or inference of the degree of a characteristic, condition, or trait. Examples of features are provided in Section V.

[0324] The cost / loss selection module 516 can select a particular cost function (also referred to as a loss function) to optimize in training the machine learning model. The cost function can have various terms for defining the accuracy of the current model. At this point, other constraints can be algorithmically injected. For example, the cost function can measure the number of misclassifications (e.g., false positives and false negatives) and have a scaling factor for each different type of misclassification, thereby providing a score that can be compared to a threshold to determine whether the current model is satisfactory. Such accuracy tests can also implicitly determine whether a set of features and a set of assays can provide a satisfactory model; if the set of features and assays does not provide a satisfactory model, a different set of features can be selected.

[0325] In some instances, the distribution of the data can influence the selection of a loss function for an unmonitored task, for example, to have technical control of the system, in which case the loss function can correspond to a distribution that is consistent with the input data.

[0326] The model selection module 518 can select the model(s) to use. Examples of such models include logistic regression, support vector machines with different kernels (e.g., linear or nonlinear kernels), neural networks (e.g., multi-layer perceptrons), and various types of decision trees (e.g., random forests, gradient trees, or gradient boosting techniques). Multiple models can be used, for example, models can be used sequentially (e.g., the output of one model feeds the input of another) or in parallel (e.g., using voting to determine the final classification). When multiple models are selected, these can be referred to as sub-models.

[0327] Cost functions are distinct from models, and distinct from features. These different parts of the architecture can have significant impacts on each other, but they are also defined by other components of the test design and their corresponding constraints. For example, a cost function can be defined by components including feature distributions, feature values, label distribution diversity, label types, label complexity, and the risks associated with different error types. Changes to specific features can result in changes to the model and / or cost function, and vice versa.

[0328] The feature selection module 520 can select a set of features to be used for the current iteration in training the machine learning model. In various embodiments, all features extracted by the feature extraction module 514 can be used, or only a portion of the features can be used. The feature weights of the selected features can be determined and used as training inputs. As part of the selection, some or all of the extracted features can undergo transformation. For example, weights can be applied to particular features based, for example, on the expected importance (probability) of the particular feature(s) relative to other feature(s). Other examples include dimensionality reduction (e.g., of a matrix), distribution analysis, normalization or regularization, and matrix decomposition (e.g., kernel-based discriminant analysis and nonnegative matrix factorization), which can provide a low-dimensional manifold corresponding to a matrix. Another example is transforming raw data or features from one type of instrument to that of another type of instrument, for example, when different samples are measured using different instruments.

[0329] The training module 522 can perform parameter optimization of the machine learning model, which may include submodels. Various optimization techniques can be used, such as gradient descent or the use of second derivatives (Hessians). In other embodiments, training can be implemented in a way that does not require Hessians or gradient calculations, such as dynamic programming or evolutionary algorithms.

[0330] The evaluation module 524 can determine whether the current model (e.g., as defined by a set of parameters) satisfies one or more criteria included in the output constraint(s). For example, the quality metric can measure the predictive accuracy of the model for a training set and / or a validation set of samples for which the labels are known. Such accuracy metrics can include sensitivity and specificity. The quality metric can be determined using values other than accuracy, such as the number of assays, the expected cost of the assay, and the time to perform the assay measurements. If the constraints are met, a final assay 529 can be provided. The final assay 529 can include, for example, a specific order for performing the assays on the test sample when an assay not in the default list is selected.

[0331] If the output constraints are not met, various items can be updated. For example, the set of selected features can be updated, or the set of selected models can be updated. Some or all upstream modules can be evaluated, checked, and alternatives suggested. Thus, feedback can be provided anywhere in the upstream pipeline. If the evaluation module 524 determines that the feature and model space has been sufficiently searched (e.g., exhausted) without satisfying the constraints, the process can decide to flow to additional modules to obtain new assays and / or sample types. Such decisions can be defined by constraints. For example, a user can only perform so many assays (and the associated time and cost), have so many samples, or perform an iterative loop (or several loops) many times. These constraints can contribute to stopping the test design of the current set of features, models, and assays in exchange for exceeding a minimum metric.

[0332] The assay identification module 526 can identify new assays to perform. If a particular assay is determined to be unimportant, the data may be discarded. The assay identification module 526 can receive certain input constraints that can be used to determine which assay or assays to select, for example, based on the cost or timing of performing the assays.

[0333] The sample identification module 528 can determine the new sample type (or portion thereof) to use. The selection can depend on which new assay(s) are being performed. Input constraints can also be provided to the sample identification module 528.

[0334] The assay identification module 526 and sample identification module 528 can be used when the assessment is that the assay and model do not meet the output constraints (e.g., accuracy). Discarding the assay can be performed in the next round of assay design, where that assay or sample type will not be used. The new assay or sample can be one that was previously measured, but for which no data was used.

[0335] In block 530, new sample types, or potentially more samples of the same type, are obtained, for example, to increase the number of samples in the cohort.

[0336] In block 532, a new assay can be performed, for example, based on a suggested assay from assay identification module 526.

[0337] The final assay 529 can specify, for example, the order of the assays in the set, the amount of data, the quality of the data, and the data throughput. The order of the assays can be optimized for cost and timing. The order and timing of the assays can be parameters that are optimized.

[0338] In some embodiments, the computer module can inform other parts of the wet lab steps. For example, some computer module(s) may precede the wet lab steps for some assay development procedures, such as when external data can be used to inform the starting point of the wet lab experiments. Furthermore, the output of the wet lab experiment components may feed the computer components with information such as cohort design and clinical questions. In turn, the computer results may feed back into the wet lab, such as the impact of the choice of cost function on cohort design.

[0339] 4. Designing a multi-analyte assay The overall process flow of the disclosed method is shown in Figure 6. In this example, the process steps are as follows:

[0340] In block 610, during operation, the system receives multiple training samples, each containing multiple classes of molecules, for which one or more labels are known. Examples of analytes are provided herein, such as cell-free DNA, cell-free RNA (e.g., miRNA or mRNA), proteins, carbohydrates, autoantibodies, or metabolites. The labels may be for a specific condition (e.g., cancer or a different classification of a specific cancer) or treatment response. Block 610 may be performed by a receiver including one or more measuring devices, such as measuring devices 151-153 in FIG. 1. The measuring devices may perform different assays. The measuring devices may convert the samples into usable features (e.g., a large library of information about each analyte from the sample) so that a computer can select the combination of input features required by a particular ML model to classify a particular biological sample.

[0341] In block 620, for each of a plurality of different assays, the system identifies a set of features operable to be input into the machine learning model for each of a plurality of training samples. The feature sets may correspond to molecular properties in the training samples. For example, the features may be read counts in different regions, methylation rates in regions, counts of different miRNAs, or concentrations of a set of proteins. Different assays may have different features. Block 620 may be performed by feature selection module 520 of FIG. 5. In FIG. 5, feature selection may occur before or after feature extraction, for example, if features are already known based on the type of assay being performed. As part of an iterative procedure, a new feature set may be identified, for example, based on results from evaluation module 524.

[0342] In block 630, for each of a plurality of training samples, the system subjects a group of classes of molecules in the training sample to a plurality of different assays to obtain a set of measurements. Each set of measurements may be derived from one assay applied to a class of molecules in the training sample. Multiple sets of measurements may be obtained for a plurality of training samples. As an example, the different assays may be lcWGS, WGBS, cf-miRNA sequencing, and protein concentration measurement. In one example, a portion contains more than one class of molecules, but only one type of assay is applied to that portion. The measurements may correspond to values obtained from the analysis of raw data (e.g., sequence reads). Examples of measurements include read counts of sequences that partially or completely overlap different genomic regions of the genome, methylation rates in a region, counts of different miRNAs, or concentrations of a set of proteins. Features may be determined from multiple measurements, such as statistics of the distribution of measurements or concatenation of measurements added together.

[0343] In block 640, the system analyzes the set of measurements to obtain training vectors for the training samples. The training vectors may include feature values for the corresponding set of assay features. Each feature value may correspond to a feature and may include one or more measurements. The training vectors may be formed using at least one feature from at least two of the N sets of features corresponding to a first subset of the plurality of different assays, where N corresponds to the number of different assays. A training vector may be determined for each sample, potentially including features from some or all of the assays, and therefore all classes of molecules. Block 640 may be performed by feature extraction module 514 of FIG. 5.

[0344] In block 650, the system operates on the training vectors using parameters of the machine learning model to obtain output labels for a plurality of training samples. Block 650 may be performed by a machine learning module that implements the machine learning model.

[0345] In block 660, the system compares the output labels to the known labels of the training samples. A comparator module can perform such comparison of labels to form a measure of error of the current state of the machine learning model. The comparator module can be part of the training module 522 of FIG. 5.

[0346] A first subset of the plurality of training samples may be identified as having a designated label, and a second subset of the plurality of training samples may be identified as not having the designated label. In one example, the designated label is a clinically diagnosed disorder, e.g., colon cancer.

[0347] In block 670, the system iteratively searches for optimal values for the parameters as part of training the machine learning model based on comparing the output labels with known labels of the training samples. Various techniques for performing the iterative search, such as gradient techniques, are described herein. Block 670 may be implemented by the training module 522 of FIG. 5.

[0348] Training the machine learning model can provide a first version of the machine learning model, after a refinement stage that may include, for example, one or more additional passes through modules 512-528. A quality metric can be determined for the first version, and the quality metric can be compared to one or more criteria, such as a threshold. The quality metric may be composed of various metrics, such as an accuracy metric, a cost metric, a time metric, etc., as described with respect to FIG. 4. Each of these metrics can be individually compared to a threshold, or it can be determined whether the metric satisfies one or more criteria. Based on the comparison(s), for example, in FIG. 5, blocks 526 and 532, it can be determined whether to select a new subset of assays for determining a set of features.

[0349] The new subset of assays may include at least one of a plurality of different assays that were not in the first subset, and / or may potentially remove assays. The new subset of assays may include at least one assay from the first subset, and a new set of features may be determined for one assay from the first subset. If the quality metrics of the new subset of assays meet one or more criteria, the new subset of assays may be output, for example, as final assays 529 of FIG. 5.

[0350] If the new subset includes a new assay not previously performed, the molecules in the training sample can be subjected to the new assay not included in the plurality of different assays to obtain a new set of measurements based on the quality metrics for the subset of new assays that do not meet one or more criteria. The new assay can be performed on a new class of molecules not in the group of classes of molecules.

[0351] In block 680, the system provides a set of machine learning model parameters and machine learning model features. The machine learning model parameters may be stored in a predetermined format or may be stored with tags identifying the number and identity of each of the parameters. The feature definitions may be obtained from settings used in feature extraction and selection, for example, as specified by the current iteration through feature extraction module 514 and feature selection module 520. Block 680 may be performed by an output module.

[0352] 5. How to identify cancer In one aspect, the present disclosure provides a method of identifying cancer in a subject, comprising: (a) providing a biological sample comprising cell-free nucleic acid (cfNA) molecules from the subject; (b) sequencing the cfNA molecules from the subject to generate a plurality of cfNA sequencing reads; (c) aligning the plurality of cfNA sequencing reads to a reference genome; (d) generating a quantitative measure of the plurality of cfNA sequencing reads at each of a first plurality of genomic regions of the reference genome to generate a first cfNA feature set, wherein the first plurality of genomic regions of the reference genome comprises at least about 15,000 distinct hypomethylated regions; and (e) applying a trained algorithm to the first cfNA feature set to generate a likelihood that the subject has the cancer.

[0353] In some embodiments, the trained algorithm includes performing dimensionality reduction by singular value decomposition. In some embodiments, the method further includes generating quantitative measures of the plurality of cfNA sequencing reads at each of a second plurality of genomic regions of the reference genome to generate a second cfNA feature set, where the second plurality of genomic regions of the reference genome include at least about 20,000 distinct protein-coding gene regions, and applying the trained algorithm to the second cfNA feature set to generate the likelihood that the subject has the cancer. In some embodiments, the method further includes generating quantitative measures of the plurality of cfNA sequencing reads at each of a third plurality of genomic regions of the reference genome to generate a third cfNA feature set, where the third plurality of genomic regions of the reference genome include contiguous, non-overlapping genomic regions of equal size, and applying the trained algorithm to the third cfNA feature set to generate the likelihood that the subject has the cancer. In some embodiments, the third plurality of non-overlapping genomic regions of the reference genome comprises at least about 60,000 distinct genomic regions. In some embodiments, the method further includes generating a report including information indicating the likelihood that the subject has the cancer. In some embodiments, the method further includes generating one or more recommended steps for treating the cancer for the subject based at least in part on the generated likelihood that the subject has the cancer. In some embodiments, the method further includes diagnosing the subject with the cancer when the likelihood that the subject has the cancer meets a predetermined criterion. In some embodiments, the predetermined criterion is that the likelihood is greater than a predetermined threshold. In some embodiments, the predetermined criterion is determined based on an accuracy metric of the diagnosis. In some embodiments, the accuracy metric is selected from the group consisting of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), precision, and area under the curve (AUC).

[0354] In some examples, the computer modules can inform other parts of the wet lab steps. For example, some computer modules may precede the wet lab steps for some assay development procedures, such as when external data is used to inform the starting point of the wet lab experiments. Additionally, the output of the wet lab experiment components can be fed into computer components such as cohort design and clinical questions. In turn, computer results can be fed back into the wet lab, such as the impact of the choice of cost function on cohort design.

[0355] 6.Results Table 2 shows the results for different analytes and the corresponding best performing models according to an embodiment of the present disclosure. [Table 2]

[0356] Samples were used that were similar across specimens.

[0357] In Table 2, SD refers to significant differences and is determined by comparing the read counts of different genes between different classification labels. This is part of dimensionality reduction. Features that are significantly different between two classifications are filtered out and then transferred to the classification. While PCA looks at groups of features that are grouped together but correlate in a certain way, SD looks at individual features unilaterally. The feature with the highest SD (e.g., gene read counts) can be used for the feature vector of interest. PCA involves projecting measurements through an initial small number of components. This is, for example, a condensed representation of many features in a smaller dimensional space.

[0358] The table was created by analyzing the results of different models with different dimensionality reductions (including no reduction) for different combinations of analytes. The table includes the best-performing models. As an example, for multi-analyte assay datasets involving proteins, there may be no need for PCA due to the low dimensionality (14), and therefore only logistic regression (LR) is used.

[0359] Of these models, LR was tried along with PCA (top 5 components) and showed significant differences in feature selection (retaining 10% of the features). PCA can be performed across samples or within a sample.

[0360] The feature columns correspond to different combinations of analytes, e.g., genes (cell-free DNA analysis) + methylation. When more than two analytes are used, there are two options: combine the features into one feature set, or run two models and output two classifications (e.g., probabilities of the classifications) and use them as a vote, e.g., majority vote or some weighted average or probability, to determine which classification has the highest score. Another example is to take the mean or mode of the predictions as opposed to looking at the scores.

[0361] Cross-validation was performed five times to obtain AUC information for the receiver operating characteristic curves in Figures 7A and 7B. Samples can be divided into five different datasets (four datasets for training and the fifth dataset for validation). Sensitivity and specificity can be determined for the four sets. Furthermore, the assignment to sets can be updated with a random number seed to provide more data. To determine sensitivity and specificity, the four classifications were reduced to four, with healthy and benign polyps as one classification and AA and CRC as the other classifications.

[0362] Figures 7A and 7B show the classification performance of different analytes.

[0363] B. Example 2: Analysis of Individual Assays for Classification of Biological Samples This example describes the analysis of multiple specimens and multiple assays to differentiate between healthy individuals, AA, and CRC stages.

[0364] Blood samples were separated into different fractions, and three classes of molecules were investigated with four assays: cell-free DNA, cell-free miRNA, and circulating proteins. Two assays were performed on cfDNA.

[0365] De-identified blood samples were obtained from healthy individuals and individuals with benign polyps, advanced adenomas (AA), and stage I-IV colorectal cancer (CRC). After plasma isolation, multiple specimens were assayed as follows: First, cell-free DNA (cfDNA) content was assessed by low-coverage whole-genome sequencing (lcWGS) and whole-genome bisulfite sequencing (WGBS). Next, cell-free microRNAs (cf-miRNAs) were assessed by small RNA sequencing. Finally, levels of circulating proteins and ATP were measured by quantitative immunoassays.

[0366] Sequenced cfDNA, WGBS, and cf-miRNA reads were aligned to the human reference genome (hg38) and analyzed as follows, as provided in more detail in the Materials and Methods section. cfDNA (lcWGS): Fragments aligned within annotated genomic regions were counted and normalized for sequencing depth to generate a 30,000-dimensional vector per sample, where each element corresponds to a gene count (e.g., the number of reads aligning to that gene in the reference genome). Samples with a high tumor fraction (>20%) were identified through manual inspection of large-scale CNVs.

[0367] WGBS: Percent methylation was calculated per sample across LINE-1 CpGs and CpG sites in target genes (56 genes).

[0368] cf-miRNA: Fragments that aligned to annotated miRNA genomic regions were counted and normalized for sequencing depth to generate a 1700-dimensional vector per sample.

[0369] Each of these sets of data can be filtered to identify measurements (e.g., reads aligned to a reference genome to obtain counts of reads for different genes). Measurements can be normalized. Normalization for each sample is described in more detail in a separate subsection for each sample. PCA analysis was performed for each sample and the results were obtained. The application of machine learning models is provided in a separate section.

[0370] 1.cf-DNA low-coverage whole-genome sequencing For a list of known genes with annotated regions, sequence read counts were determined for each of those annotated regions by counting the number of fragments that aligned to that region. Gene read counts can be normalized in various ways, for example, using a global expectation across the genome, within-sample normalization, and cross-feature normalization. Cross-feature normalization can refer to all of those features being averaged to a specified value, such as 0, a different negative value, 1, or a range between 0 and 2. With regard to cross-feature normalization, the total reads from a sample are variable and therefore may depend on the preparation process and sequencing load process. Normalization can be done to a fixed number of reads as part of global normalization.

[0371] For intra-sample normalization, it is possible to normalize by some features or selective features of some regions, especially GC bias. Therefore, the base pair composition of each region is different and can be used for normalization. In some cases, the GC number is significantly higher or lower than 50%, which has a thermodynamic impact because the base is more energetic and biases the process. Some regions give more reads than expected due to biological artifacts of sample preparation in the laboratory. Therefore, it may be necessary to correct such biases by applying different types of features / feature transformations / normalization methods.

[0372] Figures 8A and 8B show the distribution of high tumor fraction samples (i.e., greater than 20%) inferred by CNV across clinical stages, demonstrating the difference between healthy and normal. In this example, lcWGS of plasma cfDNA was able to identify CRC samples with high tumor fraction (>20%) based on genome-wide CNV. Furthermore, high tumor fraction was more frequently observed in late-stage CRC samples, but was also observed in some stage I and II samples. High tumor fraction was not observed in samples from healthy individuals or individuals with benign polyps or AA.

[0373] Figures 8A-8H show CNV plots for individuals with high (>20%) tumor fractions based on cfDNA-seq data. Note that each plot in Figures 8A-8H corresponds to a histogram of the autologous DNA copy number for a unique sample. Note also that tumor fractions can be calculated by estimation from CNV or by using open-source software such as ichorDNA. Table 3 shows the distribution of high tumor fraction cfDNA samples across clinical stages. [Table 3]

[0374] High tumor fraction samples do not necessarily correspond to samples clinically classified as late stage. In the figure, the total number of healthy controls is 26. "BP" refers to benign polyp, "AA" refers to advanced adenoma, and "Chr" refers to chromosome.

[0375] 2. Methylation Variable methylation regions (DMRs) are used for CpG sites. Regions can be dynamically assigned through discovery. It is possible to take several samples from different classes and discover which regions are most variably methylated across different classifications. Then, select a subset of variably methylated regions and use them for classification. The number of CpGs captured in the region is used. Regions tend to be of variable size. Therefore, a pre-discovery process can be performed to cluster several CpG sites together into regions. In this example, 56 genes and LINE1 elements (regions repeated across the genome) were studied. To perform classification, the methylation rates in these regions were investigated and used as features to train a machine learning model. In this example, classification utilizes essentially 57 features used in PCA. Specific regions can be selected based on regions with sufficient coverage across samples.

[0376] Figure 9 shows the CpG methylation analysis at LINE-1 sites, showing the differences between healthy and normal samples. The figure shows the methylation of all 57 regions used in PCA. Each data point shown for the normal sample is for a different gene region and methylation.

[0377] In this example, genome-wide hypomethylation at LINE-1 CpG loci was observed only in individuals with CRC. Hypomethylation was not observed in samples without CRC, e.g., from healthy individuals or individuals with benign polyps or AA. Note that each normal data point is for a different gene region and methylation. In one example, all reads that map to a region may be counted. The system may determine whether a read is at a methylated position, then sum the number of methylated CpGs (e.g., consecutively adjacent C and G bases) and the number of methylated CpGs, and calculate the ratio of the number of methylated CpGs to the number of methylated CpGs.

[0378] In this example, significance was assessed by one-way analysis of variance (ANOVA) followed by Sidak's multiple comparison test. Only significant adjusted P values are shown. CpG hypomethylation of LINE-1 was observed only in CRC cases. Polyp (benign polyp), AA, CRC (stages I-IV). 5mC, 5-methylcytosine.

[0379] The proportion of DNA fragments that align with the site and have methylation can be studied in the entire region of interest.For example, a gene region may have two CpG sites (for example, consecutive C and G bases adjacent to each other) for every 190 total reads (for example, 100 reads that align with the first CpG site and 90 reads that align with the second CpG site).All the reads that map to this region are found, and whether the read is methylated or not is observed.Then, the number of methylated CpGs is summed up, and the ratio of the number of methylated CpGs to the number of unmethylated CpGs is calculated.

[0380] 3. MicroRNA In this example, virtually all measurable microRNAs (miRNAs) (approximately 1700 in this example) were used as features. Measurements are related to the expression data of these miRNAs. The transcripts are of a certain size, and each transcript is stored, allowing the number of miRNAs found for each to be counted. For example, RNA sequences can be aligned to reference miRNA sequences, e.g., a set of 1700 sequences corresponding to known miRNAs in the human transcriptome. Each miRNA found can be used as its own feature, and across all samples, the set can all become a feature. Some samples have a feature value of 0 if no expression is detected for that miRNA.

[0381] Figure 10 shows cf-miRNA sequencing analysis to characterize microRNAs. After pooling reads from all samples, ranked by expression, the number of reads mapping to each miRNA is shown. miRNAs shown in red have been suggested in the literature as potential CRC biomarkers. Adapter-trimmed reads were mapped to adult human microRNA sequences (miRBase21) using bowtie2. Over 1800 miRNAs were detected in plasma samples with at least one read, while 375 miRNAs were present at higher abundance (detected with an average of ≥10 reads per sample).

[0382] In one example, all samples are collected and the read values are aggregated together. For each microRNA found in the sample, a large number of aggregated reads can be found. In this example, approximately 10 million aggregated reads were found to map to one single microRNA, with 300 microRNAs found in more than 1,000 reads, approximately 600 in more than 100 reads, 1,200 in 10 reads, and approximately 1,800 in only a single read. Note that microRNAs with higher expression ranks may provide better markers because they can provide more reliable signals with larger absolute changes.

[0383] The cf-miRNA profiles in individuals with CRC were inconsistent with those in healthy controls. In this example, miRNAs suggested in the literature as potential CRC biomarkers tended to be present at higher abundance compared to other miRNAs.

[0384] 4. Protein Protein data were normalized using a standard curve (14 proteins). Each of the 14 proteins has its own standard curve because it is essentially a unique immunoassay, and they are typically recombinant proteins that are highly stable in optimized buffers. Thus, a standard curve is generated that can be calculated in many ways. The concentration relationship is typically nonlinear. Samples are then run and calculated based on the expected fluorescence concentration in the primary sample. Measurements may be measured in triplicate, but can be reduced to 14 individual values, for example, by averaging or more complex statistical analysis.

[0385] Figures 11A and 11B-G show the distribution of circulating protein biomarkers. Figure 11A is a boxplot showing the levels of all circulating proteins analyzed, with outliers indicated by diamonds. Figures 11B-G show that protein levels differed significantly across tissue types using one-way ANOVA followed by Sidak's multiple comparison test. Only significant adjusted P values are shown. Proteins measured using SIMOA (Quanterix): ATP-binding cassette transporter A1 / G1 (A1G1), acylation-stimulating protein (C3a des Arg), cancer antigen 72-4 (CA72-4), carcinoembryonic antigen (CEA), cytokeratin fragment 21-1 (CYFRA21-1), and FRIL u-PA. Proteins measured by ELISA (Abcam): AACT, cathepsin D (CATD), CRP, cutaneous T cell-induced chemokine (CTACK), FAP, matrix metalloproteinase-9 (MMP9), SAA1.

[0386] In this example, circulating levels of alpha-1-antichymotrypsin (AACT), C-reactive protein (CRP), and serum amyloid A (SAA) proteins were elevated in CRC samples, whereas urokinase-type plasminogen activator (u-PA) levels were lower compared to healthy controls. Circulating levels of fibroblast activation protein (FAP) and Flt3 receptor-interacting lectin precursor (FRIL) proteins were elevated in AA samples, whereas CRP levels were lower compared to CRC samples.

[0387] In this example, differences can be observed between some of the ANOVA plots. For example, CRP appears predictable. FAP varies among different individuals. Therefore, although multi-analyte tests may show aggregation trends, evaluating each one individually can be difficult.

[0388] 5. Dimensionality reduction (e.g., PCA or significance difference) Principal component analysis (PCA) was performed for each sample. In one example, PCA was performed on protein, cell-free DNA, methylation, and microRNA data. Thus, four PCAs can be performed in that context.

[0389] In one example, 14 proteins can be considered as a single analyte. There are 14 measurements for a protein, and therefore 14 concentrations based on individual fluorescence. These are vectorized by 14. The output of the PCA can be, for example, component 1 explaining 31% of the variance, and component 2 explaining 17% of the variance. This allows one to identify which proteins contribute the most variance.

[0390] For lcWGS on cell-free DNA, use differences in gene count statistics (e.g., mean, median, etc.) to identify genes with the highest variance.

[0391] Figures 12A-12D show the output of PCA analysis of cf-DNA, CpG methylation, cf-miRNA, and protein counts as a function of tumor fraction. Figures 12E-12H show PCA of cf-DNA, CpG methylation, cf-miRNA, and protein counts as a function of specimen. High tumor fraction samples consistently have abnormal behavior across all four specimens examined.

[0392] In the example of Figures 12A-12D, PCA is used to isolate the distance between high and low tumor fractions. In Figures 12E-12H, it is sample classification of different specimens (normal, healthy, benign polyp, and colon cancer). The disclosed systems and methods can be used to maximize the differentiation between such classes. In this example, the abnormal profile across specimens indicated high TFs (inferred from cfDNA CNVs) rather than cancer stage. Each dot shown corresponds to a separate sample, and PCA is the highest component value.

[0393] Various implementations can be used for dimensionality reduction. Dimensionality reduction can be calculated using several different hypothesis tests, with several different criteria used to set thresholds for significance and inclusion. PCA or SVD (singular value decomposition) can be performed on correlation or covariance matrices rather than the data itself. Autoencoding or variational autoencoding can be used. Such filtering can filter out measurements with low variance (e.g., region counts).

[0394] 6. Conclusion lcWGS of plasma cfDNA was able to identify CRC samples with a high tumor fraction (>20%) based on genome-wide copy number variation (CNV). High tumor fractions were more frequently observed in late-stage cancer samples, but were also observed in some stage I and II patients. Abnormal signals in each of three other samples (cf-miRNA profiles) that were inconsistent with healthy controls, genome-wide hypomethylation at the LINE1 (long interspersed nucleotide sequence 1) CpG locus, and elevated levels of circulating carcinoembryonic antigen (CEA) and cytokeratin fragment 21-1 (CYFRA 21-1) protein were also observed in cancer patients. Surprisingly, the aberrant profiles across multiple samples indicated a high tumor fraction (as estimated from cfDNA CNVs) rather than cancer stage.

[0395] These data suggest that tumor fraction correlates with cancer stage but also has a large potential range in early samples. Previous literature on blood-based screening for cancer detection has shown inconsistencies in the claimed ability of different single specimens to detect early-stage cancer. Tumor fraction may explain the historical discrepancy, as we found that abnormal profiles between cfDNA CpG methylation, cf-miRNA, and circulating protein levels were more strongly associated with high tumor fraction than with late-stage disease. These findings suggest that some positive "early" findings may actually be "high tumor fraction" findings. The results further demonstrate that assaying multiple analytes from a single sample enables the development of classifiers for reliable detection of premalignant or early-stage disease at low tumor fractions. Such a multi-analyte classifier is described below.

[0396] C. Example 3: Identification of Hi-C-like structures using covariance of sequencing depth in two different genomic regions from cfDNA across multiple samples. This example describes how to identify Hi-C-like structures in two distinct genomic regions from cfDNA in a single sample to identify cell type of origin as a signature for the generation of multi-analyte models.

[0397] Genomic sequences from multiple cfDNA samples were segmented into non-overlapping bins of various lengths (e.g., 10-kb, 50-kb, and 1-Mb non-overlapping bins). The number of high-quality mapping fragments within each bin was then quantified. High-quality mapping fragments met a quality threshold. Pearson / Kendall / Spearman correlation was then used to calculate correlations between pairs of bins within the same chromosome or between different chromosomes. The structure scores calculated from the nuanced structures in the correlation matrix were used to generate the heatmap shown in Figure 13. A similar heatmap was generated using structure scores determined using Hi-C sequencing, as shown in Figure 14. The similarity between the two heatmaps suggests that the nuanced structures determined using covariance were similar to those determined by Hi-C sequencing. Potential technical biases due to GC bias, genomic DNA, and correlated structures in MNase digestion were excluded.

[0398] Genomic regions (larger bin sizes) were divided into smaller bins, and the correlation between the two larger bins was calculated using the Kolmogorov-Smirnov (KS) test. The KS test score provided information about Hi-C-like structures that could be used to distinguish between cancer and control groups.

[0399] Two-dimensional segmentation (HiCseg) was used to segment and call domains in the correlation structure in cfDNA and Hi-C. The two approaches yielded similar numbers of domains and highly overlapping domains.

[0400] Identification of cfDNA-specific co-emission patterns. Covariance structure analysis of cfDNA revealed mixed input signal patterns from multiple sources, including chromatin structure, genomic DNA, MNase digestion, and possible co-emission patterns of cfDNA. Deep learning was used to remove signals from other sources and retain only potential co-emission patterns of cfDNA.

[0401] The three-dimensional proximity of chromatin in cancer and non-cancer samples can be inferred from long-range spatially correlated fragmentation patterns. The fragmentation patterns of cfDNA from different genomic regions are not uniform and reflect the local epigenetic signatures of the genome. There is a high degree of similarity between long-range epigenetic correlation structures and higher-order chromatin organization. Therefore, long-range spatially correlated fragmentation patterns may reflect the three-dimensional proximity of chromatin. Using only fragment lengths in cfDNA, we generated a genome-wide map of in vivo higher-order chromatin organization inferred from co-fragmentation patterns. Fragments generated from endogenous physiological processes can reduce the potential for technical variability associated with random ligation, restriction enzyme digestion, and biotin ligation during Hi-C library preparation. Sample Collection and Preprocessing: Retrospective human plasma samples (>0.27 mL) were obtained from 45 patients diagnosed with colon cancer (colorectal cancer), 49 patients diagnosed with lung cancer, and 19 patients diagnosed with melanoma. One hundred samples were also obtained from patients without a current cancer diagnosis. In total, samples were collected from commercial biobanks in Southern and Northern Europe, as well as the United States. All samples were de-identified. Plasma samples were stored at -80°C and thawed before use. Cell-free DNA was extracted from 250 μL of plasma (spiked with a unique synthetic dsDNA fragment for sample tracking) using the MagMAX cell-free DNA isolation kit (Applied Biosystems) according to the manufacturer's instructions. Paired-end sequencing libraries were prepared using the NEBNext Ultra II DNA Library Preparation Kit (New England Biolabs) and sequenced on an Illumina NovaSeq 6000 sequencing system using dual indexing across multiple S2 or S4 flow cells at 2x51 base pairs.

[0402] Whole-genome sequencing data processing: Reads were demultiplexed and aligned to the human genome (GRCh38 with decoy, alt contigs, and HLA contigs) using BWA-MEM 0.7.15. PCR overlap fragments were removed using unique molecular identifiers (UMIs). Contamination was assessed for common SNPs identified by 1000 Genomes (IGSR) using a marginalized contamination model across all possible genotypes and contamination fractions.

[0403] Sequencing data were checked for quality and excluded from analysis if they met any of the following criteria: AT dropout >10 or GC dropout >2 (both calculated via Picard 2.10.5). Any samples suspected of contamination due to an expected allele fraction <0.99, unexpected genotype calls, or failed negative controls were manually inspected before inclusion in the dataset. Adapters were trimmed using Atropos with default parameters. Only properly paired, high-quality reads with uniquely mapped ends (mapping quality score >60) were used for all downstream analyses; PCR duplicates were not used. Only autosomes were used for all downstream analyses.

[0404] Hi-C library preparation: In situ Hi-C library preparation of whole blood cells and neutrophils was performed by using Arima Genomics Services.

[0405] Hi-C data processing: Raw Fastq files were uniformly processed through the Juicerbox command line tool v1.5.6. After filtering the reads, results with a mapping quality score above 30 were used to generate a Pearson correlation matrix and partition A / B. Principal component analysis (PCA) was calculated using the PCA function in scikit-learn 0.19.1 in Python 3.5. The first principal component was used to segment the partitions. For each chromosome, the partitions were grouped into two groups based on their sign. The group of partitions with a low average gene density was defined as partition B, and the other group was defined as partition A. Gene density was determined by gene number annotated by Ensemble v84. Sequencing summary statistics and associated metadata information are shown in Table 4. [Table 4]

[0406] Multi-sample cfHi-C: 500 kb bins with a mappability of less than 0.75 were removed for downstream analysis. First, each 500 kb bin was divided into 50 kb subbins. The median fragment lengths of each subbin were first aggregated into 500 kb bins, and then normalized using the mean and standard deviation for each chromosome and each sample by the z-score method. Pearson correlations between each pair of bins across all individuals were calculated.

[0407] Single-sample cfHi-C: 500 kb bins with a mappability of less than 0.75 were removed from downstream analysis. The fragment lengths of all high-quality fragments in each 500 kb bin were then determined. The distribution similarity of fragment lengths within each pair of 500 kb bins was calculated by a two-sample KS test (the ks_2samp function implemented in SciPy 1.1.0 in Python 3.6). P values were then converted to the log10 scale. Pearson correlations were then calculated for specific paired bins.

[0408] Sequence composition and mappability bias analysis: Mappability scores were generated by GEM17 for 51-bp read lengths. G+C% was calculated using the gc5 base track from the UCSC Genome Browser. For each pair of 500-kb bins, G+C% and mappability were obtained from bin 1 and bin 2. A gradient boosting machine (GBM) regression tree (GradientBoostingRegressor function implemented in scikit-learn 0.19.1 for Python 3.6) was then applied to regress the G+C% and mappability of each pixel from the matrix of cfHi-C, gDNA, and Hi-C data on the correlation coefficient score. N_estimators were varied at a depth of 5 for different model complexities. The residual values after regression were then used to calculate correlation with whole blood cell (WBC) Hi-C data at the pixel level. The r2 value was calculated to measure the goodness of fit of the model.

[0409] Tissue-of-Origin Analysis in cfHi-C: To infer tissue of origin from cfHi-C data, the cfHi-C data parcel (first PC on the correlation matrix in cfHi-C) was modeled as a linear combination of the parcels in each reference Hi-C data (first PC on the correlation matrix in cfHi-C). Eigenvalues were re-evaluated to ensure that parcel A was a positive number. Genomic regions with a mappability of less than 0.75 were filtered out. Eigenvalues across the cfHi-C and reference Hi-C panels were first transformed by quantile normalization. For each reference Hi-C dataset, only the genomic bin with the highest eigenvalue relative to the remainder of the reference Hi-C dataset (or the lowest if the eigenvalue was negative) was used for deconvolution analysis. Weights were constrained to sum to 1 so that they could be interpreted as tissue contribution to cfDNA. Quadratic programming was used to solve the constrained optimization problem. The tissue contribution fractions from cancer were summed to define the tumor fraction.

[0410] ichorCNA analysis: ichorCNA v0.1.0 with default parameters was used to calculate the tumor fraction in each cfDNA WGS sample after normalization to the group of internal healthy samples. Code and Data Availability: All analysis code was implemented in Python 3.6 and R 3.3.3. Publicly available data used in the study are shown in Table 5. Detailed summary statistics of fragment lengths at the genomic bin level for each cfDNA sample. [Table 5]

[0411] Paired-end whole genome sequencing (WGS) was performed on cfDNA from 568 different healthy individuals. An average of 395 million paired-end reads was obtained for each sample (approximately 12.8x coverage). After quality control and read filtering, an average of 310 million high-quality paired-end reads was obtained for each sample (approximately 10x coverage). Autosomes were divided into non-overlapping 500kb bins, and a normalized fragmentation score was calculated for each individual sample from the fragment lengths in each bin alone. Pearson correlation coefficients were then calculated between each pair of bins with the normalized fragmentation scores across all individuals. Similar patterns were observed between the fragmentation correlation maps of cfDNA and the sections of Hi-C experiments from whole blood cells (WBCs) from two healthy individuals (Figures 15A-C and 15D). Figure 15A-C show correlation maps generated from Hi-C, spatially correlated fragment lengths from multiple cfDNA samples, and spatially correlated fragment lengths from a single cfDNA sample. Figure 15D shows genome browser tracks of section A / B from Hi-C (WBC), multiple-sample cfDNA, and single-sample cfDNA. All comparisons were made from chromosome 14 (chr14).

[0412] To quantify the degree of similarity, Pearson correlations between Hi-C and cfDNA-inferred chromatin organization were calculated at the pixel level (genome-wide average Pearson r = 0.76, p < 2.2e-16). The pixel-level correlation coefficients shown in Hi-C were calculated from replicates of two different healthy individuals. The pixel-level correlation coefficients shown in cfDNA (multiple samples in Figure 15E and single sample in Figure 15F) were calculated by correlation with WBC individual 2.

[0413] We further called the chromatin organization estimated from Hi-C data (A / B) and cfDNA. At the compartment level, we found a high degree of agreement between Hi-C and cfDNA-derived chromatin organizations (Pearson r = 0.89, p < 2.2e-16). Hi-C-derived A / B overlapped significantly with the cfDNA results (hypergeometric test p < 2.2e-16). This approach is referred to as cfHi-C.

[0414] To extend the application of cfHi-C to the single-sample level, we divided each 500 kb bin in each sample into smaller 5 kb subbins and used the Kolmogorov-Smirnov (KS) test to measure the similarity of fragmentation score distributions between each pair of 500 kb bins. The KS test further confirmed high correlation between Hi-C and cfHi-C at both the pixel and compartment levels (Figures 16A and 16B). To eliminate potential internal library preparation and sequencing biases resulting from the patterned flow cell technology in NovaSeq, we replicated the algorithm using a publicly available external cfDNA dataset generated by the HiSeq 2000 platform (BH01). Using this dataset, we observed similar patterns in healthy cfDNA samples (Figure 15D).

[0415] To eliminate potential technical bias caused by sequence composition, we applied the locally weighted scatterplot smoothing (LOWESS) method to normalize the fragment lengths of each bin by the average G+C% value. After regressing G+C%, we observed high similarity between Hi-C and multi-sample cfHi-C in WBC (Pearson correlation r = 0.57, p < 2.2e-16, Figure 17A and Figure 17B).

[0416] As a negative control, we repeated the same steps using genomic DNA (gDNA) from primary white blood cells from 120 individuals. Before regressing G+C%, we again observed relatively high similarity between Hi-C and gDNA (Pearson correlation r = 0.40, p < 2.2e-16, Figure 17C and Figure 17D). However, after normalizing by G+C% in gDNA, low residual similarity was observed between Hi-C and gDNA (Pearson correlation r = 0.15, p < 2.2e-16, Figure 17D), and the Hi-C-like block structure was no longer observed. Figure 17E shows a boxplot of pixel-level correlations (Pearson and Spearman) with Hi-C (WBC, repeat 2) across all chromosomes represented in Figures 17A-D.

[0417] To clarify the effect of G+C% and mappability in two-dimensional space, we applied GBM regression trees to cfHi-C. For each pixel on the cfHi-C matrix, we obtained two G+C% and mappability values in the interaction pair bins. We then regressed the G+C% and mappability from the signal at each pixel of the cfHi-C matrix. After regressing the G+C% bias and mappability, we observed significant residual similarity between Hi-C in WBC and both multi-sample (Pearson correlation r = 0.28, p < 2.2e-16, n_estimator = 500, Figure 18A) and single-sample cfHi-C (Pearson correlation r = 0.36, p < 2.2e-16, n_estimator = 500, Figure 18B).

[0418] In the negative control using gDNA, no residual similarity was observed between Hi-C in WBC and both multi-sample (Pearson correlation r = 0.009, p = 0.0002, Figure 18C) and single-sample gDNA (Pearson correlation r = -0.03, p < 2.2e-16, Figure 18D) at the same range of model complexity. Furthermore, for each pair of bins in cfDNA, one of the bins was replaced with a random bin from another chromosome with the same G+C % and mappability, and the cofragmentation score was recalculated. By using the same GBM regression tree approach on the simulated cfHi-C matrix, significantly lower residual similarity to Hi-C was observed at the same range of model complexity (Pearson correlation r = 0.13, p < 2.2e-16, Figure 18E).

[0419] To demonstrate that the model retained biological signal after regressing G+C% and mappability, the same regression tree approach was applied to WBC Hi-C from another individual (replicate 1). High similarity was still observed across replicates (Pearson correlation r=0.53, p<2.2e-16, Figure 18F).

[0420] To explore the effect of model complexity on the analysis, we repeated the regression tree with different model complexities (n_estimator). Using multi-sample cfHi-C, single-sample cfHi-C, and Hi-C from different individuals, we found that even at high model complexity, it was difficult to remove correlations with Hi-C. This phenomenon did not occur with negative control samples, such as multi-sample gDNA, single-sample gDNA, and cfHi-C with permuted bins.

[0421] To rule out the possibility that the co-fragmentation patterns observed in the multi-sample cfHi-C were due to batch defects during sequencing and library preparation, one bin was randomly shuffled between individuals for each pair of bins in cfHi-C. As expected, no correlation with Hi-C was observed (Pearson correlation r = -0.0002, p = 0.74, Figures 19A and 19D). A multi-sample cfHi-C matrix was generated from samples within the same batch (18 samples). High correlation was observed between Hi-C at the pixel level (Pearson correlation r = 0.60, p < 2.2e-16, Figures 19B and 19D) and samples downsampled to the same size (Pearson correlation r = 0.63, p < 2.2e-16, Figures 19C and 19D).

[0422] To test the robustness of this approach, data from different sample sizes were randomly subsampled for multi-sample cfHi-C. A sample size of 10 yielded correlation coefficients of approximately 0.55 at the pixel level and 0.7 at the compartment level with WBC Hi-C. Saturation was achieved with sample sizes of over 80 (Figures 20A-20D).

[0423] To understand the effect of bin size, we repeated the same procedure with different bin sizes. High agreement with Hi-C experiments at different resolutions was consistently observed (Figures 21A-21H). To clarify the effect of sequencing depth in single-sample cfHi-C, the fragment numbers were downsampled to different sizes. Even at approximately 0.7x coverage, we still obtained correlation coefficients with WBC Hi-C of approximately 0.45 at the pixel level and 0.7 at the compartment level (Figures 22A and 22B).

[0424] To determine whether the observed cfHi-C signal changes across different pathological conditions, we generated additional WGS sequences at similar sequencing depths on cfDNA obtained from 45 colon cancer, 48 lung cancer, and 19 melanoma cancer patients. After standardizing eigenvalues at the compartment level across all cfHi-C samples, principal component analysis (PCA) was applied to all healthy samples and selected cancer samples containing high tumor fractions (tumor fraction >= 0.2, estimated by ichorCNA). Even at 500 kb resolution, separation between healthy and different types of cancer samples was observed (Figure 23A). Further application of a semi-supervised dimensionality reduction method, canonical correlation analysis (CCA), revealed clear separation between healthy and cancer samples (Figures 23B-23F).

[0425] To determine whether in vivo chromatin organization measured through cfDNA could be used to infer the cell types contributing to cfDNA in healthy individuals and cancer patients, we correlated the amplitude of eigenvalues observed in the Hi-C data with the amplitude of open / closed states at chromosomes. At 500 kb resolution from GM12878, we observed a significantly high correlation between DNase-seq signal intensity and eigenvalues at Hi-C compartments (Pearson correlation r = 0.8, p < 2.2e-16, Figure 24). This observation suggested that eigenvalues at the compartment level could be further used to quantify chromosome openness.

[0426] To generate a reference Hi-C panel for tissue-of-origin analysis, Hi-C data from 18 different cell types were uniformly processed across different pathological and healthy conditions. To determine whether correlation patterns were cell-specific, in situ Hi-C data were generated from neutrophil cells with 1.96 billion paired reads and 1.06 billion high-quality contacts (mapping quality score >30). Using quantile-normalized eigenvalues in cell-type-specific compartments identified from the reference Hi-C panel, approximately 80% of cfDNA was detected from different types of leukocytes, while almost no cfDNA was detected from cancer cells within cfHi-C (Figures 25A-25C). In contrast to healthy samples, increased fractions of cancer components from related cell types were observed in colon cancer, lung cancer, and melanoma samples using cfHi-C (Figures 25A and 25B).

[0427] To eliminate possible artifacts during library preparation and sequencing, we repeated the procedure using publicly available cfDNA WGS data from healthy individuals, colon cancer, lung squamous cell carcinoma, small cell lung adenocarcinoma, and breast cancer samples. Similar results were observed (Figures 25A and 25B).

[0428] To quantify the accuracy of the approach, we compared the tumor fraction estimated by cfHi-C with that estimated by ichorCNA, an orthogonal method for estimating tumor fraction by coverage using copy number variation (CNV) in cfDNA. Similar low tumor fractions were observed in healthy individuals (median tumor fraction = 0.00, mean = 0.02, Figure 25C), and significantly higher concordance with ichorCNA was observed in different cancer patients (Figure 26). To avoid confounding CNVs from late-stage cancers, we excluded genomic regions with any significant CNV signals for tissue-of-origin analysis. The results were nearly identical to those before exclusion of late-stage cancer samples.

[0429] If the long-range, spatially correlated fragmentation patterns observed in cfDNA are primarily influenced by the epigenetic landscape, similar 2D Hi-C-like patterns may be observed in different epigenetic signals. To test this hypothesis at the single-sample level, we used a modified KS test to determine the similarity between paired bins in different epigenetic signals from GM12878. High concordance was observed between Hi-C experiments from the same cell type using DNase-seq, methylation levels from whole-genome bisulfite sequencing (WGBS), H3K4me1 ChIP-seq, and H3K4me2 ChIP-seq. This observation suggests that the "virtual compartments" inferred from these epigenetic marks are a comprehensive reference panel for performing nuanced tissue-of-origin analyses.

[0430] In conclusion, these analyses demonstrate the feasibility of using cfDNA as a biomarker to monitor temporal changes in chromatin organization and cell type composition in vivo for different clinical conditions.

[0431] D. Example 4: Detection of colon, breast, pancreatic, or liver cancer This example describes the use of predictive analytics to analyze cfDNA data obtained from a subject using an artificial intelligence-based approach (to generate an output of a diagnosis of the subject having cancer, e.g., colon cancer, breast cancer, or liver cancer or pancreatic cancer).

[0432] Retrospective human plasma samples were obtained from 937 patients diagnosed with colorectal cancer (CRC), 116 patients diagnosed with breast cancer, 26 patients diagnosed with liver cancer, and 76 patients diagnosed with pancreatic cancer. In addition, a set of 605 control samples was obtained from patients without a current cancer diagnosis (but potentially with other comorbidities or undiagnosed cancer), 127 of which had confirmed negative colonoscopy. In total, samples were collected from 11 institutions and commercial biobanks from Southern and Northern Europe and the United States. All samples were de-identified.

[0433] Control samples for the CRC model included all samples except for the liver control samples (n = 524). Control samples for the breast cancer model (n = 123) included samples from the same institution contributing the breast cancer samples. The liver cancer samples were derived from a case-control study with 25 matched control samples. The control samples were HBV-positive but negative for cancer. The pancreatic cancer samples and matched controls were also obtained from a single institution. Of the 66 controls, 45 control samples had several noncancerous pathologies, including pancreatitis, CBD stones, benign strictures, and pseudocysts.

[0434] Each patient's age, sex, and cancer stage (when available) were obtained for each sample. Plasma samples collected from each patient were stored at -80°C and thawed before use.

[0435] Cell-free DNA was extracted from 250 μL of plasma (spiked with a unique synthetic double-stranded DNA (dsDNA) fragment for sample tracking) using the MagMAX Cell-Free DNA Isolation Kit (Applied Biosystems) according to the manufacturer's instructions. Paired-end sequencing libraries were prepared using the NEBNext Ultra II DNA Library Prep Kit (New England Biolabs), which includes polymerase chain reaction (PCR) amplification and unique molecular identifiers (UMIs). At least 400 million reads (median = 636 million reads) were sequenced using an Illumina NovaSeq6000 sequencing system with 2x51 base pairs across multiple S2 or S4 flow cells, except for liver cancer samples, where at least 4 million reads (median = 2.8 million reads) were sequenced.

[0436] The resulting sequencing reads were demultiplexed, adapter-trimmed, and aligned to the human reference genome (GRCh38 with decoy, alt contigs, and HLA contigs) using Burrows Wheeler aligner (BWA-MEM 0.7.15). PCR overlapping fragments, if present, were removed using fragment endpoints or unique molecular identifiers (UMIs).

[0437] For all samples except for the liver cancer experiment, sequencing data were checked for quality and excluded from further analysis if they met any of the following conditions: AT dropout greater than approximately 10 (calculated via Piccard 2.10.5), GC dropout greater than approximately 2 (calculated via Piccard 2.10.5), or a sequencing depth less than approximately 10x. Additionally, samples whose relative counts on sex chromosomes did not match the annotated sex were removed from further processing and discarded. Furthermore, any samples suspected of contamination (e.g., due to an expected allele fraction less than approximately 0.99, unexpected genotype calls, or batches with contaminated negative controls) were manually inspected before inclusion in the dataset.

[0438] A cfDNA "profile" was created for each sample by counting the number of fragments that aligned to each putative protein-coding region of the genome. This type of data representation can capture at least two types of signals: (1) somatic CNVs (where gene regions provide a sampling of the genome, allowing for capture of any consistent large-scale amplifications or deletions), and (2) epigenetic alterations of the immune system represented in the cfDNA, with variable nucleosome protection causing the observed changes in coverage.

[0439] A set of functional regions of the human genome, including putative protein-coding gene regions (with genomic coordinate ranges including both intracellular and exonic regions), was annotated to the sequencing data. Annotations for protein-coding gene regions ("gene" regions) were obtained from the Comprehensive Human Expressed Sequences (CHESS) project (v1.0). A feature set was generated from the annotated human genome regions, comprising a vector of cfDNA fragment counts corresponding to a set of genomic regions. The feature set was obtained by counting several cfDNA fragments with a mapping quality of at least 60 that overlapped each annotated gene region by at least one base, thereby generating a set of "gene features" (D = 24,152, covering 1352 Mb) for each sample.

[0440] The count feature vectors were preprocessed via the following transformations. First, counts of cfDNA fragments corresponding to sex chromosomes were removed (only autosomes were retained). Second, counts of cfDNA fragments corresponding to low-quality genomic bins were removed. Third, features were normalized for their length. Low-quality genomic bins were identified by having either an average mappability across bins of less than approximately 0.75, a GC percentage of less than approximately 30% or more than approximately 70%, or a reference genome N content of more than approximately 10%. Fourth, depth normalization was performed for the number of cfDNA fragments. For each sample depth normalization, a trimmed mean was generated by removing the bottom and top 10 percent of the bin before calculating the average counts across bins in the sample, and the trimmed mean was used as a scaling factor. GC correction was applied to the cDNA fragment counts, and a Loess regression correction was used to address GC bias. After these filtering transformations, the resulting gene feature vector had a dimension of 17,582 features covering 1172 Mb.

[0441] Cross-validation procedures can be performed as part of machine learning techniques to obtain an approximation of a model's performance on newly prospectively collected, unseen data. Such approximations can be obtained by sequentially training models on subsets of the data and testing them on a retained dataset. A k-fold cross-validation procedure can be applied, which requires randomly stratifying all data into k groups (or folds) and testing each group on models fitted to the other folds. This approach can be a general and scalable way to estimate generalization performance. However, if class labels are confounded with known covariates, such a "k-fold" cross-validation scheme can result in inflated performance that may not generalize to new datasets. A machine may simply learn to identify batches of labels and their associated distributions. This can lead to misleading results and poor generalization because the classifier learns incorrect associations between class labels and confounders in the training set and then applies them inaccurately to the test set. Cross-validation performance may overestimate generalization performance because the test set may have the same confounders, but the confounder-free prediction set may not perform well, leading to large generalization errors.

[0442] Such problems may be mitigated by performing "k-batch" validation, in which the test set is stratified to include only unseen confounders. Such "k-batch" validation may provide a more robust assessment of generalization performance for data processed at different time points. This effect may be mitigated by performing stratified validation, in which the test set includes only unseen confounders. Because short-term effects may be observed that occur simultaneously with samples processed on the same batch (e.g., a particular GC bias profile), cross-validation may include stratification by batch instead of random stratification. That is, none of the samples in the test set may be from a batch seen in training. Such an approach may be referred to as "k-batch," and validation in this manner may provide a more robust assessment of generalization performance for data on new batches.

[0443] ...

Claims

1. 1. A method of using a classifier capable of distinguishing between populations of individuals, comprising: a) assaying a plurality of classes of molecules in a biological sample using a plurality of assays, said assays providing a plurality of sets of measurements representative of said plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of the plurality of classes of molecules to be input into a machine learning model; c) creating a feature vector of features from the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of a computer system the machine learning model including the classifier, the machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified as having a specified property, and a second subset of the training biological samples identified as not having the specified property; e) identifying the population of individuals having the specified property by inputting the feature vector into the machine learning model to obtain an output classification of whether the biological sample has the specified property or not.

2. 10. The method of claim 1, wherein the multiple classes of molecules are selected from the group consisting of nucleic acids, polyamino acids, carbohydrates, or metabolites.

3. 2. The method of claim 1, wherein the multiple classes of molecules are selected from the group consisting of deoxyribonucleic acid (DNA), genomic DNA, plasmid DNA, complementary DNA (cDNA), cell-free (e.g., non-encapsulated) DNA (cfDNA), circulating tumor DNA (ctDNA), nucleosomal DNA, chromatosomal DNA, mitochondrial DNA (miDNA), artificial nucleic acid analogs, recombinant nucleic acids, plasmids, viral vectors, chromatin, and peripheral blood mononuclear cell-derived (PBMC-derived) genomic DNA.

4. 2. The method of claim 1, wherein the plurality of classes of molecules are selected from the group consisting of nucleic acids including ribonucleic acid (RNA), messenger RNA (mRNA), transfer RNA (tRNA), microRNA (mitoRNA), ribosomal RNA (rRNA), circular RNA (cRNA), alternatively spliced mRNA, small nuclear RNA (snRNA), antisense RNA, short hairpin RNA (shRNA), or small interfering RNA (siRNA).

5. 10. The method of claim 1, wherein the multiple classes of molecules are selected from the group consisting of polyamino acids, peptides, proteins, autoantibodies, or fragments thereof.

6. 10. The method of claim 1, wherein the class of molecules is selected from the group consisting of sugars, lipids, amino acids, fatty acids, phenolic compounds, or alkaloids.

7. 10. The method of claim 1, wherein the multiple classes of molecules are selected from the group consisting of at least two of cfDNA molecules, cfRNA molecules, circulating proteins, antibodies, and metabolites.

8. 2) cfDNA and cfRNA and polyamino acids; 3) cfDNA and cfRNA and small chemical molecules; 4) cfDNA, polyamino acids and small chemical molecules; 5) cfRNA, polyamino acids and small chemical molecules; 6) cfDNA and cfRNA; 7) cfDNA and polyamino acids; 8) cfDNA and small chemical molecules; 9) cfRNA and polyamino acids; 10) cfRNA and small chemical molecules; or 11) polyamino acids and small chemical molecules.

9. 10. The method of claim 1, wherein the multiple classes of molecules are cfDNA, proteins, and autoantibodies.

10. 2. The method of claim 1, wherein the plurality of assays can include at least two of whole genome sequencing (WGS), whole genome bisulfite sequencing (WGSB), EM-seq sequencing, small RNA sequencing, quantitative immunoassay, enzyme-linked immunosorbent assay (ELISA), proximity extension assay (PEA), protein microarray, mass spectrometry, low-coverage whole genome sequencing (lcWGS), selective tagging 5mC sequencing (WO 2019 / 051484), CNV calling, tumor fraction (TF) estimation, whole genome bisulfite sequencing, LINE-1 CpG methylation, 56-gene CpG methylation, cf-protein immunoquantitative ELISA, SIMOA, and cf-miRNA sequencing, and a mixture ratio of cell types or cell phenotypes derived from any of the above assays.

11. 11. The method of claim 10, wherein the whole genome bisulfite or EM-seq sequencing includes methylation analysis.

12. 10. The method of claim 1, wherein the classifier is trained and constructed according to one or more of: Linear Discriminant Analysis (LDA), Partial Least Squares (PLS), Random Forest, k-Nearest Neighbors (KNN), Support Vector Machines (SVM) with a Radial Basis Function Kernel (SVMRadial), SVM with a Linear Basis Function Kernel (SVMLinear), SVM with a Polynomial Basis Function Kernel (SVMPoly), Decision Trees, Multilayer Perceptrons, Mixtures of Experts, Sparse Factor Analysis, Hierarchical Decomposition, and a combination of linear algebraic routines and statistics.

13. The method of claim 1 , wherein the specified property is the presence of a clinically diagnosed disorder.

14. The method of claim 1 , wherein the specified property is cancer selected from the group consisting of colon cancer, liver cancer, lung cancer, pancreatic cancer, and breast cancer.

15. The method of claim 1 , wherein the specified property is responsiveness to treatment.

16. 1. A system for classifying a biological sample, comprising: a) a receiver that receives a plurality of training samples, each of the plurality of training samples having a plurality of classes of molecules, each of the plurality of training samples including one or more known labels; b) a feature selection module that identifies, for each of the plurality of training samples, a set of features corresponding to each of a plurality of different assays operable to be input into a machine learning model, the set of features corresponding to properties of molecules in the plurality of training samples; a feature selection module, for each of the plurality of training samples, the system operable to subject the plurality of classes of molecules in the training sample to the plurality of different assays to obtain sets of measurements, each set of measurements being from one assay applied to a class of molecules in the training sample, and wherein multiple sets of measurements are obtained for the plurality of training samples; c) a feature extraction module that analyzes the set of measurements for each of the plurality of training samples to obtain a training vector for the training sample, the training vector including feature values of the corresponding assay feature set, each feature value corresponding to a feature and including one or more measurements, the training vector being formed using at least one feature from at least two of the feature sets corresponding to a first subset of the plurality of different assays; and d) a machine learning module configured to operate on the training vectors using parameters of the machine learning model to obtain output labels for the plurality of training samples; e) a comparator module for comparing the output labels with the known labels of the training samples; f) a training module that iteratively searches for optimal values of the parameters as part of training the machine learning model based on comparing the output labels to the known labels of the training samples; g) an output module that provides the parameters of the machine learning model and the set of features of the machine learning model.

17. 17. The system of claim 16, wherein the machine learning module includes classification circuitry configured as a machine learning classifier selected from a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a support vector machine (SVM) classifier, a random forest (RF) classifier, a linear kernel support vector machine classifier, a first or second order polynomial kernel support vector machine classifier, a ridge regression classifier, an elastic net algorithm classifier, a sequential minimal optimization algorithm classifier, a naive Bayes algorithm classifier, and an NMF predictive algorithm classifier.

18. The system according to claim 16, wherein the system comprises means for carrying out the method according to any one of claims 1 to 15.

19. 1. A system for classifying a subject based on a multi-analyte analysis in a biological sample composition, the system comprising: (a) a computer-readable medium comprising a classifier operable to classify the subject based on the multi-analyte analysis; and (b) one or more processors for executing instructions stored on the computer-readable medium.

20. A non-transitory computer readable medium comprising machine executable code that, when executed by one or more computer processors, implements any of the methods of any one of claims 1 to 15 or elsewhere herein.

21. 19. A system comprising one or more computer processors and a computer memory coupled thereto, said computer memory comprising machine executable code that, when executed by said one or more computer processors, implements any of the methods described in any one of claims 1 to 15 or elsewhere herein.

22. 1. A method for detecting the presence of cancer in an individual, comprising: a) assaying a plurality of classes of molecules in a biological sample obtained from said individual, said assay providing a plurality of sets of measurements representative of said plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of the plurality of classes of molecules to be input into a machine learning model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of a computer system the machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified from individuals with cancer, and a second subset of the training biological samples identified from individuals without cancer; e) detecting the presence of said cancer in said individual by inputting said feature vector into said machine learning model to obtain an output classification of whether said biological sample is associated with said cancer.

23. 23. The method of claim 22, wherein the output classification comprises a detection value indicative of the presence of the cancer in the individual.

24. 23. The method of claim 22, wherein the machine learning model further outputs another classification providing a probability that the biological sample is cancer-free.

25. 23. The method of claim 22, wherein the cancer is colon cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer.

26. 1. A method for determining the prognosis of an individual with cancer, comprising: a) assaying a plurality of classes of molecules in a biological sample, said assay providing a plurality of sets of measurements representative of said plurality of classes of molecules; b) identifying a set of features corresponding to properties of said classes of molecules to be input into a machine learning model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of a computer system the machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified from individuals with a favorable cancer prognosis, and a second subset of the training biological samples identified from individuals without a favorable cancer prognosis; e) determining the prognosis of the individual with cancer by inputting the feature vector into the machine learning model to obtain an output classification of whether the biological sample is associated with the good cancer prognosis.

27. 27. The method of claim 26, wherein the cancer may be selected from colon cancer, liver cancer, lung cancer, pancreatic cancer, or breast cancer.

28. 1. A method for determining an individual's responsiveness to cancer treatment, comprising: a) assaying a plurality of classes of molecules in a biological sample, said assay providing a plurality of sets of measurements representative of said plurality of classes of molecules; b) identifying a set of features corresponding to properties of each of the plurality of classes of molecules to be input into a machine learning model; c) creating a feature vector of features from each of the plurality of sets of measurements, each feature corresponding to a feature of the set of features and including one or more measurements, the feature vector including at least one feature obtained using each set of measurements of the plurality of sets; d) loading into a memory of a computer system the machine learning model trained using training vectors obtained from training biological samples, a first subset of the training biological samples identified from individuals who will respond to the cancer treatment, and a second subset of the training biological samples identified from individuals who will not respond to the cancer treatment; e) determining the responsiveness to the cancer treatment by inputting the feature vector into the machine learning model to obtain an output classification of whether the biological sample is associated with a treatment response.

29. 29. The method of claim 28, wherein the cancer therapy is selected from an alkylating agent, a plant alkaloid, an antitumor antibiotic, an antimetabolite, a topoisomerase inhibitor, a retinoid, a checkpoint inhibitor therapy, or a VEGF inhibitor.

30. 30. The method of claim 28, wherein the output classification comprises a detection value indicative of the presence of cancer in the individual.

Citation Information

Patent Citations

  • Biomarker compositions and methods

    JP2014522993A

  • Methods and machine learning systems for predicting the likelihood or risk of having cancer

    US20180068083A1

  • Population based treatment recommender using cell free DNA

    WO2017062867A1

  • Methods for fragmentome profiling of cell-free nucleic acids

    WO2018009723A1

  • Method and system for microbial pharmacogenomics

    WO2018013865A1