Quantifying cell types present in samples

WO2026169931A1PCT designated stage Publication Date: 2026-08-13GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure US2026014177_13082026_PF_FP_ABST
    Figure US2026014177_13082026_PF_FP_ABST
Patent Text Reader

Abstract

In implementations described herein, an amount of one or more cell types present in a sample can be determined.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: GH0267WOQUANTIFYING CELL TYPES PRESENT IN SAMPLESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to and incorporates by reference in its entirety for all purposes, the patent application numbers 63 / 755,070 filed February 6, 2025; 63 / 786,593 filed April 10, 2025; 63 / 795,009 filed April 25, 2025; and 63 / 874,343 filed September 2, 2025.BACKGROUND

[0002] In tissue-localized disease, dying cells release nucleic acids into the bloodstream. Analysis of the characteristics of the nucleic acids can be used to detect and characterize a number of biological conditions. For example, circulating tumor DNA (ctDNA) is a key biomarker for cancer screening and monitoring. Quantification of immune cell- and tissue-derived nucleic acids in plasma can expand the range of conditions identifiable through liquid biopsies.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain implementations, and together with the written description, serve to explain certain principles of the methods, computer readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings which are included by way of example and not by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context indicates otherwise. It will also be understood that some or all of the figures may be schematic representations for purposes of illustration and do not necessarily depict the actual relative sizes or locations of the elements shown.

[0004] Figure 1 is a diagrammatic representation of an example environment 100 that identifies nucleic acids that correspond to classification regions of a reference sequence, where the classification regions have at least a threshold number of CpGs.

[0005] Figure 2 is a diagrammatic representation of an example architecture to determine tumor metrics based on one or more models that analyze methylation status of cell free nucleic acid molecules, according to one or more implementations.

[0006] Figure 3 is a diagrammatic representation of an example architecture to train one or more machine learning models to determine cancer metrics based on methylation status of cell-free nucleic acid molecules, according to one or more implementations.Attorney Docket No.: GH0267WO

[0007] Figure 4 is a flow diagram of an example process to determine tumor metrics related to levels of methylation of classification regions of a reference sequence, according to one or more example implementations.

[0008] Figure 5 is a flow diagram of an example process to determine an amount of a cell type present a sample, according to one or more example implementations.

[0009] Figure 6 is a flow diagram of an example process to determine amounts of a number of cell types present in a sample, according to one or more example implementations.

[0010] Figure 7 illustrates an example framework to determine cell type abundances in samples using a computational model and a reference matrix, in accordance with one or more example implementations.

[0011] Figure 8 illustrates an additional framework to determine abundances of a number of cell types present in samples, in accordance with one or more example implementations.

[0012] Figure 9 is a block diagram illustrating components of a machine, in the form of a computer system, that may read and execute instructions from one or more machine-readable media to perform any one or more methodologies described herein, in accordance with one or more example implementations.

[0013] Figure 10 is a block diagram illustrating a representative software architecture that may be used in conjunction with one or more hardware architectures described herein, in accordance with one or more example implementations.

[0014] Figure 11 is a diagram showing the samples, assays, and deconvolution methods used for analysis of methylation data derived from a plasma cell-free DNA (cfDNA) assay.

[0015] Figure 12 shows a comparison of immune cell type deconvolution in blood showing results determined using methylation data vs. RNA-seq data.

[0016] Figure 13 includes a number of graphs showing that SSM-based cell proportion estimates in blood agree with CyTOF..

[0017] Figure 14 shows immune cell type deconvolution in blood showing results determined using methylation data vs. cytometry by time of flight (CyTOF) for T cells, B cells, and NK cells.

[0018] Figure 15 shows immune cell type deconvolution in blood showing results determined using methylation data vs. cytometry by time of flight (CyTOF) for monocytes and neutrophils.

[0019] Figure 16 shows immune and tissue type frequency measurements determined from plasma – single-site methylation data vs. data derived from an epigenomics assay for hepatocytes, endothelium, and cardiomyocytes.Attorney Docket No.: GH0267WO

[0020] Figure 17 shows immune and tissue type frequency measurements determined from plasma – single-site methylation data vs. data derived from an epigenomics assay for megakaryocytes, neutrophils, T cells, B cells, and erythroid cells.

[0021] Figure 18 includes a table showing LOD90 established by simulations, for blood and tissue cell types commonly found in plasma.

[0022] Figure 19 includes pie charts showing SSM-based cell contributions to plasma cfDNA are consistent with published data.

[0023] Figure 20 includes graphs showing that a clinically validated cfDNA methylation assay reproduces immune and blood cell proportion estimates from SSM. (Top 5 cell types by correlation are shown.)

[0024] Figure 21 includes graphs showing that a clinically validated cfDNA methylation assay reproduces non-blood tissue cell type contribution estimates from SSM.

[0025] Figure 22 includes a table showing Pearson correlations of cell type proportion estimates from a cfDNA assay with SSM-based estimates.

[0026] Figure 23 includes a graphic showing counts of nucleic acids derived from different tissue types in a first lung-specific genomic region in which the nucleic acid counts across a number of tissue types were relatively low.

[0027] Figure 24 includes a graphic showing counts of nucleic acids derived from different tissue types in a second lung-specific genomic region in which nucleic acid counts across a number of tissue types were more prevalent.

[0028] Figure 25 includes a graphic showing limits of detection for lung tissue nucleic acids detected in cell-free DNA in subjects in which breast cancer was detected across a number of tumor fraction values.SUMMARY

[0029] In one aspect, a method includes obtaining a polynucleotide sample derived from a subject, determining, for individual first regions of interest from among a plurality of regions of interest, a first number of nucleic acids derived from the sample or at least a subsample that are aligned with the individual first regions of interest, where the first number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) in individual first nucleic acids, determining first quantitative measures for the individual first regions of interest based on the first number of nucleic acids, determining, by executing a first computational model and based on the first quantitative measures, an indicator of a biological condition for the subject; andAttorney Docket No.: GH0267WOdetermining, using one or more second computational models and based on the first quantitative measures for the individual first regions of interest and based on the indicator of the biological condition for the subject, a first abundance of a first cell type present in the sample or at least a subsample.

[0030] The method may also include determining, for individual second regions of interest from among the plurality of regions of interest, a second number of nucleic acids derived from the sample or at least one subsample that are aligned with the individual second regions of interest, where the second number of nucleic acids comprise at least the threshold amount of CpGs, determining second quantitative measures for the individual second regions of interest based on the second number of nucleic acids, and determining, using the one or more computational models and based on the second quantitative measures for the individual second regions of interest, a second abundance of a second cell type present in the sample or at least a subsample.

[0031] The method may also include where the first regions of interest include a first subset of the plurality of regions of interest that are differentially methylated for the first cell type, and the second regions of interest include a second subset of the plurality of regions of interest that are differentially methylated for the second cell type.

[0032] The method may also include determining, for individual control regions, an additional number of nucleic acids derived from the sample that are aligned with the individual control regions, and where the first quantitative measures for the individual first regions of interest are determined based on the first number of nucleic acids and the additional number of nucleic acids.

[0033] The method may also include where the one or more computational models implement a ridge regression algorithm.

[0034] The method may also include determining that the first number of nucleic acids satisfy one or more methylation characteristics, the one or more methylation characteristics corresponding to methylation states of the CpGs in the individual first nucleic acids.

[0035] The method may also include where the at least one subsample is obtained by partitioning the sample into a plurality of subsamples including a first subsample and a second subsample, where the first subsample includes nucleic acids with a cytosine modification in a greater proportion than the second subsample.

[0036] The method may also include where partitioning the sample into a plurality of subsamples includes separating a number of nucleic acids included in the sample into subgroups of nucleic acids on the basis of methylation level.Attorney Docket No.: GH0267WO

[0037] The method may also include where a first subgroup of the number of nucleic acids include first additional nucleic acids having a first methylation level and a second subgroup of the number of nucleic acids include second additional nucleic acids having a second methylation level.

[0038] The method may also include where the first methylation level corresponds to a first range of methylated CpGs being present in the first additional nucleic acids and the second methylation level corresponds to a second range of methylated CpGs being present in the second additional nucleic acids, the second range of methylated CpGs includes a greater number of CpGs than the first range of methylated CpGs.

[0039] The method may also include where partitioning the sample into a plurality of subsamples includes contacting the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0040] The method may also include obtaining a first plurality of training samples from a first plurality of subjects, the first plurality of training samples corresponding to the first cell type, determining, for the individual first regions of interest, an additional number of nucleic acids derived from the first plurality of training samples or at least one or more subsamples derived from the first plurality of training samples that are aligned with the individual first regions of interest, where the additional number of nucleic acids comprise at least the threshold amount of CpGs in individual additional nucleic acids, determining additional first quantitative measures for the individual first regions of interest based on the additional number of nucleic acids, and generating a first training dataset that indicates, for the first cell type, the additional quantitative measures that correspond to the individual first regions.

[0041] The method may also include obtaining a second plurality of training samples, generating a second training dataset that indicates an amount of individual cell types of a plurality of cell types present in the second plurality of training samples, and generating a combined training dataset for the one or more computational models based on the first training dataset and the second training dataset.

[0042] The method may also include where the second training dataset is produced by performing a cytometry by time of flight process with respect to the second plurality of training samples.

[0043] The method may also include where the second training dataset is produced by performing one or more ribonucleic acid (RNA) sequencing processes with respect to the second plurality of training samples.

[0044] The method may also include where the second training dataset is produced by determining, for the individual cell types of the plurality of cell types, an amount of further nucleicAttorney Docket No.: GH0267WOacids that are aligned with an additional plurality of regions of interest and that include the threshold amount of CpGs and the additional threshold number of methylated or unmethylated CpGs, determining, for the individual additional plurality of regions of interest, further quantitative measures for the individual cell types based on the further number of nucleic acids that (i) are aligned with individual additional plurality of regions of interest and (ii) correspond to the individual cell types, and generating, based on the further number of nucleic acids aligned with the individual additional plurality of regions of interest, a reference matrix that indicates an amount of the further nucleic acids that corresponds to the individual cell types and the individual additional plurality of regions of interest

[0045] The method may also include where a training process for the one or more computational models uses the reference matrix and the combined training dataset.

[0046] The method may also include includes performing a single site methylation state identification process with respect to the second plurality of training samples to determine a number of methylated or unmethylated CpGs present in individual further nucleic acids derived from the second plurality of training samples.

[0047] The method may also include determining, based on the first abundance of the first cell type present in the sample or at least a subsample, a biological condition present in the subject.

[0048] The method may also include determining, based on the first abundance of the first cell type present in the sample or at least a subsample, one or more treatments to provide to the subject in relation to a biological condition.

[0049] The method may also include where the one or more computational models were trained using training data, at least a portion of the training data being obtained from a plurality of training subjects in which the biological condition is not present.

[0050] The method may also include determining, based on the first abundance of the first cell type present in the sample or at least a subsample, an effectiveness of a treatment provided to the subject.

[0051] The method may also include where the sample is obtained from whole blood, plasma, urine, cerebrospinal fluid, buffy coat, or a tissue biopsy.

[0052] The method may also include where the first cell type includes B-cells, T-cells, neutrophils, NK cells, monocytes, megakaryocytes, erythroid cells, liver cells, endothelial cells, or cardiovascular cells.

[0053] The method may also include where the first computational model comprises a regression computational model.Attorney Docket No.: GH0267WO

[0054] The method may also include, where the first computational model is trained using a training dataset derived from additional training samples having tumor fractions within a specified range of tumor fractions.

[0055] In one aspect, a method includes obtaining a polynucleotide sample derived from a subject, determining a number of nucleic acids derived from the sample that are aligned with individual regions of interest of a plurality of regions of interest, where individual nucleic acids of the number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) and an additional threshold amount of methylated CpGs or unmethylated CpGs, determining, for individual regions of interest, quantitative measures for individual cell types of a plurality of cell types based on a subset of the number of nucleic acids that are aligned with the individual regions of interest, and determining, using one or more computational models and based on the quantitative measures and based on a reference matrix, an amount of the individual cell types present in the sample, wherein reference matrix that indicates amounts of nucleic acids that correspond to the individual cell types and the individual regions of interest.

[0056] The method may also include obtaining a plurality of training samples from a plurality of subjects, determining an additional number of nucleic acids derived from the plurality of samples that are aligned with the individual regions of interest, where individual additional nucleic acids of the additional number of nucleic acids comprise at least the threshold amount of CpGs and the additional threshold amount of methylated CpGs or unmethylated CpGs, determining, for individual regions of interest, additional quantitative measures for the individual cell types based on the additional number of nucleic acids that (i) are aligned with the individual regions of interest and (ii) correspond to the individual cell types, generating, based on the additional number of nucleic acids aligned with the individual regions of interest, the reference matrix such that the reference matrix indicates an amount of the additional nucleic acids that corresponds to the individual cell types and the individual regions of interest.

[0057] The method may also include where a training process for the one or more computational models uses the reference matrix.

[0058] The method may also include where the one or more computational models include a least squares model.

[0059] The method may also include where the one or more computational models include a partial least squares model with a non-negative value constraint.

[0060] The method may also include performing a single site methylation state identification process with respect to the plurality of samples to determine a number of methylated or unmethylated CpGs present in individual nucleic acids derived from the plurality of samples.Attorney Docket No.: GH0267WO

[0061] The method may also include determining, based on the amount of the individual cell types present in the sample, a biological condition present in the subject.

[0062] The method may also include where the one or more computational models were trained using training data obtained from a plurality of subjects in which the biological condition is not present.

[0063] The method may also include determining, based on the amount of the individual cell types present in the sample, one or more treatments to provide to the subject in relation to a biological condition.

[0064] The method may also include determining, based on the amount of the individual cell types present in the sample, an effectiveness of a treatment provided to the subject.

[0065] A computing apparatus may include a processor; and memory storing instructions that, when executed by the processor, configure the apparatus to determine, for individual first regions of interest from among a plurality of regions of interest, a first number of nucleic acids derived from a polynucleotide sample or at least a subsample that are aligned with the individual first regions of interest, where the first number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) in individual first nucleic acids, determine first quantitative measures for the individual first regions of interest based on the first number of nucleic acids, determine, by executing a first computational model and based on the first quantitative measures, an indicator of a biological condition for the subject; and determine, using one or more second computational models and based on the first quantitative measures for the individual first regions of interest and based on the indicator of the biological condition for the subject, a first abundance of a first cell type present in the sample or at least a subsample.

[0066] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, for individual second regions of interest from among the plurality of regions of interest, a second number of nucleic acids derived from the sample or at least one subsample that are aligned with the individual second regions of interest, where the second number of nucleic acids comprise at least the threshold amount of CpGs, determine second quantitative measures for the individual second regions of interest based on the second number of nucleic acids, and determine, using the one or more computational models and based on the second quantitative measures for the individual second regions of interest, a second abundance of a second cell type present in the sample or at least a subsample.

[0067] The computing apparatus may also include where the first regions of interest include a first subset of the plurality of regions of interest that are differentially methylated for the first cellAttorney Docket No.: GH0267WOtype, and the second regions of interest include a second subset of the plurality of regions of interest that are differentially methylated for the second cell type.

[0068] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, for individual control regions, an additional number of nucleic acids derived from the sample that are aligned with the individual control regions, and where the first quantitative measures for the individual first regions of interest are determined based on the first number of nucleic acids and the additional number of nucleic acids.

[0069] The computing apparatus may also include where the one or more computational models implement a ridge regression algorithm.

[0070] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine that the first number of nucleic acids satisfy one or more methylation characteristics, the one or more methylation characteristics corresponding to methylation states of the CpGs in the individual first nucleic acids.

[0071] The computing apparatus may also include where the at least one subsample is obtained by partitioning the sample into a plurality of subsamples include a first subsample and a second subsample, where the first subsample includes nucleic acids with a cytosine modification in a greater proportion than the second subsample.

[0072] The computing apparatus may also include where partitioning the sample into a plurality of subsamples includes separating a number of nucleic acids included in the sample into subgroups of nucleic acids on the basis of methylation level.

[0073] The computing apparatus may also include where a first subgroup of the number of nucleic acids include first additional nucleic acids having a first methylation level and a second subgroup of the number of nucleic acids include second additional nucleic acids having a second methylation level.

[0074] The computing apparatus may also include where the first methylation level corresponds to a first range of methylated CpGs being present in the first additional nucleic acids and the second methylation level corresponds to a second range of methylated CpGs being present in the second additional nucleic acids, the second range of methylated CpGs includes a greater number of CpGs than the first range of methylated CpGs.

[0075] The computing apparatus may also include where partitioning the sample into a plurality of subsamples includes contact the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0076] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, for the individual first regions of interest, anAttorney Docket No.: GH0267WOadditional number of nucleic acids derived from a first plurality of training samples or at least one or more subsamples derived from the first plurality of training samples that are aligned with the individual first regions of interest, where the additional number of nucleic acids comprise at least the threshold amount of CpGs in individual additional nucleic acids, the first plurality of training samples are obtained from a first plurality of subjects, and the first plurality of training samples correspond to the first cell type, determine additional first quantitative measures for the individual first regions of interest based on the additional number of nucleic acids, and generate a first training dataset that indicates, for the first cell type, the additional quantitative measures that correspond to the individual first regions.

[0077] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to generate a second training dataset that indicates an amount of individual cell types of a plurality of cell types present in a second plurality of training samples, and generate a combined training dataset for the one or more computational models based on the first training dataset and the second training dataset.

[0078] The computing apparatus may also include where the second training dataset is produced by performing a cytometry by time of flight process with respect to the second plurality of training samples.

[0079] The computing apparatus may also include where the second training dataset is produced by performing one or more ribonucleic acid (RNA) sequencing processes with respect to the second plurality of training samples.

[0080] The computing apparatus may also include where the second training dataset is produced by determining, for the individual cell types of the plurality of cell types, an amount of further nucleic acids that are aligned with an additional plurality of regions of interest and that include the threshold amount of CpGs and the additional threshold number of methylated or unmethylated CpGs, determining, for the individual additional plurality of regions of interest, further quantitative measures for the individual cell types based on the further number of nucleic acids that (i) are aligned with individual additional plurality of regions of interest and (ii) correspond to the individual cell types, and generating, based on the further number of nucleic acids aligned with the individual additional plurality of regions of interest, a reference matrix that indicates an amount of the further nucleic acids that corresponds to the individual cell types and the individual additional plurality of regions of interest.

[0081] The computing apparatus may also include where a training process for the one or more computational models uses the reference matrix and the combined training dataset.Attorney Docket No.: GH0267WO

[0082] The computing apparatus may also include where a number of methylated or unmethylated CpGs present in individual further nucleic acids derived from the second plurality of training samples are determined by performing a single site methylation state identification process with respect to the second plurality of training samples.

[0083] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, a biological condition present in the subject

[0084] The computing apparatus may also include where the one or more computational models were trained using train data, at least a portion of the training data being obtained from a plurality of training subjects in which the biological condition is not present.

[0085] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, one or more treatments to provide to the subject in relation to a biological condition.

[0086] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, an effectiveness of a treatment provided to the subject.

[0087] The computing apparatus may also include where the sample is obtained from whole blood, plasma, urine, cerebrospinal fluid, buffy coat, or a tissue biopsy.

[0088] The computing apparatus may also include where the first cell type includes B-cells, T-cells, neutrophils, NK cells, monocytes, megakaryocytes, erythroid cells, liver cells, endothelial cells, or cardiovascular cells.

[0089] The computing apparatus may also include where the first computational model comprises a regression computational model.

[0090] The computing apparatus may also include, where the first computational model is trained using a training dataset derived from additional training samples having tumor fractions within a specified range of tumor fractions.

[0091] A computing apparatus may include a processor; and memory storing instructions that, when executed by the processor, configure the apparatus to determine a number of nucleic acids derived from a polynucleotide sample derived from a subject that are aligned with individual regions of interest of a plurality of regions of interest, where individual nucleic acids of the number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) and an additional threshold amount of methylated CpGs or unmethylated CpGs, determine, forAttorney Docket No.: GH0267WOindividual regions of interest, quantitative measures for individual cell types of a plurality of cell types based on a subset of the number of nucleic acids that are aligned with the individual regions of interest, and determine, using one or more computational models and based on the quantitative measures and based on a reference matrix, an amount of the individual cell types present in the sample, wherein reference matrix that indicates amounts of nucleic acids that correspond to the individual cell types and the individual regions of interest.

[0092] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine an additional number of nucleic acids derived from a plurality of samples obtained from a plurality of training subjects that are aligned with the individual regions of interest, where individual additional nucleic acids of the additional number of nucleic acids comprise at least the threshold amount of CpGs and the additional threshold amount of methylated CpGs or unmethylated CpGs, determine, for individual regions of interest, additional quantitative measures for the individual cell types based on the additional number of nucleic acids that (i) are aligned with the individual regions of interest and (ii) correspond to the individual cell types, generate, based on the additional number of nucleic acids aligned with the individual regions of interest, a reference matrix that indicates an amount of the additional nucleic acids that corresponds to the individual cell types and the individual regions of interest.

[0093] The computing apparatus may also include where a training process for the one or more computational models uses the reference matrix.

[0094] The computing apparatus may also include where the one or more computational models include a least squares model.

[0095] The computing apparatus may also include where the one or more computational models include a partial least squares model with a non-negative value constraint.

[0096] The computing apparatus may also include where a number of methylated or unmethylated CpGs present in individual nucleic acids derived from the plurality of samples is determined by performing a single site methylation state identification process with respect to the plurality of samples.

[0097] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the amount of the individual cell types present in the sample, a biological condition present in the subject.

[0098] The computing apparatus may also include where the one or more computational models were trained using train data obtained from a plurality of subjects in which the biological condition is not present.Attorney Docket No.: GH0267WO

[0099] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the amount of the individual cell types present in the sample, one or more treatments to provide to the subject in relation to a biological condition.

[0100] The computing apparatus may also store additional instructions that, when executed by the processor, configure the apparatus to determine, based on the amount of the individual cell types present in the sample, an effectiveness of a treatment provided to the subject.

[0101] In one aspect, a non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to determine, for individual first regions of interest from among a plurality of regions of interest, a first number of nucleic acids or at least a subsample derived from a polynucleotide sample derived from a subject that are aligned with the individual first regions of interest, where the first number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) in individual first nucleic acids, determine first quantitative measures for the individual first regions of interest based on the first number of nucleic acids, determine, by executing a first computational model and based on the first quantitative measures, an indicator of a biological condition for the subject; and determine, using one or more second computational models and based on the first quantitative measures for the individual first regions of interest and based on the indicator of the biological condition for the subject, a first abundance of a first cell type present in the sample or at least a subsample.

[0102] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, for individual second regions of interest from among the plurality of regions of interest, a second number of nucleic acids derived from the sample or at least one subsample that are aligned with the individual second regions of interest, where the second number of nucleic acids comprise at least the threshold amount of CpGs, determine second quantitative measures for the individual second regions of interest based on the second number of nucleic acids, and determine, using the one or more computational models and based on the second quantitative measures for the individual second regions of interest, a second abundance of a second cell type present in the sample or at least a subsample.

[0103] The computer-readable storage medium may also include where the first regions of interest include a first subset of the plurality of regions of interest that are differentially methylated for the first cell type, and the second regions of interest include a second subset of the plurality of regions of interest that are differentially methylated for the second cell type.Attorney Docket No.: GH0267WO

[0104] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, for individual control regions, an additional number of nucleic acids derived from the sample that are aligned with the individual control regions, and where the first quantitative measures for the individual first regions of interest are determined based on the first number of nucleic acids and the additional number of nucleic acids.

[0105] The computer-readable storage medium may also include where the one or more computational models implement a ridge regression algorithm.

[0106] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine that the first number of nucleic acids satisfy one or more methylation characteristics, the one or more methylation characteristics corresponding to methylation states of the CpGs in the individual first nucleic acids.

[0107] The computer-readable storage medium may also include where the at least one subsample is obtained by partitioning the sample into a plurality of subsamples include a first subsample and a second subsample, where the first subsample includes nucleic acids with a cytosine modification in a greater proportion than the second subsample.

[0108] The computer-readable storage medium may also include where partitioning the sample into a plurality of subsamples includes separating a number of nucleic acids included in the sample into subgroups of nucleic acids on the basis of methylation level.

[0109] The computer-readable storage medium may also include where a first subgroup of the number of nucleic acids include first additional nucleic acids having a first methylation level and a second subgroup of the number of nucleic acids include second additional nucleic acids having a second methylation level.

[0110] The computer-readable storage medium may also include where the first methylation level corresponds to a first range of methylated CpGs being present in the first additional nucleic acids and the second methylation level corresponds to a second range of methylated CpGs being present in the second additional nucleic acids, the second range of methylated CpGs includes a greater number of CpGs than the first range of methylated CpGs.

[0111] The computer-readable storage medium may also include where partitioning the sample into a plurality of subsamples includes contact the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0112] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, for the individual first regions of interest, an additional number of nucleic acids derived from a first plurality of trainingAttorney Docket No.: GH0267WOsamples or at least one or more subsamples derived from the first plurality of training samples that are aligned with the individual first regions of interest, where the additional number of nucleic acids comprise at least the threshold amount of CpGs in individual additional nucleic acids, the first plurality of training samples are obtained from a first plurality of subjects, and the first plurality of training samples correspond to the first cell type, determine additional first quantitative measures for the individual first regions of interest based on the additional number of nucleic acids, and generate a first training dataset that indicates, for the first cell type, the additional quantitative measures that correspond to the individual first regions.

[0113] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to generate a second training dataset that indicates an amount of individual cell types of a plurality of cell types present in a second plurality of training samples, and generate a combined training dataset for the one or more computational models based on the first training dataset and the second training dataset.

[0114] The computer-readable storage medium may also include where the second training dataset is produced by performing a cytometry by time of flight process with respect to the second plurality of training samples.

[0115] The computer-readable storage medium may also include where the second training dataset is produced by performing one or more ribonucleic acid (RNA) sequencing processes with respect to the second plurality of training samples.

[0116] The computer-readable storage medium may also include where the second training dataset is produced by determining, for the individual cell types of the plurality of cell types, an amount of further nucleic acids that are aligned with an additional plurality of regions of interest and that include the threshold amount of CpGs and the additional threshold number of methylated or unmethylated CpGs, determining, for the individual additional plurality of regions of interest, further quantitative measures for the individual cell types based on the further number of nucleic acids that (i) are aligned with individual additional plurality of regions of interest and (ii) correspond to the individual cell types, and generating, based on the further number of nucleic acids aligned with the individual additional plurality of regions of interest, a reference matrix that indicates an amount of the further nucleic acids that corresponds to the individual cell types and the individual additional plurality of regions of interest.

[0117] The computer-readable storage medium may also include where a training process for the one or more computational models uses the reference matrix and the combined training dataset.Attorney Docket No.: GH0267WO

[0118] The computer-readable storage medium may also include where a number of methylated or unmethylated CpGs present in individual further nucleic acids derived from the second plurality of training samples are determined by performing a single site methylation state identification process with respect to the second plurality of training samples.

[0119] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, a biological condition present in the subject.

[0120] The computer-readable storage medium may also include where the one or more computational models were trained using train data, at least a portion of the training data being obtained from a plurality of training subjects in which the biological condition is not present.

[0121] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, one or more treatments to provide to the subject in relation to a biological condition.

[0122] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer the apparatus to determine, based on the first abundance of the first cell type present in the sample or at least a subsample, an effectiveness of a treatment provided to the subject.

[0123] The computer-readable storage medium may also include where the sample is obtained from whole blood, plasma, urine, cerebrospinal fluid, buffy coat, or a tissue biopsy.

[0124] The computer-readable storage medium may also include where the first cell type includes B-cells, T-cells, neutrophils, NK cells, monocytes, megakaryocytes, erythroid cells, liver cells, endothelial cells, or cardiovascular cells.

[0125] The computer-readable storage medium may also include where the first computational model comprises a regression computational model.

[0126] The computer-readable storage medium may also include, where the first computational model is trained using a training dataset derived from additional training samples having tumor fractions within a specified range of tumor fractions.

[0127] In one aspect, a non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to obtain a polynucleotide sample derived from a subject, determine a number of nucleic acids derived from the sample that are aligned with individual regions of interest of a plurality of regions of interest, where individual nucleic acids of the number of nucleic acids comprise at leastAttorney Docket No.: GH0267WOa threshold amount of cytosine-guanine dinucleotides (CpGs) and an additional threshold amount of methylated CpGs or unmethylated CpGs, determine, for individual regions of interest, quantitative measures for individual cell types of a plurality of cell types based on a subset of the number of nucleic acids that are aligned with the individual regions of interest, and determine, using one or more computational models and based on the quantitative measures and based on a reference matrix, an amount of the individual cell types present in the sample, wherein reference matrix that indicates amounts of nucleic acids that correspond to the individual cell types and the individual regions of interest.

[0128] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine an additional number of nucleic acids derived from a plurality of samples obtained from a plurality of training subjects that are aligned with the individual regions of interest, where individual additional nucleic acids of the additional number of nucleic acids comprise at least the threshold amount of CpGs and the additional threshold amount of methylated CpGs or unmethylated CpGs, determine, for individual regions of interest, additional quantitative measures for the individual cell types based on the additional number of nucleic acids that (i) are aligned with the individual regions of interest and (ii) correspond to the individual cell types, generate, based on the additional number of nucleic acids aligned with the individual regions of interest, a reference matrix that indicates an amount of the additional nucleic acids that corresponds to the individual cell types and the individual regions of interest.

[0129] The computer-readable storage medium may also include where a training process for the one or more computational models uses the reference matrix.

[0130] The computer-readable storage medium may also include where the one or more computational models include a least squares model.

[0131] The computer-readable storage medium may also include where the one or more computational models include a partial least squares model with a non-negative value constraint.

[0132] The computer-readable storage medium may also include where a number of methylated or unmethylated CpGs present in individual nucleic acids derived from the plurality of samples is determined by performing a single site methylation state identification process with respect to the plurality of samples.

[0133] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, based on the amount of the individual cell types present in the sample, a biological condition present in the subject.Attorney Docket No.: GH0267WO

[0134] The computer-readable storage medium may also include where the one or more computational models were trained using train data obtained from a plurality of subjects in which the biological condition is not present.

[0135] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, based on the amount of the individual cell types present in the sample, one or more treatments to provide to the subject in relation to a biological condition.

[0136] The computer-readable storage medium may also store additional instructions that, when executed by the computer, cause the computer to determine, based on the amount of the individual cell types present in the sample, an effectiveness of a treatment provided to the subject.

[0137] Other technical features may be readily apparent to one skilled in the art from the following figures, descriptions, and claims.DEFINITIONS

[0138] In order for the present disclosure to be more readily understood, certain terms are first defined below. Additional definitions for the following terms and other terms may be set forth through the specification. If a definition of a term set forth below is inconsistent with a definition in an application or patent that is incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.

[0139] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, a reference to “a method” includes one or more methods, and / or steps of the type described herein and / or which will become apparent to those persons of ordinary skill in the art upon reading this disclosure and so forth.

[0140] It is also to be understood that the terminology used herein is for the purpose of describing particular implementations only, and is not intended to be limiting. Further, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer readable media, and systems, the following terminology, and grammatical variants thereof, will be used in accordance with the definitions set forth below.Attorney Docket No.: GH0267WO

[0141] About: As used herein, “about” or “approximately” as applied to one or more values or elements of interest, refers to a value or element that is similar to a stated reference value or element. In certain implementations, the term “about” or “approximately” refers to a range of values or elements that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the stated reference value or element unless otherwise stated or otherwise evident from the context (except where such number would exceed 100% of a possible value or element).

[0142] Administer: As used herein, “administer” or “administering” a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means to give, apply or bring the composition into contact with the subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal and intradermal.

[0143] Adapter. As used herein, “adapter” refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that can be at least partially double-stranded and used to link to either or both ends of a given sample nucleic acid molecule. Adapters can include nucleic acid primer binding sites to permit amplification of a nucleic acid molecule flanked by adapters at both ends, and / or a sequencing primer binding site, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. Adapters can also include binding sites for capture probes, such as an oligonucleotide attached to a flow cell support or the like. Adapters can also include a nucleic acid tag as described herein. Nucleic acid tags can be positioned relative to amplification primer and sequencing primer binding sites, such that a nucleic acid tag is included in amplicons and sequence reads of a given nucleic acid molecule. The same or different adapters can be linked to the respective ends of a nucleic acid molecule. In some implementations, the same adapter is linked to the respective ends of the nucleic acid molecule except that the nucleic acid tag differs. In some implementations, the adapter is a Y-shaped adapter in which one end is blunt ended or tailed as described herein, for joining to a nucleic acid molecule, which is also blunt ended or tailed with one or more complementary nucleotides. In still other example implementations, an adapter is a bell-shaped adapter that includes a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed and C-tailed adapters.

[0144] Alignment. As used herein, “alignment” or “align” refers to determining whether at least two sequence representations have at least a threshold amount of homology. In one or more examples, the threshold amount of homology can be at least about 90%, at least about 91%, atAttorney Docket No.: GH0267WOleast about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%. In situations where two sequence representations have at least the threshold amount of homology, the two sequence representations can be referred to as being “aligned.”

[0145] Amplify. As used herein, “amplify” or “amplification” in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of the polynucleotide, starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification products or amplicons are generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.

[0146] Barcode: As used herein, “barcode” or “molecular barcode” in the context of nucleic acids refers to a nucleic acid molecule comprising a sequence that can serve as a molecular identifier. For example, individual "barcode" sequences can be added to each DNA fragment during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before the final data analysis.

[0147] Cancer Type: As used herein, “cancer type” refers to a type or subtype of cancer defined, e.g., by histopathology. Cancer type can be defined by any conventional criterion, such as on the basis of occurrence in a given tissue (e.g., blood cancers, central nervous system (CNS), brain cancers, lung cancers (small cell and non-small cell), skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, breast cancers, prostate cancers, ovarian cancers, lung cancers, intestinal cancers, soft tissue cancers, neuroendocrine cancers, gastroesophageal cancers, head and neck cancers, gynecological cancers, colorectal cancers, urothelial cancers, solid state cancers, heterogeneous cancers, homogenous cancers), unknown primary origin and the like, and / or of the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma) and / or cancers exhibiting cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptor and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether of primary or secondary origin.

[0148] Carrier Signal: As used herein, “carrier signal” refers to any intangible medium that is capable of storing, encoding, or carrying transitory or non-transitory instructions 902 for execution by the machine 900, and includes digital or analog communications signals or other intangible medium to facilitate communication of such instructions 902. Instructions 902 may be transmitted or received over the network 934 using a transitory or non-transitory transmissionAttorney Docket No.: GH0267WOmedium via a network interface device and using any one of a number of well-known transfer protocols.

[0149] Cell-Free Nucleic Acid. As used herein, “cell-free nucleic acid” refers to nucleic acids not contained within or otherwise bound to a cell or, in some implementations, nucleic acids remaining in a sample following the removal of intact cells. Cell-free nucleic acids can include, for example, all non-encapsulated nucleic acids sourced from a bodily fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis, apoptosis, or the like. Some cell-free nucleic acids are released into bodily fluid from cancer cells, e.g., circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be non-encapsulated tumor-derived fragmented DNA. A cell-free nucleic acid can have one or more epigenetic modifications, for example, a cell-free nucleic acid can be acetylated, 5-methylated, ubiquitylated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0150] Cell Type: As used herein, a “cell type” is a set of cells having a shared characteristic. For example, immune cell types can include immune cells of different origins, differentiation types, different activation types, or any combination of different origins, different differentiation types, and different activation types. Indeed, differentiation status and activation status can overlap and often change together in a given immune cell. For example, activation of an immune cell may induce differentiation of the cell. Immune cells of different activation types can include activated cells (such as cells activated by inflammatory cytokines or antigens), suppressive cells (such as T regulatory cells (Tregs), M2 macrophages, and others, or their subsets), or suppressed cells, such as cells suppressed by Tregs. Exemplary immune cell types include, but are not limited to, macrophages (including M1 macrophages and M2 macrophages); activated B cells (including regulatory B cells, memory B cells, and plasma cells); T cell subsets, such as CD4 central memory T cells, CD8 central memory T cells, naive-like T cells, naive T cells, and activated T cells (including cytotoxic T cells, regulatory T cells (Tregs), CD4 effector memory T cells, and CD8 effector memory T cells); immature myeloid cells (including myeloid-derived suppressor cells (MDSCs), low-density neutrophils, immature neutrophils, and immature granulocytes); and natural killer (NK) cells. Additional exemplary immune cell types includeAttorney Docket No.: GH0267WOneutrophils, lymphocytes, plasma cells, monocytes, macrophages, dendritic cells, mast cells, eosinophils, T cells, CD4+ T cells, B cells, megakaryocytes, CD8+ central memory cells, CD4+ central memory cells, precursor B cells, plasma cells, memory-switched B cells, plasma cells, basophils, naive B cells, memory B cells, CD8+ T cells, naive CD4+ T cells, resting CD4+ memory T cells, activated CD4+ memory T cells, follicular helper T cells, gamma delta T cells, resting NK cells; activated NK cells, MO macrophages, resting dendritic cells, activated dendritic cells, resting mast cells, and activated mast cells. In some embodiments, cell types may be distinguished based on characteristics such as one or more cell surface markers, a genetic signature (such as expression (or expression level) of a particular gene or set of genes). Cell types can also correspond to cells extracted from the same organ, tissue, or body site. For example, lung cells may include lung alveolar epithelial cells, lung bronchial epithelial cells, or a combination of both alveolar and bronchial epithelial cells”.

[0151] Cellular Nucleic Acids: As used herein, “cellular nucleic acids” means nucleic acids that are disposed within one or more cells at least at the point a sample is taken or collected from a subject, even if those nucleic acids are subsequently removed as part of a given analytical process.

[0152] Classification Region: As used herein, “classification region” refers to a genomic region that may show sequence-independent changes in neoplastic cells (e.g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects having cancer relative to cfDNA from subjects in which cancer is not present. Examples of sequenceindependent changes include, but are not limited to, changes in methylation rate (increases or decreases), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. In one or more examples, sequence-independent changes in a classification region can indicate the presence of a single form of cancer in a subject. In one or more additional examples, sequence-independent changes in a classification region can correspond to the presence of multiple forms in a subject. The classification region can be enriched by one or more probes. In addition, the classification region can be defined by a pair of primer binding sites. Further, the classification region can be defined by a predetermined beginning genomic locus and a predetermined ending genomic locus. The classification region can include from about 25 nucleotides to about 250 nucleotides, from about 50 nucleotides to about 200 nucleotides, or from about 75 nucleotides to about 150 nucleotides. For instance, classification region can be a differentially methylated region. “Differentially methylated region” or “DMR” refers to a region of DNA having a detectably different degree of methylation in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from atAttorney Docket No.: GH0267WOleast one other cell or tissue type; or having a detectably different degree of methylation in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type irrespective of the status of the biological condition of a subject; or having a detectably different degree of methylation in at least one cell or tissue type obtained from a subject having a disease or disorder relative to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region / hypermethylated target region) in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region / hypomethylated target region) in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, the classification regions comprise hypermethylated target regions and / or hypomethylated target regions.

[0153] Communications Network. As used herein, “communications network” refers to one or more portions of a network 114, 1034 that may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a W-Fi® network, another type of network, or a combination of two or more such networks. For example, a network 114, 1034 or a portion of a network may include a wireless or cellular network and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX),Attorney Docket No.: GH0267WOLong Term Evolution (LTE) standard, others defined by various standard setting organizations, other long range protocols, or other data transfer technology.

[0154] Confidence Interval: As used herein, “confidence interval” means a range of values so defined that there is a specified probability that the value of a given parameter lies within that range of values.

[0155] Control Sample: As used herein, “control sample” or “reference sample” refers to a sample obtained from individuals without known copy number variation.

[0156] Coverage: As used herein, “coverage” or “coverage metrics” refer to the number of nucleic acid molecules or sequencing reads that correspond to a particular genomic region of a reference sequence.

[0157] Deoxyribonucleic Acid or Ribonucleic Acid: As used herein, “deoxyribonucleic acid” or “DNA” refers to a natural or modified nucleotide which has a hydrogen group at the 2'-position of the sugar moiety. DNA can include a chain of nucleotides comprising four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, “ribonucleic acid” or “RNA” refers to a natural or modified nucleotide which has a hydroxyl group at the 2'-position of the sugar moiety. RNA can include a chain of nucleotides comprising four types of nucleotides: A, uracil (U), G, and C. As used herein, the term “nucleotide” refers to a natural nucleotide or a modified nucleotide. Certain pairs of nucleotides specifically bind to one another in a complementary fashion (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand made up of nucleotides that are complementary to those in the first strand, the two strands bind to form a double strand. As used herein, “nucleic acid sequencing data”, “nucleic acid sequencing information”, “sequence information”, “sequence representation”, “nucleic acid sequence”, “nucleotide sequence”, “genomic sequence”, “genetic sequence”, “fragment sequence”, “sequencing read”, or “nucleic acid sequencing read” denotes any information or data that is indicative of the order and identity of the nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment) of a nucleic acid such as DNA or RNA. It should be understood that the present teachings contemplate sequence information obtained using all available varieties of techniques, platforms or technologies, including, but not limited to: capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.Attorney Docket No.: GH0267WO

[0158] Differentially Methylated Region: As used herein, differentially methylated region” refers to a region of DNA having a delectably different degree of methylation in at least one cell or tissue type relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type; or having a detectably different degree of methylation in at least one cell or tissue type obtained from a subject having a disease or disorder relative to the degree of methylation in the same region of DNA in the same cell or tissue type obtained from a healthy subject. In some embodiments, a differentially methylated region has a detectably higher degree of methylation (e.g., a hypermethylated region) in at least one cell or tissue type, such as at least one immune cell type, relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject. In some embodiments, a differentially methylated region has a detectably lower degree of methylation (e.g., a hypomethylated region) in at least one cell or tissue type, such as at least one immune cell type, relative to the degree of methylation in the same region of DNA from at least one other cell or tissue type, such as other immune cell types and / or cell types that contribute to cfDNA in healthy individuals, or from the same cell or tissue type from a healthy subject.

[0159] Driver Mutation. As used herein, “driver mutation” means a mutation that drives cancer progression.

[0160] Epigenetic Target Regions: As used herein, “epigenetic target regions” refers to target regions that may show sequence-independent differences in different cell or tissue types (e.g., different types of immune cells) or in neoplastic cells (e.g., tumor cells and cancer cells) relative to normal cells; or that may show sequence- independent differences (i.e., in which there is no change to the nucleotide sequence, e.g., differences in methylation, nucleosome distribution, or other epigenetic features) in DNA, such as cfDNA, from different cell types or from subjects having cancer relative to DNA, such as cfDNA, from healthy subjects, or in cfDNA originating from different cell or tissue types that ordinarily do not substantially contribute to cfDNA (e.g., immune, lung, colon, etc.) relative to background cfDNA (e.g., cfDNA that originated from hematopoietic cells) Examples of sequence-independent changes include, but are not limited to, changes in methylation (increases or decreases), nucleosome distribution, cfDNA fragmentation patterns, CCCTC-binding factor (“CTCF”) binding, transcription start sites (e.g., with respect to any one of more of binding of RNA polymerase components, binding of regulatory proteins, fragmentation characteristics, and nucleosomal distribution), and regulatory protein binding regions. Epigenetic target region sets thus include, but are not limited to, hypermethylation target region sets, hypomethylation target region sets, and fragmentation variable target region sets, such as CTCFAttorney Docket No.: GH0267WObinding sites and transcription start sites. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may also be included in an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar to detection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individual positions. An epigenetic target region set is a set of epigenetic target regions.

[0161] Hypermethylation: As used herein, “hypermethylation” refers to an increased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated cytosine, at least 2 methylated cytosines, at least 3 methylated cytosines, at least 5 methylated cytosines, or at least 10 methylated cytosines.

[0162] Hypomethylation: As used herein, “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated cytosine, at most 1 methylated cytosine, at most 2 methylated cytosines, at most 3 methylated cytosines, at most 4 methylated cytosines, or at most 5 methylated cytosines.

[0163] Immunotherapy: As used herein, “immunotherapy” refers to treatment with one or more agents that act to stimulate the immune system so as to kill or at least to inhibit growth of cancer cells, and preferably to reduce further growth of the cancer, reduce the size of the cancer and / or eliminate the cancer. Some such agents bind to a target present on cancer cells; some bind to a target present on immune cells and not on cancer cells; some bind to a target present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of pathways of the immune system that maintain self-tolerance and modulate the duration and amplitude of physiological immune responses in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Example agents include antibodies against any of PD-1, PD-2, PD-L1, PD-L2, CTLA-40, 0X40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other example agents include proinflammatory cytokines, such as IL-1β, IL-6, and TNF-Attorney Docket No.: GH0267WOa. Other example agents are T-cells activated against a tumor, such as T-cells activated by expressing a chimeric antigen targeting a tumor antigen recognized by the T-cell.

[0164] Indel: As used herein, “indel” refers to a mutation that involves the insertion or deletion of nucleotides in the genome of a subject.

[0165] Limit of Detection (LoD) As used herein, “limit of detection” means the smallest amount of a substance (e.g., a nucleic acid) in a sample that can be measured by a given assay or analytical approach.

[0166] Machine-Readable Medium: As used herein, “machine-readable medium” refers to a component, device, or other tangible media able to store instructions 902 and data temporarily or permanently and may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage (e.g., erasable programmable read-only memory (EEPROM)) and / or any suitable combination thereof. The term "machine-readable medium" may be taken to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) able to store instructions 902. The term "machine-readable medium" shall also be taken to include any medium, or combination of multiple media, that is capable of storing instructions 902 (e.g., code) for execution by a machine 900, such that the instructions 902, when executed by one or more processors 904 of the machine 900, cause the machine 900 to perform any one or more of the methodologies described herein. Accordingly, a "machine-readable medium" refers to a single storage apparatus or device, as well as "cloud-based" storage systems or storage networks that include multiple storage apparatus or devices. The term "machine-readable medium" excludes signals per se.

[0167] Maximum MAF: As used herein, “maximum MAP” or “max MAF” refers to the maximum MAF (mutant allele fraction) of all somatic variants in a sample.

[0168] Methylation: As used herein, “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in a 5:A 3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-Attorney Docket No.: GH0267WOformylcytosine (5-fC), and 5-carboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3-methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.

[0169] Methylation-Dependent Nuclease: As used herein, “methylation-dependent nuclease” refers to a nuclease that preferentially cuts methylated DNA relative to unmethylated DNA. For example, a methylation-dependent nuclease may cut at or near a recognition sequence such as a restriction site in a manner dependent on methylation of at least one of the nucleobases in the recognition sequence, such as a cytosine. In some embodiments, the nucleolytic activity of the methylation-dependent nuclease is at least 10, 20, 50, or 100-fold higher on a methylated recognition site relative to an unmethylated control in a standard nucleolysis assay. Methylationdependent nucleases include methylation-dependent restriction enzymes.

[0170] Methylation-Dependent Restriction Enzyme: As used herein, “methylationdependent restriction enzyme” or “MDRE” refers to a restriction enzyme that is dependent on methylation of the DNA (e.g. cytosine methylation) i.e., the presence or absence of methyl group in a nucleotide base alters the rate at which the enzyme cleaves the target DNA. In some embodiments, the methylation dependent restriction enzymes do not cleave the DNA if a particular nucleotide base is unmethylated at the recognition sequence. For example, MspJI is a methylation dependent restriction enzyme with a recognition sequence “mCNNR(N9)” and it does not cleave DNA if the absence of the methylated cytosine (mC) in the recognition sequence.

[0171] Methylation-Sensitive Nuclease: As used herein, “methylation-sensitive nuclease” refers to a nuclease that preferentially cuts unmethylated DNA relative to methylated DNA. For example, a methylation-sensitive nuclease may cut at or near a recognition sequence such as a restriction site in a manner dependent on lack of methylation of at least one of the nucleobases in the recognition sequence, such as a cytosine. In some embodiments, the nucleolytic activity of the methylation-sensitive nuclease is at least 10, 20, 50, or 100-fold higher on an unmethylated recognition site relative to a methylated control in a standard nucleolysis assay. Methylation-sensitive nucleases include methylation- sensitive restriction enzymes.Attorney Docket No.: GH0267WO

[0172] Methylation Sensitive Restriction Enzyme: As used herein, “methylation sensitive restriction enzyme” or “MSRE” refers to a restriction enzyme that is sensitive to the methylation status of the DNA (e.g. cytosine methylation) i.e., the presence or absence of methyl group in a nucleotide base alters the rate at which the enzyme cleaves the target DNA. In some embodiments, the methylation sensitive restriction enzymes do not cleave the DNA if a particular nucleotide base is methylated at the recognition sequence. For example, Hpall is a methylation sensitive restriction enzyme with a recognition sequence “CCGG” and it does not cleave DNA if the second cytosine in the recognition sequence is methylated.

[0173] Methylation rate: As used herein, “methylation rate” refers to the probability, likelihood, or percentage that a given base (for example: cytosine residue in a CpG) is methylated on a DNA molecule at a particular genomic region analyzed in the sample. In some embodiments, the methylation rate may be applied to a defined region that comprises one or more potentially methylated bases. In some embodiments, the methylation rate refers to the percentage of CpG residues methylated in a DNA molecule. In some embodiments, the methylation rate refers to the percentage of CpG residues methylated in molecules aligned to particular genomic position or genomic region. Methylation rate can be measured by a variety of methods including, but not limited to, either using bisulfite sequencing (any single base resolution like TAPS, EM-seq, etc.) or using partitioning (DNA molecule resolution). Methylation rate can be measured in different ways. One estimation can be by counting how many DNA fragments end up in each methylation dependent partition or by counting the number of converted CpGs per fragment in the case of bisulfite sequencing or any other base-level resolution sequencing methods. In addition, in the case of methylation dependent partitioning, the rate calculation can be normalized using a set of predefined regions with known methylation state (i.e., positive control regions and / or negative control regions) or spiked- in synthetic DNA with known methylation state, deriving rate-parametrized partition distributions and estimating the rate using a maximum likelihood approach. In one or more examples, the methylation rate can be determined by determining an abundance of sequencing reads that correspond to a portion of a genomic region. The portion of the genomic region can include a number of genomic locations of the genomic region for which at least a threshold number of sequencing reads overlap.

[0174] Methylation Status: As used herein, “methylation status” or “methylation state” can refer to the presence or absence of methyl group on a DNA base (e.g. cytosine) at a particular genomic position in a nucleic acid molecule. It can also refer to the degree of methylation in a nucleic acid sequence (e.g., highly methylated, low methylated, intermediately methylated orAttorney Docket No.: GH0267WOunmethylated nucleic acid molecules). The methylation status can also refer to the number of nucleotides methylated in a particular nucleic acid molecule.

[0175] Modified Nucleotide Specific Binding Reagent: As used herein, refers to a binding reagent that is specific for, or targets, modified nucleotides. For example, a modified nucleotide can be a nucleotide that has been methylated, thus, the binding reagent can be specific for a methylated nucleotide. Examples of binding reagents include, but are not limited to, a methyl binding domain (MBD) of a methylation binding protein (“MBP”) or variants thereof, an antibody (and antibody variants e.g., single chain antibodies), aptamers, or combinations thereof. Thus, as disclosed throughout, the use of MBD can be exchanged for any other modified nucleotide specific binding reagent, provided the modified nucleotide specific binding reagent has the desired specificity and affinity for the specific modified base of interest in the selected implementation.

[0176] Mutant Allele Fraction: As used herein, “mutant allele fraction”, “mutation dose,” or “MAF” refers to the fraction of nucleic acid molecules harboring an allelic alteration or mutation at a given genomic position in a given sample. MAF is generally expressed as a fraction or a percentage. For example, an MAF can be less than about 0.5, 0.1, 0.05, or 0.01 (i.e., less than about 50%, 10%, 5%, or 1%) of all somatic variants or alleles present at a given locus.

[0177] Mutation: As used herein, “mutation” refers to a variation from a known reference sequence and includes mutations such as, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / aberrations, insertions or deletions (indels), gene fusions, transversions, translocations, frame shifts, duplications, repeat expansions, and epigenetic variants. A mutation can be a germline or somatic mutation. In some examples, a reference sequence for purposes of comparison is a wildtype genomic sequence of the species of the subject providing a test sample, typically the human genome.

[0178] Mutation Caller. As used herein, “mutation caller” means an algorithm (embodied in software or otherwise computer implemented) that is used to identify mutations in test sample data (e.g., sequence information obtained from a subject).

[0179] Mutation Count: As used herein, “mutation count” or “mutational count” refers to the number of somatic mutations in a whole genome or exome or targeted regions of a nucleic acid sample.

[0180] Negative Control Region As used herein, “negative control region”, refers to a genomic region that is expected to be unmethylated or hypomethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell.

[0181] Neoplasm: As used herein, the terms “neoplasm” and “tumor” are used interchangeably. They refer to abnormal growth of cells in a subject. A neoplasm or tumor canAttorney Docket No.: GH0267WObe benign, potentially malignant, or malignant. A malignant tumor is referred to as a cancer or a cancerous tumor.

[0182] Next Generation Sequencing-. As used herein, “next generation sequencing” or “NGS” refers to sequencing technologies having increased throughput as compared to traditional Sanger- and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequencing reads at a time. Some examples of next generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0183] Nucleic Acid Tag-. As used herein, “nucleic acid tag” refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides in length), used to distinguish nucleic acids from different samples (e.g., representing a sample index), or different nucleic acid molecules in the same sample (e.g., representing a molecular barcode), of different types, or which have undergone different processing. The nucleic acid tag comprises a predetermined, fixed, non-random, random or semi-random oligonucleotide sequence. Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or sub-samples. Nucleic acid tags can be single-stranded, doublestranded, or at least partially double-stranded. Nucleic acid tags optionally have the same length or varied lengths. Nucleic acid tags can also include double-stranded molecules having one or more blunt-ends, include 5’ or 3’ single-stranded regions (e.g., an overhang), and / or include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one end or to both ends of the other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples comprising nucleic acids bearing different molecular barcodes and / or sample indexes in which the nucleic acids are subsequently being deconvolved by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g. molecular identifier, sample identifier). Additionally, or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or sub-sample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample, or non-uniquely tagging such molecules. In the case of non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their endogenous sequence information (for example, start and / or stop positions where they map to aAttorney Docket No.: GH0267WOselected reference sequence, a sub-sequence of one or both ends of a sequence, and / or length of a sequence) in combination with at least one molecular barcode. A sufficient number of different molecular barcodes are used such that there is a low probability (e.g., less than about a 10%, less than about a 5%, less than about a 1%, or less than about a 0.1% chance) that any two molecules may have the same endogenous sequence information (e.g., start and / or stop positions, subsequences of one or both ends of a sequence, and / or lengths) and also have the same molecular barcode.

[0184] Partitioning: As used herein, “partitioning” refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on a characteristic of the nucleic acid molecules. The partitioning can be physical partitioning of molecules. Partitioning can involve separating the nucleic acid molecules into groups or sets based on the level of epigenetic feature (for e.g., methylation). For example, the nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, the methods and systems used for partitioning may be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety.

[0185] Partitioned set: As used herein, “partitioned set” or “partition” refers to a set of nucleic acid molecules partitioned into a set or group based on the differential binding affinity of the nucleic acid molecules or proteins associated with the nucleic acid molecules to a binding agent. A partitioned set may also be referred to as a subsample. The binding agent binds preferentially to the nucleic acid molecules comprising nucleotides with epigenetic modification. For example, if the epigenetic modification is methylation, the binding agent can be a methyl binding domain (MBD) protein. In some embodiments, a partitioned set can comprise nucleic acid molecules belonging to a particular level or degree of epigenetic feature (for e.g., methylation). For example, the nucleic acid molecules can be partitioned into three sets - one set for highly methylated nucleic acid molecules (first subsample, hyper partition, hyper partitioned set or hypermethylated partitioned set), a second set for low methylated nucleic acid molecules (second subsample, hypo partition, hypo partitioned set or hypomethylated partitioned set), and a third set for intermediate methylated nucleic acid molecules (third subsample, intermediate partitioned set, intermediately methylated partitioned set, residual partition, or residual partitioned set). In another example, the nucleic acid molecules can be partitioned based on the number of methylated nucleotides - one partitioned set can have nucleic acid molecules with nine methylated nucleotides, and another partitioned set can have unmethylated nucleic acid molecules (zero methylated nucleotides).Attorney Docket No.: GH0267WO

[0186] Polynucleotide: As used herein, “polynucleotide”, “nucleic acid”, “nucleic acid molecule”, “polynucleotide molecule”, or “oligonucleotide” refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleosidic linkages. A polynucleotide can comprise at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g. 3-4, to hundreds of monomeric units. Whenever a polynucleotide is represented by a sequence of letters, such as “ATGCCTG,” it will be understood that the nucleotides are in 5’3’ order from left to right and that in the case of DNA, “A” denotes deoxyadenosine, “C” denotes deoxycytidine, “G” denotes deoxyguanosine, and “T” denotes deoxythymidine, unless otherwise noted. The letters A, C, G, and T may be used to refer to the bases themselves, to nucleosides, or to nucleotides comprising the bases, as is standard in the art.

[0187] Positive Control Region: As used herein, As used herein, “positive control region”, refers to a genomic region that is expected to be methylated or hypermethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell.

[0188] Probe: As used herein, “probe” refers to a polynucleotide comprising a functionality. The functionality can be a detectable label (fluorescent), a binding moiety (biotin), or a solid support (a magnetically attractable particle or a chip). Probes can include singlestranded DNA / RNA polynucleotides or double stranded DNA polynucleotides that hybridize to target nucleic acid sequences (e.g., SureSelect® probes, Agilent Technologies). Sequence capture using probes generally depends, in part, on the number of consecutive nucleotides in at least a portion of the target nucleic acid sequence that is complementary (or nearly complementary) to the sequence of the probe. In some examples, probes can correspond to driver mutations.

[0189] Processing: As used herein, the terms “processing”, “calculating”, and “comparing” can be used interchangeably. In certain applications, the terms refer to determining a difference, e.g., a difference in number or sequence. For example, gene expression, copy number variation (CNV), indel, and / or single nucleotide variant (SNV) values or sequences can be processed.

[0190] Processor. As used herein, “processor” refers to any circuit or virtual circuit (a physical circuit emulated by logic executing on an actual processor) that manipulates data values according to control signals (e.g., "commands," "op codes," "machine code," etc.) and which produces corresponding output signals that are applied to operate a machine. A processor may, for example, be a CPU, a RISC processor, a CISC processor, a GPU, a DSP, an ASIC, a RFICAttorney Docket No.: GH0267WOor any combination thereof. A processor may further be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that may execute instructions contemporaneously.

[0191] Promoter Region As used herein, “promoter region” refers to a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, required to initiate the specific transcription of a gene.

[0192] Quantitative Measures: As used herein, “quantitative measures” refers to an absolute or relative measure. A quantitative measure can be, without limitation, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or a relative quantity (e.g., high, medium, and low). A quantitative measure can be a ratio of two quantitative measures. A quantitative measure can be a linear combination of quantitative measures. A quantitative measure may be a normalized measure.

[0193] Reference Sequence: As used herein, “reference sequence” refers to a known sequence used for purposes of comparison with experimentally determined sequences. For example, a known sequence can be an entire genome, a chromosome, or any segment thereof. A reference sequence can include at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more nucleotides. A reference sequence can align with a single contiguous sequence of a genome or chromosome or can include noncontiguous segments that align with different regions of a genome or chromosome. Example reference sequences, include, for example, human genome reference sequences, such as, hG19 and hG38.

[0194] Sample: As used herein, “sample” means anything capable of being analyzed by the methods and / or systems disclosed herein.

[0195] Sensitivity: As used herein, “sensitivity” means the probability of detecting the presence of a single nucleotide variant, an insertion, and a deletion at a given MAF and coverage and the probability of detecting the presence of a copy number variant at a given tumor fraction and coverage.

[0196] Sequencing: As used herein, “sequencing” refers to any of a number of technologies used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Example sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxyAttorney Docket No.: GH0267WOtermination sequencing, whole-genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and a combination thereof. In some implementations, sequencing can be performer by a gene analyzer such as, for example, gene analyzers commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.

[0197] Single Nucleotide Variant. As used herein, “single nucleotide variant” or “SNV” means a mutation or variation in a single nucleotide that occurs at a specific position in the genome.

[0198] Somatic Mutation: As used herein, “somatic mutation” means a mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and accordingly, are not passed on to progeny.

[0199] Specifically binds: As used herein, “specifically binds” in the context of an probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence, or replicates thereof, to form a stable probe:target hybrid, while at the same time formation of stable probe: non-target hybrids is minimized. Thus, a probe hybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a non-target sequence, to enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, incorporated by reference herein).

[0200] Subject: As used herein, “subject” refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, e.g., a mammal such as a mouse, a primate, a simian or a human. Animals include farm animals (e.g., production cattle, dairy cattle, poultry, horses,Attorney Docket No.: GH0267WOpigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual that has or is suspected of having a disease or a predisposition to the disease, or an individual that is in need of therapy or suspected of needing therapy. The terms “individual” or “patient” are intended to be interchangeable with “subject.”

[0201] For example, a subject can be an individual who has been diagnosed with having a cancer, is going to receive a cancer therapy, and / or has received at least one cancer therapy. The subject can be in remission of a cancer. As another example, the subject can be an individual who is diagnosed of having an autoimmune disease. As another example, the subject can be a female individual who is pregnant or who is planning on getting pregnant, who may have been diagnosed of or suspected of having a disease, e.g., a cancer, an auto-immune disease.

[0202] Target Region: As used herein, “target region” refers to a genomic locus targeted for identification and / or capture, for example, by using probes (e.g., through sequence complementarity). A “target region set” or “set of target regions” refers to a plurality of genomic loci targeted for identification and / or capture, for example, by using a set of probes (e.g., through sequence complementarity)..

[0203] Threshold: As used herein, “threshold” refers to a predetermined value used to characterize experimentally determined values of the same parameter for different samples depending on their relation to the threshold.

[0204] Tumor Fraction: As used herein, “tumor fraction” refers to the estimate of the fraction of nucleic acid molecules derived from a tumor in a given sample. For example, the tumor fraction of a sample can be a measure derived from the max MAF of the sample or pattern of sequencing coverage of the sample or length of the cfDNA fragments in the sample or any other selected feature of the sample. In some instances, the tumor fraction of a sample is equal to the max MAF of the sample.

[0205] Variant: As used herein, a “variant” can be referred to as an allele. A variant is usually presented at a frequency of 50% (0.5) or 100% (1), depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. Somatic variants; however, are acquired variants and usually have a frequency of < 0.5. Major and minor alleles of a genetic locus refer to nucleic acids harboring the locus in which the locus is occupied by a nucleotide of a reference sequence, and a variant nucleotide different than the reference sequence respectively. Measurements at a locus can take the form of allelic fractions (AFs), which measure the frequency with which an allele is observed in a sample.Attorney Docket No.: GH0267WODETAILED DESCRIPTION

[0206] Cancer is usually caused by the accumulation of mutations within genes of an individual's cells, at least some of which result in improperly regulated cell division. Such mutations can include single nucleotide variations (SNVs), gene fusions, insertions, transversions, translocations, and inversions. These mutations can also include copy number variations that correspond to an increase or a decrease in the number of copies of a gene within a tumor genome relative to an individual’s noncancerous cells. An extent of mutations present in cell-free nucleic acids and an amount of mutated cell-free nucleic acids of a sample can be used as biomarkers to determine tumor progression, predict patient outcome, and refine treatment choices. In various examples, the extent of mutations present in cell-free nucleic acids can be indicated by tumor cells copy number and tumor fraction for a given sample.

[0207] Additionally, cancer can be indicated by non-sequence modifications, such as methylation. Examples of methylation changes in cancer include local gains of DNA methylation in the CpG islands at the TSS of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This increased amount of methylation can be associated with an aberrant loss of transcriptional capacity of involved genes and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression.

[0208] Thus, DNA methylation profiling can be used to detect aberrant methylation in DNA of a sample. The DNA can correspond to certain genomic regions (“differentially methylated regions’’ or “DMRs”) that are normally hypermethylated or hypomethylated in a given sample type (e.g., cfDNA from the bloodstream) but which may show an abnormal degree of methylation that correlates to a neoplasm or cancer, e.g., because of unusually increased contributions of tissues to the type of sample (e.g., due to increased shedding of DNA in or around the neoplasm or cancer) and / or from extents of methylation of the genome that are altered during development or that are perturbed by disease, for example, cancer or any cancer-associated disease.

[0209] Some methods of measuring DNA methylation can make accurately determining an amount of methylation of DNA difficult. The accuracy with which DNA methylation is determined can impact the accuracy of estimates of tumor fraction for samples. Since tumor fraction can be used to determine whether a sample is derived from a subject in which a tumor is present or not, the accuracy of determination of tumor fraction estimates can impact diagnosis and / or treatment decisions for individuals.

[0210] The methods and systems described herein are directed to accurately generating information indicating the amounts of methylation of nucleic acids using data that indicates an amount of binding of nucleic acids to methyl binding domain (MBD). In various examples, theAttorney Docket No.: GH0267WOapplication is directed to systems and processes to determine an estimate for tumor fraction of a sample. In one or more examples, amounts of methylation of nucleic acids can be determined based on a strength of binding by the nucleic acids to methyl binding domain (MBD). The nucleic acids can be partitioned according to the strength of binding to MBD. Additionally, a number of cytosine-guanine (CG) regions for the nucleic acids can be determined. Amounts of methylation of classification regions of the nucleic acids can be determined based on the partition information associated with the nucleic acids and the number of cytosine-guanine regions of the nucleic acids. The classification regions can have differing amounts of methylation in tumor cells and non-tumor cells. The estimate for tumor fraction of the sample can be determined according to the amounts of methylation of the classification regions.

[0211] In at least some implementations, the methods, systems, techniques, and architectures can implement models that are configured to have at least one of parameters or weights that can be modified to more accurately fit to the methylation data provided to the models. The methods, systems, techniques, and architectures are also directed to implementing a number of optimization procedures during the training of the models to generate models that more accurately predict metrics indicating the presence or absence of tumors than other systems, methods, techniques, and architectures. Further, the methods, techniques, and processes used to generate the information used to produce the methylation data reduce the amount of noise present in the methylation data that leads to more accurate predictions of metrics that indicate the presence or absence of tumors than other methods, techniques, and processes.

[0212] Figure 1 is a diagrammatic representation of an example environment 100 that identifies nucleic acids that correspond to classification regions of a reference sequence, where the classification regions have at least a threshold number of CpGs, according to one or more implementations. In one or more examples, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, gliomas, astrocytomas, breast carcinoma, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal carcinoma, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinomas, gastrointestinal stromal tumors (GISTs), endometrial carcinoma, endometrial stromal sarcomas, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder carcinomas, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinomas, Wilms tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myeloid (CML), chronic myelomonocyticAttorney Docket No.: GH0267WO(CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, Lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphomas, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, Mantle cell lymphoma, T cell lymphomas, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T cell lymphomas, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral cavity squamous cell carcinomas, osteosarcoma, ovarian carcinoma, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasms, acinar cell carcinomas. Prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine carcinomas, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.

[0213] The environment 100 can include a sample 102. The sample 102 can be derived from a biological fluid obtained from a subject. For example, the sample 102 can be derived from blood obtained from a subject. In one or more additional examples, the sample 102 can be derived from tissue of a subject. In various examples, the sample 102 can be derived from multiple sources. To illustrate, the sample 102 can be derived from one or more fluids of a subject and / or from tissue of a subject. In one or more illustrative examples, the subject can be a mammal. In one or more additional illustrative examples, the subject can be a human. In one or more further illustrative examples, the subject can be a non-human mammal.

[0214] The sample 102 can include a number of nucleic acids 104. Individual nucleic acids 104 can include a number of regions that have at least a threshold number of cytosine molecules and guanine molecules. In one or more examples, individual nucleic acids 104 can include regions having at least a threshold number of cytosine-guanine dinucleotides. In various examples, at least a portion of the cytosine-guanine pairs included in the regions can be sequentially located in sequences of the nucleic acids 104. In one or more illustrative examples, a region of a nucleic acid having at least a threshold amount of cytosine-guanine pairs can be referred to herein as a “CG region” or a “CpG region.” In one or more examples, a CG region can include at least 200 CpG dinucleotides. In one or more illustrative examples, a CG region can include from 200 CpG dinucleotides to 5000 CpG dinucleotides, from 300 CpG dinucleotides to 3000 CpG dinucleotides, from 200 CpG dinucleotides to 2500 CpG dinucleotides, or from 500 CpG dinucleotides to 1500 CpG dinucleotides. Additionally, a CG region can have a GC percentage of at least 50% and an observed-to-expected CpG ratio of at least 60%. The observed-to-expected CpG ratio can be calculated where the observed CpG is the number of CpGs identified in a given genomic region and the expected CpGs is the number of cytosines multiplied by the number of guanines divided by the number of bases in the genomic region. The expected CpGs can also be calculated by:Attorney Docket No.: GH0267WO((number of cytosines + number of guanines) / 2)2 / length of genomic region.For example, a CG region can be determined using the techniques described by Gardiner-Garden M, Frommer M (1987). " CpG islands in vertebrate genomes". Journal of Molecular Biology. 196 (2): 261-282. and / or Saxonov S, Berg P, Brutlag DL (2006). " A genome-wide analysis of CpG dinucleotides in the human genome distinguishes two distinct classes of promoters". Proc Natl Acad Sci USA. 103 (5): 1412-1417.

[0215] In the illustrative example of Figure 1, a portion of a sequence of an example nucleic acid 104 can include a first CG region 106, a second CG region 108, and a third CG region 110. Although the illustrative example of Figure 1 illustrates a portion of a sequence of a nucleic acid 104 having three CG regions, nucleic acids 104 included in the sample 102 can have a different number of CG regions. For example, individual nucleic acids 104 included in the sample 102 can include at least 1 CG region, at least 5 CG regions, at least 10 CG regions, at least 25 CG regions, at least 50 CG regions, at least 100 CG regions, at least 250 CG regions, at least 500 CG regions, or at least 1000 CG regions.

[0216] Individual CG regions can correspond to a number of molecules with one or more methylated cytosines. In the illustrative example of Figure 1, the CG region 106 can include a molecule with a methylated cytosine 112. In the illustrative example of Figure 1, the molecule with a methylated cytosine 112 is 5-methylcytosine. Individual CG regions can also correspond to a number of molecules with an unmethylated cytosine. For example, the CG region 106 can include a molecule with an unmethylated cytosine 116. In various examples, at least a portion of the CG regions of a nucleic acid 104 can correspond to classification regions of a reference genome. Classification regions can correspond to genomic regions of a reference genome that correspond to non-sequence differences that are consistent with one or more biological conditions, such as one or more types of cancer. In at least some examples, the non-sequence differences can include one or more mutations that are consistent with one or more biological conditions. In one or more examples, a classification region can correspond to a genomic region of the reference sequence for which molecules derived from subjects having at least one form of cancer. In at least some examples, nucleic acid molecules having at least a threshold amount of methylated cytosines in at least one CG region (e.g., hypermethylated molecules) can be derived from subjects in which cancer is present and correspond to a classification. In one or more additional examples, nucleic acid molecules having less than a threshold amount of methylated cytosines (e.g., hypomethylated molecules) in at least one CG region can be derived from subjects in which cancer is present and correspond to a classification region.Attorney Docket No.: GH0267WO

[0217] In addition to the classification regions, the CG regions can include one or more positive control regions, such as positive control region 118. The positive control region 118 can be mapped to nucleic acid molecules having at least a threshold number of methylated cytosine molecules in at least one CG region and that are derived from subjects that are free of cancer and are derived from subjects in which cancer is present. In various examples, the positive control region 118 can be hypermethylated in cells derived from subjects that are free of cancer and also in cells derived from subjects in which cancer is present. The CG regions can also include one or more negative control regions, such as negative control region 120. The negative control region 120 can be mapped to nucleic acid molecules having less than a threshold number of methylated cytosine molecules in at least one CG region and that are derived from subjects that are free of cancer and also subjects in which cancer is present. In one or more illustrative examples, the negative control region 120 can be hypomethylated in subjects that are free of cancer and also in subjects in which cancer is present. In various examples, the positive control regions and the negative control regions can be used to perform normalization calculations. The normalization calculations can be performed to generate input data for one or more models that are implemented to determine tumor metrics for a given sample 102.

[0218] A first molecule separation process 122 can be performed. The first molecule separation process 122 can separate nucleic acids 104 included in the sample 102 based on an amount of methylated cytosines of the individual nucleic acids 104. In one or more examples, the first molecule separation process can separate nucleic acids 104 included in the sample 102 based on amounts of methylated cytosines included in CG regions of individual nucleic acids 104. In various examples, the first molecule separation process 122 can separate the nucleic acids 104 into a plurality of groups with individual groups corresponding to respective amounts of methylated cytosines of the nucleic acids 104.

[0219] In the illustrative example of Figure 1, the first molecule separation process 122 can be performed in relation to a first methylation threshold 124. Performing the first molecule separation process 122 with regard to the first methylation threshold 124 can produce a first partition of nucleic acids 126. In one or more examples, the first methylation threshold 124 can indicate a first threshold number of molecules with a methylated cytosine located in CG regions of the nucleic acids 104. The first molecule separation process 122 can identify a number of nucleic acids 104 having fewer molecules with a methylated cytosine in CG regions than the first methylation threshold 124. In various examples, the first methylation threshold 124 can correspond to a first methylation rate.Attorney Docket No.: GH0267WO

[0220] The first molecule separation process 122 can also be performed with respect to a second methylation threshold 128. The second methylation threshold 128 can indicate an amount of methylated cytosines in one or more genomic regions of the nucleic acids 104 that is greater than the amount of methylated cytosines in the one or more regions corresponding to the first methylation threshold 124. The second methylation threshold 128 can indicate a number of molecules with a methylated cytosine per a number of nucleic acids. In one or more additional examples, the second methylation threshold 128 can correspond to a rate of methylation of nucleic acids that is greater than the rate of methylation that corresponds to the first methylation threshold 124. Performing the first molecule separation process 122 with respect to the second methylation threshold 128 can produce a second partition of nucleic acids 130. In one or more examples, the first molecule separation process 122 can identify nucleic acids 104 having a greater amount of methylated cytosines than the first methylation threshold 124 and having a lower amount of methylated cytosines than the second methylation threshold 128 to produce the second partition of nucleic acids 130.

[0221] Additionally, the first molecule separation process 122 can also be performed with respect to a third methylation threshold 132. The third methylation threshold 132 can indicate an amount of methylated cytosines in one or more genomic regions of the nucleic acids 104 that is greater than the amount of methylated cytosines in the one or more regions corresponding to the first methylation threshold 124 and greater than the amount of methylated cytosines in the one or more regions corresponding to the second methylation threshold 128. The third methylation threshold 132 can indicate a number of molecules with a methylated cytosine per a number of nucleic acids. In one or more additional examples, the third methylation threshold 132 can correspond to a rate of methylated cytosines that is greater than the rate of methylation that corresponds to the first methylation threshold 124 and greater than the rate of methylation that corresponds to the second methylation threshold 128. Performing the first molecule separation process 122 with respect to the third methylation threshold 132 can produce a third partition of nucleic acids 134. In one or more examples, the first molecule separation process 122 can identify nucleic acids 104 having a greater amount of methylated cytosines than nucleic acids 104 included in the second partition of nucleic acids 130. In this way, the amount of methylated cytosines of nucleic acids included in the first partition 126, the second partition 130, and the third partition 134 increases from the first partition 126 to the second partition 130 and increases from the second partition 130 to the third partition 134. In one or more illustrative examples, the first partition of nucleic acids 126 can be referred to as a hypomethylation partition, the secondAttorney Docket No.: GH0267WOpartition of nucleic acids 130 can be referred to as an intermediate partition, and the third partition of nucleic acids 134 can be referred to as a hypermethylation partition.

[0222] In one or more examples, the amount of methylated cytosines of nucleic acids can correspond to a strength of binding to methyl binding domain (MBD). In these scenarios, the first partition 126, the second partition 130, and the third partition 134 can be produced based on different strengths of binding to MBD for nucleotides having different amounts of methylated cytosines. In one or more examples, the first molecule separation process 122 can include a series of washes where the nucleic acids 104 are contacted with solutions having different concentrations of sodium chloride (NaCI).

[0223] Partitioning of the nucleic acids can be performed by contacting the nucleic acids with a modified nucleotide specific binding reagent, such as a MBD of a MBP. A modified nucleotide specific binding reagent can bind to 5-methylcytosine (5mC). The modified nucleotide specific binding reagent, such as a MBD, can be coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by increasing the NaCI concentration in a series of washes. The sequences eluted from the modified nucleotide specific binding reagent are partitioned into two or more fractions (e.g., hypo, hyper) depending on which wash (e.g., NaCI concentration) eluted the sequences. Resulting partitions can include one or more of the following nucleic acid forms: double-stranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments.

[0224] The binding of the nucleic acids with the modified nucleotide specific binding reagent can be a function of the number of methylated (or modified) sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCI concentration. Salt concentrations can, in one or more implementations, range from about 100 nM to about 2500 mM NaCI. In various implementations, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising a methyl binding domain, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration, a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as a “hypomethylated” population (hypo partition). For example, the first partition 126 can be representative of the hypomethylated form of DNA is that which remains unbound at a low salt concentration. In one or more illustrative examples, the concentration of NaCI of the solution used to produce the first partition 126 can be about 100 nM, about 120 nM, about 140 nM, about 160 nM, about 180 nM, about 200 nM. or about 250 nM. The second partition 130 can be referredAttorney Docket No.: GH0267WOto as a “residual partition” or an “intermediate partition” and can be representative of intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. In one or more additional illustrative examples, the concentration of NaCI of the solution used to produce the second partition 130 can be from about 100 mM to about 500 mM, from about 100 mM to about 1000 mM, from about 100 mM to about 1500 mM, from about 250 mM to about 1000 mM, from about 250 mM to about 1500 mM, from about 500 mM to about 1500 mM, from about 250 mM to about 2000 mM, from about 500 mM to about 2000 mM, or from about 1000 mM to about 2000 mM. This is also separated from the sample. The third partition 134 can be representative of the hypermethylated form of DNA (hyper partition) and is eluted using a high salt concentration, e.g., at least about 2000 mM. In one or more further illustrative examples, the concentration of NaCI of the solution used to produce the third partition 134 can be from about 2000 mM to about 5000 mM, from about 2000 mM to about 4000 mM, from about 2000 mM to about 3500 mM, from about 2000 mM to about 3000 mM, or from about 2500 mM to about 4000 mM.

[0225] In various examples, the first partition 126 can correspond to a first range of binding strengths of nucleic acids to MBD and to a first range of methylated CG regions, and the second partition 130 can correspond to a second range of binding strengths of nucleic acids to MBD and to a second range of methylated CG regions. The first range of binding strengths can be less than the second range of binding strengths. In one or more scenarios, a first solution having a first NaCI concentration can separate a first group of nucleic acids having the first range of binding strengths from MBD, and a second solution having a second NaCI concentration can separate a second group of nucleic acids having the second range of binding strengths from MBD, with the second NaCI concentration being greater than the first NaCI concentration. Additionally, the third partition 134 can correspond to a third range of binding strengths and a third range of methylated CG regions. The third range of binding strengths can be greater than the first range of binding strengths and the second range of binding strengths. In one or more instances, a third solution having a third NaCI concentration can separate a third group of nucleic acids having the third range of binding strengths from MBD. The third NaCI concentration can be greater than the first NaCI concentration and the second NaCI concentration.

[0226] In one or more illustrative examples, a plurality of nucleic acids derived from at least one of blood or tissue of a subject can be combined with a solution including an amount of MBD to produce a nucleic acid-MBD solution. A first wash of the nucleic acid-MBD solution can be performed with a first solution including a first NaCI concentration to produce a first nucleic acid fraction and a first residual solution. The first nucleic acid fraction can include a first portionAttorney Docket No.: GH0267WOof the plurality of nucleic acids and the first residual solution can include a second portion of the plurality of nucleic acids. In one or more examples, the first portion of the plurality of nucleic acids can have a first range of binding strengths to MBD that are less than a second range of binding strengths to MBD of the second portion of the plurality of nucleic acids.

[0227] Additionally, a second wash of the first residual solution can be performed with a second solution including a second concentration of NaCI that is greater than the first concentration of NaCI to produce a second nucleic acid fraction and a second residual solution. The second nucleic acid fraction can include a first subset of the second portion of the plurality of nucleic acids and the second residual solution can include a second subset of the second portion of the plurality of nucleic acids. The first subset of the second portion of the plurality of nucleic acids can have a third range of binding strengths to MBD that are less than a fourth range of binding strengths to MBD of the second subset of the second portion of the plurality of nucleic acids. Further, a third wash of the second residual solution can be performed with a third solution including a third concentration of NaCI that is greater than the second concentration of NaCI to produce a third nucleic acid fraction that includes the second subset of the second portion of the plurality of nucleic acids.

[0228] Subsequent to the first wash, the second wash, and the third wash, a determination can be made that the first portion of the plurality of nucleic acids are associated with the first partition 126. The first portion of the plurality of nucleic acids can be attached with molecular barcodes from a first set of molecular barcodes indicating the first partition 126. In this way, a sequencing read that corresponds to the first partition 126 can be identified based on determining that the sequencing read includes the first molecular barcode. In addition, a determination can be made that the first subset of the second portion of the plurality of nucleic acids is associated with an additional partition of the plurality of partitions. In these situations, a second set of molecular barcodes different from the first set of molecular barcodes can be attached to the first subset of the second portion of the plurality of nucleic acids with the second molecular barcode indicating the additional partition. As a result, a sequencing read that corresponds to the additional partition can be identified based on determining that the sequencing read includes one or more molecular barcodes from among the second set of molecular barcodes. Further, a determination can be made that the second subset of the second portion of the plurality of nucleic acids is associated with the second partition 130. A third set of molecular barcodes different from the first set of molecular barcodes and the second set of molecular barcodes can then be attached to the second subset of the second portion of the plurality of nucleic acids where the third set of molecular barcodes indicate the second partition 130. In these instances, a sequencing read thatAttorney Docket No.: GH0267WOcorresponds to the second partition 130 can be identified based on determining that the sequencing read includes a third molecular barcode from among the third set of molecular barcodes.

[0229] In at least some examples, the first molecule separation process 122 can result in nucleic acids being present in at least one of the first partition 126, the second partition 130, or the third partition 134 having an amount of methylation that is different from the amount of methylation of the other nucleic acids in the respective partition. For example, the first partition 126 can include a number of nucleic acids having amounts of methylation that correspond to the amounts of methylation of nucleic acids included in at least one of the second partition 130 or the third partition 134. Additionally, at least one of the second partition 130 or the third partition 134 can include nucleic acids having amounts of methylation that correspond to the amounts of methylation of nucleic acids included in the first partition 126. The presence of nucleic acids in at least one of the first partition 126, the second partition 130, or the third partition 134 that do not correspond to the amounts of methylation of at least a majority of the other nucleic acids included in the respective partition can cause data noise when performing computational operations with respect to sequence reads produced from nucleic acids included in the first partition 126, the second partition 130, and the third partition 134. The data noise can result in inaccuracies with respect to calculations made based on sequence reads derived from nucleic acids included in the first partition 126, the second partition 130, and the third partition 134.

[0230] To reduce or eliminate data noise associated with nucleic acids being present in at least one of the first partition 126, the second partition 130, or the third partition 134 that have amounts of methylation that are not consistent with the amounts of methylation of at least a majority of other molecules included in the respective partitions, a second molecule separation process 136 can be performed after the first molecule separation process 122. The second molecule separation process 136 can be performed with respect to nucleic acids included in the first partition 126, nucleic acids included in the second partition 130, and nucleic acids included in the third partition 134. In one or more examples, the second molecule separation process 136 can include performing digestion of the nucleic acids included in the first partition 126 using methylation dependent restriction enzyme (MDRE) and nucleic acids included in the second partition 130 and the third partition 134 can be digested using methylation sensitive restriction enzyme (MSRE). Digestion of the nucleic acids included in the first partition 126 with MDRE can result in separation of nucleic acids included in the first partition having amounts of methylation corresponding to the second partition 130 and the third partition 134 from nucleic acids having amounts of methylation corresponding to the first partition. Additionally, digestion of nucleic acidsAttorney Docket No.: GH0267WOincluded in the second partition 130 and the third partition 134 with MSRE can result in separation of the nucleic acids having amounts of methylation corresponding to the first partition 126 from the nucleic acids of the second partition 130 and the nucleic acids of the third partition 134. By removing nucleic acids from the first partition 126 having amounts of methylation that correspond to the second partition 130 and the third partition 134 and by removing nucleic acids from the second partition 130 and the third partition 134 that have amounts of methylation that correspond to the first partition 126, an additional group of nucleic acids 138 can be produced. The additional group of nucleic acids 138 can include nucleic acids corresponding to methylation amounts of the second partition 130 and the third partition 134 with a minimal amount or no nucleic acids having amounts of methylation corresponding to the first partition 126. For example, less than 50% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 50% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 60% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 70% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 90% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 95% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 97% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 99% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, at least 99.5% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134, or at least 99.9% of the nucleic acids included in the additional group 138 can have amounts of methylation that correspond to the second partition 130 and the third partition 134.

[0231] The environment 100 can include a sequencing machine 140. In one or more examples, the sequencing machine 140 can be any of a number of sequencing machines that can perform one or more sequencing operations that amplify nucleic acids present in a sample 102. In various examples, the sequencing machine 140 can perform next-generation sequencing operations. In one or more examples, the sample 102 can include an amount of at least one bodilyAttorney Docket No.: GH0267WOfluid extracted from a subject. In one or more additional examples, the sample 102 can include a tissue sample that is obtained from a subject.

[0232] In one or more examples, prior to sequencing, the extracted polynucleotides can be partitioned into two or more partitions based on the binding strengths of polynucleotides to MBD. A blunt-end ligation can be performed on the partitioned polynucleotides and adapters, as well as tags (e.g., molecular barcodes) can be added to the partitioned polynucleotides. The tagged polynucleotides in the one or more partitions (e.g. hyper and / or intermediate partitions) can be treated with one or more methylation sensitive restriction enzymes (MSREs). In some examples, the hypo partition can be treated with one or more methylated dependent restriction enzymes (MDREs). Post the MSRE and / or MDRE treatment, the molecules can also be enriched by causing hybridization between the extracted polynucleotides and probes that correspond to target regions of a reference sequence. The enrichment process can identify thousands, hundreds of thousands, up to millions of polynucleotides that correspond to on-target regions associated with the probes.

[0233] Subsequent and / or prior to the enrichment process, the molecules can be amplified according to one or more amplification processes. The one or more amplification processes can produce thousands, up to millions of copies of individual nucleic acid molecules. In one or more examples, a portion of the unenriched polynucleotides can be amplified, in some instances, but not to the extent that the enriched polynucleotides are amplified. The one or more amplification processes can generate an amplification product that undergoes one or more sequencing operations. After performing one or more sequencing operations with respect to the sample 102, the sequencing machine 140 can produce a sequencing data 142.

[0234] The sequencing data 142 can include alphanumeric representations of the nucleic acids included in an amplification product. For example, the sequencing data 142 can include, for individual nucleic acids of the amplification product, data that corresponds to a string of letters that represent the respective chains of nucleotides that correspond to the individual nucleic acids.

[0235] The sequencing data 142 can be stored in one or more data files. For example, the sequencing data 142 can be stored in a FASTQ file that comprises a text-based sequencing data file format storing raw sequence data and quality scores. In one or more additional examples, the sequencing data 142 can be stored in a data file according to a binary base call (BCL) sequence file format. In one or more further examples, the sequencing data 142 can be stored in a BAM file. In one or more examples, the sequencing data 142 can comprise at least about one gigabyte (GB), at least about 2 GB, at least about 3 GB, at least about 4 GB, at least about 5 GB, at least about 8 GB, or at least about 10 GB. An individual sequence representation included inAttorney Docket No.: GH0267WOthe sequencing data 142 can be referred to herein as a “read” or a “sequencing read.” In various examples, individual first nucleic acids included in the pool 138 can correspond to multiple sequence representations included in the sequencing data 142 as a result of the amplification of the individual first nucleic acids. In one or more additional examples, individual second nucleic acids included in the pool 138 can correspond to a single sequence representation included in the sequencing data 142 as a result of the absence of amplification of the individual second nucleic acids.

[0236] FIG. 2 is an example architecture 200 to analyze sequencing data to determine one or more metrics indicating the presence of a tumor in subjects, in accordance with one or more implementations. The architecture 200 can include one or more sequencing machines 202 that perform one or more sequencing operations with respect to a number of samples 204. The one or more samples 204 can be obtained from subjects 206. In one or more illustrative examples, a first portion of the subjects 206 can be free of cancer. That is, a tumor is not detected in the first portion of the subjects 206. Additionally, a tumor can be present in a second portion of the subjects 206.

[0237] One or more molecule separation processes 208 can be performed with respect to the samples 204. The one or more separation processes 208 can correspond to separating nucleic acid molecules into a number of partitions based on the characteristics of the nucleic acid molecules. Examples of characteristics that can be used for partitioning nucleic acid molecules include multiple different nucleotide modifications, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. In one or more illustrative examples, a heterogeneous population of nucleic acid molecules can be partitioned into nucleic acid molecules with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include, but are not limited to, presence or absence of methylation; level of methylation, hydroxymethylation, and type of methylation (5' cytosine or 6-methyladenine).

[0238] Prior to the one or more molecule separation processes 208, nucleic acid molecules can be extracted from a sample 204. In one or more implementations, the nucleic acid molecules comprise cell-free nucleic acids (e.g., cell-free DNA). In various implementations, the sample 204 can be a sample selected from one or more of blood, plasma, serum, urine, fecal, saliva samples, combinations thereof, and / or the like. In one or more additional examples, the sample 204 can comprise a sample selected from one or more of whole blood, a blood fraction, a tissue biopsy, pleural fluid, pericardial fluid, cerebrospinal fluid, and peritoneal fluid. In one or more illustrative examples, the cell-free nucleic acid molecules can be extracted from the sampleAttorney Docket No.: GH0267WO204 where the sample 204 is obtained from a subject 206 known to have cancer (e.g., a cancer patient), or a subject 206 suspected of having cancer.

[0239] The extraction of nucleic acid molecules from the sample 204 can include implementing one or more cell lysis techniques to cleave the membranes of cells included in the sample 204 and applying one or more proteases to break down proteins included in the sample 204. The extraction of nucleic acid molecules from the sample 204 can also include a number of washing and / or elution techniques to separate the nucleic acid molecules from other components included in the sample 204. In various examples, thousands, up to millions, up to billions of nucleic acid molecules can be extracted from the sample 204 prior to being subjected to the one or more separation processes 208.

[0240] The nucleic acid molecules extracted from samples 204 can include molecules having varying levels of methylation. Methylation can occur from any one or more post-replication or transcriptional modifications. Post-replication modifications include modifications of the nucleotide cytosine, including, but not limited to, 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine. The one or more molecule separation processes 208 can separate nucleic acid molecules extracted from samples 204 into a number of partitions with individual partitions corresponding to different levels of methylation. For example, the molecule separation processes 208 can produce a first partition of nucleic acid molecules having first levels of methylation, a second partition of nucleic acid molecules having second levels of methylation, and a third partition of nucleic acid molecules having third levels of methylation. In various examples, the second levels of methylation can be greater than the first levels of methylation and the third levels of methylation can be greater than the first levels of methylation and the second levels of methylation. In one or more illustrative examples, the one or more molecule separation processes 208 can include the first molecule separation process 122 and the second molecule separation process 136 of Figure 1.

[0241] The one or more molecule separation processes 208 can produce a pool 210 that includes a portion of the nucleic acid molecules extracted from one or more samples 204 and subjected to the one or more molecule separation processes 208. For example, the pool 210 can include a number of nucleic acid molecules having the second levels of methylation and a number of nucleic acid molecules having the third levels of methylation. Thus, the nucleic acid molecules included in the pool 210 can have at least a threshold amount of methylation. In one or more illustrative examples, the nucleic acid molecules included in the pool 210 can have at least a threshold amount of methylation in CG regions of the nucleic acid molecules.Attorney Docket No.: GH0267WO

[0242] The one or more sequencing machines 202 can perform one or more sequencing operations to produce sequencing data 212 that corresponds to the pool 210. The architecture 200 can include a computing system 214 that obtains the sequencing data 212 from the one or more sequencing machines 202 and analyzes the sequencing data 212. For example, the computing system 214 can analyze the sequencing data 212 to determine one or more metrics indicating that a tumor may be present in a subject 206 that provided at least one sample 204. The computing system 214 can include one or more computing devices 216. The one or more computing devices 216 can include at least one of one or more desktop computing devices, one or more mobile computing devices, or one or more server computing devices. In various examples, at least a portion of the one or more computing devices 216 can be included in a remote computing environment, such as a cloud computing environment. In one or more examples, the computing system 214 and the sequencing machine 202 can be owned, operated, maintained, and / or controlled by a single organization. In one or more additional examples, the computing system 214 and the sequencing machine 202 can be owned, operated, maintained, and / or controlled by multiple organizations.

[0243] At operation 218, the computing system 214 can analyze the sequencing data 212. Analyzing the sequencing data 212 can include determining one or more first sequence representations 220 included in the sequencing data 212 that correspond to one or more classification regions of a reference sequence. The one or more classification regions can correspond to genomic regions of a reference sequence that are mapped to nucleic acid molecules having an amount of methylation in cfDNA obtained from subjects in which cancer is present relative to an amount of methylation of the molecules that map to the same genomic regions of the reference sequence in cfDNA obtained from subjects in which a tumor is not present. In at least some examples, the amount of methylation present in nucleic acid molecules that map to a classification region and are derived from subjects in which cancer is present is less than the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from subjects in which cancer is not present. In one or more additional examples, the amount of methylation present in nucleic acid molecules that map to a classification region and are derived from subjects in which cancer is present is greater than the amount of methylation present in nucleic acid molecules that map to the classification region and are derived from subjects in which cancer is not present. The one or more classification regions can also include at least a threshold amount of cytosine-guanine content. In various examples, the one or more classification regions can include a series of cytosine-guanine (CG) pairs in the 5’— >3’ direction (CpG sites), such as at least 3 CpG sites, at least 5 CpG sites, at least 8 CpG sites, atAttorney Docket No.: GH0267WOleast 10 CpG sites, at least 12 CpG sites, at least 15 CpG sites, at least 18 CpG sites, or at least 20 CpG sites.

[0244] In addition, the computing system 214 can analyze the sequencing data 212 to determine one or more second sequence representations 222 that correspond to one or more control regions of a reference sequence. The one or more control regions can include one or more positive control regions and / or one or more negative control regions. In various examples, a positive control region can comprise a genomic region of a reference sequence having at least a threshold amount of molecules with a methylated cytosine and including at least a threshold number of CpG sites. A positive control region can correspond to nucleic acid molecules having at least a threshold amount of methylation in one or more CG regions and that are obtained from subjects in which cancer is present and in samples obtained from subjects in which a tumor is not present. In at least some examples, the threshold amount of methylation can correspond to at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 15 or more CpGs being methylated in nucleic acid molecules. In one or more illustrative examples, positive control regions can be mapped to nucleic acid molecules that are hypermethylated in one or more CG regions and are derived from samples obtained from both subjects in which cancer is present and subjects in which cancer is not present. In one or more examples, a negative control region can comprise a genomic region of a reference sequence having less than a threshold amount of molecules with a methylated cytosine and at least a threshold number of CpG sites. A negative control region can correspond to nucleic acid molecules having less than an additional threshold amount of methylation in one or more CG regions and that are obtained from subjects in which cancer is present and in samples obtained from subjects in which a tumor is not present. In various examples, the additional threshold amount of methylation can correspond to no greater than 1, no greater than 2, no greater than 3, no greater than 4, no greater than 5, no greater than 6, or no greater than 7 CpGs being methylated in nucleic acid molecules. In one or more additional illustrative examples, negative control regions can be mapped to nucleic acid molecules that are hypomethylated in one or more CG regions and are derived from samples obtained from both subjects in which cancer is present and subjects in which cancer is not present.

[0245] In one or more illustrative examples, the first sequence representations 220 can be determined by aligning sequence representations included in the sequencing data 212 with one or more classification regions of a reference sequence. In addition, the second sequence representations 222 can be determined by aligning sequence representations included in the sequencing data 212 with one or more control regions of a reference sequence. The alignmentAttorney Docket No.: GH0267WOprocess can identify the first sequence representations 220 by determining a number of sequence representations included in the sequencing data 212 that correspond to one or more classification regions of the reference sequence. Further, the alignment process can identify the second sequence representations 222 by determining a number of sequence representations that correspond to one or more control regions of the reference sequence.

[0246] In one or more illustrative examples, the alignment process can determine an amount of homology between individual sequence representations included in the sequencing data 212 and portions of the reference sequence. The amount of homology between a given sequence representation and the reference sequence can indicate a number of positions of the reference sequence that have the same nucleotide as corresponding positions of the given sequence representation. The computing system 214 can determine that a sequence representation is aligned with a portion of a reference sequence based on determining that the sequence representation and the portion of the reference sequence have at least a threshold amount of homology. In scenarios where a sequence representation has at least the threshold amount of homology with respect to multiple portions of the reference sequence, the portion of the reference sequence having the greatest amount of homology with the sequence representation can be determined to be aligned with the sequence representation.

[0247] The amount of homology between a given sequence representation and a portion of a reference sequence can be determined using BLAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) or by using the Gap program (Wisconsin Sequence Analysis Package, Genetics Computer Group, University Research Park, Madison Wis.), using default settings, which uses the algorithm of Needleman and Wunsch (J. Mol. Biol. 48; 443-453 (1970)). The amount of homology between a sequence representation and a portion of the reference sequence can also be determined using a Burrows-Wheeler aligner (Li, H., & Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 25(14), 1754-1760).

[0248] In one or more examples, after the sequence representations included in the sequencing data 212 have been aligned with a reference sequence, the aligned sequence representations can be analyzed to identify one or more groups of sequence representations. For example, individual aligned sequence representations can correspond to individual sequencing reads that are included in the sequencing data 212. In these scenarios, the aligned sequence representations can include multiple reads that correspond to a single nucleic acid molecule included in the sample pool 210. In one or more additional examples, the aligned sequenceAttorney Docket No.: GH0267WOrepresentations can correspond to individual nucleic acid molecules included in the pool 210. In these situations, the computing system can determine a group of reads included in the sequence data 212 that correspond to an individual nucleic acid molecule included in the pool 210 based on molecular barcodes that are common to each group of sequencing reads. That is, individual nucleic acid molecules included in the pool 210 can be encoded with molecular barcodes that uniquely identify the individual nucleic acid molecules and, in at least some cases, the individual nucleic acid molecules can be represented by multiple sequencing reads included in the sequencing data 212. Accordingly, when multiple sequence representations are present in the sequencing data 212 that correspond to a single nucleic acid molecule included in the pool 210, the computing system 214 can group the multiple sequence representations together. In various examples, the groups of sequence representations that correspond to a single nucleic acid molecule included in the pool 210 can be referred to herein as “families.” Additionally, start and stop positions with respect to the reference sequence of the aligned sequence representations having a common molecular barcode can be used to group the sequence representations that correspond to individual nucleic acids included in the pool 210. In one or more illustrative examples, an individual sequence representation that represents a family of sequence representations that corresponds to a single nucleic acid molecule included in the pool 210 can be referred to herein as a “consensus sequence representation.”

[0249] At operation 224, the computing system 214 can analyze the first sequence representations and the second sequence representations 222 to generate metrics that correspond to individual classification regions. In the illustrative example of Figure 2, the computing system 214 can analyze the first sequence representations 220 and the second sequence representations 222 to generate classification region metrics 226. The classification region metrics 226 can include quantitative measures determined based on a number of first sequence representations 220 having at least a threshold amount of methylated cytosines. In one or more illustrative examples, the classification region metrics 226 can include quantitative measures determined based on a number of sequencing reads corresponding to a number of the first sequence representations 220 having at least a threshold amount of methylated cytosines located. In one or more additional illustrative examples, the classification region metrics 226 can include quantitative measures determined based on a number of nucleic acid molecules that correspond to a number of the first sequence representations 220. In various examples, the classification region metrics 226 can include quantitative measures determined based on a number of first sequence representations 220 having at least a threshold amount of methylated cytosines and a number of second sequence representations 222 that correspond to controlAttorney Docket No.: GH0267WOregions of a reference sequence. In one or more further illustrative examples, the classification region metrics 226 can include quantitative measures related to a ratio of a number of first sequence representations 220 having at least a threshold amount of methylated cytosines in relation to a number of second sequence representations. In at least some examples, the sequence representations of the second sequence representations 222 used by the computing system 214 to generate quantitative measures included in the classification region metrics 226 can include sequence representations that correspond to positive control regions of a reference sequence.

[0250] The classification region metrics 226 can also be determined by performing one or more normalization operations with respect to quantitative measures generated by the computing system 214 using at least one of the first sequence representations 220 and the second sequence representations 222. For example, a logarithm calculation can be performed with respect to quantitative measures generated by the computing system 214 using at least one of the first sequence representations 220 or the second sequence representations 222. Additionally, the classification region metrics 226 can be determined by adding a pseudocount to quantitative measures determined by the computing system 214 using at least one of the first sequence representations 220 or the second sequence representations 222. In one or more illustrative examples, the one or more normalization operations can include determining quantitative measures that correspond to a ratio of first sequence representations 220 for an individual classification region with respect to a number of second sequence representations 222 that correspond to positive control regions of a reference sequence.

[0251] In one or more illustrative examples, the computing system 214 can determine a number of the first sequence representations 220 that correspond to individual classification regions of a reference sequence and that have at least a threshold amount of methylated cytosines located in the individual classification regions. In these scenarios, the computing system 214 can determine individual classification region metrics 226 for individual classification regions. In addition, the computing system 214 can determine a number of the second sequence representations 222 that correspond to positive control regions. In at least some examples, the computing system 214 can, for individual classification regions, determine a ratio including a number of first sequence representations 220 that correspond to the individual classification region and that have at least a threshold amount of molecules with a methylated cytosine in the classification region in relation to a total number of the second sequence representations 222 that correspond to positive control regions of a reference sequence. In one or more examples, the computing system 214 can add a value of a pseudocount to the ratio to determine a classificationAttorney Docket No.: GH0267WOregion metric 226 for the individual classification region. The value of the pseudocount can be at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2. Further, the computing system 214 can perform a log base 10 operation with respect to the combination of the ratio and the pseudocount to determine a classification region metric 226 for an individual classification region. In at least some illustrative examples, the computing system 214 can determine at least a portion of the classification region metrics according to the following equation:Score of region i — log\Q( - pseudocount )^positive.control (1 ),where x, is a total number of first sequence representations 220 for an individual classification region, i, having at least a threshold amount of methylated cytosines included in the region, I, and Xpositive_controi is a total number of the second sequence representations 222 that correspond to positive control regions of a reference sequence.

[0252] At operation 228, the computing system 214 can execute a model to determine an indication of cancer based on the classification region metrics 226. In the illustrative example of Figure 2, the computing system 214 can execute a model using the classification region metrics 226 to generate model output 230. In one or more examples, the model output 230 can indicate a status of tumor detection 232 or a status of tumor not detected 234 in relation to a sample 204 provided by a subject 206. In one or more additional examples, the computing system 214 can execute a model to determine an estimate of tumor fraction 236 for a sample 204. In one or more further examples, the computing system 214 can execute a model to determine a probability of a tumor being present in a subject 206 that provided a sample 204.

[0253] In one or more examples, the model can include a classification model that implements one or more machine learning techniques. In one or more illustrative examples, the model can include a linear regression model. In various examples, the model can be executed to determine a probability of a tumor being present 238 in a subject 206 that provided a sample 204 based on the classification region metrics 226. In one or more illustrative examples, the computing system 214 can execute the model to determine weights for individual classification regions. The weights for individual classification regions can be different. For example, the computing system 214 can determine that a first weight of a first classification region metric 226 for a first classification region is different from a second weight of a second classification region metric 226 for a second classification region. In at least some illustrative examples, a probability of a tumor being present 238 in a subject 206 that provided a sample 204 can be determined by the computing system 214 by executing a model that corresponds to the following equation:Attorney Docket No.: GH0267WO1 P(cancer\ 'reqiQn scores) ' — - J q_f>- ~ w;(score of regi: -on i T)+TTl>where w, is a weight of an individual classification region, the score of the region i is calculated using Equation (1), and b is a slope corresponding to a linear regression model. In at least some examples, the probability of a tumor being present 238 can be used to generate a status of tumor detected 232 or a status of tumor not detected 234. In one or more further illustrative examples, the computing system 214 can analyze the probability of a tumor being present 238 with respect to a threshold probability to determine a status of tumor detected 232 or a status of tumor not detected 234 for a sample 204. The computing system 214 can determine that a sample 204 corresponds to the status of tumor detected 232 in response to determining that a probability of a tumor being present 238 for the sample 204 is at least the threshold probability. Additionally, the computing system 214 can determine that a sample 204 corresponds to the status of tumor not detected 234 in response to determining that a probability of a tumor being present 238 for the sample 204 is less than the threshold probability.

[0254] In one or more additional examples, the computing system 214 can execute a model that determines a maximum mutant allele fraction (MAF). In various examples, the computing system 214 can execute a model using the maximum MAF value to determine tumor fraction 236 for a sample 204. In one or more illustrative examples, the computing system 214 can execute a model using the classification region metrics 226 to determine a logit transformed maximum MAF value that can then be used by the computing system 214 to estimate tumor fraction for a sample 204. In various examples, the computing system 214 can analyze maximum MAF values to determine a probability of cancer being present 238 in a subject 206 that provided a sample 204. In various examples, a Huber regression (Huber, P. J. 1964. “Robust Estimation of a Location Parameter.” Annals of Mathematical Statistics 35 (1): 73-101) can be performed to determine a maximum MAF value based on the classification region metrics 226.

[0255] In various examples, the model output 230 can also include a tumor tissue indication 240. The tumor tissue indication 240 can indicate one or more tissues from which cancer cells that produced genomic material detected in the sample 204 originate. In one or more examples, the tumor tissue indication 240 can correspond to one or more tissues of origin for cancer cells that produced genomic material detected in the sample 204. In these scenarios, the computing system 214 can generate multiple models with individual models corresponding to a given tissue type. The output from individual models can be analyzed to determine additional metrics that indicate a tissue from which cancer cells that produced genomic material detected in one or more samples originate. In at least some examples, the output for the individual modelsAttorney Docket No.: GH0267WOcan indicate at least one of tumor fraction 236 or a probability of tumor being present 238. The computing system 214 can analyze the respective model outputs to determine the model having at least one of a greatest value for tumor fraction or a greatest probability of cancer being present. The computing system 214 can then generate a tumor tissue indication 240 that corresponds to the model having the greatest value for tumor fraction and / or a greatest probability of cancer being present.

[0256] For example, samples 204 can be obtained from subjects 206 in which different types of cancer are present. To illustrate, first samples can be obtained from a first group of subjects in which a first classification of cancer is present and second samples can be obtained from a second group of subjects in which a second classification of cancer is present. The sequencing data generated from the first samples can be analyzed by the computing system 214 to generate first metrics that correspond to classification regions for the first classification of cancer and the first metrics can be used to generate a first model that corresponds to the first classification of cancer. Additionally, the sequencing data generated from second samples can be analyzed by the computing system 214 to generate second metrics that correspond to classification regions for the second classification of cancer and the second metrics can be used to generate a second model that corresponds to the second classification of cancer. After the models have been trained, the computing system 214 can analyze sequencing data obtained from one or more additional subjects that were not included in the training subjects to determine classification region metrics for the one or more additional subjects. The classification region metrics can then be analyzed using the different tumor classification models to generate model outputs. The model outputs can be analyzed by the computing system to determine a model having the greatest values for the respective model outputs and determine the tumor tissue classification that corresponds to the model.

[0257] In one or more additional illustrative implementations, the model output 230 can also indicate methylation status for one or more genomic regions of a reference sequence. For example, the computing system 214 can analyze the classification region metrics 226 to determine a methylation status of one or more promoter regions of a reference sequence. In various examples, the one or more promoter regions can include at least one promoter region that is related to the presence of a tumor in a subject. In one or more illustrative examples, the classification region metrics 226 can indicate a number of sequence representations having at least a threshold amount of methylation with respect to the one or more promoter regions. In these scenarios, the computing system 214 can determine that a promoter region is methylated in response to determining that the number of sequence representations having at least a thresholdAttorney Docket No.: GH0267WOamount of molecules with a methylated cytosine in the promoter region is greater than a threshold number.

[0258] In still other illustrative implementations, the computing system 214 can analyze the sequencing data 212 to determine a quantitative measure that corresponds to a number of sequence representations that correspond to a promoter region and determine an additional quantitative measure that corresponds to an additional number of sequence representations that correspond to a number of positive control regions. Normalized metrics can be determined based on the quantitative measure and the additional quantitative measure. The normalized metrics can be analyzed with respect to a threshold to determine a methylation status of the promoter region. In various examples, the threshold may be different for different promoter regions. The threshold can be determined using a cancer-free dataset and the threshold value is set such that false positive rates are no greater than 5%, no greater than 4%, no greater than 3%, no greater than 2%, or no greater than 1% and at a specificity of at least 80%, at least 90%, at least 95%, at least 98%, or at least 99%. In at least some examples, the computing system 214 can analyze the promoter region with regard to the threshold in situations where at least 4 sequence representations correspond to the promoter region, where at least 5 sequence representations correspond to the promoter region, where at least 6 sequence representations correspond to the promoter region, where at least 7 sequence representations correspond to the promoter region, where at least 8 sequence representations correspond to the promoter region, where at least 9 sequence representations correspond to the promoter region, where at least 10 sequence representations correspond to the promoter region, where at least 11 sequence representations correspond to the promoter region, or where at least 12 sequence representations correspond to the promoter region.

[0259] In one or more further illustrative examples, the computing system 214 can combine results from multiple models to determine the model output 230. For example, the computing system 214 can execute models with respect to one or more epigenetic signals, such as methylation of classification regions, to determine one or more first tumor metrics. With regard to methylation, the computing system 214 can execute both a classification model, such as a logistic regression model, that produces an indication of cancer being present in a subject providing a sample and an additional model that predicts tumor fraction for a sample. In one or more examples, the epigenetic signals can also correspond to fragment lengths of sequence representations generated from samples. In addition, the computing system 214 can execute one or more additional models with respect to genomic signals to generate further tumor metrics with respect to samples. In various examples, the genomic signals can correspond to the presence ofAttorney Docket No.: GH0267WOone or more single nucleotide variants (SNVs) and / or the presence of insertions or deletions at one or more genomic regions of a reference sequence. In at least some examples, the computing system can include an integration system that combines tumor metrics generated by executing a number of models with regard to data corresponding to the genomic signals and the epigenetic signals to produce an aggregated tumor metric for a given sample. In some embodiments, the quantitative measure obtained from the model output 230 can be analyzed with respect to a threshold. In situations where the quantitative measure obtained from the model output is at least the threshold value, the computing system 214 can determine that the indication of cancer is present in the subject. In situations where the quantitative measure obtained from the model output is less than the threshold value, the computing system 214 can determine that the indication of cancer is absent or not detected in the subject. In some embodiments, the threshold used to determine whether the indication of cancer is present is calculated using a set of normal samples and is set at a particular value that provides high specificity.

[0260] In various additional implementations, the computing system 214 can determine methylation status of individual genomic regions. In one or more illustrative examples, the computing system 214 can determine methylation status of one or more promoter regions. In one or more examples, the sequencing data 212 can be analyzed to determine sequence representations that correspond to one or more genomic regions. For example, the sequencing data 212 can be analyzed to determine a number of sequence representations that correspond to one or more promoter regions. In at least some examples, the computing system 214 can determine a number of sequence representations that correspond to individual promoter regions that have at least a threshold amount of methylated cytosines.

[0261] For each genomic region and for an individual sample, the computing system 214 can determine a number of sequence representations that correspond to polynucleotide molecules having at least the threshold number of methylated cytosines in the genomic region. The computing system 214 can perform one or more normalization operations using the counts of polynucleotide molecules or sequence reads that correspond to the genomic region and have at least the threshold number of methylated cytosines to generate normalized metrics. To illustrate, the computing system 214 can divide the counts of polynucleotide molecules or reads that correspond to the genomic region and have at least the threshold number of methylated cytosines by the number of molecules or sequencing reads that correspond to a control region, such as a positive control region. In another instance, the computing system 214 can perform the normalized metrics by dividing the counts of polynucleotide molecules or reads that correspond to the genomic region and have at least the threshold number of methylated cytosinesAttorney Docket No.: GH0267WOby the number of molecules or sequencing reads in a control dataset (i.e., the control dataset comprises tumor not-detected samples) corresponding to the same genomic region and having at least the same threshold number of methylated cytosines.

[0262] The normalized metrics can be analyzed with respect to a threshold value. The threshold value can correspond to a given genomic region, such as a given promoter region. In various examples, the threshold value can be different for different promoter regions. In these scenarios, a first promoter region can have a first threshold value and a second promoter region can have a second threshold value. In situations where the normalization metric is at least the threshold value, the computing system 214 can determine that the genomic region has a first methylation status. In scenarios where the normalization metric is less than the threshold value, the computing system 214 can determine that the genomic region has a second methylation status. In one or more illustrative examples, the first methylation status can be labeled as “methylated” and the second methylation status can be labeled as “not methylated.”

[0263] The threshold value for a given genomic region can be determined based on training data obtained from samples of individuals in which cancer is not detected. In one or more examples, sequence representations obtained from the training samples can be analyzed to determine a z-score with respect to the number of polynucleotide molecules that correspond to the genomic region and that have at least the threshold amount of methylated cytosines. In one or more illustrative examples, the threshold value for a promoter region that is used to determine the normalization metrics for the promoter region can be derived from the z-score calculated based on the training samples with respect to the promoter region.

[0264] Although the illustrative example of Figure 2 describes that models can be generated to determine a number of indicators with respect to the presence or absence of cancer in a given subject, in at least some additional examples, the sequencing data 212 can be analyzed by the computing system 214 to determine indicators of the presence of cancer without training specific models. In one or more examples, the computing system 214 can determine a tumor fraction value based on sequencing data 212 generated from one or more samples obtained from a single subject in which it is unknown whether or not cancer is present in the subject. In one or more examples, the computing system 214 can determine a change in the tumor fraction value based on sequencing data 212 generated from one or more samples obtained at two or more time points from a single subject. The change in the tumor fraction value can be used to monitor the subject’s response to treatment. In one or more additional examples, a first sample can be obtained from a subject prior to or at onset of at least one of administration of a treatment or a procedure related to cancer and one or more second samples can be obtained from the subjectAttorney Docket No.: GH0267WOafter at least one of administration of a treatment or a procedure related to cancer. In one or more illustrative examples, the one or more second samples can be obtained at least one week, at least two weeks, at least three weeks, at least four weeks, at least five weeks, at least six weeks, at least eight weeks, or at least ten weeks after administration of the treatment or procedure. In at least some examples, the first sample and the second sample can be derived from at least one of a bodily fluid obtained from the subject or tissue obtained from the subject.

[0265] In one or more examples, one or more samples can be obtained from a given subject. The sequencing data 212 generated from the one or more samples can be analyzed by the computing system to determine quantitative measures for a number of classification regions. In one or more examples, the quantitative measures can correspond to an amount of sequence representations that have at least a threshold amount of overlap with one or more classification regions. In one or more additional examples, the quantitative measures can correspond to sequence representations having at least a threshold amount of methylated cytosines in CpG regions having at least a threshold amount of CG content. In various examples, the indication of cancer being present in the subject can include tumor fraction. In one or more additional examples, the indication of cancer being present in the subject can include mutant allele fraction. In at least some examples, the quantitative measures can correspond to a number of sequencing reads that correspond to a given classification region in relation to a total number of sequencing reads across a plurality of positive control regions. In one or more further examples, the indicators of cancer being present can be used to determine an output that corresponds to cancer being present or not being present in a given individual in response to analyzing the one or more indicators of cancer being present with respect to one or more thresholds. In one or more illustrative examples, tumor fraction determined from one or more samples obtained from a subject can be analyzed with respect to one or more thresholds. In instances where tumor fraction is greater than a threshold level, the computing system 214 can determine that the probability of cancer being present in the subject is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. Further, in situations where multiple samples are obtained from a subject, first quantitative measures generated from a first sample obtained from the subject can be analyzed with respect to second quantitative measures generated from a second sample obtained from the subject. In at least some examples, differences between the first quantitative measures and the second quantitative measures can be analyzed to determine an indication of treatment response in the subject.

[0266] The quantitative measures used to determine an indication of cancer being present in a subject can be determined by analyzing quantitative measures of a subset of classificationAttorney Docket No.: GH0267WOregions. In at least some examples, the subset of classification regions can be different for different subjects. In one or more illustrative examples, values of quantitative measures for a number of classification regions can be analyzed with respect to one another and ranked according to the magnitude of the value of the quantitative measures. In various examples, the classification regions for a given sample can be ranked in descending order from the one or more classification regions having the greatest value of a quantitative measure to the one or more classification regions having the least value of the quantitative measure.

[0267] In various examples, after ranking the quantitative measures of the classification regions, quantitative measures that correspond to a group of the classification regions can be removed before determining the indication of cancer being present in the subject. For example, the group of classification regions that are not used to determine the indication of cancer being present in the subject can include the 1% of classification regions having the greatest quantitative measure values, the 2% of classification regions having the greatest quantitative measure values, 3% of classification regions having the greatest quantitative measure values, 4% of classification regions having the greatest quantitative measure values, 5% of classification regions having the greatest quantitative measure values, or the 6% of classification regions having the greatest quantitative measure values. In at least some examples, a number of classification regions having relatively high quantitative measure values can be excluded from the group of classification regions used to determine the indication of cancer being present in the subject because, in at least some cases, classification regions corresponding to quantitative measure values at or near the top of the ranked list can have non-tumor origins and / or be related to sequencing artifacts. Thus, by removing the quantitative measures that correspond to these classification regions from the analysis used to determine the indication of cancer being present in the subject, the accuracy with which the indication of cancer being present in the subject can increased.

[0268] In one or more examples, after determining the group of classification regions to be used to determine an indication of cancer being present in the subject, a subset of classification regions of the group can then be determined by identifying at least 10 classification regions of the group, at least 25 classification regions of the group, at least 50 classification regions of the group, at least 75 classification regions of the group, at least 100 classification regions of the group, at least 150 classification regions of the group, at least 200 classification regions of the group, at least 250 classification regions of the group, at least 300 classification regions of the group, at least 350 classification regions of the group, at least 400 classification regions of the group, at least 450 classification regions of the group, or at least 500 classification regions of the group having the greatest values for the respective quantitative measure.Attorney Docket No.: GH0267WO

[0269] In at least some examples, one or more statistical measures, such as at least one of mean, median, or mode, can be applied to the quantitative measures of the subset of the classification regions of the group to generate an initial indication of cancer being present in the subject. In various examples, the initial indication of cancer can be modified according to a scaling factor. The scaling factor can be applied to the initial indication of cancer being present in the subject because, in at least some scenarios, the positive control regions can have different amounts of methylated CpGs. For example, at least a portion of the positive control regions can have fully methylated CpGs while other positive control regions may not be fully methylated. Additionally, in various situations, some classification regions can correspond to a high value of an indication of cancer being present in subjects, such as 90% tumor fraction, 95% tumor fraction, 99% tumor fraction, or 100% tumor fraction, but nucleic acid molecules that correspond to these classification regions may not be fully methylated. To account for these cases, the scaling factor can be applied to the initial indication of cancer being present in the subject to provide a more accurate determination of the indication. In one or more illustrative examples, the scaling factor can be determined by analyzing indications of cancer being present in subjects determined using one or more techniques described herein in relation to additional data that corresponds to additional indications of cancer being present in subjects, such as validation data or other techniques that generate data orthogonal to the indications of tumors being present in subjects described herein.

[0270] In various examples, the classification regions used to determine the quantitative measures can correspond to classification regions that correspond to one or more portions of differentially methylated regions. In one or more examples, the differentially methylated regions can include promoter regions that correspond to one or more classifications of cancer. For example, the classification regions can be determined by analyzing a number of sequencing representations across a differentially methylated region. In these scenarios, one or more portions of the differentially methylated regions that overlap with at least a threshold number of sequencing representations can be included in the classification regions. In one or more examples, the quantitative measures of the one or more portions of the differentially methylated regions can be determined based on the molecule count distribution of the differentially methylated region. For example, the quantitative measures can be determined based on the molecule count within one or more peaks of the molecule distribution of the differentially methylated region. To illustrate, in various examples, the distribution of molecules across a differentially methylated region can indicate one or more peaks where greater amounts of molecules overlap with one or more subregions within the differentially methylated region. In various examples, the one or moreAttorney Docket No.: GH0267WOgenomic regions that correspond to the one or more subregions of the differentially methylated regions that correspond to the highest amounts of sequence representations for a sample can be defined as classification regions. In at least some examples, the distribution of sequence representations can have a peak that corresponds to a subregion of the differentially methylated region having a higher number of sequence representations than other subregions of the differentially methylated region. In these scenarios, the subregion can be identified as a classification region. By determining subregions of at least a portion of the differentially methylated regions used to determine the indication of cancer being present in the subject, the amount of computing resources and memory resources used to determine the indication of cancer being present in the subject can be decreased.

[0271] To illustrate, a classification region can include one or more portions of a differentially methylated region in which at least 50% of the sequencing representations obtained from a sample overlap, at least 55% of the sequencing representations obtained from a sample overlap, at least 60% of the sequencing representations obtained from a sample overlap, at least 65% of the sequencing representations obtained from a sample overlap, at least 70% of the sequencing representations obtained from a sample overlap, at least 75% of the sequencing representations obtained from a sample overlap, at least 80% of the sequencing representations obtained from a sample overlap, at least 85% of the sequencing representations obtained from a sample overlap, at least 90% of the sequencing representations obtained from a sample overlap, at least 95% of the sequencing representations obtained from a sample overlap, or at least 99% of the sequencing representations obtained from a sample overlap. In one or more illustrative examples, the one or more portions of the differentially methylated region that comprise a classification region can be contiguous with respect to a reference sequence.

[0272] Figure 3 is a diagrammatic representation of an example framework 300 to train a computational model 302 to determine one or more tumor metrics with respect to a sample, in accordance with one or more implementations. The framework 300 can include the computing system 214. The computing system 214 can execute the computational model 302 to generate one or more model outputs 304. In one or more examples, the computational model 302 can be a machine learning model. The model output 304 can include an indication corresponding to the presence or absence of a tumor in a subject that provided a sample. In one or more illustrative examples, the model output 304 can include a tumor fraction. In one or more additional illustrative examples, the model output 304 can include a probability of cancer being present in a subject. In one or more further illustrative examples, the model output can include an indication of cancer being present in a subject or an indication of cancer not being present in a subject. In still otherAttorney Docket No.: GH0267WOillustrative examples, the model output 304 can indicate methylation status of one or more regions of nucleic acid molecules. To illustrate, the computing system 214 can execute the computational model 302 with respect to quantitative measures corresponding to a promoter region to determine an amount of methylation of the promoter region. In other illustrative examples, the model output 304 can include a tumor tissue indication of the sample.

[0273] The framework 300 can also include a sequence representation 306. In one or more examples, the sequence representation 306 can be generated based on analyzing nucleic acid molecules that are derived from a sample provided by a subject. The sequence representation 306 can include genomic regions having a number of nucleotides that correspond to a number of regions of interest. For example, the sequence representation 306 can include a sequence of nucleotides that corresponds to a first classification region 308. In addition, the sequence representation 306 can include a sequence of nucleotides that corresponds to a second classification region 310. Further, the sequence representation 306 can include a sequence of nucleotides that corresponds to a third classification region 312. In various examples, the first classification region 308, the second classification region 310, and the third classification region 312 of the sequence representation 306 can have differing amounts of methylated cytosines included in the respective classification regions 308, 310, 312. In one or more additional examples, the sequence representation 306 can include a sequence of nucleotides that corresponds to a positive control region 314 and a sequence of nucleotides that corresponds to a negative control region 316.

[0274] The computational model 302 can include a number of components that correspond to individual classification regions. In one or more examples, the components of the computational model 302 can have respective values that correspond to quantitative metrics of the respective classification regions. The quantitative metrics can indicate a number of sequence representations that correspond to the respective classification regions. In one or more examples, the computational model 302 can include a number of weights that are related to the respective components of the computational model 302. For example, the computational model 302 can include a first model component 318 that has a first weight 320. The first model component 318 can correspond to the first classification region 308. In addition, the computational model 302 can include a second model component 322 that has a second weight 324. The second model component 322 can correspond to the second classification region 310. In various examples, at least one of the first weight 320, the second weight 324, or the third weight 328 can be different from at least another one of the first weight 320, the second weight 324, or the third weight 328.Attorney Docket No.: GH0267WO

[0275] In one or more illustrative examples, a value for the first model component 318, the second model component 322, and the third model component 326 can be determined on a per sample basis. To illustrate, for different samples, the computational model 302 can determine different values for at least one of the first model component 318, the second model component 322, or the third model component 326. In various examples, the computing system 214 can determine first quantitative measures for the first classification region 308 based on sequencing data for a sample. The computing system 214 can execute the computational model 302 to determine a value for the first model component 318 based on the first quantitative measures. Additionally, the computing system 214 can determine second quantitative measures for the second classification region 310 based on sequencing data for the sample. The computing system 214 can execute the computational model 302 to determine a value for the second model component 322 based on the second quantitative measures. Further, the computing system 214 can determine third quantitative measures for the third classification region 312 based on sequencing data for the sample. The computing system 214 can execute the computational model 302 to determine a value for the third model component 326 based on the third quantitative measures. The first quantitative measures, the second quantitative measures, and the third quantitative measures can be determined based on numbers of sequence representations that have at least a threshold amount of methylation in CG regions that correspond to the first classification region 308, the second classification region 310, and the third classification region 312, respectively. In one or more additional illustrative examples, a value for the first weight 320, a value for the second weight 324, and a value for the third weight 328 can be determined on a per sample basis. For example, for different samples, the computational model 302 can determine different values for at least one of the first weight 320, the second weight 324, or the third weight 328.

[0276] In one or more examples, the computing system 214 can perform a training process to generate the computational model 302. In various examples, the training process can determine one or more features related to classification region metrics that can be used to determine the model output 304. Additionally, the training process can determine one or more parameters related to classification region metrics that can be used to determine the model output 304. For example, the training process can be used to determine the model components to include in the computational model 302 and the corresponding weights of the model components.

[0277] In the illustrative example of Figure 3, the training process can be performed using training data 330. The training data 330 can include information obtained with respect to at least a first group of subjects 332 and information obtained with respect to at least a second group ofAttorney Docket No.: GH0267WOsubjects 334. In one or more examples, the first group of subjects 332 can include subjects in which a tumor is not detected and the second group of subjects 334 can include subjects in which a tumor is detected. In various examples, the training data 330 can include characteristics related to amounts of methylation of classification regions of the sequence representation 306 for the first group of subjects 332 and the second group of subjects 334. For example, the training data 330 can indicate quantitative measures corresponding to numbers of sequence representations that have at least a threshold level of methylation for the classification regions 308, 310, 312 for the first group of subjects 332 and the second group of subjects 334. The training data 330 can also include weights for model components based on an analysis of sequencing data of the first group of subjects 332 and the second group of subjects 334. In one or more illustrative examples, the training data 330 can include values for the first weight 320, values for the second weight 324, and values for the third weight 328 based on classification region metrics determined from sequencing data obtained from samples provided by the first group of subjects 332 and the second group of subjects 334.

[0278] The training data 330 can also include information corresponding to additional characteristics of the first group of subjects 332 and the second group of subjects 334. To illustrate, the training data 330 can include medical records information, medical history information, cancer treatment history information, demographic information, genomics information, one or more combinations thereof, and the like.

[0279] In one or more examples, the computing system 214 can train the computational model 302 to determine an indication related to one or more types of cancer being present in an individual. Additionally, in various examples, the computational model 302 can comprise multiple different models, such that the computational model 302 is an ensemble model. In these situations, the computing system 214 can perform one or more training processes with respect to individual models of the ensemble model. In one or more illustrative examples, the computational model 302 can include a number of individual models that each correspond to determining model outputs for individual genomic regions, such as genes or for a specified group of genes. For example, the computational model 302 can include a number of individual models to generate maximum MAF values for individual genes or for a specified group of genes. In these scenarios, the training of the computational model 302 can have more constraints than other models used to determine indications of cancer being present in individuals because of the use of genomic information in the training process. As a result, in situations where the computational model 302 is trained using genomic information, the accuracy of the output of the computational model 302 can be increased. Further, although the training of the computational model 302 can incorporateAttorney Docket No.: GH0267WOgenomic information to determine maximum MAF values, the use of the computational model 302 to determine indications of cancer for test subjects after training can be performed without the use of genomic information and can be based on input vectors that correspond to quantitative measures determined from sequencing data obtained from the test subjects.

[0280] The computing system 214 can obtain a first training dataset 336 up to an Nth training dataset 338 to perform a training process to generate the computational model 302. In one or more examples, the first training dataset 336 can include a first portion of the training data 330 corresponding to the first group of subjects 332 and the second group of subjects 334 that is used to train the computational model 302 and the Nth training dataset 338 can include a second portion of the training data 330 corresponding to the first group of subjects 332 and the second group of subjects 334 as part of a validation process for the computational model 302. In various examples, the computational model 302 can be updated over time and undergo multiple training processes. In these scenarios, the first training dataset 336 can include a portion of the training data 330 for the first group of subjects 332 and the second group of subjects 334 that corresponds to a first period of time and the Nth training dataset 338 can include a portion of the training data 330 for the first group of subjects 332 and the second group of subjects 334 that corresponds to a second period of time.

[0281] In one or more examples, during the training process for the computational model 302, the computing system 214 can perform one or more optimization operations. In one or more illustrative examples, the computing system 214 can identify, during the training process for the computational model 302, one or more samples obtained from at least one of the first group of subjects 332 or the second group of subjects 334 that are outliers with respect to samples obtained from other subjects included in at least one of the first group of subjects 332 or the second group of subjects 334. To illustrate, the computing system 214 can determine that model output 304 generated for one or more subjects included in at least one of the first group of subjects 332 or the second group of subjects 334 has at least a threshold amount of difference with the model output 304 generated for one or more additional subjects included in at least one of the first group of subjects 332 or the second group of subjects 334. In one or more examples, the computing system 214 can identify at least one of one or more first subjects 332 or one or more second subjects 334 having model output 304 that is at least one standard deviation, at least 1.5 standard deviations, at least 2 standard deviations, at least 2.5 standard deviations, or at least 3 standard deviations different from a mean model output 304 determined for an additional group of at least one of the first group of subjects 332 or the second group of subjects 334. In various examples, the computing system 214 can apply a penalty to information generated from samplesAttorney Docket No.: GH0267WOthat correspond to subjects that are outliers with respect to information generated from samples that correspond to additional subjects.

[0282] In one or more additional examples, one or more optimization processes implemented by the computing system 214 in the training of the computational model 302 can correspond to a number of training cycles and / or a number of iterations for individual training cycles that are performed during the training process.

[0283] In one or more illustrative examples, the computing system 214 can perform at least 1000 iterations of a training process to generate the computational model 302, at least 3000 iterations of a training process to generate the computational model 302, at least 5000 iterations of a training process to generate the computational model 302, at least 8000 iterations of a training process to generate the computational model 302, at least 10,000 iterations of a training process to generate the computational model 302, at least 12,000 iterations of a training process to generate the computational model 302, or at least 15,000 iterations of a training process to generate the computational model 302. In various examples, the computing system 214 can end the training process before convergence of a loss function related to the computational model 302. In one or more examples, the number of iterations of the training process to produce the computational model 302 can correspond to a number of iterations of the training process performed before the training process is stopped and before the convergence of the loss function.

[0284] In one or more examples, a first stage of the training process implemented by the computing system 214 to generate the computational model 302 can include determining samples included in the training data 330 that include somatic mutations indicative of one or more types of cancer in relation to samples included in the training data 330 that do not include somatic mutations indicative of the one or more types of cancer. The computing system 214 can then perform a training process for the computational model 302 using the samples of the training data 330 that include one or more somatic mutations indicative of the one or more types of cancer and using a number of samples obtained from subjects in which a tumor is not detected. In various examples, at least 100 iterations of the first stage of the training process can be performed.

[0285] Further, the training process performed by the computing system 214 can include a second stage that includes predicting values of tumor metrics of samples that do not include somatic mutations with respect to the one or more types of cancer. The computing system 214 can then perform at least 100 additional iterations of the second stage of the training process to generate the computational model 302. The second stage of the training process performed by the computing system 214 to generate the computational model 302 can also include training the computational model 302 using portions of the training data 330 corresponding to samples havingAttorney Docket No.: GH0267WOsomatic mutations indicative of the one or more types of cancer, using the predicted values of samples that do not include somatic mutations indicative of the one or more types of cancer, and portions of the training data 330 that correspond to samples obtained from subjects in which a tumor is not detected. In various examples, the second stage of the training process performed by the computing system 214 to generate the computational model 302 can be performed at least 2 additional times, at least 3 additional times, at least 4 additional times, at least 5 additional times, or at least 6 additional times. After the first stage of the training process and the second stage of the training process have been completed, the computing system 214 can perform a validation process for the computational model 302 using information obtained from different samples included in the training data 330.

[0286] In one or more illustrative examples, the computing system 214 can perform a training process for multiple computational models 302. In these scenarios, individual computational models 302 trained by the computing system 214 can correspond to different tissue types that are sources of genomic material obtained from subjects included in the training data 330. In one or more examples, the individual computational models 302 trained by the computing system 214 can correspond to different classifications of cancer, such as colorectal cancer, lung cancer, pancreatic cancer, bladder cancer, breast cancer, liver cancer, skin cancer, or one or more additional classifications of cancer. In situations where the computing system 214 trains multiple computational models 302 that correspond to different classifications of cancer, the output from individual computational models 302 can be aggregated and analyzed by the computing system 214 to determine a tissue of origin for a subject.

[0287] In various examples, the individual computational models 302 that correspond to a given tissue from which genomic material included in samples is derived can have different model components. For example, a first computational model generated by the computing system 214 that corresponds to a first tissue type can have first model components that correspond to a first set of classification regions. In addition, a second computational model generated by the computing system 214 that corresponds to a second tissue type can have second model components that correspond to a second set of classification regions that has at least one classification region different from the first set of classification regions. Additionally, the weights for the individual components of the computational models that correspond to different tissue types can be different. That is, in situations where the first set of classification regions of the first computational model and the second set of classification regions of the second computational model have at least one classification region in common, the weights for the model componentAttorney Docket No.: GH0267WOthat corresponds to the at least one common classification region can be different in relation to the first computational model and the second computational model.

[0288] Additionally, one or more additional normalization processes can be performed by the computing system when generating the computational model 302. For example, in at least some scenarios, molecules treated with MBD can be partitioned differently across different samples. In one or more examples, molecules can be partitioned differently across different samples due to differences in the composition of reagents used to treat the molecules with MBD. In one or more additional examples, molecules can be partitioned differently across different samples due to at least one of equipment differences or process conditions used to treat the molecules with MBD.

[0289] To illustrate, for one or more first samples, treatment with MBD can cause first molecules having regions with first CG content to be separated into a first partition and second molecules having regions with second CG content to be separated into a second partition. In addition, for one or more second samples, treatment with MBD can cause third molecules having third CG content that is different from the first CG content to be separated into the first partition and fourth molecules having regions with fourth CG content that is different from the second CG content to be separated into the second partition. In various examples, the first molecules can be treated with MBD and separated into the first partition and the second molecules can be treated with MBD and separated into the second partition across a first cutoff range of CG content. Further, the third molecules can be treated with MBD and separated into the first partition and the fourth molecules can be treated with MBD and separated into the second partition across a second cutoff range of CG content that is different from the first cutoff range.

[0290] In one or more illustrative examples, the first cutoff range of CG content can include from 3-10 CpGs having methylated cytosines and the second cutoff range can include from 6-14 CpGs having methylated cytosines. In one or more additional illustrative examples, the first cutoff range of CG content can include from 4-9 CpGs having methylated cytosines and the second cutoff range can include from 7-13 CpGs having methylated cytosines. In one or more further illustrative examples, the first cutoff range of CG content can include from 5-8 CpGs having methylated cytosines and the second cutoff range can include from 8-12 CpGs. In still other illustrative examples, the first cutoff range of CG content can include from 4-7 CpGs and the second cutoff range can include from 6-10 CpGs. In various examples, the first cutoff range of CG content and the second cutoff range of CG content can be used to determine the threshold amount of methylated cytosines used to determine at least one of training sequencing reads or testing sequencing reads. In at least some examples, the threshold amount of methylatedAttorney Docket No.: GH0267WOcytosines can include a cutoff number that corresponds to a probability, such as at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% of individual molecules treated with MBD being separated into a given partition. In one or more examples, the threshold amount of methylated cytosines can correspond to 5 methylated cytosines, 6 methylated cytosines, 7 methylated cytosines, 8 methylated cytosines, 9 methylated cytosines, 10 methylated cytosines, 11 methylated cytosines, 12 methylated cytosines, 13 methylated cytosines, or 14 methylated cytosines.

[0291] In one or more examples, the computing system 214 can generate metrics for individual classification regions based on quantitative measures that are determined by analyzing a first number of sequencing reads to identify a first number of nucleic acid molecules having a first amount of CG content and by analyzing a second number of sequencing reads to identify a second number of nucleic acid molecules having a second amount of CG content. In at least some examples, the second number of nucleic acid molecules can be used to modify a metric determined using the first number of nucleic acid molecules to account for variations in the separation of molecules treated using MBD for different samples. In various examples, for individual classification regions, a first metric can be determined for a given sample by determining a first quantitative measure that corresponds to a number of molecules having a threshold amount of methylated cytosines and having a first amount of cytosine-guanine content in one or more partitions (for example, second partition 130 and / or third partition 134) that correspond to the individual classification region. In some embodiments, the first amount of CG content can be at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, or at least 30 CpGs in the nucleic acid molecules. In some embodiments, the first amount of CG content can be between 5-10, 5-15, 5-20, 5-30, 10-15, 10-20, 10-30, 15-20, 15-30, or 20-40 CpGs in the nucleic acid molecules.

[0292] The first metric can also be determined for a given sample by determining a second quantitative measure that corresponds to a number of molecules having a threshold amount of methylated cytosines and having the first amount of cytosine-guanine content in one or more partitions (for example, second partition 130 and / or third partition 134) that correspond to a plurality of control regions (e.g., positive control regions). To illustrate, for an individual classification region, the first metric can be determined using the first quantitative measure for the individual classification region and the second quantitative measure that corresponds to the plurality of control regions.

[0293] The normalization process can also include determining, for a given sample, a second metric for the given sample by determining one or more additional quantitative measuresAttorney Docket No.: GH0267WObased on a number of molecules in one or more partitions (e.g., second partition 130 and / or third partition 134) having at least the threshold amount of methylated cytosines and a second amount of cytosine-guanine content that correspond to the plurality of control regions, where the second amount of cytosine-guanine content is less than the first amount of cytosine-guanine content. In some embodiments, the second amount of CG content can be between 5-10, 5-15, 10-15, 10-20, or 15-20 CpGs in the nucleic acid molecules. In one or more examples, the plurality of control regions can be positive control regions and / or negative control regions. In one or more examples, the second metric can be determined using the additional quantitative measure and the second quantitative measure. In at least some examples, the second metric can be determined fora given sample by determining a ratio of the one or more additional quantitative measures with respect to the second quantitative measure. In one or more additional examples, the second metric can be determined for a given sample by determining the logarithm, such as the logarithm according to base 10, of a ratio of the one or more additional quantitative measures with respect to the second quantitative measure.

[0294] In one or more illustrative examples, the second metric for a given sample can include a combination of values, where individual values correspond to an additional quantitative measure based on a number of molecules having at least a threshold amount of methylated cytosines and a given number of CpGs for the plurality of control regions and the second quantitative measure. For example, a first additional quantitative measure can be determined based on a first number of molecules having at least the threshold amount of methylated cytosines in control regions having a first number of CpGs, such as 6, and a second additional quantitative measure can be determined based on a second number of molecules having at least the threshold amount of methylated cytosines in control regions having a second number of CpGs, such as 7. In at least some examples, more additional quantitative measures can be determined based on additional numbers of molecules having the threshold amount of methylated cytosines in control regions having additional numbers of CpGs, such as 8 CpGs, 9 CpGs, 10 CpGs, and the like up to an upper threshold of CpGs, such as 12 CpGs, 13 CpGs, or 14 CpGs. Ratios determined using the additional quantitative measures with respect to the second quantitative measures can be determined and summed to determine the second metric.

[0295] In various examples, a correlation factor can also be determined for individual classification regions in relation to different amounts of CpGs that can be used to determine the second metric. In one or more examples, the correlation factor can modify the individual additional quantitative measures and then the modified individual additional quantitative measures can be aggregated to determine the second metric. In one or more additional examples, the first metricAttorney Docket No.: GH0267WOand the second metric can be combined to determine a normalized metric that corresponds to a given classification region. In one or more illustrative examples, the second metric can be subtracted from the first metric to determine the normalized metric.

[0296] In one or more additional illustrative examples, the correlation factor for a given classification region can be determined for each of a plurality of different amounts of cytosine-guanine content, such as a first correlation factor for 6 CpGs, a second correlation factor for 7 CpGs, a third correlation factor for 8 CpGs, and so forth up to a threshold amount of CG content. In at least some examples, the correlation factor can be determined by analyzing training data using one or more linear regression techniques. For example, the training data 330 can be fit to a linear regression model for individual classification regions to determine the correlation factor. In various examples, the fitting of at least a portion of the training data 330 to the linear regression model can be performed by aggregating the additional quantitative measures for a given classification region across a range of CG content, such as 6 CpGs, 7 CpGs, up to a threshold number of CpGs, and determining a mean quantitative measure.

[0297] In one or more examples, the normalized metrics can reduce variation of quantitative measures determined for individual samples. In at least some examples, the reduction in variation can result in increased accuracy of model outputs 304 in relation to at least some model outputs 304 determined without implementing the additional normalization process to determine the normalized metric.

[0298] Figure 4 is a flowchart of an example method 400 to determine tumor metrics in a subject based on levels of methylation of classification regions, according to one or more implementations. At operation 402, the method 400 can include obtaining training sequence data including training sequencing reads derived from a plurality of samples of a plurality of subjects. Individual training sequencing reads can include a nucleotide sequence corresponding to a fragment of a nucleic acid included in a sample of the plurality of samples. Individual training sequencing reads can have a threshold amount of molecules with a methylated cytosine included in regions of the nucleotide sequence having at least a threshold cytosine-guanine content. In one or more illustrative examples, the plurality of samples can include cell-free nucleic acids. In one or more examples, methylated cytosines can be determined using at least one of sodium bisulfite conversion and sequencing, Tet-assisted bisulfite sequencing (TAB-Seq), differential enzymatic cleavage, treatment with MSRE and / or MDRE, or MBD partitioning. In one or more additional examples, methylated cytosines can be determined using one or more single molecule sequencing methods, such as nanopore DNA sequencing or those described in Eid, J., et al.Attorney Docket No.: GH0267WO(2009) Real-time DNA sequencing from single polymerase molecules. Science, 323(5910), 133–138.

[0299] In one or more examples, the training process can include obtaining, by the computing system, testing sequence data from an additional subject that is not included in the plurality of subjects. The testing sequence data can include testing sequencing reads derived from a sample of the additional subject. Individual testing sequencing reads can include a nucleotide sequence corresponding to a fragment of a nucleic acid included in the additional sample. Additionally, individual testing sequencing reads can have at least the threshold amount of molecules with a methylated cytosine included in regions of the nucleotide sequence having at least the threshold cytosine-guanine content. Based on the additional sequence data, a model can be executed to determine the indication of cancer being present in the additional subject. The testing sequencing reads can then be analyzed to determine a first quantitative measure derived from the testing sequencing reads that correspond to the individual classification regions of the plurality of classification regions. Further, the testing sequencing reads can be analyzed to determine a second quantitative measure derived from the testing sequencing reads that correspond to the individual control regions of the plurality of control regions. The metric can then be determined for the individual classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions. Subsequently, an input vector can be generated that includes the metrics for the individual classification regions. The model can use the input vector to determine the indication of cancer being present in the additional subject.

[0300] In situations where the model is trained to determine an estimate of tumor fraction, the training sequencing reads can comprise a first portion of the training sequence data and a second portion of the training sequence data includes additional training sequencing reads that are different from the training sequencing reads. In these scenarios, at least one of the first portion of the training sequence data or the second portion of the training sequence data can be analyzed to determine an individual frequency of a plurality of variants present in individual samples of the plurality of samples. With respect to individual samples, a variant of the plurality of variants having a maximum frequency can then be determined that corresponds to the individual frequency having a greatest value among individual frequencies derived from an individual sample. In one or more illustrative examples, the maximum mutant allele frequency can be determined for individual samples. In various examples, individual measures of tumor fraction for the individual samples can then be determined based on the greatest value of the individual frequencies derived from the individual sample.Attorney Docket No.: GH0267WO

[0301] In at least some examples, the training process for the model can include one or more optimization operations. For example, the training process can include determining one or more additional weights of individual samples included in the training data based on the indication of cancer for the individual samples being within a threshold confidence level. In response to determining that the indication of cancer for an individual sample is outside of the threshold confidence level, a penalty can be applied to the individual sample during the training process.

[0302] The one or more training optimization operations can also include performing, using the one or more machine learning algorithms, one or more first iterations of the training process for the model using a portion of the training data. In addition, first output data for the model can be generated based on the one or more first iterations of the training process. The first output data can correspond to one or more first additional indications of cancer being present in first individual subjects of the plurality of subjects and the first individual subjects can correspond to the portion of the training data. Further, the training process can include combining the first output data and the training data to produce additional training data and performing one or more second iterations of the training process for the model using a portion of the additional training data. Second output data can then be generated for the model based on the one or more second iterations of the training process. The second output data can indicate one or more second additional indications of cancer being present in second individual subjects of the plurality of subjects where the second individual subjects correspond to the portion of the additional training data. In one or more illustrative examples, the weights for the individual classification regions of the plurality of classification regions can be determined based on the first output data and the second output data.

[0303] Further, the training process can include determining that a number of indications of cancer are present that were determined during one or more iterations of the training process and have at least a threshold value for one or more samples included in the training data. In these scenarios, modifications to one or more weights of the model are not modified or are modified by a minimal amount. Additionally, an additional number of indications of cancer being present can be determined that were determined during the one or more iterations of the training process and are less than the threshold value for one or more additional samples included in the training data. In these scenarios, modifications to one or more additional weights of the model can be determined and the one or more additional weights are modified by more than the minimal amount.

[0304] In addition, at operation 404, the process 400 can include analyzing the training sequencing reads to determine a first quantitative measure derived from the training sequencingAttorney Docket No.: GH0267WOreads that corresponds to individual classification regions of a plurality of classification regions. In one or more examples, the first quantitative measure can be determined based on the number of training sequencing reads. In one or more additional examples, the first quantitative measure can be determined based on a number of polynucleotide molecules that correspond to the training sequencing reads. At least a portion of the individual classification regions of the plurality of classification regions correspond to genomic regions of a reference genome that have the threshold amount of molecules with a methylated cytosine in subjects in which cancer is detected and that have at least the threshold cytosine-guanine content. In various examples, the plurality of classification regions can correspond to genomic regions in which at least one mutation occurs in patients in which cancer is detected. Additionally, the plurality of classification regions can correspond to a first plurality of classification regions for a first cancer type and the model can be generated for a second cancer type based on a second plurality of classification regions that are different from the first plurality of classification regions.

[0305] At operation 406, the process 400 can include analyzing the training sequencing reads to determine a second quantitative measure derived from the training sequencing reads that correspond to a plurality of control regions. In one or more examples, the second quantitative measure can be determined based on the number of training sequencing reads. In one or more additional examples, the second quantitative measure can be determined based on a number of polynucleotide molecules that correspond to the training sequencing reads. Individual control regions of the plurality of control regions correspond to additional genomic regions of the reference genome that have at least the threshold cytosine-guanine content. Additionally, the individual control regions can have at least the threshold amount of molecules with a methylated cytosine in subjects in which cancer is detected and in additional subjects in which cancer is not detected.

[0306] Further, at operation 408, the process 400 can include determining a metric for the individual classification regions of the plurality of classification regions based on the first quantitative measure for the individual classification regions and the second quantitative measure for the plurality of control regions. In one or more examples, the metric for the individual classification regions is determined based on a scaling factor and an error correction factor. In one or more illustrative examples, the scaling factor can include a logarithmic function and the error correction factor can include a pseudocount.

[0307] At operation 410, the process 400 can include generating, by the computing device, training data that includes the metric for the individual classification regions of the plurality of classification regions for the training sequence reads. In implementations where the indicationAttorney Docket No.: GH0267WOof cancer is tumor fraction, the training data can include the individual measures of tumor fraction for the individual samples of the plurality of samples and the model can be executed with respect to individual measures of tumor fraction for the individual samples of the plurality of samples.

[0308] The process 400 can also include, at operation 412, implementing, using the training data, one or more machine learning algorithms to generate a model to determine an indication of cancer being present in subjects based on amounts of molecules with methylated cytosines in at least a portion of the plurality of classification regions. The model can determine weights for individual classification regions of the plurality of classification regions and at least a portion of the weights of the individual classification regions can be different from one another. In various examples, the one or more machine learning algorithms can include one or more classification algorithms and the indication of cancer being present corresponds to a probability of cancer being present in the additional subject. In one or more additional examples, the one or more machine learning algorithms include one or more regression algorithms and the indicator corresponds to an estimate of tumor fraction of the additional sample. In one or more illustrative examples, a limit of detection for the model to determine tumor fraction of samples can be no greater than 0.01% given 95% sensitivity, no greater than 0.05% given 95% sensitivity, no greater than 0.1% given 95% sensitivity, no greater than 0.15% given 95% sensitivity, no greater than 0.2% given 95% sensitivity, no greater than 0.25% given 95% sensitivity, or no greater than 0.3% given 95% sensitivity.

[0309] In various examples, the sequence reads provided to the model during the training process or after the training process have at least a threshold amount of methylated cytosines in classification regions. The sequence reads that satisfy the methylation levels can be produced, at least in part, using one or more molecule separation processes. The molecule separation processes can include combining a plurality of nucleic acids derived from at least one of blood or tissue of a subject with a solution including an amount of methyl binding domain (MBD) proteins to produce a nucleic acid-MBD protein solution. A plurality of washes can then be performed on the nucleic acid-MBD protein solution with a salt solution to produce a number of nucleic acid fractions. Individual nucleic acid fractions can have a threshold number of molecules with a methylated cytosine in regions of the plurality of nucleic acids having at least the threshold cytosine-guanine content. In one or more illustrative examples, a wash of the plurality of washes can be performed with a solution having a concentration of sodium chloride (NaCI) and can produce a nucleic acid fraction of the number of nucleic acid fractions having a range of binding strengths to MBD proteins.Attorney Docket No.: GH0267WO

[0310] In one or more examples, a first nucleic acid fraction can be determined to be associated with a first partition of a plurality of partitions of nucleic acids. The first partition corresponds to a first range of binding strengths to MBD proteins. Further, a first molecular barcode can be attached to nucleic acids of the first nucleic acid fraction. The first molecular barcode can be associated with the first partition. In addition, a second nucleic acid fraction can be determined that is associated with a second partition of the plurality of partitions of nucleic acids. The second partition can correspond to a second range of binding strengths to MBD proteins different from the first range of binding strengths to MBD proteins. A second molecular barcode can be attached to nucleic acids of the second nucleic acid fraction. The second molecular barcode is associated with the second partition.

[0311] In one or more additional examples, at least a portion of the number of nucleic acid fractions can be combined with an amount of restriction enzyme that cleaves molecules with one or more unmethylated cytosines to produce at least a portion of the plurality of samples used to produce the sequencing reads. In these scenarios, the threshold amount of molecules with a methylated cytosine corresponds to a minimum frequency of molecules with a methylated cytosine within a region having at least the threshold cytosine-guanine content. In one or more further examples, at least a portion of the number of nucleic acid fractions are combined with an amount of a restriction enzyme that cleaves molecules with a methylated cytosine to produce at least a portion of the plurality of samples used to produce the sequencing reads. In these situations, the threshold amount of molecules with a methylated cytosine corresponds to a maximum frequency of molecules with a methylated cytosine within a region having at least the threshold cytosine-guanine content.

[0312] Figure 5 is a flow diagram of an example process 500 to determine an amount of a cell type present in a sample, according to one or more example implementations. The process 500 can include, at 502, obtaining a polynucleotide sample derived from a subject. The sample can include a bodily sample. For example, the sample can include body tissue. In one or more examples, the sample can include tissue obtained by a biopsy procedure. In one or more additional examples, the sample can include tissue obtained by a surgical procedure that takes place as part of a treatment for a biological condition. In one or more further examples, the sample can include and / or be derived from one or more bodily fluids. To illustrate, the sample can include whole blood, platelets, serum, plasma, buffy coat, stool, red blood cells, white blood cells, endothelial cells, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, or urine. AAttorney Docket No.: GH0267WOsample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components. In one or more illustrative examples, the sample can include cell-free nucleic acids. In one or more additional illustrative examples, the subject can include a human subject. In still other examples, the subject can include a non-human mammal subject.

[0313] At 504, the process 500 can include determining a number of nucleic acids derived from the sample that are aligned with individual regions of interest of a plurality of regions of interest. In one or more examples, nucleic acids derived from the sample can be analyzed with respect to one or more criteria to determine the number of nucleic acids that are aligned with the individual regions of interest. For example, a plurality of nucleic acids can be derived from the sample and analyzed with respect to a number of cytosine-guanine dinucleotides (CpGs) present in the individual nucleic acids. In one or more illustrative examples, nucleic acids having greater than a threshold number of CpGs can be identified. In at least some examples, the threshold number of CpGs can include at least 2 CpGs, at least 3 CpGs, at least 4 CpGs, at least 5 CpGs, at least 6 CpGs, at least 7 CpGs, at least 8 CpGs, at least 9 CpGs, or at least 10 CpGs. In various examples, a subset of the plurality of nucleic acids derived from the sample that correspond to the threshold number of CpGs can comprise at least a portion of the number of nucleic acids that are aligned with individual regions of interest.

[0314] Additionally, a plurality of nucleic acids derived from the sample can be analyzed with respect to methylation states of CpGs included in the plurality of nucleic acids. For example, the plurality of nucleic acids can be analyzed with respect to a number of methylated CpGs present in the plurality of nucleic acids. Further, the plurality of nucleic acids can be analyzed with respect to a number of unmethylated CpGs present in the plurality of nucleic acids. In one or more illustrative examples, the plurality of nucleic acids can be analyzed with respect to a threshold number of methylated CpGs or a threshold number of unmethylated CpGs present in the plurality of nucleic acids. To illustrate, the plurality of nucleic acids can be analyzed to determine a subset of the plurality of nucleic acids that have at least a threshold number of unmethylated CpGs. In one or more additional illustrative examples, the plurality of nucleic acids can be analyzed to determine a subset of the plurality of nucleic acids having no greater than a threshold number of methylated CpGs. In various examples, the threshold number of methylated CpGs or unmethylated CpGs can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0315] In at least some examples, the methylation states of the CpGs present in the plurality of nucleic acids can be determined by performing a single site methylation state identification process with respect to the plurality of samples. The single site methylation stateAttorney Docket No.: GH0267WOidentification process can identify unmethylated CpGs present in the plurality of nucleic acids. In addition, the single site methylation state identification process can identify methylated CpGs present in the plurality of nucleic acids. The single site methylation state identification process can include one or more bisulfite conversion processes. In one or more illustrative examples, the single site methylation state identification process can include a Tet-assisted bisulfite conversion process.

[0316] The regions of interest can include genomic regions of a reference genome that can be indicative of the presence of one or more cell types in samples. In one or more examples, the regions of interest can comprise one or more genomic regions that are differentially methylated in samples that include one or more cell types. For example, the regions of interest can comprise one or more genomic regions that include a greater number of unmethylated CpGs than additional genomic regions that do not include the one or more cell types. In addition, the regions of interest can comprise one or more genomic regions that include a greater number of methylated CpGs than additional genomic regions that do not include the one or more cell types. In at least some examples, the genomic regions can be included in one or more assays. In one or more illustrative examples, the genomic regions can be enriched in a diagnostic test for one or more biological conditions. To illustrate, the genomic regions can be enriched in a diagnostic test for one or more types of cancer.

[0317] In addition, the process 500 can include, at 506, determining, for individual regions of interest, quantitative measures for individual cell types. The quantitative measures can be determined based on a subset of the number of nucleic acids that are aligned with an individual region of interest. For example, the quantitative measure for an individual region of interest can include a count of a subset of the nucleic acids derived from the sample that are aligned with the region of interest. In at least some examples, an individual count that indicates a respective subset of the nucleic acids can be determined for the individual regions of interest. In this way, a first count corresponding to a first subset of the nucleic acids derived from the sample can be determined that aligns with a first region of interest and a second count corresponding to a second subset of the nucleic acids derived from the sample can be determined that aligns with a second region of interest. In various examples, the first subset of nucleic acids can be distinct from the second subset of nucleic acids. That is, in one or more illustrative examples, a nucleic acid present in the first subset of nucleic acids may not be present in the second subset of nucleic acids. In one or more illustrative examples, individual quantitative measures can be determined for at least 10 regions of interest, at least 25 regions of interest, at least 50 regions of interest, at least 100 regions of interest, at least 250 regions of interest, at least 500 regions of interest, at least 1000Attorney Docket No.: GH0267WOregions of interest, at least 2500 regions of interest, at least 5000 regions of interest, or at least 10,000 regions of interest.

[0318] Further, at 508, the process 500 can include determining, using one or more computational models and based on the quantitative measures, an amount of the individual cell types present in the sample. In one or more examples, output from the one or more computational models can indicate a percentage of the number of nucleic acids corresponding to a plurality of different cell types. To illustrate, an output from the one or more computational models can indicate that a first percentage of the nucleic acids derived from the sample correspond to a first cell type and a second percentage of the nucleic acids derived from the sample correspond to a second cell type. In various examples, the output of the one or more computational models can indicate amounts of at least 2 cell types present in the sample, at least 3 cell types present in the sample, at least 4 cell types present in the sample, at least 5 cell types present in the sample, at least 6 cell types present in the sample, at least 7 cell types present in the sample, at least 8 cell types present in the sample, at least 9 cell types present in the sample, at least 10 cell types present in the sample, at least 20 cell types present in the sample, at least 30 cell types present in the sample, at least 40 cell types present in the sample, at least 50 cell types present in the sample, at least 75 cell types present in the sample, or at least 100 cell types present in the sample. In one or more illustrative examples, the one or more computational models can include a least squares model. In one or more additional illustrative examples, the one or more computational models can include a partial least squares model with a non-negative value constraint.

[0319] In one or more examples, the amounts of the individual cell types present in the sample can be used to identify one or more biological conditions present in the subject. In various examples, training samples used in the training of the one or more computational models can be derived from training subjects in which the one or more biological conditions are not present. In one or more additional examples, the amounts of the individual cell types present in the sample can be used to determine one or more treatments to provide to the subject in relation to one or more biological conditions present in the subject. In one or more further examples, the amounts of the individual cell types present in the sample can be used to determine an effectiveness of a treatment provided to the subject in relation to one or more biological conditions. In one or more illustrative examples, one or more classification models can be implemented to analyze the amounts of individual cell types present in the sample to determine one or more biological conditions present in the subject and / or to determine information related to one or more treatments administered to the subject in relation to one or more biological conditions.Attorney Docket No.: GH0267WO

[0320] In at least some examples, the amounts of one or more cell types present in a sample can indicate the impacts of one or more treatments administered to subjects during treatment for one or more biological conditions. In various examples, the amounts of one or more cell types present in a sample can indicate an amount of toxicity with respect to subjects that is produced in response to one or more treatments administered to the subjects. In one or more illustrative examples, an amount of liver cells detected in a sample can indicate a level of toxicity of one or more treatments administered to subjects to treat one or more biological conditions present in the subjects.

[0321] The one or more biological conditions can include one or more types of cancer, one or more types of autoimmune diseases, one or more types of lung disease, one or more types of liver disease, one or more types of cardiovascular disease, one or more types of neurological disorders, or one or more combinations thereof. In one or more illustrative examples, the one or more biological conditions can include Barrett's esophagus, schizophrenia, bipolar disease, dementia, cytokine release syndrome, immune effector cell-associated neurotoxicity syndrome (ICANS), rheumatoid arthritis, endometriosis, metabolic dysfunction-associated steatohepatitis (MASH), liver fibrosis, or one or more combinations thereof.

[0322] Additionally, the cell types identified by the one or more computational models can include at least one of Total lymphocytes, Total granulocytes, Total T cells, Activated T cells, Naive T cells, CD4 T cells, CD8 T cells, Activated CD8 T cells, Effector memory CD8 T cells, Central memory CD8 T cells, Effector CD8 T cells, Naive CD8 T cells, Activated CD4 T cells, Effector memory CD4 T cells, Central memory CD4 T cells, Effector CD4 T cells, Naive CD4 T cells, Regulatory T cells, Total monocytes and macrophages, Total monocytes and macrophages and dendritic cells, Monocytes, Total macrophages, M1 macrophages, M2 macrophages, Total B cells, Plasma B cells, Memory B cells, Naive B cells, Progenitor B cells, Naive or progenitor B cells, Activated B cells, NK cells, Megakaryocytes, Erythroid progenitors, Erythroblasts, Neutrophils, Eosinophils, Adipocytes, Bladder epithelial cells, Breast epithelial cells, Breast basal epithelial cells, Breast luminal epithelial cells, Colon epithelial cells, Endothelial cells, Fibroblast cells, Muscle-derived fibroblast cells, Heart-derived fibroblast cells, Gastric epithelial cells, Head and neck epithelial cells, Heart cardiomyocytes, Kidney epithelial cells, Liver hepatocytes, Lung epithelial cells, Lung alveolar epithelial cells, Lung bronchial epithelial cells, Neuron or oligodendrocyte cells, Ovary epithelial cells, Pancreas cells, Pancreas acinar cells, Pancreas ductal cells, Prostate epithelial cells, Small intestine epithelial cells, Thyroid epithelial cells, Exhausted lymphocytes, Exhausted T cells, Exhausted CD8 T cells, Exhausted CD4 T cells, or Exhausted B cells. In still other examples, the cell types identified by the one or moreAttorney Docket No.: GH0267WOcomputational models can include immune cells of any of the above populations, from cancer tumors, when compared with immune cells from healthy patients, immune cells of any of the above populations, from blood of cancer patients, when compared with immune cells from the blood of healthy patients, immune cells from diseased tissue, when compared with immune cells from healthy tissue, macrophages from diseased tissue, when compared with macrophages from non-diseased tissue, or one or more combinations thereof.

[0323] The one or more computational models can be generated by performing a training process using data derived from a number of training samples. For example, the one or more computational models can be generated by obtaining a plurality of training samples from a plurality of subjects. In one or more examples, the plurality of subjects can be free of one or more biological conditions. In at least some examples, the plurality of subjects can be free of one or more biological conditions that can be detected based on the amounts of the cell types identified in test subjects by the one or more computational models. In various examples, the plurality of subjects can be considered healthy in accordance with one or more criteria.

[0324] Additionally, the plurality of training samples can include samples that correspond to individual cell types. To illustrate, the plurality of training samples can include first training samples that correspond to B-cells, second training samples that correspond to T-cells, third training samples that correspond to neutrophils, and fourth training samples that correspond to NK cells. In at least some examples, the plurality of cell types can have at least a threshold level of purity. For example, individual groups of training samples corresponding to an individual cell type can be comprised of at least 80% cells of the individual cell type, at least 85% cells of the individual cell type, at least 90% of cells of the individual cell type, or at least 95% of cells of the individual cell type. In one or more illustrative examples, the plurality of samples can be purified with respect to one or more cell types.

[0325] The training process for the one or more computational models can also include determining an additional number of nucleic acids derived from the plurality of samples that are aligned with the individual regions of interest. In one or more illustrative examples, the additional number of nucleic acids can be aligned with the individual regions of interest and have one or more methylation characteristics. For example, the additional number of nucleic acids can have at least a threshold amount of CpGs. In one or more additional illustrative examples, the additional number of nucleic acids can have an additional threshold amount of methylated CpGs or unmethylated CpGs.

[0326] Further, additional quantitative measures can be determined for the individual regions of interest based on the additional number of nucleic acids that are aligned with theAttorney Docket No.: GH0267WOindividual regions of interest and that correspond to individual cell types. For example, the additional quantitative measures can indicate a first count of a first subset of the additional number of nucleic acids that correspond to a first cell type and are aligned with a first region of interest, a second count of a second subset of the additional number of nucleic acids that correspond to a second cell type and are aligned with the first region of interest, a third count of a third subset of the additional number of nucleic acids that correspond to the first cell type and are aligned with a second region of interest, and a fourth count of a fourth subset of the additional number of nucleic acids that correspond to a second cell type and are aligned with the second region of interest. In this way, a reference matrix can be generated, based on the additional number of nucleic acids, that indicates an amount of the additional nucleic acids that corresponds to the individual cell types and the individual regions of interest. In at least some examples, the reference matrix can be used in the training process for the one or more computational models. In various examples, the reference matrix can be used to determine individual weights to be assigned to features of the one or more computational models that correspond to individual regions of interest. Thus, some regions of interest can be weighted more heavily in determining an amount of a first cell type present in a sample than other regions of interest that are weighted more heavily in determining an amount of a second cell type present in a sample.

[0327] In one or more examples, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from training subjects in which one or more biological conditions are not present. For example, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from subjects in which a tumor has not been detected. In one or more additional examples, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from training subjects in which one or more biological conditions are present. In various examples, one or more cell types being identified in samples can be related to the one or more biological conditions present in subjects. To illustrate, amounts of one or more cell types related to treatment toxicity and / or biological reactions to treatments of a biological condition can be determined based on training samples obtained from training subjects in which the biological condition is present.

[0328] In one or more illustrative examples, training subjects in which an autoimmune related biological condition has been detected may be receiving treatment for the autoimmune related biological condition. In these scenarios, one or more computational models can be trained using training data derived from the training subjects to determine the amounts of one or moreAttorney Docket No.: GH0267WOcell types present in the training subjects. The amounts of the one or more cell types present in the training subjects can be used to determine an effectiveness of the treatment for the autoimmune related biological condition and / or to determine an amount of toxicity produced within the training subjects in response to the treatment. In one or more additional illustrative examples, training subjects in which one or more cancer types have been detected may be receiving treatment for the one or more cancer types. In these instances, one or more computational models can be trained using training data derived from the training subjects to detect the amounts of one or more cell types present in the training subjects. The amounts of the one or more cell types present in the training subjects can be used to determine an effectiveness of the treatment for the one or more types of cancer and / or to determine an amount of toxicity produced within the training subjects in response to the treatment.

[0329] In situations where one or more computational models are trained to determine an amount of one or more cell types present in subjects in which one or more biological conditions are present, the regions of interest for which the quantitative measures are determined can have some overlap such that one or more first regions of interest that are indicative of the one or more cell types can overlap with at least a portion of one or more second regions of interest that are indicative of the one or more biological conditions being present in subjects. As a result, in scenarios where a computational model is being trained to determine amounts of one or more cell types present in subjects, an amount of signal being detected in relation to the one or more biological conditions being present in subjects can be taken into account when determining the amounts of the one or more cell types. That is, in instances where the regions of interest indicative of the one or more cell types being present in samples and the regions of interest indicative of the one or more biological conditions being present in samples overlap, at least in part, the quantitative measures determined for one or more regions of interest corresponding to the presence of the one or more cell types in samples can be impacted by the amount of nucleic acids identified in samples for which the one or more biological conditions are present. In at least some examples, the nucleic acid molecules derived from samples in which the one or more biological conditions are present that are aligned with regions of interest corresponding to the one or more cell types can result in a noise signal that can impact at least one of the precision or the accuracy of the determination of the amounts of the one or more cell types present in the subjects. In these scenarios, the noise signal can be reduced by identifying regions of interest that correspond to the one or more cell types being present in subjects that do not overlap with regions of interest that are indicative of the one or more biological conditions being present in subjects. In still other examples, the noise signal can be reduced by generating a corrected signal that accounts for theAttorney Docket No.: GH0267WOnoise signal produced by the nucleic acids present in samples in which the one or more biological conditions are present. By training computational models that determine amounts of one or more cell types present in samples to minimize or reduce the noise signal generated due to one or more biological conditions being present in subjects, the implementations of the computational models described herein can more efficiently and accurately determine the amounts of cell types present in samples.

[0330] In one or more illustrative examples, one or more training processes can be performed to generate one or more computational models that can determine the amounts of one or more cell types present in samples. The one or more training processes can be performed using training data derived from training subjects in which one or more biological conditions are present. In one or more examples, the training data can be derived from subjects in which the one or more biological conditions impact the methylation characteristics of nucleic acids derived from a subject in which the one or more biological conditions are present. In at least some examples, subjects in which one or more biological conditions are present can provide samples that include nucleic acids aligned with one or more genomic regions having methylation characteristics that correspond to methylation characteristics of nucleic acids that are aligned with genomic regions that correspond to the presence of one or more cell types in samples. In one or more illustrative examples, samples derived from subjects in which one or more biological conditions are present can include nucleic acids that are hypermethylated and aligned with one or more first genomic regions. Additionally, samples derived from subjects in which one or more cell types are present can be hypermethylated and aligned with one or more second genomic regions. In situations where at least one first genomic region corresponds to or overlaps with at least one second genomic region, the hypermethylated nucleic acids related to the one or more biological conditions can impact the quantitative measures determined for the overlapping regions with respect to determining the one or more cell types present in samples. In one or more additional illustrative examples, samples derived from subjects in which one or more biological conditions are present can include nucleic acids that are hypomethylated and aligned with one or more third genomic regions. Further, samples derived from subjects in which one or more cell types are present can be hypomethylated and aligned with one or more fourth genomic regions. In situations where at least one third genomic region corresponds to or overlaps with at least one fourth genomic region, the hypomethylated nucleic acids related to the one or more biological conditions can impact the quantitative measures determined for the overlapping regions with respect to determining the one or more cell types present in samples. In at least some examples, methylation characteristics of nucleic acids can be impacted by biological conditions, such as aAttorney Docket No.: GH0267WOnumber of types of cancer, aging, a number of neurological disorders, a number of metabolic disorders, exposure to environmental pollutants, exposure to a number of chemicals, and a number of autoimmune disorders. In one or more further illustrative examples, the methylation characteristics used to determine that nucleic acids used to determine the quantitative measures can correspond to one or more partitions that separate portions of a given sample according to methylation features of the nucleic acids included in the sample. The partitions can correspond to the partitions described in relation to Figure 1. For example, partitioning a sample into a plurality of subsamples can include contacting the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0331] In various examples, one or more computational models can be trained to determine amounts of cell types present in samples using training samples derived from a number of training subjects. In at least some examples, the number of training subjects can include subjects in which one or more types of cancer are present. In one or more examples, training data for the one or more models can include quantitative measures determined based on the amounts of nucleic acids derived from the training samples that correspond to a number of genomic regions and having one or more methylation characteristics. The one or more methylated characteristics can correspond to a threshold number of methylated cytosines and / or a threshold number of unmethylated cytosines present in nucleic acids that correspond to a number of genomic regions that correspond to one or more cell types. The quantitative measures can include normalized quantitative measures that are normalized based on the number of nucleic acids derived from the training samples that correspond to a number of control regions.

[0332] In one or more examples, the quantitative measures can be modified to account for nucleic acids present in the training samples that are derived from subjects in which the one or more types of cancer are present and that overlap with the one or more genomic regions that are indicative of the one or more cell types. In various examples, a tumor indication can be determined for the training samples in which the one or more types of cancer are present. The tumor indication can include a tumor fraction or a probability of one or more types of cancer being present in the training samples. The tumor indication can be determined by one or more of the computational models described in relation to Figure 2 or Figure 3. The tumor indication can be one of the features included in the one or more computational models trained to detect amounts of the one or more cell types in addition to the quantitative measures of the genomic regions that indicate the presence of the one or more cell types in samples. The one or more computational models can be trained in such a way as to modify the quantitative measures for the regions of interest related to the one or more cell types based on the tumor indications of training samplesAttorney Docket No.: GH0267WOin which one or more types of cancer are present. As a result, the impact of the nucleic acids included in the training samples that are derived from tumors present in the training subjects can be minimized and the signal related to the one or more cell types present in samples can be more accurately determined. In still other examples, the one or more computational models can be trained such that regions of interest that correspond to the one or more cell types and that overlap with or otherwise correspond to the regions of interest that correspond to the one or more cancer types are not taken into consideration or the quantitative measures derived from the overlapping regions are penalized such that the quantitative measures derived from non-overlapping genomic regions that correspond to the one or more cell types are enhanced by the one or more computational models when determining the amounts of the one or more cell types present in samples. In one or more examples, the one or more computational models used to identify genomic regions having a signal that is minimized in relation to the tumor-derived nucleic acids can include one or more regression models. The one or more regression models can include a regression machine learning model. For example, the one or more computational models used to determine the amounts of one or more cell types present in samples while minimizing the tumor-derived signal can include a linear regression model. In one or more additional examples, the one or more computational models used to determine the amounts of one or more cell types present in samples while minimizing the tumor-derived signal can include a Huber regression model. Further, the one or more computational models used to determine the amounts of one or more cell types present in samples while minimizing the tumor-derived signal can include a ridge regression model. In still other examples, the one or more computational models used to determine the amounts of one or more cell types present in samples while minimizing the tumor-derived signal can comprise a lasso regression model.

[0333] Figure 6 is a flow diagram of an example process 600 to determine amounts of a number of cell types present in a sample, according to one or more example implementations. The process 600 can include, at 602, obtaining a polynucleotide sample derived from a subject. The sample can include a bodily sample. For example, the sample can include body tissue. In one or more examples, the sample can include tissue obtained by a biopsy procedure. In one or more additional examples, the sample can include tissue obtained by a surgical procedure that takes place as part of a treatment for a biological condition. In one or more further examples, the sample can include and / or be derived from one or more bodily fluids. To illustrate, the sample can include whole blood, platelets, serum, plasma, buffy coat, stool, red blood cells, white blood cells, endothelial cells, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, boneAttorney Docket No.: GH0267WOmarrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, or urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components. In one or more illustrative examples, the sample can include cell-free nucleic acids. In one or more additional illustrative examples, the subject can include a human subject. In still other examples, the subject can include a non-human mammal subject.

[0334] At 604, the process 600 can include determining, for individual regions of interest, a number of nucleic acids derived from the sample or at least a subsample that are aligned with the individual regions of interest. In one or more examples, nucleic acids derived from the sample can be analyzed with respect to one or more criteria to determine the number of nucleic acids that are aligned with the individual regions of interest. For example, a plurality of nucleic acids can be derived from the sample and analyzed with respect to a number of cytosine-guanine dinucleotides (CpGs) present in the individual nucleic acids. In one or more illustrative examples, nucleic acids having greater than a threshold number of CpGs can be identified. In at least some examples, the threshold number of CpGs can include at least 2 CpGs, at least 3 CpGs, at least 4 CpGs, at least 5 CpGs, at least 6 CpGs, at least 7 CpGs, at least 8 CpGs, at least 9 CpGs, or at least 10 CpGs. In various examples, a subset of the plurality of nucleic acids derived from the sample that correspond to the threshold number of CpGs can comprise at least a portion of the number of nucleic acids that are aligned with individual regions of interest.

[0335] Additionally, a plurality of nucleic acids derived from the sample can be analyzed with respect to methylation states of CpGs included in the plurality of nucleic acids. For example, the plurality of nucleic acids can be analyzed with respect to a number of methylated CpGs present in the plurality of nucleic acids. Further, the plurality of nucleic acids can be analyzed with respect to a number of unmethylated CpGs present in the plurality of nucleic acids. In one or more examples, the number of methylated CpGs or unmethylated CpGs can correspond to subsamples that are separated from the sample. To illustrate, the sample can be partitioned into a plurality of subsamples. In one or more illustrative examples, the plurality of subsamples can include a first subsample and a second subsample. In at least some examples, the first subsample can comprise nucleic acids with a cytosine modification in a greater proportion than the second subsample. In one or more illustrative examples, partitioning the sample into a plurality of subsamples can include contacting the number of nucleic acids with a methyl binding reagent immobilized on a solid support.

[0336] Further, partitioning the sample into a plurality of subsamples can include separating a number of nucleic acids included in the sample into subgroups of nucleic acids onAttorney Docket No.: GH0267WOthe basis of methylation level of the nucleic acids. To illustrate, a first subgroup of the number of nucleic acids can include first additional nucleic acids having a first methylation level and a second subgroup of the number of nucleic acids can include second additional nucleic acids having a second methylation level. In one or more illustrative examples, the first methylation level can correspond to a first range of methylated CpGs being present in the first additional nucleic acids and the second methylation level corresponds to a second range of methylated CpGs being present in the second additional nucleic acids. In various examples, the second range of methylated CpGs can comprise a greater number of CpGs than the first range of methylated CpGs.

[0337] The regions of interest can include genomic regions of a reference genome that can be indicative of the presence of one or more cell types in samples. In one or more examples, the regions of interest can comprise one or more genomic regions that are differentially methylated in samples that include one or more cell types. For example, the regions of interest can comprise one or more genomic regions that include a greater number of unmethylated CpGs than additional genomic regions that do not include the one or more cell types. In addition, the regions of interest can comprise one or more genomic regions that include a greater number of methylated CpGs than additional genomic regions that do not include the one or more cell types. In at least some examples, the genomic regions can be included in one or more assays. In one or more illustrative examples, the genomic regions can be enriched in a diagnostic test for one or more biological conditions. To illustrate, the genomic regions can be enriched in a diagnostic test for one or more types of cancer.

[0338] In addition, the process 600 can include, at 606, determining quantitative measures for the individual regions of interest based on the number of nucleic acids. The quantitative measures can be determined based on a subset of the number of nucleic acids that are aligned with an individual region of interest. For example, the quantitative measure for an individual region of interest can correspond to a count of a subset of the nucleic acids derived from the sample that are aligned with the region of interest. In at least some examples, an individual count that indicates a respective subset of the nucleic acids can be determined for the individual regions of interest. In this way, a first count corresponding to a first subset of the nucleic acids derived from the sample can be determined that aligns with a first region of interest and a second count corresponding to a second subset of the nucleic acids derived from the sample can be determined that aligns with a second region of interest. In various examples, the first subset of nucleic acids can be distinct from the second subset of nucleic acids. That is, in one or more illustrative examples, a nucleic acid present in the first subset of nucleic acids may not be presentAttorney Docket No.: GH0267WOin the second subset of nucleic acids. In one or more illustrative examples, individual quantitative measures can be determined for at least 10 regions of interest, at least 25 regions of interest, at least 50 regions of interest, at least 100 regions of interest, at least 250 regions of interest, at least 500 regions of interest, at least 1000 regions of interest, at least 2500 regions of interest, at least 5000 regions of interest, or at least 10,000 regions of interest.

[0339] In one or more additional examples, the quantitative measures can include normalized quantitative measures. For example, an additional number of nucleic acids derived from the sample can be determined that are aligned with a number of individual control regions. In these scenarios, the quantitative measures for individual regions of interest can be determined based on the counts of the nucleic acids corresponding to the individual regions of interest and the number of the additional nucleic acids that are aligned with the control regions. In one or more illustrative examples, the quantitative measures for individual regions of interest can include a ratio of a count of the number of the nucleic acids aligned with the region of interest with respect to the additional number of nucleic acids that are aligned with the control regions. In one or more additional illustrative examples, a pseudocount can be added to the ratio of the count of the number of the nucleic acids aligned with the region of interest with respect to the additional number of nucleic acids that are aligned with the control regions.

[0340] Further, at 608, the process 600 can include determining, using one or more computational models and based on the quantitative measures for the individual regions of interest, a first abundance of a cell type present in the sample or at least a subsample. In one or more examples, the one or more computational models can include one or more machine learning regression algorithms. For example, the one or more computational models can implement a ridge regression algorithm. In various examples, the one or more computational models can include multiple computational models that each determine an abundance of a different cell type in the sample or at least a subsample. To illustrate, the one or more computational models can include a first computational model to determine an abundance of a first cell type in the sample or at least a subsample and a second computational model to determine an abundance of a second cell type in the sample or at least a subsample. In one or more illustrative examples, the abundance of the second cell type can be determined by determining, for individual second regions of interest, a second number of nucleic acids derived from the sample or at least one subsample that are aligned with the individual second regions of interest and that comprise at least the threshold amount of CpGs. Second quantitative measures can be determined based on the second number of nucleic acids and used by the one or more computational models to determine an abundance of a second cell type present in the sample or at least a subsample. In one or more additionalAttorney Docket No.: GH0267WOillustrative examples, at least a portion of first regions of interest used by the first computational model to determine the abundance of the first cell type can be different from at least a portion of second regions of interest used by the second computational model to determine the abundance of the second cell type. That is, in at least some scenarios, at least a portion of the first regions of interest can be differentially methylated with respect to the first cell type in relation to methylation states of one or more additional genomic regions and at least a portion of the second regions of interest can be differentially methylated with respect to the second cell type in relation to methylation states of one or more additional genomic regions.

[0341] In one or more examples, the first abundance of the first cell type present in the sample or at least a subsample can be used to determine a biological condition present in the subject. Additionally, the first abundance of the first cell type present in the sample or at least a subsample can be used to determine one or more treatments to provide to the subject in relation to a biological condition. Further, the first abundance of the first cell type present in the sample or at least a subsample can be used to determine an effectiveness of a treatment provided to the subject. In one or more illustrative examples, multiple computational models can be used to determine the abundance of multiple cell types in the sample or at least a subsample and the abundances of the multiple cell types can be used to determine at least one of one or more biological conditions present in the subject, a treatment to provide to the subject, or an effectiveness of the treatment in relation to the one or more biological conditions. In various examples, one or more classification computational models can be implemented with respect to the abundance of one or more cell types present in the sample to determine at least one of one or more biological conditions present in the subject, a treatment to provide to the subject, or an effectiveness of the treatment in relation to the one or more biological conditions.

[0342] In at least some examples, the amounts of one or more cell types present in a sample can indicate an effectiveness of one or more treatments administered to subjects. In various examples, the amounts of one or more cell types present in a sample can indicate an amount of toxicity with respect to subjects that is produced in response to one or more treatments administered to the subjects. In one or more illustrative examples, an amount of liver cells detected in a sample can indicate a level of toxicity of one or more treatments administered to subjects to treat one or more biological conditions present in the subjects.

[0343] The one or more biological conditions can include one or more types of cancer, one or more types of autoimmune diseases, one or more types of lung disease, one or more types of liver disease, one or more types of cardiovascular disease, one or more types of neurological disorders, one or more types of gastrointestinal disease, or one or more combinations thereof. InAttorney Docket No.: GH0267WOone or more illustrative examples, the one or more biological conditions can include Barrett's esophagus, schizophrenia, bipolar disease, dementia, cytokine release syndrome, immune effector cell-associated neurotoxicity syndrome (ICANS), rheumatoid arthritis, endometriosis, metabolic dysfunction-associated steatohepatitis (MASH), liver fibrosis, or one or more combinations thereof.

[0344] Additionally, the cell types identified by the one or more computational models can include at least one of Total lymphocytes, Total granulocytes, Total T cells, Activated T cells, Naive T cells, CD4 T cells, CD8 T cells, Activated CD8 T cells, Effector memory CD8 T cells, Central memory CD8 T cells, Effector CD8 T cells, Naive CD8 T cells, Activated CD4 T cells, Effector memory CD4 T cells, Central memory CD4 T cells, Effector CD4 T cells, Naive CD4 T cells, Regulatory T cells, Total monocytes and macrophages, Total monocytes and macrophages and dendritic cells, Monocytes, Total macrophages, M1 macrophages, M2 macrophages, Total B cells, Plasma B cells, Memory B cells, Naive B cells, Progenitor B cells, Naive or progenitor B cells, Activated B cells, NK cells, Megakaryocytes, Erythroid progenitors, Erythroblasts, Neutrophils, Eosinophils, Adipocytes, Bladder epithelial cells, Breast epithelial cells, Breast basal epithelial cells, Breast luminal epithelial cells, Colon epithelial cells, Endothelial cells, Fibroblast cells, Muscle-derived fibroblast cells, Heart-derived fibroblast cells, Gastric epithelial cells, Head and neck epithelial cells, Heart cardiomyocytes, Kidney epithelial cells, Liver hepatocytes, Lung epithelial cells, Lung alveolar epithelial cells, Lung bronchial epithelial cells, Neuron or oligodendrocyte cells, Ovary epithelial cells, Pancreas cells, Pancreas acinar cells, Pancreas ductal cells, Prostate epithelial cells, Small intestine epithelial cells, Thyroid epithelial cells, Exhausted lymphocytes, Exhausted T cells, Exhausted CD8 T cells, Exhausted CD4 T cells, or Exhausted B cells. In still other examples, the cell types identified by the one or more computational models can include immune cells of any of the above populations, from cancer tumors, when compared with immune cells from healthy patients, immune cells of any of the above populations, from blood of cancer patients, when compared with immune cells from the blood of healthy patients, immune cells from diseased tissue, when compared with immune cells from healthy tissue, macrophages from diseased tissue, when compared with macrophages from non-diseased tissue, or one or more combinations thereof.

[0345] The one or more computational models can be trained by obtaining a plurality of training samples from a plurality of training subjects. In one or more illustrative examples, the plurality of training samples can correspond to the first cell type. In addition, for individual first regions of interest, an additional number of nucleic acids derived from the plurality of training samples or at least one or more subsamples derived from the plurality of training samples can beAttorney Docket No.: GH0267WOdetermined that are aligned with the individual first regions of interest. The additional number of nucleic acids can comprise at least the threshold amount of CpGs in individual additional nucleic acids. Further, additional first quantitative measures for the individual first regions of interest can be determined based on the additional number of nucleic acids. A first training dataset can then be determined that indicates, for the first cell type, the additional quantitative measures that correspond to the individual first regions.

[0346] In at least some examples, the additional number of nucleic acids derived from the plurality of training samples and used to determine the additional first quantitative measures included in the first training dataset can be determined with respect to a number of methylated CpGs present in the additional number of nucleic acids. In various examples, nucleic acids derived from the first plurality of training samples can be analyzed with respect to a number of unmethylated CpGs present in the nucleic acids. The number of methylated CpGs or unmethylated CpGs can correspond to subsamples that separate from the plurality of training samples. To illustrate, individual first training samples can be partitioned into a plurality of subsamples. In one or more illustrative examples, the plurality of subsamples for each first training sample can include a first subsample and a second subsample. In one or more examples, the first subsample can comprise nucleic acids with a cytosine modification in a greater proportion than the second subsample. In one or more illustrative examples, partitioning individual first training samples into a plurality of subsamples can include contacting the nucleic acids derived from the first training samples with a methyl binding reagent immobilized on a solid support.

[0347] In one or more additional examples, a second training dataset can be generated that is combined with the first training dataset to produce a combined training dataset for a training process of the one or more computational models. The second training dataset can indicate an amount of individual cell types of a plurality of cell types present in the plurality of training samples. In one or more illustrative examples, the second training dataset can be produced by performing a cytometry by time of flight process with respect to the second plurality of training samples In one or more additional illustrative examples, the second training dataset can be produced by performing one or more ribonucleic acid (RNA) sequencing processes with respect to the plurality of training samples In one or more further illustrative examples, the second training dataset can be generated by subjecting the training samples to a different methylation state identification technique than the partitioning methylation state identification technique used to produce the first training dataset. For example, the second training dataset can be determined by performing a single site methylation state identification process with respect to the plurality of training samplesAttorney Docket No.: GH0267WOto determine a number of methylated or unmethylated CpGs present in individual further nucleic acids derived from the plurality of training samples.

[0348] In at least some examples, the combined training dataset can be used in conjunction with a reference matrix in a training process for the one or more computational models. In one or more examples, the reference matrix can be generated by determining, for individual cell types of a plurality of cell types, an amount of further nucleic acids derived from the plurality of samples that are aligned with an additional plurality of regions of interest and that include the threshold amount of CpGs and the additional threshold number of methylated or unmethylated CpGs. That is, in various examples, the regions of interest used to produce the second training dataset can be different from at least a portion of the regions of interest used to produce the first training dataset. Additionally, for the individual regions of interest, further quantitative measures for the individual cell types can be determined based on the further number of nucleic acids that are aligned with individual additional plurality of regions of interest and that correspond to the individual cell types. The reference matrix can then be generated based on the further number of nucleic acids that correspond to the additional plurality of regions of interest. In one or more illustrative examples, the reference matrix can indicate an amount of the further nucleic acids that corresponds to the individual cell types and the individual additional plurality of regions of interest. For example, the reference matrix can indicate an amount of nucleic acids that correspond to the individual cell types and that correspond to the individual regions of interest. In at least some examples, the reference matrix can be used in the training process for the one or more computational models to determine individual weights to be assigned to features of the one or more computational models that correspond to individual regions of interest. Thus, some regions of interest can be weighted more heavily in determining an amount of the cell type present in the sample.

[0349] In one or more examples, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from training subjects in which one or more biological conditions are not present. For example, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from subjects in which a tumor has not been detected. In one or more additional examples, computational models trained to determine amounts of one or more cell types present in samples can be trained using training samples derived from training subjects in which one or more biological conditions are present. In various examples, one or more cell types being identified in samples can be related to the one or more biological conditions present in subjects. To illustrate, amounts of one or more cell types relatedAttorney Docket No.: GH0267WOto treatment toxicity and / or biological reactions to treatments of a biological condition can be determined based on training samples obtained from training subjects in which the biological condition is present.

[0350] In one or more illustrative examples, training subjects in which an autoimmune related biological condition has been detected may be receiving treatment for the autoimmune related biological condition. In these scenarios, one or more computational models can be trained using training data derived from the training subjects to determine the amounts of one or more cell types present in the training subjects. The amounts of the one or more cell types present in the training subjects can be used to determine an effectiveness of the treatment for the autoimmune related biological condition and / or to determine an amount of toxicity produced within the training subjects in response to the treatment. In one or more additional illustrative examples, training subjects in which one or more cancer types have been detected may be receiving treatment for the one or more cancer types. In these instances, one or more computational models can be trained using training data derived from the training subjects to detect the amounts of one or more cell types present in the training subjects. The amounts of the one or more cell types present in the training subjects can be used to determine an effectiveness of the treatment for the one or more types of cancer and / or to determine an amount of toxicity produced within the training subjects in response to the treatment.

[0351] In situations where one or more computational models are trained to determine an amount of one or more cell types present in subjects in which one or more biological conditions are present, the regions of interest for which the quantitative measures are determined can have some overlap such that one or more first regions of interest that are indicative of the one or more cell types can overlap with at least a portion of one or more second regions of interest that are indicative of the one or more biological conditions being present in subjects. As a result, in scenarios where a computational model is being trained to determine amounts of one or more cell types present in subjects, an amount of signal being detected in relation to the one or more biological conditions being present in subjects can be taken into account when determining the amounts of the one or more cell types. That is, in instances where the regions of interest indicative of the one or more cell types being present in samples and the regions of interest indicative of the one or more biological conditions being present in samples overlap, at least in part, the quantitative measures determined for one or more regions of interest corresponding to the presence of the one or more cell types in samples can be impacted by the amount of nucleic acids identified in samples for which the one or more biological conditions are present. In at least some examples, the nucleic acid molecules derived from samples in which the one or more biologicalAttorney Docket No.: GH0267WOconditions are present that are aligned with regions of interest corresponding to the one or more cell types can result in a noise signal that can impact at least one of the precision or the accuracy of the determination of the amounts of the one or more cell types present in the subjects. In these scenarios, the noise signal can be reduced by identifying regions of interest that correspond to the one or more cell types being present in subjects that do not overlap with regions of interest that are indicative of the one or more biological conditions being present in subjects. In still other examples, the noise signal can be reduced by generating a corrected signal that accounts for the noise signal produced by the nucleic acids present in samples in which the one or more biological conditions are present. By training computational models that determine amounts of one or more cell types present in samples to minimize or reduce the noise signal generated due to one or more biological conditions being present in subjects, the implementations of the computational models described herein can more efficiently and accurately determine the amounts of cell types present in samples.

[0352] In one or more illustrative examples, one or more training processes can be performed to generate one or more computational models that can determine the amounts of one or more cell types present in samples. The one or more training processes can be performed using training data derived from training subjects in which one or more biological conditions are present. In one or more examples, the training data can be derived from subjects in which the one or more biological conditions impact the methylation characteristics of nucleic acids derived from a subject in which the one or more biological conditions are present. In at least some examples, subjects in which one or more biological conditions are present can provide samples that include nucleic acids aligned with one or more genomic regions having methylation characteristics that correspond to methylation characteristics of nucleic acids that are aligned with genomic regions that correspond to the presence of one or more cell types in samples. In one or more illustrative examples, samples derived from subjects in which one or more biological conditions are present can include nucleic acids that are hypermethylated and aligned with one or more first genomic regions. Additionally, samples derived from subjects in which one or more cell types are present can be hypermethylated and aligned with o...

Claims

Attorney Docket No.: GH0267WOCLAIMS WHAT IS CLAIMED IS:

1. A method comprising:obtaining a polynucleotide sample derived from a subject;determining, for individual first regions of interest from among a plurality of regions of interest, a first number of nucleic acids derived from the sample or at least a subsample that are aligned with the individual first regions of interest, wherein the first number of nucleic acids comprise at least a threshold amount of cytosine-guanine dinucleotides (CpGs) in individual first nucleic acids;determining first quantitative measures for the individual first regions of interest based on the first number of nucleic acids;determining, by executing a first computational model and based on the first quantitative measures, an indicator of a biological condition for the subject; anddetermining, using one or more second computational models and based on the first quantitative measures for the individual first regions of interest and based on the indicator of the biological condition for the subject, a first abundance of a first cell type present in the sample or at least a subsample.

2. The method of claim 1, comprising:determining, for individual second regions of interest from among the plurality of regions of interest, a second number of nucleic acids derived from the sample or at least one subsample that are aligned with the individual second regions of interest, wherein the second number of nucleic acids comprise at least the threshold amount of CpGs;determining second quantitative measures for the individual second regions of interest based on the second number of nucleic acids; anddetermining, using the one or more second computational models and based on the second quantitative measures for the individual second regions of interest, a second abundance of a second cell type present in the sample or at least a subsample.

3. The method of claim 2, wherein:the individual first regions of interest include a first subset of the plurality of regions of interest that are differentially methylated for the first cell type; andthe individual second regions of interest include a second subset of the plurality of regions of interest that are differentially methylated for the second cell type.Attorney Docket No.: GH0267WO4. The method of any one of claims 1-3, comprising:determining, for individual control regions, an additional number of nucleic acids derived from the sample that are aligned with the individual control regions; andwherein the first quantitative measures for the individual first regions of interest are determined based on the first number of nucleic acids and the additional number of nucleic acids.

5. The method of any one of claims 1-4, wherein the one or more computational models implement a ridge regression algorithm.

6. The method of claim 1, comprising:determining that the first number of nucleic acids satisfy one or more methylation characteristics, the one or more methylation characteristics corresponding to methylation states of the CpGs in the individual first nucleic acids.

7. The method of claim 6, wherein the at least one subsample is obtained by partitioning the sample into a plurality of subsamples including a first subsample and a second subsample, wherein the first subsample comprises nucleic acids with a cytosine modification in a greater proportion than the second subsample.

8. The method of claim 7, wherein partitioning the sample into a plurality of subsamples comprises separating a number of nucleic acids included in the sample into subgroups of nucleic acids on the basis of methylation level.

9. The method of claim 8, wherein a first subgroup of the number of nucleic acids include first additional nucleic acids having a first methylation level and a second subgroup of the number of nucleic acids include second additional nucleic acids having a second methylation level.

10. The method of claim 9, wherein the first methylation level corresponds to a first range of methylated CpGs being present in the first additional nucleic acids and the second methylation level corresponds to a second range of methylated CpGs being present in the second additional nucleic acids, the second range of methylated CpGs comprising a greater number of CpGs than the first range of methylated CpGs.Attorney Docket No.: GH0267WO11. The method of any one of claims 7-10, wherein partitioning the sample into a plurality of subsamples comprises contacting the first number of nucleic acids with a methyl binding reagent immobilized on a solid support.

12. The method of any one of claims 1-11, comprising:obtaining a first plurality of training samples from a first plurality of subjects, the first plurality of training samples corresponding to the first cell type;determining, for the individual first regions of interest, an additional number of nucleic acids derived from the first plurality of training samples or at least one or more subsamples derived from the first plurality of training samples that are aligned with the individual first regions of interest, wherein the additional number of nucleic acids comprise at least the threshold amount of CpGs in individual additional nucleic acids;determining additional first quantitative measures for the individual first regions of interest based on the additional number of nucleic acids; andgenerating a first training dataset that indicates, for the first cell type, the additional first quantitative measures that correspond to the individual first regions.

13. The method of claim 12, wherein the additional first quantitative measures are modified based on a tumor fraction corresponding to at least a portion of the first plurality of training samples.

14. The method of claim 12, comprising:obtaining a second plurality of training samples;generating a second training dataset that indicates an amount of individual cell types of a plurality of cell types present in the second plurality of training samples; andgenerating a combined training dataset for the one or more second computational models based on the first training dataset and the second training dataset.

15. The method of claim 14, wherein the second training dataset is produced by performing a cytometry by time of flight process with respect to the second plurality of training samples.Attorney Docket No.: GH0267WO16. The method of claim 14, wherein the second training dataset is produced by performing one or more ribonucleic acid (RNA) sequencing processes with respect to the second plurality of training samples.

17. The method of claim 14, wherein the second training dataset is produced by determining, for the individual cell types of the plurality of cell types, an amount of further nucleic acids that are aligned with an additional plurality of regions of interest and that include the threshold amount of CpGs and an additional threshold number of methylated or unmethylated CpGs;determining, for the individual additional plurality of regions of interest, further quantitative measures for the individual cell types based on the further number of nucleic acids that (i) are aligned with individual additional plurality of regions of interest and (ii) correspond to the individual cell types; andgenerating, based on the further number of nucleic acids aligned with the individual additional plurality of regions of interest, a reference matrix that indicates an amount of the further nucleic acids that corresponds to the individual cell types and the individual additional plurality of regions of interest.

18. The method of claim 17, wherein a training process for the one or more second computational models uses the reference matrix and the combined training dataset.

19. The method of claim 17 or 18, comprising:performing a single site methylation state identification process with respect to the second plurality of training samples to determine a number of methylated or unmethylated CpGs present in individual further nucleic acids derived from the second plurality of training samples.

20. The method of any one of claims 1-19, comprising:determining, based on the first abundance of the first cell type present in the sample or at least a subsample, a biological condition present in the subject.

21. The method of claim 20, wherein the one or more second computational models were trained using training data, at least a portion of the training data being obtained from a plurality of training subjects in which the biological condition is not present.Attorney Docket No.: GH0267WO22. The method of any one of claims 1-19, comprising:determining, based on the first abundance of the first cell type present in the sample or at least a subsample, one or more treatments to provide to the subject in relation to a biological condition.

23. The method of any one of claims 1-19, comprising:determining, based on the first abundance of the first cell type present in the sample or at least a subsample, an effectiveness of a treatment provided to the subject.

24. The method of any one of claims 1-23, wherein the sample is obtained from whole blood, plasma, urine, cerebrospinal fluid, buffy coat, or a tissue biopsy.

25. The method of any one of claims 1-24, wherein the first cell type includes B-cells, T-cells, neutrophils, NK cells, monocytes, megakaryocytes, erythroid cells, liver cells, endothelial cells, or cardiovascular cells.

26. The method of any one of claims 1-25, comprising:determining, based on the first abundance of the first cell type present in the sample or at least a subsample, an amount of toxicity present in the subject due to one or more treatments administered in relation to one or more biological conditions.

27. The method of any one of claims 1-26, wherein the first computational model comprises a regression computational model.

28. The method of any one of claims 1-27, wherein the first computational model is trained using a training dataset derived from additional training samples having tumor fractions within a specified range of tumor fractions.

29. A method comprising:obtaining a polynucleotide sample derived from a subject;determining a number of nucleic acids derived from the sample that are aligned with individual regions of interest of a plurality of regions of interest, wherein individual nucleic acids of the number of nucleic acids comprise at least a threshold amount of cytosine-guanineAttorney Docket No.: GH0267WOdinucleotides (CpGs) and an additional threshold amount of methylated CpGs or unmethylated CpGs;determining, for individual regions of interest, quantitative measures for individual cell types of a plurality of cell types based on a subset of the number of nucleic acids that are aligned with the individual regions of interest; anddetermining, using one or more computational models and based on the quantitative measures and based on a reference matrix, an amount of the individual cell types present in the sample, wherein reference matrix that indicates amounts of nucleic acids that correspond to the individual cell types and the individual regions of interest.

30. The method of claim 29, comprising:obtaining a plurality of training samples from a plurality of subjects;determining an additional number of nucleic acids derived from the plurality of training samples that are aligned with the individual regions of interest, wherein individual additional nucleic acids of the additional number of nucleic acids comprise at least the threshold amount of CpGs and the additional threshold amount of methylated CpGs or unmethylated CpGs;determining, for individual regions of interest, additional quantitative measures for the individual cell types based on the additional number of nucleic acids that (i) are aligned with the individual regions of interest and (ii) correspond to the individual cell types;generating, based on the additional number of nucleic acids aligned with the individual regions of interest, the reference matrix such that the reference matrix indicates an amount of the additional number of nucleic acids that corresponds to the individual cell types and the individual regions of interest.

31. The method of claim 30, wherein a training process for the one or more computational models uses the reference matrix.

32. The method of any one of claims 29-31, wherein the one or more computational models include a least squares model.

33. The method of claim 31, wherein the one or more computational models include a partial least squares model with a non-negative value constraint.

34. The method of any one of claims 29-33, comprising:Attorney Docket No.: GH0267WOperforming a single site methylation state identification process with respect to the sample to determine a number of methylated or unmethylated CpGs present in individual nucleic acids derived from the sample.

35. The method of any one of claims 29-34, comprising:determining, based on the amount of the individual cell types present in the sample, a biological condition present in the subject.

36. The method of claim 35, wherein the one or more computational models were trained using training data obtained from a plurality of subjects in which the biological condition is not present.

37. The method of any one of claims 29-34, comprising:determining, based on the amount of the individual cell types present in the sample, one or more treatments to provide to the subject in relation to a biological condition.

38. The method of any one of claims 29-34, comprising:determining, based on the amount of the individual cell types present in the sample, an effectiveness of a treatment provided to the subject.

39. The method of any one of claims 27-38, comprising:determining, based on the amount of at least one individual cell type present in the sample or at least a subsample, an amount of toxicity present in the subject due to one or more treatments administered in relation to one or more biological conditions.