Multimodal sample swap detection and remediation
A multimodal approach with sample swap detection and remediation models improves cancer detection accuracy by correcting mislabeled samples, addressing the limitations of current methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GRAIL INC
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
AI Technical Summary
Current cancer detection methods are limited by low detection rates, high false positive rates, and the challenge of detecting swapped samples, which can lead to inaccurate analysis results.
A multimodal approach using multiple sample swap detection models to identify mislabeled samples, followed by a genetic similarity model to correct labeling, and a remedial workflow to address sample swaps, ensuring accurate cancer detection.
Enhances the accuracy of cancer detection by correcting sample mislabeling, thereby improving the reliability of early cancer detection and diagnosis.
Smart Images

Figure US2026012441_30072026_PF_FP_ABST
Abstract
Description
MULTIMODAL SAMPLE SWAP DETECTION AND REMEDIATION Inventors:Joerg BrednoJohn BeausangChristopher ChangJoseph MarcusDemetri NicolaouApril SaganSuba SrinivasanAlexis ThorntonCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of and priority to U.S. Provisional Application No. 63 / 749,861 filed on January 27, 2025, which is incorporated by reference.BACKGROUND
[0002] Cancer is a leading cause of death worldwide. The fatality of cancer is heightened by the fact that cancer is usually detected in later stages, limiting efficacy of treatment options for long-term survival. Current detection methods generally are specific to a particular ty pe of cancer, i.e., each cancer is separately screened for. Accordingly, each subject screening process is tailored to the t pe. For example, mammography scans are utilized in breast cancer detection, whereas colonoscopy or fecal tests have helped with colorectal cancer detection. Each varied screening method is generally not cross-applicable to other types.
[0003] Furthermore, present screening methods are encumbered by low detection rates or high false positive rates. Low detection rates often fail to detect early-stage cancers as the cancers are just developing. A high false positive rate misdiagnoses cancer-free subjects as positive for cancer status. As a result, most screening tests are only practical when they are used to test subjects who have a high risk of developing the screened cancer or have symptoms indicative of the presence of suspected cancers. As such, most screening tests have limited ability to detect cancers in the general population.
[0004] Novel research has implicated aberrant DNA methylation in many disease processes, including cancer. DNA methylation plays a role in regulating gene expression and defining tissue differentiation, cellular identity, and / or embryological lineage. Thus, aberrant DNA methylation can create issues in normal gene expression pathways or cellular identity.Atty. Docket No. 32917-65128 / WO 1 P0248-WOthereby leading to cancer or other diseases. For example, specific patterns of differentially methylated regions may be useful as molecular markers for various disease states. Detection of these differentially methylated regions may be accomplished through sequencing analysis of cell-free DNA molecules. In general, cell-free DNA molecules are DNA molecules that arise in bodily fluids. These DNA molecules are typically released due to natural cell death, active release by healthy cells, or tumor-derived DNA molecules shed from tumor cells undergoing cell death. Nonetheless, tests to detect differentially methylated regions face a number of challenges. Early cancer detection is particularly challenging due to the miniscule ratio of tumor cells to non-cancer cells in the subject. The miniscule ratio may be on the order of 1:1000, 1:10,000, or even 1:100,000. This creates a challenge of detecting small amounts of cancer signal amidst healthy signal, especially when analyzing this signal with an easily accessible sample condition, for example a blood draw to assess presence of cancer signal in blood plasma, for example in cell-free DNA.
[0005] One particular challenge arises in detecting swapped samples, i.e., samples that are mislabeled. A sample swap event may occur through human or machine error. For example, a human clinician helping to obtain a biological sample from an individual may unintentionally label the collected sample with that intended for another. In effect, the biological sample is labeled as belonging to the wrong individual, i.e., not the individual that the biological sample truly belongs to. In other examples, a machine tasked with assigning the appropriate labels may misassign a label. In any of such instances, the label for the biological sample, intended to indicate the individual of origination, is wrong. If these sample swap events go undetected, an analytics system performing early cancer detection analyses on the sample may provide a report of the analyses that misrepresent the individual. Thus, the detection of such events is key to safeguarding the accuracy of the early cancer detection analyses.
[0006] The present disclosure is directed to addressing the above-referenced challenges. The background description provided herein is for the purpose of generally- presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.Atty. Docket No. 32917-65128 / WO 2 P0248-WOSUMMARY
[0007] The invention(s) described herein this disclosure provide for improvements to cancer detection, diagnosis and treatment, in particular, providing for sample swap detection and remediation. Multimodal sample swap prediction entails applying a plurality of sample swap detection models to features of sequencing data from a sample being assessed. Each sample swap detection model is configured in a distinct mode. Each sample swap detection model outputs a prediction of whether the sample is affected by a sample swap event based on the features of the sequencing data. An aggregation model inputs the disparate predictions from the multimodal sample swap prediction models to output an aggregate sample swap prediction. Upon detecting a mislabeled sample, sample swap reversal may be performed to identify a corrected labeling without requiring an individual to submit a new sample. Sample swap reversal leverages a genetic similarity model that assesses genetic similarity between genotype signatures of a pair of samples. In some examples, a graph is constructed with nodes representing the samples, and edges connecting samples expected to belong to the same individual. The edges represent the genetic similarity between samples within the group. Graph optimization identifies a corrected group for a mislabeled sample, based on assessed genetic similarity. Remediation may be performed based on the corrections.
[0008] In some aspects, the techniques described herein relate to a method for multimodal sample swap detection, the method including: obtaining a sample labeled as belonging to an individual, the sample including sequencing data derived from sequencing nucleic acid fragments in the sample, and one or more reported covariate labels derived from clinical information on the individual; determining a plurality of genetic features from the sequencing data; applying a plurality of sample swap detection models to the plurality of genetic features for the sample and to the one or more reported covariate labels to output a plurality of predictions indicating whether the sample is mislabeled, each sample swap detection model being distinctly configured: applying an aggregation model to the plurality of predictions to determine an aggregate prediction indicating whether the sample is mislabeled; calling a sample swap event in response to determining that the aggregate prediction is above a threshold; and performing a remedial workflow to address the sample swap event.
[0009] In some aspects, the techniques described herein relate to a method, wherein the sample is a liquid biopsy sample including nucleic acid fragments shed in bodily fluid, or a tissue biopsy sample including nucleic acid fragments extracted from tissue cells.Atty. Docket No. 32917-65128 / WO 3 P0248-WO
[0010] In some aspects, the techniques described herein relate to a method, wherein the sequencing data includes sequencing data in a plurality of forms including methylation sequencing data, small variant sequencing data, copy number sequencing data, or some combination thereof, and wherein a first sample swap detection model is configured to input one or more genetic features derived from a first form of sequencing data to output a first prediction of whether the sample is mislabeled and a second sample swap detection model is configured to input one or more genetic features derived from a second form of sequencing data, that is different from the first form of sequencing data, to output a second prediction of whether the sample is mislabeled.
[0011] In some aspects, the techniques described herein relate to a method, wherein applying the plurality of sample swap detection models includes: applying one sample swap detection model to predict a label for one covariate based on the plurality of genetic features; comparing the predicted label to the reported label for the covariate; and determining the prediction of whether the sample is mislabeled based on the comparison of the predicted label to the reported label.
[0012] In some aspects, the techniques described herein relate to a method, wherein determining a plurality7of genetic features includes determining a plurality of methylation features from methylation sequencing data of the sequencing data; and wherein applying the one sample swap detection model includes applying the one sample swap detection model to the plurality of methylation features to predict the label for biological sex as the covariate.
[0013] In some aspects, the techniques described herein relate to a method, wherein determining a plurality' of genetic features includes determining a plurality' of methylation features from methylation sequencing data of the sequencing data; and wherein applying the one sample swap detection model includes applying the one sample swap detection model to the plurality of methylation features to predict the label for epigenetic age as the covariate.
[0014] In some aspects, the techniques described herein relate to a method, wherein determining a plurality of genetic features includes determining a plurality7of methylation features from methylation sequencing data of the sequencing data; and wherein applying the one sample swap detection model includes applying the one sample swap detection model to the plurality of methylation features to predict the label for ancestry as the covariate.
[0015] In some aspects, the techniques described herein relate to a method, wherein applying the aggregation model includes applying the aggregation model further to the predicted label for the covariate to determine the aggregate prediction.Atty. Docket No. 32917-65128 / WO 4 P0248-WO
[0016] In some aspects, the techniques described herein relate to a method, wherein determining the plurality of genetic features from the sequencing data includes determining a first genotype signature for the sample based on single nucleotide polymorphism (SNP) genotypes across a plurality of SNP sites in a reference genome; wherein applying the plurality of sample swap detection models includes: applying one sample swap detection model to (i) the first genotype signature for the sample and (ii) a second genotype signature for a reference sample also labeled as belonging to the individual, to output a genetic similarity score between the sample and the reference sample; comparing the genetic similarity score to a threshold score; and determining the prediction of whether the sample is mislabeled based on the comparison of the genetic similarity score to the threshold score.
[0017] In some aspects, the techniques described herein relate to a method, wherein performing the remedial workflow to address the sample swap event includes one or more of: labeling the sample as unusable; generating and transmitting a notification to a user to obtain anew sample from the individual; withholding the sample from cancer classification; withholding the sample from reporting to the individual; removing the sample from a training data set for training of a cancer classification machine-learning model; and identifying a correct origination of the sample.
[0018] In some aspects, the techniques described herein relate to a method for sample swap detection, the method including: obtaining a plurality of samples including a target sample, each sample is labeled as belonging to one of a plurality of individuals, the sample including single nucleotide polymorphism (SNP) sequencing data derived from sequencing nucleic acid fragments in the sample; determining, for each sample, a genotype signature based on SNP genotypes from the sequencing data of the sample across a plurality of SNP sites in a reference genome; generating a plurality of groups, wherein each group associates samples labeled as belonging to one of the plurality of individuals; for each of one or more pairs of samples in one group, applying a genetic similarity model to the genotype signatures of the pair of samples to output a genetic similarity score betw een the pair of samples; calling a sample swap event in response to determining that one sample of one group is mislabeled based on a comparison of the one or more genetic similarity scores to one or more other samples in the one group to a threshold score; and performing a remedial w orkflow to address the sample swap event.
[0019] In some aspects, the techniques described herein relate to a method, wherein determining the genotype signature for each sample based on the SNP genotypes from the sequencing data includes: assigning a first value for a homozygous reference genotype, aAtty. Docket No. 32917-65128 / WO 5 P0248-WOsecond value for a heterozy gous genotype, a third value for a homozy gous alternate genotype, and a fourth value for a missing genotype.
[0020] In some aspects, the techniques described herein relate to a method, wherein applying the genetic similarity model to the genotype signatures of the pair of samples includes: identifying a common subset of the plurality of SNP sites with both genotype signatures not having the fourth value for the missing genotype; identifying a count of SNP sites of the common subset of SNP sites with both genotype signatures having matching values; and determining the genetic similarity score for the pair of samples based on a ratio of the count of SNP sites having the matching values to a total number of SNP sites in the common subset.
[0021] In some aspects, the techniques described herein relate to a method, wherein determining the one sample of the one group is mislabeled includes: determining a majority of the genetic similarity scores between the one sample and other samples of the group is below' the threshold score.
[0022] In some aspects, the techniques described herein relate to a method, further including: generating a graph including a plurality of nodes representing the plurality of samples and a plurality' of edges connecting nodes of a group, wherein each edge represents the genetic similarity score between the samples connected by the edge.
[0023] In some aspects, the techniques described herein relate to a method, further including: performing a remedial workflow to address the sample swap event.
[0024] In some aspects, the techniques described herein relate to a method, wherein performing the remedial workflow to address the sample swap event includes: performing graph optimization to identify which of the other groups maximizes genetic similarity with the one sample; and relabeling the one sample to be part of a second group that maximizes genetic similarity' with the one sample.
[0025] In some aspects, the techniques described herein relate to a method, wherein performing the graph optimization includes: identify ing one other sample also determined to be mislabeled; and assessing genetic similarity of the one sample with the group of the other sample also determined to be mislabeled.
[0026] In some aspects, the techniques described herein relate to a method, wherein assessing genetic similarity of the one sample with the group of the other sample also determined to be mislabeled includes: applying the genetic similarity model to the genotype signature of a third sample in the group of the other sample determined to be mislabeled to output a genetic similarity score between the one sample and the third sample.Atty. Docket No. 32917-65128 / WO 6 P0248-WO
[0027] In some aspects, the techniques described herein relate to a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method.
[0028] In some aspects, the techniques described herein relate to a system including: one or more computer processors; and the non-transitory7computer-readable storage medium.
[0029] In some aspects, the techniques described herein relate to a treatment kit including: a collection vessel for collecting a DNA sample from a subject; optionally, one or more reagents for isolating DNA fragments in the DNA sample; optionally, one or more probes targeting one or more genomic loci determined to be indicative.BRIEF DESCRIPTION OF DRAWINGS
[0030] FIG. 1 is an exemplary flowchart describing an overall workflow of cancer classification of a sample, according to one or more embodiments.
[0031] FIG. 2 A is an exemplary flowchart describing a process of sequencing a fragment of cell-free (cf) DNA to obtain a methylation state vector, according to one or more embodiments.
[0032] FIG. 2B is an exemplary illustration of the process of FIG. 2A of sequencing a fragment of cell-free (cf) DNA to obtain a methylation state vector, according to one or more embodiments.
[0033] FIG. 3 is an exemplary flowchart describing a process of multimodal sample swap detection, according to one or more embodiments.
[0034] FIG. 4 is an exemplary flowchart of the process of deployment of a genetic similarity model, according to one or more embodiments.
[0035] FIG. 5 is an exemplary flowchart describing a process of sample swap reversal, according to one or more embodiments.
[0036] FIG. 6 is an illustrative example of the process of sample swap reversal, according to one or more embodiments.
[0037] FIG. 7 A is an exemplary flowchart describing a process of generating a control group data structure for determining anomalously methylated fragments, according to one or more embodiments.
[0038] FIG. 7B is an exemplary flowchart describing a process of determining a fragment to be anomalously methylated based on the control group data structure, according to one or more embodiments.Atty. Docket No. 32917-65128 / WO 7 P0248-WO
[0039] FIG. 8A is an exemplary flowchart describing a process of training a cancer classifier, according to one or more embodiments.
[0040] FIG. 8B illustrates an example generation of feature vectors used for training the cancer classifier, according to one or more embodiments.
[0041] FIG. 9A illustrates an exemplary flowchart of devices for sequencing nucleic acid samples according to one or more embodiments.
[0042] FIG. 9B is an exemplary block diagram of an analytics system, according to one or more embodiments.
[0043] The figures depict various embodiments for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.DETAILED DESCRIPTIONI. OVERVIEW
[0044] Early detection and classification of cancer is an important technology. Being able to detect cancer before it becomes symptomatic is beneficial to all parties involved, including patients, doctors, and loved ones. For patients, early cancer detection allows them a greater chance of a beneficial outcome; for doctors, early cancer detection allows more insight into the context and status of disease progression and availability of more pathways of treatment that may lead to a beneficial outcome; for loved ones, early cancer detection increases the likelihood of not losing friends and family to the disease.
[0045] In many instances today, early detection and classification of cancer is only achieved when a patient is being worked up for symptoms that can indicate the presence of a cancer and only after a wide variety of cost-intensive medical procedures. This process, sometimes referred to as a “diagnostic odyssey,” causes stress and anxiety for patients and loved ones over sometimes multiple months of uncertainty.
[0046] Recently, early cancer detection technology has progressed towards analyzing genetic fragments (e.g., DNA) in a sample from a person to determine if any of those genetic fragments originate from cancer cells. To increase the benefit of early cancer detection, the sample is often a relatively easy to acquire liquid sample from the person, such as blood, saliva, or urine. These new techniques allow doctors to identify a cancer presence in a patient that may not be detectable otherwise, e.g.. in conventional screening processes. For instance,Atty. Docket No. 32917-65128 / WO 8 P0248-WOconsider the example of a person at risk for breast cancer. Traditionally, this person will regularly visit their doctor for a mammogram, which creates an image of their breast tissue (e.g., taking x-ray images) that a doctor uses to identify cancerous tissue. Unfortunately, with even the highest resolution mammograms, doctors are only able to identify tumors once they are approximately a millimeter in size. This means that the cancer has been present for some time in the person and has gone undiagnosed and untreated. Visual determinations like this are typical for most cancers — that is, only identifiable once it has grown to a sufficient size to be detected by some sort of imaging technology. For many more types of cancer, there is no screening paradigm like mammography available, and even regular doctor visits cannot detect the presence of a growing and advancing cancer. For example, the person having a risk of breast cancer above might at the same time be at high risk for ovarian cancer, which is undetectable by mammography or any other currently available image-based cancer screening modality.
[0047] Cancer detection using analysis of genetic fragments in a patient, e.g., in their blood, provides for technical improvements to these issues in the conventional screening paradigm. To illustrate, cancer cells will start shedding DNA fragments into a person’s bloodstream as soon as they form. This occurs when there are very few of the cancer cells, and long before they’ve collectively achieved a size large enough to be visible with conventional imaging techniques. With the appropriate methods, therefore, a system that analyzes DNA fragments in the bloodstream (so-called "cell-free DNA” or “cfDNA”) could identify cancer presence in a person based on shed cancer DNA fragments, and, more importantly, the system could do so before the cancer is identifiable using more traditional cancer detection techniques, and even more so for cancers where no traditional cancer screening method exists.
[0048] Cancer detection based on the analysis of DNA fragments is enabled by nextgeneration sequencing (“NGS”) techniques. NGS, broadly, is a group of technologies that allows for high throughput sequencing of genetic material. The high throughput sequencing promotes maximum capturability of nucleic acid fragments in a biological sample. Without high throughput sequencing, faint signals may go undetected. As discussed in greater detail herein, NGS largely consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation includes the laboratory' methods necessary' to prepare DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in the samples, and data analysis includes processing and analyzing the genetic information in the sequencing data to identify cancer presence.Atty. Docket No. 32917-65128 / WO 9 P0248-WO
[0049] Such an analysis to identify the presence of cancer and classify its type (including organ system or tissue type and / or histologic origin) is further relevant when a patient has been diagnosed with cancer and additional information is needed for prognosis, treatment decision, and for a patient on treatment to identify residual disease, recurrence, or relapse for a patient on treatment when a treatment has not or is no longer able to control cancer growth.
[0050] While these steps of NGS may help enable early cancer detection, they also introduce their own complex, detrimental challenges to consistent early cancer detection and, therefore, any improvements to sample preparation, DNA sequencing, and / or data analysis, including the pre-processing, algorithmic processing, and summary or presentation of predications or conclusions, results in an improvement to NGS technology, cancer detection technologies broadly, and early cancer detection more generally.
[0051] To illustrate, as an example, problems introduced in (1) sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Remedying these problems would yield better genetic data for cancer detection.
[0052] Similarly, problems introduced in (2) sequencing include, for example, errors in accurate transcribing of fragments (e g., reading an “A” instead of a “C”, etc.), incorrect or difficult fragment assembly and overlap, disparate coverage uniformity', sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Again, remedying any of these problems would yield improved genetic data for cancer detection.
[0053] The problems in (3) data analysis are the most daunting and complex. The introduced challenges stem from the vast amounts of data created by NGS sequencing techniques. Sequencing data for a single sample can be on the order of hundreds of thousands (up to millions) of sequence reads, amounting to terabytes of data. Training analytical models ty pically involves collecting and processing thousands (up to tens of thousands or more) of samples with ascertained clinical cancer status, affected organ or organ group, primary site of cancer, underlying cancer biology, and other factors. Effectively and efficiently analyzing that amount of data is both procedurally and computationally demanding. For instance, analyzing NGS sequencing involves several baseline processing steps such as, e.g., aligning reads to one another, aligning and mapping reads to a reference genome, de-duplicating duplicative reads, detecting contamination of a sample, identifying and calling variant genes, identify ing and calling abnormally methylated individual genomic sites or regions, generating functional annotations, making predictions or identifications based on the called genomic methylation sites, etc. Performing any of these processes on terabytes (or more) of geneticAtty. Docket No. 32917-65128 / WO 10 P0248-WOdata is computationally expensive for even the most powerful of computer architectures, and completely impossible for a normal human mind to achieve within a clinical relevant span of time (i.e., while the cancer is still in early stages when it is most likely to be curable).Additionally, with the genetic sequencing data derived from the error-prone processes of sample preparation and sequence reading, large portions of the resulting genetic data may be low-quality or unusable for cancer identification. For example, large amounts of the genetic data may include contaminated samples, transcription errors, mismatched regions, overrepresented regions, non-informative regions, etc. and may be unsuitable for high accuracy cancer detection. Identifying and accounting for low qualify genetic data across the vast amount of genetic data obtained from NGS sequencing is also procedurally and computationally challenging to accomplish and is also not practically performable by a human mind. Overall, any process created that leads to more efficient processing of the large and structurally complex collection of sequencing data typically relied on for the analytical models used in early cancer detection enabled by NGS would be an improvement to cancer detection using NGS sequencing. Moreover, such processes as described and elaborated on herein were crafted as a direct solution to the various hurdles native to NGS technologies, and as such are non-routine and unconventional activity in this technical field.
[0054] As a further example, under (3) data analysis, accurate identification of informative DNA from NGS data to identify a cancer presence is another difficult task native to the field of early cancer detection. To be effective, algorithms are sought to compensate for, e.g., errors generated by sample preparation and sequencing, the large scale of genomic and methylation variety present in the population, and to overcome the large-scale data analysis problems accompanying NGS techniques. That is, designing a machine learning model or models, or other computational processing algorithms, that enable early cancer detection based on next generation sequencing techniques must account for the problems that those techniques create. Some of those techniques and models are discussed hereinbelow and particular improvements to state-of-the-art techniques and models are further discussed. Furthermore, such techniques are non-routine and unconventional activity in the technical field of endeavor.
[0055] One particular set of challenges arises in detecting swapped samples, i.e., samples that are mislabeled. A sample swap event may occur through human or machine error. For example, a human clinician helping to obtain a biological sample from an individual may unintentionally label the collected sample with that intended for another. In effect, the biological sample is labeled as belonging to the wrong individual, i.e., not theAtty. Docket No. 32917-65128 / WO 11 P0248-WOindividual that the biological sample truly belongs to. In other examples, a machine tasked with assigning to the appropriate labels may misassign a label. In any of such instances, the label for the biological sample, intended to indicate the individual of origination, is wrong. If these sample swap events go undetected, an analytics system performing early cancer detection analyses on the sample may provide a report of the analyses that misrepresent the individual. Thus, the detection of such events is key to safeguarding the accuracy of the early cancer detection analyses.I. A. CANCER CLASSIFICATION WORKFLOW
[0056] FIG. 1 is an exemplary flowchart describing an overall workflow 100 of cancer classification of a sample, according to one or more embodiments. The workflow 100 is by one or more entities, e.g., including a healthcare provider, a sequencing device, an analytics system, etc. Objectives of the workflow include detecting and / or monitoring cancer in subjects. From a healthcare standpoint, the workflow 100 can serve to supplement other existing diagnostic workups, e.g., cancer screening and early detection tools. The workflow 100 may serve to provide early cancer detection and / or routine cancer monitoring, minimal residual disease detection, prognosis, treatment prediction, or subtyping information to better inform treatment plans for subjects diagnosed with cancer. The overall workflow 100 may include additional / fewer steps than those shown in FIG. 1.
[0057] A healthcare provider performs sample collection 110. A subject to undergo screening, early detection or cancer classification, visits their healthcare provider. The healthcare provider collects the sample for performing cancer classification. Examples of biological samples include, but are not limited to: tissue biopsies, liquid biopsies (e.g., blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid) of the subject. In embodiments with liquid biopsies, the biological sample can be obtained from the individual in a less-invasive manner than tissue biopsies. The liquid biopsy sample may also be a small aliquot, e.g., up to 1 mL, up to 2 mL, up to 3 mL, up to 4 mL, or up to 5 mL. The sample includes genetic material belonging to the subject, which may be extracted and sequenced for screening, early detection, or cancer classification. Along with the sample, the healthcare provider may collect other information relating to the subject, e.g., biological sex, age, race, smoking status, other health metrics, any prior diagnoses, etc.
[0058] Once the sample is collected, the sample is provided to a laboratory process and sequencing device. A sequencing device performs sample sequencing 120. A lab clinician may perform one or more processing steps to the sample in preparation ofAtty. Docket No. 32917-65128 / WO 12 P0248-WOsequencing. Once prepared, the lab technician loads the sample in the sequencing device. An example of devices utilized in sequencing is further described in conjunction with FIGs.9A & 9B. The sequencing device generally extracts and isolates fragments of nucleic acid that are sequenced to determine a sequence of nucleobases corresponding to the fragments. Sequencing may also include amplification of nucleic material. Different sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing may be whole-genome sequencing or targeted sequencing with a target panel. In the context of DNA methylation, bisulfite sequencing (e.g., further described in FIGs. 2A & 2B) can determine methylation status through bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing 120 yields sequences for a plurality of nucleic acid fragments in the sample. In one or more embodiments, the sequences may include methylation state vectors, wherein each methylation state vector describes the methylation statuses for CpG sites on a fragment.
[0059] The sequencing data generated by the sequencing device is provided to an analytics system to perform analyses to generate predictions. An analytics system performs pre-analysis processing 130. An example analytics system is described in FIG. 9B. Preanalysis processing 130 may include, but not limited to, de-duplication of sequence reads, determining metrics relating to coverage, determining whether the sample is contaminated (including determining WBC contamination), removal of contaminated fragments, calling sequencing error, etc.
[0060] The analytics system performs one or more analyses 140 on the sequencing data. The analyses are statistical analyses or application of one or more trained models to predict at least a cancer status of the subject from whom the sample is derived. Different genetic features may be evaluated and considered, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), other types of genetic mutation, etc. The analyses 140 may include, anomalous methylation identification 142 (e.g., further described in FIGs. 7A & 7B), feature extraction 144 (e.g., further described FIGs. 8A & 8B), and cancer classification 146 to determine a cancer prediction (e.g., further described in FIGs. 8A & 8B), and sample swap detection 148 to detect whether sample labels have been swapped. Cancer classification generally entails inputting the extracted features to determine a cancer prediction. The cancer prediction may be a label or a value. The label may indicate a particular cancer state. A binary label may indicate presence or absence of a cancer signal; whereas a multiclass label can indicate one or more cancer signal origins from a plurality of potential cancer signal origins that are screened for. Cancer signal origins may be split basedAtty. Docket No. 32917-65128 / WO 13 P0248-WOon various characteristics of the cancer. For example, cancer signal origins may be split according to organ type and / or tumor biology type. In other examples, cancer signal origins, or other predictive information generated with the cancer prediction, may be further split according to progression (e.g., Stage I, II, III, or IV), prognosis, prediction of response to a candidate treatment, or presence of residual disease after or on treatment. A value may indicate a likelihood of a particular cancer state, e g., a likelihood of cancer, and / or a likelihood of a particular cancer signal origin.
[0061] The analytics system returns the prediction 150 to the healthcare provider. The healthcare provider may establish or adjust a treatment plan based on the cancer prediction. In some embodiments, the analytics system may return to the healthcare provider and / or the subject a diagnostic workup recommendation 160 identifying one or more diagnostic steps that might be used to confirm the test result (cancer prediction) and result in a cancer diagnosis. In such embodiments, the analytics system may store a database associating diagnostic workup options generally accepted by healthcare professionals to be useful diagnostic workup steps in the presence of a predicted cancer signal origin.Optimization of treatment is further described in Section V.D. Treatment. In some embodiments, the analytics system may leverage the cancer classification workflow for prognosis determination, treatment personalization, evaluation of treatment, monitoring cancer status, etc.LB. METHYLATION OVERVIEW
[0062] In accordance with the present description, cfDNA fragments from a subject are treated, for example by converting unmethylated cytosines to uracils, sequenced and the sequence reads compared to a reference genome to identify the methylation states at specific CpG sites within the DNA fragments. Each CpG site may be methylated or unmethylated. Identification of anomalously methylated fragments, in comparison to healthy subjects, may provide insight into a subject’s cancer status. As is well known in the art, DNA methylation changes (compared to healthy controls) can cause different effects, which may contribute to cancer. Various challenges arise in the identification of informatively methylated cfDNA fragments. First off, determining a DNA fragment to be informatively methylated can hold weight in comparison with a group of control subjects, such that if the control group is small in number, the determination loses confidence due to statistical variability within the smaller size of the control group. Additionally, among a group of control subjects, methylation status can vary which can be difficult to account for when determining a subject’s DNA fragments to be informatively methylated. On another note, methylation of a cytosine at a CpG site canAtty. Docket No. 32917-65128 / WO 14 P0248-WOcausally influence methylation at a subsequent CpG site. To encapsulate this dependency can be another challenge in itself.
[0063] Methylation can typically occur in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5 -methy lcytosine. In particular, methylation can occur at dinucleotides of cytosine and guanine referred to herein as ‘“CpG sites”. In other instances, methylation may occur at a cytosine not part of a CpG site or at another nucleotide that is not cytosine; however, these are rarer occurrences. In this present disclosure, methylation is discussed in reference to CpG sites for the sake of clarity7. Anomalous DNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a DNA fragment, if the DNA fragment comprises more than a threshold number of CpG sites wi th more than a threshold percentage of those CpG sites being methylated or unmethylated. In addition to hypermethylation or hypomethylation, an informative methylation status of a DNA fragment can further be characterized of a sequence of methylated and unmethylated CpGs that are not frequently observed in a healthy population and that is indicative for the presence of cancer or the presence of one or few cancer signal origins.
[0064] The principles described herein can be equally applicable for the detection of methylation in a non-CpG context, including non-cytosine methylation. In such embodiments, the wet laboratory7assay used to detect methylation may vary from those described herein. Further, the methylation state vectors discussed herein may contain elements that are generally sites where methylation has or has not occurred (even if those sites are not CpG sites specifically). With that substitution, the remainder of the processes described herein can be the same, and consequently the inventive concepts described herein can be applicable to those other forms of methylation.I.C. DEFINITIONS
[0065] The term “cell free nucleic acid” or “cfNA” refers to nucleic acid fragments that circulate in a subject’s body (e.g., blood) or are present in other bodily liquids like urine, cerebrovascular fluid, and originate from one or more healthy cells and / or from one or more unhealthy cells (e.g., cancer cells). The term “cell free DNA,” or “cfDNA” refers to deoxyribonucleic acid fragments present in a subject’s body outside of the cell (e.g., in blood plasma, urine, sputum, cerebrospinal fluid, etc.).
[0066] The term “genomic nucleic acid.” “genomic DNA,” or "gDNA” refers to nucleic acid molecules or deoxyribonucleic acid molecules obtained from one or more cells.Atty. Docket No. 32917-65128 / WO 15 P0248-WOIn various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or from tumor cells (e.g., a biopsy sample). In some embodiments, gDNA can be extracted from a cell derived from a blood cell lineage, such as a white blood cell.
[0067] The term "‘circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments that originate from tumor cells or other ty pes of cancer cells, and which may be released into a bodily fluid of a subject (e.g., blood, sweat, urine, or saliva) as result of biological processes such as apoptosis or necrosis of dying cells or actively released by viable tumor cells.
[0068] The term “DNA fragment,” “fragment,” or “DNA molecule” may generally refer to any deoxyribonucleic acid fragments, i.e., cfDNA, gDNA, ctDNA, etc.
[0069] The term “informative fragment.” may generally refer to any fragment having an informative genetic signature. The term “anomalously methylated fragment,” or “fragment with an anomalous methylation pattern” refers to an informative fragment that has an informative methylation pattern. Anomalous methylation of a fragment may be determined using probabilistic models to identify unexpectedness of observing a fragment’s methylation pattern in a control group.
[0070] As used herein, the term “about” or “approximately” can mean within an acceptable error range for the particular value as determined by one of ordinary' skill in the art, which can depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. “About” can mean a range of ±20%, ±10%, ±5%, or ±1% of a given value. The term “about” or “approximately” can mean within an order of magnitude, within 5-fold, or within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed. The term “about” can have the meaning as commonly understood by one of ordinary' skill in the art. The term “about” can refer to ±10%. The term “about” can refer to ±5%.
[0071] As used herein, the term “biological sample,” “patient sample,” or “sample” refers to any sample taken from a subject, which can reflect a biological state associated with the subject, and that includes gDNA or cell-free DNA. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A biological sample can be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testis), vaginal flushing fluids, pleural fluid, pericardial fluid, peritoneal fluid, ascitic fluid, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, discharge fluid from theAtty. Docket No. 32917-65128 / WO 16 P0248-WOnipple, aspiration fluid from different parts of the body (e.g., thyroid, breast), etc. A biological sample can be a stool sample. A biological sample can include any tissue or material derived from a living or dead subject. A biological sample can be a cell-free sample. A biological sample can comprise a nucleic acid (e.g., DNA or RNA) or a fragment thereof.
[0072] As used herein, the terms “control,” “control sample,” “reference,” “reference sample,” “normal,” and “normal sample” describe a sample from a subject that does not have a particular condition, or is otherwise healthy. In an example, a method as disclosed herein can be performed on a subject having a tumor, where the reference sample is a sample taken from a healthy tissue of the subject. A reference sample can be obtained from the subject, or from a database. The reference can be, e.g., a reference genome that is used to map nucleic acid fragment sequences obtained from sequencing a sample from the subject. A reference genome can refer to a haploid or diploid genome to which nucleic acid fragment sequences from the biological sample and a constitutional sample can be aligned and compared. An example of a constitutional sample can be DNA of white blood cells obtained from the subject. For a haploid genome, there can be only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified; each heterozygous locus can have two alleles, where either allele can allow a match for alignment to the locus.
[0073] As used herein, the term “cancer” or “tumor” refers to an abnormal mass of tissue in which the growth of the mass surpasses and is not coordinated with the growth of normal tissue.
[0074] As used herein, the term “genomic” refers to a characteristic of the genome of an organism. Examples of genomic characteristics include, but are not limited to, those relating to the primary nucleic acid sequence of all or a portion of the genome (e.g., the presence or absence of a nucleotide polymorphism, indel. sequence rearrangement, mutational frequency, etc.), the copy number of one or more particular nucleotide sequences within the genome (e.g., copy number, allele frequency fractions, single chromosome or entire genome ploidy, etc.), the epigenetic status of all or a portion of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome positioning, etc.), the expression profile of the organism’s genome (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).
[0075] As used herein, the phrase “healthy,” refers to a subject possessing good health. A healthy subject can demonstrate an absence of any malignant or non-malignant disease. A “healthy subject” can have other diseases or conditions, unrelated to the condition being assayed, which can normally not be considered “healthy.”Atty. Docket No. 32917-65128 / WO 17 P0248-WO
[0076] As used herein, the term ‘‘methylation” refers to a modification of deoxyribonucleic acid (DNA) where a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation tends to occur at dinucleotides of cytosine and guanine referred to herein as “CpG sites.” In other instances, methylation may occur at a cytosine not part of a CpG site or at another nucleotide that’s not cytosine; however, these are rarer occurrences. Anomalous cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of cancer status. DNA methylation anomalies (compared to healthy controls) can cause different effects, which may contribute to cancer. The principles described herein are equally applicable for the detection of methylation in a CpG context and non-CpG context, including non-cytosine methylation. Further, the methylation state vectors may contain elements that are generally vectors of sites where methylation has or has not occurred (even if those sites are not CpG sites specifically).
[0077] As used interchangeably herein, the term “methylation fragment” or “nucleic acid methylation fragment” refers to a sequence of methylation states for each CpG site in a plurality of CpG sites, determined by a methylation sequencing of nucleic acids (e.g., a nucleic acid molecule and / or a nucleic acid fragment). In a methylation fragment, a location and methylation state for each CpG site in the nucleic acid fragment is determined based on the alignment of the sequence reads (e.g., obtained from sequencing of the nucleic acids) to a reference genome. A nucleic acid methylation fragment comprises a methylation state of each CpG site in a plurality of CpG sites (e.g., a methylation state vector), which specifies the location of the nucleic acid fragment in a reference genome (e.g., as specified by the position of the first CpG site in the nucleic acid fragment using a CpG index, or another similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of a sequence read to a reference genome, based on a methylation sequencing of a nucleic acid molecule, can be performed using a CpG index. As used herein, the term “CpG index” refers to a list of each CpG site in the plurality' of CpG sites (e.g., CpG 1, CpG 2, CpG 3, etc.) in a reference genome, such as a human reference genome, yvhich can be in electronic format. The CpG index further comprises a corresponding genomic location, in the corresponding reference genome, for each respective CpG site in the CpG index. Each CpG site in each respective nucleic acid methylation fragment is thus indexed to a specific location in the respective reference genome, which can be determined using the CpG index.
[0078] The term “nucleic acid” can refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) or any hybrid or fragment thereof. The nucleic acid in the sample can be a cell-Atty. Docket No. 32917-65128 / WO 18 P0248-WOfree nucleic acid. In various embodiments, the majority of DNA in a biological sample that has been enriched for cell-free DNA (e.g.. a plasma sample obtained via a centrifugation protocol) can be cell-free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free). A biological sample can be treated to physically disrupt tissue or cell structure (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution which can further contain enzymes, buffers, salts, detergents, and the like which can be used to prepare the sample for analysis.
[0079] As used herein, the term “true positive’’ (TP) refers to a subject having a condition. “True positive” can refer to a subject that has a tumor, a cancer, a pre-cancerous condition (e.g., a pre-cancerous lesion), a localized or a metastasized cancer, or a non-malignant disease. “True positive” can refer to a subject having a condition and is identified as having the condition by an assay or method of the present disclosure, like having a cancer that will respond to a treatment, having residual disease after or during a treatment, or being likely to encounter an event like disease progression, relapse, or even cancer-related death within a pre-specified time frame. As used herein, the term “true negative” (TN) refers to a subject that does not have a condition or does not have a detectable condition. True negative can refer to a subject that does not have a disease or a detectable disease, such as a tumor, a cancer, a pre-cancerous condition (e.g., a pre-cancerous lesion), a localized or a metastasized cancer, anon-malignant disease, or a subject that is otherwise healthy. True negative can refer to a subject that does not have a condition or does not have a detectable condition, or is identified as not having the condition by an assay or method of the present disclosure.
[0080] As used herein, the term “reference genome” refers to any particular known, sequenced or characterized genome, whether partial or complete, of any organism or virus that may be used to reference identified sequences from a subject. Exemplary reference genomes used for human subjects as well as many other organisms are provided in the online genome browser hosted by the National Center for Biotechnology Information (“NCBI”) or the University of California, Santa Cruz (UCSC). A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome often is an assembled or partially assembled genomic sequence from a subject or multiple subjects. In some embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more huma subjects. The reference genome can be viewed as a representative example of a species' set of genes. In some embodiments, a reference genome comprises sequences assigned to chromosomes. Exemplary7human reference genomes include but are not limitedAtty. Docket No. 32917-65128 / WO 19 P0248-WOto NCBI build 34 (UCSC equivalent: hgl6), NCBI build 35 (UCSC equivalent: hgl7), NCBI build 36.1 (UCSC equivalent: hgl8), GRCh37 (UCSC equivalent: hgl9). and GRCh38 (UCSC equivalent: hg38).
[0081] As used herein, the term “sequence reads” or “reads” refers to nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), and sometimes are generated from both ends of nucleic acids (e.g., paired-end reads, double-end reads). In some embodiments, sequence reads (e.g., single-end or paired-end reads) can be generated from one or both strands of a targeted nucleic acid fragment. The length of the sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary’ in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of a mean, median or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp. about 95 bp, about 100 bp, about 110 bp, about 120 bp. about 130, about 140 bp, about 150 bp. about 200 bp, about 450 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In some embodiments, the sequence reads are of a mean, median or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. Nanopore sequencing, for example, can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that do not vary as much, for example, most of the sequence reads can be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g.. a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g.. about 20 to about 150) from part of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. A sequence read can be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.
[0082] As used herein, the terms “sequencing” and the like as used herein refers generally to any and all biochemical processes that may be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, sequencing dataAtty. Docket No. 32917-65128 / WO 20 P0248-WOcan include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.
[0083] As used herein, the term “sequencing depth,” is interchangeably used with the term “coverage” and refers to the number of times a locus is covered by a consensus sequence read corresponding to a unique nucleic acid target molecule aligned to the locus; e.g., the sequencing depth is equal to the number of unique nucleic acid target molecules covering the locus. The locus can be as small as a nucleotide, or as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as “Yx”, e.g., 50x, lOOx, etc., where “Y” refers to the number of times a locus is covered with a sequence corresponding to a nucleic acid target, e g., the number of times independent sequence information is obtained covering the particular locus. In some embodiments, the sequencing depth corresponds to the number of genomes that have been sequenced. Sequencing depth can also be applied to multiple loci, or the whole genome, in which case Y can refer to the mean or average number of times a locus or a haploid genome, or a whole genome, respectively, is sequenced. When a mean depth is quoted, the actual depth for different loci included in the dataset can span over a range of values. Ultra-deep sequencing can refer to at least lOOx in sequencing depth at a locus.
[0084] As used herein, the term “sensitivity ” or “true positive rate” (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify a proportion of the population that truly has a condition. For example, sensitivity' can characterize the ability' of a method to correctly identify the number of subjects within a population having cancer. In another example, sensitivity can characterize the ability' of a method to correctly identify the one or more markers indicative of cancer.
[0085] As used herein, the term “specificity” or “true negative rate” (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity' can characterize the ability' of an assay or method to correctly identify a proportion of the population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population not having cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers indicative of cancer.
[0086] As used herein, the term “subject” refers to any living or non-living organism, including but not limited to a human (e.g.. a male human, female human, fetus, pregnant female, child, or the like), a non-human animal, a plant, a bacterium, a fungus or a protist.Atty. Docket No. 32917-65128 / WO 21 P0248-WOAny human or non-human animal can serve as a subject, including but not limited to mammal, reptile, avian, amphibian, fish, ungulate, ruminant, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), swine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry', dog, cat, mouse, rat, fish, dolphin, whale, and shark. In some embodiments, a subject is a male or female of any stage (e.g., a man, a woman or a child). A subject from whom a sample is taken, or is treated by any of the methods or compositions described herein can be of any age and can be an adult, infant or child.
[0087] As used herein, the term “tissue” can correspond to a group of cells that group together as a functional unit. More than one ty pe of cell can be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells or blood cells), but also can correspond to tissue from different organisms (mother vs. fetus) or to healthy cells vs. tumor cells. The term “tissue” can generally refer to any group of cells found in the human body (e.g., heart tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some aspects, the term “tissue” or “tissue type” can be used to refer to a tissue from which a cell-free nucleic acid originates. In one example, viral nucleic acid fragments can be derived from blood tissue. In another example, viral nucleic acid fragments can be derived from tumor tissue.
[0088] The terminology used herein is for the purpose of describing particular cases only and is not intended to be limiting. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”I.D. EXAMPLE ANALYTICS SYSTEM
[0089] FIG. 9A is an exemplary flowchart of devices for sequencing nucleic acid samples according to one or more embodiments. This illustrative flowchart includes devices such as a sequencer 920 and an analytics system 900. The sequencer 920 and the analytics system 900 may work in tandem to perform one or more steps in the processes.
[0090] In various embodiments, the sequencer 920 receives an enriched nucleic acid sample 910. As shown in FIG. 9A, the sequencer 920 can include a graphical user interface 925 that enables user interactions with particular tasks (e.g., initiate sequencing or terminate sequencing) as well as one more loading stations 930 for loading a sequencing cartridge including the enriched fragment samples and / or for loading necessary' buffers for performingAtty. Docket No. 32917-65128 / WO 22 P0248-WOthe sequencing assays. Therefore, once a user of the sequencer 920 has provided the necessary’ reagents and sequencing cartridge to the loading station 930 of the sequencer 920, the user can initiate sequencing by interacting with the graphical user interface 925 of the sequencer 920. Once initiated, the sequencer 920 performs the sequencing and outputs the sequence reads of the enriched fragments from the nucleic acid sample 910.
[0091] In some embodiments, the sequencer 920 is communicatively coupled with the analytics system 900. The analytics system 900 includes some number of computing devices used for processing the sequence reads for various applications such as assessing methylation status at one or more CpG sites, variant calling or quality control. The sequencer 920 may provide the sequence reads in a BAM file format to the analytics system 900. The analytics system 900 can be communicatively coupled to the sequencer 920 through a wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analytics system 900 is configured with a processor and non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process the sequence reads or to perform one or more steps of any of the methods or processes disclosed herein.
[0092] In some embodiments, the sequence reads may be aligned to a reference genome using known methods in the art to determine alignment position information.Alignment position may generally describe a beginning position and an end position of a region in the reference genome that corresponds to a beginning nucleotide base and an end nucleotide base of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate a first CpG site and a last CpG site included in the sequence read according to the alignment to the reference genome. The alignment position information may further indicate methylation statuses and locations of all CpG sites in a given sequence read. A region in the reference genome may be associated with a gene or a segment of a gene; as such, the analytics system 900 may label a sequence read with one or more genes that align to the sequence read. In one embodiment, fragment length (or size) is determined from the beginning and end positions.
[0093] In various embodiments, for example when a paired-end sequencing process is used, a sequence read is comprised of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from the first end of a double-stranded DNA (dsDNA) molecule w hereas the second read R_2 may be sequenced from the second end of the doublestranded DNA (dsDNA). Therefore, nucleotide base pairs of the first read R 1 and second read R_2 may be aligned consistently (e.g., in opposite orientations) with nucleotide bases ofAtty. Docket No. 32917-65128 / WO 23 P0248-WOthe reference genome. Alignment position information derived from the read pair R_1 and R_2 may include a beginning position in the reference genome that corresponds to an end of a first read (e.g., R_l) and an end position in the reference genome that corresponds to an end of a second read (e.g., R_2). In other words, the beginning position and end position in the reference genome can represent the likely location within the reference genome to which the nucleic acid fragment corresponds. An output file having SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis.
[0094] Referring now to FIG. 9B, FIG. 9B is a block diagram of an analytics system 900 for processing DNA samples according to one embodiment. The analytics system implements one or more computing devices for use in analyzing DNA samples. The analytics system 900 includes a sequence processor 940, sequence database 945, model database 955, models 950, parameter database 965, and score engine 960. In some embodiments, the analytics system 900 performs some or all of the processes described throughout this disclosure.
[0095] The sequence processor 940 generates methylation state vectors for fragments from a sample. At each CpG site on a fragment, the sequence processor 940 generates a methylation state vector for each fragment specifying a location of the fragment in the reference genome, a number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment whether methylated, unmethylated, or indeterminate via the process 200 of FIG. 2A. The sequence processor 940 may store methylation state vectors for fragments in the sequence database 945. Data in the sequence database 945 may be organized such that the methylation state vectors from a sample are associated to one another.
[0096] Further, multiple different models 950 may be stored in the model database 955 or retrieved for use with test samples. In one example, a model is a trained cancer classifier for determining a cancer prediction for a test sample using a feature vector derived from informative fragments. The training and use of the cancer classifier will be further discussed in conjunction with Section III. Cancer Classifier for Determining Cancer. The analytics system 900 may train the one or more models 950 and store various trained parameters in the parameter database 965. The analytics system 900 stores the models 950 along with functions in the model database 955.
[0097] During inference, the score engine 960 uses the one or more models 950 to return outputs. The score engine 960 accesses the models 950 in the model database 955 along with trained parameters from the parameter database 965. According to each model, the score engine receives an appropriate input for the model and calculates an output basedAtty. Docket No. 32917-65128 / WO 24 P0248-WOon the received input, the parameters, and a function of each model relating the input and the output. In some use cases, the score engine 960 further calculates metrics correlating to a confidence in the calculated outputs from the model. In other use cases, the score engine 960 calculates other intermediary values for use in the model.II. SAMPLE SEQUENCING & PROCESSINGII. A. GENERATING METHYLATION STATE VECTORS FOR DNA FRAGMENTS
[0098] FIG. 2A is an exemplary flowchart describing a process 200 of sequencing a fragment of cfDNA to obtain a methylation state vector, according to one or more embodiments. In order to analyze DNA methylation, an analytics system first obtains 210 a sample from a subject comprising a plurality of cfDNA molecules. In additional embodiments, the process 200 may be applied to sequence other types of DNA molecules. The process 200 is an embodiment of sample sequencing 120 of FIG. 1.
[0099] From the sample, the analytics system can isolate 210 each cfDNA molecule. The cfDNA molecules can be treated 220 to convert unmethylated cytosines to uracils. In one embodiment, the method uses a bisulfite treatment of the DNA which converts the unmethylated cytosines to uracils without converting the methylated cytosines. For example, a commercial kit such as the EZ DNA Methylation™ - Gold, EZ DNA Methylation™ -Direct or an EZ DNA Methylation™ - Lightning kit (available from Zymo Research Corp (Irvine. CA)) is used for the bisulfite conversion. In another embodiment, the conversion of unmethylated cytosines to uracils is accomplished using an enzymatic reaction. For example, the conversion can use a commercially available kit for conversion of unmethylated cytosines to uracils, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0100] From the converted cfDNA molecules, a sequencing library can be prepared 230. During library’ preparation, molecular identifiers, including but not limited to unique molecular identifiers (UMI) or sample barcodes, can be added to the nucleic acid molecules (e.g., DNA molecules) through adapter ligation. The molecular identifiers can be short nucleic acid sequences (e.g., 4-10 base pairs) that are added to ends of DNA fragments (e.g, DNA molecules fragmented by physical shearing, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. Molecular identifiers can serve as a tag to identify sequence reads originating from a single sample or subject. UMIs represent a special case as degenerate base pairs that serve as a unique tag that can be used to identity’ sequence reads originating from a specific DNA fragment. During PCR amplification following adapter ligation, the molecular identifiers can be replicated along with the attached DNA fragment.Atty. Docket No. 32917-65128 / WO 25 P0248-WOThis can provide a way to identify sequence reads that came from the same original fragment in downstream analysis.
[0101] Optionally, the sequencing library may be enriched 235 for cfDNA molecules, or genomic regions, that are informative for cancer status using a plurality of hybridization probes. The hybridization probes are short oligonucleotides capable of hybridizing to particularly specified cfDNA molecules, or targeted regions, and enriching for those fragments or regions for subsequent sequencing and analysis. Hybridization probes may be used to perform a targeted, high-depth analysis of a set of specified CpG sites of interest to the researcher. Hybridization probes can be tiled across one or more target sequences at a coverage of IX, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more than 10X. For example, hybridization probes tiled at a coverage of 2X comprises overlapping probes such that each portion of the target sequence is hybridized to 2 independent probes. Hybridization probes can be tiled across one or more target sequences at a coverage of less than IX.
[0102] In one embodiment, the hybridization probes are designed to enrich for DNA molecules that have been treated (e.g.. using bisulfite) for conversion of unmethylated cytosines to uracils. During enrichment, hybridization probes (also referred to herein as “probes”) can be used to target and pull-down nucleic acid fragments informative for the presence or absence of cancer (or disease), cancer status, or a cancer classification (e.g., cancer class or tissue of origin). The probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA. The target strand may be the “positive” strand (e.g.. the strand transcribed into mRNA, and subsequently translated into a protein) or the complementary' “negative” strand. The probes may range in length from 10s, 100s, or 1000s of base pairs. The probes can be designed based on a methylation site panel. The probes can be designed based on a panel of targeted genes to analyze particular mutations or target regions of the genome (e.g., of the human or another organism) that are suspected to correspond to certain cancers or other ty pes of diseases. Moreover, the probes may cover overlapping portions of a target region.
[0103] Once prepared, the sequencing library or a portion thereof can be sequenced 240 to obtain a plurality of sequence reads. The sequence reads may be in a computer-readable, digital format for processing and interpretation by computer software. The sequence reads may be aligned to a reference genome to determine alignment position information. The alignment position information may indicate a beginning position and an end position of a region in the reference genome that corresponds to a beginning nucleotide base and end nucleotide base of a given sequence read. Alignment position information may also includeAtty. Docket No. 32917-65128 / WO 26 P0248-WOsequence read length, which can be determined from the beginning position and end position. A region in the reference genome may be associated with a gene or a segment of a gene. A sequence read can be comprised of a read pair denoted as Rrand R2. For example, the first read may be sequenced from a first end of a nucleic acid fragment whereas the second read R2may be sequenced from the second end of the nucleic acid fragment. Therefore, nucleotide base pairs of the first read R^ and second read R2may be aligned consistently (e.g., in opposite orientations) with nucleotide bases of the reference genome. Alignment position information derived from the read pair R and R2may include a beginning position in the reference genome that corresponds to an end of a first read (e.g., R- and an end position in the reference genome that corresponds to an end of a second read (e.g., R2). In other words, the beginning position and end position in the reference genome can represent the likely location within the reference genome to which the nucleic acid fragment corresponds. An output file having SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis such as methylation state determination.
[0104] From the sequence reads, the analytics system determines 250 a location and methylation state for each CpG site based on alignment to a reference genome. The analytics system generates 260 a methylation state vector for each fragment specifying a location of the fragment in the reference genome (e.g., as specified by the position of the first CpG site in each fragment, or another similar metric), a number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment whether methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or indeterminate (e.g., denoted as I). Observed states can be states of methylated and unmethylated, whereas an unobserved state is indeterminate. Indeterminate methylation states may originate from sequencing errors and / or disagreements between methylation states of a DNA fragment's complementary strands. The methylation state vectors may be stored in temporary or persistent computer memory for later use and processing. Further, the analytics system may remove duplicate reads or duplicate methylation state vectors from a single sample. The analytics system may determine that a certain fragment with one or more CpG sites has an indeterminate methylation status over a threshold number or percentage, and may exclude such fragments or selectively include such fragments but build a model accounting for such indeterminate methylation statuses.
[0105] FIG. 2B is an exemplary illustration of the process 200 of FIG. 2A of sequencing a cfDNA molecule to obtain a methylation state vector, according to one or moreAtty. Docket No. 32917-65128 / WO 27 P0248-WOembodiments. As an example, the analytics system receives a cfDNA molecule 212 that, in this example, contains three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 212 are methylated 214. During the treatment step 220, the cfDNA molecule 212 is converted to generate a converted cfDNA molecule 222. During the treatment 220, the second CpG site which was unmethylated has its cytosine converted to uracil. However, the first and third CpG sites were not converted.
[0106] After conversion, a sequencing library 230 is prepared and sequenced 240 to generate a sequence read 242. The analytics system aligns 250 the sequence read 242 to a reference genome 244. The reference genome 244 provides the context as to what position in a human genome the fragment cfDNA originates from. In this simplified example, the analytics system aligns 250 the sequence read 242 such that the three CpG sites correlate to CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of description). The analytics system can thus generate information both on methylation status of all CpG sites on the cfDNA molecule 212 and the position in the human genome that the CpG sites map to. As shown, the CpG sites on sequence read 242 which are methylated are read as cytosines. In this example, the cytosines appear in the sequence read 242 only in the first and third CpG site which allows one to infer that the first and third CpG sites in the original cfDNA molecule are methylated. Whereas, the second CpG site can be read as a thymine (U is converted to T during the sequencing process), and thus, one can infer that the second CpG site is unmethylated in the original cfDNA molecule. With these two pieces of information, the methylation status and location, the analytics system generates 260 a methylation state vector 252 for the fragment cfDNA 212. In this example, the resulting methylation state vector 252 is < M23, U24, M25 >, wherein M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript number corresponds to a position of each CpG site in the reference genome.
[0107] One or more alternative sequencing methods can be used for obtaining sequence reads from nucleic acids in a biological sample. The one or more sequencing methods can comprise any form of sequencing that can be used to obtain a number of sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including, but not limited to, high-throughput sequencing systems such as the Roche 454 platform, the Applied Biosystems SOLID platform, the Helicos True Single Molecule DNA sequencing technology7, the sequencing-by -hybridization platform from Affymetrix Inc., the single-molecule, realtime (SMRT) technology7of Pacific Biosciences, the sequencing-by-synthesis platforms from 454 Life Sciences, Illumina / Solexa and Helicos Biosciences, and the sequencing-by-ligationAtty. Docket No. 32917-65128 / WO 28 P0248-WOplatform from Applied Biosystems. The ION TORRENT technology from Life technologies and Nanopore sequencing can also be used to obtain sequence reads from the nucleic acids (e.g, cell-free nucleic acids) in the biological sample. Sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina’s Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 4500 (Illumina, San Diego Calif.)) can be used to obtain sequence reads from the cell-free nucleic acid obtained from a biological sample of a training subject in order to form the genotypic dataset. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one example of this type of sequencing technology7, a flow cell is used that contains an optically transparent slide with eight subject lanes on the surfaces of which are bound oligonucleotide anchors (e.g., adapter primers). A cell-free nucleic acid sample can include a signal or tag that facilitates detection. The acquisition of sequence reads from the cell-free nucleic acid obtained from the biological sample can include obtaining quantification information of the signal or tag via a variety7of techniques such as, for example, flow cytometry7, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene-chip analysis, microarray, mass spectrometry7, cytofluorimetric analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry7, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and combination thereof.
[0108] The one or more sequencing methods can comprise a whole-genome sequencing assay. A whole-genome sequencing assay can comprise a physical assay that generates sequence reads for a whole genome or a substantial portion of the whole genome which can be used to determine large variations such as copy number variations or copy number aberrations. Such a physical assay may employ whole-genome sequencing techniques or whole-exome sequencing techniques. A whole-genome sequencing assay can have an average sequencing depth of at least lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, at least 20x, at least 30x, or at least 40x across the genome of the test subject. In some embodiments, the sequencing depth is about 30,000x. The one or more sequencing methods can comprise a targeted panel sequencing assay. A targeted panel sequencing assay can have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60.000x, or at least 70,000x sequencing depth for the targeted panel of genes. The targeted panel of genes can comprise between 450 and 500 genes. The targeted panel of genes can comprise a range of 500±5 genes, a range of 500±10 genes, or a range of 500±25 genes.
[0109] The one or more sequencing methods can comprise paired-end sequencing. The one or more sequencing methods can generate a plurality of sequence reads. TheAtty. Docket No. 32917-65128 / WO 29 P0248-WOplurality of sequence reads can have an average length ranging between 10 and 700, between 50 and 400, or between 100 and 300. The one or more sequencing methods can comprise a methylation sequencing assay. The methylation sequencing can be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using a plurality of nucleic acid probes. For example, the methylation sequencing is whole-genome bisulfite sequencing (e.g. , WGBS). The methylation sequencing can be a targeted DNA methylation sequencing using a plurality of nucleic acid probes targeting the most informative regions of the methylome, a unique methylation database and prior prototype whole-genome and targeted sequencing assays.
[0110] The methylation sequencing can detect one or more 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) in respective nucleic acid methylation fragments. The methylation sequencing can comprise conversion of one or more unmethylated cytosines or one or more methylated cytosines, in respective nucleic acid methylation fragments, to a corresponding one or more uracils. The one or more uracils can be detected during the methylation sequencing as one or more corresponding thymines. The conversion of one or more unmethylated cytosines or one or more methylated cytosines can comprise a chemical conversion, an enzymatic conversion, or combinations thereof.[OHl] For example, bisulfite conversion involves converting cytosine to uracil while leaving methylated cytosines (e.g., 5-methylcytosine or 5-mC) intact. In some DNA, about 95% of cytosines may not be methylated in the DNA. and the resulting DNA fragments may include many uracils which are represented by thymines. Enzymatic conversion processes may be used to treat the nucleic acids prior to sequencing, which can be performed in various ways. One example of a bisulfite-free conversion comprises a bisulfite-free and baseresolution sequencing method. TET-assisted pyridine borane sequencing (TAPS), for nondestructive and direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosines. The methylation state of a CpG site in the corresponding plurality of CpG sites in the respective nucleic acid methylation fragment can be methylated when the CpG site is determined by the methylation sequencing to be methylated, and unmethylated when the CpG site is determined by the methylation sequencing to not be methylated.
[0112] A methylation sequencing assay (e.g. , WGBS and / or targeted methylation sequencing) can have an average sequencing depth including but not limited to up to about l,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. The methylation sequencing can have a sequencing depth that is greater than 30,000x, e.g., at least 40,000x orAtty. Docket No. 32917-65128 / WO 30 P0248-WO50,000x. A whole-genome bisulfite sequencing method can have an average sequencing depth of between 20x and 50x, and a targeted methylation sequencing method has an average effective depth of between lOOx and lOOOx, where effective depth can be the equivalent whole-genome bisulfite sequencing coverage for obtaining the same number of sequence reads obtained by targeted methylation sequencing. The methylation sequencing assay may comprise probes targeting 500 or more CpG sites, 1,000 or more CpG sites, 1,500 or more CpG sites, 2,000 or more CpG sites, 2,500 or more CpG sites, 3,000 or more CpG sites, 3,500 or more CpG sites, 4,000 or more CpG sites, 4,500 or more CpG sites, 5,000 or more CpG ,000 or more CpG sites, sites, 7,000 or more CpG sites, 8,000 or more CpG sites, 9,000 or more CpG sites, 10,000 or more CpG sites, 20,000 or more CpG sites, 30,000 or more CpG sites, 40,000 or more CpG sites, 50,000 or more CpG sites, 60,000 or more CpG sites, 70,000 or more CpG sites, 80,000 or more CpG sites, 90,000 or more CpG sites, 100,000 or more CpG sites, 150,000 or more CpG sites, etc.
[0113] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see. e.g., United States Patent Application No. 16 / 719,902, entitled “Systems and Methods for Estimating Cell Source Fractions Using Methylation Information,” filed December 18, 2019, each of which is hereby incorporated by reference. Other methods for methylation sequencing, including those disclosed herein and / or any modifications, substitutions, or combinations thereof, can be used to obtain fragment methylation patterns. A methylation sequencing can be used to identify one or more methylation state vectors, as described, for example, in United States Patent Application No.16 / 352,602, entitled “Anomalous Fragment Detection and Classification,” filed March 13, 2019, or in accordance with any of the techniques disclosed in United States Patent Application No. 15 / 931,022. entitled “Model-Based Featurization and Classification,” filed May 13, 2020, each of which is hereby incorporated by reference.
[0114] The methylation sequencing of nucleic acids and the resulting one or more methylation state vectors can be used to obtain a plurality of nucleic acid methylation fragments. Each corresponding plurality of nucleic acid methylation fragments (e.g., for each respective genotypic dataset) can comprise more than 100 nucleic acid methylation fragments. An average number of nucleic acid methylation fragments can comprise 1,000 or more nucleic acid methylation fragments, 5,000 or more nucleic acid methylation fragments, 10,000 or more nucleic acid methylation fragments, 20,000 or more nucleic acid methylation fragments, 30,000 or more nucleic acid methylation fragments, 40,000 or more nucleic acid methylation fragments, or 50,000 or more nucleic acid methylation fragments. An averageAtty. Docket No. 32917-65128 / WO 31 P0248-WOnumber of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can be between 10,000 nucleic acid methylation fragments and 50,000 nucleic acid methylation fragments. The corresponding plurality of nucleic acid methylation fragments can comprise one thousand or more, ten thousand or more, 100 thousand or more, one million or more, ten million or more, 100 million or more, 500 million or more, one billion or more, two billion or more, three billion or more, four billion or more, five billion or more, six billion or more, seven billion or more, eight billion or more, nine billion or more, or 10 billion or more nucleic acid methylation fragments. An average length of a corresponding plurality of nucleic acid methylation fragments can be between 140 and 480 nucleotides. An average length of a corresponding plurality' of nucleic acid methylation fragments can be between 100 and 600 nucleotides. An average length of a corresponding plurality of nucleic acid methylation fragments can be between 50 and 750 nucleotides. An average length of a corresponding plurality of nucleic acid methylation fragments can be between 30 and 1,000 nucleotides.
[0115] Further details regarding methods for sequencing nucleic acids and methylation sequencing data are disclosed in U.S. Patent Application No. 17 / 191,914, titled “Systems and Methods for Cancer Condition Determination Using Autoencoders,” filed March 4, 2021, which is hereby incorporated herein by reference in its entirety'.III. MULTIMODAL SAMPLE SWAP DETECTION & REMEDIATION
[0116] Multimodal sample swap detection leverages a plurality of models to generate a concordant prediction on whether a sample is mislabeled, i.e., a sample swap event. Each mode may evaluate different markers in the sequencing data, different types of sequencing data, different labels for clinical information, other related information, or some combination thereof.
[0117] The analytics system may leverage a plurality of sample swap detection models, each configured in a different mode. Each sample swap detection model is configured to input, or receive as input, features of the sequencing data and, optionally features of the clinical information collected for the individual, and to output a prediction on whether a given sample is mislabeled. In one or more embodiments, one sample swap detection model is configured to predict a covariate label based on features of the sequencing data. The sample swap detection model may generate the prediction on whether the given sample is mislabeled by comparing the predicted covariate label to a reported covariate label.Atty. Docket No. 32917-65128 / WO 32 P0248-WO
[0118] In one or more embodiments, one sample swap detection model is configured to identify genomic similarity between two samples based on features of each sample’s sequencing data. The genomic similarity can be represented by a numerical value. In some embodiments, the sample swap detection model applies the genomic similarity model pairwise to samples purported to originate, i.e., labeled as originating, from the same subject. The analytics system can compute an aggregate similarity score from the pairwise similarities, e.g., an average, a mode, a median, some other statistical measure, etc. In some instances, samples from one purported individual may include multiple samples from two or more different subjects. In such embodiments, the analytics system may use a sorting algorithm that groups samples together based on their genomic similarity, to maximize similarity across samples in one group.
[0119] The analytics system further leverages an aggregation model that aggregates the predictions to provide an aggregate prediction on calling a sample swap event. The advantage of leveraging an aggregation model comes by compounding the sample swap detectability of each individual sample swap detection model. For example, a sample swap detection model configured to identify disparity between a predicted biological sex based on the sequencing data and a reported biological sex from clinical information has a middling true positive rate (theoretically around 50%), given that half of the true positives (i.e., actual sample swap events) would go undetected if the swapped label is with another of the same biological sex. Similarly, with another sample swap detection model based on evaluating ancestral origin against a reported race or ethnicity. However, akin to the biological-sex based detection mode, the ancestry' based detection mode is limited in that the swapped label may be associated with another of the same ancestral population class. The aggregation model improves upon the individual components by aggregating the predictions output by the multimodal sample swap detection models into one aggregate prediction of a sample swap event, compounding the detectability from the disparate models. The aggregated prediction has a confidence that is greater than the individual confidences of the constituent predictions.
[0120] Upon detecting a sample swap event, the analytics system may perform a remedial workflow. The remedial workflows may include computational steps performed by the analytics system, instructions for physical steps to be performed by a clinician, or some combination thereof. The remedial workflow may include, e.g., physically discarding a sample that is identified as being swapped, or otherwise labeling the sample as unusable. However, such remedial measure is non-optimal, as the sample that is already collected and sequenced would be wasted and a replacement sample may need to be identified or procured.Atty. Docket No. 32917-65128 / WO 33 P0248-WO
[0121] In one or more embodiments, the analytics system may leverage a similarity' model configured to input features of sequencing data from two different samples to output a genetic similarity' score between the two samples, e.g., for the purpose of reversing a sample swap. So, for example, Individual A provided two biological samples BS1 and BS2. BS1 is subject to a sample swap event, i.e., BS1 is mislabeled as belonging to Individual B. The similarity model inputs features of the sequencing data from BS1 and features of the sequencing data from BS2 to output a genetic similarity score between BS1 and BS2. Given that these two samples truly originated from Individual A, the analytics system may leverage this model to identify that the genetic similarity' score between these two samples is high (i.e., satisfying a threshold specified for genetic similarity ), and thus BS1 (though labeled as originating from Individual B) is more likely to have originated from Individual A.III. A. AGGREGATION MODEL
[0122] In one or more embodiments, the analytics system leverages an aggregation model to aggregate predictions by multimodal sample swap detection models. The aggregation model generates the aggregate prediction based on the individual predictions output by the sample swap detection models. In some embodiments, the aggregation model computes a weighted average, weighting predictions differentially, e.g., based on confidence of each prediction. In some embodiments, the aggregation model takes the most restrictive prediction as the aggregate prediction. For example, if one prediction indicates sample swap, then the aggregate prediction also deems there to be a sample swap.
[0123] FIG. 3 is an exemplary flowchart describing a process 300 of multimodal sample swap detection, according to one or more embodiments. The analytics system is described as performing the process 300 of multimodal sample swap detection. In other embodiments, other systems may perform some of the steps. In other embodiments, the process 300 may include additional, fewer, or different steps. In other embodiments, one or more steps may be iterated through. The description of FIG. 3 applies to assessing whether a single sample has been affected by a sample swap event. In other embodiments, the analytics system may apply the same workflow to all samples. In other embodiments, the analytics system may apply portions of the workflow across samples concurrently.
[0124] The analytics system obtains 310 a sample including sequencing data (e.g., as output by a sequencing device) and, optionally, reported labels for clinical information associated with the individual. For example, the sequencing data may include methylation sequencing data, SNP sequencing data, small variant sequencing data, insertion and / or deletion sequencing data, other genetic sequencing data, etc. The clinical information mayAtty. Docket No. 32917-65128 / WO 34 P0248-WOinclude physical traits or characteristics about the individual, e.g., biological sex, race, ethnicity, smoking status, age, other medical diagnoses, etc.
[0125] The analytics system featurizes 320 the sequencing data for multimodal sample swap detection. Featurization may be specific to each sample swap detection model. For example, one model may evaluate a set of methylation features over a plurality of informative genomic regions. In another example, another model may evaluate features derived from SNP sequencing data.
[0126] The analytics system applies 330 a plurality of sample swap detection models to features of the sequencing data to generate a plurality' of predictions. Each sample swap detection model is configured in a distinct mode from other sample swap detection models. Each mode may evaluate different markers in the sequencing data, different types of sequencing data, different labels for clinical information, or some combination thereof. In general, each sample swap detection model outputs a prediction based on the input features, wherein the prediction is associated with an indicator of whether the sample is affected by a sample swap event. In one or more embodiments, the prediction is a binary’ prediction, i.e., detected sample swap event or not. In one or more embodiments, the prediction is a scalar, e.g., a likelihood of a sample swap event, or a confidence associated with the binary prediction.
[0127] In one or more embodiments, one or more sample swap detection models are based in predicting a covariate label and comparing the prediction to a reported covariate label from clinical information on the individual. The analytics system may leverage a machine-learning model to predict the covariate label or value based on features derived from the sequencing data. The analytics system may then compare the covariate prediction to a reported covariate label or value.
[0128] In one or more embodiments, one sample swap detection model compares a predicted biological sex against a reported biological sex. In such embodiments, a sample swap detection model may be configured to input, or receive as input, features of the sequencing data and to output a prediction of the biological sex based on the input features. The sample swap detection model may be trained as a machine-learning model. In the training, the analytics system may identify informative features for predicting the biological sex covariate. To perform training, the analytics system may obtain training samples including training samples labeled as biological male and training samples labeled as biological female. The analytics system may featurize sequencing data of the training samples according to the identified informative features. The analytics system may train theAtty. Docket No. 32917-65128 / WO 35 P0248-WOmachine-learning model to classify between biological male and biological female based on the features of the training samples. The sample swap detection model compares the predicted biological sex against the reported biological sex. If there is no discrepancy, the sample swap detection model may output that there is a negative indication or neutral indication of no detected sample swap event. If there is a discrepancy, then the sample swap detection model may output a positive indication of a sample swap event detected.Additional information describing methylation-based biological sex prediction is included in U.S. Application No. 18 / 737,779 entitled “Methylation-Based Biological Sex Prediction” filed on June 7, 2024.
[0129] In one or more embodiments, one sample swap detection model compares a predicted epigenetic age against a reported age. In such embodiments, a sample swap detection model may be configured to input, or receive as input, features of the sequencing data and to output a prediction of age based on the input features. The sample swap detection model may be trained as a machine-learning model. In the training, the analytics system may identify’ informative features for predicting the age covariate. To perform training, the analytics system may obtain training samples including training samples with ground truth ages or age ranges. The analytics system may featurize sequencing data of the training samples according to the identified informative features. The analytics system may train the machine-learning model to regress or otherw ise predict age based on the features of the training samples. In other embodiments, the machine-learning model may be trained as a classifier, to classify between different age ranges, bucketing a total age range. The sample swap detection model compares the predicted age against the reported age. In one or more embodiments, if there is a below threshold discrepancy between the predicted age and the reported age. then the sample swap detection model may output that there is a negative or neutral indication of no detected sample swap event. If there is an above threshold discrepancy, then the sample swap detection model may output a positive indication of a sample swap event detected. Additional information describing methylation-based biological sex prediction is included in U.S. Application No. 18 / 361,023 entitled “Methylation-Based Age Prediction as Feature for Cancer Classification” filed on July 28. 2023.
[0130] In one or more embodiments, one sample swap detection model compares a predicted ancestral group against a reported race or ethnicity. In such embodiments, a sample swap detection model may be configured to input, or receive as input, features of the sequencing data and to output a genetically inferred ancestry prediction based on the input features. The sample swap detection model may be trained as a machine-learning model. InAtty. Docket No. 32917-65128 / WO 36 P0248-WOthe training, the analytics system may identify informative features for predicting the ancestry covariate. To perform training, the analytics system may obtain training samples including training samples with reported races or ethnicities. The analytics system may featurize sequencing data of the training samples according to the identified informative features. The analytics system may train the machine-learning model to classify ancestry based on the features of the training samples. The sample swap detection model compares the predicted ancestry against the reported race or ethnicity. In one or more embodiments, if there is no discrepancy between the predicted ancestry and the reported race or ethnicity, then the sample swap detection model may output that there is a negative or neutral indication of no detected sample swap event. If there is a discrepancy, then the sample swap detection model may output a positive indication of a sample swap event detected.
[0131] In one or more embodiments, other sample swap detection models may consider other types of genetically correlated covariates. For example, one sample swap detection model may be based on predicting smoker status and comparing the predicted smoker status. As another example, one sample swap detection model may be based on predicting other diagnoses or health conditions.
[0132] In one or more embodiments, some sample swap detection models may leverage other reference samples to identify whether a target sample is likely derived from the same individual as the other reference samples. In such embodiments, the sample swap detection models may be agnostic of clinically -reported information.
[0133] In one or more embodiments, to determine whether two samples originate from the same biological source, the analytics system may leverage a genetic similarity model. The analytes system generates a genotype signature for each sample based on the sequencing data. For example, the analytics system may determine the genotype signature represented by genotyping of a plurality of genetic markers from the sequencing data. For example, the analytics system may generate a genotype signature for a sample based on genotypes at a plurality of common SNP sites in the genome. In some embodiments, the total number of SNP sites included in generating the genotype signature includes at least 100 SNP sites, at least 500 SNP sites, at least 1,000 SNP sites, at least 2,000 SNP sites, at least 3,000 SNP sites, at least 4,000 SNP sites, at least 5,000 SNP sites, at least 10,000 SNP sites, at least 50,000 SNP sites, at least 100,000 SNP sites, at least 500,000 SNP sites, or at least 1,000,000 SNP sites. Each of the different possible genotypes at a SNP site can be represented by a different numerical value or other classification for ease of computational reference. For example, homozygous reference can be represented by 0, heterozygous can be represented byAtty. Docket No. 32917-65128 / WO 37 P0248-WO1, homozygous alternate can be represented by 2, and missing genotyping can be represented by 3.
[0134] The analytic system inputs the genotype signature of the two samples into the genetic similarity' model to output a genetic similarity' score based on the two input genotype signatures. In some embodiments, the genetic similarity' score may be calculated based on a distance between the two genotype signatures. In one example, to compute the distance, the analytics system may compute a normalized Hamming distance: the number of loci where the encoded genotypes differ divided by the total number of common loci. Missing sites are excluded from both the numerator and denominator. In some embodiments, the analytics system may differentially weight the different ty pes of disagreement, e.g., heterozy gous-homozygous vs. homozygous-homozygous disagreements. This produces a distance in the range of [0, 1], where 0 indicates identical genotypes at all common loci and 1 indicates complete disagreement. In some embodiments, the genetic similarity' score ranges from 0 to 1, where 0 indicates complete lack of genetic similarity, and where 1 indicates genetic identity. The genetic similarity score may be a complement of the distance, e.g., if the distance is 0 indicating identical genotypes at all common loci, the genetic similarity score is 1 indicating whole similarity' across common loci.
[0135] In other embodiments, the genetic similarity' model may be trained as a machine-learning model. In such embodiments, the analy tics system may train the genetic similarity model by providing samples known to have originated from the same biological source as training input. For example, samples may be aliquoted from one original sample, thus known to be of the same biological source, but may be agnostic to the actual source. In some embodiments, the genetic similarity score may evaluate different types of genetic sites, e.g., SNPs. indels, other small variants, methylation, etc. The analytics system may leverage a threshold score to identify whether two samples are concordant (i. e. , at or above the threshold score results in a negative indication of a sample swap event) or discordant (i.e., below' the threshold score results in a positive indication of sample sw ap event). Additional information describing a genetic similarity’ scoring system is included in PCT Application No. PCT / US2024 / 052772 entitled "Systems and Methods for Assessing Similarity Between Samples Using Genoty pe Signatures” filed on October 24, 2024.
[0136] The analytics system may apply the genetic similarity model to determine genetic similarity' between all pairs of samples in a group associated with one individual. The group may include samples collected at various time points, from different sources (blood sample, or tissue biopsy sample), or for any other reason. For example, the group mayAtty. Docket No. 32917-65128 / WO 38 P0248-WOinclude a recurrent sample collected from an individual for recurrent early cancer detection screening. As another example, the group may include blood samples used in early cancer detection screening and tissue biopsy samples collected from a tumorous tissue identified in the individual. The analytics system may apply the genetic similarity' model between all pairs of the samples to output a genetic similarity score between all pairs of samples. Based on the pairwise genetic similarity scores, the analytics system may identify any sample in a group associated with an individual having a below threshold score. In one or more embodiments, the analytics system identifies a sample swap event if the similarity score is below the threshold score for at least 1, at least 2, at least 3, at least 4, or at least 5 other samples in the group. For example, if one sample in a group is identified as having below the threshold score when compared against all other samples in the group, then this is highly indicative of a sample swap event.
[0137] The analytics system applies 340 an aggregation model to the plurality of predictions output by the multimodal sample swap prediction models to generate an aggregate sample swap prediction. The aggregation model may leverage an aggregation function that disparately weights the various predictions output by the multimodal sample swap prediction models. In one or more embodiments, the aggregation model may be configured to input intermediary' predictions by one or more, including all, of the sample swap detection models. For example, the aggregation model may input a predicted age differential calculated by a methylation-based age prediction model. In another example, the aggregation model may input the genetic similarity scores between one sample under evaluation and other samples in the group. The aggregation model may be configured to output the aggregate sample swap prediction as a binary or discrete prediction - whether the sample is affected by a sample swap event or not. In other embodiments, the aggregation model may be configured to output the aggregate sample swap prediction as a numerical value indicating likelihood of the sample being affected by a sample swap event, e.g., there is a probability of 0.85 or an 85% likelihood that the sample is affected by a sample swap event.
[0138] The analytics system calls 350 a sample swap event based on the aggregate sample swap prediction. In some embodiments, the output of the aggregation model is a binary prediction of whether a sample swap event has affected a sample. In other embodiments, the output is a numerical value indicating a likelihood of a sample swap event. In such embodiments, the analytics system may use a threshold value to call the sample swap event, e.g., a score selected to hit a certain level of specificity.Atty. Docket No. 32917-65128 / WO 39 P0248-WO
[0139] The analytics system optionally performs 360 a remedial workflow to address any called sample swap events. The remedial workflow may include one or more remedial measures to remediate the sample swap event.
[0140] For example, one remedial measure may be to label the sample as unusable for downstream purposes (e.g., patient cancer classifications, model training). The sample may also be physically discarded. In another example, one remedial measure may be to transmit a request to the healthcare provider or to the patient to collect anew subsequent sample from the same patient. In other embodiments, the remedial measure may include sequencing a reserve
[0141] In another example, one remedial measure may be to remove the sample affected by the sample swap event from further downstream analyses. This may include removing training samples affected by the sample swap event prior to training of any analytical models. In another example, one remedial measure may be to reverse the sample swap event (may also be referred to as a “sample swap reversal”).
[0142] In another example, one remedial measure may be to pinpoint a cause of the sample swap event. For example, the analytics system may identify particular issues in the sample collection and processing workflow that may be causing the sample swap(s). For example, if the analytics system identifies multiple sample sw aps arising from one clinic, then the analytics system may identify that the clinic is a potential cause of the sample swaps.
[0143] Or in another example, the analytics system may identify multiple sample swaps arising due to one labeling workflow. Upon identifying potential causes, the analytics system may generate a report including the potential causes, and optionally including the sample swap detection statistics. The report may be transmitted to a client device of a human operator to further investigate and to address the potential causes.
[0144] Sample swap reversal entails identifying the true source of a sample identified as affected by a sample swap event. In one or more embodiments, the analytics system may assess all sample swap events identified in a set of samples. Typically, the sample swap events come in pairs, e.g., Label 1 for Individual 1 is accidentally placed on a sample for Individual 2 and Label 2 for Individual 2 is accidentally placed on a sample for Individual 1. If there are two sample swap events, the analytics system may assess the viability of swapping the two samples identified as affected by the two sample sw ap events.III.B. GENETIC SIMILARITY MODEL
[0145] In one or more embodiments, the analytics system may leverage a genetic similarity model to determine genetic similarity between samples. The genetic similarityAtty. Docket No. 32917-65128 / WO 40 P0248-WOmodel assesses the sequencing data of two samples to determine a genetic similarity score between the two samples. In one or more embodiments, the goal of genotype signature comparison between two samples is to assess the degree of genetic similarity or dissimilarity. The analytics system compares the two genotype signatures to determine the genetic similarity score. Thereafter, a total number of SNPs at specified locations may be identified for the first and second signatures. The specified locations may be selected so as to provide sufficient variability such that they can be used as identifying information. The total number of SNPs at common locations may correspond to the positions where both signatures have genetic variant information.
[0146] FIG. 4 is an exemplary flowchart of the process 400 of deployment of a genetic similarity’ model, according to one or more embodiments. The analytics system is described as performing the process 400 of deployment of the genetic similarity model. In other embodiments, other systems may perform some of the steps. In other embodiments, the process 400 may include additional, fewer, or different steps. In other embodiments, the genetic similarity’ model may be applied across pairs of samples.
[0147] The analytics system obtains 410 sequencing data for nucleic acid fragments in two samples. The sequencing data may include variant sequencing data for SNPs present in the nucleic acid fragments.
[0148] The analytics system determines 420 a genotype signature for each sample based on the sample’s sequencing data. In one or more embodiments, the analytics system may calculate a variant allele frequency (VAF) for each SNP position associated with the sequencing data to provide a quantitative measure of the presence of the alternate allele. Based on the VAF value, the genetic variant ty pe at each SNP position may be determined. For example, if VAF is below a first threshold (e.g.. closer to 0), the alleles at the corresponding SNP position are primarily reference alleles, indicating a homozygous reference allele. If the VAF is higher than a second threshold (e.g., closer to 1), the alleles at the corresponding SNP position are primarily alternate alleles, indicating a homozy gous alternate allele. If the VAF is between the first and second thresholds, the alleles at the SNP position are a mix of reference and alternate alleles, indicating a heterozygous allele. If the VAF cannot be accurately calculated due to insufficient coverage, the variant type for that SNP position is considered missing data. The missing data designation may indicate that the genetic information at that position is not reliable due to coverage limitations.
[0149] The analytics system may assign numeric values to represent the identified genetic variant types. More particularly, the determined genetic variant types at each SNPAtty. Docket No. 32917-65128 / WO 41 P0248-WOposition within the second subset are encoded into bytes for ease of computationally-aided comparison. Each genetic variant type is assigned a unique numeric value, and these values are combined to create a compact binary representation.
[0150] In one or more embodiments, each of the genetic variant types may be assigned a specific numeric value for encoding purposes. For instance:• Homozygous Reference Allele (HOM REF): Encoded as 0 (<00> in binary form)• Heterozygous Allele (HET): Encoded as 1 (<01> in binary form) • Homozy gous Alternate Allele (HOM_ALT): Encoded as 2 (<10> in binary7form)• Missing Data: Encoded as 3 (<11> in binary form)
[0151] For each SNP position, the assigned numeric value corresponding to the determined genetic variant is combined into a sequence representing the entire genotype signature for the sample. In one or more embodiments, encoding genetic variant types for the selected positions into bytes may result in a compact unique representation of the individual’s genetic profile: the genotype signature. This binary format requires little storage space, making it efficient for storage and data management. Additionally, the compact binary7representation allows for rapid comparison and analysis against other samples, as further described herein.
[0152] The analytics system identifies 430 common positions with genomic information in the genotype signatures of the two samples. For example, with SNP sequencing data encoded in the genotype signature, the analytics system identifies at which SNP positions the two genotype signatures have a determined genetic variant type. The analytics system may7count the total number of common positions.
[0153] As an example, consider two genotype signatures, Signature A and Signature B. Signature A has genetic variant information at SNP positions 1, 3, 5, and 7. Signature B has genetic variant information at SNP positions 2, 4, 5 and 7. SNP positions 5 and 7 are common SNPs because both Signature A and Signature B have genetic variant information at these positions. SNP positions 1, 3, and 2, 4 are not common SNPs because only one of the signatures has genetic variant information at these positions.
[0154] In one or more embodiments, the analytics system may perform an initial threshold check to determine whether the total number of common SNPs between the first and second genotype signatures is greater than a predetermined threshold. This threshold check serves as a criterion for determining if there is a sufficient level of overlap between theAtty. Docket No. 32917-65128 / WO 42 P0248-WOtwo signatures to proceed with further comparison. In one or more embodiments, the predetermined threshold value may vary based on factors such as the nature of the genetic data, the goals of the comparison, and the desired level of statistical significance. For instance, the predetermined threshold for the comparison of two genoty pe signatures associated with samples of the same ty pe (e.g., two cfDNA samples) may be higher than if the two genotype signatures were associated with samples of different types (e.g., one cfDNA sample vs. one tissue sample or two tissue samples). In one or more embodiments, as another example, the predetermined threshold may be higher and more stringent if the sample analysis was associated with a disease state decision or treatment recommendation than it may be for a research study.
[0155] Responsive to determining that the total number of common SNPs is not greater than the predetermined threshold, the genetic similarity model may output a result identifying that there is insufficient overlap between the two genotype signatures for a meaningful comparison. This may be due to factors such as low coverage of genetic data or limited shared genetic information. Accordingly, an insufficient overlap may halt the comparison from proceeding further. In contrast, responsive to determining that the total number of common SNPs was greater than the predetermined threshold, the genetic similarity model may output an indication that there was a sufficient amount of overlap between the two signatures.
[0156] The analytics system determines 440 the genetic similarity score based on a count of common positions where the genotype signatures are in genetic agreement. The genetic similarity' model identifies a total number of SNPs that have identical variant calls (e.g., homozy gous reference, homozygous alternate, and heteroz gous) between the first and the second genotype signatures. To facilitate this determination, the variant call for each position in each genotype signature may be identified, and if the variant call at a position is the same in both signatures, then the total count may be incremented.
[0157] For example, two genotype signatures may be considered, Signature X and Signature Y. The signatures share common SNP positions at positions 5 and 9. At position 5, Signature X may have a homozygous reference call, and Signature Y may also have a homozygous reference call. Because both signatures have the same variant call, the SNP at position 5 may be considered identical between signatures. Conversely, at position 9, Signature X may have a homozygous reference call, and Signature Y may have a heterozygous call. Because the signatures have different variant calls, the SNP position at position 9 may be considered not identical.Atty. Docket No. 32917-65128 / WO 43 P0248-WO
[0158] The genetic similarity' model calculates the genetic similarity' score between the two genotype signatures. The genetic similarity score provides a quantitative measure of the degree of genetic similarity' between the two signatures. It may help in assessing the degree of overlap in genetic variant types and provides insight into how closely the genetic profdes of the two samples match at the common SNP positions. The genetic similarity' score may be computed by dividing the number of identical SNPs by the total number of common SNPs. Mathematically, in some examples, the formula for computing the genetic similarity' score is:Number of Identical SNPsScore = - Total Number of Common SNPs
[0159] The genetic similarity' score ranges between 0 and 1, where a genetic similarity score of 0 indicates no genetic agreement (i.e., no identical SNPs) between the signatures, whereas a genetic similarity score of 1 indicates complete genetic agreement (all common SNPs are identical) between the signatures. As anon-limiting example of the foregoing, Signature X and Signature Y may have 5,000 common SNPs, out of which 4,000 SNPs have identical variant calls. The genetic similarity score would be calculated as: 4,000 / 5,000 = 0.8. In this example, the genetic similarity’ score is 0.8, which means that 80% of the common SNPs have identical variants.
[0160] In one or more embodiments to the foregoing, in another aspect, a kinship coefficient ( ), or “coefficient of relationship,” may be calculated. The kinship coefficient may correspond to the measure of genetic relatedness between two samples. It may be calculated based on the sharing of alleles at specific genetic markers, e.g., SNPs, and may provide insight into how closely two or more samples are related by blood, or, in this case, how likely two samples are to be from the same individual. In one or more embodiments, the kinship coefficient may be calculated by identifying, at each SNP, how many alleles are shared, which may be one of three possibilities: homozygous reference (e.g., if both individuals have the same reference allele at a given SNP), heterozygous (e.g., if one individual has the reference allele and the other has a variant allele), or homozygous alternate (e.g., if both individuals have the same variant allele at a given SNP).
[0161] The kinship coefficient may be calculated by summing the shared status at each SNP (e.g., 0 for homozygous reference, 1 for heterozygous, or 2 for homozygous variant) and then averaging across all of the total number of common SNPs. The resulting coefficient may range from 0 (indicating no genetic relatedness) to 1 (indicating complete genetic identity ). In one or more embodiments, the values may indicate different degrees ofAtty. Docket No. 32917-65128 / WO 44 P0248-WOgenetic relatedness (e.g., a value of 0.25 may suggest that a grandparent-grandchild relationship or an uncle-aunt / niece-nephew relationship, a value 0.5 may indicate a parentchild or a sibling relationship, a value of 1 may suggest identical twins). In this case, a value of 1 may indicate that the samples are from the same individual, whereas increasingly smaller values may indicate an increasingly greater likelihood that two samples are not from the same individual.
[0162] The analytics system calls 450 a positive match if the genetic similarity score satisfies a concordance threshold (e.g., at or above the concordance threshold). Conversely, if the genetic similarity' score does not satisfy the concordance threshold, then the analytics system may call a negative match. The analytics system may compare the computed genetic similarity score against the concordance threshold. The threshold value may be used as a criterion in determining whether the genetic profiles are similar enough to be considered a positive match. In one or more embodiments, the threshold value may be higher or lower based on certain factors, as described above. For instance, the predetermined threshold for the comparison of two genotype signatures associated with samples of the same type (e.g.. two cfDNA samples) may be higher than if the two genotype signatures were associated with samples of different types (e g., one cfDNA sample vs. one tissue sample). In one or more embodiments, as another example, the predetermined threshold may be higher and more stringent if the sample analysis was associated with a disease state decision or treatment recommendation than it may be for a research study.
[0163] In one or more embodiments, responsive to determining that the genetic similarity score is greater than the concordance threshold, the genetic similarity model may generate a result indicating a match betw een the genotype signatures. Conversely, responsive to determining that the genetic similarity score is less than the concordance threshold, the genetic similarity model may generate a result indicating a mismatch between the genotype signatures.
[0164] In some embodiments, situations may exist in which the genotype signature of a particular sample is compared against the genotype signatures of a plurality of other samples. At the conclusion of the comparison process, a subset of samples in the plurality may be identified as being matched with the particular sample. More particularly, a genetic similarity score generated from the comparison of the genotype signatures of the particular sample and one of the subset samples may be determined to be greater than the threshold. As a non-limiting example, the genotype signature of a first sample may be compared against the genotype signatures of 1 million other samples. After the comparison process is complete,Atty. Docket No. 32917-65128 / WO 45 P0248-WOapproximately 1000 samples contained genoty pe signatures that were matched to the genotype signature of the first sample. In this circumstance, the resultant subset of samples may then be passed to one or more different modules so that more granular comparison processes may be conducted (e.g., on the sequencing data associated with each sample) to ultimately identify a 1:1 match. In this way, the concordance threshold may act as an initial filter on a pool of data. In one or more embodiments, the threshold may be adjustable based on need and / or context (e.g., the threshold may be increased or decreased based on the desired goals of the comparison process).
[0165] In one or more embodiments, the concordance threshold value may change based on the type of samples that the genotype signatures are associated with. More particularly, different types of biological samples (e.g., blood, urine, tissue, etc.) may vary in the extent of genetic variation they contain. Some samples, such as blood, may have a relatively stable and consistent genetic profile, while others, like tumor tissue, may exhibit higher levels of genetic heterogeneity due to mutations and clonal evolution. Accordingly, a first concordance threshold may be established for the comparison of two genotype signatures that are both associated with cfDNA, whereas a second concordance threshold may be established for the comparison of samples with inherently higher genetic variation (e.g., cfDNA vs. tissue). The second concordance threshold may be more relaxed to account for the expected diversity between the samples.
[0166] The concordance threshold may in one or more embodiments be adjusted based on other factors as well. For instance, in longitudinal studies involving repeated sampling from the same participant over time, genetic changes may occur due to factors like disease progression, treatment response, natural aging, or one or more other natural variations. For such studies, the concordance threshold may be adjusted to accommodate expected genetic drift while still identify ing samples from the same individual as matching. As another example, the clinical context and specific application of the genotype signature comparison may at least in part influence the choice of the concordance threshold. For example, in diagnostic scenarios in which accurate patient identification is crucial, a stricter concordance threshold may be chosen to reduce the risk of false positive matches. On the other hand, in research studies focusing on broader population analysis, a more lenient threshold may be applied to reflect the wider range of genetic diversify between samples. In yet another example, genetic variations may also differ between different ancestral populations. If the samples being compared are from diverse ancestral backgrounds, the concordance threshold may be tailored to account for population-specific genetic variation.Atty. Docket No. 32917-65128 / WO 46 P0248-WO
[0167] In one or more embodiments, the genetic similarity' score may provide various non-binary decisions or insights that are not limited to binary ■'match" or “mismatch’" outcomes. For instance, upon analyzing the genetic profiles of the individuals associated with two samples, A and B, the concordance value may provide insight into the degree of genetic relatedness, or the likelihood that two samples are from the same individual. For example, a concordance value of 0.8 may indicate that 80% of the common SNPs between samples A and B have identical variants. Based on this, it may be concluded from the concordance value that the individuals associated with samples A and B share a significant portion of their genetic variants, suggesting a high degree of genetic similarity or relatedness, or may indicate the likelihood that the samples are from the same individual.
[0168] In yet another example, the concordance value may be used to determine whether samples A and B are associated with a single individual, and the system of the embodiments may be configured to output a confidence percentage that the samples are from the same person. In other words, the system may output a degree of genetic similarity between the two samples.III.C. SAMPLE SWAP REVERSAL
[0169] In one or more embodiments, the analytics system may leverage a genetic similarity model to reverse a sample swap event.
[0170] FIG. 5 is an exemplary flowchart describing a process 500 of sample sw ap reversal, according to one or more embodiments. The analytics system is described as performing the process 500 of sample swap reversal. In other embodiments, other systems may perform some of the steps. In other embodiments, the process 500 may include additional, fewer, or different steps. In other embodiments, one or more steps may be iterated through. The description of FIG. 5 applies to assessing whether a single sample has been affected by a sample swap event. In other embodiments, the analytics system may apply the same workflow to all samples. In other embodiments, the analytics system may apply portions of the workflow across multiple samples concurrently.
[0171] The analytics system obtains 510 samples with sequencing data, and optionally including reported clinical information on individuals from which the samples are collected from. In one or more embodiments, the samples may compose individuals of a cohort. For example, the samples may be associated with individuals that had their samples collected by one clinical site (or a region of clinical sites). In another example, the samples may be associated with individuals collected as part of one study. The analytics system may have tiered partitions of the samples, e.g., by studies at a highest level, then by clinical sites,Atty. Docket No. 32917-65128 / WO 47 P0248-WOthen by periods of collection, etc. The clinical information may include physical traits on the individuals from which the samples are collected. Samples may be further grouped for each intended individual.
[0172] The analytics system featurizes 520 each sample’s sequencing data.Featurization may include SNP featurization (e.g., as described above in the genetic similarity model), methylation featurization, other epigenetic featurization, other genetic mutation featurization, or some combination thereof. In one or more embodiments, the analytics system may featurize a sample by determining a genotype signature based on the SNP genot ping across a plurality7of SNP sites in the reference genome.
[0173] The analytics system applies 530 a genetic similarity7model to pairwise features of the sequencing data of a pair of samples to output a genetic similarity score between the sample pair. In one or more embodiments, the genetic similarity model may compute the genetic similarity7score based on SNP genotyping between the two samples. In other embodiments, the genetic similarity model may compute the genetic similarity score based on other genotyping, either alone or in conjunction with the SNP genotyping. The output similarity score indicates a concordance of the genotype signatures of the two samples.
[0174] The analytics system constructs 540 a graph comprising nodes representing samples and edges between nodes representing pairwise genetic similarity7scores for samples in each group. The analytics system constructs this graph for each group, i.e., for samples expected to belong to one individual based on labels. If there is no sample swap, the pairwise similarity7scores between samples in the group should be above a threshold score. However, if pairwise similarity7scores for one sample in the group is below the threshold score, that suggests that the sample does not truly belong to the same individual as the other samples.
[0175] The analytics system calls 550 a sample swap event based on identifying a given sample’s pairwise genetic similarity scores with others of the group being below a threshold. In examples with three or more samples for an individual, i.e., in a group, the analy tics system may identify the odd one out. However, in examples with only two samples for an individual, i.e., in a group, the analytics system would need to further investigate, as in this example it is unknown which sample’s genotype signature is true to the individual.
[0176] For a called sample swap event, the analytics system determines 560 a sample swap reversal by traversing the graph to identify7a swap that maximizes genetic similarity of the given sample with others in another group. For example, the analytics system may leverage the genetic similarity model to calculate a genetic similarity score between the given sample and at least one sample from each of the other groups (i.e., belonging to otherAtty. Docket No. 32917-65128 / WO 48 P0248-WOindividuals). The group with the sample having the highest genetic similarity' to the given sample (i.e. , above a threshold score, e.g., above 90%, 95%, or 98%) is selected as the corrected group (i.e., the corrected individual) for the sample.
[0177] In some embodiments, the analytics system may identify all called sample swap events in the cohort of samples to guide traversal of the graph. Given that sample swap events typically involve multiple samples, the analytics system may perform initial screening for sample swap reversal with the group affected by sample swap events. If there are two or more sample swap events, the analytics system may initially evaluate genetic similarity of the given sample against the other samples subject to called sample swap events. This initial screening would obviate the need to scour the entire graph. If none of the groups return a positive match for the given sample, then the analytics system may expand its search and traverse the entire graph.
[0178] In some embodiments, the analytics system may perform initial screening with one sample from each of the other groups. Groups may be identified based on proximity to the initial group, which may be defined based on samples having been received contemporaneously (e.g., from the sample site or location, or simply arriving at a lab at the same time) or processed contemporaneously (e.g., in systems in which multiple samples are handling in batches or otherwise in close physical proximity). These are instances in which a physical sample swap is more likely.
[0179] The analytics system may identify some candidate groups based on the genetic similarity scores with the individual samples of those candidate groups. The analytics system may expand the genetic similarity comparison to each of the samples in each of the candidate groups. For example, if a first candidate group has four samples and a second candidate group has six samples, then the analytics system may calculate genetic similarity scores between the given sample and all the four samples of the first candidate group, and may calculate genetic similarity scores between the given sample and all six samples of the second candidate group. The analytics system may leverage the totality' of scores to identify the better match for the given sample. For example, the analytics system may evaluate the averaged scores for each candidate group. In other examples, the analytics system may evaluate which of the candidate groups has the absolute highest genetic similarity score with the given sample.
[0180] The analytics system may leverage one or more optimization algorithms to identify corrected associations for swapped samples. The analytics system may identifyAtty. Docket No. 32917-65128 / WO 49 P0248-WOswaps that maximize a likelihood of concordance of a swapped sample with other samples of other groups.
[0181] Upon identifying a viable sample swap reversal, the analytics system relabels 570 the given sample to be associated with the other group. Relabeling the sample is advantageous in that the sample can be retained for downstream analyses. If, however, the analytics system finds no viable sample swap reversal, e.g., no group has a genetic similarity score above the threshold score, then the analytics system may deem the given sample as unusable, and withheld from downstream analyses.
[0182] FIG. 6 is an illustrative example of the process of sample swap reversal (e.g., as described in FIG. 5), according to one or more embodiments. In this example, a first individual (“Individual 1”) provided six samples for sequencing and processing, denoted by the hatch fill pattern, and a second individual (“Individual 2”) provided four samples for sequencing and processing, denoted by the white fill pattern. The square represents tissue biopsy samples, while the circle represent liquid biopsy samples, e.g., blood samples with cell-free nucleic acid fragments. Sample 3 (“S3”) collected from Individual 1 was mislabeled as belonging to Individual 2. Likewise Sample 10 ("S10”) collected from Individual 2 was mislabeled as belonging to Individual 1. In this example, the two labels were swapped between the two samples. In other examples, a sample swap event may result in just a single sample being mislabeled. In other embodiments, different types of samples may be compared, e.g., bone marrow sample, plasma sample, buffy coat sample, plasma sample, urine sample, etc.
[0183] The analytics system may group samples together based on their expected origination. Samples 1, 2, 4, 5, 6, and 10 are expected as belonging to Group 1, due to the mislabeling of Sample 10. Further, Samples 3, 7, 8, and 9 are expected as belonging to Group 2, due to the mislabeling of Sample 3.
[0184] The analytics system may compute the pairwise genetic similarity between sample pairs in each group. The analytic system may leverage a genetic similarity model to compute the pairwise genetic similarity, e.g., based on concordance of SNP genotyping between each sample pair.
[0185] The analytics system may construct a graph with the samples and the genetic similarity scores. In constructing the graph, the analytics system may initially connect nodes (i.e., samples) in one group with edges representing the calculated genetic similarity scores. In the example, if the calculated genetic similarity score (i.e., between a pair of samples) is above a threshold score to positively identify concordance between the samples, then theAtty. Docket No. 32917-65128 / WO 50 P0248-WOedge is illustrated with a solid line; whereas, if the calculated genetic similarity score is below the threshold score, then the edge is illustrated with a dashed line.
[0186] The analytics system may identify Sample 10 (“S 10") as affected by a sample swap event due to the discordant genetics between Sample 10 and other samples in Group 1. Likewise, the analytics system may identify Sample 3 (“S3"’) as affected by a sample swap event due to the discordant genetics between Sample 3 and other samples in Group 2. The confidence in the sample swap call may be further based on the percentage of discordance with other samples in the group. For example, if only one other sample in the group results in discordant genetics w ith a target sample, then the confidence in the sample swap call would be low. On the contrary, if all other samples in the group result in discordant genetics with the target sample, then the confidence in the sample swap call would be high.
[0187] To perform the sample swap reversal, the analytics system may evaluate candidate swaps. In one or more embodiments, the analytics system may evaluate swaps by swapping the group associations of two samples. Here, the analytics sy stem may evaluate swapping Sample 3 (“S3”) and Sample 10 (“S10”). Otherwise, the analytics system may evaluate swapping the group association of one sample independently, e.g., evaluate swap reversal for Sample 3 (“S3”) and swap reversal for Sample 10 (“S10”), independently.
[0188] When assessing sample swap reversal, the analytics system calculates genetic similarity score(s) between the swapped sample and sample(s) from other group(s). For example, the analytics system may calculate, via the genetic similarity model, a genetic similarity between Sample 10 and one or more samples in Group 2. If the genetic similarity score(s) show concordance, i.e., above the threshold to establish genetic concordance, then the analytics sy stem may swap association of Sample 10 to Group 2, i.e., relabeling Sample 10 as originating from Individual 2. The same workflow is applied to Sample 3, to assess a correct association for the swapped sample. The bottom pair of graphs illustrate the corrected associations, following sample swap reversal.HI D. METHYLATION-BASED BIOLOGICAL AGE PREDICTION
[0189] In one or more embodiments, the analytics system leverages, for sample swap detection, a covariate prediction model configured to predict a biological age of an individual based on the methylome of the individual.
[0190] The analytics system obtains a test sample with a plurality of nucleic acid fragments and a reported value or label for the covariate. A physician or other medical provider collects the test sample and may also obtain the reported value of one or more covariates. For this example, the covariate under analysis may be the age of the individualAtty. Docket No. 32917-65128 / WO 51 P0248-WOproviding the test sample, although any of the covariates discussed herein or grouping of those covariates may be used. In some embodiments, age may be a single value or may be an age range. For example, the individual may report an age of 47, or may report an age range of 40-50. The sample may be any ty pe of biological sample comprising nucleic acid material of the individual. In embodiments with a blood sample, the blood sample comprises at least cfDNA fragments sheared from cells.
[0191] The analytics system processes and sequences the nucleic acid fragments in the test sample to identify a methylation pattern for each nucleic acid fragment. Processing and sequencing may involve bisulfite sequencing to convert unmethylated CpG sites. In other embodiments, a sequencer performs the sequencing of the nucleic acid fragments, and the analytics system processes the sequence reads to determine the methylation pattern. The analytics system may further perform one or more processing steps to the sequence reads, e.g., de-duping copies of the same original fragment, identifying contamination fragments, identifying sequencing error, etc.
[0192] The analytics system applies the trained covariate prediction model to predict an age for the test sample based on the methylation patterns of the nucleic acid fragments of the test sample. The analytics system determines methylation features for the covariate prediction model, e.g., methylation densify; count, distribution, or percentage of highly methylated fragments overlapping a genomic region; count, distribution or percentage of highly unmethylated fragments overlapping a genomic region; other characteristics based on methylation data; or some combination thereof. The covariate prediction model is configured to input methylation features for a feature set of genomic regions and output a predicted age based on the methylation features. As with reported age, the predicted age may be a single value or an age range.
[0193] The analytics system compares the predicted age to the reported age. The comparison may be a determination of whether the predicted age matches to the reported age. For example, if the reported age was an age range of 20-30, and the predicted age was 26 or the range 20-30, then the predicted age matches to the reported age range. In some embodiments, the comparison may be a residual as a difference between the predicted age and the reported age. For example, if the reported age was 63 and the predicted age was 72, then the residual is 9 years over the reported age. The residual may also be absolute, e.g., 9 years different than the reported age.
[0194] The analytics system performs sample swap detection based on the comparison. The analytics system may utilize the residual to determine whether the sampleAtty. Docket No. 32917-65128 / WO 52 P0248-WOwas swapped, such that the sample doesn’t truly originate from the individual. The analytics system may call the sample swap if the predicted age is different than the reported age. In other embodiments, the analytics system may call the sample swap if the residual is above a threshold difference. For example, the residual threshold can be set at 10 years, so if the residual between the predicted age and the reported age is above the 10-year residual threshold, then the analytics system can call the sample swap. In yet other embodiments, the analytics system may call the sample swap based on the comparison of the predicted age and the reported age in conjunction with other analyses. For example, the analytics system may train a separate model for ancestral origin determination to determine whether a predicted ancestral origin matches to the individual's reported ancestral origin.
[0195] Upon calling a sample swap contamination event, the analytics system may perform one or more remedial measures. Example remedial measures include withholding the sample from downstream analyses. Samples not called to be sample swaps may proceed with downstream analyses. For example, upon calling a sample swap for a training sample, the analytics system may withhold the training sample from use in training one or more models or building one or more distributions. As another example, upon calling a sample swap for a test sample, the analytics system may withhold cancer prediction for the test sample. Swapped samples can also be physically discarded.HI E. METHYLATION-BASED BIOLOGICAL SEX PREDICTION
[0196] In one or more embodiments, the analytics system leverages, for sample swap detection, a covariate prediction model configured to predict a biological sex of an individual based on the methylome of the individual.
[0197] The analytics system obtains a test sample with a plurality of nucleic acid fragments and a reported value or label for the covariate. A physician or other medical provider collects the test sample and may also obtain the reported value of one or more covariates. For this example, the covariate under analysis may be the biological sex of the individual providing the test sample, although any of the covariates discussed herein or grouping of those covariates may be used.
[0198] The biological sex of the individual may one of a set of known biological sexes for the species. For example, humans generally are one of biological male, or biological female. The biological sex prediction model may output a value in the range of (0, 1), where 0 may refer to the first biological sex and 1 may refer to the second biological sex. In other embodiments, the biological sex prediction model may output a likelihood that the biological sex is biological male, or a likelihood that the biological sex is biological female.Atty. Docket No. 32917-65128 / WO 53 P0248-WO
[0199] The analytics system processes and sequences the nucleic acid fragments in the test sample to identify a methylation pattern for each nucleic acid fragment. Processing and sequencing may involve bisulfite sequencing to convert unmethylated CpG sites. In other embodiments, a sequencer performs the sequencing of the nucleic acid fragments, and the analytics sy stem processes the sequence reads to determine the methylation pattern. The analytics system may further perform one or more processing steps to the sequence reads, e.g., de-duping copies of the same original fragment, identifying contamination fragments, identify ing sequencing error, etc.
[0200] The analytics system applies the trained covariate prediction model to predict a biological sex for the test sample based on the methylation patterns of the nucleic acid fragments of the test sample. The analytics system determines methylation features for the covariate prediction model, e.g., methylation densify; count, distribution, or percentage of highly methylated fragments overlapping a genomic region; count, distribution or percentage of highly unmethylated fragments overlapping a genomic region; other characteristics based on methylation data; or some combination thereof. The covariate prediction model is configured to input methylation features for a feature set of genomic regions and output a predicted biological sex based on the methylation features. As with the reported biological sex, the predicted biological sex may be one of the two labels.
[0201] The analytics system compares the predicted biological sex to the reported biological sex. The comparison may be a determination of whether the predicted biological sex matches to the reported biological sex.
[0202] The analytics system performs sample swap detection based on the comparison. If the reported biological sex does not match the predicted biological sex, then the analytics system may detect a sample swap event.
[0203] Upon calling a sample swap contamination event, the analytics system may perform one or more remedial measures. Example remedial measures include withholding the sample from downstream analyses. Samples not called to be sample swaps may proceed with downstream analyses. For example, upon calling a sample swap for a training sample, the analytics system may withhold the training sample from use in training one or more models or building one or more distributions. As another example, upon calling a sample swap for a test sample, the analytics system may withhold cancer prediction for the test sample. Swapped samples can also be physically discarded.Atty. Docket No. 32917-65128 / WO 54 P0248-WOIV. CANCER CLASSIFICATION FOR DETERMINING CANCER
[0204] Cancer classification involves extracting genetic features and applying one or more models to the extracted features to determine a cancer prediction. The analytics system aggregates extracted features into a feature vector which can then be input into a trained cancer prediction model to determine a cancer prediction based on the input feature vector. The cancer prediction may comprise one or more labels and / or one or more values. One label may be binary, indicating a presence or absence of cancer in the test subject. Another label may be multiclass, indicating one or more particular cancer signal origins from a plurality of screened cancer signal origins. One value may indicate a likelihood of presence of cancer. Another value may indicate a likelihood of absence of cancer. Yet another value may¬ otherwise indicate another prognosis of the cancer. For example, the value may quantify a progression and / or potential response to treatment of the cancer.
[0205] In one or more embodiments, the feature vectors input into the cancer classifier are based on a set of informative fragments (also referred to as ‘‘anomalously-methylated” or “unusual fragments of extreme methylation” (UFXM)) determined from the test sample.
[0206] In some embodiments, a cancer classifier may be a machine-learned model which is a computation model comprising a plurality of classification parameters and a function representing a relation between the feature vector as input and the cancer prediction as output. Inputting the feature vector into the function with the classification parameters yields the cancer prediction. The machine-learned model may be trained using training samples derived from subjects with known cancer diagnoses. The training samples may be divided into cohorts of vary ing labels. For example, there may be a cohort of training samples for each cancer signal origin.
[0207] A machine-learned model may be trained through any combination of machine-learning techniques including, but not limited to, Linear Regression, Logistic Regression, Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Decision Trees (e.g., ID3, C4.5, CART), Random Forest, Neural Networks (e.g., Multi-Layer Perceptron, Convolutional Neural Networks. Recurrent Neural Networks), Naive Bay es Classifier, Gradient Boosting Machines (e g., XGBoost, LightGBM, CatBoost), AdaBoost, K-Means Clustering, Hierarchical Clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Gaussian Mixture Models (GMM), PCA (Principal Component Analysis), t-SNE (t-Distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), ICA (Independent Component Analysis), Autoencoders (e.g.,Atty. Docket No. 32917-65128 / WO 55 P0248-WOVariational Autoencoders, Denoising Autoencoders), Self-Organizing Maps (SOM), Q-Leaming, Deep Q Network (DQN), SARSA (State-Action-Reward-State-Action), Monte Carlo Methods, Temporal Difference Learning (TD Learning), Policy Gradients (e.g., REINFORCE, Actor-Critic), Proximal Policy Optimization (PPO), Advantage Actor-Critic (A2C), Soft Actor-Critic (SAC), Deep Deterministic Policy Gradient (DDPG), Hidden Markov Models (HMM), Conditional Random Fields (CRF), Latent Dirichlet Allocation (LDA), Restricted Boltzmann Machines (RBM), Genetic Algorithms, and Swarm Intelligence Algorithms (e g., Particle Swarm Optimization, Ant Colony Optimization), Expectation Maximization (EM).IV. A. IDENTIFYING INFORMATIVE FRAGMENTS
[0208] The analytics system can determine informative fragments for a sample using the sample’s methylation state vectors. For each fragment in a sample, the analytics system can determine whether the fragment is an informative fragment using the methylation state vector corresponding to the fragment. In some embodiments, the analytics system calculates a p-value score for each methylation state vector describing a probability of observing that methylation state vector or other methylation state vectors even less probable in the healthy control group. The process for calculating a p-value score is further discussed below in Section III. A.i. P-Value Filtering. The analytics system may determine fragments with a methylation state vector having below a threshold p-value score as informative fragments. In some embodiments, the analytics system further labels fragments with at least some number of CpG sites that have over some threshold percentage of methylation or unmethylation as hypermethylated and hypomethylated fragments, respectively. A hypermethylated fragment or a hypomethylated fragment may also be referred to as an unusual fragment with extreme methylation (UFXM). In other embodiments, the analytics system may implement various other probabilistic models for determining informative fragments. Examples of other probabilistic models include a mixture model, a deep probabilistic model, etc. In some embodiments, the analytics system may use any combination of the processes described below for identifying informative fragments. With the identified informative fragments, the analytics system may filter the set of methylation state vectors for a sample for use in other processes, e g., for use in training and deploying a cancer classifier.IV. A.i. P-VALUE FILTERING
[0209] In some embodiments, the analytics system calculates a p-value score for each methylation state vector compared to methylation state vectors from fragments in a healthy control group. The p-value score can describe a probability of observing the methylationAtty. Docket No. 32917-65128 / WO 56 P0248-WOstatus matching that methylation state vector or other methylation state vectors even less probable in the healthy control group. In order to determine a DNA fragment to be anomalously methylated, the analytics system can use a healthy control group with a majority of fragments that are normally methylated. When conducting this probabilistic analysis for determining informative fragments, the determination can hold weight in comparison with the group of control subjects that make up the healthy control group. To bolster robustness in the healthy control group, the analytics system may select some threshold number of healthy subjects to source samples including DNA fragments. FIG. 7A below describes the method of generating a data structure for a healthy control group with which the analytics system may calculate p-value scores. FIG. 7B describes the method of calculating a p-value score with the generated data structure.
[0210] FIG. 7A is a flowchart describing a process 700 of generating a data structure for a healthy control group, according to an embodiment. To create a healthy control group data structure, the analytics system can receive a plurality of DNA fragments (e.g., cfDNA) from a plurality of healthy subjects. The analytics system can generate 705 a methylation state vector for each fragment, for example via the process 200 of FIG. 2A.
[0211] With each fragment’s methylation state vector, the analytics system can subdivide 710 the methylation state vector into strings of CpG sites. In some embodiments, the analytics system subdivides 710 the methylation state vector such that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of length less than or equal to 3 would result in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 being subdivided into strings of length less than or equal to 4 can result in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If a methylation state vector is shorter than or the same length as the specified string length, then the methylation state vector may be converted into a single string containing all of the CpG sites of the vector.
[0212] The analytics system tallies 715 the strings by counting, for each possible CpG site and possibility of methylation states in the vector, the number of strings present in the control group having the specified CpG site as the first CpG site in the string and having that possibility of methylation states. For example, at a given CpG site and considering string lengths of 3, there are 2A3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analytics system tallies 710 how many occurrences of each methylation state vector possibility come up in the control group.Atty. Docket No. 32917-65128 / WO 57 P0248-WOContinuing this example, this may involve tallying the following quantities: < Mx, Mx+i, Mx+2 >, < Mx, Mx+i, Ux+2 >, . . .. < Ux, Ux+i. Ux+2 > for each starting CpG site in the reference genome. The analytics system creates 715 the data structure storing the tallied counts for each starting CpG site and string possibility.
[0213] There are several benefits to setting an upper limit on string length. First, depending on the maximum length for a string, the size of the data structure created by the analytics system can dramatically increase in size. For instance, a maximum string length of 4 means that every CpG site has at the very least 2A4 numbers to tally for strings of length 4. Increasing the maximum string length to 5 means that every' CpG site has an additional 2A4 or 16 numbers to tally, doubling the numbers to tally (and computer memory required) compared to the prior string length. Reducing string size can help keep the data structure creation and performance (e.g., use for later accessing as described below), in terms of computational and storage, reasonable. Second, a statistical consideration to limiting the maximum string length can be to avoid overfitting dow nstream models that use the string counts. If long strings of CpG sites do not, biologically, have a strong effect on the outcome (e.g., predictions of anomalousness that predictive of the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic as it uses a significant amount of data that may not be available, and thus can be too sparse for a model to perform appropriately. For example, calculating a probability of anomalousness / cancer conditioned on the prior 100 CpG sites can use counts of strings in the data structure of length 100. ideally some matching exactly the prior 100 methylation states. If only sparse counts of strings of length 100 are available, there can be insufficient data to determine whether a given string of length of 100 in a test sample is anomalous or not.
[0214] FIG. 7B is a flowchart describing a process 730 for identifying anomalously methylated fragments from a subject, according to an embodiment. In process 730, the analytics system generates 740 methylation state vectors from cfDNA fragments of the subject, e.g., via the process 200 of FIG. 2A. The analytics system can handle each methylation state vector as follows.
[0215] For a given methylation state vector, the analytics system enumerates 745 all possibilities of methylation state vectors having the same starting CpG site and same length (i.e., set of CpG sites) in the methylation state vector. As each methylation state is generally either methylated or unmethylated there can be effectively two possible states at each CpG site, and thus the count of distinct possibilities of methylation state vectors can depend on a power of 2, such that a methylation state vector of length n would be associated with 2nAtty. Docket No. 32917-65128 / WO 58 P0248-WOpossibilities of methylation state vectors. With methylation state vectors inclusive of indeterminate states for one or more CpG sites, the analytics system may enumerate 730 possibilities of methylation state vectors considering only CpG sites that have observed states.
[0216] The analytics system calculates 750 the probability' of observing each possibility of methylation state vector for the identified starting CpG site and methylation state vector length by accessing the healthy control group data structure. In some embodiments, calculating the probability of observing a given possibility uses a Markov chain probability' to model the joint probability7calculation. The Markov model can be trained, at least in part, based upon evaluation of a methylation state of each CpG site in the corresponding plurality of CpG sites of the respective fragment (e.g.. nucleic acid methylation fragment) across those nucleic acid methylation fragments in a healthy noncancer cohort dataset that have the corresponding plurality of CpG sites. For example, a Markov model (e.g. , a Hidden Markov Model or HMM) is used to determine the probability that a sequence of methylation states (comprising, e g, “M” or “U”) can be observed for a nucleic acid methylation fragment in a plurality of nucleic acid methylation fragments, given a set of probabilities that determine, for each state in the sequence, the likelihood of observing the next state in the sequence. The set of probabilities can be obtained by training the HMM. Such training can involve computing statistical parameters (e.g., the probability that a first state can transition to a second state (the transition probability) and / or the probability that a given methylation state can be observed for a respective CpG site (the emission probability)), given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). HMMs can be trained using supervised training (e.g., using samples where the underlying sequence as well as the observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). In other embodiments, calculation methods other than Markov chain probabilities are used to determine the probability of observing each possibility of methylation state vector. For example, such calculation method can include a learned representation. The p-value threshold can be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold can be 0.05. The p-value threshold can be less than 0.01, less than 0.001, or less than 0.0001.
[0217] The analytics system calculates 755 a p-value score for the methylation state vector using the calculated probabilities for each possibility. In some embodiments, this includes identifying the calculated probability corresponding to the possibility that matchesAtty. Docket No. 32917-65128 / WO 59 P0248-WOthe methylation state vector in question. Specifically, this can be the possibility having the same set of CpG sites, or similarly the same starting CpG site and length as the methylation state vector. The analytics system can sum the calculated probabilities of any possibilities having probabilities less than or equal to the identified probability to generate the p-value score.
[0218] This p-value can represent the probability of observing the methylation state vector of the fragment or other methylation state vectors even less probable in the healthy control group. A low p-value score can, thereby, generally correspond to a methylation state vector which is rare in a healthy subject, and which causes the fragment to be labeled anomalously methylated, relative to the healthy control group. A high p-value score can generally relate to a methylation state vector that is expected to be present, in a relative sense, in a healthy subject. If the healthy control group is a non-cancerous group, for example, a low p-value can indicate that the fragment is anomalous methylated relative to the non-cancer group, and therefore possibly indicative of the presence of cancer in the test subject.
[0219] As above, the analytics system can calculate p-value scores for each of a plurality of methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which of the fragments are anomalously methylated, the analytics system may filter 765 the set of methylation state vectors based on their p-value scores. In some embodiments, filtering is performed by comparing the p-values scores against a threshold and keeping only those fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001 , 0.0001 , or similar.
[0220] According to example results from the process 700, the analytics system can yield a median (range) of 2,800 (1,500-12,000) fragments with anomalous methylation patterns for participants without cancer in training, and a median (range) of 3,000 (1.200-420,000) fragments with anomalous methylation patterns for participants with cancer in training. These filtered sets of fragments with anomalous methylation patterns may be used for the downstream analyses as described below in Sections III.B & III.C.
[0221] In some embodiments, the analytics system uses 760 a sliding window to determine possibilities of methylation state vectors and calculate p-values. Rather than enumerating possibilities and calculating p-values for entire methylation state vectors, the analytics system can enumerate possibilities and calculates p-values for only a window of sequential CpG sites, where the window is shorter in length (of CpG sites) than at least some fragments (otherwise, the window would serve no purpose). The window length may be static, user determined, dynamic, or otherwise selected.Atty. Docket No. 32917-65128 / WO 60 P0248-WO
[0222] In calculating p-values for a methylation state vector larger than the window, the window can identify the sequential set of CpG sites from the vector within the window starting from the first CpG site in the vector. The analytics system can calculate a p-value score for the window including the first CpG site. The analytics system can then “slide” the window to the second CpG site in the vector, and calculates another p-value score for the second window. Thus, for a window size I and methylation vector length m, each methylation state vector can generate m-l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows can be taken as the overall p-value score for the methylation state vector. In other embodiments, the analytics system aggregates the p-value scores for the methylation state vectors to generate an overall p-value score.
[0223] Using the sliding window can help to reduce the number of enumerated possibilities of methylation state vectors and their corresponding probability calculations that would otherwise need to be performed. To give a realistic example, it can be for fragments to have upwards of 54 CpG sites. Instead of computing probabilities for 2A54 (~1.8xlOA16) possibilities to generate a single p-score, the analytics system can instead use a window of size 5 (for example) which results in 50 p-value calculations for each of the 50 windows of the methylation state vector for that fragment. Each of the 50 calculations can enumerate 2A5 (32) possibilities of methylation state vectors, which total results in 50x2A5 (1.6xlOA3) probability calculations. This can result in a vast reduction of calculations to be performed, with no meaningful hit to the accurate i den ti fi cation of informative fragments.
[0224] In embodiments with indeterminate states, the analytics system may calculate a p-value score summing out CpG sites with indeterminates states in a fragment’s methylation state vector. The analytics system can identify all possibilities that have consensus with the all methylation states of the methylation state vector excluding the indeterminate states. The analytics system may assign the probability to the methylation state vector as a sum of the probabilities of the identified possibilities. As an example, the analytics system can calculate a probability of a methylation state vector of < Mi, b, Us > as a sum of the probabilities for the possibilities of methylation state vectors of < Mi, M2, U3 > and < Mi, U2, Us > since methylation states for CpG sites 1 and 3 are observed and in consensus with the fragment’s methylation states at CpG sites 1 and 3. This method of summing out CpG sites with indeterminate states can use calculations of probabilities of possibilities up to 2Ai, wherein i denotes the number of indeterminate states in the methylation state vector. In additional embodiments, a dynamic programming algorithm mayAtty. Docket No. 32917-65128 / WO 61 P0248-WObe implemented to calculate the probability of a methylation state vector with one or more indeterminate states. Advantageously, the dynamic programming algorithm operates in linear computational time.
[0225] In some embodiments, the computational burden of calculating probabilities and / or p-value scores may be further reduced by caching at least some calculations. For example, the analytics system may cache in transitory or persistent memory’ calculations of probabilities for possibilities of methylation state vectors (or windows thereof). If other fragments have the same CpG sites, caching the possibility probabilities can allow for efficient calculation of p-score values without needing to re-calculate the underlying possibility probabilities. Equivalently, the analytics system may calculate p-value scores for each of the possibilities of methylation state vectors associated with a set of CpG sites from vector (or window thereof). The analytics system may cache the p-value scores for use in determining the p-value scores of other fragments including the same CpG sites. Generally, the p-value scores of possibilities of methylation state vectors having the same CpG sites may be used to determine the p-value score of a different one of the possibilities from the same set of CpG sites.
[0226] One or more nucleic acid methylation fragments can be filtered prior to training region models or cancer classifier. Filtering nucleic acid methylation fragments can comprise removing, from the corresponding plurality of nucleic acid methylation fragments, each respective nucleic acid methylation fragment that fails to satisfy one or more selection criteria (e.g., below or above one selection criteria). The one or more selection criteria can comprise a p-value threshold. The output p-value of the respective nucleic acid methylation fragment can be determined, at least in part, based upon a comparison of the corresponding methylation pattern of the respective nucleic acid methylation fragment to a corresponding distribution of methylation patterns of those nucleic acid methylation fragments in a healthy noncancer cohort dataset that have the corresponding plurality of CpG sites of the respective nucleic acid methylation fragment.
[0227] Filtering a plurality of nucleic acid methylation fragments can comprise removing each respective nucleic acid methylation fragment that fails to satisfy a p-value threshold. The filter can be applied to the methylation pattern of each respective nucleic acid methylation fragment using the methylation patterns observed across the first plurality of nucleic acid methylation fragments. Each respective methylation pattern of each respective nucleic acid methylation fragment (e.g. , Fragment One. ... , Fragment N) can comprise a corresponding one or more methylation sites (e.g., CpG sites) identified with a methylationAtty. Docket No. 32917-65128 / WO 62 P0248-WOsite identifier and a corresponding methylation pattern, represented as a sequence of 1 's and 0's, where each “1” represents a methylated CpG site in the one or more CpG sites and each “0” represents an unmethylated CpG site in the one or more CpG sites. The methylation patterns observed across the first plurality of nucleic acid methylation fragments can be used to build a methylation state distribution for the CpG site states collectively represented by the first plurality of nucleic acid methylation fragments (e.g. , CpG site A, CpG site B, ... , CpG site ZZZ). Further details regarding processing of nucleic acid methylation fragments are disclosed in U.S. Patent Application No. 17 / 191,914, titled ‘‘Systems and Methods for Cancer Condition Determination Using Autoencoders,” filed March 4, 2021, which is hereby incorporated herein by reference in its entirety .
[0228] The respective nucleic acid methylation fragment may fail to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has an anomalous methylation score that is less than an anomalous methylation score threshold. In this situation, the anomalous methylation score can be determined by a mixture model. For example, a mixture model can detect an anomalous methylation pattern in a nucleic acid methylation fragment by determining the likelihood of a methylation state vector (e.g., a methylation pattern) for the respective nucleic acid methylation fragment based on the number of possible methylation state vectors of the same length and at the same corresponding genomic location. This can be executed by generating a plurality of possible methylation states for vectors of a specified length at each genomic location in a reference genome. Using the plurality of possible methylation states, the number of total possible methylation states and subsequently the probability of each predicted methylation state at the genomic location can be determined. The likelihood of a sample nucleic acid methylation fragment corresponding to a genomic location within the reference genome can then be determined by matching the sample nucleic acid methylation fragment to a predicted (e.g., possible) methylation state and retrieving the calculated probability7of the predicted methylation state. An anomalous methylation score can then be calculated based on the probability of the sample nucleic acid methylation fragment.
[0229] The respective nucleic acid methylation fragment can fail to satisfy a selection criterion in the one or more selection criteria when the respective nucleic acid methylation fragment has less than a threshold number of residues. The threshold number of residues can be between 10 and 50, between 50 and 100, between 100 and 150, or more than 150. The threshold number of residues can be a fixed value between 20 and 90. The respective nucleic acid methylation fragment may fail to satisfy a selection criterion in the one or more selectionAtty. Docket No. 32917-65128 / WO 63 P0248-WOcriteria when the respective nucleic acid methylation fragment has less than a threshold number of CpG sites. The threshold number of CpG sites can be 4. 5, 6, 7, 8, 9, or 10. The respective nucleic acid methylation fragment can fail to satisfy a selection criterion in the one or more selection criteria when a genomic start position and a genomic end position of the respective nucleic acid methylation fragment indicates that the respective nucleic acid methylation fragment represents less than a threshold number of nucleotides in a human genome reference sequence.
[0230] The filtering can remove a nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments that has the same corresponding methylation pattern and the same corresponding genomic start position and genomic end position as another nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments. This filtering step can remove redundant fragments that are exact duplicates, including, in some instances, PCR duplicates. The filtering can remove a nucleic acid methylation fragment that has the same corresponding genomic start position and genomic end position and less than a threshold number of different methylation states as another nucleic acid methylation fragment in the corresponding plurality of nucleic acid methylation fragments. The threshold number of different methylation states used for retention of a nucleic acid methylation fragment can be 1, 2, 3, 4, 5, or more than 5. For example, a first nucleic acid methylation fragment having the same corresponding genomic start and end position as a second nucleic acid methylation fragment but having at least 1 , at least 2, at least 3, at least 4, or at least 5 different methylation states at a respective CpG site (e.g., aligned to a reference genome) is retained. As another example, a first nucleic acid methylation fragment having the same methylation state vector (e.g, methylation pattern) but different corresponding genomic start and end positions as a second nucleic acid methylation fragment is also retained.
[0231] The filtering can remove assay artifacts in the plurality of nucleic acid methylation fragments. The removal of assay artifacts can comprise removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences that failed to undergo conversion during bisulfite conversion. The filtering can remove contaminants (e.g., due to sequencing, nucleic acid isolation, and / or sample preparation).
[0232] The filtering can remove a subset of methylation fragments from the plurality of methylation fragments based on mutual information filtering of the respective methylation fragments against the cancer state across the plurality of training subjects. For example.Atty. Docket No. 32917-65128 / WO 64 P0248-WOmutual information can provide a measure of the mutual dependence between two conditions of interest sampled simultaneously. Mutual information can be determined by selecting an independent set of CpG sites (e.g. , within all or a portion of a nucleic acid methylation fragment) from one or more datasets and comparing the probability7of the methylation states for the set of CpG sites between two sample groups (e.g., subsets and / or groups of genotypic datasets, biological samples, and / or subjects). A mutual information score can denote the probability of the methylation pattern for a first condition versus a second condition at the respective region in the respective frame of the sliding window, thus indicating the discriminative power of the respective region. A mutual information score can be similarly calculated for each region in each frame of the sliding window as it progresses across the selected sets of CpG sites and / or the selected genomic regions. Further details regarding mutual information filtering are disclosed in U.S. Patent Application 17 / 119,606, titled “Cancer Classification using Patch Convolutional Neural Networks,” filed December 11, 2020, which is hereby incorporated herein by reference in its entirety.IV. A.n. HYPERMETHYLATED FRAGMENTS AND HYPOMETHYLATED FRAGMENTS
[0233] In some embodiments, the analytics system identifies 770 hypomethylated fragments or hypermethylated fragments from the filtered set as informative fragments. The analytics system identifies hypermethylated fragments having over a threshold number of CpG sites and over a threshold percentage of the CpG sites methylated. The analytics system identifies hypomethylated fragments having over the threshold number of CpG sites and over a threshold percentage of CpG sites unmethylated. Example thresholds for length of fragments (or CpG sites) include more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Example percentage thresholds of methylation or unmethylation include more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%.IV. B. TRAINING OF CANCER CLASSIFIER
[0234] FIG. 8A is a flowchart describing a process 800 of training a cancer classifier, according to an embodiment. The analytics system obtains 810 a plurality of training samples each having a set of informative fragments and a label of a cancer signal origin. The plurality of training samples can include any combination of samples from healthy subjects with a general label of “non-cancer,” samples from subjects with a general label of “cancer” or a specific label (e.g., “breast cancer,” “lung cancer,” etc.). The training samples from subjects for one cancer signal origin may be termed a cohort for that cancer signal origin or a cancer signal origin cohort.Atty. Docket No. 32917-65128 / WO 65 P0248-WO
[0235] The analytics system determines 820, for each training sample, a feature vector based on the set of informative fragments of the training sample. The analytics system can calculate an anomaly score for each CpG site in an initial set of CpG sites. The initial set of CpG sites may be all CpG sites in the human genome or some portion thereof - which may be on the order of 104, 105, 106, 107, 108, etc. In one embodiment, the analytics system defines the anomaly score for the feature vector with a binary scoring based on whether there is an informative fragment in the set of informative fragments that encompasses the CpG site. In another embodiment, the analytes system defines the anomaly score based on a count of informative fragments overlapping the CpG site. In one example, the analytics system may use a trinary scoring assigning a first score for lack of presence of informative fragments, a second score for presence of a few informative fragments, and a third score for presence of more than a few informative fragments. For example, the analytics system counts 5 informative fragment in a sample that overlap the CpG site and calculates an anomaly score based on the count of 5.
[0236] Once all anomaly scores are determined for a training sample, the analytics system can determine the feature vector as a vector of elements including, for each element, one of the anomaly scores associated with one of the CpG sites in an initial set. The analytics system can normalize the anomaly scores of the feature vector based on a coverage of the sample. Here, coverage can refer to a median or average sequencing depth over all CpG sites covered by the initial set of CpG sites used in the classifier, or based on the set of informative fragments for a given training sample.
[0237] As an example, reference is now made to FIG. 8B illustrating a matrix of training feature vectors 822. In this example, the analytics system has identified CpG sites [K] 826 for consideration in generating feature vectors for the cancer classifier. The analytics system selects training samples [N] 824. The analytics system determines a first anomaly score 828 for a first arbitrary CpG site [kl] to be used in the feature vector for a training sample [nl]. The analytics system checks each informative fragment in the set of informative fragments. If the analytics system identifies at least one informative fragment that includes the first CpG site, then the analytics system determines the first anomaly score 828 for the first CpG site as 1, as illustrated in FIG. 8B. Considering a second arbitrary CpG site [k2], the analytics system similarly checks the set of informative fragments for at least one that includes the second CpG site [k2]. If the analytics system does not find any such informative fragment that includes the second CpG site, the analytics system determines a second anomaly score 829 for the second CpG site [k2] to be 0, as illustrated in FIG. 8B. Once theAtty. Docket No. 32917-65128 / WO 66 P0248-WOanalytics system determines all the anomaly scores for the initial set of CpG sites, the analytics system determines the feature vector for the first training sample [nl] including the anomaly scores with the feature vector including the first anomaly score 828 of 1 for the first CpG site [kl] and the second anomaly score 829 of 0 for the second CpG site [k2] and subsequent anomaly scores, thus forming a feature vector [1, 0, ...].
[0238] Additional approaches to featurization of a sample can be found in: U.S. Application No. 15 / 931,022 entitled "Model-Based Featunzation and Classification;” U.S. Application No. 16 / 579,805 entitled “Mixture Model for Targeted Sequencing;” U.S.Application No. 16 / 352,602 entitled “Anomalous Fragment Detection and Classification;” and U.S. Application No. 16 / 723,716 entitled “Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples;” all of which are incorporated by reference in their entirety. The various featurization approaches may generate distinct features to be included in a sample’s feature vector.
[0239] In some embodiments, each classifier may be trained on different matrices of training feature vectors. For example, one classifier may be trained on a first matrix covering a first set of genomic regions, whereas another classifier may be trained on a second matrix covering a second, differing set of genomic regions. In another example, one classifier may be trained on a first matrix with features determined according to one featurization approach, whereas another classifier may be trained on a second matrix with features determined according to another featurization approach.
[0240] The analytics system may further limit the CpG sites considered for use in the cancer classifier. The analytics system computes 830, for each CpG site in the initial set of CpG sites, an information gain based on the feature vectors of the training samples. From step 820, each training sample has a feature vector that may contain an anomaly score all CpG sites in the initial set of CpG sites which could include up to all CpG sites in the human genome. However, some CpG sites in the initial set of CpG sites may not be as informative as others in distinguishing between cancer signal origins, or may be duplicative with other CpG sites.
[0241] In one embodiment, the analytics system computes 830 an information gain for each cancer signal origin and for each CpG site in the initial set to determine whether to include that CpG site in the classifier. The information gain is computed for training samples with a given cancer signal origin compared to all other samples. For example, tw o random variables ‘informative fragment’ (‘IF’) and ‘cancer signal' (‘CS’) are used. In one embodiment, IF is a binary variable indicating whether there is an informative fragmentAtty. Docket No. 32917-65128 / WO 67 P0248-WOoverlapping a given CpG site in a given samples as determined for the anomaly score / feature vector above. CS is a random variable indicating whether the cancer signal has a particular origin, for example a particular organ or organ group, or a particular cancer biology7. The analytics system computes the mutual information with respect to CS given IF. That is, how many bits of information about the cancer signal origin are gained if it is known whether there is an informative fragment overlapping a particular CpG site. In practice, for a first cancer signal origin, the analytics system computes pairwise mutual information gain against each other cancer signal origin and sums the mutual information gain across all the other cancer signal origins.
[0242] For a given cancer signal origin, the analytics system can use this information to rank CpG sites based on how cancer specific they are. This procedure can be repeated for all cancer signal origins under consideration. If a particular region is commonly anomalously methylated in training samples of a given cancer but not in training samples of other cancer signal origins or in healthy training samples, then CpG sites overlapped by those informative fragments can have high information gains for the given cancer signal origin. The ranked CpG sites for each cancer signal origin can be greedily added (selected) 840 to a selected set of CpG sites based on their rank for use in the cancer classifier.
[0243] In additional embodiments, the analytics system may consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion may be that the selected CpG sites are above a threshold separation from other selected CpG sites. For example, the selected CpG sites are to be over a threshold number of base pairs away from any other selected CpG site (e.g., 100 base pairs), such that CpG sites that are within the threshold separation are not both selected for consideration in the cancer classifier.
[0244] In one embodiment, according to the selected set of CpG sites from the initial set, the analytics system may modify 850 the feature vectors of the training samples as needed. For example, the analytics system may truncate feature vectors to remove anomaly scores corresponding to CpG sites not in the selected set of CpG sites.
[0245] With the feature vectors of the training samples, the analytics system may train the cancer classifier in any of a number of ways. The feature vectors may correspond to the initial set of CpG sites from step 820 or to the selected set of CpG sites from step 850. In one embodiment, the analytics system trains 860 a binary cancer classifier to distinguish between cancer and non-cancer based on the feature vectors of the training samples. In this manner, the analytics system uses training samples that include both non-cancer samplesAtty. Docket No. 32917-65128 / WO 68 P0248-WOfrom healthy subjects and cancer samples from subjects. Each training sample can have one of the two labels “cancer” or “non-cancer.” In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.
[0246] In another embodiment, the analytics system trains 870 a multi class cancer classifier to distinguish between many cancer signal origins (also referred to as CSO labels). CSO can include one or more organs, organ groups, classes for cancer biology, cellular lineage of a cancer cell, or cancer drivers like a viral status. To do so, the analytics system can use the cancer signal origin cohorts and may also include or not include a non-cancer cohort. In this multiclass embodiment, the cancer classifier is trained to determine a cancer prediction (or, more specifically, a CSO prediction) that comprises a prediction value for each of the cancer signal origins being classified for. The prediction values may correspond to a likelihood that a given training sample (and during inference, a test sample) has each of the cancer signal origins. In one implementation, the prediction values are scored between 0 and 100, wherein the cumulation of the prediction values equals 100. For example, the cancer classifier returns a cancer prediction including a prediction value for the affected organ or organ group being breast, lung, and non-cancer. For example, the classifier can return a cancer prediction that a test samples signal origin is 65% likelihood of breast, 25% likelihood of lung, and 10% likelihood of non-cancer. The analytics system may further evaluate the prediction values to generate a prediction of a presence of one or more cancers in the sample, also may be referred to as a CSO prediction indicating one or more CSO labels, e.g., a first CSO prediction with the highest prediction value, a second CSO prediction with the second highest prediction value, etc. Continuing with the example above and given the percentages, in this example the system may determine that the sample's cancer signal origin is breast given that the CSO breast has the highest likelihood.
[0247] In both embodiments, the analytics system trains the cancer classifier by inputting sets of training samples with their feature vectors into the cancer classifier and adjusting classification parameters so that a function of the classifier accurately relates the training feature vectors to their corresponding label. The analytics system may group the training samples into sets of one or more training samples for iterative batch training of the cancer classifier. After inputting all sets of training samples including their training feature vectors and adjusting the classification parameters, the cancer classifier can be sufficiently trained to label test samples according to their feature vector within some margin of error. The analytics system may train the cancer classifier according to any one of a number of methods. As an example, the binary cancer classifier may be a L2-regularized logisticAtty. Docket No. 32917-65128 / WO 69 P0248-WOregression classifier that is trained using a log-loss function. As another example, the multicancer classifier may be a multinomial logistic regression. In practice either type of cancer classifier may be trained using other techniques. These techniques are numerous including potential use of kernel methods, random forest classifier, a mixture model, an autoencoder model, machine learning algorithms such as multilayer neural networks, etc. In some embodiments, such models may include 1,000 or more parameters, 2,000 or more parameters, 3,000 or more parameters, 4.000 or more parameters, 5,000 or more parameters, 10,000 or more parameters, 15,000 or more parameters, 20,000 or more parameters, 25,000 or more parameters, 50,000 or more parameters, 100,000 or more parameters, 150,000 or more parameters, 200,000 or more parameters, 250,000 or more parameters, 300,000 or more parameters, 350,000 or more parameters, 400.000 or more parameters, 450,000 or more parameters, or 500,000 or more parameters.
[0248] The classifier can include a logistic regression algorithm, a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm.IV. C. DEPLOYMENT OF CANCER CLASSIFIER
[0249] During use of the cancer classifier, the analytics system can obtain a test sample from a subject of unknown cancer signal origin. The analytics system may process the test sample comprised of DNA molecules with any combination of the processes 200 and 330 to achieve a set of informative fragments. The analytics system can determine a test feature vector for use by the cancer classifier according to similar principles discussed in the process 800. The analytics system can calculate an anomaly score for each CpG site in a plurality of CpG sites in use by the cancer classifier. For example, the cancer classifier receives as input feature vectors inclusive of anomaly scores for 1,000 selected CpG sites. The analytics system can thus determine a test feature vector inclusive of anomaly scores for the 1,000 selected CpG sites based on the set of informative fragments. The analytics system can calculate the anomaly scores in the same manner as the training samples. In some embodiments, the analytics system defines the anomaly score as a binary score based on whether there is a hypermethylated or hypomethylated fragment in the set of informative fragments that encompasses the CpG site.
[0250] The analytics system can then input the test feature vector into the cancer classifier. The function of the cancer classifier can then generate a cancer prediction based on the classification parameters trained in the process 800 and the test feature vector. In theAtty. Docket No. 32917-65128 / WO 70 P0248-WOfirst manner, the cancer prediction can be binary' and selected from a group consisting of “cancer’ or non-cancer;” in the second manner, the cancer prediction is selected from a group of many potential cancer signal origins that could identify affected organ or organ group and at the same time the underlying cancer biology like a histologic type, the cellular lineage of the cancer cell or origin, or the presence of a cancer driver like infection wi th an oncovirus and “non-cancer.” In additional embodiments, the cancer prediction has prediction values for each of the many cancer signal origins. Moreover, the analytics system may determine that the test sample's cancer signal has most likely one of the cancer signal origins. Following the example above with the cancer prediction for a test sample as 65% likelihood of originating from the organ breast, 25% likelihood of originating from the organ lung, and 10% likelihood of non-cancer, the analytics system may determine that the test sample is most likely to have a cancer of origin breast. In another example, where the cancer prediction is binary as 60% likelihood of non-cancer and 40% likelihood of cancer, the analytics system determines that the test sample is most likely to not have a cancer signal. In additional embodiments, the cancer prediction with the highest likelihood may still be compared against a threshold (e.g..40%, 50%, 60%, 70%) in order to call the test subject as having a cancer signal or that cancer signal origin present. If the cancer prediction w ith the highest likelihood does not surpass that threshold, the analytics system may return an inconclusive result, or report the presence of a cancer signal, while not reporting a cancer signal origin
[0251] In additional embodiments, the analytics system chains a cancer classifier trained in step 860 of the process 800 with one or more additional cancer classifiers trained in step 870 of the process 800. The analytics system can input the test feature vector into the cancer classifier trained as a binary' classifier in step 860 of the process 800. The analytics system can receive an output of a cancer prediction. The cancer prediction may be binary as to whether the test subject likely has or likely does not have cancer and if a cancer signal is detected in the sample. In other implementations, the cancer prediction includes prediction values that describe likelihood of cancer and likelihood of non-cancer. For example, the cancer prediction has a cancer prediction value of 85% and the non-cancer prediction value of 15%. The analytics system may determine the test subject to likely have cancer. Once the analytics system determines a test subject is likely to have cancer, the analytics system may input the test feature vector into a multiclass cancer classifier trained to distinguish between different cancer signal origins. The multiclass cancer classifier can receive the test feature vector and returns one or more cancer signal predictions of a cancer signal origin of the plurality of cancer signals in one or multiple categories of cancer signal origins. ForAtty. Docket No. 32917-65128 / WO 71 P0248-WOexample, the multiclass cancer classifier provides a cancer prediction specifying that the test subject is most likely to have a cancer signal originating in a predicted organ or organ group like ovaries or Fallopian tube cancer while at the same time providing a cancer signal origin prediction that this cancer signal has a cell of Mullerian lineage as origin. In another implementation, the multiclass cancer classifier provides a prediction value for each cancer signal origin of the plurality of cancer signal origins and cancer signal origin classifiers. For example, a cancer prediction may include a cancer signal origin for head and neck of 40%, a cancer signal origin for anus of 20%, and cancer signal origin of cervix of 20% for the organ of organ group, and in parallel a prediction of a cancer signal origin being 80% a carcinoma associated with a Human Papilloma Virus (HPV) infection, 10% having origin in a squamous cell carcinoma, and 10% having origin in a cancer of cells of Mullerian lineage.
[0252] According to generalized embodiment of binary cancer classification, the analytics system can determine a cancer score for a test sample based on the test sample’s sequencing data (e g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analytics system can compare the cancer score for the test sample against a binary' threshold cutoff for predicting whether the test sample likely has cancer. The binary threshold cutoff can be tuned using CSO thresholding based on one or more CSO predictions. The analytics system may further generate a feature vector for the test sample for use in the multiclass cancer classifier to determine a cancer signal origin prediction indicating one or more likely organ or organ groups as cancer signal origins, and in parallel indicating one or more histologic types, cellular lineages, or oncogenic drives as cancer signal origin.
[0253] The classifier may be used to determine the disease state of a test subject, e.g., a subject whose disease status is unknown. The method can include obtaining a test genomic data construct (e.g., single time point test data), in electronic form, that includes a value for each genomic characteristic in the plurality of genomic characteristics of a corresponding plurality7of nucleic acid fragments in a biological sample obtained from a test subject. The method can then include applying the test genomic data construct to the test classifier to thereby determine the state of the disease condition in the test subject. The test subject may not be previously diagnosed with the disease condition.
[0254] The classifier can be a temporal classifier that uses at least (i) a first test genomic data construct generated from a first biological sample acquired from a test subject at a first point in time, and (ii) a second test genomic data construct generated from a second biological sample acquired from a test subject at a second point in time.Atty. Docket No. 32917-65128 / WO 72 P0248-WO
[0255] The trained classifier can be used to determine the disease state of a test subject, e.g., a subject whose disease status is unknown. In this case, the method can include obtaining a test time-series data set, in electronic form, for a test subject, where the test timeseries data set includes, for each respective time point in a plurality of time points, a corresponding test genotypic data construct including values for the plurality of genotypic characteristics of a corresponding plurality of nucleic acid fragments in a corresponding biological sample obtained from the test subject at the respective time point, and for each respective pair of consecutive time points in the plurality of time points, an indication of the length of time between the respective pair of consecutive time points. The method can then include applying the test genotypic data construct to the test classifier to thereby determine the state of the disease condition in the test subject. The test subject may not be previously diagnosed with the disease condition.V. APPLICATIONS
[0256] In some embodiments, the methods, analytics systems and / or classifier of the present invention can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor therapeutic response or effectiveness, determine a presence or monitor minimum residual disease (MRD), or any combination thereof. For example, as described herein, a classifier can be used to generate a probability score (e.g., from 0 to 1) or scaled score or percent (e.g., from 0 to 100) describing a likelihood that a test feature vector is from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether or not the subject has a detectable cancer signal.
[0257] In one or more embodiments, the likelihood or probability score can be assessed at multiple different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). For example, the individual being monitored may be subject to recurrent sample draws. The biological sample drawn from the individual may be sequenced, e.g., by one or more sequencing devices, to measure sequencing data from the sample. The analytics system may perform various analyses, including any of the cancer classification analyses described elsewhere. In the context of monitoring progression, recurrent sample collection, sequencing, and analysis can track change in cancer signal in the individual over time. If cancer signal is increasing over time, the analytics system may infer the tumor is malignant and progressing. If the cancer signal plateaus, the analytics system may infer the tumor is benign or stagnant. In the context of treatment assessment, the analytics system may collect samples, sequence the samples, and analyze the samples to predict cancer signal before versus after treatment. If theAtty. Docket No. 32917-65128 / WO 73 P0248-WOcancer signal decreases following treatment administration, then the analytics system may infer the treatment is successful. The analytics system may further compare the change in cancer signal across a population to assess efficacy individual-to-individual. In one or more embodiments, upon detecting a significant cancer signal, the analytics system may transmit a notification to the individual (e.g., for viewing on their personal computing device) to visit a healthcare provider for diagnostic workup.
[0258] In still other embodiments, the likelihood or probability score can be used to make or influence a clinical decision (e.g., diagnostic workup of a cancer signal detection to diagnose a cancer, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe an appropriate treatment. In other embodiments, the prediction results may include a CSO prediction comprising one or more predicted CSO classes. For example, the CSO prediction may include an organ or organ group, a tumor biology, a cell of origin, or some combination thereof. Such results (or combination of results) may inform diagnostic workup options, treatment options and follow-up strategies between the healthcare provider and the subject.V. A. MULTIMODAL SAMPLE SWAP DETECTION & REMEDIATION
[0259] Multimodal sample swap detection and remediation is useful for identifying and remediation of sample swaps before cancer classifier training or after classification analyses of patient samples. If undetected sample swaps proceed through to cancer classification, insights (including classifier training results) from inaccurately labeled samples would be baked into the classifier and would affect the accuracy of future classifier results. Similarly, if undetected sample swaps proceed through to cancer classification on patient samples, results from samples belonging to other individuals would be reported. Such could lead to misinformed diagnostic workups, or could cause needless anxiety to a patient.Moreover, results of cancer classification are personal health data, and providing another’s personal health data can lead to privacy concerns. Accordingly, implementing a robust sample swap detection pipeline counteracts the above issues.
[0260] Conventionally, sample swaps would be detected by a single or a sequence of independent concordance tests. Each sample might “pass” or “fail” a concordance test, for example a comparison of assay-identified biological sex against clinically-reported biological sex. A sample may be used in a clinical study, product development, or commercial operation of a test if and only if it passes all these independent concordance tests. However, the sequential concordance screening does not empower reversing a sample swap. Moreover,Atty. Docket No. 32917-65128 / WO 74 P0248-WOthe sequential use does not usefully combine information across the disparate models into one comprehensive, more confident call.
[0261] Accordingly, the multimodal sample swap detection process described herein aims to combine multiple checks, potentially from multiple assay readouts into a multi-modal test of sample integrity. Information from copy numbers (i.e., copies of genes), small variants (e.g., single nucleotide polymorphisms), and DNA methylation (e.g., for age and disease prediction) can be combined to generate a quantitative metric for a sample that identifies the likelihood with which the sample represents known clinical information, the likelihood with which it is concordant to other samples presumed to originate from the same patient, or some combination thereof.
[0262] Furthermore, the sample swap reversal process, leveraging the graph-based optimization, is useful in identifying sample swap events and proposing reversals to the sample swaps. Applying the sample swap reversal can serve to retain called swapped samples for use in clinical studies or for product development of a molecular test, which would otherwise be deemed unusable due to the sample swap event.V.B. DETECTION OF CANCER
[0263] In some embodiments, the methods and / or classifier of the present invention are used to detect the presence or absence of cancer in a subject unsuspected of having cancer. For example, a classifier (e.g., as described above in Section III) can be used to determine a cancer prediction describing a likelihood that a test feature vector is from a subject that has cancer.
[0264] In one embodiment, a cancer prediction is a likelihood (e.g., scored between 0 and 100) for whether the test sample has cancer (i.e., binary classification). Thus, the analytics system may determine a threshold for determining whether a test subject has cancer. For example, a cancer prediction of greater than or equal to 60 can indicate that the subject has cancer. In still other embodiments, a cancer prediction greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95 indicates that the subject has cancer. In other embodiments, the cancer prediction can indicate the severity of disease. For example, a cancer prediction of 80 may indicate a more severe form, or later stage, of cancer compared to a cancer prediction below 80 (e.g., a probability score of 70). Similarly, an increase in the cancer prediction over time (e.g.. determined by classifying test feature vectors from multiple samples from the same subject taken at two or more timeAtty. Docket No. 32917-65128 / WO 75 P0248-WOpoints) can indicate disease progression or a decrease in the cancer prediction over time can indicate successful treatment.
[0265] In another embodiment, a cancer prediction comprises many prediction values, wherein each of a plurality' of different cancer signal origins being classified (i.e., multiclass classification) for has a prediction value (e.g., scored between 0 and 100). The prediction values may correspond to a likelihood that a given training sample (and during inference, training sample) has each of the cancer signal origins. The analytics system may identify the cancer signal origin that has the highest prediction value and indicate that the test subject likely has a cancer from this signal origin. In other embodiments, the analytics system further compares the highest prediction value to a threshold value (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to determine that the test subject likely has a cancer of that signal origin. In other embodiments, a prediction value can also indicate the severity' of disease. For example, a prediction value greater than 80 may indicate a more severe form, or later stage, of cancer compared to a prediction value of 60. Similarly, an increase in the prediction value over time (e.g., determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) can indicate disease progression or a decrease in the prediction value over time can indicate successful treatment.
[0266] In some embodiments, the methods and / or classifier of the present invention are used to determine a cancer source of origin prediction in a subject suspected of having cancer. For example, one or more CSO classifiers (e.g., as described above in Section IV) can be used to determine a cancer source of origin prediction to aid in diagnostic workup options.
[0267] According to aspects of the invention, the methods and systems of the present invention can be trained to detect or classify multiple cancer indications. For example, the methods, systems and classifiers of the present invention can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different ty pes of cancer.
[0268] Examples of cancers that can be detected using the methods, systems and classifiers of the present invention include carcinoma, lymphoma, blastoma, sarcoma, and leukemia or lymphoid malignancies. More particular examples of such cancers include, but are not limited to, squamous cell cancer (e.g., epithelial squamous cell cancer), skin carcinoma, melanoma, lung cancer, including small-cell lung cancer, non-small cell lung cancer (“NSCLC’), adenocarcinoma of the lung and squamous carcinoma of the lung, cancer of the peritoneum, gastric or stomach cancer including gastrointestinal cancer, pancreaticAtty. Docket No. 32917-65128 / WO 76 P0248-WOcancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high grade serous ovarian carcinoma), liver cancer (e.g., hepatocellular carcinoma (HCC)), hepatoma, hepatic carcinoma, bladder cancer (e.g., urothelial bladder cancer), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2 positive, HER2 negative, and triple negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine carcinoma, salivary gland carcinoma, kidney or renal cancer (e.g., renal cell carcinoma, nephroblastoma or Wilms’ tumor), prostate cancer, vulval cancer, thyroid cancer, anal carcinoma, penile carcinoma, head and neck cancer, esophageal carcinoma, and nasopharyngeal carcinoma (NPC).Additional examples of cancers include, without limitation, retinoblastoma, thecoma, arrhenoblastoma, hematological malignancies, including but not limited to non-Hodgkin's lymphoma (NHL), multiple myeloma and acute hematological malignancies, endometriosis, fibrosarcoma, choriocarcinoma, lary ngeal carcinomas, Kaposi's sarcoma, Schwannoma, oligodendroglioma, neuroblastomas, rhabdomyosarcoma, osteogenic sarcoma, leiomyosarcoma, and urinary’ tract carcinomas.
[0269] In some embodiments, the cancer is one or more of anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head & neck cancer, hepatobiliary’ cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, thyroid cancer, uterine cancer, or any combination thereof.
[0270] In some embodiments, the one or more cancer can be a ‘'high-signal” cancer (defined as cancers with greater than 50% 5 -year cancer-specific mortality’), such as anorectal, colorectal, esophageal, head & neck, hepatobiliary, lung, ovarian, and pancreatic cancers, as well as lymphoma and multiple myeloma. High-signal cancers tend to be more aggressive and typically have an above-average cell-free nucleic acid concentration in test samples obtained from a patient.V. C . CANCER AND TREATMENT MONITORING
[0271] In some embodiments, the cancer prediction can be assessed at multiple different time points (e.g., or before or after treatment) to make a patient prognosis, predict response to a candidate treatment, to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). For example, the present invention includes methods that involve obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point, determining a first cancer prediction therefrom (as described herein), obtaining a second test sample (e.g., a second plasma cfDNA sample) fromAtty. Docket No. 32917-65128 / WO 77 P0248-WOthe cancer patient at a second time point, and determining a second cancer prediction therefrom (as described herein).
[0272] In certain embodiments, the first time point is before a cancer treatment (e.g., before a resection surgery or a therapeutic intervention), and the second time point is after a cancer treatment (e.g., after a resection surgery7or therapeutic intervention), and the classifier is utilized to monitor the effectiveness of the treatment. For example, if the second cancer prediction decreases compared to the first cancer prediction, then the treatment is considered to have been successful. However, if the second cancer prediction increases compared to the first cancer prediction, then the treatment is considered to have not been successful. In other embodiments, both the first and second time points are before a cancer treatment (e.g., before a resection surgery or a therapeutic intervention). In still other embodiments, both the first and the second time points are after a cancer treatment (e.g., after a resection surgery or a therapeutic intervention). In still other embodiments, cfDNA samples may be obtained from a cancer patient at a first and second time point and analyzed, e g., to monitor cancer progression, to determine if a cancer is in remission (e.g., after treatment), to monitor or detect residual disease or recurrence of disease, or to monitor treatment (e.g., therapeutic) efficacy.
[0273] Those of skill in the art will readily appreciate that test samples can be obtained from a cancer patient over any desired set of time points and analyzed in accordance with the methods of the invention to monitor a cancer state in the patient. In some embodiments, the first and second time points are separated by an amount of time that ranges from about 15 minutes up to about 30 years, such as about 30 minutes, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, such as about 1, 2, 3. 4, 5, 10, 15, 20, 25 or about 50 days, or such as about 1, 2, 3, 4. 5, 6, 7, 8, 9, 10, 11, or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or about 30 years. In other embodiments, test samples can be obtained from the patient at least once every 5 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.V.D. TREATMENT
[0274] In still another embodiment, the cancer prediction can be used to make or influence a clinical decision (e.g., diagnosis of cancer, treatment selection, assessment ofAtty. Docket No. 32917-65128 / WO 78 P0248-WOtreatment effectiveness, etc.). For example, in one embodiment, if the cancer prediction (e.g., for cancer or for a particular cancer signal origin) exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., a resection surgery, radiation therapy, chemotherapy, and / or immunotherapy).
[0275] A classifier (as described herein) can be used to determine a cancer prediction that a sample feature vector is from a subject that has cancer. In one embodiment, an appropriate treatment (e.g., resection surgery or therapeutic) is prescribed when the cancer prediction exceeds a threshold. For example, in one embodiment, if the cancer prediction is greater than or equal to 60 one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75. greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more appropriate treatments are prescribed. In other embodiments, the cancer prediction can indicate the severity of disease. An appropriate treatment matching the severity of the disease may then be prescribed.
[0276] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of a chemotherapy agent, a targeted cancer therapy agent, a differentiating therapy agent, a hormone therapy agent, and an immunotherapy agent. For example, the treatment can be one or more chemotherapy agents selected from the group consisting of alkylating agents, antimetabolites, anthracy clines, anti-tumor antibiotics, cytoskeletal disruptors (taxans), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapy agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HD AC) inhibitors, retinoic receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiating therapy agents including retinoids, such as tretinoin, alitretinoin and bexarotene. In some embodiments, the treatment is one or more hormone therapy agents selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapy agents selected from the group comprising monoclonal antibody therapies such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapies and adjuvants, such as BCG, interleukin-2 (IL-2), and interferon-alfa. immunomodulating drugs, for instance, thalidomide and lenalidomide (REVLIMID). It is within the capabilities of a skilled physician orAtty. Docket No. 32917-65128 / WO 79 P0248-WOoncologist to select an appropriate cancer therapeutic agent based on characteristics such as the type of tumor, cancer stage, previous exposure to cancer treatment or therapeutic agent, and other characteristics of the cancer.V.E. KIT IMPLEMENTATION
[0277] Also disclosed herein are kits for performing the methods described above including the methods relating to the cancer classifier. The kits may include one or more collection vessels for collecting a sample from the subject comprising genetic material. The sample can include blood, plasma, serum, urine, fecal, saliva, other ty pes of bodily fluids, or any combination thereof. Such kits can include reagents for isolating nucleic acids from the sample. The reagents can further include reagents for sequencing the nucleic acids including buffers and detection agents. In one or more embodiments, the kits may include one or more sequencing panels comprising probes for targeting particular genomic regions, particular mutations, particular genetic variants, or some combination thereof. In other embodiments, samples collected via the kit are provided to a sequencing laboratory that may use the sequencing panels to sequence the nucleic acids in the sample. The WBC contamination detection may be applied to various configurations of the kit, to minimize WBC contamination potentially originating from components of the kit. For example, experiments may be run comparing types of collection vessels. WBC contamination can be assessed and compared between the types of collection vessels to identify an optimal type that minimizes WBC contamination.
[0278] A kit can further include instructions for use of the reagents included in the kit. For example, a kit can include instructions for collecting the sample, extracting the nucleic acid from the test sample. Example instructions can be the order in which reagents are to be added, centrifugal speeds to be used to isolate nucleic acids from the test sample, how to amplify nucleic acids, how to sequence nucleic acids, or any combination thereof. The instructions may further illuminate how to operate a computing device as the analytics system 200, for the purposes of performing the steps of any of the methods described.
[0279] In addition to the above components, the kit may include computer-readable storage media storing computer software for performing the various methods described throughout the disclosure. One form in which these instructions can be present is as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert. Y et another means would be a computer readable medium, e.g., diskette, CD, hard-drive, network data storage, on which the instructions have been stored in the form of computer code. Yet another meansAtty. Docket No. 32917-65128 / WO 80 P0248-WOthat can be present is a website address or QR code which can be used via the internet to access the information at a removed site.VI. ADDITIONAL CONSIDERATIONS
[0280] The foregoing detailed description of embodiments refers to the accompanying drawings, which illustrate specific embodiments of the present disclosure. Other embodiments having different structures and operations do not depart from the scope of the present disclosure. The term "the invention” or the like is used with reference to certain specific examples of the many alternative aspects or embodiments of the applicants’ invention set forth in this specification, and neither its use nor its absence is intended to limit the scope of the applicants’ invention or the scope of the claims.
[0281] Embodiments of the invention may also relate to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, and / or it may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may¬ be stored in a non-transitory. tangible computer readable storage medium, or any t pe of media suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing systems referred to in the speci fi cation may include a single processor or may be architectures employing multiple processor designs for increased computing capability.
[0282] Any of the steps, operations, or processes described herein as being performed by the analytics system may be performed or implemented with one or more hardware or software modules of the apparatus, alone or in combination with other computing devices. In one embodiment, a software module is implemented with a computer program product comprising a computer-readable medium containing computer program code, which can be executed by a computer processor for performing any or all of the steps, operations, or processes described.Atty. Docket No. 32917-65128 / WO 81 P0248-WO
Claims
CLAIMSWlIAT IS CLAIMED IS:
1. A method for multimodal sample swap detection, the method comprising: obtaining a sample labeled as belonging to an individual, the sample comprising sequencing data derived from sequencing nucleic acid fragments in the sample, and one or more reported covariate labels derived from clinical information on the individual;determining a plurality of genetic features from the sequencing data; applying a plurality of sample swap detection models to the plurality of genetic features for the sample and to the one or more reported covariate labels to output a plurality7of predictions indicating whether the sample is mislabeled, each sample swap detection model being distinctly configured; applying an aggregation model to the plurality of predictions to determine an aggregate prediction indicating whether the sample is mislabeled; calling a sample swap event in response to determining that the aggregate prediction is above a threshold; andperforming a remedial workflow to address the sample swap event.
2. The method of claim 1, wherein the sample is a liquid biopsy sample comprising nucleic acid fragments shed in bodily fluid, or a tissue biopsy sample comprising nucleic acid fragments extracted from tissue cells.
3. The method of claim 1 ,wherein the sequencing data comprises sequencing data in a plurality of forms including methylation sequencing data, small variant sequencing data, copy number sequencing data, or some combination thereof, and wherein a first sample swap detection model is configured to input one or more genetic features derived from a first form of sequencing data to output a first prediction of whether the sample is mislabeled and a second sample swap detection model is configured to input one or more genetic features derived from a second form of sequencing data, that is different from the first form of sequencing data, to output a second prediction of whether the sample is mislabeled.Atty. Docket No. 32917-65128 / WO 82 P0248-WO4. The method of claim 1 , wherein applying the plurality of sample swap detection models comprises:applying one sample swap detection model to predict a label for one covariate based on the plurality7of genetic features;comparing the predicted label to the reported label for the covariate; and determining the prediction of whether the sample is mislabeled based on the comparison of the predicted label to the reported label.
5. The method of claim 4,wherein determining a plurality of genetic features comprises determining a plurality7of methylation features from methylation sequencing data of the sequencing data; andwherein applying the one sample swap detection model comprises applying the one sample swap detection model to the plurality7of methylation features to predict the label for biological sex as the covariate.
6. The method of claim 4.wherein determining a plurality of genetic features comprises determining a plurality' of methylation features from methylation sequencing data of the sequencing data; andwherein applying the one sample swap detection model comprises applying the one sample swap detection model to the plurality7of methylation features to predict the label for epigenetic age as the covariate.
7. The method of claim 4,wherein determining a plurality' of genetic features comprises determining a plurality7of methylation features from methylation sequencing data of the sequencing data; andwherein applying the one sample swap detection model comprises applying the one sample swap detection model to the plurality7of methylation features to predict the label for ancestry as the covariate.Atty. Docket No. 32917-65128 / WO 83 P0248-WO8. The method of claim 4, wherein applying the aggregation model comprises applying the aggregation model further to the predicted label for the covariate to determine the aggregate prediction.
9. The method of claim 1,wherein determining the plurality of genetic features from the sequencing data comprises determining a first genotype signature for the sample based on single nucleotide polymorphism (SNP) genotypes across a plurality of SNP sites in a reference genome;wherein applying the plurality of sample swap detection models comprises: applying one sample swap detection model to (i) the first genoty pe signature for the sample and (ii) a second genotype signature for a reference sample also labeled as belonging to the individual, to output a genetic similarity score between the sample and the reference sample;comparing the genetic similarity score to a threshold score; anddetermining the prediction of whether the sample is mislabeled based on the comparison of the genetic similarity score to the threshold score.
10. The method of claim 1, wherein performing the remedial workflow to address the sample swap event comprises one or more of:labeling the sample as unusable;generating and transmitting a notification to a user to obtain anew sample from the individual;withholding the sample from cancer classification;withholding the sample from reporting to the individual;removing the sample from a training data set for training of a cancer classification machine-learning model; andidentifying a correct origination of the sample.
11. A method for sample swap detection, the method comprising:obtaining a plurality of samples including a target sample, each sample is labeled as belonging to one of a plurality of individuals, the sample comprisingAtty. Docket No. 32917-65128 / WO 84 P0248-WOsingle nucleotide polymorphism (SNP) sequencing data derived from sequencing nucleic acid fragments in the sample;determining, for each sample, a genotype signature based on SNP genotypes from the sequencing data of the sample across a plurality of SNP sites in a reference genome;generating a plurality of groups, wherein each group associates samples labeled as belonging to one of the plurality of individuals;for each of one or more pairs of samples in one group, applying a genetic similarity model to the genotype signatures of the pair of samples to output a genetic similarity score between the pair of samples; and calling a sample swap event in response to determining that one sample of one group is mislabeled based on a comparison of the one or more genetic similarity scores to one or more other samples in the one group to a threshold score.
12. The method of claim 11, wherein determining the genotype signature for each sample based on the SNP genotypes from the sequencing data comprises:assigning a first value for a homozygous reference genotype, a second value for a heterozygous genotype, a third value for a homozy gous alternate genoty pe, and a fourth value for a missing genotype.
13. The method of claim 12, wherein applying the genetic similarity model to the genotype signatures of the pair of samples comprises:identifying a common subset of the plurality of SNP sites with both genotype signatures not having the fourth value for the missing genotype; identifying a count of SNP sites of the common subset of SNP sites with both genotype signatures having matching values; anddetermining the genetic similarity score for the pair of samples based on a ratio of the count of SNP sites having the matching values to a total number of SNP sites in the common subset.
14. The method of claim 11, wherein determining the one sample of the one group is mislabeled comprises:Atty. Docket No. 32917-65128 / WO 85 P0248-WOdetermining a majority of the genetic similarity scores between the one sample and other samples of the group is below the threshold score.
15. The method of claim 11, further comprising:generating a graph comprising a plurality of nodes representing the plurality of samples and a plurality7of edges connecting nodes of a group, wherein each edge represents the genetic similarity score between the samples connected by the edge.
16. The method of claim 11, further comprising:performing a remedial workflow to address the sample swap event.
17. The method of claim 16, wherein performing the remedial workflow to address the sample swap event comprises:performing graph optimization to identify which of the other groups maximizes genetic similarity with the one sample; andrelabeling the one sample to be part of a second group that maximizes genetic similarity with the one sample.
18. The method of claim 17, wherein performing the graph optimization comprises:identifying one other sample also determined to be mislabeled; and assessing genetic similarity7of the one sample with the group of the other sample also determined to be mislabeled.
19. The method of claim 18, wherein assessing genetic similarity of the one sample with the group of the other sample also determined to be mislabeled comprises:applying the genetic similarity model to the genotype signature of a third sample in the group of the other sample determined to be mislabeled to output a genetic similarity score between the one sample and the third sample.
20. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-19.Atty. Docket No. 32917-65128 / WO 86 P0248-WO21. A system comprising:one or more computer processors; andthe non-transitory computer-readable storage medium of claim 20.
22. A treatment kit comprising:a collection vessel for collecting a DNA sample from a subject;optionally, one or more reagents for isolating DNA fragments in the DNA sample; optionally, one or more probes targeting one or more genomic loci determined to be indicative of cancer status; andthe non-transitory' computer-readable storage medium of claim 20.Atty. Docket No. 32917-65128 / WO 87 P0248-WO