Cancer detection through integrated analysis of whole-genome sequencing

The method uses WGS and machine learning to enhance ctDNA detection by accurately identifying somatic mutations, addressing inaccuracies in current NGS technologies and improving cancer prognosis through early detection of minimal residual disease.

JP2026516667APending Publication Date: 2026-05-26PERSONAL GENOME DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PERSONAL GENOME DIAGNOSTICS INC
Filing Date
2024-04-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Current NGS technologies face challenges in reliably detecting somatic mutations, particularly subclonal or low-purity mutations from tumor samples, and distinguishing them from germline mutations or technical artifacts, leading to inaccuracies in cancer detection, especially in low-abundance cell-free circulating tumor DNA (ctDNA).

Method used

A computer implementation method using whole-genome sequencing (WGS) generates sequence reads from tumor, non-cancerous, and non-tissue samples, applies a classification machine learning model to score candidate somatic variants, and determines ctDNA status, thereby improving sensitivity and specificity in detecting ctDNA.

Benefits of technology

The method enhances the detection of ctDNA, enabling early identification of minimal residual disease and improving patient prognosis by accurately distinguishing somatic mutations from germline mutations and technical artifacts, thus guiding treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516667000001_ABST
    Figure 2026516667000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a technique for identifying tumor-specific mutations through integrated analysis of next-generation sequencing data using machine learning models. In certain embodiments, a computer implementation method is provided, comprising: generating sequence reads from one or more samples collected from the same patient; generating variant calling files by analyzing the sequence reads corresponding to each of the one or more samples; generating a list of candidate somatic variants by comparing the variant calling files; generating a score for each of the candidate somatic variants in the list of candidate somatic variants using a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model; determining the ctDNA status for the patient based on the score, wherein the ctDNA status is either positive or negative; and generating a report providing the ctDNA status for the patient.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority and benefit of U.S. Provisional Application No. 63 / 496,643, filed on 17 April 2023, and U.S. Provisional Application No. 63 / 501,219, filed on 10 May 2023, the entire contents of each of these applications being incorporated herein by reference for all purposes.

[0002] This disclosure relates to a cancer detection technology that utilizes machine learning models to identify tumor-specific mutations through integrated analysis of next-generation sequencing data. [Background technology]

[0003] Next-generation sequencing (NGS) technology, with its massive parallel sequencing capabilities, has revolutionized routine diagnostics for detecting mutations in clinical laboratories worldwide. Whole-genome sequencing (WGS) is a comprehensive NGS method for analyzing the entire genome (sequencing all or virtually all of the 3 billion DNA base pairs that make up the entire genome by determining the order of nucleotides (A, C, G, T)). The purpose of WGS is typically to search for genetic abnormalities (e.g., single-nucleotide variants, deletions, insertions, and structural variants). Because the entire genome is sequenced, changes in non-coding or intronic regions of the genome can also be determined. WGS has had a particularly significant impact in the field of oncology, for detecting tumor-specific (somatic) mutations and assisting oncologists in making decisions regarding patient diagnosis and treatment management.

[0004] In addition to standard high-coverage NGS approaches, typically used low-coverage WGS (1x to 10x) and very-low-coverage WGS (less than 1x coverage) have been developed for the analysis of low-quality / enriched DNA samples, such as cell-free circulating tumor DNA (ctDNA) in blood or plasma samples. Low-coverage and very-low-coverage WGS can accurately assess common genetic mutations and large subchromosomal and whole-chromosome events using approximately 0.4x sequencing coverage on circulating tumor DNA (ctDNA).

[0005] Cell-free DNA (cfDNA) is DNA released by apoptotic or necrotic cells that circulates throughout an individual's body. cfDNA can be isolated from blood, plasma, sputum, saliva, cerebrospinal fluid, surgical drain fluid, urine, cystic fluid, and other sources. cfDNA isolated from non-cancerous individuals primarily consists of leukocyte-derived DNA; however, individuals with cancer may also possess ctDNA. As a tumor grows in a person's body, small fragments of DNA from the tumor may be found circulating in the person's bloodstream. This ctDNA carries information such as tumor-specific mutations and structural changes. For decades, researchers and clinicians have used ctDNA from the bloodstream of cancer patients to facilitate therapy selection, identify drug resistance, and monitor treatment responses by detecting oncological signals through the measurement of genomic instability. For example, one way clinicians monitor the effectiveness of therapy and predict cancer recurrence is to detect and measure ctDNA levels before, during, and after surgical and therapeutic procedures. This practice is often referred to by physicians as minimal or molecular residual disease (MRD) surveillance.

[0006] Despite these applications, challenges remain regarding the reliability of NGS and WGS for detecting somatic mutations, particularly subclonal or low-purity mutations originating from tumor samples. Furthermore, challenges in distinguishing somatic mutations from germline mutations or technical artifacts have raised concerns about the overall accuracy of NGS methods. [Overview of the Initiative]

[0007] In various embodiments, a computer implementation method is provided, comprising: generating sequence reads from tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples collected from the same patient, wherein the sequence reads are generated using whole-genome sequencing (WGS); generating tumor variant calling files, non-cancerous variant calling files, and non-tissue variant calling files by analyzing the sequence reads corresponding to the tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue samples, respectively; generating a list of somatic variants by comparing the tumor variant calling files with the non-cancerous variant calling files; generating a list of candidate somatic variants by comparing the list of somatic variants with the non-tissue variant calling files; generating a score for each of the candidate somatic variants in the list of candidate somatic variants by a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model; determining the ctDNA status for the patient based on the score, wherein the ctDNA status is either positive or negative; and generating a report providing the ctDNA status for the patient.

[0008] In some embodiments, a tumor nucleic acid sample is any body tissue or fluid containing nucleic acids that are considered cancer-positive, a non-cancerous sample is any body tissue or fluid containing nucleic acids that are considered not cancerous, and a non-tissue sample is any fluid containing nucleic acids that are considered to contain cell-free DNA and circulating tumor DNA.

[0009] In some embodiments, the tumor nucleic acid sample is cancer-positive tissue, the non-cancerous nucleic acid sample is leukocytes, and the non-tissue nucleic acid sample is plasma.

[0010] The computer implementation method according to claim 3, wherein in some embodiments the non-tissue nucleic acid sample is circulating tumor DNA.

[0011] In some embodiments, non-cancerous nucleic acid samples and non-tissue nucleic acid samples are collected from the same whole blood sample.

[0012] In some embodiments, tumor nucleic acid samples are sequenced to a depth of at least 50 times, non-cancerous nucleic acid samples are sequenced to a depth of at least 30 times, and non-tissue nucleic acid samples are sequenced to a depth of at least 20 times.

[0013] In some embodiments, tumor nucleic acid samples are sequenced to a depth of 80x, non-cancerous nucleic acid samples to a depth of 40x, and non-tissue nucleic acid samples to a depth of 30x.

[0014] In some embodiments, the patient is diagnosed with cancer, undergoes surgery to remove one or more tumors, and receives therapeutic treatment after the surgery.

[0015] In some embodiments, the therapeutic treatment is adjuvant chemotherapy.

[0016] In some embodiments, the patient is diagnosed with colorectal cancer, head and neck cancer, lung cancer, breast cancer, or melanoma.

[0017] In some embodiments, the patient is diagnosed with colorectal cancer.

[0018] In some embodiments, tumor nucleic acid samples, non-cancerous samples, and non-tissue samples are collected (i) before surgery, (ii) during surgery, (iii) approximately 3 to 65 days after surgery and before therapeutic treatment, (iv) approximately every 6 months for up to 3 years after surgery and after therapeutic treatment, or (v) any combination thereof.

[0019] In some embodiments, the tumor variant call file and the non-cancerous variant call file are filtered using a set of filtering criteria, the set of filtering criteria including removing (i) variants annotated as having low confidence, (ii) variants annotated as indels, (iii) variants observed in genomic databases, (iv) variants overlapping simple tandem repeat tracts, (v) variants at genomic positions having less than 10-fold coverage, (vi) variants at genomic positions where the number of alternate alleles is less than 4 in the tumor nucleic acid sample or greater than 1 in the non-cancerous nucleic acid sample, (vii) variants having a variant allele frequency less than 0.05, or (viii) any combination thereof.

[0020] In some embodiments, the list of candidate somatic variants includes substitutions, small indels, chromosomal rearrangements, copy number variations, microsatellite instabilities, or any combination thereof.

[0021] In some embodiments, the list of candidate somatic variants includes at least 40,000 to at least 70,000 somatic variants.

[0022] In some embodiments, each candidate somatic variant on the list of candidate somatic variants has at least 50 corresponding features.

[0023] In some embodiments, the features include quality metrics output from sequencing, alignment, and variant calling.

[0024] In some embodiments, the sequencing features include the quality score of any given base in the sequence reads, the alignment features include the alignment quality, read quality, strand information, metrics regarding the complexity of the region within the genome, or any combination thereof, and the variant calling features include the variant confidence score, the quality of the base variant, or any combination thereof.

[0025] In some embodiments, before generating the score, the classification model uses a set of non-cancerous donor samples to filter a list of candidate somatic variants to generate a filtered list of candidate somatic variants.

[0026] In some embodiments, the classification machine learning model is a random forest classifier that includes a set of trees having at least 500 decision trees, each of the trees generates a score for an input candidate somatic variant, the random forest classifier averages the scores generated by each of the trees to determine a final score, compares the final score with a predetermined threshold to determine whether the ctDNA status of the non-tissue nucleic acid sample is positive or negative, the set of trees considers at least 50 features associated with the candidate somatic variants, and each tree makes a class prediction considering a different subset of features from at least 50 features.

[0027] In some embodiments, the predetermined threshold is the maximum normalized score + 1 standard deviation of the cohort of reference variants.

[0028] In some embodiments, the final score is greater than or equal to the predetermined threshold, the ctDNA status is positive, the final score is less than the predetermined threshold, and the ctDNA status is negative.

[0029] In some embodiments, the ctDNA status is determined by normalizing the score and comparing the normalized score with the maximum normalized score + 1 standard deviation, and if the normalized score is greater than or equal to the maximum normalized score, the ctDNA status is positive.

[0030] In some embodiments, the ctDNA status represents the postoperative ctDNA status.

[0031] In some embodiments, ctDNA status correlates with clinicopathological risk factors for predicting survival, where the clinicopathological risk factors predict the risk of recurrence, and these clinicopathological risk factors include the depth of tumor invasion and the extent of tumor spread to adjacent lymph nodes.

[0032] In some embodiments, the report includes a correlation between ctDNA status and clinicopathological risk factors, and the report further explains the patient's relapse risk and predicted survival rate based on the patient's ctDNA status and clinicopathological risk factors.

[0033] In various embodiments, a computer implementation method is provided, comprising: generating sequence reads from a non-tissue nucleic acid sample collected from a patient, wherein the sequence reads are generated using whole-genome sequencing (WGS); generating a non-tissue variant calling file by analyzing the sequence reads corresponding to the non-tissue sample; generating a list of candidate somatic variants by comparing a list of somatic variants with the non-tissue variant calling file; generating a score for each of the candidate somatic variants in the list of candidate somatic variants by a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model; determining the ctDNA status for the patient based on the score, wherein the ctDNA status is either positive or negative; and generating a report providing the ctDNA status for the patient.

[0034] In various embodiments, a computer implementation method comprising accessing a labeled training dataset, wherein the labeled training dataset includes ground truth true positive variants and associated features collected from patients with cancer, and ground truth false positive variants and associated features collected from patients without cancer, and training a classification model using the labeled training dataset to generate scores, wherein training is an iterative process starting at a first node of a first tree, inputting a portion of the labeled training data into the classification model, and inputting several variants from a portion of the labeled training dataset. A computer implementation method is provided which includes generating, which includes randomly selecting features of a variant, determining which variant features from several variant features provide the best binary partition, the determination being based on a subset of variant features that minimize an objective function, and assigning the determined variant features to a first node; repeating the iterative process for several iterations or epochs at the second and subsequent nodes of the classification model; repeating the iterative process at the first node of the second and subsequent trees until all variant features have been assigned to the tree; and outputting the trained classification model.

[0035] In some embodiments, a system is provided which includes one or more processors and a memory coupled to the one or more processors for storing a plurality of instructions, wherein when an instruction is executed by one or more processors, the system causes one or more processors to perform one of the methods disclosed herein.

[0036] In some embodiments, a computer program product is provided which is tangibly embodied in non-temporary computer-readable memory containing instructions, and which, when an instruction is executed by one or more processors, causes one or more processors to perform one of the methods disclosed herein.

[0037] The terms and expressions used are for illustrative purposes only, not to limit, and the use of such terms and expressions is not intended to exclude any equivalents of the exhibited and described features or any part thereof, but it should be recognized that various modifications are possible within the scope of the claimed invention. Accordingly, although the invention is specifically disclosed by embodiments and any features, it should be understood that modifications and variations of the concepts disclosed herein can be made by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawing]

[0038] The drawings illustrate, but do not limit, certain embodiments of the technology. For clarity and ease of explanation, the drawings are not made to scale, and in some examples, various aspects may be exaggerated or enlarged to facilitate understanding of a particular embodiment.

[0039] [Figure 1] This paper presents statistical data related to postoperative ACT treatment for outcomes in patients with stage III colon cancer. [Figure 2] This document illustrates computing environments in various embodiments. [Figure 3] This document describes exemplary sample processing and bioinformatics workflows for detecting ctDNA in non-tissue samples using various embodiments. [Figure 4] An exemplary block diagram of a machine learning pipeline is shown, which includes several subsystems that work together to train, validate, and implement one or more machine learning models in various embodiments. [Figure 5A-5B] This document illustrates exemplary workflows for using a machine learning pipeline during the inference phase (Figure 5A) and the training of a classification model (Figure 5B) in various embodiments. [Figure 6] This document provides illustrative examples of random forest machine learning models in various embodiments. [Figure 7]Examples of computing environments for implementing the disclosed technology, in various embodiments, are shown. [Figure 8] This section outlines the PROVENC3 study. Figure 8A shows the PROVENC3 study population and the main exclusion criteria for the final analysis. Figure 8B is an exemplary schematic diagram of the PROVENCE3 study design, showing the number of patients analyzed for each research topic under various embodiments. [Figure 9] Dot graphs (Figure 9A) and box plots (Figure 9B) of postoperative ctDNA status and cfDNA concentration are shown for various embodiments. [Figure 10A] This diagram illustrates exemplary tumor-based detection of ctDNA via integrated WGS analysis and machine learning modeling techniques using various embodiments. [Figure 10B] This paper presents analytical sensitivity studies conducted using a devised reference model derived from commercially available cell lines, in various embodiments. [Figure 10C] This document presents analytical specificity studies conducted according to various embodiments. [Figure 10D] The estimated tumor fractions across 45 independent runs evaluated for the PROVENC3 clinical study (n=45 runs, CV=7.2%) demonstrate high reproducibility with an externally designed reference control sample (SeraSeq gDNA TMB-mix score 26). [Figure 11A] This graph shows the analytical specificity across non-cancerous donor plasma samples. [Figure 11B] This graph shows the analytical sensitivity of the method according to various embodiments. [Figure 12] The results of evaluating ctDNA status across multiple solid tumor types using various embodiments are presented. [Figures 13A-13E]This study demonstrates that in ACT-treated stage III colon cancer across various embodiments, postoperative ctDNA detection is independently associated with 3-year recurrence. Figure 13A shows Kaplan-Meier estimates of TTR stratified by postoperative ctDNA status, and Figure 13B shows the proportion of patients at risk of 3-year recurrence. Figure 13C shows Kaplan-Meier estimates of TTR stratified by clinicopathological risk, and Figure 13D shows the proportion of patients at risk of 3-year recurrence. Figure 13E shows Kaplan-Meier estimates of TTR stratified by clinicopathological risk and ctDNA status. [Figure 13F] This shows the percentage of patients at risk of recurrence after 3 years. Abbreviations: ctDNA, circulating tumor DNA; ACT, adjuvant chemotherapy; TTR, time to recurrence; ctDNA-pos, ctDNA positive; ctDNA-neg, ctDNA negative; n, number; Clin.path., clinicopathological; HR, hazard ratio. [Figure 14] Kaplan-Meier estimates of Cox regression analysis, including confidence intervals, for clinicopathologically low-risk (Figure 14A) and high-risk (Figure 14B) groups stratified by postoperative ctDNA status, are shown for various embodiments. [Figure 15] This figure shows the time to recurrence based on ctDNA status in various embodiments. Figure 15A shows Kaplan-Meier estimates of TTR stratified by postoperative ctDNA status for all patients experiencing disease recurrence. Figure 15B is a swimmer plot for all patients with recurrence (n=47). Figure 15C is a Kaplan-Meier estimate of TTR stratified by postoperative ctDNA status. Figure 15D shows the proportion of patients at risk of recurrence at 3 years. Abbreviations: TTR, time to recurrence; ctDNA, circulating tumor DNA; ACT, adjuvant chemotherapy; HR, hazard ratio.

[0040] term Where used herein, the singular “a,” “an,” and “the” include plural references unless the context clearly indicates otherwise. For example, a reference to “method” includes one or more methods and / or steps of the type described herein, which would be obvious to those skilled in the art by reading this disclosure, etc. In addition, the term “nucleic acid” includes multiple nucleic acids, including mixtures thereof.

[0041] The terms “about” and “approximately” are used interchangeably and mean within an acceptable margin of error of a particular value as determined by those skilled in the art, and thus partially depend on how the value is measured or determined, for example, the limits of the measuring system. For example, the terms “substantially,” “approximately,” or “about” may be replaced with “within [percentage]” of a given, where the percentages include 0.1, 1, 5, and 10 percent. Where a particular value is stated in this application and claims, unless otherwise stated, the term “about” means within an acceptable margin of error of that particular value.

[0042] As used herein, the term “allele” refers to any alternative form of a gene at a particular locus. One or more alternative forms may exist, all of which may be associated with a single trait or characteristic at a particular locus. In the diploid cells of an organism, the alleles of a given gene may be located at a specific position or at a locus on a chromosome (multiple loci). Different gene sequences between different alleles at each locus are referred to as “variants,” “polymorphisms,” or “mutations.” The term “single nucleotide polymorphism (SNP)” is used throughout as interchangeable with “single nucleotide variant (SNV).”

[0043] As used herein, the terms “allele frequency” or “allelic frequency” generally refer to the relative frequency of alleles (e.g., gene variants) in a sample, expressed, for example, as a fraction or percentage. In some cases, allele frequency may refer to the relative frequency of alleles (e.g., gene variants) in a sample, such as a CFNA sample. In some cases, allele frequency may refer to the relative frequency of alleles (e.g., gene variants) in a sample, such as a CFNA standard. The allele frequency of a mutant allele may refer to the frequency of the mutant allele relative to the wild-type allele in a sample, such as a cell-free nucleic acid sample. For example, if a sample contains 100 copies, of which 5 are mutant alleles and 95 are wild-type alleles, the allele frequency of that mutant allele is approximately 5 / 100 or approximately 5%. A sample that does not contain any copies of the mutant allele (e.g., an allele frequency of approximately 0%) can be used, for example, as a negative control. A negative control might be a sample in which the mutant allele is not expected to be detected. A sample containing the mutant allele at approximately 50% allele frequency may, for example, represent a germline heterozygous mutation.

[0044] Cancer refers to an abnormal state or condition characterized by rapidly growing cells. Rapidly growing cells can be classified as pathological (i.e., characterizing or constituting a disease state) or non-pathological (i.e., a deviation from normal but not associated with a disease state). Generally, cancer is associated with the presence of one or more tumors (i.e., abnormal cell masses). In addition, cancer cells can spread locally or to other parts of the body via the bloodstream and lymphatic system. Examples of cancer include malignant tumors of various organ systems such as lung cancer, breast cancer, thyroid cancer, lymphatic system cancer, gastrointestinal cancer, and genitourinary cancer. Cancer can also refer to adenocarcinomas, including malignant tumors such as colon cancer, renal cell carcinoma, prostate cancer and / or testicular tumors, non-small cell lung cancer, small intestine cancer, and esophageal cancer. Carcinomas are malignant tumors of epithelial or endocrine tissue, including respiratory carcinomas, gastrointestinal carcinomas, genitourinary carcinomas, testicular carcinomas, breast carcinomas, prostate carcinomas, endocrine carcinomas, and melanomas. "Adenocarcinoma" refers to carcinomas originating from glandular tissue or carcinomas in which tumor cells form recognizable glandular structures. "Sarcoma" refers to malignant tumors of mesenchymal origin. "Melanoma" refers to tumors arising from melanocytes. Melanoma most commonly occurs in the skin and is often observed to metastasize extensively.

[0045] The terms “cell-free nucleic acid” or “CFNA” refer to extracellular nucleic acids and circulating free nucleic acids. Therefore, the terms “cell-free nucleic acid,” “cell-free nucleic acid,” and “circulating free nucleic acid” are used interchangeably. Extracellular nucleic acids can be found in biological sources such as blood, urine, and feces. CFNA may refer to cell-free DNA (cfDNA), circulating free DNA (cfDNA), cell-free RNA (cfRNA), or circulating free RNA (cfRNA). CFNA may result from the detachment of nucleic acids from cells undergoing apoptosis or necrosis. Previous studies have demonstrated that CFNA, e.g., cfDNA, may be present at steady-state levels and may increase with cytotoxicity or necrosis. In some cases, CFNA detaches from abnormal or unhealthy cells, such as tumor cells. CFDNA, which is detached from tumor cells, commonly referred to as ctDNA in some cases, can be distinguished from cfDNA detached from normal or non-cancer cells by using genomic information, such as by identifying genetic modifications including mutations and / or structural changes that distinguish normal and abnormal cells, as well as additional discriminant factors such as polynucleotide length, terminal position, and base modifications (e.g., methylation, hydroxymethylation, formylation, carboxylation, etc.). In some cases, CFNA is detached from fetal-related cells into maternal circulation. In some cases, CFNA may originate from pathogens infecting a host, such as a subject (e.g., a patient).

[0046] The terms nucleic acid or nucleotide refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and their polymers, in either single-stranded or double-stranded forms. Unless otherwise specified, the term encompasses nucleic acids containing known analogues of natural nucleotides having properties equivalent to a reference nucleic acid. Nucleic acid sequences may include combinations of deoxyribonucleic acid and ribonucleic acid. Such deoxyribonucleic acid and ribonucleic acid include both naturally occurring molecules and synthetic analogues. Nucleic acids also encompass all forms of sequences, including but not limited to single-stranded, double-stranded, hairpin, stem, and loop structures.

[0047] The terms “variant” or “variant,” when applied to alleles or sequences, generally refer to alleles or sequences that do not encode the most common phenotype in a particular natural population. The terms “variant allele” and “variant allele” can be used interchangeably. In some cases, a variant allele may refer to an allele that is present at a lower frequency in a population than the wild-type allele. In some cases, a variant allele or sequence may refer to an allele or sequence that has mutated from a wild-type sequence to a variant sequence, exhibiting a phenotype associated with a disease state and / or drug resistance state. Variant alleles and sequences may differ from wild-type alleles and sequences by only one base, but can differ by up to several bases or more. The term “variant,” when applied to genes, generally refers to one or more sequence variations in a gene, including point mutations, SNPs, insertions, deletions, substitutions, translocations, copy number variations, or other genetic variations, changes, or sequence modifications.

[0048] The terms “polynucleotide,” “nucleic acid,” and “nucleotide” are used interchangeably. They refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or their analogues. The following are non-exclusive examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci (may be multiple) defined by binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cfDNA and cell-free RNA (cfRNA), nucleic acid probes, and primers. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogues. Modifications to the nucleotide structure, where present, may be conferred before or after the assembly of the polymer, and the sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, for example, by conjugation with labeling components.

[0049] As used herein, the terms “standard” or “reference” generally refer to a substance prepared to a specific, predefined standard that can be used, for example, to evaluate a particular aspect of an assay. A standard or reference is preferably reproducible, consistent, and reliable. These aspects may include, but are not limited to, performance metrics, such as accuracy, specificity, sensitivity, linearity, reproducibility, detection limit, and / or quantification limit. Standards or references may be used for assay development, assay validation, and / or assay optimization. Standards can be used to evaluate quantitative and qualitative aspects of an assay. It should be understood that a standard may be used in any application where a defined reference is needed and / or useful. In some aspects, applications may include monitoring, comparing, and / or otherwise evaluating QC samples / controls, assay controls (products), packing samples, training samples, and / or lot-to-lot performance for a given assay.

[0050] Generally, the term “sequence variant” refers to any variation of a sequence relative to one or more reference sequences. Typically, sequence variants occur less frequently than the reference sequences in a given population of individuals for which the reference sequences are known. In some cases, the reference sequence is a single known reference sequence, such as the genome sequence of a single individual. In other cases, the reference sequence is a consensus sequence formed by aligning multiple known sequences, such as the genome sequences of multiple individuals that serve as a reference population, or multiple sequencing reads of polynucleotides from the same individual. In some cases, sequence variants occur less frequently in a population (sometimes referred to as “rare” sequence variants). For example, in non-tissue samples, sequence variants may occur at frequencies of approximately 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or lower. In some non-tissue sample examples, sequence variants occur at frequencies of approximately 0.1% or less. In tissues, sequence variants may occur at frequencies of approximately 100%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or lower. Sequence variants can be any sequence that deviates from the reference sequence. Sequence variations can consist of changes, insertions, or deletions of a single nucleotide or multiple nucleotides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). If a sequence variant contains two or more nucleotide differences, the different nucleotides may be contiguous or discontinuous. Non-limiting examples of sequence variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (INDELs), copy number variants (CNVs), loss of heterozygosity (LOH), microsatellite instability (MSI), variable number tandem repeats (VNTRs), and retrotransposon-based insertion polymorphisms.Further examples of sequence variant types include those occurring within short tandem repeats (STRs) and simple sequence repeats (SSRs), or those resulting from amplified fragment length polymorphisms (AFLPs) or differences in detectable epigenetic marks (e.g., methylation differences). In some embodiments, sequence variants may refer to chromosomal rearrangements, including but not limited to translocations or fusion genes, or rearrangements of multiple genes resulting, for example, from chromothripsis.

[0051] The term "wild-type," when used in reference to an allele or sequence, refers to the allele or sequence that encodes the most common phenotype in a particular natural population. In some cases, the wild-type allele may refer to the allele that is most frequently present in the population. In other cases, the wild-type allele or sequence refers to the allele or sequence associated with a normal state versus an abnormal state, such as a disease state.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art. Any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of the present invention, and preferred methods and materials are described herein. [Modes for carrying out the invention]

[0053] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the Disclosure. Rather, the following description of preferred exemplary embodiments will provide a useful explanation for implementing various embodiments for those skilled in the art. It should be understood that various modifications can be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.

[0054] To provide a thorough understanding of the embodiments, specific details are provided below. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in the form of block diagrams to avoid obscuring the embodiments with unnecessary details. In other examples, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments.

[0055] Furthermore, it should be noted that individual embodiments may be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts or diagrams describe operations as a continuous process, many operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operation is complete, but it may have additional steps not included in the diagram. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, its termination may correspond to the return of the function to the calling function or the main function.

[0056] I. Introduction Cancer is a complex group of diseases characterized by the uncontrolled proliferation and spread of abnormal cells. Advances in medicine have made cancer increasingly treatable, especially when detected early. Among the many approaches used to treat cancer are surgery, chemotherapy, radiation therapy, targeted therapy, immunotherapy, and hormone therapy. In addition to primary approaches to treating cancer (e.g., surgery), secondary treatment options are becoming more common when treating cancer patients in efforts to reduce the likelihood of cancer recurrence. An example of this practice is for patients with stage III colon cancer, where standard clinical guidelines recommend adjuvant chemotherapy (ACT) after surgery as standard care. As shown in Figure 1, recent studies have shown that approximately 50%–55% (below) of subjects treated with surgery + ACT were cured by surgery alone and therefore received ACT unnecessarily. On the other hand, about 30–35% of subjects who received ACT may experience cancer recurrence, resulting in only 15–20% of subjects with stage III colon cancer benefiting from ACT (below). Therefore, there is a need for prognostic biomarkers that can identify patients who are more likely to benefit from ACT and at higher risk of relapse, from among patients with relapses that do not improve with ACT. Doing so would allow clinicians to reduce patient overtreatment and explore potentially more beneficial alternative second-line therapies.

[0057] Recent observational and interventional studies in non-metastatic colorectal cancer have shown that the detection of postoperative cell-free circulating tumor DNA (ctDNA) in the blood indicates the presence of minimal residual disease (MRD) and a high prognosis for the development of cancer recurrence. Therefore, ctDNA analysis is a promising approach to guide treatment decisions in stage III colorectal cancer and other cancers with similar ACT treatment paradigms. Cell-free ctDNA is small, random fragments of DNA isolated from tumors and circulating in human blood. Regarding postoperative detection, ctDNA may originate from a small number of cancer cells that may remain in the subject after surgical intervention. Therefore, early detection of MRD is important for demonstrating the effectiveness of initial treatment, assessing the risk of recurrence, and adjusting the treatment plan accordingly.

[0058] Detecting ctDNA in patients with early-stage cancer or low tumor burden can be challenging due to the low abundance of ctDNA, often present at levels less than 0.10% of total cell-free DNA. Furthermore, when evaluating a single landmark time point after surgery, radiotherapy, or systemic therapy, sensitivity for detecting patients who eventually recur may be less than 50%, compared to surveillance trials where sensitivity often exceeds 80%. In summary, clinical data highlight the ongoing unmet need for technologies that enable detection of low levels of ctDNA for improved clinical sensitivity in identifying high-risk patients with early lesions who may benefit from additional interventions.

[0059] Prior art methods use fixed gene panels with specific genetic mutations or probes designed to detect specific mutations. Both approaches have limitations in their clinical performance and usefulness. For example, if ctDNA is isolated from tumors with random, unpredictable fragments, and these random ctDNA fragments are not complementary to the sequences detected by the gene panel / probe, the targeted gene panel and probe become useless. In addition, the fabrication of patient-specific panels for ctDNA detection is also costly, time-consuming, and impractical. Achieving high sensitivity without compromising specificity can be challenging with NGS approaches. Current next-generation sequencing (NGS)-based technologies for ctDNA detection rely on analyzing various cell-free DNA features to improve sensitivity. Furthermore, the specificity associated with NGS technologies can be susceptible to sequencing errors, background noise, and other artificial errors.

[0060] Several approaches have been attempted to overcome the challenges currently faced by NGS technology. One approach is "tumor deinformatics," in which only plasma-derived cfDNA specimens are evaluated for the presence and levels of ctDNA. An example of a tumor deinformatics method involves fixed panels for the analysis of sequence changes and methylation loci. However, the sensitivity of this method depends on the changes present across a given panel content, due to the lack of prior knowledge of which specific locations within the tumor are mutated. Another approach is a "tumor-informed" approach. However, this requires the fabrication of patient-specific panels to detect and quantify ctDNA. This introduces some operational and technical complexity to the assay workflow, primarily resulting in extended turnaround times of several weeks. Attempts have been made to eliminate the need for patient-specific panels by developing fixed panels whose content represents common regions that have changed across specific, pre-specified tumor types; however, these methods are also limited to detecting changes in the regions included in the panel, thus limiting their sensitivity.

[0061] Attempts have been made to extend fixed panels across the entire human genome; however, these require specialized ctDNA detection algorithms that are not fully utilized to maximize analytical sensitivity and specificity. Previously proposed methods for optimizing the ctDNA signal-to-background noise ratio either result in a reduction of the actual ctDNA signal or create workflow inefficiencies to maintain analytical performance. Specifically, these methods involve inefficient redundancy sequencing of independent cfDNA replications to improve specificity, detection of somatic changes sufficient to allow analysis of mutation signatures associated with pre-determined mutagenic processes, and establishment of thresholds for ctDNA detection based on observed ctDNA levels, which, together, do not maximize technical performance across the range of tumor-specific changes identified for each patient's tumor.

[0062] Other difficulties, problems, and challenges may be associated with the underlying cancer. For example, because colorectal cancer is a heterogeneous disease, the individual genetic makeup and tumor location of patients make it difficult to predict the prognosis of ACT. In addition, prognostic biomarkers for colorectal cancer may have small effect sizes, making it difficult to identify their significance and predict their impact on patient outcomes. Furthermore, the biology of many cancers, including colorectal cancer, is complex and not fully understood, which makes it more difficult to identify and validate prognostic biomarkers. Moreover, developing and validating prognostic biomarkers can be expensive and time-consuming, which may limit their availability and use in clinical settings.

[0063] To address and overcome the challenges described above and others, this disclosure describes an innovative method for detecting cancer using WGS analysis of matched tumor tissue, non-cancerous, and non-tissue samples as both test and reference samples in the development and implementation of a gene analysis assay for training and evaluating the performance of the assay. Highly reliable tumor-specific somatic variants are identified from patient-matched tumor and non-cancerous variant datasets and then used to compare them to a non-tissue (e.g., plasma) variant dataset via a tumor information-based approach. Non-tissue variants are then filtered and scored via a pre-trained machine learning model to determine whether circulating tumor DNA (ctDNA) is present (based on the variant score) and the relevance level within a given total cell-free DNA (cfDNA) to the distribution of variant scores observed from a reference cohort. Non-tissue variants and their corresponding variant scores may also be used for other downstream applications.

[0064] Because of the challenges mentioned above (sequencing errors, background noise, artifact errors, etc.), overcoming these fundamental limitations of the NGS approach is not straightforward, and therefore using WGS is not an obvious approach. As described, the disclosed method overcomes the challenges of background noise, artifact errors, and germline mutations by first comparing WGS tumor samples, both obtained from the same patient, with WGS non-cancerous samples. In doing so, a tumor-specific profile is obtained that is free from noise, artifacts, and germline mutations, leaving only somatic tumor-associated mutations. Furthermore, the patient's own tumor-specific mutations are compared to the patient's non-tissue (e.g., plasma) variant profile to generate a patient-specific list of candidate somatic variants. The addition of one or more machine learning models that leverage high-quality candidate somatic variants is also not obvious in the art or has not been previously described. At least one of the machine learning models further filters each candidate somatic variant to generate a variant score, which is then used to determine the presence or absence of ctDNA and estimate the ctDNA level. This particular method significantly improves the specificity, sensitivity, and reproducibility of detecting ctDNA and ultimately MRD, enabling even early detection of cancer and thus improving patient survival outcomes.

[0065] II. Computer Environment Figure 2 shows a computing environment 200 according to an aspect of the present disclosure. The computing environment 200 includes client devices 205, a data repository 210, a minimal residual disease (MRD) detector platform 215, and a sequencer 275, all connected to each other by a network 220. While Figure 2 shows a specific arrangement of the client devices 205, data repository 210, MRD detector platform 215, and network 220, the present disclosure intends any suitable arrangement of the client devices 205, data repository 210, MRD detector platform 215, sequencer 275, and network 220. For example, two or more client devices 205, data repository 210, MRD detector platform 215, and sequencer 275 may be directly connected to each other, bypassing the network 220. For another example, two or more client devices 205, data repository 210, MRD detector platform 215, and sequencer 275 may coexist physically or logically with each other, either entirely or partially. Furthermore, while Figure 2 shows a specific number of client devices 205, a data repository 210, an MRD detector platform 215, a sequencing determiner 275, and a network 220, this disclosure intends any suitable number of client devices 205, a data repository 210, an MRD detector platform 215, a sequencing determiner 275, and a network 220. For example, and not limited to, a computing environment 200 may include multiple client devices 205, a data repository 210, an MRD detector platform 215, a sequencing determiner 275, and a network 215.

[0066] This disclosure envisions any type of network 220 well known to those skilled in the art, which may support data communication using any of the various available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Switching), AppleTalk®, etc. As mere examples, network 220 may be a local area network (LAN), an Ethernet-based network, Token Ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating under any of the IEEE 1002.11 protocol suite, Bluetooth®, and / or any other wireless protocols), and / or any combination of these and / or other networks.

[0067] Link 225 may connect the client device 205, the data repository 210, and the MRD detector platform 215 to or to the network 220. This disclosure envisions any suitable link 225. In certain embodiments, one or more links 225 include one or more wired (e.g., digital subscriber line (DSL) or data over cable service interface specification (DOCSIS)), wireless (e.g., Wi-Fi or World Wide Interoperability for Microwave Access (WiMAX)), or optical (e.g., synchronous optical network (SONET) or synchronous digital tier (SDH)) links. In certain embodiments, one or more links 225 each include an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, part of the Internet, part of the PSTN, cellular technology-based network, satellite communication technology-based network, another link 225, or a combination of two or more such links 225. Link 225 does not necessarily have to be the same throughout the computing environment 200. One or more first links 225 may differ from one or more second links 225 in one or more respects.

[0068] The client device 205 is an electronic device comprising hardware, software, or embedded logical components, or a combination of two or more such components, which can interact with the data repository 210 and the MRD detector platform 215 with respect to appropriate product target discovery functions in accordance with the technology of this disclosure. The client device may include several types of computing systems, such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, game systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices may run various types and versions of software applications and operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®) including various mobile operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX®, or UNIX®-like operating systems, Linux® or Linux®-like operating systems such as Google Chrome® OS). Portable handheld devices may include mobile phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), and personal digital assistants (PDAs). Wearable devices may include Google Glass® head-mounted displays and other devices. Client device 205 may be capable of running various applications, such as various internet-related applications and communication applications (e.g., email applications, short message service (SMS) applications), and may use various communication protocols. This disclosure envisions any suitable client device 205 configured to generate and output product target discovery content to the user.For example, a user may use client device 205 to run one or more applications, which may then generate one or more discover or store requests that can be served in accordance with the teachings of this disclosure. Client device 205 may provide an interface 230 (e.g., a graphical user interface) that allows a user of client device 205 to interact with it. Client device 205 may also output information to the user through this interface 230. Figure 2 shows only one client device 205, but any number of client devices 205 may be supported.

[0069] The data repository 210 is a data storage entity (or sometimes an entity) in which data is specifically partitioned for analysis or reporting purposes. The data repository 210 can be used to store data and other information for use by the MRD detector platform 215 and client devices 205. For example, one or more of the data repositories 210(a) and 210(b) may be used to store data and information used as input to the MRD detector platform 215 for generating prognosis predictions about a patient. In some examples, the data and information relate to various sequencing and variant calling files for at least two or more samples obtained from the same patient, generated by performing WGS. The data may also include any other information used by the MRD detector platform 215 when the MRD assay is functioning. The data repository 210 can reside in various locations, including the server 235. For example, a data repository used by the server 235 may be local to the server 235, or it may be remote from the server 235 and communicate with the server 235 via a network-based or dedicated connection on the network 220. The data repositories 210(a) and 210(b) may be of different or the same type. In a particular example, a data repository may be a database, which is an organized collection of data stored and electronically accessed from one or more storage devices, such as one or more servers 235. One or more servers 235 may be configured to run a database application that provides database services to other computer programs or computing devices in the computing environment (e.g., client device 205 and MRD detector platform 215), as defined by the client-server model. One or more of these databases may be adapted to enable the storage, updating, and retrieval of data to and from the databases in response to SQL-like commands or similar programming languages ​​used to manage the databases and perform various operations on the data within them.

[0070] The MRD detector platform 215 comprises a set of tools 240 for analyzing and visualizing data (i.e., data stored in the data repository 210). The MRD detector platform 215 is used to identify high-risk patients with early lesions, such as patients with MRD, and to perform a process to predict whether patients will benefit from second-line therapeutic drugs. In the configuration shown in Figure 2, the set of tools 240 includes three processors: a candidate somatic variant generator 245, a ctDNA predictor 250, and a prognosis predictor 255. The candidate somatic variant generator 245 is involved by itself in loading, processing, and storing data accessed from the data repository 210 for use by the ctDNA predictor 250 and / or the prognosis predictor 255. The ctDNA predictor 250 uses processed data from the candidate somatic variant generator 245 (e.g., high-confidence candidate somatic variant calls) to generate variant scores for candidate somatic variants, classify patient samples (e.g., non-tissue samples such as plasma) as ctDNA+ or ctDNA-, and / or estimate the ctDNA levels in the patient samples. The prognosis predictor 255 uses the candidate somatic variant scores generated by the ctDNA predictor 250 to output predictions about whether the patient is at low or high risk of cancer recurrence and whether the patient will benefit from disease therapy. The MRD detector platform 215 can reside in various locations, including the server 235. For example, the MRD detector platform 215 used by the server 235 may be local to the server 235 or may be remote from the server 235 and communicate with the server 235 via a network-based or dedicated connection of the network 220. The MRD detector platform 215 may have different or the same configuration. One or more servers 235 may be configured to run discovery applications that provide discovery services to other computer programs or computing devices (e.g., client devices 205) in the computing environment, as defined by the client-server model.

[0071] In various examples, server 235 may be adapted to run one or more services or software applications that enable one or more embodiments described in this disclosure. In certain examples, server 235 may also provide other services or software applications, which may include non-virtual and virtual environments. In some examples, these services may be provided to users of client device 205 as web-based or cloud services, such as under a Software as a Service (SaaS) model. Users operating client device 205 may then interact with server 235 using one or more client applications and utilize the services provided by these components (e.g., database and rescue applications). In the configuration shown in Figure 2, server 235 may include one or more components 260, 265, and 270 that implement the functions performed by server 235. These components may include one or more processors, hardware components, or software components that can be run by a combination thereof. It should be understood that multiple different device configurations are possible and may differ from the computing environment 200. Therefore, the example shown in Figure 2 is an example of a computing environment (e.g., a distributed system for implementing an exemplary computing system) and is not intended to be limiting.

[0072] Server 235 may consist of one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mount servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. Server 235 may also include other computing architectures with virtualization, such as one or more virtual machines running a virtual operating system, or one or more flexible pools of logical memory devices that can be virtualized to maintain the server's virtual memory devices. In various examples, Server 235 may be adapted to run one or more services or software applications that provide the functions described in the foregoing disclosure.

[0073] The computing system in Server 235 may run one or more operating systems, including any of those considered above, and any commercially available server operating systems. Server 235 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Java® servers, and database servers. Examples of database servers include, but are not limited to, those commercially available from Oracle®, Microsoft®, Sybase®, IBM® (International Business Machines), and others.

[0074] In some implementations, the server 235 may include one or more applications for analyzing and integrating data feeds and / or data updates received from users of the client computing device 205. For example, the data feeds and / or data updates may include, but are not limited to, real-time events related to sensor data applications, biological system monitoring, in vivo feeds, in silico feeds, or real-time updates received from public research, user research, one or more third-party sources, and data streams (continuous, batch, or periodic). The server 235 may also include one or more applications for displaying the data feeds, data updates, and / or real-time events via one or more display devices on the client computing device 205.

[0075] The sequencer 275 is a sequencing device, which is any machine capable of sequencing one or more nucleic acid molecules to generate raw sequencing data (e.g., reads). Library-prepared nucleic acid samples may be pooled and loaded into lanes of a sequencing flow cell. The flow cell may be loaded into the sequencer 275 and imaged to generate sequence data. For example, a reagent interacting with the nucleic acid sample may fluoresce at a specific wavelength in response to an excitation beam, thereby returning a signal for imaging. For example, the fluorescent component may be generated by a fluorescently tagged nucleic acid that hybridizes to a fluorescently tagged nucleotide incorporated into an oligonucleotide using a complementary molecule of the component or a polymerase. As will be understood by those skilled in the art, the wavelength at which the dyes of the sample are excited and the wavelength at which they fluoresce will depend on the absorption and emission spectra of the particular dye. The sequencer 275 may optionally include or be operably coupled to its own dedicated sequencer computer having its own input / output mechanism, one or more processors, and memory. In addition or alternatively, the sequencing determiner 275 may be operably coupled to a server 235 or a client device 205 via a network 220. The client device 205 may access raw sequencing data files from a data repository 210 and execute commands to analyze or communicate the sequence data to the network 220.

[0076] III. Workflow for detecting ctDNA in non-tissue samples Figure 3 shows an exemplary sample preparation and computational workflow 300 for detecting cancer using WGS data to address the limitations of current technology. The computational portion of workflow 300 (e.g., corresponding to the candidate somatic cell variant generator 245 in relation to Figure 2) analyzes the WGS data to enable detection of ctDNA at low levels, thereby providing improved clinical sensitivity. Briefly, the sample preparation workflow includes access to / acquisition of samples (e.g., tumor, normal (non-cancerous), and non-tissue samples), DNA isolation, library preparation, and sequencing. Experimental procedures may be performed in a laboratory by a qualified researcher, while bioinformatics procedures may be performed on an electronic device of a client (e.g., a researcher, clinician, etc.) that includes hardware, software, or embedded logical components, or a combination of two or more such components. Client devices may include several types of computing systems such as portable handheld devices, general-purpose computers such as personal computers and laptops, workstation computers, wearable devices, game systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices may run various types and versions of software applications and operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android®, BlackBerry®, Palm OS®), including various mobile operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX®, or UNIX®-like operating systems, Linux® or Linux®-like operating systems such as Google Chrome® OS). Portable handheld devices may include mobile phones, smartphones (e.g., iPhone®), tablets (e.g., iPad®), and personal digital assistants (PDAs).Wearable devices may include Google Glass® head-mounted displays and other devices. Client devices may be capable of running various applications such as various internet-related apps and communication applications (e.g., email applications, short message service (SMS) applications) and may use various communication protocols. This disclosure envisions any suitable client device configured to perform the biometric information workflow shown in Figure 3.

[0077] Obtaining a specimen In the sample processing portion of workflow 300 (top), at least two samples are obtained from a single patient. The sample or biological sample may be a cell-containing fluid or tissue. The sample may include, but is not limited to, amniotic fluid, tissue biopsy, blood, blood cells, bone marrow, fine-needle biopsy sample, ascites, amniotic fluid, plasma, pleural fluid, saliva, semen, serum, tissue, or tissue homogenate, frozen or paraffin-sectioned tissue. Methods for obtaining the sample include, but are not limited to, biofilm, aspiration, tissue section, swab, blood collection or other fluid, surgical or needle biopsy, etc. At least two samples obtained from the same patient may be nucleic acid samples (e.g., DNA and / or RNA in both natural and synthetic forms).

[0078] Samples can be obtained from non-cancerous subjects or subjects with disease (e.g., solid tumors or malignant tumors). As shown in Figure 3, at least two samples may be tumor samples (e.g., cancer-positive samples), normal samples (generally any body tissue or fluid containing nucleic acids without cancer (e.g., lymphocytes, saliva, buccal cells, or other tissues and fluids)), and non-tissue samples. All samples (e.g., tumor, normal, and non-tissue) are collected from the same patient. Figure 3 specifically shows patient samples including plasma samples, but some non-limiting examples include additional cfDNA non-tissue samples such as sputum, saliva, cerebrospinal fluid, surgical drain fluid, urine, and cystic fluid. In some cases, only two samples (e.g., tumor samples and whole blood samples) may be collected, for example, because plasma samples can be isolated from blood samples that leave leukocytes as normal, cancer-free samples. In other examples, more than two samples are collected from the same patient, for example, three samples including a tumor sample, a normal sample (e.g., a cancer-free sample obtained from any tissue or fluid), and a whole blood sample for plasma isolation. The tumor sample and / or normal sample may be a tissue sample or a body fluid sample. In addition, one or more whole blood samples may be collected from the patient at a single point in time or at multiple points in time, such as over the course of treatment. For example, at least two whole blood samples may be collected at the first appointment. Then, one or more whole blood samples may be collected at second and subsequent points in time.

[0079] Tumor specimens may be obtained as pre-prepared formalin-fixed paraffin-embedded (FFPE) specimens (e.g., tissue). A portion of the FFPE tumor specimen may first be sectioned and stained in a histopathology laboratory (or any other laboratory suitable for tissue preparation and staining) before DNA isolation.305 The processes of tissue / cell fixation, embedding, sectioning, staining, and imaging are well known in the art and any suitable method may be used. Briefly, the specimen (e.g., tissue) may first be fixed with a fixative to preserve the specimen and slow its degradation. The fixed specimen may then be embedded in paraffin, for example, in preparation for tissue sectioning. The fixed and / or embedded specimen may be sectioned into slices of appropriate thickness, for example, using a cryostat. The sectioned specimen may be mounted on a slide and various staining methods may be performed to better visualize the relevant structures. Examples of staining methods that may be used include histopathological staining, histochemical staining, hematoxylin and eosin (H&E) staining, trichrome staining, periodate Schiff, silver staining, iron staining, and immunohistochemistry (IHC).

[0080] Following sectioning and staining of the tumor specimen 305, the stained images are evaluated / analyzed by a pathologist 310. The pathologist may evaluate the specimen by indicating the desired features (e.g., tissue degeneration, tissue damage, cancer-positive / negative, etc.) and manually annotate it. If the tumor specimen is deemed acceptable after the pathological evaluation 310, the tumor specimen may be sent for experimental processing such as DNA isolation.

[0081] As described above, normal and non-tissue samples (e.g., plasma) can be collected from a single or multiple samples collected from the same patient as the tumor sample.315 As an example of single sample collection, a whole blood sample can be collected from a patient using venipuncture by other routine methods known in the art. A non-tissue sample, for example, may be a plasma sample. Plasma is separated from the blood sample by adding an anticoagulant to the blood sample and separating the plasma from the blood cells by centrifugation of the blood sample at a sufficient rate. The plasma sample may contain nucleic acids (e.g., cell-free DNA, ctDNA) associated with the patient's MRD. The remaining fractions separated from the plasma include blood cells (e.g., leukocytes (monocytes, lymphocytes, neutrophils, eosinophils, basophils, and macrophages)), red blood cells (erythrocytes), platelets, and the buffy coat fraction (e.g., including leukocytes and platelets), all of which can be used as normal samples. As an example of when normal and non-tissue samples (e.g., plasma) are collected from different biological samples from the same patient, a normal sample may generally be any body tissue or fluid containing nucleic acids that are not considered cancerous. Non-tissue samples may be collected from any biological sample containing cell-free DNA and / or ctDNA, such as plasma, sputum, saliva, cerebrospinal fluid, surgical drain fluid, urine, and cystic fluid, to give some non-limiting examples.

[0082] In certain embodiments, the tumor sample may include, for example, cell-free nucleic acids (containing DNA or RNA) or nucleic acids isolated from tumor tissue samples such as biopsy or excised tissue. In certain embodiments, the normal sample may include nucleic acids isolated from any non-tumor tissue of the patient, for example, patient lymphocytes or cells obtained via an oral swab. Cell-free nucleic acids may be fragments of DNA or ribonucleic acid (RNA) present in the patient's bloodstream. For example, circulating cell-free nucleic acids are one or more fragments of DNA obtained from a non-tissue sample of the patient (e.g., plasma, saliva, urine, etc.).

[0083] Where used herein, “patient” and “subject” are used interchangeably and refer to a mammal such as a human or a non-human primate, and the mammalian subject may be of any age. In any of the methods described herein, the subject may be suspected of having a disease, diagnosed with a disease, or receiving treatment for a disease. For example, the subject may be suspected of having cancer, diagnosed with cancer, or receiving treatment for cancer. In one embodiment, the subject may be suspected of having colon cancer, diagnosed with colon cancer, or receiving treatment for colon cancer. The subject may also include a living human being receiving medical treatment for a disease or condition. This includes people who are not defined as having a disease but are being investigated for signs of the disease. In some embodiments, the patient has undergone surgery to remove a cancerous tumor (e.g., a colon cancer tumor) and may or may not have received postoperative ACT. In other embodiments, postoperative ctDNA, indicating the presence of MRD, a strong prognostic factor for cancer, may be detected.

[0084] When at least two samples are collected from the same patient, the samples are ready for DNA isolation. The DNA can be isolated from FFPE tumor tissue samples to produce purified tumor DNA 325, the DNA can be isolated from the Buffy fraction or leukocyte (WBC) layer of a blood sample to produce purified germline DNA 335, and the DNA can be isolated from the plasma layer 340 of a blood sample to produce cfDNA 345. Germline DNA 335 is a normal, non-cancerous sample. In some cases, normal (germline) DNA 335 and plasma cfDNA 345 may not be collected from the same sample (e.g., the same whole blood collection), but instead from two different samples collected from the same patient. For example, germline DNA 335 can generally be collected from any biological sample that is not considered cancerous, while cfDNA 345 can be collected from any biological sample that is considered to contain cfDNA and / or ctDNA, such as plasma, sputum, saliva, cerebrospinal fluid, surgical drain fluid, urine, and cystic fluid.

[0085] Various methods for isolating DNA from a sample (e.g., cells, tissues, non-tissues, etc.) are known in the art. One method for isolating DNA may involve using a reagent kit (e.g., tubes and DNA extraction reagents, etc.). The kit may include probes for hybrid capture, as well as tools for library preparation such as any useful reagents and protocols for fragmentation, adapter ligation, purification / isolation, etc. Using the kit or other techniques known in the art, a sample containing DNA can be obtained. Other methods for isolating / extracting DNA from a sample involve disrupting and lysing the starting material, followed by the removal of proteins and other contaminants, and finally the recovery of the DNA. Cell lysis procedures and reagents are known in the art and can generally be carried out by chemical (e.g., detergents, hypotonic solutions, enzymatic procedures, etc.), physical (e.g., French press, sonication, etc.), or electrolytic lysis methods. Protein removal can be achieved, for example, by digestion with proteinase K, followed by salting out, organic extraction, gradient separation, or binding of DNA to a solid support (either anion exchange or silica technology). DNA can be recovered by precipitation using ethanol or isopropanol. The choice of method depends on many factors, including, for example, the amount of sample, the required amount and molecular weight of DNA, the purity required for downstream use, and time and cost. Isolated / extracted sample DNA may be whole-genome DNA, circulating cell-free DNA, ctDNA, mitochondrial DNA, circular DNA, etc. As shown in Figure 3, tumor DNA 325 and germline DNA 335 are whole-genome DNA samples, while DNA isolated from non-tissue (e.g., plasma) is cfDNA 345. The amount of DNA isolated from a sample may depend on several factors, such as the type of sample (tissue vs. cell vs. low-concentration cfDNA), sample size, and sample quality. DNA isolation from tumor tissue samples can yield at least 200 ng of DNA. DNA isolated from normal samples can yield at least 50 ng of DNA, and DNA isolated from non-tissue samples (e.g., plasma) can yield at least 10 ng of cfDNA.In some cases, to ensure that sufficient (e.g., at least 10 ng) of cfDNA is isolated, more than one whole blood sample is collected from the patient, and the DNA isolated from more than one whole blood sample is pooled. For example, at least two 10 mL volumes of whole blood are collected from the patient, and enough plasma is obtained to isolate at least 10 ng of cfDNA.

[0086] Other examples of methods for isolating DNA from tumor, normal, and non-tissue samples include the QIAmp system from Qiagen (Venlo, Netherlands), the Triton / heat / phenol protocol (THP), blunt-end ligation-mediated whole-genome amplification (BL-WGA), or the NucleoSpin system from Macherey-Nagel, GmbH & Co. KG (Duren, Germany). See also Xue, 2009, Optimizing the yield and utility of circulating cell-free DNA from plasma and serum, Clin Chim Acta 404(2):100-104. Also see Li, 2006, Whole genome amplification of plasma-circulating DNA enables expanded screening for allelic imbalances in plasma, J Mol Diag 8(1):22-30. Both are incorporated by reference.

[0087] In some cases, if the amount of nucleic acid is determined to be insufficient for analysis, amplification may be used to increase the amount of nucleic acid. Amplification refers to the production of additional copies of a nucleic acid sequence and is generally performed using polymerase chain reaction (PCR) or other techniques known in the art (e.g., Dieffenbach and Dveksler, PCR Primer, a Laboratory Manual, 1995, Cold Spring Harbor Press, Plainview, NY). PCR refers to the method by KBMullis (U.S. Patent Nos. 4,683,195 and 4,683,202, incorporated herein by reference) for increasing the concentration of a nucleic acid sequence segment in a mixture of genomic DNA without cloning or purification.

[0088] Whole genome sequencing All three obtained nucleic acid samples (e.g., tumor, normal, and non-tissue) are sequenced using any suitable whole-genome sequencing (WGS) method. Nucleic acids may be amplified before sequencing. Sequencing data is obtained from WGS, and the sequencing data includes sequence reads.

[0089] 1. DNA fragmentation and library preparation After DNA isolation, the isolated DNA (tumor 325, germline 335, and cfDNA 345) undergoes library preparation 350. Whole-genomic DNA (tumor 325 and germline 335) is fragmented into multiple shorter double-stranded DNA target fragments, although cfDNA from non-tissue samples may not be fragmented. Generally, DNA fragmentation can be performed physically or enzymatically. For example, physical fragmentation can be performed by acoustic shear, sonication, microwave irradiation, or hydrodynamic shear. Acoustic shear and sonication are the main physical methods used to shear DNA. For example, the Covaris® instrument (Woburn, MA) is an acoustic device for disrupting DNA to 100 bp–5 kb. Covaris also manufactures tubes (gTubes) for processing 6–20 kb samples for Mate vs. libraries. Another example is the Bioruptor® (Denville, NJ), an ultrasonic device used to shear chromatin, DNA, and destroy tissue. It can shear small amounts of DNA into lengths of 150 bp to 1 kb. Another example is the Hydroshear® from Digilab (Marlborough, MA), which uses hydrodynamic forces to shear DNA. Nebulizers, such as those manufactured by Life Technologies (Grand Island, NY), can also be used to spray a liquid using compressed air and shear DNA into 100 bp to 3 kb fragments in seconds. Because spraying can result in sample loss, in some cases it may not be the preferred fragmentation method for limited sample volumes. Sonication and acoustic shearing may be better fragmentation methods for smaller sample volumes because the total amount of DNA from the sample can be more efficiently retained. Other known or developed physical fragmentation devices and methods can also be used.

[0090] Various enzymatic methods can also be used to fragment DNA. For example, DNA can be treated with DNase I, or a combination of nonspecific nucleases such as maltose-binding protein (MBP)-T7 Endo I and Vibrio vulnificus nuclease (Vvn). The combination of nonspecific nuclease and T7 Endo works synergistically to produce nonspecific nics and counternics, generating fragments that dissociate 8 nucleotides or less from the nic sites. In another example, DNA can be treated with NEBNext® dsDNA Fragmentase® (NEB, Ipswich, MA). NEBNext® dsDNA Fragmentase produces dsDNA disruption in a time-dependent manner, yielding DNA fragments of 50 to 1,000 bp depending on the reaction time. NEBNext dsDNA Fragmentase contains two enzymes: one randomly generates nicks on dsDNA, and the other recognizes the nick sites and cleaves the DNA strand opposite the nick, producing dsDNA disruption. The resulting DNA fragments contain short overhangs, 5' phosphate groups, and 3' hydroxyl groups.

[0091] In some cases, a whole-genome DNA sample is fragmented into specific size ranges of target fragments. For example, a whole-genome DNA sample may be fragmented into fragments within the following ranges: approximately 25–100 bp, approximately 25–150 bp, approximately 50–200 bp, approximately 25–200 bp, approximately 50–250 bp, approximately 25–250 bp, approximately 50–300 bp, approximately 25–300 bp, approximately 50–500 bp, approximately 25–500 bp, approximately 150–250 bp, approximately 100–500 bp, approximately 200–800 bp, approximately 500–1300 bp, approximately 750–2500 bp, approximately 1000–2800 bp, approximately 500–3000 bp, approximately 800–5000 bp, or any other size range within these ranges. For example, a whole-genome DNA sample may be fragmented into fragments of approximately 300–800 bp. In some cases, the fragments can be about 25 bp larger or smaller. After fragmentation, the DNA fragments may have blunt ends.

[0092] A DNA library is prepared using DNA fragments (or unfragmented cfDNA) generated by the above process. A DNA library is a set of polynucleotide molecules (e.g., a sample of nucleic acids) that are prepared, assembled, and / or modified for a specific process, including, but not limited to, immobilization on a solid phase (e.g., a solid support, flow cell, or beads), enrichment, amplification, cloning, detection, and / or nucleic acid sequencing. A DNA library can be prepared before or during a sequencing process. A DNA library (e.g., a sequencing library) can be prepared by preferred methods known in the art. A DNA library can be prepared by targeted or untargeted preparation processes.

[0093] A DNA library is modified to contain one or more polynucleotides of a known composition, non-limiting examples of which include identifiers (e.g., tags, indexing tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer-binding sites, annealing sites), preferred integration sites (e.g., transposons, viral integration sites), modified nucleotides, or combinations thereof. Polynucleotides of known sequences can be added to preferred locations, e.g., the 5' end, the 3' end, or within the nucleic acid sequence. The polynucleotides of known sequences can be the same or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a 5' known sequence may hybridize to a first plurality of oligonucleotides, while a 3' known sequence may hybridize to a second plurality of oligonucleotides. A DNA library may include chromosome-specific tags, capture sequences, labels, and / or adapters. A DNA library may contain one or more detectable labels. These one or more detectable labels may be incorporated into the DNA library at the 5' end, 3' end, and / or any nucleotide position within the nucleic acid in the library. The DNA library may also contain hybridized oligonucleotides, which are labeled probes that can be added before immobilization on a solid phase.

[0094] Ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego Calif.). These methods often utilize adapter designs that allow for the incorporation of index sequences (e.g., sample index sequences for identifying sample origins for nucleic acid sequences) in the initial ligation step, and can often be used to prepare samples for single-read sequencing, paired-end sequencing, and multiplex sequencing. For example, nucleic acids (e.g., fragmented or unfragmented nucleic acids) are obtained by repairing them using fill-in reactions, exonuclease reactions, or a combination thereof. The resulting blunt-end repaired nucleic acid can then be extended with a single nucleotide complementary to a single nucleotide overhang on the 3' end of the adapter / primer. Any nucleotide can be used as the extension / overhang nucleotide.

[0095] DNA library preparation involves ligating an adapter oligonucleotide to a sample DNA fragment or ctDNA. The adapter sequence is bound to the template nucleic acid molecule by an enzyme. The enzyme may be a ligase or a polymerase. The ligase may be any enzyme capable of ligating an oligonucleotide (RNA or DNA) to the template nucleic acid molecule. Preferred ligases include T4 DNA ligase and T4 RNA ligase, commercially available from New England Biolabs (Ipswich, MA). Methods for using the ligase are well known in the art. The polymerase may be any enzyme capable of adding nucleotides to the 3' and 5' ends of the template nucleic acid molecule.

[0096] Adapter oligonucleotides are often complementary to flow cell anchors and are sometimes used to immobilize nucleic acid libraries on solid supports, such as the inner surface of a flow cell. Adapter oligonucleotides may include identifiers, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-end sequencing primers, pair-end sequencing primers, multiplex sequencing primers, etc.), or combinations thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). Adapter oligonucleotides may include one or more of the following: primer-annealing polynucleotides (e.g., for annealing to flow cell-bound oligonucleotides and / or free amplification primers), index polynucleotides (e.g., sample index sequences for tracking nucleic acids from different samples, also called sample IDs), and barcode polynucleotides (e.g., single-molecule barcodes (SMBs) for tracking individual molecules of sample nucleic acids to be amplified before sequencing, also called molecular barcodes). The primer-annealing components of an adapter oligonucleotide include one or more universal sequences (e.g., sequences complementary to one or more universal amplification primers). The index polynucleotide (e.g., sample index, sample ID) is a component of the adapter oligonucleotide and / or a component of the universal amplification primer sequence.

[0097] Adapter oligonucleotides can be used in combination with amplification primers (e.g., universal amplification primers) to generate library constructs containing one or more of the following: universal sequences, molecular barcodes, sample ID sequences, spacer sequences, and sample nucleic acid sequences. When used in combination with universal amplification primers, adapter oligonucleotides are designed to generate library constructs containing one or more ordered combinations of the universal sequences, molecular barcodes, sample ID sequences, spacer sequences, and sample nucleic acid sequences. For example, a library construct may contain a first universal sequence, followed by a second universal sequence, followed by a first molecular barcode, followed by a spacer sequence, followed by a template sequence (e.g., sample nucleic acid sequence), followed by a spacer sequence, followed by a second molecular barcode, followed by a third universal sequence, followed by a sample ID, followed by a fourth universal sequence. In addition or alternatively, when used in combination with amplification primers (e.g., universal amplification primers), adapter oligonucleotides are designed to generate library constructs for differentiating each strand of a template molecule (e.g., a sample nucleic acid molecule). In some cases, the adapter oligonucleotide is a double-stranded adapter oligonucleotide.

[0098] A universal sequence is a specific nucleotide sequence incorporated into two or more nucleic acid molecules or two or more subsets of nucleic acid molecules, where the universal sequence is the same for all molecules or subsets of molecules into which it is incorporated. Universal sequences are often designed to hybridize and / or amplify into multiple different sequences using a single universal primer complementary to the universal sequence. Two or more (e.g., pairs) of universal sequences and / or universal primers may be used. Universal primers often contain universal sequences. In some examples, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.

[0099] Optionally, a DNA library or a portion thereof may be amplified (e.g., by a PCR-based method). For example, a sequencing method may include amplification of a DNA library. The DNA library may be amplified before or after immobilization on beads or a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification involves a process of amplifying or increasing the number of nucleic acid templates and / or complementary strands present (e.g., in a nucleic acid library) by producing one or more copies of the template and / or complementary strand. Amplification can be carried out by a preferred method. The DNA library may be amplified by thermal circulation, isothermal amplification, or rolling circle amplification. In a particular sequencing method, the DNA library is added to a flow cell and immobilized by hybridization to an anchor under preferred conditions. This type of nucleic acid amplification is often called solid-phase amplification. During solid-phase amplification, all or part of the amplification product is synthesized by extension starting from an immobilized primer. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support. In some cases, modified nucleic acids (e.g., nucleic acids modified by the addition of adapters) are amplified.

[0100] 2. Whole genome sequencing Library-prepared nucleic acids (e.g., tumor, normal, cfDNA) are sequenced using a machine capable of sequencing nucleic acids (e.g., a sequencer described in relation to Figure 2).360 Examples of sequencing include, but are not limited to, NovaSeq, HiSeq, Genome Analyzer IIx, MiSeq, HiScanSQ, 454 DNA sequencers, GS FLX+, GS Junior System, OLiD next-generation sequencing platform, Ion PGM System, Ion Proton System, Ion S5, Ion S5xl, CEQ8000, RS system, Sequel system, nanopore sequencers, DNBSEQ-G50, DNBSEQ-G400, DNBSEQ-T7, Ultima Genomics UG100, etc. In certain cases, complete or substantially complete sequences are obtained, and sometimes partial sequences are obtained.

[0101] Any suitable nucleic acid sequencing method can be used, non-limiting examples of which include Maxim and Gilbert, linkage termination, synthesis sequencing, ligation sequencing, mass spectrometry sequencing, microscope-based techniques, or combinations thereof. In some embodiments, first-generation techniques such as Sanger sequencing methods, including automated Sanger sequencing methods, including microfluidic Sanger sequencing, can be used in the methods provided herein. In some embodiments, sequencing techniques including the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)) can be used. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally include cloned amplified DNA templates or single DNA molecules, which are sequenced in a large-scale parallel manner, sometimes within a flow cell. Next-generation (e.g., second and third-generation) sequencing techniques that can sequence DNA in a large-scale parallel manner can be used in the methods described herein and are collectively referred to herein as “large-scale parallel sequencing” (MPS). In certain embodiments, a non-targeting approach is used in which most or all nucleic acids in the sample are randomly sequenced, amplified, and / or captured.

[0102] Other suitable sequencing techniques include Pacific Biosciences' single-molecule real-time (SMRT) technique (in SMRT, each of the four DNA bases is bound to one of four different fluorescent dyes. These dyes are phosphorylated. A single DNA polymerase is immobilized with a single molecule of template single-stranded DNA at the bottom of a zero-mode waveguide (ZMW) where the fluorescent tag is excited to produce a fluorescent signal, and the fluorescent tag is cleaved. Detection of the corresponding fluorescence of the dye indicates which base has been incorporated), nanopore sequencing (as described in Soni & Meller, 2007, Progress toward ultrafast DNA sequence using solid-state nanopores, ClinChem 53(11):1996-2001, where DNA passes through nanopores and each base is determined by the change in current across the pore), chemosensitive field-effect transistor (chemPET) array sequencing (as described in, e.g., USPub. 2009 / 0026082), and (e.g., Modrianakis, E.N. and Beer M., in Base sequence) Electron microscopy sequencing can be used, as described in "determination in nucleic acids with the electron microscope, III. Chemistry and microscopy of guanine-labeled DNA, PNAS 53:564-71 (1965)."

[0103] In some embodiments, WGS is performed on a prepared DNA library sample. WGS is based on the amplification of DNA on a solid surface using folded PCR and anchor primers. Genomic DNA is fragmented, and adapters are attached to the 5′ and 3′ ends of the fragments. The DNA fragments bound to the surface of the flow cell channel are extended and bridged. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification followed by denaturation can create millions of clusters of approximately 1,000 copies of single-stranded DNA molecules of the same template within each channel of the flow cell. Sequential sequencing is performed using primers, DNA polymerase, and four fluorophore-labeled, reversibly terminated nucleotides. After nucleotide incorporation, a laser is used to excite the fluorophores, capture an image, and record the identity of the first base. The 3′ terminator and fluorophores are removed from each incorporation, and the incorporation, detection, and specific steps are repeated. The sequencing techniques of this method are described in U.S. Patents 7,960,120, 7,835,871, 7,232,656, 7,598,035, 6,911,345, 6,833,246, 6,828,100, 6,306,597, 6,210,891, U.S. Publication 2011 / 0009278, U.S. Publication 2007 / 0114362, U.S. Publication 2006 / 0292611, and U.S. Publication 2006 / 0024681, each of which is incorporated in whole by reference.

[0104] The WGS method described above allows for sequencing of samples at different depths. For example, WGS may be performed at 80 times the depth for tumor DNA sample 325, 40 times the depth for normal (e.g., germline) DNA sample 335, 30 times the depth for non-tissue cfDNA sample 345, and 20 times or more the depth for external control samples.

[0105] Sequencing methods (e.g., WGS) generate a large number of reads. As used herein, a “read” (e.g., “read”, “sequence read”) is a short nucleotide sequence produced by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment (“single-end read”), and sometimes from both ends of a nucleic acid fragment (e.g., paired-end read, double-end read). The length of a sequence read is often associated with a particular sequencing technique. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). Sequencing reads may have mean, median, average, or absolute lengths ranging from about 15 bp to about 1000 bp. For example, sequencing reads can be approximately 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, 20 bp, 25 bp, 50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or approximately 1000 bp, or any integer value between approximately 15 bp and 1000 bp. Sequencing reads and their associated quality scores are stored in a file known as a FASTQ file or FASTA file. Typically, a FASTQ file can contain approximately 1 million to 5 million reads per sample, however, more or fewer reads may be generated depending on the sample. Non-limiting examples include: (i) a FASTQ file of a tumor sample may contain approximately 2 billion to 4 billion reads per sample; (ii) a FASTQ file of a normal (non-cancerous) sample may contain approximately 800 million to 1.5 billion reads per sample; and (iii) a FASTQ file of a non-tissue (e.g., plasma) sample may contain approximately 800 million to 2 billion reads per sample.

[0106] 3. Bioinformatics Workflow In some embodiments, sequence reads are generated, acquired, collected, assembled, manipulated, transformed, processed, and / or provided by a sequence subsystem. A machine including a sequence subsystem can be a suitable machine and / or apparatus for sequencing nucleic acids using sequencing techniques known in the art. In some embodiments, the sequence subsystem can perform alignment, assembly, fragmentation, complementary strand, reverse complementary strand, and / or error checking (e.g., error-correct sequence read). Sequence reads are processed using a sequence processing subsystem to obtain sequence read data. Processing of sequence reads includes read alignment, mapping, and filtering. To carry out all these processing steps, a bioinformatics workflow includes steps including demultiplexing 365, reference genome alignment 370, variant calling 375 to identify whole-genome cfDNA variants 380 and whole-genome somatic variants 385, a ctDNA algorithm 390, and ctDNA percentage values ​​395.

[0107] As described above, the sequencing output is a FASTQ file containing all reads for a single sample. Part of the process of generating the FASTQ file is demultiplexing (e.g., sorting) all different library samples pooled together in a single flow cell lane into their own FASTQ files. In a typical WGS sequencing run, multiple library samples (e.g., 4, 12, 16, etc.) are combined and loaded into a single lane of the sequencing flow cell. This is because, during library preparation, each DNA fragment in the sample had a corresponding unique barcode ligated to the fragment. Thus, when multiple libraries are pooled for sequencing, the barcodes allow the samples to be distinguished from one another. The barcodes are also used to sort each sample into its own sequencing FASTQ file (i.e., demultiplexing).

[0108] The alignment of reads to a reference genome (e.g., the human reference genome) 370 involves mapping any number of reads to a specific nucleic acid region (e.g., a chromosome or a portion thereof) and is referred to as counting. As used herein, the term “reference genome” can refer to any known sequenced or characterized genome, partial or complete, of any organism or virus that can be used to reference a specific sequence from a subject. For example, reference genomes used for human subjects and many other organisms can be found at the National Center for Biotechnology Information at the World Wide Web URL ncbi.nlm.nih.gov.

[0109] Any suitable mapping / alignment method (e.g., process, algorithm, program, software, subsystem, or combination thereof) can be used. Non-exclusive examples of computer algorithms that can be used to align sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE1, BOWTIE2, ELAND, MAQ, PROBEMATCH, SOAP, BWA, or SEQMAP, or variations or combinations thereof. The terms “aligned,” “alignment,” or “to align” generally refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identity) or a partial match. Alignment can be performed manually or by computer (e.g., software, program, subsystem, or algorithm), non-exclusive examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Sequence read alignment can be a 100% sequence match. In some cases, the alignment is a sequence match of less than 100% (i.e., an incomplete match, a partial match, or a partial alignment). In some embodiments, the alignment is a match of approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75%. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes one, two, three, four, or five mismatches. Two or more sequences can be aligned using either strand (e.g., a sense strand or an antisense strand). In certain embodiments, one nucleic acid sequence is aligned with the reverse complementary strand of another nucleic acid sequence. The alignment results are stored in an alignment file (e.g., BAM).

[0110] As a quality control step, all alignment files are filtered, and non-primary alignment records, reads mapped to inappropriate pairs, and reads with six or more edits may be removed. Individual bases are excluded if their Phred base quality is less than 30 in tumor samples and less than 20 in normal samples. Where used herein, the term “less than” includes all integers and rational numbers. For example, less than 30 includes 29.9, 29.8, 29.7, 29.6, 29.5, 29.4, 29.3, 29.2, 29.1, 29.0, 25, 20, 15, 10, 5, and 0.

[0111] During a reference genome alignment 370, variations between the sample and the reference genome may be identified. The process of comparing sequence data with the reference is called variant calling 375. As described herein, a variant includes naturally occurring changes to DNA sequences not found in the reference sequence, and changes can be classified as benign, likely benign, variant of unknown significance, likely pathogenic, or pathogenic variant. Furthermore, variants can include both germline variants (e.g., variants present in all somatic cells) and somatic variants (variants that occur during an individual's lifetime, such as when an individual develops cancer). Examples of variants include small sequence variants (less than 50 base pairs) such as single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), and small structural variants (SVs) (e.g., deletions, insertions, insertions and deletions, sometimes referred to as indels), as well as larger (more than 50 base pairs) SVs such as chromosomal rearrangements (e.g., translocations and inversions) and copy number changes. SNVs / SNPs are the result of single-point mutations that can cause synonymous changes (a nucleotide change that does not alter the encoded amino acid), missense changes (a nucleotide change that alters the encoded amino acid), or nonsense changes (the resulting amino acid change converts the encoded codon to a stop codon). Furthermore, variants can occur in both coding and non-coding regions of the genome and can be detected by WGS, as opposed to targeted gene panels and target-specific probes.

[0112] Variant calling 375 uses one or more variant calling tools to examine alignment / mapped sequencing data and a reference genome side by side to determine the presence of sequence variations (single nucleotide changes and small indels). The variant calling tools may extract candidate variants from the alignment data, score several individual metrics for each variant, and apply these scores both individually and in combination to identify good sequence variations and exclude sequence artifacts. In some embodiments, at least one or more substitutions, small indels, and larger changes such as rearrangements, copy number variations, and microsatellite instability can be determined from the sequencing data. Structural changes may be detected using any preferred technology / variant calling tool, for example, MuTect, Strelka, and / or JointSNVMix2.

[0113] A list of detected variants and their characteristics (e.g., variant type) are annotated and deposited into a variant file (e.g., Variant Calling Format (VCF)). The output VCF file from Variant Calling 375 may be accessed by the ctDNA algorithm 390 or by a machine learning pipeline (described in Section III) to determine variant scores (e.g., importance scores). A VCF file for a single sample may contain approximately 1,500 to 800,000 variants, but depending on the sample, more or fewer variants may be found.

[0114] Identify and filter tumor-specific somatic changes to generate candidate changes. Tumor sequencing data and normal sequencing data can be compared to a reference human genome using any preferred method such as those described above to identify somatic alterations and their associated features (e.g., coverage, variant allele fractions, quality score, reliability score). A preferred reference human genome may include a published human genome (e.g., hg18 or hg36), sequence data from sequencing of a relevant sample (e.g., a patient's non-tumor DNA), or several other reference materials, such as a “gold standard” sequence obtained by Sanger sequencing of the nucleic acid in question. Variant calling analysis (e.g., patient sample for reference to the human genome) can identify various chromosomal alterations (e.g., rearrangements or amplifications), genomic signatures (e.g., microsatellite instability), and sequence variations (single nucleotide substitutions and small indels).

[0115] Variants identified as tumors and normal (germline) variants may be filtered using a set of criteria. Filtering criteria may include removing: (i) variants annotated as low reliability, (ii) variants annotated as indels, (iii) variants observed in genome databases (e.g., 1000 genomes or gnomAD germline database), (iv) variants overlapping simple tandem repeat tracks (e.g., UCSC simple tandem repeat track), (v) variants with locations having less than 10x coverage, (vi) variants with fewer than 4 alternative alleles in tumors or more than 1 in normals, (vii) variants with variant allele frequencies less than 0.05, or any combination thereof.

[0116] In addition, or alternatively, stricter filtering criteria may be applied to variants having cytosine substituted with thymine or guanine substituted with adenine, which may be associated with technical artifacts prior to analysis. Variants with these substitution patterns are removed if the variant allele frequency is less than 0.20 or the number of substitute alleles is less than 10. The finally filtered tumor variants and their characteristics, as well as normal (germline) variants and their characteristics, are stored in a VCF file.

[0117] To identify highly reliable, whole-genome tumor-specific somatic variants (380), any germline mutations that may be present in the tumor variant VCF file are removed. This is achieved by comparing the patient's tumor-identified variants with their non-tumor references (e.g., sequence data from normal / germline DNA of the same patient). Germline mutations, which are mutations present in all cells of the patient, are considered background noise or false-positive tumor mutations. When tumor sequencing data is compared only to the reference human genome, the resulting VCF file will contain both somatic and germline mutations. By filtering out germline mutations, the invocation of candidate somatic variants is significantly more likely to represent the patient's tumor somatic mutation profile. Such a profile cannot be achieved by performing WGS on tumor samples only, nor can a purified, reliable whole-genome tumor-specific profile be obtained from gene panels or targeted probes.

[0118] Candidate somatic variant calls are compared to a set of reference non-cancerous plasma donors. If a candidate somatic variant is present in at least 10% of the non-cancerous donors, or if any one of the non-cancerous donors contains the variant at at least 25% of the variant allele frequencies, that variant is filtered out. The number of candidate somatic variant calls can range from approximately 1,500 to 800,000 variants. However, more or fewer candidate somatic variant calls may be identified based on the sample. This step can be performed separately from or externally to the machine learning model. Alternatively, this step can be incorporated into the machine learning model so that thresholds can be fine-tuned by training the machine learning model.

[0119] Identification and filtering of ctDNA changes to determine candidate changes Variant calls 375 also generate whole-genome cfDNA variants 385 from the patient's non-tissue sequencing data file. First, whole-genome cfDNA variants 385 can be identified by comparing the non-tissue cfDNA sequencing data file to a reference human genome. The unfiltered whole-genome cfDNA variants 385 can be compared to a filtered list of candidate somatic variant calls, and only candidate somatic changes found in both the cfDNA variant list and the candidate somatic list can be selected to generate a final list of candidate somatic variant calls specific to the patient's MRD tumor profile. The final list of candidate somatic variants may contain approximately 40,000 to 70,000 variants. However, more or fewer candidate somatic variant calls may be identified based on the sample.

[0120] The final list of candidate somatic variant calls can be input into a ctDNA algorithm 390 (e.g., a ctDNA predictor 250, as described with respect to Figure 2) to predict whether a patient's non-tissue sample is ctDNA+ or ctDNA-. The ctDNA status can also be used to estimate the level of ctDNA present 393. In some embodiments, the ctDNA algorithm 390 includes a pre-trained machine learning model (MLM) that filters the final list of candidate somatic variant calls and generates a variant score for each of the candidate somatic variant calls. The variant score for each candidate somatic variant is between 0 and 1 (inclusive) and is determined using a set of features corresponding to each of the candidate somatic variant calls. Features refer to any kind of quality feature that can be output from sequencing, alignment, variant calls, or any combination thereof. For example, features may include a quality score for any given base in the sequence data, alignment quality, read quality, strand information, and metrics from a FASTQ file such as a metric for the complexity of a region in the genome (e.g., repeating regions and other regions prone to NGS sequencing errors). With respect to variant invocation, features may include a confidence or probability score output by the variant invocationr when the variant is identified, and / or the base quality of the variant. Variants that have a score above a given threshold (e.g., 0.25), more than 0 different duplicate variants, 30 or more tumor variant read mapping quality scores, a variant mismatch mean of 5 or less, and any numerical value for mean fragment size. A pre-trained MLM may also generate variant scores for a reference cohort of variants identified from non-cancerous samples using the same methods as described above.

[0121] When a variant score is assigned to each candidate somatic cell variant call, all variant scores for non-tissue samples are summed and divided by the total number of candidate somatic cell variants to obtain a normalized variant score. The normalized variant score can be used as a primary measure for cancer detection (e.g., whether a non-tissue sample is ctDNA+ or ctDNA-). A non-tissue sample is considered ctDNA+ if the normalized variant score is greater than or equal to one standard deviation of the maximum normalized variant score plus the reference cohort variant.

[0122] The ctDNA level in non-tissue samples is determined by dividing the total number of different duplicate variant reads for which the variant has a score greater than 0.25 by the sum of the products of (1) different duplicate reads per observed variant and (2) the median genome-wide coverage of different duplicate reads and the total number of unobserved candidate somatic variants, to obtain the estimated ctDNA fraction (as a percentage). In other words, the estimated ctDNA level represents the proportion of total cfDNA collected from the patient.

[0123] As a quality control check, the ctDNA algorithm 390 can also perform an SNP quality control check 396 to confirm that datasets obtained from tumor, normal, and non-tissue samples originate from the same patient based on detected SNPs and their associated allele fractions. This step ensures that no sample swaps occurred at any point in the preparation or analysis of the sample sets. An SNP quality control (QC) report 399 may be generated, and an exemplary summary of quality control metrics for the SNP check that may be displayed in the SNP QC report 399 is provided in Table 1. [Table 1]

[0124] Table 1 shows quality control metrics for SNP checks with a limit of blank (LoB) study, a limit of detection (LoD) study, accuracy / clinical confirmation study, and external controls. The objective of LoB studies is to determine the highest apparent concentration of ctDNA expected to be found when replication is tested in samples that do not contain ctDNA (e.g., normal, non-cancerous tissue, buffy coat blood fraction). The objective of LoD studies is to determine the lowest concentration of ctDNA that is likely to be reliably distinguishable from LoB studies. In other words, LoD determines the lowest viable concentration at which ctDNA can be detected in artificial tumor samples (e.g., synthetically generated) at various concentrations. As shown, 100% of replications passed the SNP check with a threshold of 0.8, and SNPs could be accurately identified with a median of 0.98 MutPct (e.g., variant allele frequency) for both tumor and plasma. The objective of accuracy / clinical confirmation studies is to determine the accuracy of the analysis of detected ctDNA (e.g., the closeness of agreement between true results and test results) by evaluating the agreement of sequencing and variant calling with orthogonal studies. As shown, 100% of the replicas passed the SNP check with a threshold of 0.8, with a median of 0.97 MutPct for both tumor and plasma. The objective of DNA input guard banding studies is to determine the range in which the amount of DNA input can deviate from the recommended input amount and still produce accurate results. In some cases, the range may be ±20% of the recommended input amount. With a threshold of 0.8, 0% of the DNA input studies passed the SNP QC check, indicating that these samples did not originate from the same patient.

[0125] Raw sequencing files (e.g., FASTQ files), processed sequencing files (e.g., alignment / mapping files), and variant call files generated from the sample processing and computation workflow 300 may be stored in storage devices such as servers, databases, or data repositories, as shown in Figure 2. Files may be stored on local, remote, and / or cloud servers. Each file may be stored associated with a subject identifier and date (e.g., the date the sample was collected and / or the date the file was generated). During analysis, one or more files may be further sent to another system (e.g., a machine learning pipeline or deployment system, as described in more detail herein).

[0126] IV. Training and Implementation of Machine Learning Models Figure 4 shows a block diagram of an exemplary machine learning pipeline 400, including several subsystems that work together to train, validate, and implement one or more machine learning models in various embodiments. The machine learning pipeline 400 may be run as part of the ctDNA predictor 250 or prognostic predictor 255 of the MRD detector platform 215 described in Figure 2. The machine learning pipeline 400 comprises a data subsystem 405 for collecting, generating, preprocessing, and labeling training and validation datasets 410, and for collecting, generating, setting, or implementing model hyperparameters 440; a training and validation subsystem 415 for facilitating the training and validation of one or more machine learning algorithms 420, and the generation of one or more machine learning models 430; and an inference subsystem 425 for deploying and implementing one or more trained machine learning models 430 independently or in combination with one or more downstream applications 435 for further processes (e.g., providing a diagnosis or managing a treatment).

[0127] As used herein, a machine learning algorithm (also described herein as simply an algorithm or multiple algorithms) is a procedure performed on a dataset (e.g., training and validation datasets) to extract features from the dataset, perform pattern recognition on the dataset, learn from the dataset, and / or fit to the dataset. Examples of machine learning algorithms include linear and logistic regression, decision trees, random forests, support vector machines, principal component analysis, Apriori algorithms, gradient descent algorithms, hidden Markov models, artificial neural networks, k-means clustering, and k-nearest neighbors. As used herein, a machine learning model (also described herein as a simple model or multiple models) is the output of a machine learning algorithm and consists of model parameters and a predictive algorithm. In other words, a machine learning model is a program saved after running a machine learning algorithm on training data and represents the rules, numbers, and other algorithm-specific data structures necessary for inference. For example, a linear regression algorithm may yield a model consisting of a vector of coefficients having specific values; a decision tree algorithm may yield a model consisting of a tree of if-then statements having specific values; a random forest algorithm may yield a random forest model, which is a collection of decision trees for classification or regression; or neural networks, backpropagation, and gradient descent algorithms together may yield a model consisting of a graph structure having a vector or matrix of weights having specific values.

[0128] A. Data subsystem 405 The data subsystem 405 is used to collect, generate, preprocess, and label data used by the training and validation subsystem 415 to train and validate one or more machine learning algorithms 420. The data subsystem 405 comprises training and validation datasets 410 and model hyperparameters 440. Raw data can be obtained through public or commercial databases. For example, the data subsystem 405 may access and load paired sequencing data and variant data from a data repository, such as the data repository 210 shown in Figure 2. Paired sequencing data and variant data can be generated by performing WGS and analysis on biological samples obtained from the same patient. The paired sequencing data and variant data accessed by the data subsystem 405 may include sets of sequences or sequence reads containing mutations and / or structural changes. The data subsystem 405 may also access WGS and variant files for sets of longitudinal samples collected from the same patient across treatment plans. The acquired raw data can be further preprocessed to generate the training and validation datasets 410.

[0129] Preprocessing can be implemented by a data subsystem 405 that acts as a bridge between raw data acquisition and effective model training. The primary purpose of preprocessing is to transform raw data into a format suitable and efficient for analysis, ensuring that the data fed into machine learning algorithms is clean, consistent, and relevant. This step can be useful because raw data often contains various problems such as missing values, noise, irrelevant information, and inconsistencies that can significantly hinder model performance. By standardizing and cleaning the data beforehand, preprocessing helps improve the accuracy and efficiency of subsequent analysis by making the data more representative of the underlying problem the model is trying to solve.

[0130] Raw data preprocessing may include data synthesis and / or data augmentation. Different data synthesis and / or data augmentation techniques may be implemented by the data subsystem 405 to generate preprocessed data used in the training and validation subsystem 415. Data synthesis involves creating entirely new data points from scratch. This technique may be used when real-world data is insufficient, too sensitive to use, or the cost and logistical barriers to obtaining more real-world data are too high. The synthesized data must be realistic enough to effectively train machine learning models, but clear enough to comply with regulations where necessary (e.g., privacy regulations (such as the US Health Insurance Portability and Accountability Act) and ethical guidelines). New data examples may be generated using techniques such as generative adversarial networks (GANs) or variational autoencoders (VAEs). These models attempt to learn the distribution of real-world data and produce new data examples that are statistically similar but not identical. Data augmentation, on the other hand, refers to techniques used to artificially expand the size of a dataset by creating modified versions of existing data examples. The primary goal of data augmentation is to increase the variation in the data to make the model more robust to changes that may be encountered in the real world, thereby improving its ability to generalize from training data to unseen data.

[0131] Other raw data preprocessing techniques include data cleaning, normalization, feature extraction, and dimensionality reduction. Data cleaning may include removing duplicates, filling in missing values, or filtering outliers to improve data quality. Normalization involves scaling numbers to a common scale without distorting differences in value ranges, which helps prevent model bias caused by the inherent scale of features. Feature extraction involves transforming input data into a set of usable features, reducing the dimensionality of the data in the process as much as possible. For example, raw sequencing data may include initial output generated by a sequencing machine from a sequencing assay. This initial output is typically in the form of raw sequence reads, which are short nucleotide sequences (e.g., DNA or RNA) representing fragments of the genome or transcriptome being sequenced. Feature extraction can transform raw sequencing data into a set of features including coverage, variant allele fractions, quality scores, and / or confidence scores. For example, WGS and analytical assays produce a variety of different sequencing, alignment, mapping, variant calling, and quality control files, each containing all types of features that describe the features or characteristics of the sequencing, alignment / mapping, variant calling, and quality control files. Extracted sequencing features may include metrics from the FASTQ file such as the quality score of any given base in the sequence data, alignment quality, read quality, and metrics regarding the complexity of regions within the genome (e.g., repeating regions and other regions prone to NGS sequencing errors). Variant calling features may also be extracted, including the confidence or probability score output by the variant caller when the variant is identified, and / or the base quality of the variant. The number of features depends on the project's needs; for example, about 10 to about 500 features may be extracted. In some examples, the extracted features include at least 62 predetermined features. It should be understood that more or fewer features may be considered.

[0132] By using dimensionality reduction techniques such as Principal Component Analysis (PCA) to obtain a set of principal variables, the number of variables under consideration can be reduced. These techniques not only reduce the computational load of the model but also help mitigate problems such as overfitting by simplifying the data without losing important information.

[0133] In examples where the machine learning pipeline 400 is used for supervised or semi-supervised learning of machine learning models, labeling techniques can be implemented as part of data preprocessing. Since labels serve as a definitive guide that the model uses to learn the relationship between input features and desired outputs, the quality and accuracy of data labeling directly impacts the model's performance. Accurate and consistent labeling is particularly important in complex areas such as cancer detection and medical diagnosis, as it provides ground truth or target outcomes to which the model's predictions are compared and adjusted during training. Effective labeling ensures that the model is trained with accurate and clear examples, and therefore improves its ability to generalize from training data to real-world scenarios. In some examples, ground truth values ​​are provided within the raw data.

[0134] In some cases, ground truth values ​​(labels) are provided within the raw data. For example, if the raw data includes sequencing data, the labels may include variant types. Many different variant types may be contained in variant files accessed and loaded by data subsystem 405. For example, variants may include benign, likely benign, variant of unknown significance, likely pathogenic, or pathogenic variants. Variants may include germline variants, somatic variants, or combinations thereof. Different structural variants may include small structural variants (less than 50 base pairs) such as single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), and small structural sequence variants (SVs) (e.g., deletions, insertions, insertions and deletions, sometimes referred to as indels), and larger SVs (e.g., more than 50 base pairs) such as chromosomal rearrangements (e.g., translocations and inversions). In some cases, variant types may be larger variations such as substitutions, small indels, and rearrangements, copy number variations, and microsatellite instability.

[0135] Labeling techniques can vary significantly depending on the type of data and the specific requirements of the project. Manual labeling, where human annotators label the data, is one method that can be used. This approach can be useful when detailed understanding and judgment are required, such as labeling medical data or classifying text data where context and nuance are important. However, manual labeling is time-consuming and can be inconsistent, especially with a large number of annotators. To mitigate this, semi-automated labeling tools can be used as part of the data subsystem 405 to pre-label data using algorithms, which human annotators can then evaluate and correct as needed. Another approach is active learning, a technique that iteratively labels new data using a model under development. This model suggests labels for new data points, and human annotators can evaluate and adjust certain predictions, such as the most uncertain ones. This technique optimizes labeling efforts by concentrating human resources on subsets of data, e.g., the most ambiguous cases, and improves efficiency and label quality through continuous refinement.

[0136] For example, if the raw data includes sequencing data, the label may include whether the variant is a true positive variant or a false positive variant. True positive variants / variants can be obtained from clinical FFPE tissue, cell lines, plasma samples from patients with cancer or patients with relapse after cancer treatment, or any combination thereof. False positive variants / variants can be obtained from non-cancerous normal FFPE tissue, cells, non-cancerous samples, or plasma samples from patients without relapse after cancer treatment, or any combination thereof. If a variant is partially labeled or remains unlabeled, the user may update the variant's label or create annotations indicating which parts of the input data should be labeled.

[0137] The training and validation datasets 410 may include raw data and / or preprocessed data. Typically, the training and validation datasets 410 are divided into at least three subsets of data: training, validation, and test. The training subset is used to fit the model, which is configured to perform inferences based on the training data. The validation subset, on the other hand, is used to tune hyperparameters and prevent overfitting to the training data. Finally, the test subset serves as a new, invisible dataset of the model, used to simulate real-world applications and evaluate the performance of the final model. The division process ensures that the model can perform well not only on the training data but also on the new, invisible data, thereby validating and testing the model's generalization ability.

[0138] Various techniques can be used to effectively partition data with the aim of maintaining a good representation of the overall dataset within each subset. Simple random partitioning (e.g., 70 / 20 / 10%, 80 / 10 / 10%, or 60 / 25 / 15%) is the simplest approach, where examples from the data are randomly assigned to each of the three sets. However, more sophisticated techniques may be needed to preserve the underlying distribution of the data. For example, stratified sampling can be used to ensure that each partition reflects the overall distribution of a particular variable, which is particularly useful when a particular category or outcome is underrepresented. Another technique, k-fold cross-validation, involves rotating the validation set across different subsets of the data, maximizing the use of data available for training while still retaining portions for validation. These techniques help achieve more robust and reliable model evaluations and are useful for developing predictive models that perform consistently across datasets.

[0139] The data subsystem 405 can also be used to collect, generate, set, or implement model hyperparameters 440 for the training and validation subsystem 415. Hyperparameters control the overall behavior of the model. Unlike model parameters 445, which are learned automatically during training, model hyperparameters 440 are external settings that must be determined before training begins. Model hyperparameters 440 can have a significant impact on the performance of the model. For example, in a neural network, model hyperparameters 440 may include, among other things, the learning speed, the number of layers, the number of neurons per layer, and / or activation features, while in a random forest, model hyperparameters 440 may include the number of decision trees in the forest, the maximum depth of each decision tree, the minimum number of samples that must be in each leaf node, the maximum number of features to consider when looking for the best split, and / or bootstrap parameters. These settings can determine how quickly the model learns, its ability to generalize to data not visible from the training data, and its overall complexity. It is important to set the hyperparameters correctly, as inappropriate values ​​can result in a model that underfits or overfits the data. Underfitting occurs when the model is too simple to learn the underlying patterns in the data, while overfitting occurs when the model is too complex and learns noise in the training data as if it were a signal. Many different variant types may be contained in the variant file accessed and loaded by the data generator 405. For example, variants may include benign, likely benign, variants of unknown significance, likely pathogenic, or pathogenic variants. Variants may include germline variants, somatic variants, or combinations thereof.Different structural variants may include small structural variants (less than 50 base pairs) such as single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), and small structural sequence variants (SVs) (e.g., deletions, insertions, insertions and deletions, sometimes referred to as indels), and larger (more than 50 base pairs) SVs such as chromosomal rearrangements (e.g., translocations and inversions). In some embodiments, variants may be larger changes such as substitutions, small indels, and rearrangements, copy number variations, and microsatellite instability.

[0140] B. Training and Verification Subsystem 415 The training and validation subsystem 415 consists of a combination of dedicated hardware and software to efficiently handle the computational demands required for training, validating, and testing machine learning algorithms / models. On the hardware side, high-performance GPUs (graphics processing units) can be used for their ability to perform parallel processing, significantly accelerating the training of complex models, particularly deep learning networks. CPUs (central processing units), while generally slower for this task, can also be used for training less complex models or when parallel processing is less critical. TPUs (tensor processing units), specifically designed for tensor computation, offer another level of optimization for machine learning tasks. In some examples, field-programmable gate arrays (FPGAs), or specially designed FPGAs, can be used to perform training, validation, and / or testing tasks.

[0141] Training is the initial stage of developing a machine learning model 430, where the model learns to make predictions, classifications, or decisions based on the training data provided from the training and validation datasets 410. During this stage, the model iteratively adjusts its internal model parameters 445 to achieve pre-defined optimization conditions. In the supervised machine learning training process, pre-defined optimization conditions can be achieved by minimizing the difference between the model output (e.g., prediction, classification, or decision) in the training data and the ground truth labels. In some cases, pre-defined optimization conditions can be achieved when a pre-defined fixed number of iterations or epochs (completely traversing the training dataset) is reached. In some cases, pre-defined optimization conditions are achieved when performance on the validation dataset stops improving or begins to deteriorate. In some cases, pre-defined optimization conditions are achieved when a convergence criterion is met, such as when the change in model parameters falls below a certain threshold between iterations. This process, known as fitting, is fundamental because it directly affects the accuracy and effectiveness of the model.

[0142] In an exemplary training phase performed by the training and validation subsystem 415, a training subset of the data is input to the machine learning algorithm 420 to find a set of model parameters 445 (e.g., weights, coefficients, trees, feature importance, and / or biases) that minimize or maximize an objective function (e.g., loss function, cost function, contrast loss function, cross-entropy loss function, out-of-bag (OOB) score, etc.). To train the machine learning algorithm 420 to achieve accurate predictions, it is necessary to minimize the “error” (e.g., the difference between the predicted label and the ground truth label). To minimize the error, the model parameters can be configured to update incrementally by minimizing the objective function over the training phases (“optimization”). Various different techniques can be used to perform optimization. For example, to train a machine learning algorithm such as a neural network, optimization can be performed using backpropagation. The current error is typically backpropagated to the previous layer and used to modify the weights and biases in such a way that the error is minimized. The weights are modified using an optimization function. Other techniques such as random feedback, direct feedback alignment (DFA), indirect feedback alignment (IFA), and Hebbian learning can also be used to update the model parameters 445 in a way that minimizes or maximizes the objective function. This cycle is repeated until the desired state (e.g., a predetermined minimum value of the objective function) is reached.

[0143] The training phase is driven by three main components: the model architecture (which defines the structure of the algorithm), the training data (which provides examples for learning), and the learning algorithm (which instructs the model on how to tune its model parameters). The goal is for the model to capture underlying patterns in the data without having to memorize specific examples, and to perform well on new, unseen data.

[0144] A model architecture is the specific arrangement and structure of the various components and / or layers that make up a model. In the context of a neural network, the model architecture may include the configuration of the layers in the neural network, such as the number of layers, the type of layers (e.g., convolutional, recursive, fully connected), the number of neurons in each layer, and the connections between these layers. In the context of a random forest consisting of a collection of decision trees, the model architecture may include the configuration of features used by the decision trees, the voting scheme, and hyperparameters such as the number of trees in the forest, the maximum depth of each tree, the minimum number of samples required to split a node, and the maximum number of features to consider when searching for the best split. In some examples, a model architecture is configured to perform multiple tasks. For example, the first component of the model architecture may be configured to perform a feature selection function, and the second component of the model architecture may be configured to perform a feature scoring function. Different components may correspond to different algorithms or models, and the model architecture can be a collection of multiple components.

[0145] Model architecture also encompasses the selection and arrangement of features and algorithms used in various models, such as decision trees or linear regression. This architecture determines how input data is processed and transformed through various computational steps to produce output. Model architecture directly impacts the model's ability to learn effectively and efficiently from data, and influences how the model adapts to the specific complexity and nuances of the data it is designed to handle and perform tasks such as classification, regression, or prediction well.

[0146] The model architecture can encompass a wide range of algorithms 420 suitable for different types of tasks and data types. Examples of algorithms 420 include, but are not limited to, linear regression, logistic regression, decision trees, support vector machines, naive Bayes algorithms, Bayesian classifiers, linear classifiers, K-nearest neighbors, K-means algorithms, random forests, dimensionality reduction algorithms, grid search algorithms, genetic algorithms, AdaBoosting algorithms, gradient boost machines, and artificial neural networks such as convolutional neural networks ("CNNs"), initial neural networks, U-Nets, V-Nets, residual neural networks ("ResNets"), transformative neural networks, recurrent neural networks, generative adversarial networks (GANs), or other variations of deep neural networks ("DNNs") (e.g., multi-label n-binary DNN classifiers or multi-class DNN classifiers). These algorithms can be implemented using various machine learning libraries and frameworks such as TensorFlow, PyTorch, Keras, and scikit-learn, which provide a wide range of tools and features to facilitate model building, training, validation, and testing. For example, the ctDNA algorithm 390 described in Figure 3 could be a random forest algorithm. In some examples, the ctDNA algorithm 390 could be a combination of different algorithms, such as a grid search algorithm and a random forest algorithm.

[0147] A learning algorithm is an overall method or procedure used to tune model parameters to fit data. It determines how the model learns from the data provided during training. This includes steps or rules that the algorithm follows to process the input data and tune the model's internal parameters (e.g., weights in a neural network) based on the output of the objective function. Examples of learning algorithms include gradient descent, backpropagation in neural networks, and splitting criteria for decision trees.

[0148] The training and validation subsystem 415 allows various techniques to be used to train the machine learning model 430 using learning algorithms, depending on the type of model and the specific task. For supervised learning models where the training data includes both inputs and expected outputs (e.g., ground truth labels), gradient descent is a possible method. This technique iteratively adjusts the model parameters 445 to minimize or maximize an objective function (e.g., loss function, cost function, contrast loss function, etc.). The objective function is a way of measuring how well the model's predictions match the actual labels or outcomes in the training data. It quantifies the error between the predicted value and the true value and presents this error as a single real number. The goal of training is to minimize this error, indicating that the model's predictions are, on average, close to the true data. Common examples of loss functions include the mean squared error for regression tasks and the cross-entropy loss for classification tasks.

[0149] The tuning of model parameters 445 is performed by an optimization function or algorithm, which refers to a specific method used to minimize (or maximize) the objective function. The optimization function is the engine behind the learning algorithm, guiding how the model parameters 445 are tuned during training. It determines the strategy to use when investigating the best weights to minimize (or maximize) the objective function. Gradient descent is a prime example of an optimization algorithm, including variants such as stochastic gradient descent (SGD), mini-batch gradient descent, and advanced versions like Adam or RMSprop, which offer various ways to adjust the learning rate or leverage the momentum of change. For example, when training a neural network, backpropagation may be used with gradient descent to update the network's weights based on the error rate obtained in the previous epoch (cycle through the complete training dataset). Another technique in supervised learning is the use of decision trees, where a decision tree-like model is built by splitting the training dataset into subsets based on attribute value tests. This process is repeated in each derived subset in a recursive manner called recursive splitting. When training a random forest, a set of decision trees can be trained collectively to minimize Gini impurities or entropy, leading to accurate classification.

[0150] In unsupervised learning, where training data does not contain labels, various techniques are used. Clustering is one method of grouping data into clusters that maximize similarity between data within the same cluster and differences between data in other clusters. For example, the K-means algorithm assigns each data point to the nearest cluster by minimizing the sum of the distances between the data point and the centroid of each of its clusters. Another technique, principal component analysis (PCA), involves reducing the dimensionality of the data by transforming the data into a new set of variables, principal components, which are uncorrelated and ordered, such that the first few retain most of the deformations present in all of the original variables. These techniques can help reveal hidden structures or patterns in the data, which may be essential for feature reduction, anomaly detection, or preparing the data for further supervised learning tasks.

[0151] Validation is another stage in developing a machine learning model 430, in which the model is checked for performance flaws and the hyperparameters 440 are optimized based on validation data provided from training and validation datasets 410. Validation data helps evaluate the model's performance, such as accuracy, precision, or recall, and measure how well the model is likely to perform in real-world scenarios. On the other hand, hyperparameter optimization involves adjusting settings that govern the model's learning process (e.g., learning speed, number of layers, size of layers in the neural network) to find the combination that yields the best performance on the validation data. One optimization technique is grid search, in which a predefined set of hyperparameter values ​​is systematically evaluated. The model is trained with each combination of these values, and the combination that produces the best performance on the validation set is selected. While thorough, grid search can be computationally expensive and impractical when the hyperparameter space is large. A more efficient alternative optimization technique is random search, which randomly uses and samples combinations of hyperparameters from a defined distribution. This approach can, in some examples, find good combinations of hyperparameter values ​​faster than grid search. Advanced methods such as Bayesian optimization, genetic algorithms, and gradient-based optimization can also be used to find optimal hyperparameters more effectively. These techniques model the hyperparameter space and use statistical methods to intelligently explore the space and find hyperparameters that result in improved model performance.

[0152] An exemplary validation process involves iterative operations where validation subsets of the data are input into the trained algorithm using validation techniques such as Kx cross-validation, leave-one-out cross-validation, leave-one-group-out cross-validation, and nested cross-validation, fine-tuning the hyperparameters until the optimal set of hyperparameters is finally found. In some examples, a 5x cross-validation technique may be used to avoid overfitting of the trained algorithm and / or restrict the number of features selected per split to the square root of the total number of input features. In some examples, the training dataset is split into five equally sized cohorts (or approximately equally sized), and every four of these cohorts is used to train the algorithm and generate five models (e.g., Model 1 is trained and generated using cohorts 1, 2, 3, and 4; Model 2 is trained and generated using cohorts 1, 2, 3, and 5; Model 3 is trained and generated using cohorts 1, 2, 4, and 5; Model 4 is trained and generated using cohorts 1, 3, 4, and 5; and Model 5 is trained and generated using cohorts 2, 3, 4, and 5). Each model is evaluated (or validated) using the unused cohorts in training (e.g., for Model 5, cohort 1 is used for validation). The overall performance of the training can be evaluated by the average performance of the five models. K-fold cross-validation utilizes the entire dataset for both training and evaluation, reducing the variance in performance estimates, thus providing a more robust estimate of model performance compared to a single training / validation split.

[0153] Once the machine learning model is trained and validated, it undergoes a final evaluation using test data provided from the training and validation dataset 410. This is a separate subset of the training and validation dataset 410, generally not used during the training or validation phases. This step is important because it provides an unbiased evaluation of the model's performance when simulating real-world operation. The test dataset acts as new, invisible data for the model, mimicking how the model will perform when deployed to actual use. During testing, the model's predictions are compared against the true values ​​of the test dataset using various performance metrics such as accuracy, precision, recall, and mean squared error, depending on the nature of the problem (classification or regression). This process helps validate the model's generalizability, i.e., its ability to perform well across different data samples and environments, highlighting potential issues such as overfitting or underfitting, and ensuring the model is robust and reliable for practical applications. The machine learning model 430 is fully validated and tested once its output predictions are deemed acceptable by user-defined acceptance parameters. Acceptance parameters can be determined using correlation techniques such as the Bland-Altman method and Spearman's rank correlation coefficient, and performance metrics such as error, accuracy, precision, recall, and receiver operating characteristic curve (ROC) can be calculated.

[0154] C. Inference subsystem 425 The inference subsystem 425 consists of various components for deploying the machine learning model 430 into a production environment. Deploying the machine learning model 430 involves moving the model from a development environment (e.g., the training and validation subsystem 415, where it is trained, validated, and tested) to a production environment where it can perform inference on real-world data (e.g., input data 450). This step typically begins with the post-trained model, including parameters and configurations such as the final architecture and hyperparameters.

[0155] Once deployed, the model is ready to receive input data 450 and return an output (e.g., inference 455). In some examples, the model exists as a component of a larger system or service (e.g., including an additional downstream application 435). In some examples, the model 430 and / or inference 455 can be used by the downstream application 435 to provide further information. For example, inference 455 can be used to determine whether a particular treatment should be administered to a patient. The downstream application can be configured to produce an output 460. In some examples, the output 460 includes a report containing the information generated by inference 455 and the downstream application 435.

[0156] In an exemplary inference subsystem 425, the input data 450 includes sequencing and variant files generated from one or more biological samples from a patient diagnosed with a disease (e.g., cancer). The input data 450 may further include clinical data about the same patient providing information about the type / stage of the disease, past, present, and / or future treatment plans, whether the patient has had disease relapses, and any other information related to the patient. In some examples, the input data 450 includes clinicopathological risk factors related to distinguishing patients, whether they have a very low or very high risk of developing cancer relapse within a particular time period (e.g., 3 years). The sequencing and variant files may be generated by performing WGS and variant calls on one or more biological samples collected from the patient by the sample processing and bioinformatics workflow 300, as described with respect to Figure 3. One or more biological samples may be a single non-tissue sample obtained from the patient (e.g., a plasma sample, or other samples such as sputum, saliva, cerebrospinal fluid, surgical drain, urine, or cystic fluid). One or more biological samples may also include tumor samples and non-cancerous samples (e.g., leukocytes or buffy coat, or tissue samples from a part known or determined to be non-cancerous). In some examples, one or more biological samples may further include a set of reference samples obtained for a non-cancerous subject or donor. Multiple rounds of samples may be collected and used as input data 450. In some examples, one or more samples may be collected at any point between pre-surgery and 3 years post-surgery. For example, one or more samples may be collected (i) pre-surgery, (ii) approximately 3 to 65 days after surgery and before therapeutic intervention, and / or (iii) approximately every 6 months for up to 3 years post-surgery and after therapeutic intervention. In some examples, tumor samples may be collected during surgery. In some examples, non-cancerous samples may be collected from non-tissue samples collected during surgery or at a different time than surgery.

[0157] In some examples, the input data 450 may be preprocessed before being input to the model 430 in order to achieve faster model performance. For example, the input data 450 may be preprocessed by the candidate somatic cell variant generator 245 processor of the MRD detector platform 215 as described with respect to Figure 2. The input data 450 may also be preprocessed by the sample preparation and bioinformatics workflow 300 as described with respect to Figure 3. Preprocessing may reduce the dimensionality of the input data 450 and thus save computing time and resources in the inference stage to generate the inference 455 (e.g., requiring less computer memory).

[0158] To manage and maintain their performance, deployed models can also be continuously monitored to ensure they perform as expected over time. This involves tracking the model's predictive accuracy, response time, and other operational metrics. In addition, models may require retraining or updating based on new data or changes in conditions. This can be useful because machine learning models can drift over time due to changes in the underlying data on which they make predictions, a phenomenon known as model drift. Therefore, maintaining machine learning models in a production environment often involves setting up performance monitoring mechanisms, periodic evaluation against new test data, and potentially periodic updating and retraining of the models to ensure they remain effective and accurate when making predictions.

[0159] Identifying patients who will benefit from V.ACT Figure 5A shows an exemplary workflow 500 for determining the status of a non-tissue sample as ctDNA positive or negative. The processes shown in Figure 5A, and also in Figure 5B, may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or combinations thereof (e.g., the computing environment 200 described in relation to Figure 2). The software may be stored in a non-temporary storage medium (e.g., on a memory device). The methods presented in Figure 5A and described below are illustrative and intended to be non-limiting. Figure 5A shows various processing steps occurring in a particular sequence or order, but this is not intended to be limiting.

[0160] In block 505, sequence reads from tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples are generated using whole-genome sequencing (WGS). Tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples may be obtained from the same patient at the same or different time points. For example, tumor samples, non-cancerous samples, and non-tissue samples may be collected at different time points during the patient's treatment; for example, samples may be collected (i) before surgery, (ii) during surgery, and (iii) before therapeutic treatment (e.g., adjuvant chemotherapy (ACT)), and approximately 3 to 65 days after surgery. The patient may have been previously diagnosed with cancer and undergone surgery to remove one or more tumors. The patient may preferably have been diagnosed with cancer, possibly colon cancer, but other cancer types (e.g., head and neck cancer, lung cancer, breast cancer, melanoma cancer, etc.) may be considered. It may be unclear whether the patient has a low or high risk of cancer recurrence after surgery, and therefore it may be unclear whether a second-line treatment option would be beneficial. In some embodiments, the preferred option for secondary treatment is adjuvant chemotherapy (ACT), however, other secondary treatment options may be considered. In some examples, the non-tissue nucleic acid sample comprises cell-free nucleic acids extracted from the plasma sample, and the plasma sample is isolated by adding an anticoagulant to the blood sample and centrifugating the blood sample at a rate sufficient to separate the plasma from the blood cells.

[0161] In some cases, non-tissue samples contain nucleic acids, such as cell-free DNA, released by cells undergoing apoptosis or necrosis. Furthermore, non-tissue samples may also contain ctDNA in very low abundances (e.g., often present at levels less than 0.10% of total cell-free DNA). Depending on when during treatment and / or how the patient responds to surgery, non-tissue samples may be ctDNA+ or ctDNA-. Optionally, non-cancerous samples may be obtained from non-plasma fractions, for example, as blood cells, leukocytes, buffy coat fraction, etc. In addition, or instead, non-cancerous samples may be collected during surgery using tumor samples. In this example, non-cancerous samples may be any body tissue or fluid containing nucleic acids that are not considered cancerous. Tumor samples may be collected during surgery as tissues, cells, plasma, blood, cell-free DNA, circulating tumor DNA, or any combination thereof.

[0162] Block 510 generates tumor variant calling files, non-cancerous variant calling files, and non-tissue variant calling files. Generation can be performed by analyzing the sequence reads corresponding to the tumor nucleic acid sample, non-cancerous nucleic acid sample, and non-tissue nucleic acid sample, respectively. The analysis can be performed by the sample handling and calculation workflow 300 described with respect to Figure 3. The variant calling file, or VCF file, contains a list of all detected variants, their characteristics (e.g., variant type), and quality features of a single sample. The first variant calling file is the result of comparing the experimental sample (e.g., tumor nucleic acid sample, non-cancerous nucleic acid sample, and plasma sample) with a reference or "gold standard" genome, and the differences / modifications identified between the experimental sample and the reference are re-encoded in the corresponding VCF file.

[0163] In block 515, the tumor variant call file is compared with the non-cancerous variant call file to generate a list of somatic variants. In some cases, variants in the non-cancerous variant call file are treated as "germline variants" that do not have a useful effect in determining true positive mutations in non-tissue samples. "Germline variants" are excluded or deleted from the tumor variant call file. The remaining variants in the tumor variant call file are somatic variants.

[0164] Block 520 generates a list of candidate somatic cell variants by comparing a list of somatic cell variants with a non-tissue variant calling file. In some examples, only variants appearing in both the list of somatic cell variants and the non-tissue variant calling file are considered as candidate somatic cell variants. In some examples, other criteria or variant files may be used to generate a list of candidate somatic cell variants. The list of candidate somatic cell variants may include substitutions, small indels, chromosomal rearrangements, copy number variations, microsatellite instability, or any combination thereof. In addition, the set of candidate somatic cell variants also retains information about their properties (e.g., variant type) and their quality features.

[0165] In some cases, SNP quality control checks may be performed to confirm that datasets obtained from tumor, normal, and non-tissue samples originate from the same patient based on the detected SNPs and their associated allele fractions. This step ensures that no sample swaps occurred at any point in the preparation or analysis of the sample sets.

[0166] In block 525, a score for each candidate somatic cell variant in the list of candidate somatic cell variants may be generated using a classification machine learning model. The score may be generated based on multiple classifications generated by the classification machine learning model. In some examples, the score includes a variant score for each candidate somatic cell variant. In various embodiments, the classification machine learning model is a random forest classification model containing a set of decision trees (see, for example, Figure 6). In some examples, the classification machine learning model is configured to generate variant scores by filtering candidate somatic cell variants through a series of yes / no questions and assigning a variant score (e.g., confidence / probability score) to each variant. For example, a first candidate somatic cell variant (e.g., input) is input to the classification model along with its corresponding features. When the input traverses tree 1 corresponding to the first feature, the yes / no question may be whether the value of the input feature meets the threshold of the first feature. If yes, the input variant may receive a score of 1. This process is repeated for each feature of the input variant until all trees in the set have generated scores. Assuming a random forest contains 62 trees and an input variant is scored as yes (e.g., 1) by 58 trees, the final variant score is 58 / 62 = 0.94. The final variant score can be compared to a predetermined threshold to determine whether the input variant is removed. In some examples, the individual trees that make up the forest set may have different relative weights depending on how predictive their particular features are for the overall classification.

[0167] In some examples, a classification machine learning model is a collection of multiple models configured to perform variant selection before inputting a candidate somatic cell variant random forest classification model. Variant selection may be performed based on a search model or a selection model. In some examples, the search or selection model is also pre-trained using the process described with respect to Figure 4. Candidate somatic cell variants are searched, and a subset of variants is selected using the search or selection model.

[0168] In some cases, variant scores generated from a classification model can be used to determine the status of ctDNA in non-tissue samples (e.g., present or absent) and to estimate the level of ctDNA in non-tissue samples. To determine the status of ctDNA in non-tissue samples, all variant scores for the non-tissue samples are summed and divided by the total number of candidate somatic variants to obtain a normalized variant score. The normalized variant score can be used as a primary measure for cancer detection (e.g., whether the non-tissue sample is ctDNA+ or ctDNA-). A non-tissue sample is considered ctDNA+ if the normalized variant score is greater than or equal to one standard deviation of the maximum normalized variant score plus the reference cohort variant.

[0169] In block 530, ctDNA status is determined for a patient's non-tissue nucleic acid sample based on a score. ctDNA status can be either positive or negative. The ctDNA status is determined by dividing the total number of different duplicate variant reads for which a variant has a score greater than 0.25 by the sum of the products of (1) the number of different duplicate reads per observed variant and (2) the median genome-wide coverage of different duplicate reads, and the total number of unobserved candidate somatic variants, to obtain the estimated ctDNA fraction (as a percentage). In other words, the estimated ctDNA fraction within the total cfDNA collected from the patient's non-tissue is compared to the ctDNA distribution observed from a reference cohort of healthy (e.g., non-cancerous) individuals to determine whether the status is positive or negative.

[0170] Block 535 generates a report to provide the patient's ctDNA status. In some examples, the report may include other information, such as the patient's configured genome using sequence reads, or some or all variants in tumor variant calling files, non-cancerous variant calling files, and / or non-tissue variant calling files.

[0171] Figure 5B shows an exemplary workflow for training a classification model, more specifically a random forest classification model.

[0172] In block 540, labeled training datasets are accessed. Labelled training datasets include thousands of ground truth true positive mutations and WGS of their associated features from clinical FFPE tissues, cell lines, plasma cases from patients with cancer, or any combination thereof, and their corresponding features. In addition, labeled training datasets may also include thousands of ground truth false positive mutations and WGS of their associated features from healthy (e.g., non-cancerous) normal FFPE tissues, cells, plasma cases from non-cancerous samples, or any combination thereof, and their corresponding features. True / false positive mutations per sample may include one or more examples of substitutions, small indels, rearrangements, copy number variations, microsatellite instability, or any combination thereof. Furthermore, sample data may include sequencing results and variant calls generated by diluting samples by different dilution levels to achieve various DNA concentrations and sequencing the diluted samples. For example, biological samples (e.g., tissue samples, non-cancerous samples, and / or non-tissue samples) may have dilutions of approximately 0.01, approximately 0.001, and approximately 5 × 10⁻⁶. -4 , about 2×10 -4 , about 1×10 -4 , about 5×10 -5 , or approximately 1 x 10 -5 It can be diluted to the dilution level.

[0173] 1. Training and Implementation of a Random Forest Machine Learning Model Figure 6 shows exemplary examples of the Random Forest Machine Learning Model 600 in various embodiments. For example, the Random Forest Machine Learning Model 600 may be a classification model implemented in a system, for example, as part of the ctDNA predictor 250 of the MRD detector platform 215 described with respect to Figure 2 and / or as the ctDNA algorithm 390 described with respect to Figure 3. The Random Forest Machine Learning Model 600 takes a dataset 605 (e.g., variants and their corresponding features) as input and applies different combinations of features 615 to a decision tree 610 to generate a score for each variant. The scores can then be used by the voting scheme 620 of the Random Forest Machine Learning Model 600 to determine the output 630. In the training phase, the dataset 605 contains training data. In the inference phase, the dataset 605 contains real-world data. For example, the real-world data may contain patient-specific variants generated before and after the filtering step, as shown in Figure 3.

[0174] The dataset 605 may include sequencing data corresponding to the variants. In some examples, the dataset 605 includes paired sequencing data and variant data. The paired sequencing data and variant data may have different sequencing coverage or depth and may be obtained by sequencing nucleic acids. For example, a tissue sample may have a sequencing depth (e.g., about 80-fold) because there is a large abundance of nucleic acids that can be isolated and / or sequenced from such a tissue sample. Non-cancerous (e.g., normal) samples such as tissues, cells, white blood cells, buffy coat cells, etc. may be sequenced at a different depth (e.g., about 40-fold), while other samples (e.g., non-tissue samples or plasma samples) may be sequenced at a sequencing depth of about 30-fold (e.g., about 30-fold) different from the above depth due to the limited abundance of nucleic acid material. The difference in sequencing depth can affect the overall quality of the sequencing results and variant calls. Regarding obtaining the dataset 605, the same set of sequencing depths may be used in the training phase and the inference phase. In some examples, different sets of sequencing depths are used in the training phase and the inference phase. Thus, in some examples, the paired sequencing and variant data of the dataset 605 accessed by the data generator 405 may also be generated by diluting the samples to various DNA concentrations with different dilution factor levels and sequencing the diluted samples. For example, a biological sample (e.g., a tissue sample, a non-cancerous sample, and / or a non-tissue sample) may have a DNA concentration of about 0 to about 1×10 -10. More specifically, the data sample may have a DNA concentration diluted at a dilution level of about 0.01, about 0.001, about 5×10 -4 , about 2×10 -4 , about 1×10 -4 , about 5×10 -5 , or about 1×10 -5 .

[0175] Each decision tree 610 is a decision aid tool that uses a binary tree graph to make decisions and / or predict their possible outcomes. When training a random forest, each decision tree is built independently based on a random subset of the training data and a random subset of features ("bootstrap"). When building each decision tree 610, instead of considering all features of the data points (e.g., variants) for each split, a random subset of features (N) is used. i Feature 615) is generally selected, which helps introduce randomness and diversity among the decision trees in the forest. In addition to the random selection of Ni feature 615, each decision tree is also trained on bootstrap samples of training data that may be randomly replaced samples of the same size as the original dataset. This means that some samples (e.g., variants) may be included multiple times, while others may be omitted when training a particular decision tree. Each decision tree can be thought of as embodying several yes / no questions to assess the probability that a variant is a true positive variant exhibiting a positive ctDNA status. Each tree generates its own variant score, independently of the other trees in the collective model. The random forest can later use a voting scheme 610 (e.g., majority or soft voting) to aggregate the decision trees 610 and determine the final classification, final score, or ctDNA status of the samples associated with the dataset 605. Training of the random forest machine learning model 600 and / or decision trees 610 can be carried out using the training and validation subsystem 415 described with respect to Figure 4. The number n of decision trees can be a hyperparameter provided before training. It should be understood that each decision tree 610 does not need to be a balanced tree with an equal number of nodes on the left and right branches, and can have depths different from the depth shown in Figure 4.

[0176] Each data point in dataset 605 (e.g., a variant with a corresponding feature) traverses each of the decision trees 610 that make up the random forest model. The random forest machine learning model 600 may contain at least several hundred decision trees (e.g., n ≥ 500 or n ≥ 1,000), and each decision tree contributes weakly to classification, but collectively, the random forest machine learning model 600 is a powerful classifier. For example, each decision tree 610 may consider about 10 to 500 features, and each decision tree may consider different subsets of features (e.g., N i Features 615) may be considered. In some examples, the total number of features that a decision tree considers for each variant may be 62. In some examples, the number is at least 62. It should be understood that more or fewer features may be considered.

[0177] In some examples, each decision tree 610 generates a score for each variant in the dataset 605, and the score is a value between [0, 1]. In some examples, the score is a binary score of either 0 or 1 (i.e., a classification score). In some examples, each decision tree 610 is configured to generate a score for all variants in the dataset 605, and the score is a value between [0, 1], or a binary score of either 0 or 1.

[0178] If a feature used for splitting is missing from dataset 605, various techniques can be used to correct the missing feature. For example, surrogate splitting can be used when a feature is missing for a data point in the training subset of the data, and the decision tree is configured to make a decision using another feature that correlates with the missing feature. A surrogate feature is typically the feature that best mimics the split that would have occurred if the missing feature were available. If a suitable alternative feature is not available, the decision tree may use the most common value of the missing feature in the training data, or it may use a default value. If a feature is missing during the inference phase, imputation can be used to replace the missing value with an alternative value. The alternative value may be configured during training to be the mean, median, or mode of the feature in the training dataset. Surrogate splitting can be used to select another feature that correlates with the missing feature and split. In some examples, a random forest machine learning model may be configured to have a default path to handle missing features.

[0179] The scores generated by decision trees 610 can be aggregated based on a voting scheme 620. In some examples, the voting scheme 620 includes majority voting. In a majority voting scheme, each decision tree in the random forest generates a classification score for a given variant, and the final classification is the class that receives the most “votes” from each individual tree (e.g., 0 or 1). In some examples, the voting scheme 620 includes soft voting, which calculates the average score from all decision trees and / or selects the class with the highest average probability as the final classification. The final classification can be provided as output 630. In some examples, the random forest machine learning model is configured to generate a final score for each subject based on all variants in dataset 605, where the score is the final classification normalized across all variants in dataset 605. The final score can also be provided as part of output 630. Some or all of the input data in dataset 605 can also be provided as part of output 630.

[0180] Referring back to Figure 5B, in box 545, the classification model (e.g., a random forest classification model) is trained on several trees. Training is an iterative process that begins at the first node of the first tree and involves initially inputting a portion of the labeled training data into the classification model. The portion of the labeled training data is replaced and randomly sampled to create a subset of the training data (e.g., also known as bootstrapping resampling). The subset may be, for example, about 66% of the total training dataset. At each node: (i) features of several variants are selected from the portion of the labeled training dataset, (ii) an objective function is used to determine which of the variant features provides the best binary split from the number of variant features. The best split is based on which feature variant features minimize the objective function, (iii) the first node is assigned the determined variant features, and (iv) the iterative process is repeated on the second and subsequent nodes of the first tree for several iterations or epochs until the first tree is generated. This process, steps (i) to (iv), is repeated for the first node of the second and subsequent trees until all variant features are assigned to the tree. In other words, the random forest tree is constructed from the parameters / features (e.g., variant features) of the data it is trained on. The types of features that can be used include FASTQ quality score, alignment score, read coverage, chain bias, etc. Furthermore, several optimal numbers of variant features are preferably discovered. Random forest execution time is fast and can handle unbalanced and missing data.

[0181] In the case of random forests, generally, the number of features in a variant is less than the number of predictor variables. When running a random forest, as a new input (e.g., a variant) enters the system, it traverses all the trees. The result can be either the mean or a weighted mean of all the terminal nodes reached. For many predictors, the set of eligible predictors differs from node to node. As the number of features in a variant decreases, both the correlation between trees and the strength of individual trees decrease.

[0182] In block 550, the trained classification model is the output that generates variant scores for variants in the labeled training dataset. The classification model can be subjected to various filtering and scoring techniques to ensure that only reliable variants are considered. Furthermore, the filtering and scoring techniques can act as pass-through criteria with minimum or ideal ranges to ensure that variations of high-quality candidates are considered. In other words, a classification model trained with various filters and thresholds will robustly remove any false-positive variants and low-quality variants.

[0183] After variant scores are generated by the classification model, they can be used to determine the status of ctDNA in non-tissue samples (e.g., present or absent) and to estimate the level of ctDNA in non-tissue samples. To determine the status of ctDNA, all variant scores for non-tissue samples are summed and divided by the total number of candidate somatic variants to obtain a normalized variant score. The normalized variant score can be used as a primary measure for cancer detection (e.g., whether the non-tissue sample is ctDNA+ or ctDNA-). A non-tissue sample is considered ctDNA+ if the normalized variant score is greater than or equal to one standard deviation of the maximum normalized variant score + reference cohort variant.

[0184] The ctDNA level in non-tissue samples is determined by dividing the total number of different duplicate variant reads for which the variant has a score greater than 0.25 by the sum of the products of (1) different duplicate reads per observed variant and (2) the median genome-wide coverage of different duplicate reads and the total number of unobserved candidate somatic variants, to obtain the estimated ctDNA fraction (as a percentage). In other words, the estimated ctDNA level represents the proportion of total cfDNA collected from the patient.

[0185] VI. Computing Environments: Machines, Software, and Interfaces Certain processes and methods described herein (e.g., mapping, counting, normalization, scaling, adjustment, classification, and / or determination of sequence reads, counting, levels and / or profiles, ctDNA detection and analysis) are performed within a computing environment including computers, microprocessors, software, modules, other machines such as sequencers, or combinations thereof. The methods described herein are typically computer implementations, and one or more parts or steps of the methods are performed by one or more processors (e.g., microprocessors), computers, systems, devices, or machines (e.g., microprocessor-controlled machines). Suitable computers, systems, devices, machines, and computer program products often include or are used in conjunction with computer-readable storage media. Non-exclusive examples of computer-readable storage media include memory, hard disks, CD-ROMs, and flash memory devices. Computer-readable storage media are generally computer hardware and are often non-temporary computer-readable storage media. Computer-readable storage media are not computer-readable transmission media, the latter being transmission signals themselves.

[0186] Figure 7 shows a non-limiting example of a computing environment 710 in which various systems, methods, processes, and data structures described herein may be implemented. Computing environment 710 is only one example of a suitable computing environment and is not intended to imply any limitations on the scope or functionality of the systems, methods, and data structures described herein. Computing environment 710 should not be construed as having any dependencies or requirements on any one or combination of components exemplified in computing environment 710. A subset of the systems, methods, and data structures shown in Figure 7 may be available in certain embodiments. The systems, methods, and data structures described herein operate in a number of other general-purpose or special-purpose computing system environments or configurations. Examples of known computing systems, environments, and / or configurations that may be suitable include, but are not limited to, personal computers, server computers, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.

[0187] The computing environment 710 includes a computing device 720 (a computer or other type of machine, such as an array determination machine, photocell, photomultiplier tube, optical reader, sensor, etc.) which includes a processing unit 721, system memory 722, and a system bus 723 that operably connects various system components, including the system memory 722, to the processing unit 721. There may be one or more processing units 721, such that the processor of the computing device 720 includes a single central processing unit (CPU) or multiple processing units commonly referred to as a parallel processing environment. The computing device 720 may be a conventional computer, a distributed computer, or any other type of computer.

[0188] The system bus 723 may be one of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of the various bus architectures. System memory may also be simply referred to as memory and includes read-only memory (ROM) 724 and random access memory (RAM) 725. A basic input / output system (BIOS) 726, which includes basic routine operations that help transfer information between elements within the computing device 720, such as during startup, is stored in ROM 724. The computing device 720 may further include a hard disk drive 727 for reading from and writing to a hard disk (not shown), a magnetic disk drive 728 for reading from and writing to a removable magnetic disk 729, and an optical disk drive 730 for reading from and writing to a removable optical disk 731 such as a CD-ROM or other optical media.

[0189] The hard disk drive 727, the magnetic disk drive 728, and the optical disk drive 730 are connected to the system bus 723 by hard disk drive interfaces 732, magnetic disk drive interfaces 733, and optical disk drive interfaces 734, respectively. The drives and their associated computer-readable media provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device 720. Any type of computer-readable media capable of storing computer-accessible data, such as magnetic cassettes, flash memory cards, digital video discs, Bernoulli cartridges, random access memory (RAM), and read-only memory (ROM), may be used in the operating environment.

[0190] Some program modules, including an operating system 735, one or more application programs 736, other program modules 737, and program data 738, may be stored on a hard disk 727, magnetic disk 728, optical disk 730, ROM 724, or RAM 725. The user may input commands and information to the computing device 720 via input devices such as a keyboard 740 and a pointing device (e.g., a mouse) 742. Other input devices (not shown) may include a microphone, joystick, gamepad, satellite antenna, scanner, etc. These and other input devices are often connected to the processing unit 721 via a serial port interface 746 coupled to the system bus 723, but may be connected via other interfaces such as a parallel port, game port, or Universal Serial Bus (USB). A monitor 747 or other type of display device is also connected to the system bus 723 via an interface such as a video adapter 748. In addition to the monitor 747, the computer typically includes other peripheral output devices (not shown) such as speakers and printers.

[0191] The computing device 720 may operate in a network environment using logical connections to one or more remote computers, such as remote computers 749. These logical connections may be achieved by communication devices coupled to the computing device 720, by parts of the computing device 720, or in other ways. The remote computers 749 may be another computer, server, router, network PC, client, peer device, or other common network node, typically including many or all of the above elements for the computing device 720, although only memory storage devices are shown in Figure 7. The logical connections shown in Figure 7 include local area networks (LANs) 751 and wide area networks (WANs) 752. Such networking environments are common in office networks, enterprise-wide computer networks, intranets, and the internet, all of which are types of networks.

[0192] When used in a LAN network environment, the computing device 720 is connected to the LAN 751 via a network interface or adapter 753, which is a type of communication device. When used in a WAN network environment, the computing device 720 often includes a modem 754, a type of communication device, or any other type of communication device for establishing communication via the WAN 752. The modem 754, which may be internal or external, is connected to the system bus 723 via a serial port interface 746. In a network environment, program modules described for the computing device 720 or a part thereof may be stored in a remote memory storage device. The network connections shown are non-limiting examples, and it should be understood that other communication devices may be used to establish communication links between computers. [Examples]

[0193] VII. Examples Exemplary examples of the present invention are provided in the embodiments and further illustrate the advantages and features of the invention, but are not intended to limit the scope of the invention. These examples are typical of what may be used, but alternatively, other procedures, methodologies, or techniques known to those skilled in the art may be used.

[0194] Purpose and relevance of translation Adjuvant chemotherapy (ACT) after surgery is standard care practice for patients with stage III colon cancer. ACT decisions for non-metastatic colon cancer are currently based on clinicopathological risk factors. All patients with stage III colon cancer are eligible for ACT, even though more than 50% are cured by surgery alone. Furthermore, only 15-20% of patients who receive ACT benefit from it, while all patients are at risk of developing significant side effects. Therefore, there is an urgent and unmet need to identify these stage III colon cancer patients who are truly at risk of recurrence after surgery and who could benefit from ACT. Fluid biopsy detection of circulating tumor DNA (ctDNA) after resection of the primary tumor makes it possible to identify patients with micrometastatic disease at high risk of experiencing disease recurrence.

[0195] ctDNA-based minimal residual disease (MRD) detection is a potent prognostic biomarker for disease recurrence in stage II and III colorectal cancer. Postoperative MRD detection is technically challenging due to the extremely low levels of ctDNA. Tumor information-based WGS approaches are promising for MRD testing, given their ability to track thousands of tumor-specific mutations without requiring personalized assay development. However, the clinical performance of these methods is not fully established. Herein, a novel tumor information-based WGS approach for detecting MRD is described herein. The Future Netherlands Colorectal Cancer Cohort (PLCRC) sub-study PROVENC3 aimed to determine the clinical validity of postoperative ctDNA status for predicting recurrence within 3 years in patients with stage III colorectal cancer treated with ACT.

[0196] The PROVENC3 study determined the clinical validity of a novel whole-genome sequencing-based ctDNA detection assay in adjuvant-chemotherapy-treated stage III colon cancer patients. Combining ctDNA test results with established clinicopathological risk factors allowed for differentiation of patients into groups with either a very low or very high risk of recurrence within three years. These data have broad implications for modifying current clinical practice treatment plans and enable the design of ctDNA-inducible intervention (gradual palliative care) elevation trials aimed at improving disease management in patients with stage III colon cancer.

[0197] Overview of the experimental design Blood samples were collected before surgery, after surgery, and after ACT. Plasma ctDNA detection based on tumor information was performed via integrated whole-genome sequencing (WGS) analysis of formalin-fixed paraffin-embedded tumor tissue DNA (80x), leukocyte germline DNA (40x), and plasma cell-free DNA (30x).

[0198] Materials and methods 1. PROVENC3 Research Group: PLCRC Infrastructure Mentally healthy patients aged 18 years or older diagnosed with colorectal cancer were recruited from both academic and non-academic hospitals in the Netherlands to participate in the ongoing Future Netherlands Colorectal Cancer Cohort (PLCRC, NCT02070146). Informed consent regarding the collection of long-term clinical and survival data was required for participation in the PLCRC. Patients were then given the option to consent to: 1) completing 205 health-related quality of life, functional outcome, and workability questionnaires; 2) biobanking of tumor and normal tissue; 3) blood sample collection; and 4) participation in studies conducted within the cohort infrastructure. Treatment-naive non-metastatic colorectal cancer (CRC) patients who gave informed consent for the PLCRC and additional blood sampling were included in the observational PLCRC sub-study, MEDOCC (Molecular Early Detection of Colorectal Cancer). Patients with stage III colon cancer who initiated adjuvant chemotherapy (ACT) after surgery and whose postoperative blood was available were included in the PLCRC-MEDOCC sub-study PROVENC3 at 26 hospitals from 2016 to 2021. One patient with stage III rectal cancer treated as colon cancer (cap(ox) / no radiation) was also included in the cohort. Clinical data were collected via the Netherlands Cancer Registry and through field visits by the PLCRC research team.

[0199] The PLCRC study was conducted in accordance with the Declaration of Helsinki and approved by the Central Committee on Research Involving Human Subjects (CCMO: NL47888.041.14). All patients signed written informed consent for participation in the study and for the collection of blood and tissue samples for translational research. The PLCRC sub-study PROVENC3 was approved by the Institutional Review Board (IRB) of the Netherlands Cancer Institute, Amsterdam, the Netherlands (Protocol CFMPB472).

[0200] 2. Sample collection and processing Formalin-fixed paraffin-embedded (FFPE) tumor blocks were requested through PALGA, the Dutch national network and registry for histopathology and cytopathology. Hematoxylin and eosin (H&E) slides were evaluated by pathologists, and the "tumor region" was outlined on the slide for macroscopic dissection. DNA was isolated from the FFPE slides using the QIAGEN AllPrep DNA / RNA FFPE kit (QIAGEN, Hilden, Germany) and stored for only a short period at -20°C or 4°C before shipment. DNA quality and quantity were measured using the Qubit dsDNA High-Sensitivity Assay (Thermo Fisher Scientific, USA) with Nanodrop One (Isogen, Ijsselstein, The Netherlands) and Qubit 3.0 Fluorometer (Molecular Probes, Leiden, The Netherlands).

[0201] Blood samples were collected before surgery, before the initiation of adjuvant chemotherapy after surgery, after the completion of adjuvant chemotherapy, and every 6 months for up to 3 years. Blood was collected at participating hospitals using cell-stabilized BCT tubes (Streck, La Vista, NE) and shipped to the Netherlands Cancer Institute. Cell-free plasma and leukocytes (WBCs) were separated by centrifugation at 1,700xg for 10 minutes, followed by 20,000xg for 10 minutes, and then stored at -80°C until further processing. Cell-free DNA (cfDNA) was isolated from available plasma using the QIAsymphony DSP Circulating DNA Kit (QIAGEN, Hilden, Germany) with a fixed elution volume of 60 μL. Genomic DNA was isolated from WBCs using the QIAsymphony DSP DNA Midi Kit (QIAGEN, Hilden, Germany) and a 1 mL blood protocol. cfDNA and genomic DNA from WBCs were stored at -20°C until further processing. The Qubit dsDNA High-Sensitivity Assay (Thermo Fisher, Waltham, MA) was used to quantify DNA yield for next-generation sequencing. After deidentification and blinding of samples, they were shipped to Personal Genome Diagnostics (Labcorp, Baltimore, MD) for sample testing and analysis. Postoperative ctDNA was evaluated for all patients in the cohort. Preoperative ctDNA was evaluated for 18 of 22 postoperative ctDNA-positive patients with available blood samples, and 33 patients were randomly selected from the remaining cohort. Post-ACT ctDNA was evaluated for 13 of 22 postoperative ctDNA-positive patients with available blood samples.

[0202] 3. Non-cancerous donor plasma and commercial cell line cohorts All patients provided written informed consent, and the study was conducted in accordance with the Declaration of Helsinki. Non-cancerous donor plasma samples were obtained from Discovery Life Sciences (Alabama, USA) with the approval of the Institutional Review Board. Human tumor and normal cells were obtained from previously characterized cell lines using ATCC (Virginia, USA) (COLO-829, HCC-1187, HCC-1143, HCC-1954) and SeraCare (Massachusetts, USA) (SeraSeq gDNA TMB-mix score 26). cfDNA was isolated from plasma using the Qiagen Circulating Nucleic Acid kit (Qiagen, Germany), and its concentration was assessed using the Qubit dsDNA High-Sensitivity Assay (Thermo Fisher, USA). Genomic DNA was isolated from cell line samples using the QIAamp DNA Blood Mini Kit (Qiagen, Germany), and its concentration was evaluated using the Qubit dsDNA Broad Range Assay (Thermo Fisher, USA).

[0203] 4. NGS analysis of tumor tissue-derived and germline genomic DNA Genomic DNA was quantified using the Qubit dsDNA Broad Range Assay (Thermo Fisher, USA), and up to 400 ng of DNA was sheared to a target fragment size of approximately 450 base pairs (bp) using Covaris Focused Sonication (Covaris, USA). In addition, genomic DNA derived from FFPE tumor tissue was repaired using PreCR Repair Mix (New England Biolabs, USA). Whole-genome next-generation sequencing libraries were prepared from fragmented genomic DNA via end repair, A-tailing, and adapter ligation with the KAPA HyperPrep reagent kit, according to the manufacturer's protocol (Roche, USA). These libraries were then amplified and pooled via 7-cycle polymerase chain reaction (PCR) and sequenced with 150 bp paired end reads at a target depth of 80x for tumor samples and 40x for germline samples using the Illumina NovaSeq6000 platform (Illumina, USA). After demultiplexing using bcl2fastq (Illumina, USA), the FASTQ files were aligned to the GRCh38 human reference genome using BWA-MEM (v0.7.15). PCR duplication was marked using Novosort (v1.03.01), and base quality score recalibration was performed using GATK BQSR (v4.1.0). The aligned BAM files were subjected to single nucleotide variant (SNV) analysis using MuTect2 (GATK v4.0.5.1), Strelka2 (v2.9.3), and Lancet (v1.0.7). SNVs were annotated with high confidence if reported by at least two variant callers.

[0204] 5. NGS analysis of plasma-derived cfDNA and artificial DNA Fragmented and matched cfDNA and artificial DNA obtained from tumor and germline cell lines were quantified using the Qubit dsDNA High-Sensitivity Assay (Thermo Fisher, USA). Whole-genome next-generation sequencing libraries were prepared from cell-free or artificial DNA using 10 ng of DNA targets via end repair, A-tailing, and adapter ligation with custom molecular barcoding adapters. These libraries were then amplified via 5 cycles of PCR, pooled, and sequenced with 150 bp paired end reads at a 30-fold target depth using the Illumina NovaSeq6000 platform (Illumina, USA). After demultiplexing, FASTQ files were quality-trimmed using Trimmomatic (v0.33) and aligned to the hg19 human reference genome using BWA-MEM2 (v2.2.1). Somatic variant identification was performed using VariantDx (v11.0.0), which has demonstrated high accuracy in distinguishing technical artifacts to enable somatic mutation detection and SNV analysis.

[0205] 6. Detection of tumor DNA via WGS analysis First, to ensure that tumor, germline, and plasma WGS datasets originated from the same subject, analysis was performed across 10,000 common single nucleotide polymorphisms. Next, quality control analysis was performed using Picard (v2.18.14), requiring a median insertion size of ≥150 bp and a sequencing depth of ≥20x for cfDNA samples, ≥40x for tumor samples, and ≥20x for germline samples. Tumor-specific single nucleotide variants were filtered into a candidate somatic mutation set by removing the following: (1) Variants observed in the 1000 genome (phase 3) or gnomAD(r2.0.1) population database, (2) Variants overlapping with 296 hg19 UCSC simple tandem repeat tracks, (3) Locations in tumors with a depth of less than 10x or consistent with normal, (4) Locations in tumors with fewer than 4 alternative alleles or more than 1 alternative allele in the corresponding germline, and (5) Variants with a tumor variant allele frequency (VAF) of less than 0.05 (stricter filtering was applied to T>C / A>G variants, which were removed if tumor VAF was less than 0.20 or alternative alleles were less than 10). Additional variant filtering was performed via blacklist generation, further removing variants if they were present in any non-cancerous donor with (1) more than 10% VAF, or (2) variants with more than 25% VAF across a cohort of 20 quadrupled non-cancerous donor plasma samples (total n=80). The final set of candidate tumor-specific variants was then compared to the unfiltered variant results of matched test samples. Candidate tumor-specific SNVs identified in the test samples were scored (in the range of 0 to 1) independently of the PROVENC3 cohort using a random forest machine learning algorithm trained with the caret package (v6.0.90) in the R statistical computing environment (v4.1.1). To avoid overfitting, model training utilized 5x cross-validation, and the number of variables selected per split procedure (hyperparameter mtry) was limited to the square root of the total number of input features.Variants present in well-paired mapped fragments with a random forest score > 0.25 were further evaluated, requiring alternative read mapping quality ≥ 30 and read-based mutation rate ≤ 5. Individual variant random forest scores were then aggregated and normalized based on the total number of tumor-specific SNVs evaluated. The normalized random forest scores (NRFS) were then compared to a non-cancerous donor cohort, and a cutoff of one standard deviation above the observed maximum NRFS was required to report individual test samples as having evidence of tumor-specific variants. The estimated tumor fraction (referred to as "aggregate ctDNA VAF") was then calculated for each positive test sample based on observed aggregate variant allele observations as the percentage of total unique coverage for all individual tumor-specific variants evaluated. The analytical sensitivity of the tumor-informed WGS approach was evaluated using five commercially available cell lines across four log tumor content ranges, demonstrating detection limits of 95% for 0.005% tumor content and 50% for 0.001% tumor content. Furthermore, the observed tumor fractions were also highly correlated with the reference tumor fractions (Pearson correlation coefficient = 0.96, p < 0.001). Analytical specificity was determined through the analysis of 119 non-cancerous donor plasma samples evaluated against 17 reference whole-genome somatic mutation datasets, demonstrating a specificity of 99.6% (2,015 / 2,023). Finally, analysis of externally prepared reference control samples showed highly reproducible results for estimated tumor fractions across 24 independent runs evaluated for the PROVENC3 clinical study (n=45, CV=7.2%). See Figures 15A–15D.

[0206] 7.Statistical analysis Differences in baseline features between the compared groups were analyzed using Fisher's exact test for the classification variable and the Mann-Whitney test for the continuous variable. In the complete cohort, postoperative ctDNA positive versus postoperative ctDNA negative, the primary outcome measure was "time to recurrence" as defined by Cohen et al. The only event considered was recurrence. For the time to event analysis, patients were censored at the last point in time with available follow-up information without reporting a recurrence, or at 36 months if the FU was longer. Patients with no event but less than one year of available follow-up were excluded from the analysis. Kaplan-Meier estimators and fitted Cox regression models were used for univariate time-event analysis. The clinicopathological variables evaluated were selected based on clinical relevance: clinicopathological risk status (low risk = T1-3N1, high risk = T4 and / or N2), T status determined by pathological report (T1-3, T4), N status determined by pathological report (N1, N2), and microsatellite instability (MSI) status determined by next-generation sequencing of the primary tumor (stable, unstable). After stratifying each clinicopathological covariate based on postoperative ctDNA status (clinopathological risk + ctDNA status, T status + ctDNA status, N status + ctDNA status, MSI status + ctDNA status), changes in hazard ratios were also evaluated. Kaplan-Meier estimation curves were also fitted to these models.

[0207] Furthermore, we evaluated whether postoperative ctDNA status provided an additional independent predictor for recurrence in addition to clinicopathological variables. Several (multivariate) Cox regression models were fitted. First, the added value of ctDNA status was determined by fitting multivariate models combining clinicopathological risk factors and ctDNA status and performing likelihood ratio tests among them (LRT1: clinicopathological risk vs. clinicopathological risk + ctDNA status, LRT2: clinicopathological risk + MSI status vs. clinicopathological risk + MSI status + ctDNA status, LRT3: T status + N status vs. T status + N status + ctDNA status, LRT4: T status + N status + MSI status vs. T status + N status + MSI status + ctDNA status). Secondly, we evaluated the independent predictors of each variable in the models by independently exploring the hazard ratios for each variable in the two best-case models obtained from likelihood ratio tests: (Model 1: Clinicopathological risk + MSI status + ctDNA status; Model 2: T status + N status + MSI status + ctDNA status). All statistical and survival analyses were performed using the R package "Survival" (R version 4.2.1) for survival.

[0208] 8. Training Machine Learning Models The random forest model was trained using the caret package (v6.0.90) within the R statistical computing environment (v4.1.1). A set of 62 features was provided for model training, consisting of 1,000 true positive variants and 1,000 false positive variants from the COLO-829 cell line across dilution levels of 0.01, 0.001, 5×10⁻⁴, 2×10⁻⁴, 1×10⁻⁴, 5×10⁻⁵, and 1×10⁻⁵. To avoid overfitting, model training utilized 5x cross-validation, and the number of variables selected per split procedure (hyperparameter mtry) was limited to the square root of the total number of input features.

[0209] The feature set includes: ('ATcount', 'AverageQualityScore', 'BaseFrom.ToAtoG', 'BaseFrom.ToAtoT', 'BaseFrom.ToCtoA', 'BaseFrom.ToCtoG', 'BaseFrom.ToCtoT', 'BaseFrom.ToGtoA', 'BaseFrom.ToGtoC', 'BaseFrom.ToGtoT', 'BaseFrom.ToTtoA', 'BaseFrom.ToTtoC', ' BaseFrom.ToTtoG', 'DistinctCoverage', 'DistinctNoOlapMuts', 'DistinctOlap1Mut', 'DistinctOlapMuts', 'DistinctOlapReads', 'DistinctPairs', 'DistMutPairsEORA', 'DistMutPairsEORAplusB', 'DistMutPairsEORB', 'Dust','DustRaw', 'EORMutPct', 'F1R2Mut ', 'F2R1Mut', 'Forward', 'GCcount', 'MaskedMutPct', 'MaskedPairs', 'MMAvg', 'MMTumor', 'MutCountCov', 'MutMMAvg', 'MutPct', ' NonMutFwd', 'NonMutRev', 'NoOlapMut', 'NumAlleles', 'Olap1Mut', 'OlapMuts', 'OlapReads', 'PolyMut', 'PolyN', 'PolyNN', 'PolyN NN', 'ProperPairs', 'ProperPairsPct', 'Reverse', 'RMDScore', 'RptMask', 'SBFisherLeft', 'SBFisherRight', 'SBFisherTwotail', 'SBMutProportion', 'SBNonMutProportion', 'SBPropDelta', 'TumMutRMSMAPQ', 'TumNERDistMean', 'TumNERDistSD', 'TumRMSMAPQ').

[0210] Example 2: Overview of the PROVENC3 study Figure 8A provides an overview of the PROVENC3 study population, along with the main exclusion criteria used to obtain the final patient population used for the final analysis. Initially, we evaluated to include 268 stage III colon cancer patients who had undergone ACT and whose postoperative blood was available. Patients whose blood was collected on postoperative day 1 or 2 (14) and / or whose primary tumor sample was unavailable (13) were excluded. After the first filtering step, 243 patients remained for whom postoperative blood and tumor tissue were available. Of these 243 patients, those with insufficient isolated DNA from tumor tissue (3), insufficient total isolated DNA for assay (5), insufficient postoperative cfDNA sample (5), or whose variant profile did not pass quality control (19) were also excluded from the study. As a result, the study included 209 stage III ACT-treated colon cancer patients with a median follow-up of 40 months (interquartile range (IQR) 29–48 months).

[0211] Figure 8B shows an exemplary schematic approach to the PROVENC3 (Prognostic value for early notification by ctDNA in stage 3 colon cancer) study. After informed consent, blood samples were collected pre- and post-operatively. Tumor samples were collected during surgery for FFPE preparation, and blood samples were collected 3–65 days post-surgery. The mean (median) number of days for collection was 13 days, and the interquartile range (IQR) for collection was 4–20 days post-surgery. Post-surgery, all patients underwent ACT, and blood samples were collected every 6 months and post-ACT for up to 3 years. All tissue and blood samples were sent to the Central Laboratory (Netherlands Cancer Institute), and clinical data were collected in the Netherlands Cancer registry by IKNL.

[0212] The clinical sensitivity of the ctDNA detection test was evaluated using blood samples collected from 149 patients before surgery. Of these 149 patients, 134 (90%) were found to be ctDNA positive, highlighting the high sensitivity of the ctDNA test.

[0213] Blood samples collected from 209 patients post-surgery were used to determine the prognostic value of post-operative ctDNA status and correlated with clinicopathological risk factors to predict the risk of recurrence. In total, 28 out of 209 patients (13%) were post-operative ctDNA positive. The median post-operative aggregate ctDNA variant allele frequency (VAF) was 0.035% (range 0.01% to 3.13%). As shown in Table 2, none of the baseline clinicopathological features evaluated were significantly associated with post-operative ctDNA status. [Table 2-1] [Table 2-2]

[0214] Clinicopathological risk factors were based on the patient's tumor pathological stage (T status) and lymph node pathological status (N status). T status was assessed on a scale of 1 to 4. T1 status indicates that the tumor is confined to the inner lining of the intestine. T2 means that the tumor has grown into the muscular layer of the intestinal wall. T3 means that the tumor has grown into the outer lining of the intestinal wall but has not grown through it. T4 means that the tumor has grown into the outer lining of the intestinal wall and has spread to other tissues and / or organs. Patients with a clinicopathological risk factor of pT4 are considered to be at high risk of recurrence, and patients with a clinicopathological risk factor of pT1 are considered to be at low risk of recurrence. N status was assessed on a scale of 1 and 2. As shown, N1 status is divided into three stages: N1a, N1b, and N1c. N1a means that cancer cells are present in one nearby lymph node, N1b means that cancer cells are present in two or three nearby lymph nodes, and N1c means that cancer is not present in nearby lymph nodes, but cancer cells are present in the tissue near the tumor. N2 is divided into two stages, N2a and N2b. N2a means that cancer cells are present in four to six nearby lymph nodes, and N2b means that cancer cells are present in seven or more nearby lymph nodes. Patients with a pN2 clinicopathological risk are considered to be at high risk of recurrence, and patients with a 3N1 clinicopathological risk factor are considered to be at low risk of recurrence.

[0215] As shown in Figure 8B, blood samples from 47 of the 209 patients were used to evaluate the location of recurrence and determine whether there was a difference in the time to recurrence based on the postoperative ctDNA status. Furthermore, blood samples from 170 patients were evaluated by ACT for postoperative prognostic ctDNA status and ctDNA clearance.

[0216] Figure 9 shows dot graphs (Figure 9A) and box plots (Figure 9B) for postoperative ctDNA status and cfDNA concentration. ctDNA-positive patients experiencing relapse were not underestimated among blood samples with higher cfDNA levels extracted in week 1 (P=0.5). Figure 9A shows summary cfDNA concentrations (plasma in ng / mL) plotted against the time of postoperative blood collection for 47 patients who experienced relapse. Figure 9B shows that cfDNA concentrations were similar between ctDNA-positive and ctDNA-negative patients, whether blood was collected in the first week postoperatively (P=0.96) or from week 2 onwards (P=0.98). The relapse rate was similar between patients whose blood was collected in week 1 (26%) and those whose blood was collected from week 2 onwards (26%). The y-axis shows cfDNA concentration on a logarithmic scale. Abbreviations: cfDNA, cell-free DNA; ng, nanograms; mL, milliliters.

[0217] Example 3: Analytical study to validate the ctDNA detection assay Figure 10A shows an exemplary schematic diagram of tumor-information-based detection of ctDNA. To detect ctDNA, Labcorp Plasma Detect® was used, utilizing integrated WGS analysis of patient-fit formalin-fixed paraffin-embedded (FFPE) tumor tissue DNA, leukocyte-derived DNA, and plasma cell-free DNA (cfDNA). Samples were sequenced to the depths of tumor tissue DNA (80x), germline DNA (40x), and plasma cell-free DNA (30x). WGS identified a median of 5,108 highly reliable tumor-specific single nucleotide variants (IQR 3,776–7,411) per patient used for plasma ctDNA detection. Variant scores were generated using machine learning techniques (e.g., random forest classification). Furthermore, useful features from candidate somatic cell variants were used to determine whether plasma samples were ctDNA+ or ctDNA- and to estimate the level of ctDNA in the plasma samples.

[0218] Figure 10B shows the results of an analytical study validating the assay workflow described in Figure 10A. A devised reference model derived from five commercially available cell lines, including lung cancer (n=1), breast cancer (n=3), and melanoma (n=1), was used. The devised samples were generated from three cell lines (COLO-829, HCC-1187, and HCC-1143) and evaluated in triplicate at tumor content levels of 10%, 1%, 0.10%, 0.05%, 0.02%, 0.01%, 0.005%, and 0.001%. Additional sample series (HCC-1187, HCC-1954) were generated and evaluated in triplicate with externally prepared reference control samples (SeraSeq gDNA TMB-Mix score 26) evaluated at tumor contents of 0.05%, 0.01%, 0.005%, 0.001%, and 0.001%, respectively, to increase the number of data points close to the expected detection limits. These analyses demonstrated detection limits of 95% for 0.005% tumor content and 50% for 0.001% tumor content.

[0219] Figure 10C shows the results of an analytical specificity study for the assay shown in Figure 10A. The assay was found to have a specificity of 99.6% (2,015 / 2,023) across 119 non-cancerous donor plasma samples evaluated against 17 reference whole-genome somatic mutation datasets.

[0220] Figure 10D shows the results of an analytical reproducibility study for the assay shown in Figure 10A. High reproducibility was observed (CV=7.2%) across 24 independent runs evaluated in the PROVENC3 clinical study using externally prepared reference control samples.

[0221] Table 3 below lists the blank limit (LoB), limit of detection (LoD), and analytical study designs for the clinical confirmatory studies. LoB studies were conducted using sets of non-cancerous donor plasma to determine the specificity of ctDNA detection. LoD studies were conducted using cell line titration to determine the minimum level at which ctDNA could be confidently identified. Clinical confirmatory studies were conducted using preoperative plasma to test the accuracy of ctDNA-positive calling in sets of clinical samples. [Table 3]

[0222] Figure 11A shows the specificity evaluation of non-cancerous donor plasma samples (n=119) to 17 reference somatic mutation datasets in one embodiment of the claimed invention. The reference somatic mutation datasets included 14 clinical FFPE tumors and 3 tumor cell line samples, including head and neck, colorectal, breast, and lung cancers. The reference tumors and normal samples, as well as non-cancerous donor plasma cases, were prepared, sequenced, and analyzed as detailed in the Examples.

[0223] Figure 11B shows the results of an analytical sensitivity study relating to one embodiment of the claimed invention. The assay described in the claims demonstrated high sensitivity to ctDNA detection at a detection limit (95%) of 0.005% tumor content using a devised reference model derived from commercially available cell lines including lung cancer, breast cancer, and melanoma. The observed tumor percentages also correlated highly with the reference tumor fraction (Pearson correlation coefficient = 0.96, p < 0.001). Cell line materials were titrated to normal levels of 0.05%, 0.01%, 0.005%, 0.001%, and 0%, sheared to 160–170 bp, and subjected to double-sided bead washing to simulate cfDNA. The prepared cell lines were prepared, sequenced, and analyzed as detailed in the examples.

[0224] Figure 12 shows the results of ctDNA detection across multiple solid tumors in one embodiment of the claimed invention. ctDNA was detected in 71% (10 / 14) of clinical samples (diamonds) at significantly above the background patient-specific reference levels established across an independent non-cancerous donor cohort (circles). These clinical samples were obtained prior to surgical intervention across patients with stage II and IV colorectal tumors and head and neck tumors. Tumor DNA and matched normal DNA samples from each patient were prepared, sequenced, and analyzed as detailed in the Examples.

[0225] Example 4: Detection of ctDNA after surgery Figure 13A shows Kaplan-Meier estimates of time to recurrence (TTR) stratified by postoperative ctDNA status. Patients treated with ACT with positive postoperative ctDNA had a higher risk of recurrence compared to patients with undetectable postoperative ctDNA (hazard ratio (HR) 6.3 [95% confidence interval (CI): 3.5–11.3], P<10). -8 There was no difference in the median follow-up between patients with ctDNA-positive and ctDNA-negative status after surgery (P=0.2). Figure 13B shows the proportion of patients at risk of recurrence at 3 years. Of the 209 patients evaluated, 181 were found to be ctDNA-negative and 28 were found to be ctDNA-positive. Among ctDNA-negative patients, 83% were disease-free and no disease was detected at 3 years. On the other hand, 17% of ctDNA-negative patients showed recurrence. Among ctDNA-positive patients, only 36% were disease-free, and 64% showed recurrence, indicating that patients with a positive ctDNA status after surgery are at high risk of recurrence.

[0226] Next, the prognostic value of ctDNA was evaluated in the context of established clinicopathological risk stratification factors for recurrence in stage III colon cancer. High-risk patients had the (pT4, pT4 and / or pN2, pN2) risk factor, and low-risk patients had the (pT1, pT1-3N1) risk factor, where "T" refers to tumor status and "N" refers to lymph node status. Figure 13C shows the Kaplan-Meier estimates of TTR stratified by clinicopathological risk (HR 3.5 [95% CI: 2.0-6.5], P<10). -4 Table 4 shows the univariate analysis of postoperative ctDNA status and clinicopathological risk factors. Figure 13D shows the proportion of patients at risk of recurrence at 3 years. 127 patients were at low risk of recurrence; 83% of low-risk patients did not experience recurrence, while 13% did. 82 patients were found to be at high risk of recurrence. Of these, 60% did not experience recurrence, and 40% experienced disease recurrence. [Table 4]

[0227] As shown in Figure 13E and Table 5, multivariate analysis defined four groups according to the combination of clinicopathological risk and postoperative ctDNA status. The recurrence risk in clinicopathologically high-risk patients was further increased when the patient was ctDNA-positive, and the recurrence risk in clinicopathologically low-risk patients was further decreased when the patient was ctDNA-negative. Therefore, as shown in Figure 13F, a significant survival difference exists between clinicopathologically high-risk ctDNA-positive patients and clinicopathologically low-risk ctDNA-negative patients (3-year recurrence risk 82% vs. 7%) (HR 28.9 [95% CI: 10.6~78.2], P<10). -10 ). [Table 5-1] [Table 5-2]

[0228] To further support Figures 13E and 13D, Figure 14 shows Kaplan-Meier estimates of Cox regression analysis for clinicopathological low-risk (Figure 14A) and high-risk (Figure 14B) groups stratified by postoperative ctDNA status with confidence intervals. Patients with a positive postoperative ctDNA status are at higher risk of postoperative recurrence and are more likely to recur earlier compared to high-risk ctDNA-negative patients.

[0229] Furthermore, the values ​​of postoperative ctDNA status, which was added to multivariate Cox models including clinicopathological risk and MSS status, were evaluated by performing likelihood ratio tests (LRTs) to include or exclude ctDNA status. As shown in Table 6, the inclusion of ctDNA status significantly improved the models (LRT P<10⁻⁷). Multivariate Cox regression models were fitted to include different clinicopathological variables. In addition, four likelihood ratio tests were performed to evaluate the added value of ctDNA status in each model. It was found that ctDNA status corresponds to the postoperative time point. [Table 6]

[0230] As shown in Table 7, ctDNA status was the strongest independent predictor of relapse (HR 6.8) in models that included clinicopathological risk (HR 4.0) and microsatellite stability (MSS) status (HR 0.7, ns). [Table 7]

[0231] In summary, Tables 4–7 show the effects of tumor (T), nodule (N), and MSS status as independent risk factors in multivariate models, with postoperative ctDNA status remaining the strongest predictor of recurrence.

[0232] Time until recurrence Figure 15 shows the estimated time to relapse based on ctDNA status. Figures 16A and 16B show that among 47 (22%) of 209 patients who experienced relapse, ctDNA-positive patients tended to have a shorter time to relapse compared to ctDNA-negative patients. Figure 15B also shows that among 28 postoperative ctDNA-positive patients, 7 (25%) remained disease-free for at least 36 months postoperatively, suggesting they may benefit from ACT treatment. Next, postoperative ctDNA analysis was performed on 170 of the 209 patients. As shown in Figures 16C and 16D, ctDNA-positive patients had a higher risk of developing relapse (HR 7.9 [95% CI: 3.55~15.9], P<10⁻⁸). Further details on postoperative ctDNA status, ctDNA clearance, and relapse location are provided in Tables 2 and 4.

[0233] Consideration Stratification based on clinicopathological risk and postoperative ctDNA status can lead to common ACT decisions. Gradual easing or withholding of adjuvant therapy in postoperative ctDNA-negative patients with low clinicopathological risk should be evaluated in clinical trials along with appropriate MRD surveillance. The clinical sensitivity of ctDNA for detecting disease recurrence in the PROVENC3 study indicates that ctDNA is detectable approximately 6–10 months before clinically detected recurrence. This provides an opportunity to evaluate interventions in studies designed for this selected patient population.

[0234] In conclusion, the PROVENC3 study demonstrates the strong potential of a WGS-based plasma ctDNA approach to MRD trials that informs tumors, enabling the design of robust clinical practice that will transform interventional ctDNA-inducible studies to improve disease management in patients with stage III colon cancer.

[0235] VIII. Additional Considerations The implementation of the above techniques, blocks, steps, and means can be carried out in numerous ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the above functions, and / or a combination thereof.

[0236] Furthermore, it should be noted that embodiments may be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts can describe operations as a series of processes, many operations can be performed in parallel or simultaneously. Moreover, the order of operations can be rearranged. A process terminates when its operation is complete, but it may have additional steps not shown in the diagram. A process can correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, its termination may correspond to the return of the function to the calling function or the main function.

[0237] Furthermore, embodiments can be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented by software, firmware, middleware, scripting languages, and / or microcode, program code or code segments for performing the required tasks can be stored in a machine-readable medium such as a storage medium. Code segments or machine-executable instructions can represent procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, scripts, classes, or any combination of instructions, data structures, and / or program statements. Code segments can be coupled to other code segments or hardware circuits by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc., can be passed, transferred, or transmitted via any preferred means, including memory sharing, message passing, ticket passing, network transmission, etc.

[0238] In firmware and / or software implementations, the methodology may be implemented in modules (e.g., procedures, functions, etc.) that perform the functions described herein. Any machine-readable medium that tangibly embodies instructions may be used when implementing the methodology described herein. For example, software code may be stored in memory. Memory may be implemented within or outside the processor. As used herein, the term “memory” means any type of long-term, short-term, volatile, non-volatile, or other storage medium, and is not limited to any particular type of memory or number of memories, or the type of media in which the memory is stored.

[0239] Furthermore, as disclosed herein, the terms “storage medium,” “storage device,” or “memory” may refer to one or more memories for storing data, including read-only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage medium, optical storage medium, flash memory device, and / or other machine-readable media for storing information. The term “machine-readable media” includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage media that contain or can carry instructions and / or data.

[0240] While the principles of this disclosure are described above in relation to specific apparatuses and methods, it should be clearly understood that this description is for illustrative purposes only and does not limit the scope of this disclosure.

[0241] Publications cited herein and materials from which they are cited are incorporated herein by reference.

Claims

1. A computer implementation method, The process involves generating sequence reads from tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples collected from the same patient, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the tumor nucleic acid sample, the non-cancerous nucleic acid sample, and the non-tissue sample, respectively, a tumor variant calling file, a non-cancerous variant calling file, and a non-tissue variant calling file are generated. The tumor variant calling file is compared with the non-cancerous variant calling file to generate a list of somatic variants, The process involves comparing the list of somatic cell variants with the non-tissue variant calling file to generate a list of candidate somatic cell variants, The process involves generating a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants using a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, and determine whether the ctDNA status is positive or negative. A computer implementation method comprising generating a report providing the ctDNA status for the patient.

2. The computer implementation method according to claim 1, wherein the tumor nucleic acid sample is any body tissue or body fluid containing nucleic acids deemed to be cancer-positive, the non-cancerous sample is any body tissue or body fluid containing nucleic acids deemed not to be cancerous, and the non-tissue sample is any body fluid containing nucleic acids deemed to contain cell-free DNA and circulating tumor DNA.

3. The computer implementation method according to claim 2, wherein the tumor nucleic acid sample is cancer-positive tissue, the non-cancerous nucleic acid sample is leukocytes, and the non-tissue nucleic acid sample is plasma.

4. The computer implementation method according to claim 3, wherein the non-tissue nucleic acid sample is circulating tumor DNA.

5. The computer implementation method according to claim 3, wherein the non-cancerous nucleic acid sample and the non-tissue nucleic acid sample are collected from the same whole blood sample.

6. The computer implementation method according to claim 2, wherein the tumor nucleic acid sample is sequenced to at least 50 times its original depth, the non-cancerous nucleic acid sample is sequenced to at least 30 times its original depth, and the non-tissue nucleic acid sample is sequenced to at least 20 times its original depth.

7. The computer implementation method according to claim 6, wherein the tumor nucleic acid sample is sequenced to a depth of 80 times, the non-cancerous nucleic acid sample is sequenced to a depth of 40 times, and the non-tissue nucleic acid sample is sequenced to a depth of 30 times.

8. The computer implementation method according to claim 1, wherein the patient is diagnosed with cancer, undergoes surgery to remove one or more tumors, and receives therapeutic treatment after the surgery.

9. The computer implementation method according to claim 8, wherein the therapeutic treatment is adjuvant chemotherapy.

10. The computer implementation method according to claim 8, wherein the patient is diagnosed with colorectal cancer, head and neck cancer, lung cancer, breast cancer, or melanoma.

11. The computer implementation method according to claim 10, wherein the patient is diagnosed with colorectal cancer.

12. The computer implementation method according to claim 1, wherein the tumor nucleic acid sample, non-cancerous sample, and non-tissue sample are collected (i) before surgery, (ii) during surgery, (iii) approximately 3 to 65 days after surgery and before therapeutic treatment, (iv) approximately every 6 months for up to 3 years after surgery and after the therapeutic treatment, or (v) any combination thereof.

13. The computer implementation method according to claim 1, wherein the tumor variant calling file and the non-cancerous variant calling file are filtered using a set of filtering criteria, the set of filtering criteria comprising removing (i) variants annotated as low reliability, (ii) variants annotated as indels, (iii) variants observed in a genome database, (iv) variants that overlap simple tandem repeat tracks, (v) variants at genomic locations with coverage less than 10x, (vi) variants at genomic locations with fewer than 4 alternative alleles in the tumor nucleic acid sample or more than 1 in the non-cancerous nucleic acid sample, (vii) variants with a variant allele frequency of less than 0.05, or (viiii) any combination thereof.

14. The computer implementation method according to claim 1, wherein the list of candidate somatic cell variants includes substitutions, small indels, chromosomal rearrangements, copy number variations, microsatellite instability, or any combination thereof.

15. The computer implementation method according to claim 14, wherein the list of candidate somatic cell variants includes at least 40,000 to at least 70,000 somatic cell variants.

16. The computer implementation method according to claim 15, wherein each candidate somatic cell variant on the list of candidate somatic cell variants has at least 50 corresponding features.

17. The computer implementation method according to claim 16, wherein the features include a quality metric output from array determination, alignment, and variant invocation.

18. The computer implementation method according to claim 17, wherein the sequencing feature includes a quality score for any given base in the sequence read, the alignment feature includes alignment quality, read quality, strand information, a metric relating to the complexity of the region in the genome, or any combination thereof, and the variant calling feature includes a variant confidence score, base variant quality, or any combination thereof.

19. The computer implementation method according to claim 1, wherein, before generating the score, the classification model uses a set of non-cancerous donor samples to filter the list of candidate somatic cell variants to generate a filtered list of candidate somatic cell variants.

20. The aforementioned classification machine learning model is a random forest classifier that includes a set of trees having at least 500 decision trees, Each of the aforementioned trees generates a score for the input candidate somatic cell variant. The random forest classifier determines the final score by averaging the scores generated by each of the trees. The final score is compared with a predetermined threshold to determine whether the ctDNA status of the non-tissue nucleic acid sample is positive or negative. The aforementioned tree set takes into account at least 50 features associated with the candidate somatic cell variant, The computer implementation method according to claim 1, wherein each tree makes a class prediction by considering a different subset of features from the at least 50 features.

21. The computer implementation method according to claim 20, wherein the predetermined threshold is the maximum normalized score plus one standard deviation of the cohort of the reference variant.

22. The computer implementation method according to claim 20, wherein the final score is greater than or equal to the predetermined threshold and the ctDNA status is positive, and the final score is less than the predetermined threshold and the ctDNA status is negative.

23. The computer implementation method according to claim 1, wherein the ctDNA status is determined by normalizing the score and comparing the normalized score with the maximum normalized score + 1 standard deviation, and if the normalized score is greater than or equal to the maximum normalized score, the ctDNA status is positive.

24. The computer implementation method according to claim 23, wherein the ctDNA status represents the postoperative ctDNA status.

25. The computer implementation method according to claim 1, wherein the ctDNA status correlates with a clinicopathological risk factor for predicting survival rate, the clinicopathological risk factor predicts recurrence risk, and the clinicopathological risk factor includes the depth of tumor invasion and the spread of the tumor to adjacent lymph nodes.

26. The computer implementation method according to claim 25, wherein the correlation between the ctDNA status and the clinicopathological risk factors is included in the report, and the report further explains the patient's recurrence risk and predicted survival rate based on the patient's ctDNA status and clinicopathological risk factors.

27. A computer implementation method, The process involves generating sequence reads from non-tissue nucleic acid samples collected from patients, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the non-tissue sample, a non-tissue variant call file is generated, The process involves comparing a list of somatic cell variants with the non-tissue variant calling file to generate a list of candidate somatic cell variants, The process involves generating a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants using a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, and determine whether the ctDNA status is positive or negative. A computer implementation method comprising generating a report providing the ctDNA status for the patient.

28. A computer implementation method, Accessing a labeled training dataset, wherein the labeled training dataset includes ground truth true positive variants and associated features collected from patients with cancer, and ground truth false positive variants and associated features collected from non-cancer patients. Training a classification model using the labeled training dataset to generate a score, wherein the training is an iterative process starting at the first node of the first tree. Inputting a portion of the labeled training data into the classification model, Randomly selecting features of several variants from the aforementioned labeled training dataset, The determination of which variant features provide the best binary partition from among the features of several variants, wherein the determination is based on a subset of variant features that minimize the objective function. The generation process includes assigning the determined variant features to the first node, For several iterations or epochs, the iterative process is repeated in the second and subsequent nodes of the classification model, The iterative process is repeated at the first node of the second and subsequent trees until all variant features are assigned to the tree, A computer implementation method that includes outputting a trained classification model.

29. It is a system, One or more processors, One or more computer-readable media for storing instructions, wherein when the instructions are executed by the one or more processors, the system The process involves generating sequence reads from tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples collected from the same patient, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the tumor nucleic acid sample, the non-cancerous nucleic acid sample, and the non-tissue sample, respectively, a tumor variant calling file, a non-cancerous variant calling file, and a non-tissue variant calling file are generated. The tumor variant calling file is compared with the non-cancerous variant calling file to generate a list of somatic variants. The list of somatic cell variants is compared with the non-tissue variant calling file to generate a list of candidate somatic cell variants. A classification machine learning model generates a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, which is either positive or negative, and A system comprising one or more computer-readable media for causing an operation to be performed, including generating a report providing the ctDNA status for the patient.

30. One or more non-temporary computer-readable media for storing instructions, wherein when the instructions are executed by one or more processors, the system... The process involves generating sequence reads from tumor nucleic acid samples, non-cancerous nucleic acid samples, and non-tissue nucleic acid samples collected from the same patient, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the tumor nucleic acid sample, the non-cancerous nucleic acid sample, and the non-tissue sample, respectively, a tumor variant calling file, a non-cancerous variant calling file, and a non-tissue variant calling file are generated. The tumor variant calling file is compared with the non-cancerous variant calling file to generate a list of somatic variants, The process involves comparing the list of somatic cell variants with the non-tissue variant calling file to generate a list of candidate somatic cell variants, The process involves generating a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants using a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, and determine whether the ctDNA status is positive or negative. One or more non-temporary computer-readable media that perform an operation including generating a report providing the ctDNA status for the patient.

31. It is a system, One or more processors, One or more computer-readable media for storing instructions, wherein when the instructions are executed by the one or more processors, the system The process of generating sequence reads from non-tissue nucleic acid samples collected from patients, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the non-tissue sample, a non-tissue variant calling file is generated. A list of somatic cell variants is compared with the non-tissue variant calling file to generate a list of candidate somatic cell variants. A classification machine learning model generates a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, which is either positive or negative, and A system comprising one or more computer-readable media for causing an operation to be performed, including generating a report providing the ctDNA status for the patient.

32. It is a system, One or more processors, One or more computer-readable media for storing instructions, wherein when the instructions are executed by the one or more processors, the system Accessing a labeled training dataset, wherein the labeled training dataset includes ground truth true positive variants and associated features collected from patients with cancer, and ground truth false positive variants and associated features collected from non-cancer patients. Training a classification model using the labeled training dataset to generate a score, wherein the training is an iterative process starting at the first node of the first tree. Inputting a portion of the labeled training data into the classification model, Randomly selecting features of several variants from the aforementioned labeled training dataset, The determination of which variant features provide the best binary partition from among the features of several variants, wherein the determination is based on a subset of variant features that minimize the objective function. The generation process includes assigning the determined variant features to the first node, For several iterations or epochs, the iterative process is repeated in the second and subsequent nodes of the classification model. The iterative process is repeated at the first node of the second and subsequent trees until all variant features are assigned to the tree, and A system comprising one or more computer-readable media that perform operations, including outputting a trained classification model.

33. One or more non-temporary computer-readable media for storing instructions, wherein when the instructions are executed by one or more processors, the system... The process involves generating sequence reads from non-tissue nucleic acid samples collected from patients, wherein the sequence reads are generated using whole-genome sequencing (WGS). By analyzing the sequence reads corresponding to the non-tissue sample, a non-tissue variant call file is generated, The process involves comparing a list of somatic cell variants with the non-tissue variant calling file to generate a list of candidate somatic cell variants, The process involves generating a score for each of the candidate somatic cell variants in the list of candidate somatic cell variants using a classification machine learning model, wherein the score is generated based on a plurality of classifications generated by the classification machine learning model. Based on the score, determine the ctDNA status of the patient, and determine whether the ctDNA status is positive or negative. One or more non-temporary computer-readable media that perform an operation including generating a report providing the ctDNA status for the patient.

34. One or more non-temporary computer-readable media for storing instructions, wherein when the instructions are executed by one or more processors, the system... Accessing a labeled training dataset, wherein the labeled training dataset includes ground truth true positive variants and associated features collected from patients with cancer, and ground truth false positive variants and associated features collected from non-cancer patients. Training a classification model using the labeled training dataset to generate a score, wherein the training is an iterative process starting at the first node of the first tree. Inputting a portion of the labeled training data into the classification model, Randomly selecting features of several variants from the aforementioned labeled training dataset, The determination of which variant features provide the best binary partition from among the features of several variants, wherein the determination is based on a subset of variant features that minimize the objective function. The generation process includes assigning the determined variant features to the first node, For several iterations or epochs, the iterative process is repeated in the second and subsequent nodes of the classification model, The iterative process is repeated at the first node of the second and subsequent trees until all variant features are assigned to the tree, One or more non-temporary computer-readable media that perform operations, including outputting a trained classification model.