Systems and methods for determining tumor fraction in cell-free nucleic acid

Through computer systems, the cell-free nucleic acids in liquid biological samples are analyzed, and the frequency of ctDNA variants is identified and quantified, which solves the problem of low-level ctDNA detection in the prior art, and realizes early diagnosis and treatment decision support for cancer.

CN112218957BActive Publication Date: 2025-08-08SDG OPS +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980037052.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-16
Filing Date
2019-04-16
Publication Date
2025-08-08
Estimated Expiration
2039-04-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively and economically detect and quantify the content of low-level circulating tumor DNA (ctDNA) in subject samples, especially in blood samples, affecting cancer diagnosis and treatment decisions.

Method used

Cell-free nucleic acids in liquid biological samples were analyzed by computer systems, variant frequency was identified using sequence reads, and compared with reference frequency, tumor scores were calculated, including evaluation of single nucleotide variants, insertion deletions, copy number variants, and epigenetic modifications.

Benefits of technology

It realizes efficient and accurate detection of low-level ctDNA, supports early diagnosis and treatment decisions for cancer, and provides dynamic monitoring of disease progression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112218957B_ABST
    Figure CN112218957B_ABST
Patent Text Reader

Abstract

The present invention discloses various systems and methods for determining a tumor score in a cell-free nucleic acid of a liquid biological sample from a subject. A plurality of sequence reads are obtained using the biological sample. The plurality of sequence reads are used to identify support for each variant in a variant set, thereby determining an observed frequency for each variant in the variant set. For each individual variant in the variant set, a corresponding reference frequency for the individual variant is obtained in a reference set, wherein each corresponding reference frequency in the reference set is for another variant in an abnormal solid tissue sample obtained from the subject. The observed frequency of each individual variant in the variant set is evaluated against the observed frequency of the individual variant in the reference set, thereby determining the tumor score in the cell-free nucleic acid of the liquid biological sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 658,479, filed April 16, 2018, entitled “Systems and Methods for Classifying Subjects Using the Frequency of Variants in Cell-Free Nucleic Acids,” the contents of which are incorporated herein by reference. Technical Field

[0003] Described herein are methods for determining a tumor fraction in a subject's cell-free nucleic acid to inform various improved classifiers for cancer classification, including detecting cancer at lower tumor fractions. Background Art

[0004] The human genome contains approximately three billion base pairs. Various large-scale sequencing technologies, such as next-generation sequencing (NGS), offer the opportunity to sequence at a cost of less than $1 per million base pairs, and have already achieved costs below ten cents per million base pairs. These sequencing technologies have enabled the identification of single nucleotide variants (SNVs), small insertions and deletions (indels), and large copy number variants (CNVs) in a variety of abnormal somatic tissues, such as tumor samples.

[0005] This type of analysis of multiple somatic variants in a variety of abnormal somatic tissues provides a basis for understanding the molecular perturbations that underlie the vast differences in individual disease phenotypes or treatment responses. However, the identification of these variants and the frequency of these variants may vary from subject to subject and, furthermore, may change in any given subject as the disease progresses. Furthermore, the multiple variants associated with various diseases, such as cancer, require deep sequencing of nucleotides in a biological sample, such as a tissue biopsy or blood drawn from a subject, due to the rarity of some variants. For example, detecting DNA originating from multiple tumor cells from a blood sample is difficult because circulating tumor DNA (ctDNA) is present at very low levels relative to other molecules in cfDNA extracted from the blood. However, knowledge of ctDNA levels in a subject, no matter how low, has the potential to inform treatment decisions and improve prognosis and diagnosis.

[0006] Against this backdrop, there is a need in the art for robust, efficient, and cost-effective techniques for determining ctDNA in a subject that can detect even very low levels of ctDNA. Summary of the Invention

[0007] The present invention provides multiple technical solutions (eg, computer systems, methods, and non-transitory computer-readable storage media) for solving the aforementioned identified problems by detecting ctDNA.

[0008] The following is an overview of the present invention to provide a basic understanding of some aspects of the present invention. This overview is not an extensive overview of the present invention. It is not intended to identify the multiple key elements of the present invention or to describe the scope of the present invention. Its sole purpose is to present some concepts of the present invention in a simplified form as a prelude to the more detailed description presented later.

[0009] The various embodiments of the systems, methods, and devices within the scope of the appended claims each have several aspects, no single aspect of which is solely responsible for the many desirable properties described herein. Without limiting the scope of the appended claims, some prominent features are described herein. After considering this discussion, and particularly after reading the section entitled "Detailed Description," one will understand how the various features of the various embodiments may be used.

[0010] One aspect of the present disclosure provides a method for determining a tumor score in cell-free nucleic acid from a liquid biological sample from a subject. The method includes: performing the following steps in a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors: obtaining a plurality of first sequence reads in electronic form from the liquid biological sample from the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules; using the plurality of first sequence reads to identify support for each variant in a first set of variants, thereby determining an observed frequency for each variant in the first set of variants; obtaining, for each individual variant in the first set of variants, a corresponding reference frequency for the individual variant in a first reference set; each corresponding reference frequency in the first reference set is for a different variant in a first abnormal solid tissue sample obtained from the subject; and evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue sample, thereby determining a first tumor score in the cell-free nucleic acid from the liquid biological sample from the subject.

[0011] In some embodiments, a variant in the first set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the copy number of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or any abnormal epigenetic modification (e.g., an abnormal methylation pattern) associated with a predetermined genomic location.

[0012] In some embodiments, when another sequence read among the multiple first sequence reads includes all or part of a first variant in the first variant set, the individual sequence read is considered to support the first variant; when another sequence read among the multiple first sequence reads does not include the first variant in the first variant set, the individual sequence read is considered to not support the first variant; and a number of multiple sequence reads among the multiple first sequence reads that support the first variant is used to determine the observed frequency of the first variant, and a variant frequency of the first variant in the liquid biological sample is estimated from the observed frequency of the first variant.

[0013] In some embodiments, the subject is human. In some embodiments, the subject has cancer from a single primary site. In some embodiments, the subject has cancer originating from two or more different organs.

[0014] In some embodiments, the subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0015] In some embodiments, the subject has a predetermined stage of breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, bladder cancer, or gastric cancer.

[0016] In some embodiments, the first abnormal solid tissue sample is a tumor sample.

[0017] In some embodiments, the first variant set consists of a single variant for a single genetic variation located at a single site in the genome of the subject. In alternative embodiments, the first variant set consists of a first variant for a first genetic variation located at a first site in the genome of the subject and a second variant for a second genetic variation located at a second site in the genome of the subject. In further alternative embodiments, the first variant set consists of a first variant for a first genetic variation located at a first site in the genome of the subject, a second variant for a second genetic variation located at a second site in the genome of the subject, and a third variant for a third genetic variation located at a third site in the genome of the subject.

[0018] In some embodiments, the first set of variants consists of between 2 and 20 variants, or consists of between 2 and 200 variants, or includes 1000 or more variants, or includes 5000 or more variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject.

[0019] In some embodiments, the step of using the multiple sequence reads to identify support for each variant in a variant set includes: aligning a sequence read from the multiple first sequence reads with a region in a reference genome, or with a lookup table of multiple variants, to determine whether the sequence read includes all or a portion of a first variant.

[0020] In some embodiments, the step of using the plurality of sequence reads to identify support for each variant in a set of variants comprises aligning a sequence read from the plurality of first sequence reads with each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome.

[0021] In some embodiments, the subject has stage II, III, or IV breast cancer, and the step of evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue determines that the first tumor fraction of the cell-free nucleic acid is less than 1×10 -3 .

[0022] In some embodiments, the method further comprises the steps of: using the plurality of first sequence reads to identify support for each variant in a second variant set, thereby determining an observed frequency for each variant in the second variant set; for each individual variant in the second variant set, obtaining a corresponding reference frequency for the individual variant in a second reference set, wherein each corresponding reference frequency in the second reference set is for another variant in a second abnormal solid tissue sample obtained from the subject; and evaluating the observed frequency of each individual variant in the second variant set against the observed frequency of the individual variant in the second reference set, thereby determining a second tumor fraction in the cell-free nucleic acid of the liquid biological sample from the subject. In some such embodiments, an individual sequence read is considered to support a variant in the second variant set when the individual sequence read in the plurality of first sequence reads includes all or a portion of the variant in the second variant set; and an individual sequence read is considered to not support the variant in the second variant set when the individual sequence read in the plurality of first sequence reads does not include the variant in the second variant set. In some such embodiments, the first abnormal tissue sample is composed of a first tumor fraction and the second abnormal tissue sample is composed of a second tumor fraction from the same tumor of the subject. In some such embodiments, the first abnormal tissue sample is a first cancer type and the second abnormal tissue sample is a second cancer type. In some such embodiments, the first cancer type is the same as the second cancer type. In various alternative embodiments, the first cancer type is different from the second cancer type. In some such embodiments, the first cancer type and the second cancer type are each selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer.

[0023] In some embodiments, the frequency of each variant in the first reference set is obtained by a plurality of second sequence reads collectively obtained from the first abnormal entity tissue sample. In some such embodiments, more than 1000 sequence reads, or more than 3000 sequence reads, or more than 5000 sequence reads are collectively obtained from the first abnormal entity tissue sample. In some such embodiments, the method further comprises: analyzing the plurality of second sequence reads obtained from the first abnormal entity tissue sample against a panel of variant candidates. In some such embodiments, the panel of variant candidates comprises between 100 and 1000 variants.

[0024] In some embodiments, the plurality of second sequence reads obtained from the first abnormal solid tissue sample represent whole genome data for each cell. In some embodiments, the plurality of second sequence reads obtained from the first abnormal solid tissue sample have an average coverage of at least 10-fold, at least 100-fold, or at least 2000-fold.

[0025] In some embodiments, the liquid biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0026] In some embodiments, the liquid biological sample is composed of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0027] In some embodiments, the step of evaluating the observed frequency of each individual variant in the first variant set against a corresponding reference frequency of the individual variant in the first reference set comprises evaluating a cumulative density function or a cumulative distribution function for the individual variant using the observed frequency and the reference frequency for the individual variant over a range of possible tumor fractions. In some embodiments, a cumulative density function is used, and the range is from 0% to 110%. In some such embodiments, the first tumor fraction is taken as a median of the cumulative density function.

[0028] In some embodiments, a cumulative distribution function is used.

[0029] In some embodiments, the cumulative distribution function has the following form:

[0030]

[0031] where x = a 2i , the observed number of the plurality of sequence reads supporting the individual variant in the liquid biological sample; p = t*f 1i , where t is the estimated first tumor fraction, and f 1i is the observed frequency of the individual variant in the first set of variants; and n=d 2i , a total number of sequence reads from the biological sample that map to the genomic location corresponding to the individual variant.

[0032] In some embodiments, the cumulative distribution function has the following form:

[0033]

[0034] where x = a 2i , the observed number of the plurality of sequence reads supporting the individual variant k in the liquid biological sample; p k =t*f 1i , where t is the estimated first tumor fraction, and f 1i is the observed frequency of the individual variant k in the first set of variants; and n k =d 2i , a total number of sequence reads from the biological sample that map to the genomic position corresponding to the individual variant k.

[0035] In some embodiments, the cumulative density function or the cumulative distribution function is derived under the assumption of negative binomial distribution.

[0036] In some embodiments, the method further comprises: repeating the step of obtaining the plurality of first sequence reads at each of a plurality of time points over a period of time from a separate biological sample obtained from the subject at each separate time point, wherein the separate biological sample comprises a plurality of cell-free nucleic acid molecules, thereby obtaining a corresponding plurality of first sequence reads for the subject at each separate time point. Furthermore, in many such embodiments, for each of the plurality of time points, determining support for each variant in the first set of variants in the corresponding plurality of first sequence reads for the subject at each separate time point, thereby determining an observed frequency for each variant in the first set of variants between the plurality of sequence reads in the corresponding plurality of first sequence reads that support or do not support the variant at each of the plurality of time points. The observed frequency of each variant in the first set of variants at each of the plurality of time points is evaluated against the observed frequency of the variant in the first abnormal solid tissue, thereby determining the status or progression of a condition in the subject during the period of time as an increase or decrease in the first tumor score over the period of time. In some such embodiments, the period is a period of several months (e.g., less than 4 months, between 1 month and 4 months, etc.), and each of the plurality of time points is a different time point in the period of several months. In some such embodiments, the period is a period of several years (between 2 and 10 years), and each of the plurality of time points is a different time point in the period of several years. In some such embodiments, the period is a period of several hours (e.g., between 1 hour and 6 hours), and each of the plurality of time points is a different time point in the period of several hours.

[0037] In some embodiments, the method further comprises changing a diagnosis of the subject upon observing that the first tumor score of the subject changes by a threshold amount within the time period (e.g., a change of 10 percent, 20 percent, 30 percent relative to a reference amount at a first measurement time point).

[0038] In some embodiments, the method further comprises: changing a prognostic profile of the subject upon observing that the first tumor score of the subject changes by a threshold amount within the time period (e.g., a change of 10 percent, 20 percent, or 30 percent relative to a reference amount at a first measurement time point).

[0039] In some embodiments, the method further comprises changing a treatment of the subject when the first tumor score of the subject is observed to change by a threshold amount during the time period (e.g., a change of 10 percent, 20 percent, 30 percent relative to a reference amount at a first measurement time point).

[0040] In some embodiments, the condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some embodiments, the condition is a stage of a cancer (e.g., a stage of breast cancer, a stage of lung cancer, a stage of prostate cancer, a stage of colorectal cancer, a stage of kidney cancer, a stage of uterine cancer, a stage of pancreatic cancer, a stage of esophageal cancer, a stage of lymphoma, a stage of head and neck cancer, a stage of ovarian cancer, a stage of hepatobiliary cancer, a stage of melanoma, a stage of cervical cancer, a stage of multiple myeloma, a stage of leukemia, a stage of thyroid cancer, a stage of bladder cancer, or a stage of gastric cancer).

[0041] In some embodiments, the condition is a predetermined subtype of cancer.

[0042] In some embodiments, the method further comprises the steps of applying the plurality of first sequence reads to a trained classifier to obtain a classifier result, wherein the result of the trained classifier indicates whether the subject has a first cancer condition; and when the first tumor score is between 0.003 and 1.0 and the result of the trained classifier indicates that the subject has the first cancer condition, using the result of the trained classifier as a basis for a diagnosis or prognosis of the subject for the first cancer condition. In some embodiments, the first cancer condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some embodiments, the first cancer condition is a subtype of cancer (e.g., the cancer is a subtype of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer).

[0043] In some such embodiments, the first tumor score is between 0.003 and 1.0, and the first cancer condition is a tissue of origin of a cancer.

[0044] In some embodiments, the trained classifier is a neural network, a support vector machine, a decision tree, an unsupervised clustering model, a supervised clustering model, or a regression model.

[0045] Another aspect of the present disclosure provides a computer system comprising: one or more processors; a memory storing one or more programs for execution by the one or more processors; the one or more programs comprising a plurality of instructions for determining a tumor fraction in a cell-free nucleic acid sample from a subject using a method comprising: obtaining a plurality of first sequence reads in electronic form from the liquid biological sample from the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules. The method further comprises: using the plurality of first sequence reads to identify support for each variant in a first set of variants, thereby determining an observed frequency for each variant in the first set of variants. The method further comprises: for each individual variant in the first set of variants, obtaining a corresponding reference frequency for the individual variant in a first reference set, wherein each corresponding reference frequency in the first reference set is for a different variant in a first abnormal solid tissue sample obtained from the subject. The method further comprises evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue, thereby determining a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

[0046] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs for determining a tumor fraction in a cell-free nucleic acid of a liquid biological sample from a subject. The one or more programs are configured to be executed by a computer. The one or more programs include multiple instructions for obtaining a plurality of first sequence reads in electronic form from the liquid biological sample from the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules. The one or more programs further include multiple instructions for using the multiple first sequence reads to identify support for each variant in a first set of variants, thereby determining an observed frequency for each variant in the first set of variants. The one or more programs further include multiple instructions for obtaining, for each individual variant in the first set of variants, a corresponding reference frequency for the individual variant in a first reference set, wherein each corresponding reference frequency in the first reference set is for a different variant in a first abnormal solid tissue sample obtained from the subject. The one or more programs further comprise instructions for evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue to thereby determine a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

[0047] Another aspect of the present disclosure provides a method for determining a tumor score in cell-free nucleic acid from a liquid biological sample from a subject. The method comprises: obtaining, in a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors, a plurality of sequence reads in electronic form from the liquid biological sample from the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules. The method further comprises: using the plurality of sequence reads to identify support for each variant in a set of variants, thereby determining an observed frequency for each variant in the first set of variants. The method further comprises: determining the observed frequency of the variant having the Nth highest allele frequency in the set of variants as the tumor score in the cell-free nucleic acid from the liquid biological sample from the subject, where N is a positive integer other than 1 (e.g., 1, 2, 3, 4, 5, etc.).

[0048] In some embodiments, a variant in the set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the number of copies of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an abnormal epigenetic modification pattern (e.g., methylation pattern) associated with a predetermined genomic location.

[0049] In some embodiments, when another one of the multiple sequence reads includes all or a portion of a first variant in the variant set, the individual sequence read is considered to support the first variant; when another one of the multiple sequence reads does not include the first variant in the variant set, the individual sequence read is considered to not support the first variant; and a number of multiple sequence reads in the multiple sequence reads that support the first variant is used to determine the observed frequency of the first variant, and a variant frequency of the first variant in the liquid biological sample is estimated from the observed frequency of the first variant.

[0050] In some embodiments, the subject has cancer from a single primary site. In some embodiments, the subject has cancer originating from two or more different organs. In some embodiments, the subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0051] In some embodiments, the variant set comprises five or more variants, wherein each individual variant in the variant set is located at a different site in the genome of the subject. In some embodiments, the variant set consists of between 3 and 20 variants, wherein each variant in the variant set is a different genetic variation in the genome of the subject.

[0052] In some embodiments, the variant set is comprised of between 2 and 200 variants, wherein each variant in the variant set is a different genetic variation in the genome of the subject. In some embodiments, the variant set comprises 1000 variants, wherein each variant in the variant set is a different genetic variation in the genome of the subject.

[0053] In some embodiments, the step of using the multiple sequence reads to identify support for each variant in a variant set includes: aligning a sequence read from the multiple sequence reads with a region in a reference genome to determine whether the sequence read includes all or a portion of a first variant.

[0054] In some embodiments, the step of using the plurality of sequence reads to identify support for each variant in a set of variants includes comparing a sequence read from the plurality of sequence reads to a lookup table of multiple variants to determine whether the sequence read includes all or a portion of a first variant.

[0055] In some embodiments, the step of using the plurality of sequence reads to identify support for each variant in a set of variants comprises aligning a sequence read from the plurality of first sequence reads with each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome.

[0056] In some embodiments, the liquid biological sample comprises or consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0057] In some embodiments, the method further comprises the steps of: repeating the step of obtaining a plurality of sequence reads at each respective time point in a plurality of time points over a period of time from a separate biological sample taken from the subject at each respective time point, wherein the respective biological sample comprises a plurality of cell-free nucleic acid molecules, thereby obtaining a corresponding plurality of sequence reads for the subject at each respective time point; and determining support for the variant having the Nth highest allele frequency in the set of variants in the original calling step for each respective time point, thereby determining the status or progression of a condition in the subject during the period of time in the form of an increase or decrease in the allele frequency of the variant over the period of time.

[0058] In some embodiments, the period is a period of several months (e.g., between 1 month and 4 months), and each of the plurality of time points is a different time point in the period of several months. In some embodiments, the period is a period of several years (e.g., between 2 and 10 years), and each of the plurality of time points is a different time point in the period of several years. In some embodiments, the period is a period of several hours (e.g., between 1 hour and 6 hours), and each of the plurality of time points is a different time point in the period of several hours.

[0059] In some embodiments, the method further comprises changing a diagnosis of the subject when the allele frequency of the variant is observed to change by a threshold amount within the time period (e.g., a change of 10 percent, 20 percent, 30 percent relative to a reference amount at a first measurement time point).

[0060] In some embodiments, the method further comprises changing a prognosis of the subject when the allele frequency of the variant is observed to change by a threshold amount during the period (e.g., a change of 10 percent, 20 percent, 30 percent relative to a reference amount at a first measurement time point).

[0061] In some embodiments, the method further comprises changing a treatment of the subject when the allele frequency of the variant is observed to change by a threshold amount within the time period (e.g., a change of 10 percent, 20 percent, 30 percent relative to a reference amount at a first measurement time point).

[0062] In some embodiments, the condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some embodiments, the condition is a stage of a cancer (e.g., a stage of breast cancer, a stage of lung cancer, a stage of prostate cancer, a stage of colorectal cancer, a stage of kidney cancer, a stage of uterine cancer, a stage of pancreatic cancer, a stage of esophageal cancer, a stage of lymphoma, a stage of head and neck cancer, a stage of ovarian cancer, a stage of hepatobiliary cancer, a stage of melanoma, a stage of cervical cancer, a stage of multiple myeloma, a stage of leukemia, a stage of thyroid cancer, a stage of bladder cancer, or a stage of gastric cancer). In some embodiments, the condition is a predetermined subtype of a cancer.

[0063] In some embodiments, the method further comprises the steps of applying the plurality of sequence reads to a trained classifier to obtain a classifier result, wherein the result of the trained classifier indicates whether the subject has a first cancer condition; and when the tumor score is between 0.003 and 1.0 and the result of the trained classifier indicates that the subject has the first cancer condition, using the result of the trained classifier as a basis for diagnosing the subject with the first cancer condition. In some such embodiments, the first cancer condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some such embodiments, the first cancer condition is a subtype of a cancer (e.g., a subtype of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer). In some such embodiments, the first tumor score is between 0.003 and 1.0, and the first cancer condition is a tissue of origin of a cancer. In some embodiments, the trained classifier is a neural network, a support vector machine, a decision tree, an unsupervised clustering model, a supervised clustering model, or a regression model.

[0064] Another aspect of the present disclosure provides a computer system comprising: one or more processors; and a memory storing one or more programs executed by the one or more processors. The one or more programs include a plurality of instructions for determining a tumor score in a cell-free nucleic acid of a liquid biological sample of a subject by a method comprising: obtaining a plurality of sequence reads in electronic form from the liquid biological sample of the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules. The method further comprises: using the plurality of sequence reads to identify support for each variant in a set of variants, thereby determining an observed frequency for each variant in the first set of variants. The method further comprises: considering the observed frequency of the variant having the Nth highest allele frequency in the set of variants as the tumor score in the cell-free nucleic acid of the liquid biological sample of the subject, where N is a positive integer other than 1.

[0065] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs for determining a tumor score in a cell-free nucleic acid of a liquid biological sample of a subject. The one or more programs are configured to be executed by a computer. The one or more programs include multiple instructions for obtaining a plurality of sequence reads in electronic form from the liquid biological sample of the subject, wherein the liquid biological sample includes a plurality of cell-free nucleic acid molecules. The one or more programs further include multiple instructions for using the multiple sequence reads to identify support for each variant in a set of variants, thereby determining an observed frequency for each variant in the first set of variants. The one or more programs further include multiple instructions for considering the observed frequency of the variant having the Nth highest allele frequency in the set of variants as the tumor score in the cell-free nucleic acid of the liquid biological sample of the subject, where N is a positive integer other than 1.

[0066] Incorporation by Reference

[0067] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The various embodiments disclosed herein are illustrated by way of example and not limitation in the figures of the accompanying drawings. Like reference numerals refer to corresponding parts throughout the several views of the drawings.

[0069] Figure 1A and 1B An example block diagram depicting a computing device is illustrated according to some embodiments of the present disclosure;

[0070] Figure 2A 、 2B 2C, 2D, 2E and 2F illustrate an example flow chart of a method for subject-based classification according to some embodiments of the present disclosure;

[0071] Figure 3According to some embodiments of the present disclosure, a box plot is described, wherein, for each individual cancer type, the ctDNA score for a plurality of subjects having the individual cancer type is provided, wherein for each individual subject, the y-axis provides an estimated ctDNA score based on a matched pair comparison of an observed frequency of each variant in a set of variants in a biological sample (e.g., blood) from the individual subject and a corresponding reference frequency of each such variant obtained from an abnormal tissue sample (e.g., tumor fraction) from the individual subject;

[0072] Figure 4 According to some embodiments of the present disclosure, the incidence of cancer as a function of cancer stage is described. Figure 3 a graphical representation of ctDNA fractions for a plurality of subjects having any of a plurality of cancers described;

[0073] Figure 5 According to some embodiments of the present disclosure, a graphical representation of ctDNA scores for multiple subjects as a function of breast cancer stage is illustrated, divided into three categories: subjects whose cell-free DNA is sufficient to identify a variant found in a paired tumor of such subject without prior knowledge of the variant in the paired tumor; subjects whose cell-free DNA supports a variant found in a paired tumor; and subjects whose cell-free DNA does not support a variant found in a paired tumor cancer.

[0074] Figure 6 According to some embodiments of the present disclosure, the ability to detect cancer in a plurality of subjects as a function of their cfDNA fraction is described;

[0075] Figure 7A and 7B According to some embodiments of the present disclosure, the ability to identify breast cancer as a function of cfDNA score, classifier, and breast cancer subtype is described;

[0076] Figure 8 According to some embodiments of the present disclosure, a method for determining the presence of a cfDNA fraction as a function of the cfDNA fraction is described in detail. Figure 3 The accuracy of the WGBS multi-class classifier for a subject population with a spectrum of different cancers identified in

[15] ;

[0077] Figure 9 According to some embodiments of the present disclosure, detailing a percentage of subjects exhibiting a minimum ctDNA score as a function of clinical stage;

[0078] Figure 10According to some embodiments of the present disclosure, a positive correlation between tumor size and ctDNA score across all stages of cancer is demonstrated;

[0079] Figure 11 According to some embodiments of the present disclosure, the correlation of ctDNA score with the Ki67 marker for proliferation is described;

[0080] Figure 12 A flow chart illustrating a method for preparing a nucleic acid sample for sequencing according to some embodiments of the present disclosure;

[0081] Figure 13 is a diagrammatic representation of a process for obtaining a plurality of sequence reads according to some embodiments of the present disclosure;

[0082] Figure 14 is a flow chart of a method for determining variants of a plurality of sequence reads according to some embodiments of the present disclosure;

[0083] Figure 15 A flowchart of a method for obtaining a methylation state vector for identifying multiple variants according to some embodiments of the present disclosure;

[0084] Figure 16 According to some embodiments of the present disclosure, a cumulative density function is provided over a range of estimated shedding rates for a trial;

[0085] Figure 17 Describing the concordance of multiple tumor score measurements between a tumor pair embodiment and a second highest allele embodiment of the present disclosure;

[0086] Figure 18 Describing details of a CCGA study as a basis for determining a cell-free tumor fraction according to some embodiments of the present disclosure;

[0087] Figure 19A and 19B According to one embodiment of the present disclosure, Figure 18 The training set ( Figure 19A , N = 1,416) and a test set ( Figure 19B , N = 847), provided the use Figure 18 The summarized sensitivity information of the plurality of patterns trained using the training set;

[0088] Figure 19C and 19D According to one embodiment of the present disclosure, a method for classifying tumors by their origin is provided. Figure 18 The training set ( Figure 19C ) and the test set ( Figure 19D ) in the tumor fraction; and

[0089] Figure 20A and 20B According to one embodiment of the present disclosure, the cfDNA tumor score is calculated by comparing the results of cfDNA WGS and tumor WGS, wherein the results of the cfDNA WGS and tumor WGS are based on the stage of a total set of breast cancer, colorectal cancer, lung cancer and other cancers ( Figure 20A ), and by cancer type ( Figure 20B ). DETAILED DESCRIPTION

[0090] Several embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks have not been described in detail to avoid unnecessarily obscuring aspects of the various embodiments.

[0091] The various embodiments described herein provide various technical solutions for determining a tumor score in a subject. Such information can be used to determine the cancer status of the subject, including, for example, classifying the subject's tissue of origin. A plurality of sequence reads are obtained from a biological sample of a subject. The biological sample includes cell-free nucleic acid. Therefore, the plurality of sequence reads are for cell-free nucleic acid. The plurality of sequence reads are used to identify support for each variant in a variant set, thereby determining an observed frequency for each variant. The plurality of observed variant frequencies are compared with corresponding reference frequencies for each variant in a reference set. Each such reference frequency is a frequency of another variant in an abnormal tissue sample (e.g., a tumor) from the subject. In this way, the tumor score of the subject is determined. In some embodiments, the tumor score is used in combination with a classifier to classify a cancer condition of the subject.

[0092] Figure 3Provide a basis for multiple embodiments of the present disclosure. Typically, the observed frequencies of the multiple variants in the variant set obtained from the cell-free nucleic acid of the biological sample are less than the observed reference frequencies for such multiple variants in the reference set. Without intending to be limited to any particular theory, it is assumed that the source of the cell-free nucleic acid containing such multiple variants is the degradation or decomposition of multiple cancer cells in the abnormal tissue. Therefore, in some embodiments, it is assumed that the cell-free nucleic acids in the multiple biological samples containing such multiple variants in the multiple disclosed variant sets of the present disclosure are represented as ctDNA, or "circulating tumor DNA" (ctDNA), a small portion of the cell-free nucleic acid (cfDNA), which are used as the basis for determining the observed frequency of each variant. Therefore, it is expected that the observed frequencies of the multiple variants in the variant set obtained from the cell-free nucleic acid of the biological sample are less than the observed reference frequencies for such multiple variants in the reference set. Figure 3 The data summarized in support this contention and indicate that different cancer types have different ratios of the observed frequencies of the multiple variants in the variant set for multiple specific subjects to the observed reference frequencies for such multiple variants in a reference abnormal tissue of the same multiple specific subjects.

[0093] Figure 3 A box plot is provided in which, for each cancer type studied in a CCGA cohort, there are multiple individuals and an estimate of the ctDNA fraction for each individual is on the y-axis. Figure 3 An overview of the distribution of ctDNA fractions observed for each individual cancer type is shown for two categories of subjects for each cancer type: (i) subjects with the individual cancer type in whom there is no measured evidence of a variant (in the plurality of sequence reads from the plurality of cell-free biological samples) in their cfDNA; and (ii) subjects with the individual cancer type who have no measured evidence of a variant (in the plurality of sequence reads from the plurality of cell-free biological samples). Figure 3 and (ii) the subjects have the respective cancer type in which there is measured evidence of a variant (in the plurality of sequence reads from the plurality of cell-free biological samples) in their cfDNA ( Figure 3 For each particular cancer type, a first distribution of the measured ctDNAs of the plurality of subjects in the true class forms a first box ( Figure 3 A second distribution of the expected ctDNA of the multiple subjects in the error category forms a second box ( Figure 3), where the 25th and 75th quartiles define each box, and the whiskers of each box indicate the extreme values. The black line in each box is the median tumor score estimate for all individuals in a particular class of a particular cancer type. For example, referring to kidney cancer, there is a median ctDNA score for those subjects in the false class, while there is a different median ctDNA score for those subjects in the true class.

[0094] Figure 3 It is shown that in the CCGA population studied, there is a large dynamic range in the shedding rates (ctDNA scores) of different cancers. Example 12 below provides multiple details of the CCGA population. The observed large dynamic range can be used to provide a basis for establishing multiple meaningful and useful thresholds from the observed frequencies of the multiple variants in the reference set. That is, for example, given the observed frequencies of multiple variants in the abnormal tissue of a particular subject, and optionally information about the expected ctDNA scores of multiple subjects with a particular condition, a threshold for the particular cancer subject is determined and evaluated for the observed frequencies of the multiple variants in a variant set for the particular subject in order to classify the subject as having or not having the condition. For example, reference Figure 3A threshold of 0.01 can be used to analyze whether a subject has kidney cancer. In this example, an abnormal tissue, such as a tumor, is obtained from a patient and used to determine a reference frequency for each individual variant in a first reference set. In fact, in some embodiments, the frequencies of each possible variant are used to define the plurality of variants in the reference set. Next, cell-free nucleic acid is obtained from a biological sample other than the abnormal tissue, and the variant frequencies of multiple identical variants in the reference set are determined from multiple sequence reads of the cell-free nucleic acid in the biological sample, thereby forming the observed ctDNA frequency for each individual variant in the first variant set. Next, a comparison of the ctDNA frequencies used to determine whether the threshold condition of 0.01 is met with the reference frequencies provides a basis for determining whether the subject has kidney cancer. For example, if the comparison indicates a ctDNA score above 0.01, then the subject does not have kidney cancer. On the other hand, an observation of a ctDNA score formed from the observed frequencies of each individual variant in the first variant set of approximately 1e-03 is consistent with the finding of kidney cancer. Furthermore, in some embodiments, rather than using the systems and methods of the present disclosure to indicate whether a subject has a particular condition on an absolute binary basis, a likelihood or probability that a subject has a particular condition is provided. In such embodiments, the comparison of the observed frequency of each individual variant in the first set of variants with a corresponding reference frequency for the individual variant in a first reference set is used to determine how far the observed frequency of each individual variant in the first set of variants is from the corresponding reference frequency for the individual variant, and the probability or likelihood that a subject has a particular condition is determined based on this distance or a function of this distance.

[0095] exist Figure 3 In , the method used to calculate the ctDNA score is a Bayesian approach. For example, consider the example of a tumor sequencing set (reference set) for multiple variants of a different cancer type and the paired cell-free DNA for a collection of multiple subjects with that cancer type. If none of the multiple tumor variants for the individual cancer type can be paired with the cfDNA of any of the subjects in the collection of subjects, then in the absence of any supporting sequencing data, the collection of subjects can still be used to estimate what the ctDNA score is for a different cancer type, thereby providing an upper limit on how much signal is available, even if it is missed in the cell-free nucleic acid assay. For Figure 3According to the calculation, on average, there are 3000 sequence reads obtained from the cell-free DNA of a biological sample from each subject. If none of the plurality of sequence reads supports (includes) the variant in the 3000 sequence reads, and the subject belongs to Figure 3 If the variant is not included in the error category, then this information can still be used to estimate the possible potential frequency (posterior probability) of the variant in the biological sample, even if none of the multiple sequence reads support (include) the variant. Such an analysis forms the basis for Figure 3 The plurality of subjects in the error categories for each cancer shown provide the basis for the estimated ctDNA scores. That is, Figure 3 The plurality of grey boxes in the figure represent the plurality of biological samples, wherein no single variant associated with the cancer has been measured in the plurality of individual biological samples. These biological samples are used to independently estimate the possible potential ctDNA fraction for the cancer. Figure 3 It is demonstrated that such multiple biological samples, with respect to the median, produce a reduced median ctDNA fraction relative to the median of the ctDNA fraction calculated using a population of samples of the same cancer, wherein multiple variants associated with the cancer have been observed in the multiple biological samples.

[0096] definition

[0097] As disclosed herein, the term "biological sample" refers to any sample obtained from a subject that reflects a biological condition associated with the subject, including cell-free DNA. Examples of biological samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0098] As disclosed herein, the terms "cell-free nucleic acid," "cell-free DNA," and "cfDNA" interchangeably refer to nucleic acid fragments that circulate within a subject's body (e.g., bloodstream) and originate from one or more healthy cells and / or from one or more cancer cells.

[0099] As disclosed herein, the term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments originating from abnormal tissue, such as cells of a tumor or other cancer type, that may be released into a subject's bloodstream due to biological processes, such as apoptosis or necrosis of dying cells, or active release by living tumor cells.

[0100] As disclosed herein, the term "cell-free nucleic acid" refers to nucleic acid molecules that can be found in body fluids outside of cells, such as blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of a subject. Cell-free nucleic acid is interchangeably referred to as circulating nucleic acid. Examples of such cell-free nucleic acid include, but are not limited to, RNA, mitochondrial DNA, or genomic DNA.

[0101] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted into a methyl group to form 5-methylcytosine. In particular, methylation tends to occur at dinucleotides of cytosine and guanine, referred to herein as "CpG sites". In other examples, methylation may occur on a portion of a non-CpG site of a cytosine, or on other nucleotides other than cytosine; however, these situations rarely occur. In this disclosure, for clarity, methylation is discussed with reference to CpG sites. Abnormal cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate the state of cancer. As is well known in the art, abnormal DNA methylation (compared to healthy controls) can cause different effects, which may lead to cancer.

[0102] As used herein, the term "methylation index" may refer to the ratio of a plurality of sequence reads showing methylation at the site to the total number of a plurality of reads covering the site for each genomic site (e.g., a CpG site, which is a region of DNA in which a cytosine nucleotide in a linear sequence of bases in the 5'→3' direction is followed by a guanine nucleotide). The "methylation density" of a region may be the number of a plurality of reads at a plurality of sites within a region showing methylation divided by the total number of a plurality of reads covering the plurality of sites in the region. The plurality of sites may have a plurality of specific characteristics (e.g., the plurality of sites may be CpG sites). The "CpG methylation density" of a region may be the number of a plurality of reads showing CpG methylation divided by the total number of a plurality of reads covering a plurality of CpG sites in the region (e.g., a specific CpG site, a plurality of CpG sites within a CpG island, or a larger region). For example, the methylation density of each 100-kb data box (bin) in the human genome can be determined as the ratio of all CpG sites covered by multiple sequence reads mapped to the 100-kb region from the total number of multiple unconverted cytosines (which may correspond to methylated cytosines) at multiple CpG sites. In some embodiments, this analysis is performed on the size of other data boxes, such as 50-kb or 1-Mb. In some embodiments, a region can be an entire genome or a chromosome or a portion of a chromosome (e.g., a chromosome arm). When a region includes only one CpG site, a methylation index of the CpG site can be the same as the methylation density of the region. The "ratio of methylated cytosines" can refer to the number of multiple cytosine sites "C's" that show methylation (e.g., unconverted after bisulfite conversion) in the region divided by the total number of multiple analyzed cytosine residues, for example, including multiple cytosines outside the CpG background. The methylation index, the methylation density, and the ratio of methylated cytosine are examples of “methylation level”.

[0103] As disclosed herein, the terms "nucleic acid" and "nucleic acid molecule" are used interchangeably. These terms refer to any compositional form of nucleic acid, for example, deoxyribonucleic acid (DNA, e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), and / or DNA analogs (e.g., containing base analogs, sugar analogs, and / or a non-natural backbone, etc.), all of which can be in single-stranded or double-stranded form. Unless otherwise limited, a nucleic acid can include known analogs of natural nucleotides, some of which can function in a manner similar to naturally occurring nucleotides. A nucleic acid can be in any form (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.) that can be used to perform the various processes described herein. In some embodiments, a nucleic acid can be derived from a single chromosome or a fragment of a single chromosome (e.g., a nucleic acid sample can be derived from a chromosome from a sample obtained from a diploid organism). In certain embodiments, a nucleic acid includes multiple nucleosomes, multiple fragments or portions of multiple nucleosomes, or multiple nucleosome-like structures. Nucleic acids sometimes include proteins (e.g., histones, DNA-binding proteins, etc.). Nucleic acids analyzed by the various processes described herein are sometimes substantially isolated and substantially unassociated with proteins or other molecules. Nucleic acids also include derivatives, variants, and analogs of DNA synthesized, replicated, or amplified from single-stranded ("sense" or "antisense," "positive" or "negative," "positive reading frame" or "reverse reading frame") and double-stranded polynucleotides. Deoxyribonucleic acids include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. Nucleic acids can be prepared using nucleic acid obtained from a subject as a template.

[0104] As disclosed herein, the term "reference genome" refers to any specific known, sequenced, or characterized genome of any organism or virus that can be used to reference multiple identification sequences from a subject, whether partial or complete. Multiple exemplary reference genomes for human subjects and many other organisms are provided in the online genome browser hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). A "genome" refers to the complete genetic information of an organism or virus expressed as multiple nucleic acid sequences. As used herein, a reference sequence or reference genome is typically an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. The reference genome can be considered as a representative example of a species' gene set. In some embodiments, a reference genome includes multiple sequences assigned to multiple chromosomes. Several exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0105] As disclosed herein, the terms "regions of a reference genome," "genomic regions," or "genomic regions" refer to any portion of a reference genome, whether continuous or discontinuous. It may also be referred to as, for example, a bin, a partition, a genome portion, a portion of a reference genome, a portion of a chromosome, etc. In some embodiments, a genomic segment is based on a specific length of genomic sequence. In some embodiments, a method may include analysis of multiple mapped sequences of multiple genomic regions. Multiple genomic regions may be approximately the same length, or the multiple genomic segments may be of different lengths. In some embodiments, the lengths of multiple genomic regions are approximately the same. In some embodiments, multiple genomic regions of different lengths are adjusted or weighted. In some embodiments, a genomic region is about 10 kilobases (kb) to about 500kb, about 20kb to about 400kb, about 30kb to about 300kb, about 40kb to about 200kb, and sometimes about 50kb to about 100kb. In some embodiments, a genomic region is about 100kb to about 200kb. A genomic region is not limited to a continuous extended sequence. Thus, multiple genomic regions can be composed of multiple continuous and / or discontinuous sequences. A genomic region is not limited to a single chromosome. In some embodiments, a genomic region includes all or part of a chromosome, or all or part of two or more chromosomes. In some embodiments, multiple genomic regions can span one, two, or more complete chromosomes. Furthermore, the multiple genomic regions can span multiple joined or unjoined portions of multiple chromosomes.

[0106] As disclosed herein, the term "sequence read" or "read" refers to a plurality of nucleotide sequences generated by any sequencing process described herein or known in the art. Multiple reads can be generated from one end of multiple nucleic acid fragments ("single-end reads"), and sometimes from both ends of a nucleic acid (e.g., double-end reads, double-end reads). The length of the sequence reads is often related to the specific sequencing technology. For example, a number of high-throughput methods provide a plurality of sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, an average, median, or mean length of the plurality of sequence reads is about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, about 200 bp, about 210 bp, about 220 bp, about 230 bp, about 240 bp, about 250 bp, about 260 bp, about 270 bp, about 280 bp, about 290 bp, about 300 bp, about 310 bp, about 320 bp, about 330 bp, about 340 bp, about 350 bp, about 360 bp, about 370 bp, about 380 bp, about 390 bp, about 400 bp, about 410 bp, about 420 bp, about 430 bp, about 440 bp, about 450 bp, about 460 bp, about 470 bp, about 480 bp, about 490 bp, about 500 bp, about About 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp. In some embodiments, an average, median, or mean length of the plurality of sequence reads is about 1000 bp or longer. For example, nanopore sequencing provides a plurality of sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina's parallel sequencing method can provide a plurality of sequence reads that do not vary much, for example, a majority of sequence reads can be less than 200 bp.

[0107] As disclosed herein, the terms "sequencing" or "sequence determination" and the like are used herein to generally refer to any or all biochemical processes that can be used to determine the sequence of a biological macromolecule, such as a nucleic acid or protein. For example, sequence data can include all or a portion of the nucleotide bases of a nucleic acid molecule, such as a DNA fragment.

[0108] As disclosed herein, the term "single nucleotide variant" or "SNV" refers to a substitution of a nucleotide for a different nucleotide at a position (e.g., site) in a nucleotide sequence, such as a sequence read from an individual. A substitution from a first nucleobase X to a second nucleobase Y can be represented as "X>Y." For example, a SNV from cytosine to thymine is represented as "C>T."

[0109] As disclosed herein, the term "subject" refers to any living or non-living organism, including but not limited to humans (e.g., male, female, fetus, pregnant female, child, etc.), non-human animals, plants, bacteria, fungi, or protozoa. Any human or non-human animal can be a subject, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cattle), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), porcines (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursids (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, a subject is male or female (e.g., man, woman, or child) at any stage of life.

[0110] Exemplary System Embodiments

[0111] Having now provided an overview of some aspects of the present disclosure and some definitions used herein, we will now describe in more detail an exemplary system with reference to FIG1 . FIG1 is a block diagram illustrating a system 100 according to some embodiments. In some embodiments, a device 100 includes: one or more processing units (CPUs) 102 (also referred to as processors); one or more network interfaces 104; a user interface 106; a non-persistent memory 111; a persistent memory 112; and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 may optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communications between the various system components. The non-persistent memory 111 typically includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, ROM, EEPROM, or flash memory, while the persistent memory 112 typically includes a CD-ROM, digital versatile disc (DVD) or other optical storage, a magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU(s) 102. The persistent memory 112 and the non-volatile memory device(s) in the non-persistent memory 112 include non-transitory computer-readable storage media. In some embodiments, the non-persistent memory 111 or alternatively the non-transitory computer-readable storage media stores the following programs, modules, and data structures, or a subset thereof, sometimes in conjunction with the persistent memory 112:

[0112] An optional operating system 116, including multiple processes for handling various basic system services and for performing various hardware-related tasks;

[0113] An optional network communication module (or instructions) 118 for connecting the system 100 to other devices or a communication network;

[0114] a disease condition monitoring module 120 for classifying a subject, and / or assessing a status of a disease condition of a subject, and / or determining or monitoring a ctDNA tumor score of a subject;

[0115] One or more data constructs 122 for a dataset of one or more abnormal tissue samples from a subject, each such data construct 122 comprising a plurality of second sequence reads 126;

[0116] One or more reference sets 128 , each respective reference set 128 being specific to a corresponding data construct 122 of an abnormal tissue sample and including an identification of each variant 130 in a set of variants and a reference frequency 132 for each such variant;

[0117] a biological sample sequence library 134 comprising a separate data construct 138 for each corresponding biological sample from the subject, the corresponding biological sample comprising a plurality of cell-free nucleic acid molecules, the separate data construct 138 comprising a plurality of first sequence reads 140 for such plurality of cell-free nucleic acid molecules; and

[0118] A variant set database 136 includes a variant set 142 for each corresponding biological sample, each such variant set 142 including a set of variants 144, each variant including an indication of support for the first variant in the corresponding biological sample.

[0119] In some embodiments, one or more of the elements identified above are stored in one or more of the previously mentioned storage devices and in a set of instructions for performing the above-mentioned functions. A plurality of the modules, data or programs (e.g., sets of instructions) identified above do not need to be implemented as a plurality of individual software programs, processes, data sets or modules, and therefore, various subsets of these modules or data may be combined or otherwise rearranged in various embodiments. In some embodiments, the non-persistent memory 111 optionally stores a subset of the plurality of modules and data structures identified above. Furthermore, in some embodiments, the memory stores a plurality of additional modules and data structures not described above. In some embodiments, one or more of the elements identified above are stored in a computer system other than the visualization system 100, and the computer system is addressable by the visualization system 100 so that the visualization system 100 can retrieve all or part of such data when needed.

[0120] While FIG1 depicts a "system 100," this illustration is intended more as a functional description of various features that may be present in various computer systems than as a schematic diagram of the various embodiments described herein. In practice, and as one skilled in the art will appreciate, several items shown separately may be combined, and some items may be separated. Furthermore, while FIG1 depicts certain data and modules in the non-persistent memory 111, some or all of these data and modules may be located in the persistent memory 112.

[0121] Exemplary Method Example - Discovering a Condition Based on Sequencing of Abnormal Tissue

[0122] While a system according to the present disclosure has been disclosed with reference to FIG. 1 , a method according to the present disclosure will now be described in detail with reference to FIG. 2 .

[0123] refer to Figure 2A In blocks 202 through 208, in some embodiments, a method for determining a tumor fraction in cell-free nucleic acid from a liquid biological sample of a subject is performed in a computer system, such as system 100 of FIG. 1 , having one or more processors 102 and memory 111 / 112 storing one or more programs, such as a disease monitoring module 120, executed by the one or more processors. In some such embodiments, a plurality of first sequence reads 140 are obtained in electronic form from a biological sample of the subject, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules.

[0124] Referring to block 204, in some embodiments, the subject is a human or a mammal. In some embodiments, the subject is any living or non-living organism, including but not limited to humans (e.g., male, female, fetus, pregnant female, child, etc.), non-human animals, plants, bacteria, fungi, or protists. In some embodiments, the subject is a mammal, reptile, bird, amphibian, fish, ungulate, ruminant, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), porcine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry, dog, cat, mouse, rat, fish, dolphin, whale, and shark. In some embodiments, a subject is male or female (e.g., man, woman, or child) at any stage of life.

[0125] In some embodiments, the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject (block 206). In many such embodiments, the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject, as well as other components of the subject (e.g., solid tissue, etc.).

[0126] In some embodiments, the biological sample consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject (block 208). In many of these embodiments, the biological sample is limited to blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject and does not include other components of the subject (e.g., solid tissue, etc.).

[0127] In some embodiments, the biological sample is processed to extract cell-free nucleic acids ready for sequencing analysis. By way of non-limiting example, in some embodiments, cell-free nucleic acids are extracted from a blood sample collected from a subject in a K2 EDTA tube. Within two hours of collection, multiple samples are processed by first spinning the blood at 1000g for ten minutes, followed by a second spin of the plasma at 2000g for ten minutes. The plasma is then stored at -80°C in 1 ml aliquots. In this way, an appropriate amount of plasma (e.g., 1 to 5 ml) is prepared from the biological sample for the various purposes of cell-free nucleic acid extraction. In some such embodiments, the cell-free nucleic acids are extracted using the QIAamp circulating nucleic acid reagent set (Qiagen) and eluted into DNA suspension buffer (Sigma). In some embodiments, the purified cell-free nucleic acids are stored at -20°C until use. For example, see Swanton et al., 2017, "Phylogenetic ctDNA analysis describes the evolution of early lung cancer," Nature, 545(7655): 446-451, which is incorporated herein by reference. For sequencing purposes, other equivalent methods can be used to prepare cell-free nucleic acids from a variety of biological methods, and all such methods are within the scope of the present disclosure.

[0128] In some embodiments, the cell-free nucleic acid obtained from a biological sample is any form of nucleic acid as defined herein, or a combination thereof. For example, in some embodiments, the cell-free nucleic acid obtained from a biological sample is a mixture of RNA and DNA.

[0129] Any form of sequencing can be used to obtain the plurality of sequence reads 140 from the cell-free nucleic acid obtained from the biological sample, including but not limited to a plurality of high-throughput sequencing systems, such as the Roche 454 platform, the SOLID platform from Applied Biosystems, the true single molecule DNA sequencing technology from Helicos, the hybridization sequencing platform from Affymetrix, the single molecule real-time (SMRT) technology from Pacific Biosciences, the sequencing-by-synthesis platforms from 454 Life Sciences, Illumina / Solexa, and Helicos, and the ligation sequencing platform from Applied Biosystems. ION TORRENT technology from Life Sciences and nanopore sequencing can also be used to obtain the plurality of sequence reads 140 from the cell-free nucleic acid obtained from the biological sample.

[0130] In some embodiments, sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego, CA)) are used to obtain a plurality of sequence reads 140 from the cell-free nucleic acid obtained from the biological sample. In some such embodiments, millions of cell-free nucleic acid (e.g., DNA) fragments are sequenced in parallel. In one example of this type of sequencing technology, a flow cell is used that contains an optically transparent slide having eight separate lanes on a surface to which a plurality of oligonucleotide anchors (e.g., adapter primers) are bound. A flow cell is often a solid support configured to hold and / or allow a reagent solution to pass sequentially through a plurality of bound analytes. In some examples, flow cells are planar, optically transparent, typically on the millimeter or submillimeter scale, and often have multiple channels or lanes in which analyte / reagent interactions occur. In some embodiments, a cell-free nucleic acid sample can include a signal or label that facilitates detection. In some such embodiments, obtaining a plurality of sequence reads 140 from the cell-free nucleic acid obtained from the biological sample comprises obtaining quantitative information of the signal or tag by various techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarray, mass spectrometry, cytofluorescence analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field levitation, sequencing, and combinations thereof.

[0131] In some embodiments, a plurality of sequence reads 140 are obtained in the manner described in the exemplary assay protocol disclosed in Example 10. In some embodiments, steps are taken to determine that each such read represents a unique nucleic acid fragment in the cell-free nucleic acid in the biological sample. Depending on the sequencing method used, each such unique nucleic acid fragment can be represented by a number of sequence reads. In typical examples, multiple sequencing techniques such as barcoding are used to resolve the redundancy of the multiple sequence reads for the multiple unique nucleic acid fragments in the cell-free nucleic acid. Thus, the number of sequence reads for a particular allele represents the number of unique nucleic acid fragments in the cell-free nucleic acid in the biological sample that map to different portions of the genome of the species represented by the individual allele, rather than the actual total number of sequence reads in the multiple sequence reads that map to the individual allele. See Kircher et al., 2012, Nucleic Acids Research 40, No. 1e3, incorporated herein by reference, for example, regarding disclosure of barcoding. In some embodiments, such mapping only allows for perfect matches. In some embodiments, such mapping allows for some false positives. In some embodiments, a program such as Bowtie 2 is used to perform such mapping. See, for example, Langmead and Salzberg, 2012, Nat Methods 9, pp. 357-359, for disclosure of such mapping.

[0132] In some embodiments, the plurality of first sequence reads obtained from the cell-free nucleic acid of a biological sample at block 202 includes more than ten sequence reads of the cell-free nucleic acid, more than one hundred sequence reads of the cell-free nucleic acid, more than five hundred sequence reads of the cell-free nucleic acid, more than one thousand sequence reads of the cell-free nucleic acid, more than two thousand sequence reads of the cell-free nucleic acid, between more than two thousand five hundred and five thousand sequence reads of the cell-free nucleic acid, or more than five thousand sequence reads of the cell-free nucleic acid. In some embodiments, each of the sequence reads is a different portion of the cell-free nucleic acid. In some embodiments, a sequence read 140 in the plurality of first sequence reads is all of the cell-free nucleic acid, or is the same portion as another sequence read in the plurality of first sequence reads.

[0133] refer to Figure 2AIn blocks 210 to 216 of the method, the plurality of first sequence reads 140 are used to identify support 146 for each variant 144 in a first set of variants 142, thereby determining an observed frequency for each variant in the first set of variants. In some embodiments, each variant 144 in the first set of variants 142 is obtained from the plurality of first sequence reads after performing noise modeling, junction modeling using white blood cells (WBCs), and / or artifact modeling of marginal variants, as disclosed in U.S. patent application Ser. No. 16 / 201,912, filed on November 27, 2018, entitled "Models for Targeted Sequencing," which is incorporated herein by reference.

[0134] refer to Figure 2B In block 219, in some embodiments, another sequence read 140 of the plurality of first sequence reads is considered to support a first variant 144 in the first variant set 142 when the individual sequence read (i) encompasses or is located in a genomic location associated with the first variant and (ii) contains all or a portion of the first variant. Another sequence read of the plurality of first sequence reads is considered not to support the first variant in the first variant set when the individual sequence read (i) encompasses or is located in a genomic location associated with the first variant and (ii) does not contain all or a portion of the first variant. For example, consider an example where a first variant is associated with a particular genomic location. The sequence reads encompassing or located in the particular genomic location are evaluated to determine whether they support the variant. In other words, the sequence reads that uniquely map to the particular genomic location are evaluated to determine whether they support the variant. If a sequence read encompasses or is located in a genomic location and encodes the variant, the sequence read is considered to support the variant. For example, in the case where the variant is a single nucleotide variation, sequence reads that (i) encompass the genomic position corresponding to the single nucleotide variation and (ii) have the single nucleotide variation are considered to support the variant. In another example, in the case where the variant is an insertion that is longer than an average length of the plurality of sequence reads, sequence reads that are in the genomic position corresponding to the variation (e.g., mapped to the genomic site where the insertion is to be incorporated) and (ii) have all or a portion of the insertion are considered to support the variant.

[0135] In some embodiments, the plurality of first sequence reads 140 are used to identify support for each variant 144 in a first set of variants by aligning each sequence read 140 in the plurality of first sequence reads to a region in a reference genome, thereby determining whether the sequence read contains all or a portion of a first variant 144 (block 214). The alignment of a sequence read 140 to a region in a reference genome involves pairing multiple sequences from one or more sequence reads 140 with multiple sequences of the reference genome based on full or partial identity between the multiple sequences. Multiple alignments can be performed manually or using a computer algorithm, examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis Pipeline. The alignment of a sequence read to the reference genome can be a 100% sequence match. In some embodiments, an alignment is less than a 100% sequence match (e.g., an imperfect match, a partial match, a partial alignment). In some embodiments, an alignment includes a mismatch. In some embodiments, an alignment includes 1, 2, 3, 4, or 5 mismatches. In some embodiments, a plurality of such mismatches indicate and support a variant 144 in a first variant set. For example, in an example where a variant 144 is a single nucleotide variant located at a specific position in the genome, an alignment of a sequence read containing the variant with the genome is expected to have a mismatch between the sequence read and the genome at a position in the genome associated with the single nucleotide variant. Two or more sequences can be aligned using any of the strands. In some embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid.

[0136] In alternative embodiments, the plurality of first sequence reads 140 are used to identify support for each variant 144 in a first set of variants by aligning a sequence read 140 from the plurality of first sequence reads to a lookup table of variants, thereby determining whether the sequence read contains all or a portion of a first variant 144 (block 214). Thus, in such examples, rather than using each sequence read 140 to find an alignment anywhere in the entire genome of a subject, each sequence read 140 is aligned to each of the plurality of sequences in a lookup table, where each such sequence in the lookup table represents a variant 144 in the first set of variants 142. As an example, consider again the example of a variant being a single nucleotide variant associated with a particular position in the genome. In this example, for the variant, the lookup table would include a portion of the sequence of the genome near the relevant position of the genome. In some examples, the size of this portion may depend on the type of sequencing method used to generate the plurality of sequence reads 140. As a non-limiting example, 50 bases flanking the 3' side of the position in the genome associated with the single nucleotide variant and 50 bases flanking the 5' side of the position in the genome associated with the single nucleotide variant are used to represent the variant in the lookup table. In some embodiments, as discussed below with respect to block 218, in some examples, the variant is some other type of variant, for example, an insertion mutation associated with a specific position in the genome. In many such examples, in the lookup table, the variant is represented by a portion of the genome sufficient to align with all or a significant portion of a sequence read containing the insertion mutation.

[0137] In some embodiments, the plurality of first sequence reads 140 are used to identify support 146 for each variant 144 in the first set of variants 142 using a variant calling process, such as HaplotypeCaller. See, for example, McKenna et al., 2010, "Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data," Genome Research 20:1297-303; and Van der Auwera, 2013, "From FastQ data to high-confidence variant calls: A best-practice pipeline for the Genome Analysis Toolkit," Current Protocols In Bioinformatics 43:11.10.1-11.10.33, each of which is incorporated herein by reference.

[0138] In some embodiments, the plurality of first sequence reads 140 are used to identify support 146 for each variant 144 in the first set of variants 142 using VarScan. See, for example, Koboldt et al., 2012, "VarScan 2: Discovery of somatic mutations and copy number alterations in cancer by exome sequencing," Genome Research, PMID: 22300766; and Koboldt et al., 2009, "VarScan: Variant detection in massively parallel sequencing of individual and pooled samples," Bioinformatics 25(17): 2283-5, each of which is incorporated herein by reference.

[0139] In some embodiments, the plurality of first sequence reads 140 are used to identify support 146 for each variant in the first set of variants 142 using Strelka. See, for example, Kim et al., 2017, "Strelka2: Rapid and Accurate Variant Calling for Clinical Sequencing Applications," bioRxiv doi:10.1101 / 192872, which is incorporated herein by reference.

[0140] In some embodiments, the plurality of first sequence reads 140 are used to identify support 146 for each variant in the first set of variants 142 using SomaticSniper. See, for example, Larson et al., 2012, "SomaticSniper: Identification of somatic point mutations in whole-genome sequencing data," Bioinformatics 28(3), pp. 311-317, each of which is incorporated herein by reference.

[0141] In some embodiments, according to Example 11, the plurality of first sequence reads 140 are used to identify support 146 for each variant in the first set of variants 142. In some embodiments, the plurality of sequence reads 140 are pre-processed using one or more methods, such as normalization, correction for GC bias, and / or correction for bias due to PCR over-amplification, to correct for biases or errors.

[0142] In some embodiments, the UMIs and endpoint positions of a plurality of sequence reads collected according to the present disclosure are used to define a plurality of bags of possible PCR repeats, which are collapsed (thereby obtaining an average split coverage) and stitched into a plurality of high-precision fragment sequences. Thus, in a plurality of such embodiments, the "coverage" reported for a plurality of sequence reads is the average split coverage for such bags. In some embodiments, a plurality of candidate variants are generated using a De Bruijn assembler and scored using a noise model trained on a population of non-smoking participants under 35 years of age who have not been diagnosed with cancer, which is used to measure the technical variation from the sequencing assay. The noise model provides a calibrated quality score estimated based on the support for each variant, thereby allowing the plurality of candidate variants to be filtered into a high-quality subset of variants in which pure technical variants are unlikely to occur. For example, in the example of targeted sequencing such as ART sequencing, some embodiments of the present disclosure use the noise model and heuristic algorithm for identifying multiple variants, as disclosed in U.S. patent application No. 16 / 201912, entitled "Model for Targeted Sequencing," filed on November 27, 2018. In the example of whole genome sequencing, some embodiments of the present disclosure use the noise model and heuristic algorithm for identifying multiple variants disclosed in U.S. patent application No. 16 / 352,214, entitled "Identifying Copy Number Abnormalities," filed on March 13, 2019. Multiple candidate variants are further filtered for multiple DNA damage artifacts that are clustered near multiple ends of multiple reads and occur in a subset of multiple samples. Multiple variants that are estimated to have a phred score of 60 or higher and are unlikely to be technical artifacts are considered multiple variants in some embodiments. Variants estimated to have a phred score of 40 or higher, 45 or higher, 50 or higher, 55 or higher, 60 or higher, 65 or higher, or 70 or higher that are unlikely to be technical artifacts are considered variants in some embodiments.

[0143] Multiple epigenetic features such as methylation are used as multiple variants. In some embodiments, according to example 13, and as U.S. patent application No. 62 / 642,480, entitled "Methylation Fragment Abnormal Detection" filed on March 13, 2018, which is incorporated herein by reference, the multiple first sequence reads 140 are used to identify support 146 for each variant in the first variant set 142 by determining one or more methylation state vectors. In a plurality of such embodiments, 5-cytosine methylation occurs in a CpG context. A method for determining methylation state is by bisulfite conversion sequencing (BS-seq). In the case of BS-seq, unmethylated cytosine is converted to uracil bases, which are read as thymine in sequencing. Therefore, in some embodiments, an epigenetic pattern such as the methylation state at one or more nucleotide positions is used as a basis for determining a variant allele, wherein the ctDNA score is determined for the variant allele. In some embodiments, the methylation may include a methylation index of a CpG site, a methylation density of multiple CpG sites in a region (e.g., including 2 or more, 3 or more, 4 or more, 5 or more, or 6 or more CpG sites), a distribution of multiple CpG sites across a continuous region, a pattern or degree of methylation for each individual CpG site in a region containing more than one CpG site, and / or non-CpG methylation. "DNA methylation" in mammalian genomes may refer to the addition of a methyl group at position 5 of the heterocycle of cytosine in CpG dinucleotides (e.g., to produce 5-methylcytosine). Methylation of cytosine may occur in multiple cytosines in other sequence contexts, for example, 5'-CHG-3' and 5'-CHH-3', where H is adenine, cytosine, or thymine. Cytosine methylation may also be in the form of 5-hydroxymethylcytosine. Methylation of DNA may include methylation of non-cytosine nucleotides, such as N6-methyladenine. In some embodiments, the cell-free nucleic acid fragments are treated to convert unmethylated cytosines to uracils. In one embodiment, the method uses a bisulfite treatment of DNA that converts unmethylated cytosines to uracils without converting methylated cytosines. For example, a DNA methylation reaction such as EZ DNA methylation reaction is performed. TM -Gold, EZ DNA methylation TM - Direct or EZ DNA methylation TMA commercial kit, such as the Zymo Lightning kit (available from Zymo Research, Inc. (Irvine, CA), is used to perform the bisulfite conversion method. In another embodiment, the conversion of unmethylated cytosines to uracils is accomplished using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracils, such as APOBEC-Seq (NEBiolabs, Ipswich, MA), or by using techniques disclosed in Schutsky et al., 2018, “Nondestructive base-resolution sequencing of 5-hydroxymethylcytosine using DNA deaminases,” Nature Biotechnology 36, 1083-1090, or Liu et al., 2019, “Direct bisulfite-free detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution,” Nature Biotechnology 37, pp. 424-429. A sequencing library is prepared from the multiple transformed cell-free nucleic acid fragments. Alternatively, the sequencing library is enriched with multiple cell-free nucleic acid fragments or multiple genomic regions, which can use multiple hybridization probes to provide information for cell origin. The multiple hybridization probes are multiple short oligonucleotides, which hybridize with multiple specially designated cell-free nucleic acid fragments or target regions and enrich these fragments or regions for subsequent sequencing and analysis. In some embodiments, multiple hybridization probes are used to perform targeted and high-depth analysis of a group of designated CpG sites that provide information for cell origin. Once prepared, the sequencing library or a portion of the sequencing library is sequenced to obtain multiple sequence reads. In multiple alternative embodiments, whole genome bisulfite sequencing is performed as described for the CCGA study in Example 12 (WGBS; 34 times).

[0144] Methylation sequencing is used to identify multiple variants. In some embodiments, whole genome bisulfite sequencing (WGSB) or targeted bisulfite sequencing is used to obtain the multiple sequence reads 140. For example, in some embodiments, a WGBS with a coverage of 34x from the CCGA study described in Example 12 is used. In some embodiments, the coverage of such a WGBS is 100x or less, 50x or less, or between 30x and 200x. In various exemplary embodiments, multiple unique molecular indices (UMIs) and multiple endpoint positions of a sequence read are used to define multiple possible PCR replicates, which are split into multiple packets to achieve such coverage statistics. In some embodiments, a single sequence read from each packet is used in the disclosed analysis. In some implementations, this single sequence read is a consensus sequence read. In some embodiments, this single sequence read is any sequence read in a packet. Thus, in this manner, 100-fold refers to the number of unique fragments covering each allele position, rather than the number of sequence reads covering each allele position, as such multiple sequence reads may include multiple PCR replicates. Such multiple sequence reads from the multiple split packets can be used to detect multiple sequencing variations (e.g., single nucleotide variants, insertions, deletions) or multiple copy number variations. In some embodiments where multiple sequence reads are used to identify multiple single nucleotide variants, multiple C->T or T->C variants between non-cancer and cancer cannot be used because multiple unmethylated uracils are converted to multiple uracils, which are read as thymines during sequencing; for example, by including a variant noise filter in a noise model for variant identification. In some embodiments, the noise model is modified to include one or more parameters to account for the strand origin of a sequence read (e.g., whether the read is from the forward strand or the reverse strand of the original target molecule). Multiple additional factors can be taken into account, including but not limited to trinucleotide background, the position of the variant in the fragment, and different types of other covariates. In some embodiments where multiple sequence reads are used to identify multiple single nucleotide variants, in fact, multiple C->T or T->C variants between non-cancer and cancer can be used as long as the bisulfite treatment of the DNA converts the multiple unmethylated cytosines to multiple uracils but does not convert multiple methylated cytosines. TM -Gold, EZ DNA methylation TM - Direct or EZ DNA methylation TMThis can be achieved by using a commercial kit such as the Lightning kit (available from Limo Research, Irvine, CA) or by combining the techniques disclosed in Schutsky et al., 2018, "Nondestructive base-resolution sequencing of 5-hydroxymethylcytosine using DNA deaminases," Nature Biotechnology 36, 1083-1090, or Liu et al., 2019, "Direct bisulfite-free detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution," Nature Biotechnology 37, pp. 424-429, for the bisulfite conversion method. A sequence library is prepared from the plurality of converted cell-free nucleic acid fragments. Optionally, the sequence library is enriched for a plurality of cell-free nucleic acid fragments or a plurality of genomic regions, which can provide information about cell origin using a plurality of hybridization probes. The plurality of hybridization probes are short oligonucleotides that hybridize to a plurality of specifically designated cell-free nucleic acid fragments or target regions and enrich these fragments or regions for subsequent sequencing and analysis. In some embodiments, the plurality of hybridization probes are used to perform a targeted and high-depth analysis of a set of designated CpG sites that provide information about cell origin. Once prepared, the sequencing library or a portion of the sequencing library is sequenced to obtain a plurality of sequence reads. In alternative embodiments, whole genome bisulfite sequencing is performed as described for the CCGA study in Example 12 (WGBS; 34x).

[0145] Whole-Genome Plasma Assays. In some embodiments, the subject is human, and the plurality of first sequence reads 140 obtained from the biological sample are part of a whole-genome plasma assay.

[0146] In some such embodiments, the whole-genome plasma assay is performed using cfDNA extracted from two tubes of plasma using a modified QIAamp circulating nucleic acid kit (Qiagen; Germantown, MD). Genomic DNA (gDNA) from buffy coats was extracted using Qiagen's DNEasy Blood and Tissue Kit and quantified using NanoDrop (Thermo Fisher Scientific; Waltham, MA). The extracted gDNA was fragmented using a Covaris E220 ultrasonic disruptor (Woburn, MA) and size-selected using Agencourt AMPure XP magnetic beads (Beckman Coulter; Beverly, MA). Plasma cfDNA (up to 75 ng) and buffy coat gDNA (75 ng) were used to construct next-generation sequencing (NGS) libraries. Adapters include a set of 218 unique molecular index (UMI) sequences to reduce measurement and sequencing errors. A small portion (4 μl of the 25 μl) of each amplified library was diluted and quantified using the AccuClear Ultra-High Sensitivity dsDNA Quantification Kit (Biotium; Fremont, CA). The remainder was used in a targeted sequencing protocol (see below). Three or four diluted libraries were normalized, pooled, and clustered in a flow cell before sequencing on an ILlumina HiSeq X (30x).

[0147] The plurality of sequence reads 140 are compared to the entire human genome to identify a plurality of variants. In some embodiments, the plurality of first sequence reads 140 obtained from the biological sample have at least 30 times coverage of a target gene panel, at least 40 times coverage of a target gene panel, at least 50 times coverage of a target gene panel, at least 60 times coverage of a target gene panel, or at least 70 times coverage of a target gene panel. In some such embodiments, the target gene panel is between 450 and 550 genes. In some embodiments, the target gene panel is within the range of 500 ± 5 genes, within the range of 500 ± 10 genes, or within the range of 500 ± 25 genes. In some embodiments, the whole genome assay plasma searches for changes in multiple somatic copy numbers (SCNA) or multiple fragmented features in the genome.

[0148] Targeted Plasma Assays. In some embodiments, the subject is human, and the plurality of first sequence reads 140 obtained from the biological sample are part of a targeted plasma assay.

[0149] In some such embodiments, as part of the ART assay disclosed in Example 12, the multiple amplified libraries (see whole genome plasma assay above) are used for targeted enrichment using a panel of 507 cancer-related genes. Up to 3.5 micrograms of each library undergo hybridization-based capture. The multiple enriched libraries are quantified using the Yekuklelle Ultra High Sensitivity dsDNA Quantification Kit. Three or four enriched libraries are normalized, merged, aggregated in a flow cell, and sequenced on an Illumina HiSeq X (150-bp paired end sequencing, 60,000x).

[0150] The plurality of sequence reads 140 obtained in this manner are compared to a target gene panel of the targeted plasma assay to identify a plurality of variants. In some such embodiments, the target gene panel is between 450 and 550 genes. In some embodiments, the target gene panel is within a range of 500 ± 5 genes, within a range of 500 ± 10 genes, or within a range of 500 ± 25 genes. In some embodiments, the plurality of first sequence reads 140 obtained from the biological sample have at least 50,000-fold coverage of the target gene panel, at least 55,000-fold coverage of the target gene panel, at least 60,000-fold coverage of the target gene panel, or at least 70,000-fold coverage of the target gene panel. In some such embodiments, the targeted plasma assay looks for multiple single nucleotide variants in the target gene panel, multiple insertions in the target gene panel, multiple deletions in the target gene panel, multiple somatic copy number alterations (SCNAs) in the target gene panel, multiple aberrant methylation patterns, or rearrangements affecting the target gene panel.

[0151] Targeted leukocyte assay. In some embodiments, the subject is human, and the plurality of first sequence reads 140 obtained from the biological sample are part of a targeted leukocyte assay. That is, the biological sample is a plurality of leukocytes from the subject, and the plurality of sequence reads 140 are compared with a target gene panel of the targeted leukocyte assay to identify a plurality of variants. In some such embodiments, the target gene panel is between 450 and 550 genes. In some embodiments, the target gene panel is within the range of 500±5 genes, within the range of 500±10 genes, or within the range of 500±25 genes. In some embodiments, the plurality of first sequence reads 140 obtained from the biological sample have at least 50,000 times coverage of the target gene panel, at least 55,000 times coverage of the target gene panel, at least 60,000 times coverage of the target gene panel, or at least 70,000 times coverage of the target gene panel. In some such embodiments, the targeted leukocyte assay looks for multiple single nucleotide variants in the target gene panel, multiple insertions in the target gene panel, multiple deletions in the target gene panel, or multiple somatic copy number alterations (SCNAs) in the target gene panel.

[0152] Whole-genome leukocyte assay. In some embodiments, the subject is human, and the plurality of first sequence reads 140 obtained from the biological sample are part of a whole-genome leukocyte assay. That is, the biological sample is a plurality of leukocytes from the subject, and the plurality of sequence reads 140 are compared with the entire human genome to identify a plurality of variants. In some embodiments, the plurality of first sequence reads 140 obtained from the biological sample have at least 30 times coverage of a target gene panel, at least 40 times coverage of a target gene panel, at least 50 times coverage of a target gene panel, at least 60 times coverage of a target gene panel, or at least 70 times coverage of a target gene panel. In some such embodiments, the target gene panel is between 450 and 550 genes. In some embodiments, the target gene panel is within the range of 500 ± 5 genes, within the range of 500 ± 10 genes, or within the range of 500 ± 25 genes. In some embodiments, the whole genome leukocyte assay looks for multiple somatic copy number alterations (SCNAs) or multiple segmental signatures in the genome.

[0153] Whole genome bisulfite sequencing assay. In some embodiments, the subject is human, and the plurality of first sequence reads 140 are obtained by performing bisulfite sequencing, and a plurality of variants are evaluated on a genome-wide basis. In some embodiments, the whole genome bisulfite sequencing assay searches for variants of a plurality of methylation patterns in the genome. For example, see Example 13. See also U.S. Patent Application No. 62 / 642,480, entitled “Methylation Fragment Abnormal Detection,” filed on March 13, 2018, which is incorporated herein by reference.

[0154] In some embodiments, reference Figure 2A Block 216 of the present invention uses the plurality of first sequence reads 140 to identify support 146 for each variant 144 in a first set of variants by comparing each sequence read 140 in the plurality of first sequence reads with each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome (e.g., a reference genome). In some examples, such an embodiment is used to populate the lookup table with multiple hotspots in the genome. Thus, instead of comparing each sequence read by searching for comparisons in the entire genome, each sequence read is only compared to those portions of the genome that are relevant to the condition of interest, such as multiple genes within the genome. For example, consider an example in which a mutation in a particular gene is associated with a clinical condition. In this example, according to an embodiment of block 216, the genomic sequence of the gene may be included as an entry in the lookup table, and multiple sequence reads 140 may be compared to this entry to identify support for the gene on a variant pair. In a variation of this example, each known mutation of the gene can be listed as a separate entry in the lookup table, and each sequence read 140 in the plurality of first sequence reads can be aligned with each of these separate entries to determine whether there is a match between the sequence read and one of the plurality of mutations in the plurality of genes, thereby identifying support for a variant 144 in the set of variants.

[0155] In yet another example, only those variants found in an abnormal tissue (e.g., a tumor) of a particular subject (and sufficient genomic sequences near such variants) are included in the lookup table. In this way, compared to multiple embodiments in which the multiple sequence reads are compared to the entirety of a reference genome to identify support for such variants, the process of pairing with one of the multiple variants in the tumor to identify support for the variant is greatly accelerated. For example, consider an example in which sequencing of a particular tumor of a subject identifies three variants. In this example, the three variants are provided as multiple separate entries in the lookup table, and each of the multiple sequence reads 140 in the multiple first sequence reads is independently compared to each of the multiple entries in the lookup table to determine whether they are compared to one of the multiple variants, thereby supporting the variant.

[0156] In some embodiments, the lookup table consists of a single entry, wherein the single entry is a variant that has been identified in an abnormal tissue of a subject. In some embodiments, the lookup table consists of two entries, wherein each entry represents a variant that has been identified in an abnormal tissue of a subject. In some embodiments, the lookup table consists of three entries, wherein each entry represents a variant that has been identified in an abnormal tissue of a subject. In some embodiments, the lookup table consists of between three and ten entries, wherein each entry represents a variant that has been identified in an abnormal tissue of a subject.

[0157] In some embodiments, the lookup table includes between two and one thousand entries, where each entry represents a different gene in the human genome.

[0158] refer to Figure 2A In block 218, in some embodiments, a variant 144 in the first set of variants 142 is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the number of copies of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an aberrant methylation pattern associated with a predetermined genomic location. By way of example, in some embodiments, a variant 144 is a single nucleotide variant associated with a specific gene in the genome, and is therefore associated with the genomic location of the specific gene in the genome.

[0159] In some embodiments, the variant set 142 includes more than one variant type. For example, in some embodiments, the variant set 142 includes a single nucleotide variant associated with a genomic position and a deletion mutation associated with another genomic position in a genome.

[0160] In some embodiments, a variant 144 is any form of somatic mutation.

[0161] In some embodiments, each of the plurality of variants 144 in the first variant set is also found in the reference set 128. In some embodiments, there is a one-to-one correspondence between the plurality of variants 144 in the variant set 142 and the plurality of variants 130 in the reference set 128. In many such embodiments, the variant set 142 includes the identified support 146 for the plurality of samples in the biological sample (e.g., blood) of the subject, and the reference set 128 includes the reference frequencies 132 of such plurality of variants in the abnormal tissue (e.g., tumor) of the subject.

[0162] In some embodiments, the first variant set 142 consists of a single variant 144 for a single genetic variation at a single site in the genome of the subject (block 220). For example, consider an example where a specific single nucleotide variant is found in a specific gene in a certain percentage of the sequence reads mapped to the specific gene from the abnormal tissue (e.g., tumor) of a subject. In this example, the variant set 142 will also include the specific single nucleotide variant and any support 146 identified for the specific single nucleotide variant in the specific gene found in the first sequence reads obtained from the biological sample (e.g., blood) of the subject.

[0163] In some embodiments, the first variant set 142 is comprised of a first variant 144-1 for a first genetic variation located at a first site in the genome of the subject and a second variant 144-2 for a second genetic variation located at a second site in the genome of the subject (block 222). For example, consider an example in which a first variant is found in a first gene at some estimated percentage (e.g., greater than 1%, greater than 2%, greater than 5%) or the number of sequence reads mapped to the first gene in an abnormal tissue (e.g., a tumor) of a subject; and a second variant is found in a second gene at some estimated percentage or the number of sequence reads mapped to the second gene in the abnormal tissue. In this example, the variant set 142 would include the first variant and any identified support 146 for the first variant found in the first sequence reads obtained from the biological sample (e.g., blood) of the subject. The variant set 142 will also include the second variant and any identified support 146 for the second variant in the second gene found in the plurality of first sequence reads obtained from the biological sample.

[0164] In some embodiments, a variant is included in the reference set when at least one sequence read from the abnormal tissue supports a variant. In a plurality of such embodiments, a sequence read from the abnormal tissue supports the variant when the sequence read (i) maps to a genomic position associated with a variant and (ii) includes the variant. In some embodiments, a variant is included in the reference set when at least two sequence reads from the abnormal tissue support a variant. In some embodiments, a variant is included in the reference set when at least 2 sequence reads, at least 5 sequence reads, at least 10 sequence reads, at least 100 sequence reads, at least 200 sequence reads, or at least 1000 sequence reads from the abnormal tissue support a variant.

[0165] In some embodiments, the first variant set 142 is comprised of a first variant 144-1 for a first genetic variation located at a first site in the genome of the subject, a second variant 144-2 for a second genetic variation located at a second site in the genome of the subject, and a third variant 144-1 for a third genetic variation located at a third site in the genome of the subject (block 224). For example, consider an example in which a first variant is found in a first gene at some estimated percentage or number of sequence reads of an abnormal tissue (e.g., a tumor) of a subject that includes the first gene; a second variant is found in a second gene at some estimated percentage or number of sequence reads of the abnormal tissue that includes the second gene; and a third variant is found in a third gene at some estimated percentage or number of sequence reads of the abnormal tissue that includes the third gene. In this example, the variant set 142 will include the first variant and any identified support 146 for the first variant found in the plurality of first sequence reads obtained from the biological sample (e.g., blood) of the subject. The variant set 142 will also include the second variant and any identified support 146 for the third variant in the third gene found in the plurality of first sequence reads obtained from the biological sample.

[0166] In some embodiments, the first variant set 142 consists of between 2 and 20 variants, wherein each variant in the first variant set represents a different genetic variation at a different site in the genome of the subject (block 226). In some embodiments, the first variant set 142 consists of between 2 and 20 variants, wherein each variant in the first variant set represents a different genetic variation in the genome of the subject (block 226). In some embodiments, each individual variant in the first variant set is also found in an estimated percentage (e.g., greater than 1%, greater than 2%, greater than 5%) or in the number of sequence reads mapping to the genomic location of the individual variant in an abnormal tissue (e.g., a tumor) of a subject. In some embodiments, the first variant set 142 consists of between 1 and 10 variants, wherein each variant in the first variant set represents a different genetic variation (optionally at a different site) in the genome of the subject. In some embodiments, the first variant set 142 consists of between 1 and 100 variants, wherein each variant in the first variant set represents a different genetic variation in the genome of the subject (optionally at a different site). In some embodiments, the first variant set 142 consists of between 2 and 100 variants, wherein each variant in the first variant set represents a different genetic variation in the genome of the subject (optionally at a different site). In some embodiments, the first variant set 142 consists of between 1 and 1000 variants, wherein each variant in the first variant set represents a different genetic variation in the genome of the subject (optionally at a different site).

[0167] In some embodiments, a first variant and a second variant in the set of variants are associated with the same site in the genome of a subject. For example, the first and second variants may represent two different abnormal alleles of the same gene.

[0168] For each individual variant in the first set of variants, a corresponding reference frequency for the individual variant in a first reference set is obtained, wherein each corresponding reference frequency in the first reference set is for another variant in a first abnormal solid tissue sample obtained from the subject. Figure 2BAt block 228, in various disclosed methods, the observed frequency (e.g., support 146) of each individual variant 144 in the first set of variants 142 is compared to a corresponding reference frequency 132 for the individual variant in a first reference set 128. Each corresponding reference frequency 132 in the first reference set 128 is a frequency of another variant 130 in a first abnormal tissue sample obtained from the subject.

[0169] refer to Figure 2BIn block 230, in some embodiments, the first abnormal tissue sample is a tumor sample, or a small portion of the tumor sample. In some embodiments, the first abnormal tissue sample is adrenocortical carcinoma, childhood adrenocortical carcinoma, AIDS-related cancer tumors, Kaposi's sarcoma, anal cancer-related tumors, appendix cancer-related tumors, astrocytomas, childhood (brain cancer) tumors, atypical teratoma / rhabdomyosarcoma, central nervous system (brain cancer) tumors, basal cell carcinoma of the skin, bile duct carcinoma, or cancer-related tumors, bladder cancer tumors, childhood bladder cancer tumors, bone cancer (e.g., Ewing's sarcoma and osteosarcoma and malignant fibrous histiocytoma) tissue, brain tumors, breast cancer tissue, childhood breast cancer tissue, childhood bronchial tumors, Burkitt's lymphoma tissue, carcinoid tumors (gastrointestinal), childhood carcinoid tumors, carcinoma of unknown primary, carcinoma of unknown primary in children, childhood cardiac (heart) tumors, central nervous system (e.g., brain cancer such as atypical teratoma / rhabdomyosarcoma in children), embryonal tumors, childhood germ cell tumors, cervical cancer tissue, childhood cervical cancer tissue, cholangiocarcinoma tissue, childhood chordoma tissue, chronic myeloproliferative neoplasms, colorectal cancer tumors, colorectal cancer in children Cancer tumors, childhood craniopharyngioma tissue, ductal carcinoma in situ (DCIS), childhood embryonal tumors, endometrial cancer (uterine cancer) tissue, childhood ependymoma tissue, esophageal cancer tissue, childhood esophageal cancer tissue, olfactory neuroblastoma (head and neck cancer) tissue, childhood extracranial germ cell tumor, extragonadal germ cell tumor, eye cancer tissue, intraocular melanoma, retinoblastoma, fallopian tube cancer tissue, gallbladder cancer tissue, gastric (stomach) cancer tissue, childhood gastric (stomach) cancer tissue, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), childhood gastrointestinal stromal tumors, germ cell tumors (e.g., childhood central nervous system germ cell tumors, childhood extracranial germ cell tumors, extragonadal germ cell tumors, ovarian germ cell tumors,or testicular cancer tissue), head and neck cancer tissue, childhood heart tumors, hepatocellular carcinoma (HCC) tissue, islet cell tumors (pancreatic neuroendocrine tumors), kidney or renal cell carcinoma (RCC) tissue, laryngeal cancer tissue, leukemia, liver cancer tissue, lung cancer (non-small cell and small cell) tissue, childhood lung cancer tissue, male breast cancer tissue, malignant fibrous histiocytoma and osteosarcoma of bone, melanoma, childhood melanoma, intraocular melanoma, childhood intraocular melanoma, Merkel cell carcinoma, malignant mesothelioma, childhood mesothelioma, metastatic cancer tissue, metastatic squamous cell neck cancer tissue with a hidden primary, melanoma with a NUT gene alteration Tissues of nasal and paranasal sinus cancer, nasopharyngeal cancer (NPC), neuroblastoma, non-small cell lung cancer, oral cancer, lip and oral cancer, oropharyngeal cancer, osteosarcoma and malignant fibrous histiocytoma of bone, ovarian cancer, pediatric ovarian cancer, pancreatic cancer, pediatric pancreatic cancer, papilloma (larynx) tissue, Paraganglioma tissue, childhood paraganglioma tissue, paranasal sinus and nasal cavity cancer tissue, parathyroid cancer tissue, penile cancer tissue, pharyngeal cancer tissue, pheochromocytoma tissue, childhood pheochromocytoma tissue, pituitary adenoma, plasmacytoma / multiple myeloma, pleuropulmonary blastoma, primary central nervous system (CNS) lymphoma, primary peritoneal cancer tissue, prostate cancer tissue, rectal cancer tissue, retinoblastoma, childhood rhabdomyosarcoma, salivary gland cancer tissue, sarcoma (e.g., childhood hemangioma, osteosarcoma, uterine sarcoma, etc.), Sézary syndrome (lymphoma) tissue, skin cancer tissue, childhood skin cancer group tissue, small cell lung cancer tissue, small intestinal cancer tissue, squamous cell carcinoma of the skin, squamous cell neck cancer with a cryptic primary, cutaneous T-cell lymphoma, testicular cancer tissue, testicular cancer tissue in children, laryngeal cancer (e.g., nasopharyngeal cancer, oropharyngeal cancer, hypopharyngeal cancer) tissue, thymoma or thymic carcinoma, thyroid cancer tissue, transitional cell carcinoma of the renal pelvis and ureter, cancer of unknown primary, ureter or renal pelvis tissue, transitional cell (kidney (renal cell)) cancer tissue, urethral cancer tissue, endometrial cancer tissue, uterine sarcoma tissue, vaginal cancer tissue, vaginal cancer tissue in children, hemangioma, vulvar cancer tissue, Wilms' tumor or other childhood kidney tumors.

[0170] In some embodiments, the plurality of sequence reads from the first abnormal tissue sample are a plurality of formalin-fixed and paraffin-embedded (FFPE) tumor tissue sections, which are scraped and sent to the Genomic Services Laboratory at HudsonAlpha Institute for Biotechnology (Huntsville, Alabama), where DNA is extracted from the scrapings and converted into a plurality of NGS libraries for whole-genome sequencing on an Illumina HiSeq X (30x). For each tissue scraping, a corresponding tube of buffy coat is sent to HudsonAlpha for extraction, library preparation, and whole-genome sequencing on an Illumina HiSeq X (60x). The sequencing data is then analyzed according to the present disclosure.

[0171] With reference to blocks 234 to 240, in some embodiments, the frequency (reference frequency 132) of each variant 130 in the first reference set 128 is obtained from a plurality of second sequence reads (a plurality of reference sequence reads) 126 obtained from the first abnormal tissue sample (block 234). In some embodiments, the frequency of an additional variant is a measure of the proportion of a plurality of cells in the first abnormal tissue of the subject to which the variant belongs. For example, see Lu et al., 2015 "Allele frequencies of somatic mutations in individuals reveal characteristics of cancer-associated genes", Acta Biochim Biophys Sin. 47(8), 657-680, which is incorporated herein by reference so as to disclose the steps for determining the frequencies of a plurality of somatic variants in an abnormal tissue according to some embodiments.

[0172] In some embodiments, the frequency of the individual variant 130 is determined by first identifying the plurality of sequence reads that may have a separate variant 130. For example, if the individual variant is a single nucleotide variant, the plurality of variants from the first abnormal tissue that map to the genomic location corresponding to the individual variant are identified. Next, the proportion of these identified sequence reads that include the variant represents the frequency of the individual variant. Thus, if there are 200 sequence reads from the abnormal tissue that map to the genomic location associated with the variant, and 50 of these sequence reads include the allele for the variant, while the remaining 150 sequence reads have a wild-type allele instead of the allele for the variant, then the frequency for the individual variant 130 is 25 percent. In some such embodiments, multiple steps are taken to ensure that each such sequence read represents a unique nucleic acid fragment in the abnormal tissue. Depending on the sequencing method used, each such unique nucleic acid fragment can be represented by the number of multiple sequence reads. In typical examples, multiple sequencing technologies such as barcodes are used to address the redundancy of multiple sequence reads of multiple unique nucleic acid fragments in the abnormal solid tissue sample, so that the number of the multiple sequence reads for a particular allele represents the number of multiple unique nucleic acid fragments in the abnormal solid tissue sample that are mapped to different parts of the genome of the species represented by the individual alleles, rather than the actual total number of multiple sequence reads in the multiple sequence reads that are mapped to the individual alleles. See Kircher et al., 2012, Nucleic Acids Research 40, No. 1e3, which is incorporated herein by reference, for example, regarding disclosure of barcodes.

[0173] In some embodiments, more than 1000, 2000, 3000, 4000, 5000, 10,000, 20,000, 100,000, or one million reference sequence reads 126 are obtained from the abnormal tissue. In some embodiments, the plurality of reference sequence reads 126 obtained from the abnormal tissue provide a coverage of 1 fold or more, 2 fold or more, 5 fold or more, 10 fold or more, or 50 fold or more of at least 2%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 98%, or at least 99% of the genome of the subject. In some embodiments, the plurality of reference sequence reads 126 obtained from the abnormal tissue provide a coverage of 1 fold or more, 2 fold or more, 5 fold or more, 10 fold or more, or 50 fold or more for at least three genes, at least five genes, at least ten genes, at least twenty genes, at least thirty genes, at least forty genes, at least sixty genes, at least seventy genes, at least eighty genes, at least 200 genes, at least 300 genes, at least 400 genes, at least 500 genes, or at least 1000 genes of the genome of the subject.

[0174] In some embodiments, the plurality of reference sequence reads 126 obtained from the first abnormal tissue are analyzed (aligned) against a panel of variant candidates. For example, in some embodiments, the panel of variant candidates includes a plurality of sequences of a plurality of variant candidates for at least three genes, at least five genes, at least ten genes, at least twenty genes, at least thirty genes, at least forty genes, at least fifty genes, at least sixty genes, at least seventy genes, at least eighty genes, at least ninety genes, at least 200 genes, at least 300 genes, at least 400 genes, at least 500 genes, or at least 1000 genes of the subject. To perform such an analysis, alignment of a particular reference sequence read 126 with the sequence of a variant candidate from the panel of variant candidates involves pairing the sequence of the reference sequence read 126 with the sequence of the variant candidate to see if there is full or partial identity between the plurality of sequences. Multiple such comparisons (analysis) can be done manually or by a computer algorithm, with multiple examples including the computer program Efficient Local Alignment of Nucleotide Data (ELAND) being distributed as part of the Illumina Genomics analysis pipeline. In some embodiments, the reference sequence read 126 is considered matched to the sequence of the variant candidate in the variant candidate panel when 100% of the reference sequence reads 126 are paired with a corresponding portion of the sequence of the variant candidate. In some embodiments, the reference sequence read 126 is considered matched to the sequence of the variant candidate in the variant candidate panel when 100% of the sequence of the variant candidate 126 is paired with a corresponding portion of the sequence of the reference sequence read 126. In some embodiments, an alignment is a sequence pairing less than 100% (e.g., an imperfect pairing, partial pairing, partial alignment). In some embodiments, an alignment includes an error pairing. In some embodiments, an alignment includes 1, 2, 3, 4, or 5 error pairs. Two or more sequences can be aligned using any of the above methods. In some embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.

[0175] refer to Figure 2CIn blocks 244 to 246, in some embodiments, the plurality of reference sequence reads 126 obtained from the first abnormal tissue sample represent the whole genome data for the individual cells. In some such embodiments, an average coverage of the plurality of reference sequence reads 126 obtained from the first abnormal tissue sample is at least 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, at least 20-fold, at least 30-fold, or at least 40-fold of the entire genome of the subject. In some embodiments, the average coverage of the plurality of second sequence reads of the entire first reference set 128 is at least 10-fold, at least 100-fold, or at least 2000-fold.

[0176] refer to Figure 2CIn block 248, in some embodiments, an additional sequence read 126 in the plurality of second sequence reads is considered to support a first variant 130 in the reference set 128 when the additional sequence read (i) maps to a portion of the genome associated with the first variant, and (ii) the additional sequence read 126 contains all or a portion of the first variant 130. An additional sequence read 126 in the plurality of second sequence reads is considered to not support a first variant 130 in the reference set 128 when the additional sequence read 126 (i) maps to a portion of the genome associated with the first variant 130 (corresponding to the genomic location of the first variant), and (ii) does not contain the first variant 130. For example, consider an example in which the variant is a single nucleotide variant associated with a predetermined genomic location. In this example, a sequence read 126 supports the variant when the first variant maps to the predetermined genomic location and contains the single nucleotide variant. In practice, in order to determine whether the first variant contains this single nucleotide variant, the sequence read also includes the 5' and 3' sequences on both sides of this single nucleotide variant in the genome of the species of the subject, so that the sequence read is mapped to the genome, thereby determining whether it is mapped to the genomic position corresponding to the variant. Next, consider that the variant is an example of 38 insertion bases inserted into a specific gene. When a sequence read contains the 38 insertion bases (and the 5' and 3' regions on both sides of this insertion in the specific gene), the sequence read will support this variant. In some examples, when the sequence read includes less than all of the variant, the sequence read may still support this variant. For example, the sequence read may terminate at 25 bases of the 38 insertion bases. However, the region of the sequence read on both sides of this variant can be paired with the gene and the first 25 bases of the insertion, and therefore, the sequence read can be considered to support the variant. Further, the number of sequence reads 126 in the plurality of second sequence reads that support a first variant 130 in the reference set 128 relative to the number of sequence reads 126 in the plurality of second sequence reads that do not support the first variant 130 determines the observed frequency (support 132) of the first variant 130. For example, again considering an example where a variant is associated with a particular genomic location in the genome, wherein the plurality of second reference sequence reads 126 from the first abnormal sample consists of 1000 sequence reads, but only 100 of the 1000 sequence reads cover (map to, are associated with) the genomic location associated with the variant. The 100 sequence reads covering the genomic location associated with the variant are analyzed to see whether they support or do not support the variant.These sequence reads that contain all or part of the variant in the 100 sequence reads are considered to support the variant, and these sequence reads that do not contain the variant in the 100 sequence reads are considered to not support the variant.Other 900 sequence reads do not meet the support or support of the variant because they do not cover the genomic position related to the variant in question. Further, considering an example, 3 of the 100 sequence reads contain all or part of the variant, and are considered to support the variant, while the remaining 97 sequence reads in the 100 sequence reads do not contain the variant, and therefore do not support the variant. According to this embodiment, in this example, the observed frequency (support 146) for the first variant is 3 / 100 or three percent.

[0177] refer to Figure 2D At block 256, the method continues by evaluating the observed frequency of each individual variant in the first set of variants 142 against the observed frequency of the individual variant in the first reference set 128 of the first abnormal solid tissue, thereby determining a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

[0178] In some embodiments, the first tumor score is used to classify a subject by considering the subject as having a first condition when the observed frequency (support 14) of each variant 144 in the first set of variants 142 meets a first threshold, wherein the first threshold is determined by a frequency of each variant 130 in the first reference set 128 of the first abnormal tissue sample. For example, reference Figure 2DIn block 258, in some embodiments, the evaluating step of block 256 includes calculating a single estimated ctDNA fraction in the cfDNA of the subject from the observed frequency (support 146) of each variant 144 in the first variant set 142 of the plurality of first sequence reads. Further, in many such embodiments, the first threshold is a single expected ctDNA fraction in the cfDNA of the subject determined from the frequency (reference frequency 132) of each variant 130 in the reference set 128 for the first abnormal tissue sample. For example, consider the example of comparing a single variant in the evaluating step of block 256. Thus, the support 146 for this variant in the variant set 142 from the biological sample (e.g., blood) is compared to the reference frequency 132 of the same variant in the reference set 128 for the abnormal tissue. It is assumed that the sole source of the single variant in the cell-free nucleic acid is from the abnormal tissue. Therefore, under this assumption, the single estimated ctDNA fraction is calculated based on the ratio of the support 146 for the variant in the variant set to the reference frequency 132 for the same variant in the reference set. For example, if the support 146 for the variant is 3 out of 100 sequence reads in the variant set 142, and the reference frequency 132 for the same variant in the reference set 128 is 0.10, then the single estimated ctDNA fraction is (3 / 100) / (0.10), or 0.3.

[0179] Next, consider an example in which two variants, a first variant and a second variant, are compared in the evaluation step of block 256. The support 146 for the first variant in the variant set 142 from the biological sample (e.g., blood) is compared to the reference frequency 132 for the same variant in the reference set 128 for the abnormal tissue. Similarly, the support 146 for the second variant in the variant set 142 from the biological sample is compared to the reference frequency 132 for the same variant in the reference set 128. It is assumed that the only source of the first and second variants in the cell-free nucleic acid is from the abnormal tissue. Therefore, under this assumption, a ratio for the first variant is calculated based on the support 146 for the first variant in the variant set 142 versus the reference frequency 132 for the first variant in the reference set. For example, if the support 146 for the first variant is 3 out of 100 sequence reads in the variant set 142, and the reference frequency 132 for the first variant in the reference set 128 is 0.10, then the ratio for the first variant is (3 / 100) / (0.10), or 0.3. Further, a ratio for the second variant is calculated based on the support 146 for the second variant in the variant set 142 versus the reference frequency 132 for the second variant in the reference set. For example, if the support 146 for the second variant is 5 out of 85 sequence reads in the variant set 142, and the reference frequency 132 for the first variant in the reference set 128 is 0.12, then the ratio for the second variant is (5 / 85) / (0.12), or 0.49.

[0180] In some embodiments, more than one variant is compared in the evaluation step of block 256, and for each such variant, a ratio is calculated between the observed support for each variant in the biological sample and the frequency of the same variant in the variant set. For example, in some embodiments, more than two variants are compared in the evaluation step of block 256. In a number of such embodiments, the above examples are expanded in that a ratio is calculated between the observed support for each variant in the biological sample and the frequency of the same variant in the variant set for each such variant. Indeed, in some embodiments, between 2 and 200 variants are compared in the comparison step of block 228. In some embodiments, more than 25, 50, 100, 200, 300, 400, 500, 1000, 2000, or 5000 variants are compared in the evaluation step of block 256.

[0181] Thus, a number of somatic variants k is observed from the first abnormal tissue sample, where k is a positive integer (e.g., 2, 3, greater than 20, greater than 100, greater than 200, etc.). This can be expressed as a number of variant frequencies (supporting the variant a) for each variant in the reference set. 1i The number of the plurality of sequence reads 126 mapped to the plurality of sequence reads 126d corresponding to the genomic location of the variant 1i A k-length vector f1=(f 11 ,f 12 ,…,f 1k ), where each component of f1 is f 1i It takes a value between 0 and 1. This forms the reference variant 128 .

[0182] Further, a plurality of sequence reads overlapping the k variants represented by the vector f1 are scanned from the biological sample comprising a plurality of cell-free nucleic acids from the subject. For each individual variant position i among the k variant positions, a plurality of sequence reads mapped to the genomic position corresponding to the variant position i is determined 140 (d 2i ) and the total number of said variants (a 2i ) The total number of paired sequence reads 140. The plurality of measurements d 2i and a 2i Is a non-negative integer value, from which a is obtained 2i Divide by d 2i The quotient f 2i For the individual quotients f for the plurality of variants using the entire reference set of sequence reads 140 measured from the biological sample comprising modules of cell-free nucleic acids from the subject, 2i It can be expressed as the k-length vector f2 = (f 21 ,f 22 ,…,f 2k ).

[0183] The goal is to determine a single estimated ctDNA for the subject from the observed frequency (support 146) of each variant 144 in the first set of variants 142 of the plurality of first sequence reads according to block 256. In other words, the goal is to determine the single estimated ctDNA using the fraction of mutation reads provided from the first abnormal tissue sample (e.g., tumor) to the biological sample (e.g., blood) comprising cell-free nucleic acid. The plurality of vectors f1 and f2 summarize the plurality of sequence read counts measured from the plurality of individual tissues (the first abnormal tissue and the biological sample comprising cell-free nucleic acid) from which an underlying rate is inferred. In some embodiments, variants that are clearly not associated with cancer are excluded from the analysis. In other words, they are excluded from the k variants considered.

[0184] In some embodiments, the plurality of sequence reads 126 from the abnormal tissue sample are assumed to be generated according to a Poisson process. For each variant i in k, there is an observed 2i The actual supporting sequence read counts, and f 1i Multiply by d 2i For example, for variant 1, consider an example where a 21 is 100 and d 21 is 1000, meaning that among the 1000 sequence reads 140 measured from the biological sample containing cell-free nucleic acid and overlapping with the genomic position corresponding to variant 1, 100 of the plurality of sequence reads 140 support the variant. Further assuming that, from the first abnormal tissue, the frequency (f 11 ) is 0.25. Therefore, it is expected that f 11 (0.25) multiplied by d 21 (1000) or 250 read counts. Thus, in some embodiments, a cumulative distribution function (binomial cumulative probability function) is estimated by estimating the data conditioned on t (the plurality of proportional mutant sequence reads are provided from the first abnormal tissue sample to the biological sample containing the cell-free nucleic acid) and D(t), thereby estimating a plurality of single estimated ctDNA fractions corresponding to the 5th, 50th (median), and 95th percentiles, or any other desired percentiles. The a for one of the k variants considered is observed in the cell-free DNA biological sample. 2i Further, the variant frequency f of the first abnormal tissue for the individual variant in the first abnormal tissue sample can be calculated. 1i with d2i =(t) / t ... 1i analysis) to (ii) the actual observed number (a) of multiple reads supporting said variant i in said tissue containing cell-free DNA. 2i ), and this can be used to estimate a cumulative distribution function (a probability distribution) that provides an estimate of the value of each trial value of t (where, in some embodiments, t is sampled from anywhere between 0 and 110 percent). Thus, with reference to Figure 16 , using the cumulative distribution function over a range of values for t to calculate a likelihood for the individual trial values of t for a particular allele i.

[0185] In some embodiments, for a single variant, the cumulative distribution function has a value ranging from 0 (zero probability) to 1 (100% probability) and has the form

[0186]

[0187] Where x = a 2i , the number of sequence reads from the biological sample that are paired with the genomic position corresponding to the variant i and support the variant allele at this position; p = t*f 1i , where t is the single estimated ctDNA fraction, and f 1i is (a) the ratio of the number of sequence reads from the first abnormal tissue that are paired with the genomic position corresponding to the variant i and support the variant allele at this position; and n=d 2i , a total number of sequence reads from the biological sample that map to the genomic position corresponding to the variant position i.

[0188] From this we can see that reference Figure 16, a median value for t (the most likely value for t) based on a probability distribution for t within the numerical range of 0% to 110% for t, a 5th percentile value for t (the lowest value for t, the lowest critical value for t) based on a probability distribution for t within the numerical range of 0% to 110% for t (1604), and a 95th percentile value for t (the highest value for t, the highest critical value for t) based on a probability distribution for t within the numerical range of 0% to 110% for t (1606) can be calculated. Figure 16 In FIG, solid line 1610 represents a cumulative density function, and line 1608 represents the cumulative distribution function. In some embodiments, the cumulative distribution function is used to calculate multiple percentile values for t. The 95th percentile value means that it is extremely rare for an observed fraction of sequence reads supporting a variant allele over the total number of sequence reads overlapping the allele position to exceed the 95th percentile for t, and 95% of the product (time) values for t are expected to be less than the 95th percentile for t (in Figure 16 about 28% of the total population).

[0189] Other cutoffs may be used, for example, the 2nd percentile and the 98th percentile.

[0190] The above discussion relates to how to calculate t from a single variant. However, as discussed herein, in more common embodiments, multiple variants are sampled so that each variant produces an independent likelihood (probability of t) within the range of values considered for t (e.g., 0 to 100%). Thus, the cumulative distribution function provides: a first probability of t based on the observed and expected values for variant 1 at a particular trial value of t, a second probability of t based on the observed and expected values for variant 2 at the particular trial value of t, and so on. To arrive at the cumulative likelihood of t at the particular trial value of t, each of the multiple component probabilities (the first probability of t based on the observed and expected values for variant 1 at the particular trial value of t, the second probability of t based on the observed and expected values for variant 2 at the particular trial value of t, and so on) are combined and used to calculate the cumulative distribution function. In other words, using data from any number of variants, a cumulative probability of t can be drawn. Figure 16The cumulative distribution function 1608 of k is based on the assumption that they are multiple independent observations of the same underlying single estimated ctDNA score. In some embodiments, when expressing the multiple probabilities in a logarithmic space to arrive at the calculated probability for the trial value of t, the multiple probabilities provided by each individual variant in the set of k variants for a particular trial value of t are combined by adding together. For example:

[0191]

[0192] where k refers to the kth allele and the sum is for all k variants. Alternatively, in some embodiments, when expressing the multiple probabilities on a natural scale to arrive at the calculated probability for the trial value of t, the multiple probabilities provided by each individual variant in the set of k variants for a particular trial value of t are combined by multiplying together.

[0193] In some embodiments, for each variant k, the Poisson model of the likelihood of t over the experimental range of t is calculated individually, thereby calculating multiple Poisson models, one for each variant. Then, for each experimental value of t sampled, the multiple Poisson models are combined (e.g., summed in log space, or multiplied if on the natural scale) to obtain a likelihood of the experimental value of t for each experimental value of t sampled. Thus, each point on line 1608 is summarized into the k variants, where k is a positive integer (e.g., 2 or more, 20 or more, 1000 or more). In this way, the most concise interpretation of the tumor score is provided.

[0194] In some embodiments, the single estimated ctDNA fraction is obtained according to a median for t obtained from a distribution of multiple possibilities for t within the sampled range of values of t using the cumulative density function.

[0195] Importantly, in instances where zero supporting reads 140 are observed in the biological sample for the k variants, this framework is able to estimate multiple confidence intervals based on a single estimated ctDNA score.

[0196] Thus, the tumor fraction of the cell-free DNA is estimated conditioned on the readout information for the set of variants between (i) the biological sample containing the cell-free nucleic acid and (ii) the first abnormal tissue sample. Thus, in this embodiment, only those variants represented in the reference set of variants 128 and the set of variants 142 for the biological sample are used to calculate the single estimated ctDNA fraction for the subject.

[0197] In alternative embodiments, a negative binomial distribution is assumed instead of a Poisson distribution to calculate Figure 16 The cumulative distribution function 1608 of .

[0198] In some embodiments, multiple observed sequence reads are corrected for background copy number. For example, multiple sequence reads supporting multiple variants of multiple chromosomes or multiple parts of multiple chromosomes that are repeated in the subject are corrected for this duplication. This can be done by normalizing before running this inference, or allowing the numerical value of more than one ctDNA score. Allowing more than one ctDNA score can also assess the heterogeneity within / across tumors. Therefore, in some embodiments, the assumption that each variant represents an independent observation of the single estimated ctDNA score is corrected for background copy number.

[0199] As a reference Figure 3 As another example, as discussed above, in some embodiments, the single expected ctDNA fraction in the cfDNA is between 0.5x10 -4 to 1.5x10 -4 In some embodiments, the single expected ctDNA fraction in the cfDNA is between 0.5 x 10 -3 Up to 1x10 -2 and the first condition is renal cancer, uterine cancer, thyroid cancer, prostate cancer, breast cancer, bladder cancer, gastric cancer, cervical cancer, or a combination thereof. In some embodiments, the single expected ctDNA fraction in the cfDNA is between 1x10 -2 to 0.8, and the first condition is lung cancer, esophageal cancer, head and neck cancer, colorectal cancer, anorectal cancer, ovarian cancer, hepatobiliary cancer, pancreatic cancer, or lymphoma.

[0200] In some embodiments, when the observed frequency (support 14) of each variant 144 in the first variant set 142 meets a first threshold, the subject is classified by considering the subject as having a first condition. In some embodiments, the first threshold is determined based on quantification of the reference frequencies of the multiple variants in the variant set. In some embodiments, for example, the observed frequency (support 14) of each variant 144 in the first variant set 142 is normalized by the reference frequencies for the multiple corresponding variants in the variant set as discussed above with reference to block 258 to achieve a circulating tumor nucleic acid score for the subject. For example, in some embodiments, the observed frequency (support 14) of each variant 144 in the first variant set 142 is divided by the reference frequencies for the multiple corresponding variants in the variant set as discussed above with reference to block 258 to achieve the circulating tumor nucleic acid score for the subject. In this way, the first threshold is determined by a frequency of each variant 130 in the first reference set 128 of the first abnormal tissue sample.

[0201] In some embodiments, a population of subjects with a similar condition is used to improve the first threshold associated with a condition. For example, consider an example where the first condition is a stage of cancer and is not related to the type of cancer. Figure 4 Describe the dropout rate (ctDNA fraction) within a subject population. Figure 4 Each point in represents the ctDNA score for a different subject in a population of subjects divided into one of four cancer stages (I, II, III, and IV). For each individual subject, the ctDNA score (tumor score) is plotted based on the ratio of the support for the set of variants in the set variants 142 collected from a biological sample of the subject (e.g., according to blocks 202 and 210 of FIG. 2 ) to the reference frequency 132 of these same variants in the reference set 128 obtained from a tumor from the same subject for the individual subject (e.g., according to the disclosure outlined in block 228 of FIG. 2 ). Figure 4 It is shown that there is a range of ctDNA scores for each cancer subject, but the median ctDNA score generally increases with the stage of cancer. Figure 4 Motivation is provided for determining the first threshold based on quantification of the reference frequencies of the plurality of variants in the variant set. That is, Figure 4The possibility of using the multiple observed frequencies of multiple variants in the abnormal tissue of a particular subject, and optionally information about expected ctDNA of multiple subjects with a particular cancer stage or type, is described to determine a first threshold value for the subject with the particular cancer that can be evaluated against the observed frequencies of the multiple variants in a set of variants in a biological sample for the particular subject to classify the subject as having or not having the condition (e.g., a clinical stage of a particular cancer). Thus, reference is made to Figure 4 A first threshold of 0.05 can be used to analyze whether a subject has stage I of a particular cancer. In such an example, an abnormal tissue, such as a tumor, is obtained from a subject and used to determine a reference frequency for each individual variant in a first reference set (e.g., as per block 228 of FIG. 2 ). In some embodiments, the frequencies of each possible variant are used to identify the plurality of variants for the reference set. Further, cell-free nucleic acid is obtained from a biological sample from the same subject, excluding the abnormal tissue (e.g., as per block 202 ), and the variant frequencies of the plurality of identical variants in the reference set are determined from a plurality of sequence reads of the cell-free nucleic acid in the biological sample (e.g., as per block 210 ). The variant frequencies of these variants in the biological sample (support 146 ) are normalized by the reference frequencies of the plurality of identical variants in the abnormal tissue (e.g., by obtaining a ratio, etc.), thereby forming the observed ctDNA score for the biological sample (e.g., as per block 258 of FIG. 2 ). Here, the first threshold is determined by a frequency of each variant 130 in the first reference set 128 of the first abnormal tissue sample, as these frequencies form the basis of the denominator of the ratio as discussed above in connection with block 258 of FIG. 2 . Figure 4 As a guide, determining whether the ctDNA of a particular biological sample meets the threshold condition of 0.05 provides a basis for determining whether the subject has stage I cancer in this example. Figure 4 , the subject is considered to have a more advanced cancer when the comparison of the observed frequency (support 146) of each variant 144 in the variant set with the reference frequencies of the same variants in the reference set 128 indicates that the ctDNA score is above 0.05, because Figure 4 Few stage I patients in the population had a ctDNA score above 0.05. On the other hand, observing a ctDNA score below 0.001 was consistent with finding that the subject had stage I of a particular cancer because Figure 4Relatively few subjects in the population with stage II, III, or IV disease have a ctDNA score below 0.001. This is merely an example, and is discussed in more detail in Example 1 below. Figure 3 It was shown that more precise thresholds can be defined when the type of cancer a subject suffers from is known.

[0202] Block 260 provides a specific embodiment, wherein the evaluating step of block 256 includes calculating a single estimated circulating tumor DNA (ctDNA) fraction in the cell-free DNA (cfDNA) of the subject from the observed frequency (support 146) of each variant 144 in the first set of variants 142, wherein the single estimated circulating tumor DNA (ctDNA) fraction exceeds 1×10 -3 , and the first condition is stage II, stage III, or stage IV breast cancer, the observed frequency of each first variant 144 in the first variant set 142 meets a threshold. This threshold is limited by Figure 5 , as discussed in Example 2 below. Figure 5 , each point is the ctDNA score for a separate subject with breast cancer. In some embodiments, the method for calculating the cfDNA score for each subject includes obtaining a plurality of first sequence reads 140 in electronic form from a biological sample of each subject in a population, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules. The plurality of first sequence reads 140 are used to identify support for each variant in a variant set 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the variant set 142. In some embodiments, the observed frequency (support 146) of each individual variant 144 in the variant set 142 is compared with a corresponding reference frequency 132 for the individual variant in a reference set 128. Each corresponding reference frequency 132 in the reference set 128 is a frequency of another variant in a first abnormal tissue sample obtained from the subject. In this way, in some embodiments, the ctDNA score for each subject is determined. In addition to plotting the ctDNA score for each subject, Figure 5 The subjects were divided into groups according to the stage of breast cancer. For each tumor stage, the tumor score observed was Figure 5 Point out a great dynamic range. Figure 5 If the circulating tumor DNA (ctDNA) fraction exceeds 1x10 -3 , then the subject may have breast cancer stage II, III or IV, because Figure 5Few subjects with stage 0 or stage 1 breast cancer had more than 1x10 -3 Of course, due to Figure 5 It also showed that a large number of Phase III subjects had less than 1x10 -3 Therefore, additional tests may be required to determine a definitive classification of a breast cancer subject. Thus, the disclosed methods support examples where the subject has stage II, stage III, or stage IV breast cancer and the evaluation step of block 256 determines that the first tumor fraction of the cell-free nucleic acid is less than 1x10 -3 .

[0203] With reference to block 262, in some embodiments, the disclosed methods are used to assess a tumor score in a subject having cancer from a common primary site. For example, with reference to block 264, in some embodiments, the disclosed methods are used to assess a tumor score in a subject having breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0204] refer to Figure 2E At block 268, in some embodiments, the plurality of disclosed methods are used to assess a tumor score in a subject having a predetermined stage of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer.

[0205] refer to Figure 2E Block 270, in some embodiments, the disclosed methods are used to assess a tumor score in a subject having a predetermined subtype of a cancer. In some such embodiments, reference is made to Figure 2E Block 272, wherein the cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer.

[0206] refer to Figure 2E and 2FIn blocks 274 to 286, the disclosed methods are not limited to analyzing a single abnormal tissue, or analyzing a single abnormal tissue at a single time point. In some embodiments, the disclosure of blocks 202 to 272 is extended to multiple tumor samples and multiple tumor scores from a patient corresponding to inter- / intra-tumor heterogeneity. In other words, the disclosed methods can be used to calculate an additional ctDNA score for a second biological sample relative to a second abnormal tissue. For example, in some embodiments, the first abnormal tissue sample discussed above in conjunction with blocks 202 to 272 of FIG. 2 is of a first cancer type, and the second abnormal tissue sample is of a second cancer type (block 278). In other embodiments, the first abnormal tissue sample discussed above in conjunction with blocks 202 to 272 of FIG. 2 is from a tumor at a first time point, and the second abnormal tissue sample is from the same tumor at a second time point. In yet other embodiments, the abnormal tissue of the subject is heterogeneous, and the first abnormal tissue sample is a first section of the abnormal tissue, and the second abnormal tissue sample is a second section of the same abnormal tissue collected simultaneously with the first section.

[0207] In more detail, referring to block 274, the plurality of first sequence reads 140 are used to identify support for each variant 144 in a second set of variants 142, thereby determining an observed frequency for each variant 144 in the second set of variants. For each individual variant 144 in the second set of variants 142, a corresponding reference frequency 132 for the individual variant is obtained in a second reference set 128, wherein each corresponding reference frequency in the second reference set is for another variant in a second abnormal solid tissue sample obtained from the subject. In some such embodiments, the evaluating step of block 256 further includes using the observed frequency of each individual variant in the second set of variants against the observed frequency of the individual variant in the second reference set to determine a second tumor fraction in the cell-free nucleic acid of the liquid biological sample from the subject. In this way, the ctDNA score of a biological sample can be first calculated relative to the first abnormal tissue (e.g., to determine whether the subject has a first condition, to monitor the progression of the first abnormal tissue over a period of time, to monitor tumor heterogeneity, etc.), and a different ctDNA score of the biological sample can be calculated relative to the second abnormal tissue (e.g., to determine whether the subject has a second condition, to monitor the progression of the second abnormal tissue over a period of time, to monitor tumor heterogeneity, etc.).

[0208] refer to Figure 2FIn block 276, in some embodiments, another one of the plurality of first sequence reads 140 is considered to support a variant 144 in the second variant 142 when the individual sequence read 140 (i) maps to a genomic location corresponding to the variant and (ii) contains all or a portion of the variant. Another one of the plurality of first sequence reads 140 is considered to not support a variant 144 in the second variant 142 when the individual sequence read 140 (i) maps to a genomic location corresponding to the variant and (ii) does not contain all or a portion of the variant.

[0209] refer to Figure 2F At block 278, in some embodiments, the first abnormal tissue sample is comprised of a first tumor fraction and the second abnormal tissue sample is comprised of a second tumor fraction from a common (same) tumor of the subject.

[0210] refer to Figure 2F At block 280, in some embodiments, the first abnormal tissue sample is a first cancer type and the second abnormal tissue sample is a second cancer type. The first cancer type may be the same as the second cancer type (block 282). Alternatively, the first cancer type may be different from the second cancer type (block 284). In some embodiments, the first cancer type and the second cancer type are each selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer (block 286).

[0211] Exemplary Method Example - Assessing the aggressiveness of a known condition in a subject based on the variation in the ctDNA fraction in cfDNA over time.

[0212] Another aspect of the present disclosure provides a method for assessing a status of a condition in a subject. The method includes: in a computer system 100 having one or more processors 102 and a memory 111 / 112 storing one or more programs executed by the one or more processors, performing the steps of: for each individual time point within a period of time, obtaining a corresponding dataset 138 in electronic form from a biological sample of the subject obtained at each individual time point, the dataset 138 comprising a plurality of corresponding first sequence reads 140 for the individual biological sample, thereby obtaining a plurality of datasets for the subject (e.g., as described in block 202). Each individual biological sample comprises a plurality of cell-free nucleic acid molecules. In some embodiments, the plurality of cell-free nucleic acid molecules is obtained from a specific biological sample, as discussed above in connection with any of blocks 202 through 208 of FIG. 2 . In some embodiments, the plurality of sequence reads for the plurality of cell-free nucleic acid molecules of a specific biological sample is obtained, as discussed above in connection with any of blocks 202 through 208 of FIG. 2 .

[0213] The method further includes, for each individual dataset (e.g., data construct 138) in the plurality of individual datasets, determining support for each variant 144 in the variant set 142 (e.g., as disclosed in blocks 210 to 226 of FIG. 2 ). Considering an additional sequence read 140 in the plurality of first sequence reads in the individual dataset to support a variant 144 in the variant set 142 if the additional sequence read 140 (i) maps to a genomic location corresponding to the variant and (ii) contains all or a portion of the variant 144. Considering an additional sequence read 140 in the plurality of first sequence reads in the individual dataset to not support a variant in the variant set 142 if the additional sequence read 140 (i) maps to a genomic location corresponding to the variant and (ii) does not contain all or a portion of the variant 144. In this way, at each of the multiple time points, an observed frequency of each variant 144 in the variant set 142 is determined using the multiple sequence reads in the multiple first sequence reads in the individual data sets 138 that support or do not support each variant in the variant set 142.

[0214] In some embodiments, the plurality of sequence reads 140 are used to find support for a plurality of variants 144 in the set of variants 142 by utilizing a B-score classifier to identify a plurality of variants using the plurality of sequence reads 140. The B-score classifier is described in U.S. Patent Publication No. 62 / 642,461, entitled “Methods and Systems for Selecting, Managing, and Analyzing High-Dimensional Data,” filed on March 13, 2018, which is incorporated herein by reference, and is further described in detail in Example 3.

[0215] In some embodiments, the plurality of sequence reads 140 are used to find support for a plurality of variants 144 in the set of variants 142 by utilizing an M-score classifier to identify a plurality of variants using the plurality of sequence reads 140. The M-score classifier is described in U.S. patent application Ser. No. 62 / 642,480, entitled “Methylated Segment Outlier Detection,” filed on March 13, 2018, which is incorporated herein by reference.

[0216] In some embodiments, the plurality of sequence reads 140 are used to find support for the plurality of variants 144 in the variant set 142 by identifying a plurality of variations using the plurality of techniques disclosed in any one of blocks 210 to 216 described above in conjunction with FIG. 2 .

[0217] The method further includes evaluating the observed frequency (e.g., support 146) of each variant 144 in the set of variants 142 against the observed frequency of the individual variant in the first abnormal solid tissue (e.g., as determined in the first example of block 210) at each of the plurality of time points, thereby determining the status or progression of a disease condition in the subject during the time period in the form of an increase or decrease in the first tumor score over the time period.

[0218] In some embodiments, the time period is calibrated to enable measurement of changes in ctDNA over a period of approximately hours (e.g., to measure the success rate of a surgical procedure to remove abnormal tissue from a subject), weeks / months (e.g., to monitor the success rate of a treatment for a subject), or years (e.g., to monitor disease remission in a subject). Thus, in some embodiments, the time period is a period of several months, and each of the multiple time points is a different time point in the period of several months. In some such embodiments, the period of several months is less than four months. In some embodiments, the time period is a period of several years, and each of the multiple time points is a different time point in the period of several years. In some such embodiments, the period of several years is between 2 and 10 years. In some embodiments, the time period is a period of several hours, and each of the multiple time points is a different time point in the period of several hours. In some such embodiments, the period of several hours is between 1 hour and 6 hours.

[0219] In some embodiments, the step of evaluating the observed frequency of each variant 144 in the set of variants 142 at each of the plurality of time points against the observed frequency of the individual variants in the first abnormal solid tissue comprises calculating, from the observed frequency of each variant 144 in the set of variants 142 at each time point in the plurality of time points, a single estimated circulating tumor DNA (ctDNA) fraction in the subject's cell-free DNA (cfDNA) in the manner described above in conjunction with block 256. In some such embodiments, the method further comprises changing a diagnosis of the subject when the single estimated circulating tumor ctDNA fraction in the subject's cfDNA is observed to change by a threshold amount during the time period. For example, in some embodiments, the ctDNA fraction at each time point during the time period is a number between 0 and 1, and the diagnosis of the subject is changed when the ctDNA fraction changes by a predetermined value during the time period. In one example, the diagnosis of the subject is downgraded when the ctDNA score increases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent during the time period, indicating that the subject has a more aggressive form of the disease condition and / or a later stage of the disease condition compared to the initial diagnosis. In another example, the diagnosis of the subject is upgraded when the ctDNA score decreases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent during the time period, indicating that the subject has a less aggressive form of the disease condition and / or an earlier stage of the disease condition compared to the initial diagnosis.

[0220] In some embodiments, the method further comprises: changing a prognostic profile of the subject when each of the single estimated ctDNA fractions in the subject's cfDNA is observed to change by a threshold amount during the time period. For example, in some embodiments, the ctDNA fraction at each time point during the time period is a number between 0 and 1, and the prognostic profile of the subject is changed when the ctDNA fraction changes by a predetermined value during the time period. In one example, when the ctDNA fraction increases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent during the time period, the prognostic profile of the subject is downgraded, indicating that the subject has a decreased likelihood of recovery from the disease condition. In another example, when the ctDNA fraction decreases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent during the time period, the prognostic profile of the subject is upgraded, indicating that the subject has an increased likelihood of recovery from the disease condition.

[0221] In some embodiments, the method further comprises: changing a treatment for the subject when the individual estimated ctDNA fractions in the subject's cfDNA are observed to change by a threshold amount over the period. For example, in one example, the treatment regimen for the subject is changed to a more aggressive treatment when the ctDNA fraction increases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent over the period. In another example, the treatment regimen for the subject is changed to a less aggressive treatment when the ctDNA fraction decreases by more than 2 percent, more than 3 percent, more than 4 percent, more than 5 percent, more than 10 percent, or more than 20 percent over the period.

[0222] In some embodiments, the condition is a disease, such as cancer. For example, in some embodiments, the disease is a cancer, and the cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0223] In some embodiments, the condition is a stage of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer.

[0224] In some embodiments, the disease condition is a predetermined subtype of cancer, wherein the cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer.

[0225] In some embodiments, each individual variant in the set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the number of copies of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an aberrant methylation pattern associated with a predetermined genomic location.

[0226] In some embodiments, the abnormal tissue is a tumor. In some embodiments, the first abnormal tissue sample is one of the plurality of abnormal tissues described above with reference to block 230 of FIG. 2 .

[0227] In some embodiments, the variant set 142 is comprised of a single variant 144, which is a single genetic variation located at a single site in the subject's genome. In some embodiments, the variant set 142 is comprised of a first variant and a second variant, wherein the first variant is a first genetic variation located at a first site in the subject's genome, and the second variant is a second genetic variation located at a second site in the subject's genome.

[0228] In some embodiments, the variant set 142 consists of a first variant, a second variant, and a third variant, wherein the first variant is a first genetic variation located at a first site in the genome of the subject, the second variant is a second genetic variation located at a second site in the genome of the subject, and the third variant is a third genetic variation located at a third site in the genome of the subject.

[0229] In some embodiments, the variant set 142 is composed of between 2 and 20 variants 144, wherein each variant 144 in the variant set is a different genetic variation in the genome of the subject (optionally located at a different site). In some embodiments, the variant set 142 includes 30 variants 144, 50 variants 144, 75 variants 144, 100 variants 144, 125 variants 144, 250 variants 144, 500 variants 144, 750 variants 144, 1000 variants 144, 2500 variants 144, or 5000 variants 144, wherein each variant 144 in the variant set is a different genetic variation in the genome of the subject (optionally located at a different site).

[0230] In some embodiments, for each individual dataset in the plurality of individual datasets, determining support for each variant 144 in a variant set 142 comprises aligning a sequence read 140 in the plurality of first sequence reads in another dataset to a region in a reference genome to determine whether the sequence read contains all or a portion of a variant in the variant set. For example, see Figure 2A Block 212 and the same disclosure presented above.

[0231] In some embodiments, for each individual dataset in the plurality of individual datasets, determining support for each variant 144 in a variant set 142 comprises comparing a sequence read 140 from the plurality of first sequence reads of another dataset to a lookup table of variants to determine whether the sequence read contains all or a portion of a variant in the variant dataset. For example, see Figure 2A Block 214 and the same disclosure presented above.

[0232] In some embodiments, for each individual dataset in the plurality of individual datasets, determining support for each variant 144 in a variant set 142 comprises aligning a sequence read 140 in the plurality of first sequence reads in one of the other datasets to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a reference genome. For example, see Figure 2A Block 216 and the same disclosure presented above.

[0233] In some embodiments, the subject is a human subject. In some embodiments, the subject is a mammal. In some embodiments, the subject is any of the species disclosed above in connection with block 204 of FIG. 2 .

[0234] In some embodiments, the individual biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid of the subject. That is, the biological sample is a mixture of the subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid and other components of the subject.

[0235] In some embodiments, the individual biological sample is composed of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid of the subject. That is, the biological sample is the blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid of the subject, and is free of other components of the subject.

[0236] Exemplary Method Example - Use of Multiple Results of a Classifier Using Tumor Scores

[0237] Another aspect of the present disclosure provides a method for classifying a subject. The method includes: in a computer system 100 having one or more processors 102 and a memory 111 / 112, the memory 111 / 112 storing one or more programs (e.g., a condition monitoring module 120) executed by the one or more processors, obtaining a dataset (e.g., a data construct 138) in electronic form in the computer system, the dataset comprising a plurality of first sequence reads 140 from a biological sample of the subject. Here, the biological sample comprises a plurality of cell-free nucleic acid molecules. In some embodiments, the plurality of first sequence reads are obtained in any manner combined with blocks 202 to 208. And, in many such embodiments, the plurality of first sequence reads 140 are used to identify support for each variant 144 in a first variant set 142, thereby determining an observed frequency for each variant in the first variant set in the manner disclosed above with reference to any of blocks 210 to 226 disclosed above in conjunction with FIG. 2 . Furthermore, for each individual variant 144 in the first variant set 132, a corresponding reference frequency for the individual variant in a first reference set 128 is obtained in the manner disclosed above with reference to blocks 228 through 248 of FIG2 , wherein each corresponding reference frequency in the first reference set is for another variant in a first abnormal solid tissue sample obtained from the subject. The method further discloses: evaluating the observed frequency of each individual variant in the first variant set 142 against the observed frequency of the individual variant in the first reference set 128 of the first abnormal solid tissue, thereby determining a first tumor fraction in the cell-free nucleic acid of the liquid biological sample from the subject in any of the manners disclosed above with reference to blocks 256 through 272 of FIG2 .

[0238] The method further includes applying the plurality of first sequence reads (or dimensionally reduced data from the plurality of sequence reads, such as a plurality of principal components) to a classifier to obtain a classifier result. The result of the classifier indicates whether the subject has a first cancer condition. Further, before performing the instant method, the classifier is trained using data other than the tumor score observed in the cell-free DNA (cfDNA) of the plurality of subjects. In some embodiments, when the first tumor score is between 0.003 and 1.0 and the result of the trained classifier indicates that the subject has the first cancer condition, the result of the trained classifier is used as a basis for the diagnosis or prognosis of the subject for the first cancer condition. As used herein, the term "trained classifier" refers to a model (e.g., a machine learning algorithm, such as logistic regression, neural network, regression, support vector machine, clustering algorithm, decision tree, etc.) with multiple fixed (locked) parameters (weights) and thresholds that can be prepared for application to multiple samples that have not been seen before.

[0239] In some embodiments, the estimated tumor fraction in the cfDNA of the subject is determined using the techniques disclosed above and in the examples below with reference to FIG. 2 .

[0240] In some embodiments, the first condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof).

[0241] In some embodiments, the first condition is a subtype of cancer. In some such embodiments, the cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or stomach cancer.

[0242] In some embodiments, the estimated tumor score is between 0.003 and 1.0, and the first cancer condition is a tissue of origin of a cancer.

[0243] In some embodiments, calculating the estimated tumor fraction in the cfDNA includes using the dataset to identify support for each variant 144 in a set of variants 142, wherein an additional sequence read 140 in the plurality of first sequence reads is considered to support a variant 144 in the set of variants 142 if the additional sequence read 140 (i) maps to the portion of the genome corresponding to the variant and (ii) contains all or a portion of the variant 144, and an additional sequence read 140 in the plurality of first sequence reads is considered not to support a variant 144 in the set of variants 142 if the additional sequence read 140 (i) maps to the portion of the genome corresponding to the variant and (ii) does not contain the additional variant 144. In this manner, an observed frequency for each variant 144 in the set of variants 142 is determined from the plurality of sequence reads in the plurality of first sequence reads that support or do not support each variant in the set of variants.

[0244] In some embodiments, the plurality of sequence reads 140 are used to find support for a plurality of variants 144 in the set of variants 142 by utilizing a B-score classifier to identify a plurality of variants using the plurality of sequence reads 140. The B-score classifier is described in U.S. Patent Publication No. 62 / 642,461, entitled “Methods and Systems for Selecting, Managing, and Analyzing High-Dimensional Data,” filed as such, which is incorporated herein by reference, and is further described in detail in Example 3.

[0245] In some embodiments, the plurality of sequence reads 140 are used to find support for a plurality of variants 144 in the set of variants 142 by utilizing an M-score classifier to identify a plurality of variants using the plurality of sequence reads 140. The M-score classifier is described in U.S. patent application Ser. No. 62 / 642,480, entitled “Methylated Segment Outlier Detection,” filed on March 13, 2018, which is incorporated herein by reference.

[0246] In some embodiments, the plurality of sequence reads 140 are used to find support for the plurality of variants 144 in the variant set 142 by identifying a plurality of variations using the plurality of techniques disclosed in any one of blocks 210 to 216 described above in conjunction with FIG. 2 .

[0247] Further, in many such embodiments, a single estimated tumor fraction in the cfDNA of the subject is calculated from the observed frequency of each variant in the set of variants. For example, see the disclosure of block 258 of FIG. 2 for disclosure regarding calculating the single estimated tumor fraction in the cfDNA.

[0248] In some such embodiments, a variant in the set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the number of copies of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an aberrant methylation pattern associated with a predetermined genomic location.

[0249] In some embodiments, the abnormal tissue is all or part of a tumor. In some embodiments, the abnormal tissue sample is any of the abnormal tissues described above in conjunction with block 230 .

[0250] In some embodiments, the set of variants 142 consists of a single variant 144, which is a single genetic variation located at a single site in the genome of the subject.

[0251] In some embodiments, the variant set 142 consists of a first variant 144 and a second variant 144, wherein the first variant 144 is a first genetic variation located at a first site in the genome of the subject, and the second variant 144 is a second genetic variation located at a second site in the genome of the subject.

[0252] In some embodiments, the variant set 142 consists of a first variant 144, a second variant 144, and a third variant 144, wherein the first variant 144 is a first genetic variation located at a first site in the genome of the subject, the second variant 144 is a second genetic variation located at a second site in the genome of the subject, and the third variant is a third genetic variation located at a third site in the genome of the subject.

[0253] In some embodiments, the variant set 142 is comprised of between 2 and 20 variants, wherein each variant 144 in the variant set 142 is a different genetic variation in the genome of the subject (optionally located at a different site). In some embodiments, the variant set comprises 40 variants, 50 variants, 75 variants, 100 variants, 200 variants, 500 variants, 1000 variants, 2000 variants, or 5000 variants, and each variant in the variant set is a different genetic variation in the genome of the subject (optionally located at a different site).

[0254] In some embodiments, the single estimated tumor fraction in the cfDNA is between 0.5x10 -4 to 1.5x10 -4 In some embodiments, the single estimated tumor fraction in the cfDNA is between 0.5 x 10 -3 Up to 1x10 -2 and the first condition is renal cancer, uterine cancer, thyroid cancer, prostate cancer, breast cancer, bladder cancer, gastric cancer, cervical cancer, or a combination thereof. In some embodiments, the single estimated tumor fraction in the cfDNA is between 1×10 -2 to 0.8, and the first condition is lung cancer, esophageal cancer, head and neck cancer, colorectal cancer, anorectal cancer, ovarian cancer, hepatobiliary cancer, pancreatic cancer, lymphoma, or a combination thereof.

[0255] In some embodiments, the step of using the plurality of first sequence reads to identify support for each variant in a variant set includes aligning another sequence read 140 in the plurality of first sequence reads to a region in a reference genome to determine whether the individual sequence read 140 contains all or a portion of a variant in the variant set. For example, see Figure 2A Block 212 and the same disclosure presented above.

[0256] In some embodiments, using the plurality of first sequence reads to identify support for each variant 144 in a variant set 142 includes comparing another sequence read 140 in the plurality of first sequence reads to a lookup table of variants to determine whether the sequence read contains all or a portion of a variant in the variant set. For example, see Figure 2A Block 214 and the same disclosure presented above.

[0257] In some embodiments, using the plurality of first sequence reads to identify support for each variant 144 in a variant set 142 comprises aligning another sequence read 140 in the plurality of first sequence reads to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome. For example, see Figure 2A Block 216 and the same disclosure presented above.

[0258] In some embodiments, the subject is a human subject. In some embodiments, the subject is a mammal. In some embodiments, the subject is any of the species disclosed above in connection with block 204 of FIG. 2 .

[0259] In some embodiments, the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject. That is, the biological sample is a mixture of the subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid and one or more other components of the subject.

[0260] In some embodiments, the individual biological sample is composed of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject. That is, the biological sample is the blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, and / or peritoneal fluid of the subject, and is free of other components of the subject.

[0261] In some embodiments, the classifier uses a B-score classifier as described in US Patent Publication No. 62 / 642,461, filed as "Method and System for Selecting, Managing, and Analyzing High-Dimensional Data," which is incorporated herein by reference.

[0262] In some embodiments, the classifier uses the M-score classifier described in U.S. Patent Application No. 62 / 642,480, filed on March 13, 2018, entitled “Methylated Segment Abnormality Detection,” which is incorporated herein by reference.

[0263] In some embodiments, the classifier is a neural network or a conventional neural network. See Vincent et al., 2010, "Stacked denoising autoencoders: Learning useful representations in deep networks using local denoising criteria," J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, "Exploring strategies for training deep neural networks," J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Foundations of Artificial Neural Networks, MIT, each of which is incorporated herein by reference.

[0264] In some embodiments, the classifier is a support vector machine (SVM). Several SVMs are described in Cristianini and Shawe-Taylor, 2000, "Introduction to Support Vector Machines", Cambridge University Press, Cambridge; Boser et al., 1992, "A training algorithm for optimal margin classifiers", Proceedings of the Fifth Annual ACM Symposium on Computational Learning Theory, ACM Press, Pittsburgh, PA, pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, NY; Mount, 2001, Bioinformatics: Sequence and Genome Analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, 2nd ed., 2001, John Wiley & Sons, pp. 259, 262-265; and Hastie, 2001, Elements of Statistical Learning, Springer, NY; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a specific binary labeled data set with a hyperplane that maximizes the distance to the labeled data. For situations where linear separation is not possible, SVMs can be combined with "kernel" techniques, which automatically implement a nonlinear mapping to a feature space. The hyperplane in the feature space discovered by the SVM corresponds to a nonlinear decision boundary in the input space.

[0265] In some embodiments, the classifier is a decision tree. Decision trees are generally described in Duda, 2001, Pattern Classification, John Wiley & Sons, New York, pp. 395-396, which is incorporated herein by reference. Tree-based methods divide the feature space into a set of rectangles and then fit a model (such as a constant) in each rectangle. In some embodiments, the decision tree is a random forest regression. A specific algorithm that can be used is a classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and various random forests. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, New York, pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Various random forests are described in Breiman, 1999, "Random Forests—The Signature of Randomness," Technical Report 567, Department of Statistics, University of California, Berkeley, mid-September 1999, which is incorporated herein by reference in its entirety.

[0266] In some embodiments, the classifier is an unsupervised clustering model. In some embodiments, the classifier is a supervised clustering model. The performance of clustering is described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, New York, NY (hereinafter "Duda 1973"), pages 211 to 216, which is incorporated herein by reference in its entirety. As described in Section 6.7 of Duda 1973, the problem of clustering is described as one of finding natural groupings in a data set. In order to identify multiple natural groupings, two problems are solved. First, determine a method for measuring the similarity (or dissimilarity) between two samples. This measure (measure of similarity) is used to ensure that multiple samples in one cluster are more similar to each other than they are to multiple samples in other clusters. Second, determine a mechanism for using the similarity measure to divide the data into multiple clusters. Several measures of similarity are discussed in Section 6.7 of Duda 1973, which states that one way to begin a clustering investigation is to define a distance function and compute a distance matrix between all pairs of examples in a training set. If the distance is a good similarity measure, then the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. However, as stated on page 215 of Duda 1973, clustering can be performed without using a distance metric. For example, a non-metric similarity function s(x, x') can be used to compare two vectors x and x'. Conventionally, s(x, x') is a symmetric function that is large when x and x' are "similar" to some extent. An example of a non-metric similarity function s(x, x') is provided on page 218 of Duda 1973. Once a method has been chosen for measuring the "similarity" or "dissimilarity" between points in a data set, clustering requires a criterion function that measures the quality of a cluster for any partition of the data. Partitions of the data that are polarized by the criterion function are used to cluster the data. See Duda 1973, p. 217. Criterion functions are discussed in Duda 1973, section 6.8. More recently, Duda et al., Pattern Classification, 2nd ed., John Wiley & Sons, New York, NY, has been published. Pages 537-563 describe the clustering process in detail. More information on various clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster Analysis (3D Edition), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Inference in Cluster Analysis, Prentice Hall, Upper Saddle River, NJ, each of which is incorporated herein by reference.Specific exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using the nearest neighbor algorithm, the farthest neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum of squares algorithm), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. In some embodiments, the clustering includes unsupervised clustering, in which there is no preconceived notion of what clusters should be formed when clustering the training set.

[0267] In some embodiments, the classifier is a regression model, for example, a multi-class logit model as described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the classifier uses a regression model as disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York.

[0268] An Alternative Method for Determining Tumor Score Without Tumor Pairing The methods disclosed above in conjunction with FIG. 2 require the use of a reference set 128 from an abnormal tissue, such as a tumor tissue, from the subject. Another aspect of the present disclosure provides a method for determining tumor score in cell-free nucleic acid from a liquid biological sample from a subject that does not require pairing multiple allele frequencies to a corresponding tumor sample. This reference-free method includes obtaining a plurality of sequence reads in electronic form from the liquid biological sample from the subject in a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors, wherein the computer system obtains a plurality of sequence reads from the liquid biological sample from the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules. In some embodiments, any of the methods for obtaining sequence reads disclosed above in conjunction with blocks 202 through 208 of FIG. 2 is used.

[0269] The method further includes using the plurality of sequence reads to identify support for each variant in a set of variants, thereby determining an observed frequency for each variant in the first set of variants. In some embodiments, any of the methods disclosed above in connection with blocks 210 to 226 for using a plurality of sequence reads to identify support for each variant in a set of variants, thereby determining an observed frequency for each variant in the set of variants, is used.

[0270] The method further comprises considering the observed frequency of the variant having the Nth highest allele frequency in the set of variants as the tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject, where N is a positive integer other than 1 (e.g., 1, 2, 3, 4, 5, etc.). Figure 17 Provides a comparison of the tumor fraction estimated from tumor variant coverage with the reference-free tumor fraction estimated from cfDNA alone. Figure 17 The reference-free TF estimated from multiple small variants re-identified in the cfDNA of one aspect of the present disclosure, wherein N is set to 2 (y-axis) is compared with the TF (x-axis) estimated from the tumor mutation coverage in the cfDNA assessed by using the pairing method described above in conjunction with Figure 2. In order to estimate the reference-free TF, multiple somatic variants are identified again from multiple ART assay sequence reads of the CCGA population described in Example 12. After performing noise modeling, joint modeling using white blood cells (WBC), and artifact modeling of marginal variants, multiple variants are filtered, as disclosed in U.S. patent application No. 16 / 201,912, entitled "Model for Targeted Sequencing," filed on November 27, 2018, which is incorporated herein by reference. Furthermore, variant attribution is performed on multiple variants. For example, see U.S. patent application No. 16 / 201,912, entitled "Model for Targeted Sequencing," filed on November 27, 2018, which is incorporated herein by reference. Among variants identified as somatic and not attributable to WBC origin, the tumor fraction was estimated as the second-ranked variant allele frequency (af_max2). Figure 17 In

[15] , multiple results were faceted by whether tumor evidence (at least one tumor mutation read in cfDNA, true) was compared to no tumor evidence in cfDNA (false). Figure 17 The no-reference method of the instant aspect of the disclosure was shown to agree with the paired method of FIG. 2 down to a tumor fraction of approximately 1 / 1000 in multiple estimates for multiple samples with evidence of positive reads.

[0271] In some embodiments, a variant in the set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the copy number of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an abnormal methylation pattern associated with a predetermined genomic location.

[0272] In some embodiments, another of the multiple sequence reads is considered to support a first variant in the variant set when the individual sequence read contains all or a portion of the first variant, and another of the multiple sequence reads is considered to not support a first variant in the variant set when the individual sequence read does not contain the first variant, and a number of multiple sequence reads in the multiple sequence reads that support the first variant is paired with a number of multiple sequence reads in the multiple sequence reads that do not support the first variant to determine the observed frequency of the first variant, and the observed frequency of the first variant can estimate the variant frequency of the first variant in the liquid biological sample.

[0273] In some embodiments, the subject has a cancer originating from a single primary site. In some embodiments, the subject has a cancer originating from two or more different organs. In some embodiments, the subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0274] In some embodiments, the variant set comprises five or more variants, and each individual variant in the variant set is located at a different site in the genome of the subject. In some embodiments, the variant set consists of between 3 and 20 variants, and each variant in the variant set is a different genetic variation in the genome of the subject.

[0275] In some embodiments, the variant set consists of between 2 and 200 variants, and each variant in the variant set is a different genetic variation in the genome of the subject. In some embodiments, the variant set comprises 1000 variants, and each variant in the variant set is a different genetic variation in the genome of the subject.

[0276] In some embodiments, the step of using the multiple sequence reads to identify support for each variant in a variant set includes: aligning a sequence read from the multiple sequence reads with a region in a reference genome to determine whether the sequence read contains all or a portion of a first variant.

[0277] In some embodiments, the step of using the plurality of sequence reads to identify support for each variant in a set of variants includes comparing a sequence read from the plurality of sequence reads to a lookup table of multiple variants to determine whether the sequence read contains all or a portion of a first variant.

[0278] In some embodiments, the step of using the plurality of sequence reads to identify support for each variant in a set of variants includes: aligning a sequence read from the plurality of first sequence reads with each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome.

[0279] In some embodiments, the liquid biological sample comprises or consists of blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0280] In some embodiments, the method further comprises: repeating the step of obtaining a plurality of sequence reads at each respective time point in a plurality of time points over a period of time from a separate biological sample of the subject taken at each respective time point, wherein the respective biological sample comprises a plurality of cell-free nucleic acid molecules, thereby obtaining a corresponding plurality of sequence reads for the subject at each respective time point; and determining support for the variant having the Nth highest allele frequency in the set of variants of the original calling step for each respective time point, thereby determining the status or progression of a condition in the subject during the period of time in the form of an increase or decrease in the allele frequency of the variant over the period of time.

[0281] In some embodiments, the period is a period of several months (e.g., between 1 month and 4 months), and each of the multiple time points is a different time point in the period of several months. In some embodiments, the period is a period of several years (e.g., between 2 and 10 years), and each of the multiple time points is a different time point in the period of several years. In some embodiments, the period is a period of several hours (e.g., between 1 hour and 6 hours), and each of the multiple time points is a different time point in the period of several hours.

[0282] In some embodiments, the method further comprises changing a diagnosis of the subject when the allele frequency of the variant is observed to change by a threshold amount within the time period (e.g., such as a change of 10 percent, 20 percent, 30 percent relative to a reference amount at the first measurement time point).

[0283] In some embodiments, the method further comprises changing a prognosis of the subject when the allele frequency of the variant is observed to change by a threshold amount during the period (e.g., a change of 10%, 20%, or 30% relative to a reference amount at the first measurement time point).

[0284] In some embodiments, the method further comprises changing a treatment of the subject when the allele frequency of the variant is observed to change by a threshold amount within the time period (e.g., such as a change of 10 percent, 20 percent, 30 percent relative to a reference amount at the first measurement time point).

[0285] In some embodiments, the condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some embodiments, the condition is a stage of a cancer (e.g., a stage of breast cancer, a stage of lung cancer, a stage of prostate cancer, a stage of colorectal cancer, a stage of kidney cancer, a stage of uterine cancer, a stage of pancreatic cancer, a stage of esophageal cancer, a stage of lymphoma, a stage of head and neck cancer, a stage of ovarian cancer, a stage of hepatobiliary cancer, a stage of melanoma, a stage of cervical cancer, a stage of multiple myeloma, a stage of leukemia, a stage of thyroid cancer, a stage of bladder cancer, or a stage of gastric cancer). In some embodiments, the condition is a predetermined subtype of a cancer.

[0286] In some embodiments, the method further comprises: applying the plurality of sequence reads to a trained classifier to obtain a classifier result, wherein the trained classifier result indicates whether the subject has a first cancer condition; and when the tumor score is between 0.003 and 1.0 and the trained classifier result indicates that the subject has the first cancer condition, using the trained classifier result as a basis for diagnosing the subject with the first cancer condition. In some such embodiments, the first cancer condition is a cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof). In some such embodiments, the first cancer condition is a subtype of a cancer (e.g., a subtype of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer). In some such embodiments, the first tumor score is between 0.003 and 1.0, and the first cancer condition is a tissue of origin of a cancer. In some embodiments, the trained classifier is a neural network, a support vector machine, a decision tree, an unsupervised clustering model, a supervised clustering model, or a regression model.

[0287] Example 1

[0288] Median increased ctDNA scores according to cancer stage

[0289] refer to Figure 4 , regardless of the type of cancer the multiple subjects have, the multiple subjects are grouped by cancer stage I, II, III and IV. Figure 4In the figure, the x-axis indicates the cancer stage of each subject, and the y-axis indicates the observed ctDNA fraction for each subject. The method for calculating the cfDNA fraction for each subject includes obtaining a plurality of first sequence reads 140 in electronic form from a biological sample from each subject in a population, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules. The plurality of first sequence reads 140 are used to identify support for each variant 144 in a variant set 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the variant set 142. The observed frequency (support 146) for each individual variant 144 in the variant set 142 is compared to a corresponding reference frequency 132 for the individual variant in a reference set 128. Each corresponding reference frequency 132 in the reference set 128 is a frequency of a different variant in a first abnormal tissue sample obtained from the subject. Figure 4 Subjects without a plurality of positive reads are not included, meaning that the subjects do not have a plurality of sequence reads 140 supporting the plurality of variants observed in the paired reference set for such subjects.

[0290] In the example where the variant set consists of a single variant, the step of comparing the observed frequency of each individual variant 144 in the variant set 142 with a corresponding reference frequency 132 for the individual variant in a reference set 128 includes: obtaining a ratio of the frequency of the variant in the variant set (obtained from multiple sequence reads of cfDNA in the biological sample) to the frequency of the same variant in the reference set (obtained from multiple sequence reads of DNA in the abnormal tissue).

[0291] In the example where the variant set consists of two or more variants, comparing the observed frequency of each individual variant 144 in the variant set 142 with a corresponding reference frequency 132 for the individual variant in a reference set 128 includes taking a ratio of the frequency of each individual variant in the variant set (obtained from a plurality of sequence reads of cfDNA from the biological sample) to the frequency of the same variant in the reference set (obtained from a plurality of sequence reads of DNA from the abnormal tissue), i.e., there is a one-to-one correspondence between the variants in the reference set and the variants in the variant set.

[0292] therefore, Figure 4An analysis is provided of how ctDNA fraction varies with cancer stage in multiple subjects having multiple cell-free sequence reads supporting their underlying cancer, regardless of cancer type. Figure 4 There was evidence that higher tumor fractions (larger ctDNA fractions) were found in the cfDNA as the disease was more severe according to clinical stage (stages 1 to 4). Figure 4 This appears to be a general situation for the entire CCGA population (see Example 12 for more details on the CCGA population), but there are several violations (outliers) to this trend. Figure 4 Multiple such outliers in are suggestive and are best explained by clinical misclassification. Figure 4 A base component of the underlying disease is shown, which is the normally expected tumor fraction ratio in the cfDNA. Figure 4 It also shows that period 4 has some individuals with very low dropout rates, indicating that there are different sub-states within period 4.

[0293] Figure 4 Illustrating the multiple observed frequencies of the multiple variants from the reference set, multiple off-rates (ctDNA scores) can be used as a basis for establishing multiple meaningful and useful thresholds. That is, for example, given multiple specific frequencies of multiple variants in the abnormal tissue of a specific subject, and optionally information about the expected ctDNA scores for multiple subjects at a specific stage of cancer, a threshold for the specific subject can be determined and evaluated against the observed frequencies of the multiple variants in a variant set for the specific subject in order to classify the subject as having or not having the condition (e.g., a clinical stage of a specific cancer). For example, reference Figure 4, a threshold of 0.05 can be used to analyze whether a subject has stage I of a particular cancer. In this example, an abnormal tissue, such as a tumor, can be obtained from a patient and used to determine a reference frequency for each individual variant in a first reference set. In fact, in some embodiments, the frequencies of various possible variants can be used to define the multiple variants of the variant set. Next, cell-free nucleic acid is obtained from a biological sample other than the abnormal tissue from the same subject, and the variant frequencies of multiple identical variants in the reference set are determined from multiple sequence reads of the cell-free nucleic acid, thereby forming the observed ctDNA frequency for each individual variant in the first variant set. Then, a comparison of the ctDNA score used to determine whether the threshold condition of 0.05 is met with the multiple reference frequencies provides a basis for determining whether the subject has stage I cancer. For example, if the comparison indicates that the ctDNA score is above 0.05, it indicates that the subject has a more advanced stage of cancer. In another aspect, observing a ctDNA score formed from the observed frequency of each individual variant in the first set of variants of less than 0.001 is consistent with finding that the subject has stage 1 of a particular cancer.

[0294] Example 2

[0295] Ability to detect ctDNA as a function of breast cancer stage

[0296] exist Figure 5 In the example 12, each point is the ctDNA score for an individual subject with breast cancer in the CCGA population described in Example 12 below, using WGS sequencing. The method for calculating the cfDNA score for each subject includes obtaining a plurality of first sequence reads 140 in electronic form from a biological sample of each subject in a population, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules. The plurality of first sequence reads 140 are used to identify support for each variant 144 in a variant set 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the variant set 142. The observed frequency (support 146) of each individual variant 144 in the variant set 142 is compared to a corresponding reference frequency 132 for the individual variant in a reference set 128. Each such corresponding reference frequency 132 in the reference set 128 is a frequency of another variant in a first abnormal tissue sample obtained from the subject.

[0297] In addition to plotting the ctDNA fraction for each subject, Figure 5The subjects are divided according to the stage of breast cancer, and each subject is annotated according to one of three different categories. The first category (red triangles) represents situations where the plurality of sequence reads 140 from a biological sample of the subject provide sufficient basis to independently identify at least one variant 144 that is paired with one of the plurality of variants in the reference set. Thus, in many such embodiments, the abnormal tissue sample (e.g., tumor) does not utilize the plurality of cell-free DNA variants, or vice versa; rather, the targeted assay based on the plurality of sequence reads from the biological sample (e.g., blood) independently identifies the variants without relying on sequencing data from the tumor. The second category (blue triangles) represents a read-based evidence analysis of a tumor variant, where cfDNA is observed to have a plurality of sequence reads supporting at least one variant identified by direct tumor sequencing of the tumor. The third category (black circles) indicates that there is no evidence that the plurality of cfDNA sequence reads have a plurality of variants that are paired with the plurality of variants observed directly in the abnormal tissue (breast cancer tissue).

[0298] Figure 5 A very large dynamic range of tumor fractions was observed within each tumor stage. Figure 5 It is further indicated that when the tumor fraction is 1% or higher, the assay uses an estimable confidence interval for detecting breast cancer. Between 1.0% and 0.1%, the assay's performance decreases. For many black dots, the confidence interval is always zero, meaning that for many of these individuals, it is confident that these individual samples will not exceed the tumor fraction. Therefore, for Figure 5 In an analysis of stage II breast cancer in the CFDNA cohort, a large cohort of subjects with stage II breast cancer had a tumor fraction below the cutoffs for detection by the assay. In other words, a large number of subjects with stage II breast cancer had a low dropout rate, indicating that the discrimination of ctDNA in the cfDNA for many of these subjects was below the cutoffs for detection.

[0299] Example 3

[0300] Ability to detect ctDNA as a function of cfDNA fraction

[0301] Figure 6An estimate of how many individuals in the CCGA population described in Example 12 below using WGS sequencing were classified as having cancer using one of three different classifiers (Y-axis) as a function of the cfDNA score (X-axis) is provided. That is, multiple subjects were divided into one of eight data bins on the X-axis based on the cfDNA score, and then, for each of the three different classifiers, the mean set range of the sensitivity of each such data bin at 95% specificity was plotted on the Y-axis. Figure 6 For each cfDNA data bin in , the three different classifiers are, from left to right (using the data bin (0, 0.000316] for illustration), “A score” 602, “B score” 604, and “M score” 606.

[0302] The A score classifier described herein is a classifier of tumor mutation burden based on targeted sequencing analysis of multiple non-synonymous mutations. For example, a classification score (e.g., an "A score") can be calculated using logistic regression for tumor mutation burden data, wherein an estimate of the tumor mutation burden for each individual is obtained from the targeted cfDNA assay. In some embodiments, a tumor mutation burden can be estimated based on the total number of multiple variants for each individual, the individual being identified as multiple candidate variants in the cfDNA, identified by noise modeling and conjugation, and / or found to be non-synonymous in any gene annotation that overlaps with the multiple variants. The number of tumor mutation burdens of a training set can be input to a penalized logistic regression classifier to determine a cutoff for achieving 95% specificity using cross-validation. Figure 6 An example showing the effectiveness of such cross-validation can be found in, for example, R. Chaudhary et al., 2017, “Estimating tumor mutational burden using next-generation sequencing assays,” Journal of Clinical Oncology, 35(5), suppl. e14529, online preprint, which is incorporated herein by reference in its entirety.

[0303] The B-score classifier is described in U.S. Patent Publication No. 62 / 642,461, entitled "Methods and Systems for Selecting, Managing, and Analyzing High-Dimensional Data," filed as 62 / 642,461, which is incorporated herein by reference. According to the B-score method, for a plurality of regions of low variability, a first set of sequence reads from a plurality of nucleic acid samples from a plurality of healthy subjects in a reference group of a plurality of healthy subjects is analyzed. Thus, each sequence read in the first set of sequence reads from a plurality of nucleic acid samples from each healthy subject is compared with a region in the reference genome. Thus, a training set of sequence reads from a plurality of sequence reads of a plurality of nucleic acid molecules from a plurality of subjects in a training group is selected. Each training read in the sequence set is compared with a region in the plurality of low variability regions in the reference genome identified from the reference set. The training set includes a plurality of sequence reads from a plurality of nucleic acid samples from a plurality of healthy subjects, and a plurality of sequence reads from a plurality of nucleic acid samples from a plurality of diseased subjects known to have cancer. The plurality of nucleic acid samples from the sequence group are of the same or similar type as the plurality of nucleic acid samples from the reference group of healthy subjects. Thus, the number of sequence reads from the training set is used to determine one or more parameters reflecting differences between the plurality of sequence reads from the plurality of nucleic acid samples from the plurality of healthy subjects and the plurality of sequence reads from the plurality of nucleic acid molecules from the plurality of diseased subjects in the training set. Next, a test set of sequence reads associated with the plurality of nucleic acid molecules is received, the test set of sequence reads comprising a plurality of cfNA fragments from a test subject whose status with respect to the cancer is unknown, and a likelihood of the test subject having the cancer is determined based on the one or more parameters.

[0304] The M-score classifier is described in U.S. patent application Ser. No. 62 / 642,480, filed Mar. 13, 2018, entitled “Methylated Fragment Abnormal Detection,” which is incorporated herein by reference.

[0305] Figure 6 It is noted that at cfDNA fractions above 3%, all three classifiers detected the multiple individuals with the cancer. For lower cfDNA fractions, the sensitivity of the M-score classifier was statistically significantly improved relative to the B-score classifier in the (0.00316, 0.01] interval. Therefore, for intermediate dropout rates, the M-score classifier appears to be more advantageous. For lower dropout rates, when cfDNA is below 00.316, no classifier appears to be suitable. Therefore, Figure 6This provides motivation for further improving the cancer detection classifier. On the X-axis, a comma between two values indicates a range, parentheses indicate "excluding," and square brackets indicate "inclusive." For cfDNA fractions of 3% or higher, the multiple classifiers each had a sensitivity of 95% or higher and a false positive rate of 5%.

[0306] Example 4

[0307] Ability to identify breast cancer based on a cfDNA score function, sequencing protocol, and breast cancer subtype

[0308] Figure 7A and 7B Describe in detail the sensitivity of a breast cancer identification classifier using whole genome bisulfite sequencing (WGBS) ( Figure 7A ) and whole genome sequencing (WGS) ( Figure 7B ) for variant identification, thereby identifying multiple subjects as having breast cancer or not based on a cfDNA score function for four different subtypes of breast cancer using the CCGA population described in Example 12 below, wherein the four different subtypes of breast cancer are HER2+ (solid circles), HR+ / HER2- (open circles), other / missing (solid squares), and TNBC (open squares). Figure 7 demonstrates that different types of variant identification methods have differences in classifier sensitivity given a breast cancer subtype (e.g., HER2+ versus hormone receptor+ (HR+). Figure 7 further indicates that more aggressive cancers have better signal availability for HER2+ than less aggressive forms of breast cancer. For example, see Figure 7A 7 , the sensitivity is the assignment of cancer to non-cancer. For FIG7 , no ctDNA-free information is needed to classify multiple subjects as “having cancer” and “not having breast cancer” based on WGSB and WGS data, respectively. FIG7 demonstrates that the cancer detection classifier has a higher ctDNA score for these cancers.

[0309] Example 5

[0310] Accuracy of a whole-genome bisulfite-sequencing multi-class cancer type classifier as a function of cfDNA score

[0311] Figure 8 Detailed description of the accuracy of a multi-class classifier according to a cDNA score function for the CCGA subject population that has been sequenced using whole genome bisulfite sequencing (WGBS) (Example 12 below), which spans Figure 3For more details about WGBS, see, for example, Example 13. See also U.S. Patent Application No. 62 / 642,480, filed on March 13, 2018, entitled "Detection of Methylated Fragment Abnormalities," which is incorporated herein by reference. Figure 8 As illustrated, the population is divided into eight bins of different cfDNA scores, and the accuracy of the WGBS classifier for each such bin and the number of subjects in the population for each such bin is provided, wherein the accuracy is defined as the ability to place the correct cancer for a particular subject within the top two cancer class probabilities. Figure 8 It was shown that a threshold of ctDNA score level was required to achieve correct assignment using the WGBS multi-class cancer type classifier.

[0312] Example 6

[0313] Proportion of subjects showing a minimum ctDNA score as a function of clinical stage

[0314] Figure 9 The number of samples in the CCGA population that exhibit a minimum ctDNA score across all cancers represented by the population is specified. Figure 9 The method for calculating the cfDNA score for each subject disclosed in the present invention includes obtaining a plurality of first sequence reads 140 in electronic form from a biological sample of each subject in the population, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules. The plurality of first sequence reads 140 are used to identify support for each variant 144 in a set of variants 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the set of variants 142. The observed frequency (support 146) for each individual variant 144 in the set of variants 142 is compared to a corresponding reference frequency 132 for the individual variant in a reference set 128 to determine the ctDNA score for each subject. Each such corresponding reference frequency 132 in the reference set 128 is a frequency of another variant in a first abnormal tissue sample obtained from the subject.

[0315] Figure 9Proportions for multiple subjects in the population are disclosed, showing a ctDNA score of 0.01 climbing from just above 0.00 for all stage I cancers represented by the population (N=157 subjects with stage I cancer in the population) to approximately 0.75 for all stage II cancers represented by the population (N=59 subjects with stage IV cancer in the population). Figure 9 It is illustrated that the available information in the ctDNA scores of multiple cancer patients can be used to classify the conditions of multiple subjects according to the present disclosure, including the methods described in Figure 2. Examples 1 to 6 collectively demonstrate that the methods of the present disclosure are capable of classifying multiple subjects, evaluating the performance of multiple classifiers based on ctDNA scores, and assessing the quality of signals given a fixed ctDNA score across different cancer types. Advantageously, Examples 1 to 6 collectively demonstrate that the disclosed systems and methods are capable of detecting more aggressive forms of cancer, which is highly desirable.

[0316] Example 7

[0317] In silico models combining ctDNA scores with tumor-derived digital pathology

[0318] Examples 1 to 6 illustrate that the ctDNA score determined according to various methods of the present disclosure can be combined with information obtained from digital pathology to provide a model for predicting the aggressiveness of a particular cancer. Thus, the present disclosure demonstrates the utility of various models that take into account the ctDNA score and further include digital pathology to determine the aggressiveness of a particular cancer condition in a particular subject. The cfDNA score for a subject is determined by obtaining a plurality of first sequence reads 140 in electronic form from a biological sample of the subject, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules (e.g., from the subject's blood). According to various teachings of the present disclosure, the plurality of first sequence reads 140 are used to identify support for each variant 144 in a variant set 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the variant set 142. The observed frequency (support 146) for each individual variant 144 in the variant set 142 is compared to a corresponding reference frequency 132 for the individual variant in a reference set 128 to determine the ctDNA score for the subject. A plurality of such reference frequencies 132 are obtained from a plurality of sequence reads obtained from a tumor or a tumor portion of the subject. Furthermore, one or more sections of the tumor or tumor portion are analyzed using a plurality of computer vision techniques to estimate density, immune cell infiltration, cell necrosis, and / or proliferation rate, or other parameters related to the aggressiveness of a cancer. This information is then combined with the ctDNA and input into a classifier that assesses the aggressiveness of the cancer in the subject and / or any other status associated with the cancer.

[0319] Example 8

[0320] Figure 10 The positive correlation between ctDNA scores and tumor size across all stages of cancer using the CCGA cohort described in Example 12 is demonstrated. Because tumor size in many cases is positively correlated with cancer aggressiveness, Example 8 provides additional support for using cfDNA scores to classify multiple subjects in accordance with the present disclosure, including the methods disclosed in conjunction with FIG2 , the additional embodiments disclosed below, and the claims of the present disclosure.

[0321] Example 9

[0322] Correlation of ctDNA scores with the Ki67 marker for proliferation

[0323] Ki-67 is a nuclear protein associated with cell proliferation. See Gerdes et al., 1983, "Generation of mouse monoclonal antibodies reactive with a human nuclear antigen involved in cell proliferation," Int. J. Cancer 31(1), 13-20, incorporated herein by reference. One method for analyzing a subject for the Ki-67 antigen is immunohistochemical assessment. It has been shown that the Ki-67 nuclear antigen is expressed in certain phases of the cell cycle, namely, S, G1, G2, and M phases, but not in G0. For example, see Gerdes et al., 1984, "Cell cycle analysis of human nuclear antigens associated with cell proliferation defined by the monoclonal antibody Ki-67," J Immunol. 133(4), 1710–1715; and Scholzen and Gerdes, 2000, "Ki-67 protein: from the known and the unknown," J Cell Physiol. 182(3), 311–322, each of which is incorporated herein by reference. In multiple samples from normal breast tissue, Ki-67 has been found to be expressed at low levels (<3% of cells) in ER-negative cells but not in ER-positive cells. For example, see Urruticoecha et al., 2005, "Proliferation marker Ki-67 in early breast cancer," J Clin Oncol. 23: 7212–7220, which is incorporated herein by reference. By immunostaining using the monoclonal antibody Ki-67, the growth fraction of the tumor cell population can be assessed.

[0324] For this example, immunohistochemical staining was performed, and the proportion of multiple malignant cells that stained positive for the nuclear antigen Ki-67 was quantitatively and visually assessed using an optical microscope. Using the anti-human Ki-67 monoclonal antibody, multiple Ki-67 values were obtained based on the percentage of multiple positively labeled malignant cells. Figure 11 In the present invention, the Ki-67 percentage score is defined as the percentage of multiple positively stained tumor cells out of the total number of malignant cells evaluated. See Inwald, 2013, "Ki-67 is a prognostic parameter in breast cancer patients: results from a large population-based cancer registry cohort", Breast Cancer Res. Treat. 139(2):539-552, which is incorporated herein by reference.

[0325] exist Figure 11In the present invention, the cfDNA fraction for each specific subject in a CCGA cohort exhibiting solid invasive cancers, as described in Example 12 below, is determined by obtaining a plurality of first sequence reads 140 in electronic form from a biological sample of the subject, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules (e.g., blood from the subject). The plurality of first sequence reads 140 are used to identify support for each variant 144 in a set of variants 142 of the biological sample, thereby determining an observed frequency (support 146) for each variant 144 in the set of variants 142. The observed frequency (support 146) for each individual variant 144 in the set of variants 142 is compared to a corresponding reference frequency 132 for the individual variant in a reference set 128 to determine the ctDNA fraction for the subject. The plurality of such reference frequencies 132 are obtained from a plurality of sequence reads obtained from a tumor or a portion of a tumor of the subject from which the plurality of Ki-67 values were obtained.

[0326] exist Figure 11 In the leftmost column, multiple samples with a Ki-67 score greater than 10 show that there are many samples in the tail with a dropout rate greater than 0.1000. This suggests that the combination of Ki-67 and the ctDNA score can provide a basis for diagnosing a condition in a subject, such as a more aggressive form of breast cancer.

[0327] Example 10

[0328] Obtain multiple sequence reads

[0329] Figure 12 FIG1 is a flow chart of a method 1200 for preparing a nucleic acid sample for sequencing according to one embodiment. The method includes, but is not limited to, the following steps. For example, any step of method 1200 may include a quantification substep for quality control, or other laboratory assay procedures known to those skilled in the art.

[0330] In block 1202, a nucleic acid sample (DNA or RNA) is extracted from a subject. The sample may be any subset of the human genome, including the whole genome. The sample may be extracted from a subject known to have or suspected of having cancer. The sample may include blood, plasma, serum, urine, feces, saliva, other types of body fluids, or any combination thereof. In some embodiments, methods for extracting a blood sample (e.g., syringe or finger prick) may be less invasive than procedures for obtaining a tissue biopsy, which may require surgery. The extracted sample may include cfDNA and / or ctDNA. For healthy individuals, the human body can naturally clear cfDNA and other cellular debris. If a subject has a cancer or disease, ctDNA in an extracted sample may be present at a detectable level for diagnosis.

[0331] At block 1204, a sequenced library is prepared. During library preparation, multiple unique molecular indices (UMIs) can be added to the multiple nucleic acid molecules (e.g., DNA molecules) through adapter ligation. The multiple UMIs are multiple short nucleic acid sequences (e.g., 4 to 10 base pairs) that are added to the ends of the multiple DNA fragments during adapter ligation. In some embodiments, the multiple UMIs are multiple degenerate base pairs that serve as a unique tag that can be used to identify multiple sequence reads originating from a specific DNA fragment. During PCR amplification after adapter ligation, the multiple UMIs are replicated along with the attached DNA fragments. This provides a method for identifying multiple sequence reads originating from the same original fragment in downstream analysis.

[0332] In block 1206, a plurality of target DNA sequences are enriched from the library. During enrichment, a plurality of hybridization probes (also referred to herein as "probes") are used to target and pull down a plurality of nucleic acid fragments that provide information on the presence or absence of cancer (or disease), the state of the cancer, or a cancer classification (e.g., cancer type or tissue of origin). For a particular workflow, the plurality of probes can be designed to bind (or hybridize) to a target (complementary) strand of DNA. The target strand can be a "positive" strand (e.g., this strand is transcribed into mRNA and subsequently translated into a protein) or a "negative" strand. The plurality of probes can range in length from tens, hundreds, or thousands of base pairs. In some embodiments, the plurality of probes are designed based on a panel of genes to analyze a plurality of specific mutations or a plurality of target regions of the genome corresponding to certain cancers or other disease types (e.g., humans or other organisms). Furthermore, the plurality of probes can cover a plurality of overlapping portions of a target region.

[0333] Figure 13 A diagrammatic representation of a process for obtaining multiple sequence reads according to one embodiment. Figure 13 An example of a nucleic acid fragment segment 1300 from the sample is depicted. Here, the nucleic acid fragment 1300 can be a single-stranded nucleic acid fragment, such as a single strand. In some embodiments, the nucleic acid fragment 1300 is a double-stranded cfDNA fragment. This illustrative example depicts three regions 1305A, 1305B, and 1305C of the nucleic acid fragment that can be targeted by different probes. Specifically, each of the three regions 165A, 165B, and 165C includes an overlapping position on the nucleic acid fragment 160. An exemplary overlapping position is Figure 13 1305B, and is depicted as a cytosine ("C") nucleotide base 1302. The cytosine nucleotide base 1302 is located near a first edge of region 1305A, in the middle of region 1305B, and near a second edge of region 1305C.

[0334] In some embodiments, one or more (or all) of the plurality of probes are designed based on a gene panel to analyze a plurality of specific mutations or a plurality of target regions of the genome (e.g., of a human or other organism) suspected of corresponding to certain cancers or other disease types. By using a target gene panel, rather than sequencing all the expressed genes of a genome, also known as "whole exome sequencing," the method 1200 can be used to increase the sequencing depth of the plurality of target regions, where depth refers to a count of the number of times a particular target sequence in the sample is sequenced. The increase in sequencing depth reduces the amount of nucleic acid sample that needs to be input.

[0335] The nucleic acid sample 1300 is hybridized using one or more probes to generate an understanding of a target sequence 1370. Figure 13 As shown, the target sequence 1370 is the nucleotide base sequence of the region 1305 targeted by a hybridization probe. The target region 1307 can also be referred to as a hybrid nucleic acid segment. For example, the target region 1370A corresponds to the region 1305A targeted by a first hybridization probe, the target region 1370B corresponds to the region 1305B targeted by a second hybridization probe, and the target region 1370C corresponds to the region 1305C targeted by a third hybridization probe. Considering that the cytosine nucleotide base 1302 is located at a different position in each region 1305A to C targeted by a hybridization probe, each target sequence 1370 includes a nucleotide base corresponding to the cytosine nucleotide base 1302 located at a specific position on the target sequence 1370.

[0336] After a hybridization step, the multiple hybrid nucleic acid fragments are captured and can also be amplified using PCR. For example, the multiple target sequences 1370 can be enriched to obtain multiple enriched sequences 1380 that can be subsequently sequenced. In some embodiments, each enriched sequence 1380 is copied from a target sequence 1370. The multiple enriched sequences 1380A and 1380C amplified from target sequences 1370A and 1370C, respectively, also include thymine nucleotide bases near the edge of each sequence read 180A or 180C. As used below, the mutant nucleotide sequence (e.g., thymine nucleotide base) that mutates relative to the reference allele (e.g., cytosine nucleotide base 1302) in the enriched sequence 1380 is considered to be an alternative allele. In addition, each enriched sequence 1380B amplified from target sequence 1370B includes the cytosine nucleotide base close to or located in the middle of each enriched sequence 1380B.

[0337] At block 1208, a plurality of sequence reads are generated from the plurality of enriched DNA sequences, e.g., Figure 13 The plurality of enriched sequences 180 shown are shown. A plurality of sequencing data can be obtained from the plurality of enriched DNA sequences using tools known in the art. For example, the method 1200 may include next-generation sequencing (NGS) techniques, including sequencing by synthesis (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (ion-pour sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore), or paired-end sequencing. In some embodiments, sequencing by synthesis is used with a plurality of reversible dye terminators for massively parallel sequencing.

[0338] In some embodiments, the plurality of sequence reads are compared to a reference genome using a plurality of methods known in the art to determine information about the alignment positions. The alignment position information may indicate a starting position and an ending position of a region of the reference genome corresponding to a starting nucleotide base and an ending nucleotide base of a particular sequence read. The alignment position information may also include a sequence read length determined by the starting position and the ending position. A region in the reference genome may be associated with a gene or a fragment of a gene.

[0339] In various embodiments, a sequence read includes a read pair represented as R1 and R2. In some examples, the first read R1 is sequenced from a first end of a nucleic acid fragment, and the second read R2 is sequenced from a second end of the nucleic acid fragment. Therefore, in the example, the multiple nucleotide base pairs of the first read R1 and the second read R2 are consistent with the multiple nucleotide bases of the reference genome in comparison (for example, in opposite directions). The comparison position information derived from the read pairs R1 and R2 may include: a starting position corresponding to an end of a first read (for example, R1) in the reference genome, and an end position corresponding to an end of a second read (for example, R2) in the reference genome. In other words, the starting position and the end position in the reference genome represent possible positions corresponding to the nucleic acid fragment in the reference genome. An output file having a SAM (sequence alignment map) format or a BAM (binary) format can be generated, and the output can be used for further analysis, for example, the variant identification described above in conjunction with Figure 2 and Example 11.

[0340] Example 11

[0341] Identifying variants

[0342] Figure 14FIG14 is a flow chart of a method 1400 for determining multiple variants of multiple sequence reads according to one embodiment. In some embodiments, variant identification (e.g., for multiple SNVs and / or multiple insertions and deletions (indels)) based on input sequencing data is performed, as discussed above in conjunction with FIG2 and Example 10. In step 1402, multiple aligned sequence reads of the input sequencing data are split. In some embodiments, splitting the multiple sequence reads includes using multiple UMIs and, optionally, alignment position information of the sequence data from an output file (e.g., from the method described in Example 10) to split the multiple sequence reads into a consensus sequence for determining the most likely sequence of a nucleic acid fragment or a portion of the nucleic acid fragment. In some embodiments, the unique sequence tags are approximately 4 to 20 nucleic acids in length. Because the multiple UMIs are replicated along with the multiple linked nucleic acid fragments through enrichment and PCR, it is possible to determine that some sequence reads originate from the same molecule of a nucleic acid sample. In some embodiments, multiple sequence reads with identical or similar alignment position information (e.g., start and end positions at a threshold offset) and a common UMI are split, and a split read (also referred to herein as a consensus read) is generated to represent the nucleic acid fragment. If the corresponding split read pair has a common UMI, the consensus read is designated as a "duplex," indicating that both the positive and negative strands of the source nucleic acid molecule are captured; otherwise, the split read is designated as a "non-duplex." In some embodiments, as an alternative to or in addition to splitting the multiple sequence reads, other types of error correction can be performed on the multiple sequence reads.

[0343] In step 1405, the plurality of split reads are stitched together based on the corresponding alignment position information. In some embodiments, the alignment position information between a first read and a second read is compared to determine whether the nucleotide base pairs of the first and second reads overlap in the reference genome. In one use case, in response to determining that an overlap (e.g., a specific number of nucleotide bases) between the first and second reads is greater than a threshold length (e.g., a threshold number of nucleotide bases), the first and second reads are designated as "stitched"; otherwise, the plurality of split reads are designated as "unstitched." In some embodiments, if the overlap is greater than the threshold length and if the overlap is not a sliding overlap, the first and second reads are stitched together. For example, a sliding overlap can include a homopolymer run (e.g., a single repeated nucleotide base), a dinucleotide run (e.g., a two-nucleotide base sequence), or a trinucleotide run (e.g., a three-nucleotide base sequence), wherein the homopolymer run, dinucleotide run, or trinucleotide run all have at least a threshold length of a plurality of base pairs.

[0344] In step 1410, the plurality of reads are assembled into a plurality of paths. In some embodiments, this involves assembling the plurality of reads to generate a directed graph, such as a de Bruijn graph, for a target region (e.g., a gene). The directed graph has a plurality of unidirectional edges representing sequences of k nucleotide bases (also referred to herein as "k-mers") in the target region, and the plurality of edges are connected by a plurality of vertices (or nodes). The plurality of split reads are aligned to a directed graph such that any one of the plurality of split reads can be sequentially represented by a subset of the plurality of edges and corresponding plurality of vertices.

[0345] In some embodiments, multiple parameter sets describing multiple directed graphs and processing multiple directed graphs are confirmed. The parameter set may include a count of successfully comparing multiple k-mers from multiple split reads to a k-mer represented by a node or edge in the directed graph. In some embodiments, the multiple directed graphs and multiple corresponding parameter sets may be stored so that subsequent retrieval updates multiple diagrams or generates multiple new diagrams. For example, a compressed version of a directed graph can be generated based on the parameter set (e.g., or an existing diagram can be modified). In one use case, in order to filter out data of a directed graph with a lower importance level, multiple nodes or multiple edges with a count less than a threshold are removed (e.g., "trimmed" or "pruned"), while maintaining multiple nodes or multiple edges with multiple counts greater than or equal to the threshold.

[0346] In step 1415, the variant identifier 240 generates multiple candidate variants from the multiple assembly paths. In one embodiment, multiple candidate variants are generated by comparing a directed graph (compressed by pruning multiple edges or multiple nodes in step 1410) with a reference sequence of a target region of a genome. The multiple edges of the directed graph can be compared with the reference sequence, and the multiple genomic positions of the multiple error-pairing edges and the nucleotide bases of the multiple error-pairings adjacent to the multiple edges are recorded as the positions of multiple candidate variants. In some embodiments, the multiple genomic positions of the multiple error-pairing edges and the nucleotide bases of the multiple error-pairings on the left and right sides of the multiple edges are recorded as the positions of multiple identified variants. In addition, multiple candidate variants can be generated based on the sequencing depth of a target region. In particular, multiple variants in multiple target regions with a greater sequencing depth can be identified more confidently, for example, because more sequence reads help resolve (for example, using redundancy) multiple error-pairings or other base pair variations between multiple sequences.

[0347] In some embodiments, a model is used to generate multiple candidate variants to determine the expected noise rate for multiple sequences from a subject. Although in some embodiments, one or more different types of models are used, the model may be a Bayesian hierarchical model. Furthermore, a Bayesian hierarchical model may be one of many possible model architectures that can be used to generate multiple candidate variants and are related to each other, because they all model position-specific noise information to improve the sensitivity / specificity of variant identification. More specifically, the model can be trained using multiple samples from multiple healthy individuals to model the multiple expected noise rates for the position of each sequence read.

[0348] Further, multiple different models can be used for post-training of the application. In one example, a first model is trained to model multiple SNV noise rates, and a second model is trained to model multiple insertion and deletion noise rates. Further, multiple parameters of the model can be used to determine a likelihood of one or more true positives in a sequence read. A quality score (e.g., on a logarithmic scale) can be determined based on the likelihood. In one example, the quality score is a Phred quality score Q = -10log 10 P, where P is the probability of an erroneous candidate variant identification (e.g., a false positive). Other models, such as a junctional model, use the output of one or more Bayesian hierarchical models to determine the expected noise of multiple nucleotide mutations in multiple sequence reads from different samples.

[0349] At step 1420, one or more types of models or filters are used to filter the candidate variants. In one embodiment, the candidate variants are scored using a junction model, a marginal variant prediction model, or corresponding true positive likelihood or quality scores. Additionally, a marginal filter and / or a non-synonymous filter are used to filter the marginal variants and / or non-synonymous mutations, respectively.

[0350] At step 1425, the plurality of filtered candidate variants are output. In some embodiments, some or all of the plurality of determined candidate variants are output along with a corresponding score from the plurality of filtering steps.

[0351] Example 12

[0352] Cell-Free Genome Atlas (CCGA) community

[0353] Multiple subjects from the CCGA [NCT02889978] were used in multiple examples of the present disclosure. The CCGA is a prospective, multicenter, case-control, and observational study with longitudinal follow-up. The study recruited 9,977 of 15,000 participants who were demographically balanced at 141 sites. Blood was collected from subjects who were defined as newly diagnosed therapy-naive cancer (C, cases) and subjects who were not diagnosed with cancer (non-cancer [NC], controls) at the time of enrollment. This pre-planned substudy included 1628 cases and 1172 controls, spanning 20 tumor types and all clinical stages. Prior to analysis, multiple samples were divided into a training set (1,785) and a test set (1,015). Multiple samples were selected to ensure a pre-specified distribution of multiple cancer types and non-cancer across sites in each population, and to match the frequencies of cancer and non-cancer samples by sex and age. Figure 18 Demographic information for multiple participants included in the final analysis is provided.

[0354] Cell-free DNA was isolated from plasma, while genomic DNA (gDNA) was isolated from white blood cells (WBCs) and tumor tissue using multiple standard methods. Three different high-intensity sequencing methods were used for cfDNA analysis: (i) whole-genome bisulfite sequencing (WGBS; 30x depth) of cfDNA, in which multiple aberrantly methylated fragments were used to generate multiple normalized scores; (ii) paired cfDNA and WBC whole-genome sequencing (WGS; 30x depth), in which a novel machine learning algorithm generated multiple cancer-associated signal scores, and junction analysis identified multiple shared events; and (iii) paired cfDNA and WBC targeted sequencing (507-gene panel; 60,000x depth, referred to herein as the "ART" assay), in which a junction caller removed multiple somatic variants derived from WBCs and residual technical noise. Targeted sequencing of WBC gDNA was performed to identify clonal hematopoiesis (CH). WGS of tumor tissue gDNA was performed to identify multiple somatic variants, which were used to calculate the cfDNA tumor fraction.

[0355] In the targeted assay, multiple cfDNA somatic variants (SNVs / indels) paired with non-tumor WBCs accounted for 76% of all variants in NC and 65% in C. Consistent with somatic mosaicism (e.g., clonal hematopoiesis), multiple variants paired with WBC increased with age; some were atypical loss-of-function mutations not previously reported. After removing WBC variants, multiple canonical driver somatic variants were highly specific for C (e.g., in EGFR and PIK3CA, 0 NC had variants paired with 11 and 30 for C, respectively). Similarly, 8 NCs detected by WGS had multiple somatic copy number alterations (SCNAs), 4 of which originated from WBC. The CCGA WGBS data showed useful information on high- and low-level CpGs (1:2 ratio); a subset of these was used to calculate multiple methylation scores. A consistent "cancer-like" signal (representing potentially undiagnosed cancer) was observed in <1% of NC participants across all assays. An increasing trend was observed in NC versus stage I to III versus stage IV (mean ± SD nonsyn.SNVs / indels per Mb [mean ± SD] NC: 1.01 ± 0.86, stage I to III: 2.43 ± 3.98; stage IV: 6.45 ± 6.79; WGS score NC: 0.00 ± 0.08, stage I to III: 0.27 ± 0.98; IV: 1.95 ± 2.33, methylation score NC: 0 ± 0.50; stage I to III: 1.02 ± 1.77; IV: 3.94 ± 1.70). These data demonstrate the feasibility of achieving > 99% specificity for aggressive cancers and support the promise of cfDNA analysis for early cancer detection.

[0356] More information on the CCGA assay is disclosed in Klein et al., “Development of a comprehensive cell-free DNA (cfDNA) assay for early detection of multiple tumor types: the Circulating Cell-Free Genome Atlas (CCGA) study,” 2018 ASCO Annual Meeting, June 1-5, 2018, abstract 12021 #134, Chicago, IL, which is incorporated herein by reference.

[0357] Methylation alleles were determined using WGBS. For each sample, the WGBS fragment set was reduced to a subset of aberrant fragments with extreme methylation states (UFXM). Fragments that occur at high frequencies in individuals without cancer or have unstable methylation are unlikely to produce highly discriminatory features for classifying cancer status. A statistical model of typical fragments was used using an independent reference set of 108 cancer-free, non-smoking participants from the CCGA study (age: 58±14 years, 79 [73%] women). These samples were used to train a Markov chain model (level 3) that estimates the probability of a specific sequence of CpG methylation status in a fragment. This model was shown to be calibrated within a normal range of fragments (p-value ≥ 0.001) and was used to reject multiple fragments as insufficiently aberrant using a p-value from the Markov model p≥0.001.

[0358] A further data reduction step selected only fragments with coverage of at least 5 CpGs and an average methylation of >0.9 (hypermethylated) or <0.1 (hypomethylated) per fragment. This process resulted in a median (range) of 2,800 (1,500-12,000) UFXM fragments for participants without cancer during training, and a median (range) of 3,000 (1,200-220,000) UFXM fragments for participants with cancer during training. Since this data reduction process only uses data from the reference set, this stage only needs to be applied once to each sample.

[0359] At selected loci within the genome, an approximate log-ratio score for cancer status was constructed independently for hypermethylated and hypomethylated UFXM. First, for each sample at the locus, a binary feature was generated: 0 if no UFXM fragment overlapped with the locus within the sample; 1 if a UFXM fragment overlapped with the locus. Then, from the sample with cancer (C c ) or no cancer (C nc ) of multiple participants to calculate the number of positive values (1s) in multiple samples. Then, the logarithmic ratio score is constructed as: log(C c +1)-log(C nc +1), where a regular term is added to the count and the number of items with each group (N c and N nc ) is a normalization term related to the total number of samples in , because it is a constant (log[N nc +2]-log[N cMultiple scores were constructed at multiple positions of all CpG sites in the genome, resulting in approximately 25M loci with multiple assigned scores: one score for UFXM hypermethylated segments and one score for UFXM hypomethylated segments.

[0360] Given a locus-specific log ratio score, multiple UFXM segments in a sample were scored by taking the maximum of all log ratios for the locus within the segment and pairing it with a methylation category of hypermethylated or hypomethylated. This resulted in a score for each UFXM segment in a sample.

[0361] By separately deriving the scores for a subset of the most highly ranked fragments in each sample for both hypermethylated and hypomethylated fragments, these fragment-level scores in a sample are reduced to a small set of features for each sample. In this way, a small set of useful features is used to capture information about the most informative fragments in each sample. In a sample with a low cfDNA tumor fraction, only a few fragments are expected to have abnormal information content.

[0362] In each category of fragments, the ranking 1, 2, 4, ..., 64 (2 i , i in 0:6), resulting in 14 features (7 and 7). In order to adjust the ordering depth of the sample, the ordering process is regarded as a function that maps multiple orders to multiple scores, and interpolates between the multiple observed scores to obtain multiple scores corresponding to the adjusted ordering. The multiple orders are adjusted in a linear proportion to the relevant sample depth: if the relevant sample depth is x, then multiple interpolated scores are obtained when x is multiplied by the multiple initial orderings (for example, for x=1.1, we obtain the scores in the order floor(1.1), floor(2.2), ..., floor(x·2 i )). Each sample is then assigned a set of 14 adjusted extreme ranking scores for further classification.

[0363] Given the feature vector, a kernel function logistic regression classifier is used to capture multiple potential nonlinearities when predicting cancer / non-cancer status from the multiple features. Specifically, an iso-radial basis function (power 2) is used as the kernel function with a scale parameter γ, and an L2 regularization parameter λ (by dividing by m 2A regularized kernel logistic regression classifier (KLR) was trained using the γ and λ parameters (adjusted for γ, where m is the number of samples, so λ naturally scales with the amount of training data). γ and λ were optimized using internal cross-validation to holdout log loss on the specified training data, and were optimized using a grid search in the range 1-0.01 (γ) and 1000-10 (λ) in 7 multiplication steps, starting from the maximum value and halving the parameter at each step. The median of the multiple best parameters during the internal cross-validation folds was 0.125 for γ and 125 for λ.

[0364] To evaluate the performance of this extreme ranking score classifier process on the CCGA substudy dataset, cross-validation was applied to the training set, where the samples were divided into 10 folds. Each fold was left out, and the extreme ranking score (ERS) classifier was trained on the remaining 9 / 10 of the data (using internal cross-validation within these folds to optimize γ and λ). The log-ratio scores used in characterization only accessed data from the training folds. The output scores from each left-out fold were combined and used to construct a receiver operating characteristic (ROC) curve for performance.

[0365] Sensitivity estimates.

[0366] Figure 19A and 19B Provides information on model sensitivity using Figure 18 The training set is trained and compared with Figure 18 The training set ( Figure 19A , N=1,416) and the test set ( Figure 19B , N = 1,416), and were divided according to tumor origin. Figure 19C and 19D Provide the training set divided according to tumor origin ( Figure 19C ) and the test set ( Figure 19D ) in the tumor fraction. In more detail, for Figure 19A and 19B , when in the training set ( Figure 19A ) and the test set ( Figure 19B) are presented for each tumor type in the training and test sets (x-axis) at 98% specificity (y-axis) when analyzed by WGBS (left-hand strip, blue), WGS illustrating CH (middle strip, orange), and the targeted assay illustrating CH (left-hand strip, gray). Error bars represent 95% confidence intervals. The number of samples per cancer type is indicated in brackets. Multiple myeloma and leukemia from post hoc assays are presented separately. Figure 19C and 19D Provides the Figure 19C ) and the test set ( Figure 19D Figure 3. Box plot of the cfDNA tumor fraction (y-axis) for a subset of participants with sequenceable tumor-normal tissue and at least one mutant cfDNA read (as indicated in parentheses) for each tumor type (x-axis) in the WT ... Figure 19D It is established that the methods of the present disclosure can be used to detect a tumor fraction in a subject's cell-free nucleic acid even when the tumor fraction f is 0.100 or less, in some examples, when the subject's tumor fraction f is 0.050 or less, 0.050 or less, 0.040 or less, 0.030 or less, or even 0.020 or less.

[0367] Figure 20A and 20B The results of cfDNA WGS were compared with those of tumor WGS according to the stage of breast, colorectal, lung and other cancers. Figure 20A ), and by each cancer type ( Figure 20B ) were compared to the cfDNA tumor fraction calculated using a WGBS dataset. Multiple samples with at least one mutation read in cfDNA are presented. Tumor scores for each participant are indicated by triangles (training set) and circles (test set), with the color of the symbol indicating detection by WGBS at 98% specificity (detected: blue; not detected: orange). Figure 20A All samples without breast, lung, or colorectal cancer were included. Figure 20B These included two neuroendocrine, two mesothelioma, two gastrointestinal stromal tumors, one anal, and four adenocarcinomas (not otherwise specified) of unknown origin.

[0368] Example 13

[0369] Generation of methylation state vector

[0370] Figure 15 A flow chart of a process 1500 for sequencing a fragment of cfDNA to obtain a methylation state vector is described according to one embodiment of the present disclosure.

[0371] Referring to step 1502, the plurality of cfDNA fragments are obtained from the biological sample (e.g., as described above in conjunction with FIG. 2 ). Referring to step 1520, the plurality of cfDNA fragments are treated to convert a plurality of unmethylated cytosines into a plurality of uracils. In one embodiment, the DNA is subjected to a bisulfite treatment that converts the plurality of unmethylated cytosines in the fragments of the cfDNA into a plurality of uracils without converting a plurality of methylated cytosines. For example, in some embodiments, a methylation reaction such as EZ DNA is performed. TM -Gold, EZ DNA methylation TM - Direct or EZ DNA methylation TM A commercial kit, such as the Rapid kit (available from Limo Research, Inc., Irvine, CA), is used for the bisulfite conversion. In other embodiments, an enzymatic reaction is used to convert unmethylated cytosines to uracils. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracils, such as APOBEC-Seq (New England Biolabs, Ipswich, MA).

[0372] A sequence library is prepared from the multiple converted cfDNA fragments (step 1530). Optionally, for cfDNA fragments or multiple genomic regions, 1535 enriches the sequencing library, which can use multiple hybridization probes to provide information for cancer status. The multiple hybridization probes are multiple short oligonucleotides that can hybridize with multiple specifically designated cfDNA fragments or target regions and enrich these fragments or regions for subsequent sequencing and analysis. Multiple hybridization probes can be used for targeted and high-depth analysis of a set of designated CpG sites of interest to researchers. Once prepared, the sequencing library or a portion of the sequencing library is sequenced to obtain multiple sequence reads (1540). The multiple sequence reads can be in a computer-readable digital format for processing and interpretation by computer software.

[0373] From the plurality of sequence reads, a position and methylation state for each CpG site is determined based on alignment of the plurality of sequence reads to a reference genome (1550). A methylation state vector for each fragment specifies a position of the fragment in the reference genome (e.g., as specified by the position of the first CpG site in each fragment, or another similar metric), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment (1560).

[0374] Additional Examples

[0375] Using tumor scores to evaluate the performance of a classifier. Another aspect of the present disclosure provides a method for evaluating the performance of a classifier. The method includes obtaining a separate dataset in electronic form comprising a plurality of first sequence reads from a separate biological sample of a corresponding subject, for each subject in a plurality of subjects, thereby obtaining a plurality of datasets. The individual biological sample of each corresponding subject includes a plurality of cell-free nucleic acid molecules from the corresponding subject. Each individual dataset in the plurality of datasets is applied to a classifier, thereby obtaining a corresponding classifier result for the individual dataset. The result of the classifier indicates whether the corresponding subject in the plurality of subjects suffers from a first cancer condition. Furthermore, before applying the plurality of datasets described above, the classifier is trained using data outside the plurality of datasets.

[0376] In some embodiments, an estimated tumor fraction in the cell-free DNA of each of the plurality of subjects is estimated using the dataset corresponding to the subject. A performance of the computer is then calculated as a function of the estimated tumor fraction across the plurality of subjects by comparing the result of the classifier for each individual subject with a clinical observation for the individual subject, the clinical observation being derived independently of the estimated tumor fraction for the individual subject based on the classifier.

[0377] Conditioned on sequencing tumors of an entire population, not just a single subject. The present disclosure provides a method for classifying a subject. The method includes: obtaining, in a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors, a plurality of first sequence reads in electronic form from a biological sample of the subject in the computer system, wherein the biological sample comprises a plurality of cell-free nucleic acid molecules. The plurality of first sequence reads of the plurality of cell-free nucleic acid molecules are used to identify support for each variant in a first set of variants. Another sequence read in the plurality of first sequence reads is considered support for a variant in the first set of variants if the individual sequence read contains all or a portion of the variant. Another sequence read in the plurality of first sequence reads is considered not support for a variant in the first set of variants if the individual sequence read does not contain all or a portion of the variant. In this manner, an observed frequency for each variant in the first set of variants is determined from the plurality of sequence reads in the plurality of first sequence reads that support or do not support each variant in the first set of variants. The observed frequency of each variant in the first variant set is compared to a corresponding reference frequency in a first reference set. Each corresponding reference frequency in the first reference set is a frequency of the corresponding variant across a plurality of first abnormal tissue samples of a common (same) first category. Next, the subject is classified. This classification step includes: when the observed frequency of each variant in the first variant set meets a first threshold, the subject is considered to have a first condition associated with the plurality of first abnormal tissue samples. The first threshold is determined by each reference frequency in the first reference set.

[0378] In some embodiments, the first condition is cancer from a common primary site. In some embodiments, the first condition is cancer from two or more common primary sites. The first condition is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

[0379] In some embodiments, the first condition is a predetermined stage of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer.

[0380] In some embodiments, the first condition is a predetermined subtype of cancer (e.g., breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, or gastric cancer).

[0381] In some embodiments, a variant in the first set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the copy number of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an abnormal methylation pattern associated with a predetermined genomic location.

[0382] In some embodiments, the plurality of first abnormal tissue samples are a plurality of tumor samples. In some embodiments, the first variant set consists of a single variant, which is a single genetic variant located at a single site in the genome of the subject. In some embodiments, the first variant set consists of a first variant and a second variant, wherein the first variant is a first genetic variant located at a first site in the genome of the subject, and the second variant is a second genetic variant located at a second site in the genome of the subject.

[0383] In some embodiments, the first variant set consists of a first variant, a second variant, and a third variant, wherein the first variant is a first genetic variation located at a first site in the genome of the subject, the second variant is a second genetic variation located at a second site in the genome of the subject, and the third variant is a third genetic variation located at a third site in the genome of the subject.

[0384] In some embodiments, the first set of variants consists of between 2 and 20 variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject (optionally located at a different site). In some embodiments, the first set of variants consists of between 2 and 200 variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject (optionally located at a different site).

[0385] In some embodiments, the set of variants comprises 200 variants, comprises 300 variants, comprises 400 variants, comprises 500 variants, comprises 750 variants, comprises 1000 variants, comprises 2000 variants, or comprises 5000 variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject (optionally located at a different site).

[0386] In some embodiments, the comparing step comprises calculating a single estimated ctDNA in the cfDNA of the human subject from the observed frequency of each variant in the first set of variants. In many such embodiments, the first threshold is a single expected ctDNA fraction in the cfDNA of the human subject determined from the values of each reference frequency in the first reference set. In some such embodiments, the single expected ctDNA fraction in the cfDNA is between 0.5x10 -4 to 1.5x10 -4 and said first condition is melanoma. In some such embodiments, said single expected ctDNA fraction in said cfDNA is between 0.5 x 10 -3 Up to 1x10 -2 and the first condition is renal cancer, uterine cancer, thyroid cancer, prostate cancer, breast cancer, bladder cancer, gastric cancer, cervical cancer, or a combination thereof. In some such embodiments, the single expected ctDNA fraction in the cfDNA is between 1x10 -2 to 0.8, and the first condition is lung cancer, esophageal cancer, head and neck cancer, colorectal cancer, anorectal cancer, ovarian cancer, hepatobiliary cancer, pancreatic cancer, lymphoma, or a combination thereof.

[0387] In some embodiments, the using step comprises aligning another of the plurality of first sequence reads to a region in a reference genome to determine whether the individual sequence read contains all or a portion of a variant. In some embodiments, the using step comprises aligning another of the plurality of first sequence reads to a lookup table of variants to determine whether the individual sequence read contains all or a portion of a variant. In some embodiments, the using step comprises aligning a sequence read in the plurality of first sequence reads to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome.

[0388] In some embodiments, the comparing step comprises calculating a single estimated circulating tumor DNA (ctDNA) fraction in the cell-free DNA (cfDNA) of the human subject from the observed frequency of each variant in the first set of variants. In many such embodiments, when the single estimated circulating tumor DNA (ctDNA) fraction exceeds 1x10 -3 , and the first condition is stage II, stage III, or stage IV breast cancer, the observed frequency of each variant in the first set of variants meets the first threshold.

[0389] In some embodiments, the method further includes: using the plurality of first sequence reads to identify support for each variant in a second set of variants, wherein another sequence read in the plurality of first sequence reads is considered to support a variant in the second set of variants when the individual sequence read contains all or a portion of the second variant, and another sequence read in the plurality of first sequence reads is considered not to support a variant in the second set of variants when the individual sequence read does not contain the individual second variant. In this manner, an observed frequency of each variant in the second set of variants is determined from the plurality of sequence reads in the plurality of first sequence reads that support or do not support a variant in the second set of variants. The observed frequency of each variant in the second set of variants is compared to a corresponding second reference frequency in a second reference set. In a plurality of such embodiments, each corresponding second reference frequency in the second reference set is a frequency of the corresponding variant across a plurality of second abnormal tissue samples of a common (same) second category. Further, in many such embodiments, said step of classifying said human subject further comprises: deeming said human subject to have a second condition associated with said plurality of second abnormal tissue samples when said observed frequency of each variant in said second set of variants satisfies a second threshold, wherein said second threshold is determined by each reference frequency in said second reference set.

[0390] In some embodiments, the subject is a human subject. In some embodiments, the biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid of the subject.

[0391] Another aspect of the present disclosure provides a computer system comprising one or more processors and a memory storing one or more programs executed by the one or more processors. The one or more programs include a plurality of instructions for classifying a subject using a method. The method includes: (A) obtaining a plurality of first sequence reads in electronic form from a biological sample of the subject, wherein the biological sample includes a plurality of cell-free nucleic acid molecules. The method further includes: (B) using the plurality of first sequence reads of the plurality of cell-free nucleic acid molecules to identify support for each variant in a first variant set. Here, another sequence read in the plurality of first sequence reads is considered to support a variant in the first variant set if the individual sequence read contains all or a portion of the variant, and another sequence read in the plurality of first sequence reads is considered to not support a variant in the first variant set if the individual sequence read does not contain the variant. In this manner, an observed frequency of each variant in the first variant set is determined from the plurality of sequence reads in the plurality of first sequence reads that support or do not support each variant in the first variant set. The observed frequency of each variant in the first set of variants is compared to a corresponding reference frequency in a first reference set. In many such embodiments, each corresponding reference frequency in the first reference set is a frequency of the corresponding variant across a plurality of first abnormal tissue samples of a common (same) first category. Next, the subject is classified. This classification step includes: when the observed frequency of each variant in the first set of variants meets a first threshold, the subject is deemed to have a first condition associated with the plurality of first abnormal tissue samples, wherein the first threshold is determined by each reference frequency in the first reference set.

[0392] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs for classifying a subject. The one or more programs are configured to be executed by a computer. The one or more programs include instructions for obtaining a plurality of first sequence reads in electronic form from a biological sample of the subject, wherein the biological sample includes a plurality of cell-free nucleic acid molecules. The one or more programs further include instructions for using the plurality of first sequence reads of the plurality of cell-free nucleic acid molecules to identify support for each variant in a first variant set. In many such embodiments, another sequence read in the plurality of first sequence reads is considered to support a variant in the first variant set when the individual sequence read contains all or a portion of the variant, and another sequence read in the plurality of first sequence reads is considered to not support a variant in the first variant set when the individual sequence read does not contain the variant. In this manner, an observed frequency for each variant in the first variant set is determined from the plurality of sequence reads in the plurality of first sequence reads that support or do not support each variant in the first variant set. The observed frequency for each variant in the first variant set is compared to a corresponding reference frequency in a first reference set. In many such embodiments, each corresponding reference frequency in the first reference set is a frequency of the corresponding variant across a plurality of first abnormal tissue samples of a common (same) first category. The one or more programs further include instructions for: classifying the subject. The classification step includes: deeming the subject to have a first condition associated with the plurality of first abnormal tissue samples when the observed frequency of each variant in the first variant set satisfies a first threshold. Here, the first threshold is determined by each reference frequency in the first reference set.

[0393] The cfDNA tumor fraction is estimated from read information from an individual without the need for direct analysis of the abnormal tissue.

[0394] In alternative embodiments, the cfDNA tumor fraction is not estimated using sequence reads 126 from an abnormal tissue. In some such embodiments, sequence reads 140 from a biological sample containing the cell-free nucleic acid are used to identify tumor-derived features (e.g., small variants). The potential tumor fraction is then estimated conditional on the observed frequency of one of these variants.

[0395] In some such embodiments, to ensure that a particular mutation is a suitable surrogate for the single estimated ctDNA fraction in the subject's cfDNA, the variant selected is one having a frequency other than the highest, based on the assumption that the variant has a high probability of not originating from the abnormal tissue. To illustrate, consider an example in which the cell-free nucleic acid of a biological sample is sequenced and a first variant 130-1 having a first frequency 132-1 and a second variant 130-2 having a second frequency 132-2 are found, where the first reference frequency 132-1 is higher than the second reference frequency 132-2. In this example, only the second variant 132-2 is assumed to be a suitable surrogate for a condition associated with an unmeasured abnormal tissue in the particular subject.

[0396] In some such embodiments, to ensure that a particular mutation is a suitable surrogate for a single estimated ctDNA fraction in the subject's cfDNA, variants known to be unassociated with the condition under study (e.g., variants frequently associated with leukocytes) are excluded from consideration.

[0397] In some embodiments, an additional variant 144 is used to estimate the tumor fraction based on the individual variant 144 having the second highest frequency of all variants observed in the biological sample (e.g., a blood sample) containing the cell-free nucleic acid. For example, if the frequency of this variant (the number of multiple observed sequence reads covering the position of the variant in the genome supporting the variant divided by the total number of multiple observed sequence reads covering the position of the variant in the genome) is ten percent, then the single estimated ctDNA fraction in the cfDNA of the subject is ten percent.

[0398] In some embodiments, an individual variant 144 of the first set of variants 142 is used to estimate a tumor fraction based on the individual variant having the third highest frequency of all variants observed in the biological sample (e.g., a blood sample) containing the cell-free nucleic acid. For example, if the frequency of the individual variant (the number of observed sequence reads supporting the variant covering the position of the variant in the genome divided by the total number of observed sequence reads covering the position of the variant in the genome) is 10 percent, then the single estimated ctDNA fraction in the cfDNA of the subject is 10 percent.

[0399] In some such embodiments, embodiments that do not utilize the abnormal tissue sample, but rather utilize only the biological sample containing the cell-free nucleic acid, can be used to calculate single estimated tumor fractions of less than about one percent.

[0400] In some embodiments, the variant ranked second highest in frequency is used as a proxy for the true tumor fraction (the single estimated ctDNA in the subject's cfDNA). For example, if the frequency of such a variant (the number of multiple observed sequence reads covering the position of the variant in the genome that support the variant divided by the total number of multiple observed sequence reads covering the position of the variant in the genome) is ten percent, then the single estimated ctDNA fraction in the subject's cfDNA is ten percent.

[0401] In some embodiments, the variant ranked third highest in frequency is used as a proxy for the true tumor fraction (the single estimated ctDNA in the subject's cfDNA). For example, if the frequency of such a variant (the number of observed sequence reads supporting the variant covering the position of the variant in the genome divided by the total number of observed sequence reads covering the position of the variant in the genome) is ten percent, then the single estimated ctDNA fraction in the subject's cfDNA is ten percent.

[0402] In some embodiments, the single estimated ctDNA in the cfDNA of the subject from the biological sample containing cell-free nucleic acid is used as a reference base for multiple biological samples obtained from the same subject at multiple subsequent time points to determine changes in the tumor score of the subject over time.

[0403] in conclusion

[0404] A plurality of examples may be provided as a single instance for multiple components, operations or structures described herein. Finally, the boundaries between various components, operations and data stores are arbitrary to some extent, and multiple specific operations are described in the context of a specific illustrative configuration. Other allocations of functions are conceivable and may fall within the scope of the described (multiple) embodiments. Typically, structures and functions presented as multiple individual components in the example configurations may be implemented as a combined structure or assembly. Similarly, structures and functions presented as a single assembly may be implemented as multiple individual components. These and other variations, modifications, additions and improvements all fall within the scope of the described (multiple) embodiments.

[0405] It should also be understood that although multiple terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of this disclosure, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject. The first subject and the second subject are both subjects, but they are not the same subject.

[0406] The terms used in this disclosure are for the purpose of describing a plurality of specific embodiments only and are not intended to limit the present invention. As used in the description of the present invention and the appended claims, the singular forms "a", "an" and "said" are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "and / or" used herein refer to and encompass any and all possible combinations of one or more associated listed items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements and / or components, but do not exclude the presence or increase of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0407] As used herein, the term "if" may be interpreted to mean "when," or "upon," or "in response to determining," or "in response to detecting," depending on the context. Likewise, the phrases "if it is determined" or "if [a stated condition or event] is detected" may be interpreted to mean "upon determination," or "in response to determining," or "upon detection of (the stated condition or event)," or "in response to detecting (the stated condition or event)," depending on the context.

[0408] The foregoing description includes a plurality of example systems, methods, techniques, instruction sequences, and computer program products that embody a plurality of illustrative embodiments. For purposes of explanation, numerous specific details have been set forth to provide an understanding of the various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that the various embodiments of the subject matter of the invention may be practiced without these specific details. Generally, a plurality of well-known instruction examples, experimental procedures, structures, and techniques have not been shown in detail.

[0409] For purposes of illustration, the foregoing description has been described with reference to several specific embodiments. However, the above illustrative discussion is ...

Claims

1. A method for determining tumor fraction in cell-free nucleic acid of a liquid biological sample of a subject, characterized by: The method comprises: In a computer system having one or more processors and a memory, wherein the memory stores one or more programs executed by the one or more processors, the following steps are performed in the computer system: (A) obtaining a plurality of first sequence reads in electronic form from the liquid biological sample of the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules; (B) using the plurality of first sequence reads to identify support for each variant in a first variant set, wherein the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to a region in a reference genome to determine whether the sequence read comprises all or a portion of the first variant; and the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome, thereby determining an observed frequency of each variant in the first variant set when another sequence read in the plurality of first sequence reads comprises the first when all or a portion of the first variant is present, the individual sequence read is considered to support the first variant in the first variant set; and when another sequence read among the plurality of first sequence reads does not include the first variant, the individual sequence read is considered to not support the first variant in the first variant set; and a number of multiple sequence reads among the plurality of first sequence reads supporting the first variant is used to determine the observed frequency of the first variant, and the variant frequency of the first variant in the liquid biological sample is estimated based on the observed frequency of the first variant; (C) for each individual variant in the first set of variants, obtaining a corresponding reference frequency for the individual variant in a first reference set, wherein each corresponding reference frequency in the first reference set is for another variant in a first abnormal solid tissue sample obtained from the subject; and (D) evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue to thereby determine a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

2. The method according to claim 1, wherein: A variant in the first set of variants is a single nucleotide variant associated with a predetermined genomic location, an insertion mutation associated with a predetermined genomic location, a deletion mutation associated with a predetermined genomic location, an alteration in the copy number of a cell, a nucleic acid rearrangement associated with a predetermined genomic site, or an abnormal methylation pattern associated with a predetermined genomic location.

3. The method according to claim 1, wherein: The subject is a human.

4. The method according to claim 1, wherein: The subject has cancer from a single primary site.

5. The method according to claim 1, wherein: The subject has cancer originating from two or more different organs.

6. The method according to claim 1, wherein: The subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, or a combination thereof.

7. The method according to claim 1, wherein: The subject has a predetermined stage of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, bladder cancer, or gastric cancer.

8. The method according to claim 1, wherein: The first abnormal solid tissue sample is a tumor sample.

9. The method according to claim 1, wherein: The first variant set consists of a single variant for a single genetic variation located at a single site in the genome of the subject.

10. The method according to claim 1, wherein: The first variant set consists of a first variant for a first genetic variation located at a first site in the genome of the subject and a second variant for a second genetic variation located at a second site in the genome of the subject.

11. The method according to claim 1, wherein: The first variant set consists of: a first variant for a first genetic variation located at a first site in the genome of the subject; a second variant for a second genetic variation located at a second site in the genome of the subject; and A third variant for a third genetic variation located at a third site in the genome of the subject.

12. The method according to claim 1, wherein: The first set of variants consists of between 2 and 20 variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject.

13. The method according to claim 1, wherein: The first set of variants consists of between 2 and 200 variants, wherein each variant in the first set of variants is a different genetic variation in the genome of the subject.

14. The method according to claim 1, wherein: The first variant set comprises 1000 variants, wherein each variant in the first variant set is a different genetic variation in the genome of the subject.

15. The method according to claim 1, wherein: The first variant set comprises 5,000 variants, wherein each variant in the first variant set is a different genetic variation in the genome of the subject.

16. The method according to claim 1, wherein: The using step (B) comprises comparing a sequence read from the plurality of first sequence reads to a lookup table of a plurality of variants to determine whether the sequence read comprises all or a portion of a first variant.

17. The method according to claim 1, wherein: The subject has stage II, stage III, or stage IV breast cancer, and the evaluating step (D) determines that the first tumor fraction of the cell-free nucleic acid is less than 1×10 -3 .

18. The method according to claim 1, wherein: The method further comprises the steps of: using the plurality of first sequence reads to identify support for each variant in a second set of variants, thereby determining an observed frequency for each variant in the second set of variants; for each individual variant in the second set of variants, obtaining a corresponding reference frequency for the individual variant in a second reference set, wherein each corresponding reference frequency in the second reference set is for another variant in a second abnormal solid tissue sample obtained from the subject; and The observed frequency of each individual variant in the second set of variants is evaluated against the observed frequency of the individual variant in the second reference set to determine a second tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

19. The method according to claim 18, wherein: When another sequence read in the plurality of first sequence reads comprises all or a portion of a variant in the second set of variants, treating the individual sequence read as supporting the variant; and When another sequence read among the plurality of first sequence reads does not include a variant in the second set of variants, the individual sequence read is considered to not support the variant.

20. The method of claim 18, wherein: The first abnormal solid tissue sample is composed of a first tumor fraction, and the second abnormal solid tissue sample is composed of a second tumor fraction from the same tumor of the subject.

21. The method of claim 18, wherein: The first abnormal solid tissue sample is of a first cancer type, and the second abnormal solid tissue sample is of a second cancer type.

22. The method according to claim 21, wherein: The first cancer type is the same as the second cancer type.

23. The method of claim 21, wherein: The first cancer type is different from the second cancer type.

24. The method of claim 23, wherein: The first cancer type and the second cancer type are each selected from the group consisting of breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, and gastric cancer.

25. The method of claim 1, wherein: The frequency of each variant in the first reference set is obtained by collectively obtaining a plurality of second sequence reads from the first abnormal solid tissue sample.

26. The method of claim 25, wherein: More than 1,000 sequence reads are obtained from the first abnormal entity tissue sample.

27. The method of claim 25, wherein: More than 3,000 sequence reads are obtained from the first abnormal entity tissue sample.

28. The method of claim 25, wherein: More than 5,000 sequence reads are obtained from the first abnormal entity tissue sample.

29. The method of claim 25, wherein: The method further comprises analyzing the plurality of second sequence reads obtained from the first abnormal solid tissue sample against a panel of variant candidates.

30. The method of claim 29, wherein: The variant candidate panel contains between 100 and 1000 variants.

31. The method of claim 25, wherein: The plurality of second sequence reads obtained from the first abnormal solid tissue sample represent whole genome data for each cell.

32. The method of claim 31, wherein: An average coverage of the plurality of second sequence reads obtained from the first abnormal solid tissue sample is at least 10-fold.

33. The method of claim 31, wherein: An average coverage of the plurality of second sequence reads obtained from the first abnormal solid tissue sample is at least 100-fold.

34. The method of claim 31, wherein: An average coverage of the plurality of second sequence reads obtained from the first abnormal solid tissue sample is at least 2000-fold.

35. The method of claim 1, wherein: The liquid biological sample comprises blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid or peritoneal fluid of the subject.

36. The method of claim 1, wherein: The liquid biological sample is composed of the subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid or peritoneal fluid.

37. The method of claim 1, wherein: The step (C) of evaluating the observed frequency of each individual variant in the first variant set against a corresponding reference frequency of the individual variant in the first reference set comprises evaluating a cumulative density function or a cumulative distribution function for the individual variant using the observed frequency and the reference frequency for the individual variant within a range of possible tumor fractions.

38. The method of claim 37, wherein: Use a cumulative density function.

39. The method of claim 38, wherein: The range is from 0% to 110%.

40. The method according to claim 38 or 39, wherein: The first tumor fraction is considered as a median of the cumulative density function.

41. The method of claim 37, wherein: Use a cumulative distribution function.

42. The method of claim 41, wherein: The cumulative distribution function has the following form: in: x=a 2i , an observed number of said plurality of sequence reads supporting said individual variant in said liquid biological sample; p=t*f 1i , where t is the possible first tumor fraction, and f 1i is the observed frequency of the individual variant in the first set of variants; and n=d 2i , a total number of sequence reads from the biological sample that map to the genomic location corresponding to the individual variant.

43. The method of claim 41, wherein: The cumulative distribution function has the following form: in: x=a 2i , an observed number of said plurality of sequence reads supporting said individual variant k in said liquid biological sample; p k =t*f 1i , where t is the possible first tumor fraction, and f 1i is the observed frequency of the individual variant k in the first set of variants; and n k =d 2i , a total number of sequence reads from the biological sample that map to the genomic position corresponding to the individual variant k.

44. The method of claim 37, wherein: The cumulative density function or the cumulative distribution function is derived under the assumption of negative binomial distribution.

45. The method of claim 1, wherein: The subject has a tumor fraction f of 0.100 or less.

46. The method of claim 1, wherein: The subject has a tumor fraction f of 0.050 or less.

47. A computer system, characterized in that: The computer system comprises: one or more processors; a memory storing one or more programs executed by the one or more processors; The one or more programs comprise a plurality of instructions for determining a tumor fraction in cell-free nucleic acid of a liquid biological sample of a subject by a method comprising the steps of: (A) obtaining a plurality of first sequence reads in electronic form from the liquid biological sample of the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules; (B) using the plurality of first sequence reads to identify support for each variant in a first variant set, wherein the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to a region in a reference genome to determine whether the sequence read comprises all or a portion of a first variant; and the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome, thereby determining an observed frequency for each variant in the first variant set, and treating the individual sequence read as supporting the first variant in the first variant set when another sequence read in the plurality of first sequence reads comprises all or a portion of the first variant; and, When another sequence read in the plurality of first sequence reads does not include the first variant, the individual sequence read is considered to not support the first variant in the first variant set; and a number of a plurality of sequence reads in the plurality of first sequence reads supporting the first variant and a number of a plurality of sequence reads in the plurality of first sequence reads not supporting the first variant to determine the observed frequency of the first variant, and estimating a variant frequency of the first variant in the liquid biological sample from the observed frequency of the first variant; (C) for each individual variant in the first set of variants, obtaining a corresponding reference frequency for the individual variant in a first reference set, wherein each corresponding reference frequency in the first reference set is for another variant in a first abnormal solid tissue sample obtained from the subject; and (D) evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue to thereby determine a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

48. A non-transitory computer-readable storage medium storing one or more programs for determining a tumor fraction in cell-free nucleic acid of a liquid biological sample of a subject, the one or more programs configured for execution by a computer, characterized in that: The one or more programs include multiple instructions for the following steps: (A) obtaining a plurality of first sequence reads in electronic form from the liquid biological sample of the subject, wherein the liquid biological sample comprises a plurality of cell-free nucleic acid molecules; (B) using the plurality of first sequence reads to identify support for each variant in a first variant set, wherein the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to a region in a reference genome to determine whether the sequence read comprises all or a portion of a first variant; and the use of step (B) comprises: aligning a sequence read in the plurality of first sequence reads to each entry in a lookup table, wherein each entry in the lookup table represents a different portion of a genome, thereby determining an observed frequency for each variant in the first variant set, and treating the individual sequence read as supporting the first variant in the first variant set when another sequence read in the plurality of first sequence reads comprises all or a portion of the first variant; and, When another sequence read in the plurality of first sequence reads does not include the first variant, the individual sequence read is considered to not support the first variant in the first variant set; and a number of a plurality of sequence reads in the plurality of first sequence reads supporting the first variant and a number of a plurality of sequence reads in the plurality of first sequence reads not supporting the first variant to determine the observed frequency of the first variant, and estimating a variant frequency of the first variant in the liquid biological sample from the observed frequency of the first variant; (C) for each individual variant in the first set of variants, obtaining a corresponding reference frequency for the individual variant in a first reference set, wherein each corresponding reference frequency in the first reference set is for another variant in a first abnormal solid tissue sample obtained from the subject; and (D) evaluating the observed frequency of each individual variant in the first set of variants against the observed frequency of the individual variant in the first reference set in the first abnormal solid tissue to thereby determine a first tumor fraction in the cell-free nucleic acid of the liquid biological sample of the subject.

Citation Information

Patent Citations

  • Models for Targeted Sequencing

    US20190164627A1

  • Identifying copy number aberrations

    US20190287646A1

  • Methods for using mosaicism in nucleic acids sampled distal to their origin

    WO2016070131A1