Methods and systems for differentiating somatic and germline variants

By employing a beta-binomial distribution to model expected germline mutation allele counts and classify variants based on p-values, the method addresses the challenges of distinguishing somatic and germline variants in cell-free DNA, achieving improved accuracy in variant origin determination.

JP2025081596APending Publication Date: 2025-05-27GUARDANT HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025027106
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-09-20
Filing Date
2025-02-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Current methods for distinguishing somatic and germline variants in nucleic acid samples, particularly in cell-free DNA, face challenges in accurately modeling the variance in nucleic acid molecule counts and adjusting for factors like allelic imbalance due to copy number variation or loss of heterozygosity.

Method used

The method involves determining quantitative measurements for nucleic acid variants, identifying associated variables, and generating a statistical model for expected germline mutation allele counts using a beta-binomial distribution. This model estimates the probability value (p-value) for classifying variants as somatic or germline based on deviations from expected germline mutant allele fractions.

Benefits of technology

This approach enables accurate classification of somatic and germline variants by effectively modeling the variance in nucleic acid molecule counts and adjusting for genetic and environmental factors, thereby improving the accuracy of variant origin determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081596000001
    Figure 2025081596000001
  • Figure 2025081596000002
    Figure 2025081596000002
  • Figure 2025081596000003
    Figure 2025081596000003
Patent Text Reader

Abstract

To provide methods and systems for differentiating somatic and germline variants.SOLUTION: A method of the present invention comprises: determining a quantitative measure for a nucleic acid variant comprising the total allele count and minor allele count for the nucleic acid variant; identifying an associated variable of the nucleic acid variant; determining a quantitative value for the associated variable; generating a statistical model for the expected germline mutant allele count at a genomic locus of the nucleic acid variant; generating a probability value (p-value) for the nucleic acid variant based on at least one of at least in part of the statistical model, quantitative value and quantitative measure; and classifying the nucleic acid variant.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 62 / 561,048, filed Sep. 20, 2017, which is incorporated by reference in its entirety. [Background technology]

[0002] background An important aspect of cancer genomics is precisely identifying the origin of genetic alterations for appropriate treatment of patients. Recent studies have found that in >2% of patients with advanced cancer, unidentified germline alterations were found incidentally during next generation sequencing (NGS) for targetable somatic alterations. However, tissue-based NGS may not be able to accurately distinguish between germline and somatic variants without comparison to normal tissue. In plasma, somatic variants typically occur at mutant allele fractions (MAFs) that can be 1-2 orders of magnitude lower than germline variants, and thus liquid biopsies can accurately assign germline / somatic origin. However, certain factors, such as allelic imbalance from copy number variation (CNV) or loss of heterozygosity (LOH), can skew germline MAFs from the expected range for germline MAFs. Thus, there is a need for methods that can take these factors into account when determining the origin of a variant. Summary of the Invention [Means for solving the problem]

[0003] Abstract The present disclosure provides methods and systems for distinguishing somatic and germline variants in a sample of nucleic acid molecules, such as cell-free deoxyribonucleic acid (cfDNA). Such methods can use common single nucleotide polymorphisms (SNPs) to model local germline allele count behavior and can distinguish somatic variants based on MAF deviation from the observed germline MAF.

[0004] In one aspect, the disclosure provides a method for identifying somatic or germline origin of a nucleic acid variant from a sample of nucleic acid molecules (e.g., a tissue sample, a sample of cell-free DNA, and / or the like). The method includes (a) determining one or more quantitative measurements for the nucleic acid variant from the nucleic acid sample. The quantitative measurements include total allele count and minor allele count for the nucleic acid variant. The method also includes (b) identifying at least one associated variable of the nucleic acid variant from the nucleic acid sample, and (c) determining a quantitative value for the associated variable of the nucleic acid variant. The method further includes (d) generating a statistical model for expected germline mutation allele count at a genomic locus of the nucleic acid variant, and (e) generating a probability value (p-value) for the nucleic acid variant based on at least one of the statistical model for expected germline allele count, the quantitative value for the associated variable of the nucleic acid variant, and the quantitative measurement for the nucleic acid variant. Moreover, the method also includes (f) classifying the nucleic acid variant as (i) somatic in origin when the p-value for the nucleic acid variant is below a threshold, or (ii) germline in origin when the p-value for the nucleic acid variant is at or above a threshold.

[0005] In one aspect, the disclosure provides a method of identifying somatic or germline origin of a nucleic acid variant from a sample of cell-free nucleic acid molecules (e.g., cell-free deoxyribonucleic acid (cfDNA) molecules), comprising: (a) determining a plurality of quantitative measurements for the nucleic acid variants from the sample of cell-free nucleic acid molecules, the plurality of quantitative measurements comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying associated variables of the nucleic acid variants from the sample of cell-free nucleic acid molecules; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) determining a quantitative value for the associated variable of the nucleic acid variant that is expected at a genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on the statistical model for expected germline mutant allele counts, the quantitative value for the associated variable of the nucleic acid variant, and at least one of a plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0006] In some embodiments, the method further comprises obtaining a sample of cell-free nucleic acid molecules from a subject.In some embodiments, the method further comprises receiving sequencing information generated from the sample of cell-free nucleic acid molecules, the sequencing information comprises cell-free nucleic acid sequencing reads comprising nucleic acid variants and associated variables of the nucleic acid variants, and the associated variables comprise at least one heterozygous single nucleotide polymorphism (het SNP) in the genomic region defined for the nucleic acid variants.In some embodiments, the method further comprises sequencing nucleic acid from the sample of cell-free nucleic acid molecules to generate sequencing information, and a plurality of quantitative measurements for the nucleic acid variants and quantitative values ​​for the associated variables are determined from the sequencing information.

[0007] In some embodiments, the method further comprises determining a plurality of quantitative measurements for the nucleic acid variants, identifying associated variables for the nucleic acid variants, and determining quantitative values ​​for the associated variables from sequencing information generated from the sample of cell-free nucleic acid molecules. In some embodiments, the method of any of claims 1 to 5 further comprises generating a predetermined threshold value using a beta-binomial model of expected germline mutant allele counts for the nucleic acids of the sample of cell-free nucleic acid molecules. In some embodiments, the method further comprises classifying somatic or germline origins of the plurality of nucleic acid variants from a plurality of genomic loci within the sample of cell-free nucleic acid molecules.

[0008] In some embodiments, the associated variable of the nucleic acid variant comprises at least one heterozygous single nucleotide polymorphism (het SNP). In some embodiments, the associated variable of the nucleic acid variant comprises at least two het SNPs. In some embodiments, the associated variable of the nucleic acid variant comprises a genomic locus that is linked to the genomic locus that contains the nucleic acid variant.

[0009] In some embodiments, the method further comprises determining one or more mean and / or variance values ​​of mutant allele counts for the associated variables of the nucleic acid variants. In some embodiments, the method further comprises determining an average quantitative value for the associated variables of the nucleic acid variants. In some embodiments, the associated variables of the nucleic acid variants include one or more of heterozygous single nucleotide polymorphisms (het SNPs), GC content measurements, probe-specific bias measurements, fragment length values, sequencing statistics measurements, copy number breakpoints, and clinical data for the subject. In some embodiments, the method further comprises determining the mean and / or variance values ​​of the associated variables of the nucleic acid variants.

[0010] In some embodiments, the method further comprises determining a local germline folded mutant allele fraction (MAF), μbin, for the nucleic acid variant, where bin is a gene or another defined genomic region that contains the nucleic acid variant, and the folded MAF is min(MAF,1-MAF). In some embodiments, the defined genomic region is within about 10 of the nucleic acid variant. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a population allele frequency (AF) of greater than about 0.001. In some embodiments, the associated variables of the nucleic acid variants include at least one non-oncogenic single nucleotide polymorphism (SNP). In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a mutant allele fraction (MAF) of less than about 0.9.

[0011] In some embodiments, the associated variable comprises at least one heterozygous single nucleotide polymorphism (SNP) in the genomic region defined for the nucleic acid variant, and the method comprises estimating a beta binomial distribution parameter using: (x,y)~beta binomial(μ bin ,ρ), where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each germline heterozygous SNP; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each germline heterozygous SNP; μ bin p = an estimate of the average mutant allele count of heterozygous SNPs in a bin, where the bin is a genomic region defined for the nucleic acid variant, and ρ = an estimate of the dispersion parameter. In some embodiments, the method further comprises calculating upper and lower bounds on the p-value. In some embodiments, the method further comprises: p-value = 2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin , ρ,B)) for calculating a two-sided p-value for the nucleic acid variant, bb = beta-binomial probability, x' = random variable distributed with beta-binomial, A = mutant allele count of the nucleic acid variant, and B = total molecular count of the nucleic acid variant. In some embodiments, ρ comprises a median of at least one set of ρ values ​​from a past sample set. In some embodiments, the method further comprises replacing the median ρ parameter with a function of the GC content of the nucleic acid variant. In some embodiments, the method further comprises: bin In some embodiments, the method further comprises determining a maximum likelihood estimate of μ binIn some embodiments, the method further comprises determining a mean estimate of p. In some embodiments, the method further comprises determining a maximum likelihood estimate of p. In some embodiments, the method further comprises determining a variance estimate of p. In some embodiments, the method further comprises generating a report in electronic and / or paper format that provides an indication of the classification of the nucleic acid variants, either somatic or germline origin.

[0012] In another aspect, the disclosure provides a method, when executed by at least one electronic processor, comprising: (a) determining a plurality of quantitative measurements for nucleic acid variants from sequencing information generated from a sample of cell-free nucleic acid molecules (e.g., cell-free deoxyribonucleic acid (cfDNA) molecules), the plurality of quantitative measurements including total and minor allele counts for the nucleic acid variants; (b) identifying associated variables of the nucleic acid variants from the sequencing information; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) determining a quantitative value for an expected germline mutation allele count at a genomic locus where the nucleic acid variant lies. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on the statistical model for expected germline mutant allele counts, the quantitative value for the associated variable of the nucleic acid variant, and at least one of the plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0013] In some embodiments, the predetermined threshold is generated using a beta-binomial model of expected germline mutant allele counts for the sample of cell-free nucleic acid molecules (e.g., cfDNA molecules). In some embodiments, the nucleic acid variant associated variable comprises at least one heterozygous single nucleotide polymorphism (het SNP). In some embodiments, the nucleic acid variant associated variable comprises at least two het SNPs. In some embodiments, the nucleic acid variant associated variable comprises a genomic locus linked to the genomic locus containing the nucleic acid variant. In some embodiments, one or more mean and / or variance values ​​of mutant allele counts are determined for the nucleic acid variant associated variable. In some embodiments, at least one of the plurality of quantitative measurements comprises the number of nucleic acid molecules of the sample of cell-free nucleic acid molecules that contain the nucleic acid variant. In some embodiments, the nucleic acid variant associated variable comprises one or more of a heterozygous single nucleotide polymorphism (het SNP), a GC content measurement, a probe-specific bias measurement, a fragment length value, a sequencing statistical measurement, a copy number breakpoint, and clinical data for the subject.

[0014] In some embodiments, a local germline folded mutant allele fraction (MAF), μbin, is determined for a nucleic acid variant, where bin is a gene or another defined genomic region that contains the nucleic acid variant, and the folded MAF is min(MAF,1-MAF). In some embodiments, the defined genomic region is within about 10 of the nucleic acid variant. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a population allele frequency (AF) of greater than about 0.001. In some embodiments, the associated variables of the nucleic acid variants include at least one non-oncogenic single nucleotide polymorphism (SNP). In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a mutant allele fraction (MAF) of less than about 0.9.

[0015] In some embodiments, the associated variable comprises at least one heterozygous single nucleotide polymorphism (SNP) within the genomic region defined for the nucleic acid variant, and the beta binomial distribution parameters are estimated using: (x,y) ~ beta binomial (μ bin ,ρ), where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each of the at least one germline heterozygous SNP; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each of the at least one germline heterozygous SNP; μ bin p = estimate of the mutant allele count of heterozygous SNPs in a bin, where the bin is a genomic region defined for the nucleic acid variant, and ρ = estimate of the dispersion parameter. In some embodiments, upper and lower bounds on the p-value are calculated. In some embodiments, a two-sided p-value for a nucleic acid variant is calculated as p-value = 2*min(Pr bb (x'>x|μ bin ,ρ,B),Pr bb (x' <x|μ bin ,ρ,B)) where Pr bb = beta-binomial probability, x' = a random variable distributed with beta-binomial, A = mutant allele count of the nucleic acid variant, and B = total molecular count of the nucleic acid variant.

[0016] In another aspect, the disclosure provides a method, when executed by at least one electronic processor, comprising: (a) determining a plurality of quantitative measurements for nucleic acid variants from sequencing information generated from a sample of nucleic acid molecules (e.g., a sample of cell-free deoxyribonucleic acid (cfDNA) molecules), the plurality of quantitative measurements including total and minor allele counts for the nucleic acid variants; (b) identifying associated variables of the nucleic acid variants from the sequencing information; (c) determining quantitative values ​​for the associated variables of the nucleic acid variants; and (d) generating a statistical model for expected germline mutation allele counts at a genomic locus at which the nucleic acid variant lies. (e) generating a probability value (p-value) for the nucleic acid variant based, at least in part, on at least one of a statistical model for expected germline mutant allele counts, a quantitative value for an associated variable of the nucleic acid variant, and a plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0017] In some embodiments, the system comprises a nucleic acid sequencing device operably connected to the controller, the nucleic acid sequencing device configured to provide sequencing information from the nucleic acid of a sample of nucleic acid molecules (e.g., acellular nucleic acid molecules). In some embodiments, the system comprises a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the nucleic acid of the sample to be sequenced by the nucleic acid sequencing device. In some embodiments, the system comprises a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify the nucleic acid of the sample. In some embodiments, the system comprises a material transport component operably connected to the controller, the material transport component configured to transport one or more materials between the nucleic acid sequencing device and the sample preparation component.

[0018] In some embodiments, the pre-defined threshold is generated using a beta-binomial model of expected germline mutant allele counts for the nucleic acid of the sample (e.g., cfDNA molecule). In some embodiments, the associated variable of the nucleic acid variant comprises at least one heterozygous single nucleotide polymorphism (het SNP). In some embodiments, the associated variable of the nucleic acid variant comprises at least two het SNPs. In some embodiments, the associated variable of the nucleic acid variant comprises a genomic locus that is linked to the genomic locus that contains the nucleic acid variant.

[0019] In some embodiments, the mean and / or variance of one or more mutant allele counts are determined for the associated variables of the nucleic acid variant. In some embodiments, the p-value is used to classify the nucleic acid variant. In some embodiments, at least one of the plurality of quantitative measurements comprises the number of nucleic acid molecules of the sample of cell-free nucleic acid molecules that contain the nucleic acid variant. In some embodiments, the associated variables comprise one or more of heterozygous single nucleotide polymorphisms (het SNPs), GC content measurements, probe-specific bias measurements, fragment length values, sequencing statistics measurements, copy number breakpoints, and clinical data for the subject.

[0020] In some embodiments, a local germline folded mutant allele fraction (MAF), μbin, is determined for a nucleic acid variant, where bin is a gene or another defined genomic region that contains the nucleic acid variant, and the folded MAF is min(MAF,1-MAF). In some embodiments, the defined genomic region is within about 10 of the nucleic acid variant. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a population allele frequency (AF) of greater than about 0.001. In some embodiments, the associated variables of the nucleic acid variants include at least one non-oncogenic single nucleotide polymorphism (SNP). In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) with a mutant allele fraction (MAF) of less than about 0.9.

[0021] In some embodiments, the associated variable comprises at least one heterozygous SNP within the genomic region defined for the nucleic acid variant, and the beta binomial distribution parameters are estimated using: (x,y) ~ beta binomial (μ bin ,ρ), where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each germline heterozygous SNP; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each germline heterozygous SNP; μ bin p = estimate of the mutant allele count of heterozygous SNPs in a bin, where the bin is a genomic region defined for the nucleic acid variant, and ρ = estimate of the dispersion parameter. In some embodiments, upper and lower bounds on the p-value are calculated. In some embodiments, a two-sided p-value for a nucleic acid variant is calculated as p-value = 2*min(Pr bb (x'>x|μ bin ,ρ,B),Pr bb (x' <x|μ bin ,ρ,B)) where Pr bb = beta-binomial probability, x' = a random variable distributed with beta-binomial, A = mutant allele count of the nucleic acid variant, and B = total molecular count of the nucleic acid variant.

[0022] In another aspect, the disclosure provides a method of identifying somatic or germline origin of a nucleic acid variant from a sample of cell-free deoxyribonucleic acid (cfDNA) molecules, comprising the steps of: (a) determining a mutant allele count (A) and a total molecular count (B) of a nucleic acid variant from the sample of cfDNA molecules; (b) identifying at least one germline heterozygous single nucleotide polymorphism (SNP) within a genomic region defined for the nucleic acid variant; (c) determining a total molecular count (y) and a mutant allele count (y) of the at least one germline heterozygous SNP; and (d) determining a mutant allele count (y) of the at least one germline heterozygous SNP from the sample of cfDNA molecules. bin and determining an estimate of ρ from a beta binomial distribution, bin ,ρ), where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each germline heterozygous SNP; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each germline heterozygous SNP; μ bin = an estimate of the mutant allele count of a germline heterozygous SNP within a bin, where the bin is a genomic region defined for the nucleic acid variant; and ρ = an estimate of the dispersion parameter; and (ii) calculating a two-sided p-value from the equation: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B)), where Pr bbwhere x'=a beta-binomial probability, x'=a random variable distributed with a beta-binomial distribution, A=the mutant allele count of the nucleic acid variant, and B=the total molecular count of the nucleic acid variant; and (e) classifying the nucleic acid variant as (i) of somatic origin when the p-value is below a predetermined threshold, or (ii) of germline origin when the p-value is at or above a predetermined threshold.

[0023] In some embodiments, p comprises a median value of at least one set of p values ​​from a past sample set. bin In some embodiments, the method includes determining a maximum likelihood estimate of μ bin In some embodiments, the method comprises determining a mean estimate of p. In some embodiments, the method comprises determining a maximum likelihood estimate of p. In some embodiments, the method comprises determining a variance estimate of p. In some embodiments, the method further comprises generating a report in electronic and / or paper format that provides an indication of the classification of the nucleic acid variants, either somatic or germline origin.

[0024] In another aspect, the disclosure provides a system comprising: a communications interface that obtains sequencing information generated from nucleic acids of a sample of nucleic acid molecules (e.g., a sample of cell-free deoxyribonucleic acid (cfDNA) molecules) over a communications network; and a computer in communication with the communications interface, the computer comprising at least one computer processor and a non-transitory computer readable medium comprising machine executable code, which when executed by the at least one computer processor, performs the steps of: (a) determining a plurality of quantitative measures for nucleic acid variants from the sequencing information, the plurality of quantitative measures including total allele counts and minor allele counts for the nucleic acid variants; and (b) identifying associated variables of the nucleic acid variants from the sequencing information. (c) determining a quantitative value for the associated variable of the nucleic acid variant; (d) generating a statistical model for expected germline mutant allele counts at a genomic locus where the nucleic acid variant is located; (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on the statistical model for expected germline mutant allele counts, the quantitative value for the associated variable of the nucleic acid variant, and at least one of the plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0025] In some embodiments, the sequencing information is provided by a nucleic acid sequencing device. In some embodiments, the nucleic acid sequencing device performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, sequencing by ligation, or sequencing by hybridization of nucleic acids to generate sequencing information. In some embodiments, the nucleic acid sequencing device generates sequencing information using a clonal single molecule array derived from a sequencing library. In some embodiments, the nucleic acid sequencing device comprises a chip having an array of microwells for sequencing the sequencing library and generating sequencing information. In some embodiments, the non-transitory computer readable medium comprises a memory, a hard drive, or a memory or hard drive of a computer server. In some embodiments, the communication network comprises one or more computer servers capable of distributed computing. In some embodiments, the distributed computing is cloud computing. In some embodiments, the computer is part of a computer server located at a remote location from the nucleic acid sequencing device. In some embodiments, the system further includes an electronic display in communication with the computer via a network, the electronic display including a user interface for displaying results responsive to implementing at least a portion of (a)-(f). In some embodiments, the user interface is a graphical user interface (GUI) or a web-based user interface. In some embodiments, the electronic display is part of a personal computer. In some embodiments, the electronic display is part of an Internet-enabled computer. In some embodiments, the Internet-enabled computer is located remotely from the computer. In some embodiments, the non-transitory computer-readable medium comprises a memory, a hard drive, or a memory or a hard drive of a computer server.In some embodiments, the communications network includes a telecommunications network, the Internet, an extranet, or an intranet.

[0026] In another aspect, the disclosure provides a method of treating a disease in a subject, the method comprising administering one or more customized therapies to the subject, thereby treating the disease in the subject, the customized therapies comprising: (a) determining one or more quantitative measures for a nucleic acid variant from a sample of nucleic acid molecules (e.g., a sample of cell-free DNA), the quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variant; (b) identifying at least one associated variable of the nucleic acid variant from the sample of nucleic acid molecules; (c) determining a quantitative value for the associated variable of the nucleic acid variant; (d) generating a statistical model for expected germline mutation allele counts at a genomic locus of the nucleic acid variant; and (e) determining a quantitative value for the associated variable of the nucleic acid variant. (f) generating a probability value (p-value) for the nucleic acid variant based on at least one of a statistical model for the germline allele counts generated, a quantitative value for an associated variable of the nucleic acid variant, and a quantitative measure for the nucleic acid variant; (i) classifying the nucleic acid variant as being of somatic origin when the p-value for the nucleic acid variant is below a threshold, or (ii) as being of germline origin when the p-value for the nucleic acid variant is at or above a threshold; (g) comparing the classified nucleic acid variant to one or more comparator results indexed with the one or more therapies; and (h) identifying one or more customized therapies for treating the disease in the subject when a substantial match exists between the classified nucleic acid variant and the comparator result.

[0027] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Thus, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0028] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Thus, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. In an embodiment of the present invention, for example, the following items are provided: (Item 1) 1. A method for identifying somatic or germline origin of nucleic acid variants from a sample of cell-free deoxyribonucleic acid (cfDNA) molecules, comprising: (a) determining a plurality of quantitative measures for the nucleic acid variants from the cfDNA sample, the plurality of quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying associated variables of said nucleic acid variants from said sample of cfDNA molecules; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) generating a statistical model for expected germline mutant allele counts at the genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on a statistical model for the expected germline mutant allele count, a quantitative value for an associated variable of the nucleic acid variant, and at least one of a plurality of quantitative measures for the nucleic acid variant; (f) classifying the nucleic acid variant as (i) somatic in origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) germline in origin when the p-value for the nucleic acid variant is at or above the predetermined threshold; A method comprising: (Item 2) 2. The method of claim 1, further comprising obtaining a sample of the cfDNA molecules from a subject. (Item 3) 3. The method of claim 1 or 2, further comprising receiving sequencing information generated from the cfDNA sample, the sequencing information comprising cfDNA sequencing reads comprising the nucleic acid variant and an associated variable of the nucleic acid variant, the associated variable comprising at least one heterozygous single nucleotide polymorphism (het SNP) within a genomic region defined for the nucleic acid variant. (Item 4) 20. The method of claim 19, further comprising sequencing nucleic acids from the cfDNA sample to generate sequencing information, wherein a plurality of quantitative measures for the nucleic acid variants and quantitative values ​​for the associated variables are determined from the sequencing information. (Item 5) 20. The method of claim 19, further comprising determining a plurality of quantitative measurements for the nucleic acid variants, identifying associated variables of the nucleic acid variants, and determining quantitative values ​​for the associated variables from sequencing information generated from the sample of cfDNA molecules. (Item 6) The method of any preceding item, further comprising generating the predetermined threshold using a beta-binomial model of expected germline mutant allele counts for nucleic acids of the sample of cfDNA molecules. (Item 7) 20. The method of any preceding item, further comprising classifying somatic or germline origin of the plurality of nucleic acid variants from a plurality of genomic loci within the sample of cfDNA molecules. (Item 8) The method of any of the preceding claims, wherein the associated variable of the nucleic acid variant comprises at least one heterozygous single nucleotide polymorphism (het SNP). (Item 9) 9. The method of claim 8, wherein the associated variable of the nucleic acid variant comprises at least two het SNPs. (Item 10) The method of any of the preceding items, wherein the associated variable of the nucleic acid variant comprises a genomic locus that is linked to the genomic locus that contains the nucleic acid variant. (Item 11) The method of any preceding claim, further comprising determining the mean and / or variance of one or more mutant allele counts for the associated variables of the nucleic acid variants. (Item 12) 20. The method of any preceding claim, further comprising determining a mean quantitative value for the associated variable of the nucleic acid variants. (Item 13) The method of any of the preceding items, wherein the associated variables of the nucleic acid variants comprise one or more of a heterozygous single nucleotide polymorphism (het SNP), a GC content measurement, a probe specific bias measurement, a fragment length value, a sequencing statistic measurement, a copy number breakpoint, and clinical data regarding the subject. (Item 14) 2. The method of any of the preceding claims, further comprising determining the mean and / or variance of the associated variables of the nucleic acid variants. (Item 15) 2. The method of any preceding claim, further comprising determining a local germline folded mutant allele fraction (MAF), μ bin, for said nucleic acid variant, where bin is a gene or another defined genomic region that contains said nucleic acid variant, and the folded MAF is min(MAF,1-MAF). (Item 16) The defined genomic region comprises about 10 of the nucleic acid variants. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 16. The method according to item 15, wherein the region is within base pairs. (Item 17) The method of any of the preceding items, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) comprising a population allele frequency (AF) of greater than about 0.001. (Item 18) The method of any of the preceding items, wherein the associated variable of the nucleic acid variant comprises at least one non-oncogenic single nucleotide polymorphism (SNP). (Item 19) The method of any of the preceding items, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) comprising a mutant allele fraction (MAF) of less than about 0.9. (Item 20) The associated variable comprises at least one heterozygous single nucleotide polymorphism (SNP) within a genomic region defined for the nucleic acid variant, and the method further comprises estimating a beta binomial distribution parameter using: (x,y)~Beta binomial (μ bin ,ρ) During the ceremony, y=a vector of total molecular counts of said germline heterozygous SNPs, with one entry for each germline heterozygous SNP identified in (b); a vector of x=min(mutant allele count of said germline heterozygous SNP, y-mutant allele count of said germline heterozygous SNP), with one entry for each germline heterozygous SNP identified in (b); μ bin = an estimate of the average mutant allele count of heterozygous SNPs within a bin, the bin being a genomic region defined for the nucleic acid variant; and ρ = an estimate of the dispersion parameter. 2. The method according to any of the preceding items. (Item 21) Calculating a two-sided p-value for said nucleic acid variant using: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) During the ceremony, Pr bb = Beta binomial probability, x' = a random variable distributed with the beta binomial, A=the mutant allele count of the nucleic acid variant; B=total molecular count of the nucleic acid variant; The method according to item 20. (Item 22) 21. The method of claim 20, wherein p comprises a median of at least one set of p values ​​from a historical sample set. (Item 23) 23. The method of claim 22, further comprising replacing the median ρ parameter with a function of the GC content of the nucleic acid variant. (Item 24) μ bin 25. The method of claim 20, further comprising determining a maximum likelihood estimate of μ bin 21. The method of claim 20, further comprising determining a mean value estimate of (Item 26) Item 21. The method of item 20, further comprising the step of determining a maximum likelihood estimate of p. (Item 27) Item 21. The method of item 20, further comprising the step of determining a variance estimate of p. (Item 28) 2. The method of any preceding claim, further comprising the step of calculating upper and lower bounds on the p-value. (Item 29) When executed by at least one electronic processor, (a) determining a plurality of quantitative measures for nucleic acid variants from sequencing information generated from a cell-free deoxyribonucleic acid (cfDNA) sample, the plurality of quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying associated variables of said nucleic acid variants from said sequencing information; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) generating a statistical model for expected germline mutant allele counts at the genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on a statistical model for the expected germline mutant allele count, a quantitative value for an associated variable of the nucleic acid variant, and at least one of a plurality of quantitative measures for the nucleic acid variant; (f) classifying the nucleic acid variant as (i) somatic in origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) germline in origin when the p-value for the nucleic acid variant is at or above the predetermined threshold; A non-transitory computer readable medium comprising computer executable instructions implementing a method, comprising: (Item 30) 30. The non-transient computer-readable medium of claim 29, wherein the predetermined threshold is generated using a beta-binomial model of expected germline mutant allele counts for nucleic acids of the cfDNA sample. (Item 31) 31. The non-transitory computer-readable medium of any one of items 29-30, wherein the associated variable of the nucleic acid variant comprises at least one heterozygous single nucleotide polymorphism (het SNP). (Item 32) 32. The non-transient computer-readable medium of claim 31, wherein the associated variable of the nucleic acid variant comprises at least two het SNPs. (Item 33) 33. The non-transient computer readable medium of any one of paragraphs 29-32, wherein the associated variable of the nucleic acid variant comprises a genomic locus linked to the genomic locus comprising the nucleic acid variant. (Item 34) 34. The non-transitory computer readable medium of any one of items 29-33, wherein a mean and / or variance of one or more mutant allele counts is determined for associated variables of the nucleic acid variants. (Item 35) 35. The non-transitory computer-readable medium of any one of items 29-34, wherein at least one of the plurality of quantitative measures comprises a number of nucleic acid molecules of the cfDNA sample that contain the nucleic acid variant. (Item 36) The non-transitory computer-readable medium of any one of claims 29 to 35, wherein the associated variables of the nucleic acid variants include one or more of a heterozygous single nucleotide polymorphism (het SNP), a GC content measure, a probe-specific bias measure, a fragment length value, a sequencing statistic measure, a copy number breakpoint, and clinical data about the subject. 37. The non-transient computer readable medium of any one of items 29 to 36, wherein a local germline folded mutant allele fraction (MAF), μ bin, is determined for the nucleic acid variant, where bin is a gene or another defined genomic region that contains the nucleic acid variant, and the folded MAF is min(MAF,1-MAF). (Item 38) The defined genomic region comprises about 10 of the nucleic acid variants. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 38. The non-transitory computer-readable medium of Item 37, wherein the region is within base pairs. (Item 39) 39. The non-transitory computer-readable medium of any one of items 29 to 38, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) having a population allele frequency (AF) of greater than about 0.001. (Item 40) 40. The non-transitory computer-readable medium of any one of claims 29 to 39, wherein the associated variable comprises at least one non-oncogenic single nucleotide polymorphism (SNP). (Item 41) 41. The non-transitory computer-readable medium of any one of claims 29 to 40, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) comprising a mutant allele fraction (MAF) of less than about 0.9. (Item 42) The associated variable comprises at least one heterozygous single nucleotide polymorphism (SNP) within the genomic region defined for the nucleic acid variant, and a beta binomial distribution parameter is estimated using: (x,y)~Beta binomial (μ bin ,ρ) During the ceremony, y=a vector of total molecular counts of said germline heterozygous SNPs, with one entry for each germline heterozygous SNP identified in (b); a vector of x=min(mutant allele count of said germline heterozygous SNP, y-mutant allele count of said germline heterozygous SNP), with one entry for each germline heterozygous SNP identified in (b); μ bin = an estimate of the mutant allele count of a heterozygous SNP within a bin, the bin being a genomic region defined for the nucleic acid variant, ρ = Estimate of the variance parameter, 42. The non-transitory computer-readable medium according to any one of items 29 to 41. (Item 43) 43. The non-transitory computer-readable medium of any one of claims 29 to 42, wherein an upper bound and a lower bound on the p-value are calculated. (Item 44) The two-sided p-value for the nucleic acid variant is calculated using: p-value=2*min(Pr bb (x'>x|μ bin ,ρ,B),Pr bb (x' <x|μ bin ,ρ,B) During the ceremony, Pr bb = Beta binomial probability, x' = a random variable distributed with the beta binomial, A=the mutant allele count of the nucleic acid variant; B=total molecular count of the nucleic acid variant; Item 44. The non-transitory computer-readable medium of item 43. (Item 45) When executed by at least one electronic processor, (a) determining a plurality of quantitative measures for nucleic acid variants from sequencing information generated from a cell-free deoxyribonucleic acid (cfDNA) sample, the plurality of quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying associated variables of said nucleic acid variants from said sequencing information; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) generating a statistical model for expected germline mutant allele counts at the genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on a statistical model for the expected germline mutant allele count, a quantitative value for an associated variable of the nucleic acid variant, and at least one of a plurality of quantitative measures for the nucleic acid variant; (f) classifying the nucleic acid variant as (i) somatic in origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) germline in origin when the p-value for the nucleic acid variant is at or above the predetermined threshold; A system comprising a controller that comprises or has access to a non-transitory computer-readable medium that contains computer-executable instructions for performing a method, comprising: (Item 46) 46. ​​The system of claim 45, further comprising a nucleic acid sequencing device operably connected to the controller, the nucleic acid sequencing device configured to provide sequencing information from nucleic acids of the cfDNA sample. (Item 47) 47. The system of claim 45 or 46, further comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare nucleic acids of the cfDNA sample to be sequenced by a nucleic acid sequencing device. (Item 48) 48. The system of any one of items 45 to 47, further comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify nucleic acids of the cfDNA sample. (Item 49) 49. The system of any one of items 45 to 48, further comprising a material transport component operably connected to the controller, the material transport component configured to transport one or more materials between the nucleic acid sequencing device and the sample preparation component. (Item 50) 50. The system of any one of items 45 to 49, wherein the predetermined threshold is generated using a beta-binomial model of expected germline mutant allele counts for nucleic acids of the cfDNA sample. (Item 51) 51. The system of any one of items 45-50, wherein the associated variable of the nucleic acid variant comprises at least one heterozygous single nucleotide polymorphism (het SNP). (Item 52) 52. The system of claim 51, wherein the associated variable of the nucleic acid variant comprises at least two het SNPs. (Item 53) The system according to any one of claims 45 to 52, wherein the associated variable of the nucleic acid variant comprises a genomic locus linked to the genomic locus containing the nucleic acid variant. 54. The system of any one of items 45 to 53, wherein the mean and / or variance of one or more mutant allele counts is determined for associated variables of the nucleic acid variants. (Item 55) 55. The system according to any one of items 45 to 54, wherein the p-value is used to classify the nucleic acid variant. (Item 56) 56. The system of any one of items 45 to 55, wherein at least one of the plurality of quantitative measurements comprises a number of nucleic acid molecules in the cfDNA sample that contain the nucleic acid variant. (Item 57) 57. The system of any one of items 45-56, wherein the associated variables include one or more of a heterozygous single nucleotide polymorphism (het SNP), a GC content measure, a probe-specific bias measure, a fragment length value, a sequencing statistic measure, a copy number breakpoint, and clinical data about the subject. (Item 58) 58. The system of any one of items 45 to 57, wherein a local germline folded mutant allele fraction (MAF), μ bin, is determined for the nucleic acid variant, where bin is a gene or another defined genomic region that contains the nucleic acid variant, and the folded MAF is min(MAF,1-MAF). (Item 59) The defined genomic region comprises about 10 of the nucleic acid variants. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 59. The system according to any one of items 45 to 58, wherein the region is within base pairs. (Item 60) 60. The system of any one of items 45 to 59, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) having a population allele frequency (AF) of greater than about 0.001. (Item 61) 61. The system of any one of items 45 to 60, wherein the associated variable of the nucleic acid variant comprises at least one non-oncogenic single nucleotide polymorphism (SNP). (Item 62) 62. The system of any one of items 45 to 61, wherein the associated variable of the nucleic acid variant comprises at least one single nucleotide polymorphism (SNP) comprising a mutant allele fraction (MAF) of less than about 0.9. (Item 63) The associated variable comprises at least one heterozygous SNP within the genomic region defined for the nucleic acid variant, and a beta binomial distribution parameter is estimated using: (x,y)~Beta binomial (μ bin ,ρ) During the ceremony, y=a vector of total molecular counts of said germline heterozygous SNPs, with one entry for each germline heterozygous SNP identified in (b); a vector of x=min(mutant allele count of said germline heterozygous SNP, y-mutant allele count of said germline heterozygous SNP), with one entry for each germline heterozygous SNP identified in (b); μ bin = an estimate of the mutant allele count of the heterozygous SNP within a bin, the bin being a genomic region defined for the nucleic acid variant, ρ = Estimate of the variance parameter, A system described in any one of items 45 to 62. (Item 64) The two-sided p-value for the nucleic acid variant is calculated using: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) During the ceremony, Pr bb = Beta binomial probability, x' = a random variable distributed with the beta binomial, A=the mutant allele count of the nucleic acid variant; B=total molecular count of the nucleic acid variant; Item 64. The system according to item 63. (Item 65) 65. The system of any one of items 45 to 64, wherein an upper bound and a lower bound on the p-value are calculated. (Item 66) 1. A method for identifying somatic or germline origin of nucleic acid variants from a sample of cell-free deoxyribonucleic acid (cfDNA) molecules, comprising: (a) determining a mutant allele count (A) and a total molecule count (B) of said nucleic acid variants from said sample of cfDNA molecules; (b) identifying at least one germline heterozygous single nucleotide polymorphism (SNP) within the genomic region defined for said nucleic acid variant; (c) determining the total molecular count (y) and mutant allele count of said at least one germline heterozygous SNP; (d) (i)μ bin determining estimates of ρ and ρ from a beta binomial distribution, (x,y)~Beta binomial (μ bin ,ρ) During the ceremony, y=a vector of total molecular counts of said germline heterozygous SNPs, with one entry for each germline heterozygous SNP identified in (b); a vector of x=min(mutant allele count of said germline heterozygous SNP, y-mutant allele count of said germline heterozygous SNP), with one entry for each germline heterozygous SNP identified in (b); μ bin = an estimate of the mutant allele count of a germline heterozygous SNP within a bin, the bin being a genomic region defined for the nucleic acid variant, ρ = Estimate of the variance parameter, Steps and (ii) calculating a two-sided p-value from the following equation: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) During the ceremony, Pr bb = Beta binomial probability, x' = a random variable distributed with the beta binomial distribution, A=the mutant allele count of the nucleic acid variant; B=total molecular count of the nucleic acid variant; Steps and calculating a probability value (p-value) for said nucleic acid variant by (e) classifying the nucleic acid variant as (i) of somatic origin when the p-value is below a predetermined threshold, or (ii) of germline origin when the p-value is at or above the predetermined threshold; A method comprising: (Item 67) Item 67. The method of item 66, wherein p comprises a median of at least one set of p values ​​from historical sample sets. (Item 68) μ bin 68. The method of claim 66 or 67, comprising determining a maximum likelihood estimate of (Item 69) μ bin 69. The method according to any one of items 66 to 68, comprising a step of determining an average value estimate of (Item 70) 70. The method of any one of items 66 to 69, comprising a step of determining a maximum likelihood estimate of p. (Item 71) 71. The method according to any one of items 66 to 70, comprising a step of determining a variance estimate of p. (Item 72) a communications interface for obtaining sequencing information generated from nucleic acids of the cell-free deoxyribonucleic acid (cfDNA) sample over a communications network; and a computer in communication with the communication interface, the computer comprising at least one computer processor and a non-transitory computer readable medium containing machine executable code; A system comprising: The machine executable code, when executed by at least one computer processor, (a) determining a plurality of quantitative measures for nucleic acid variants from the sequencing information, the plurality of quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying associated variables of said nucleic acid variants from said sequencing information; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) generating a statistical model for expected germline mutant allele counts at the genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on a statistical model for the expected germline mutant allele count, a quantitative value for an associated variable of the nucleic acid variant, and at least one of a plurality of quantitative measures for the nucleic acid variant; (f) classifying the nucleic acid variant as (i) somatic in origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) germline in origin when the p-value for the nucleic acid variant is at or above the predetermined threshold; A system implementing the method, comprising: (Item 73) 73. The system of claim 72, wherein the sequencing information is provided by a nucleic acid sequencing device. (Item 74) 74. The system of claim 73, wherein the nucleic acid sequencing device performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, sequencing by ligation, or sequencing by hybridization of the nucleic acid and generates the sequencing information. (Item 75) 74. The system of claim 73, wherein the nucleic acid sequencing device generates the sequencing information using a clonal single molecule array derived from a sequencing library. (Item 76) 74. The system of claim 73, wherein the nucleic acid sequencing device comprises a chip having an array of microwells for sequencing a sequencing library and generating the sequencing information. (Item 77) 77. The system of any one of claims 72 to 76, wherein the non-transitory computer-readable medium comprises a memory, a hard drive, or a memory or hard drive of a computer server. (Item 78) 77. The system of any one of claims 72 to 76, wherein the communications network comprises one or more computer servers capable of distributed computing. (Item 79) Item 79. The system of item 78, wherein the distributed computing is cloud computing. (Item 80) 80. The system of any one of items 72 to 79, wherein the computer is part of a computer server located remotely from the nucleic acid sequencing device. (Item 81) 81. The system of any one of items 72 to 80, further comprising an electronic display in communication with the computer via a network, the electronic display including a user interface for displaying results responsive to implementing at least a portion of (a)-(f). (Item 82) Item 82. The system of item 81, wherein the user interface is a graphical user interface (GUI) or a web-based user interface. (Item 83) Item 82. The system of item 81, wherein the electronic display is part of a personal computer. (Item 84) Item 82. The system of item 81, wherein the electronic display is part of an Internet-enabled computer. (Item 85) 85. The system of claim 84, wherein the Internet-enabled computer is located at a location remote from the computer. (Item 86) 86. The system of any one of items 72 to 85, wherein the non-transitory computer-readable medium comprises a memory, a hard drive, or a memory or hard drive of a computer server. (Item 87) The system of any one of items 72 to 86, wherein the communication network includes a telecommunications network, the Internet, an extranet, or an intranet. (Item 88) The method of claim 1 or claim 66, wherein the method further comprises generating a report in electronic and / or paper format providing an indication of the classification of the nucleic acid variants as either somatic or germline in origin. (Item 89) 1. A method of treating a disease in a subject, the method comprising administering one or more customized therapies to the subject, thereby treating the disease in the subject, the customized therapies comprising: (a) determining one or more quantitative measures for nucleic acid variants from a sample of cell-free deoxyribonucleic acid (cfDNA) molecules, the quantitative measures comprising total allele counts and minor allele counts for the nucleic acid variants; (b) identifying at least one associated variable of said nucleic acid variants from said sample of cfDNA molecules; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) generating a statistical model for expected germline mutant allele counts at the genomic locus of the nucleic acid variant. (e) generating a probability value (p-value) for the nucleic acid variant based on at least one of a statistical model for expected germline allele counts, a quantitative value for an associated variable of the nucleic acid variant, and the quantitative measure for the nucleic acid variant; (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above the threshold; (g) comparing the classified nucleic acid variants to one or more comparator results indexed with one or more therapies; (h) identifying one or more customized therapies for treating disease in the subject when a substantial match exists between the classified nucleic acid variants and the comparator results; The method is identified by (Item 90) 90. The method of claim 89, wherein the disease is cancer. [Brief description of the drawings]

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain embodiments and, together with the written description, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not by way of limitation. It should be understood that, unless otherwise indicated by the context, like reference numbers identify like components throughout the drawings. It should also be understood that some or all of the drawings may be schematic for illustrative purposes and do not necessarily depict the actual relative size or location of the elements shown.

[0030] [Figure 1] FIG. 1 is a flowchart representation of a method for distinguishing somatic and germline variants in a sample of nucleic acid molecules, according to an embodiment of the present disclosure.

[0031] [Diagram 2] FIG. 2 is a flowchart representation of a method for distinguishing somatic and germline variants in a sample of nucleic acid molecules using a beta-binomial distribution, according to an embodiment of the present disclosure.

[0032] [Diagram 3] FIG. 3 is a graphical representation of the decision boundary for discriminating between germline / somatic variants using the beta binomial distribution.

[0033] [Figure 4] FIG. 4 is a schematic diagram of an exemplary system suitable for use with some embodiments of the present disclosure.

[0034] [Figure 5A] FIG. 5A is a graphical representation of mutant allele fraction (MAF) versus genomic position for the T790M variant and six common germline heterozygous SNPs in the EGFR gene.

[0035] [Figure 5B]FIG. 5B is a graphical representation of min(MAF,1-MAF) versus genomic position for the T790M variant and six common germline heterozygous SNPs in the EGFR gene. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036] definition In order that this disclosure may be more readily understood, certain terms are first defined below. Additional definitions for the following terms and other terms may be set forth throughout the specification. In the event that a definition of a term set forth below conflicts with a definition within an application or patent incorporated by reference, the definition set forth in this application shall be used to understand the meaning of the term.

[0037] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to a "method" includes one or more methods, and / or steps, of the type described herein and / or that will be apparent to those skilled in the art upon perusal of this disclosure, and so forth.

[0038] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Moreover, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variations thereof will be used in accordance with the definitions set forth below.

[0039] About: As used herein, "about" or "approximately" as applied to one or more values ​​or elements of interest refers to a value or element that is similar to the stated reference value or element. In an embodiment, the term "about" or "approximately" refers to a range of values ​​or elements within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater or less) of the stated reference value or element, unless otherwise stated or otherwise clear from the context (except where such number would exceed 100% of the possible values ​​or elements).

[0040] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides long) that is typically at least partially double-stranded and used to ligate to one or both ends of a given sample nucleic acid molecule. The adapter can include a nucleic acid primer binding site to allow amplification of the nucleic acid molecule flanked by the adapters, and / or a sequencing primer binding site, including a primer binding site for sequencing applications, such as various next generation sequencing (NGS) applications. The adapter can also include a binding site for a capture probe, such as an oligonucleotide, that is attached to a flow cell support or equivalent. The adapter can also include a nucleic acid tag, as described herein. The nucleic acid tag is typically positioned relative to the amplification primer and sequencing primer binding sites such that the nucleic acid tag is included in the amplicon and sequencing reads of a given nucleic acid molecule. The same or different adapters can be ligated to separate ends of a nucleic acid molecule. In some embodiments, the same adapter is ligated to separate ends of a nucleic acid molecule, except that the nucleic acid tag is different. In some embodiments, the adaptor is a Y-shaped adaptor, one end of which is blunt or terminated with one or more complementary nucleotides for joining to a nucleic acid molecule, as described herein. In yet other exemplary embodiments, the adaptor is a bell-shaped adaptor, including a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other examples of adaptors include T-terminated and C-terminated adaptors.

[0041] Amplification: As used herein, "amplifying" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide or portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.

[0042] Associated variables: As used herein, the term "associated variables" refers to variables that are associated with nucleic acid variants and are used in estimating expected germline mutant allele counts. Such variables can include, but are not limited to, germline heterozygous SNPs, GC content measurements, probe-specific bias measurements, fragment length values, sequencing statistics measurements, copy number breakpoints, clinical data from subjects, or any combination thereof.

[0043] Cancer type: As used herein, "cancer type" refers to a type or subtype of cancer as defined, for example, by histopathology. A cancer type can refer to the occurrence within a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, mouth cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, solid cancer, heterogeneous cancer, homogenous cancer), undifferentiated cancer, or a combination of both. Cancers can be defined by any conventional criteria, such as based on cancers of known primary origin and equivalent, and / or of the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or exhibiting cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether they are primary or secondary in origin.

[0044] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise bound to a cell, or in some embodiments, nucleic acid that remains in a sample after removal of intact cells. Cell-free nucleic acid can include, for example, all non-encapsulated nucleic acid derived from bodily fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acid can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acid can be released into bodily fluids through bodily fluid secretions or cell death processes, such as cell necrosis, apoptosis, or the like. Cell-free nucleic acid can be found in efferosomes or exosomes when efferosomes or exosomes incorporate cell-free nucleic acid released into other cellular fluids. Some cell-free nucleic acid is released into fluids from cancer cells, e.g., circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be non-encapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acid is fetal DNA that circulates freely in maternal bloodstream, also called cell-free fetal DNA (cffDNA). Cell-free nucleic acid can have one or more epigenetic modifications, e.g., cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0045] Cellular Nucleic Acid: As used herein, "cellular nucleic acid" refers to nucleic acids that are located within one or more cells from which they originated, at least at the time the sample is taken or collected from a subject, even if those nucleic acids are subsequently removed (e.g., via cell lysis) as part of a given analytical process.

[0046] Common germline heterozygous SNP: As used herein, the term "common germline heterozygous SNP" refers to a germline heterozygous single nucleotide polymorphism (SNP) obtained from an external population database (e.g., ExAC) and / or any historical sample set, such that the heterozygous SNP has at least a particular population allele frequency (AF), where the particular population AF can be any value between 0 and 1.

[0047] Comparator result: As used herein, "comparator result" refers to a result or set of results that a given test sample or test result can be compared to to identify one or more likely properties of the test sample or result and / or one or more possible prognostic outcomes and / or one or more customized therapies for the subject from which the test sample was taken or otherwise derived.Comparator results are typically obtained from a set of reference samples (e.g., from a subject having the same disease or cancer type as the test subject).

[0048] Copy number breakpoint: As used herein, the term "copy number breakpoint" refers to a genomic locus at which the copy numbers (CN) of two neighboring genomic regions (within the same chromosome) on either side of that genomic locus differ.

[0049] Copy number variant: As used herein, "copy number variant," "CNV," or "copy number polymorphism" refers to the phenomenon whereby segments of the genome are repeated and the number of repeats within the genome varies between individuals within a population under consideration and between two conditions or states of an individual (e.g., CNVs can vary in an individual before and after undergoing a therapy).

[0050] Coverage: As used herein, the terms "coverage," "total molecule count," or "total allele count" are used interchangeably. They refer to the total number of DNA molecules at a particular genomic location in a given sample.

[0051] Customized Therapy: As used herein, "customized therapy" refers to a therapy that is associated with a desired therapeutic outcome for a subject or population of subjects having a given classified nucleic acid variant.

[0052] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to a natural or modified nucleotide that has a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a chain of nucleotides that includes four types of nucleotides: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to a natural or modified nucleotide that has a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a chain of nucleotides that includes four types of nucleotides: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to a natural or modified nucleotide. A pair of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand, which consists of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence," or "nucleic acid sequencing read" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment) of a nucleic acid, such as DNA or RNA.It should be understood that the present teachings contemplate sequence information obtained using all available variations of techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide discrimination systems, pyrosequencing, ion or pH-based detection systems, and electronic signature-based systems.

[0053] Expected germline mutant allele count: As used herein, the term "expected germline mutant allele count" refers to the expected mutant allele count of a germline SNP at a genomic locus of a nucleic acid variant. For example, the expected germline mutant allele count can be estimated by a statistical distribution. The statistical distribution can be, but is not limited to, a beta-binomial distribution. The distribution is used to determine the expected mutant allele count within a germline heterozygous SNP at that locus. For example, when a beta-binomial distribution is used to determine the expected germline mutant allele count at a particular genomic locus, the distribution of the expected mutant allele count is parameterized by the mean estimate (μ), variance estimate (ρ), and coverage at that genomic locus.

[0054] Germline Mutant: As used herein, the terms "germline mutant" or "germline variant" are used interchangeably and refer to a heritable mutation (i.e., not occurring after conception). A germline mutant may be a unique mutation that can be inherited by offspring and may be present in all somatic and germline cells in the offspring.

[0055] Historical Sample Set: As used herein, the term "historical sample set" refers to a set of samples obtained from normal subjects (not having disease / cancer), subjects with any disease or cancer, subjects with a particular cancer type, and / or subjects undergoing or having undergone a particular therapy.

[0056] Indel: As used herein, "indel" refers to a mutation involving the insertion or deletion of nucleotides in the genome of a subject.

[0057] Mutant allele count: As used herein, the term "mutant allele count" refers to the number of DNA molecules that carry a mutant allele at a particular genomic locus.

[0058] Minor allele count: As used herein, "minor allele count" refers to the number of minor alleles (e.g., not the most common alleles) occurring in a given population of nucleic acids, such as a sample obtained from a subject. Genetic variants with low minor allele counts are typically present in a sample in relatively low numbers.

[0059] Mutant allele fraction: As used herein, "mutant allele fraction", "mutant dose", or "MAF" refers to the fraction of nucleic acid molecules that carry an allelic alteration or mutation at a given genomic location / locus in a given sample. MAF is generally expressed as a fraction or percentage. For example, the MAF of a somatic variant may be less than 0.15.

[0060] Mutant: As used herein, "mutant" refers to a variation from a known reference sequence, including, for example, mutations such as single nucleotide variants (SNVs) and insertions or deletions (indels). Mutants can be germline or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genome sequence of the species of interest from which the test sample is provided, typically the human genome.

[0061] Mutation caller: As used herein, "mutation caller" means an algorithm (typically embodied in software or otherwise computer-implemented) used to identify mutations in test sample data (e.g., sequence information obtained from a subject).

[0062] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to the abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is referred to as a cancer or cancerous tumor.

[0063] Next generation sequencing: As used herein, "next generation sequencing" or "NGS" refers to sequencing technologies that have increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequence reads at once. Some examples of next generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0064] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, less than about 50 nucleotides, or less than about 10 nucleotides in length) that is used to distinguish between nucleic acids from different samples (e.g., representing a sample index) or different nucleic acid molecules of different types or that have been subjected to different processes in the same sample (e.g., representing a molecular barcode). Such nucleic acid tags may be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can optionally have the same length or variable lengths. Nucleic acid tags can also have one or more blunt ends, include double-stranded molecules, include 5' or 3' single-stranded regions (e.g., overhangs), and / or include one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to allow for pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indexes, where the nucleic acids are subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers or indexes. Such nucleic acid tags, identifiers, or indexes may include one or more barcodes. Additionally or alternatively, nucleic acid tags can be used as molecular identifiers or indexes (e.g., to distinguish between amplicons of different molecules or different parent molecules in the same sample or subsample). This includes, for example, uniquely tagging each different nucleic acid molecule in a given sample, or non-uniquely tagging such molecules.For non-unique tagging applications, a limited number of tags (e.g., barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their endogenous sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length of the sequence, in combination with at least one barcode) as they map to a selected reference genome. Typically, a sufficient number of different nucleic acid tags are used such that there is a low probability (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance) that any two molecules will have the same endogenous sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length) and also have the same nucleic acid tag (e.g., barcode). Alternatively, the nucleic acid tag may only include endogenous sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length). Some nucleic acid tags include multiple molecular identifiers to label samples, forms of nucleic acid molecules in the samples, and nucleic acid molecules within forms that have the same intrinsic sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length). Such nucleic acid tags may be referenced using the exemplary format "A1i," where the capital letter indicates the sample type, the Arabic numerals indicate the form of the molecule in the sample, and the lowercase Roman numerals indicate the molecule within the form.

[0065] Polynucleotide: As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to hundreds of monomeric units. Whenever a polynucleotide is represented by a sequence of letters, such as "ATGCCTG", it is understood that the nucleotides are in 5'→3' order from left to right, and that in the case of DNA, "A" denotes deoxyadenosine, "C" denotes deoxycytidine, "G" denotes deoxyguanosine, and "T" denotes deoxythymidine, unless otherwise noted. The letters A, C, G, and T may be used to refer to the bases themselves, nucleosides, including bases, or nucleotides, as is standard in the art.

[0066] Reference sequence: As used herein, "reference sequence" refers to a known sequence that is used for the purpose of comparison with an experimentally determined sequence. For example, the known sequence can be an entire genome, a chromosome, or any section thereof. The reference typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1,000, or more than 1,000 nucleotides. The reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise non-contiguous sections that align with different regions of a genome or chromosome. Examples of reference sequences include human genomes, such as, for example, hG19 and hG38.

[0067] Sample: As used herein, "sample" means anything capable of being analyzed by the methods and / or systems disclosed herein.

[0068] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., identity and order of monomeric units) of a biological molecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single base sequencing, and nucleotide sequencing. The sequencing methods include base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification with low denaturation temperature PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, short-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, synthesis sequencing, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof.In some embodiments, sequencing can be performed by a genetic analyzer, such as the genetic analyzers available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.

[0069] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomeric units (eg, nucleotides) within that polymer.

[0070] Single Nucleotide Polymorphism: As used herein, the terms "single nucleotide polymorphism" or "SNP" are used interchangeably. They refer to variations in a single base that occur at specific locations in the genome, with each variation present to some appreciable degree in the population (e.g., greater than about 1%).

[0071] Single Base Variant: As used herein, "single base variant" or "SNV" refers to a mutation or variation in a single base that occurs at a specific location in the genome.

[0072] Somatic Mutant: As used herein, the terms "somatic mutant" or "somatic variant" are used interchangeably. They refer to mutations in the genome that arise after conception. Somatic mutations can occur in any cell of the body, except germ cells, and therefore are not inherited by offspring.

[0073] Subject: As used herein, "subject" refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., avian) species, or other organism, such as a plant. More specifically, a subject can be a mammal, such as a vertebrate, e.g., a mouse, a primate, an ape, or a human. Animals include livestock (e.g., beef cattle, dairy cattle, poultry, horses, pigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of therapy or suspected of needing therapy. The terms "individual" or "patient" are intended to be synonymous with "subject."

[0074] For example, the subject can be an individual who has been diagnosed with cancer, who is going to receive cancer therapy, and / or who has received at least one cancer therapy. The subject can be in remission of cancer. As another example, the subject can be an individual who has been diagnosed with an autoimmune disease. As another example, the subject can be a female individual who is pregnant or planning to become pregnant and may be diagnosed with or suspected of having a disease, such as cancer, an autoimmune disease.

[0075] Substantial matching: As used herein, "substantial matching" means that at least a first value or element is at least approximately equal to at least a second value or element. In an embodiment, for example, a customized therapy is identified when there is at least a substantial or approximate match between the classified nucleic acid variant and the comparator result.

[0076] Threshold: As used herein, "threshold" refers to a predefined value used to characterize experimentally determined values ​​of the same parameter for different samples depending on its relationship to the threshold. For example, a threshold for a p-value may refer to any predefined value between 0 and 1 and is used to identify the origin of a nucleic acid variant.

[0077] Variant: As used herein, "variant" can refer to an allele. Variants are usually present at a frequency of 50% (0.5) or 100% (1), depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. However, somatic variants are acquired variants and usually have a frequency of less than about 0.5. Dominant and recessive alleles of a locus refer to nucleic acids in which the locus is occupied by nucleotides of a reference sequence and variant nucleotides that differ from the reference sequence, respectively. Measurements at a locus can take the form of allele fraction (AF), which is a measurement of the frequency at which an allele is observed in a sample. Detailed Description I. Overview

[0078] The present disclosure provides a method and system for using statistical models such as beta-binomial models to classify or identify nucleic acid variants in a sample of nucleic acid molecules as somatic or germline origin. In some embodiments, the methods and systems of the present disclosure are suitable for analyzing cell-free nucleic acids such as cell-free DNA (cfDNA). Many solutions available for distinguishing somatic and germline variants using sequencing data from tumor tissue may rely on the availability of matched paired tumors, and normal tissues may not therefore be applicable to data obtained from cell-free nucleic acids. Solutions for analyzing cfDNA samples may include thresholding on mutant allele fraction (MAF) or applying Poisson statistical models to determine germline or somatic status. However, such approaches may not accurately model the variance found in cfDNA molecule counts, and therefore somatic / germline distinction based on these approaches may not be optimally accurate. The methods and systems disclosed herein can accurately model the variance found in nucleic acid molecule counts (such as in cfDNA) and can distinguish somatic and germline variants with high accuracy. The methods and systems disclosed herein can statistically model local germline mutant allele count behavior (e.g., germline mutant allele count behavior within a genomic region relative to a nucleic acid variant) using parameters such as common germline single nucleotide polymorphisms (SNPs) and distinguish somatic variants based on MAF deviation from the observed germline MAF.

[0079] In one aspect, the disclosure provides a method of identifying somatic or germline origin of a nucleic acid variant from a sample of cell-free deoxyribonucleic acid (cfDNA) molecules, comprising: (a) determining a plurality of quantitative measurements for the nucleic acid variant from the cfDNA sample, the plurality of quantitative measurements comprising a total allele count and a minor allele count for the nucleic acid variant; (b) identifying associated variables of the nucleic acid variant from the cfDNA sample; (c) determining a quantitative value for the associated variable of the nucleic acid variant; and (d) determining an expected germline mutation at a genomic locus at which the nucleic acid variant lies. (e) generating a probability value (p-value) for the nucleic acid variant based at least in part on the statistical model for expected germline mutation allele counts, the quantitative value for the associated variable of the nucleic acid variant, and at least one of the plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0080] 1 illustrates an exemplary embodiment of a method 100 for identifying somatic and germline variants in a sample of nucleic acid molecules. Once nucleic acid variants are identified from the nucleic acid molecules in the sample, quantitative values ​​and associated variables related to the nucleic acid variants can be established to provide input values ​​for implementing statistical models. Nucleic acid variants may be identified or detected by any known method, including, but not limited to, the methods described in U.S. Patent Nos. 9,598,731, 9,834,822, 9,840,743, and 9,902,992 (each of which is incorporated herein by reference in its entirety).

[0081] In operation 102, quantitative values ​​for the nucleic acid variants may be measured and determined. These values ​​may include, but are not limited to, mutant allele counts and / or total molecular counts of the nucleic acid variants.

[0082] Another input value required for the model may be a quantitative value for an associated variable. In operation 104, at least one associated variable may be identified. The associated variable may be used in estimating the expected germline mutant allele count at the genomic locus of the nucleic acid variant. Such associated variables may include, but are not limited to, germline heterozygous SNPs, GC content measurements, probe-specific bias measurements, fragment length values, sequencing statistics measurements, copy number breakpoints, clinical data from the subject, or any combination thereof.

[0083] In some embodiments, the associated variables may be within a genomic region (also referred to as a "bin") defined for the nucleic acid variant. In some embodiments, the bin may be a gene that contains the nucleic acid variant. In some embodiments, the bin can be a genomic region defined for the nucleic acid variant. In some embodiments, a bin (defined genomic region) is within about 10 of the nucleic acid variant. 1 , 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , or 10 10In some embodiments, the bin is within "N" bases of the nucleic acid variant, where N is about 1, about 5, about 10, about 25, about 50, about 100, about 250, about 500, about 1,000, about 5,000, about 10,000, about 50,000, about 100,000, about 500,000, about 1,000,000, or more than about 1,000,000 bases. In some embodiments, N can be up to 3,000,000 bases. For example, a bin is within 10 of a nucleic acid variant. 5 10 bases. In some embodiments, the associated variables of the nucleic acid variants include genomic loci linked to the genomic locus containing the nucleic acid variant. In some embodiments, the associated variables include at least 1, at least 2, at least 5, at least 10, or more than 10 heterozygous SNPs. In some embodiments, the associated variables of the nucleic acid variants include at least one SNP that includes a population allele frequency (AF) of at least 0.00001, at least 0.0001, at least 0.001, at least 0.002, at least 0.005, at least 0.01, at least 0.02, at least 0.05, at least 0.1, at least 0.2, at least 0.5, at least 0.75, or at least 0.99. In some embodiments, the associated variables of the nucleic acid variants include at least one SNP that includes a population allele frequency (AF) value between 0 and 1. In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) that comprises a mutant allele fraction (MAF) of less than 0.9. In some embodiments, the associated variables of the nucleic acid variants include at least one single nucleotide polymorphism (SNP) that comprises a mutant allele fraction (MAF) of between 0 and about 1. In some embodiments, the associated variables of the nucleic acid variants include at least one heterozygous SNP, and the heterozygous SNP can be a common germline heterozygous SNP.

[0084] In some embodiments, associated variables are within copy number breakpoints.Instead of having fixed width bins or bins defined by gene annotation, associated variables can be identified within bins bounded by copy number breakpoints, so that the bins of each nucleic acid variant are as wide as possible without overlapping any copy number breakpoints.In some embodiments, associated variables include heterozygous SNPs within copy number breakpoints.

[0085] In operation 106, a quantitative value for the associated variable of the nucleic acid variant may be determined. The quantitative value of the associated variable may be used as an input in applying a statistical model to estimate an expected germline mutant allele count at the genomic locus of the nucleic acid variant. In some embodiments, the quantitative value for the associated variable comprises a mutant allele count and / or a total molecular count of the associated variable. In some embodiments, the method further comprises determining a MAF. In some embodiments, the MAF is adjusted to a reduced scale, referred to herein as the "folded MAF" of the associated variable, where the folded MAF = min(MAF,1-MAF). In some embodiments, the method comprises determining a folded mutant allele count of the associated variable, where the folded mutant allele count = min(mutant allele count, total molecular count - mutant allele count). In some embodiments, the quantitative value comprises one or more allele counts identified in the associated variable of the nucleic acid variant. In some embodiments, the method comprises determining the mean and / or variance of one or more allele counts identified in the nucleic acid variant associated variables. In some embodiments, the method comprises determining the mean quantitative value for the nucleic acid variant associated variables. In some embodiments, the method comprises determining the mean and / or variance for the nucleic acid variant associated variables. In some embodiments, the nucleic acid variant associated variables comprise at least one non-oncogenic SNP.

[0086] In operation 108, the determined quantitative values ​​may be processed using a statistical model, such as a beta-binomial model. The distribution generated from the statistical model may be used to determine mutant allele counts that may be expected within germline heterozygous SNPs at that locus. For example, if a beta-binomial distribution is used to determine expected germline mutant allele counts at a particular genomic locus, the distribution of expected germline mutant allele counts may be parameterized by a set of statistical parameters corresponding to the beta-binomial distribution at that genomic locus, e.g., mean estimate (μ), variance estimate (ρ), and coverage. In some embodiments, the method includes determining a μ for a nucleic acid variant. bin determining μ bin is an estimate of the mutant allele count of heterozygous SNPs in a bin.

[0087] In some embodiments, the associated variable comprises at least one heterozygous single nucleotide polymorphism (SNP) in the genomic region defined for the nucleic acid variant, and the method comprises estimating a beta binomial distribution parameter using: (x,y)~Beta binomial (μ bin ,ρ) where y=a vector of total molecular counts of germline heterozygous SNPs, with one entry for each germline heterozygous SNP considered; x=a vector of min(mutant allele counts of germline heterozygous SNPs, y-mutant allele counts of germline heterozygous SNPs), with one entry for each germline heterozygous SNP considered; μ bin = estimate of the mutant allele count of heterozygous SNPs within a bin, where a bin is a defined genomic region for nucleic acid variants, and ρ = estimate of the dispersion parameter.

[0088] In an embodiment, x and y can be represented as vectors, with one entry for each germline heterozygous SNP. This is the case when two or more germline heterozygous SNPs are considered in the model. For example, when two germline heterozygous SNPs are considered, y can be expressed as y 1 (SNP 1 (total molecule counts for y 2 (het SNP 2 (total molecule counts with respect to x). Similarly, x is expressed as a vector of 1 (het SNP 1 ) and x 2 (het SNP 2 In some embodiments, only one germline heterozygous SNP may be considered. In these cases, the values ​​for x and y may be represented as a vector with only one entry, or alternatively, as y=total molecular count of heterozygous SNPs and x=min(mutant allele count of heterozygous SNPs, y-mutant allele count of heterozygous SNPs).

[0089] In some embodiments, p comprises a median of at least one set of p values ​​from a past sample set. In some embodiments, the method comprises replacing the median p parameter with a function of the GC content of the nucleic acid variants. In some embodiments, the method comprises: bin In some embodiments, the method includes determining a maximum likelihood estimate of μ bin In some embodiments, the method includes determining a mean estimate of p. In some embodiments, the method includes determining a maximum likelihood estimate of p. In some embodiments, the method includes determining a variance estimate of p.

[0090] In some embodiments, rather than being modeled as a fixed number, the dispersion parameter (ρ) can be modeled as a function of the GC content of the local genomic context (e.g., the genomic context of a bin). The function can be estimated from a past sample set, and the median value of ρ in the above equation can be replaced by the value of this function at the GC content level of the variant.

[0091] In operation 110, a probability value (p-value) for the nucleic acid variant may be determined based, at least in part, on at least one of a statistical model for expected germline mutant allele counts, a quantitative value for an associated variable of the nucleic acid variant, and a quantitative measure for the nucleic acid variant. In some embodiments, the method includes calculating a two-sided p-value for the nucleic acid variant using: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) In the formula, Pr bb = beta-binomial probability, x' = a random variable distributed with beta-binomial, A = mutant allele count of the nucleic acid variant, and B = total molecular count of the nucleic acid variant.

[0092] In operation 112, a nucleic acid variant may be classified as (i) of somatic origin when the p-value of the nucleic acid variant is below a threshold value, or (ii) of germline origin when the p-value of the nucleic acid variant is at or above a threshold value. The threshold value can be any value that can distinguish between germline variants and somatic variants. The threshold value can be determined from experimental data. For example, the threshold value can be any value between 0 and 1. In some embodiments, the threshold value is at least 10 -50 , at least 10 -40 , at least 10 -30 , at least 10 -20 , at least 10 -10, at least 10 -5 , at least 0.01, at least 0.01, at least 0.1, at least 0.2, at least 0.5, at least 0.75, or at least 0.99. In some embodiments, the method includes generating the threshold value using a beta-binomial model of expected germline mutant allele counts for nucleic acids in the sample.

[0093] In some embodiments, the methods comprise typing the somatic or germline origin of a plurality of nucleic acid variants from a plurality of genomic loci in a nucleic acid sample.

[0094] The methods and systems disclosed herein generally include obtaining sequence information from a nucleic acid in a sample taken from a subject. In some embodiments, the method further includes receiving sequencing information generated from a nucleic acid sample, the sequencing information includes sequencing reads from the nucleic acid including a nucleic acid variant and an associated variable of the nucleic acid variant, the associated variable including at least one heterozygous single nucleotide polymorphism (SNP) in a genomic region defined for the nucleic acid variant. In some embodiments, the method further includes sequencing the nucleic acid from the sample to generate sequencing information, and a quantitative measurement value is determined from the sequencing information. In some embodiments, the method includes determining a quantitative measurement value for the nucleic acid variant, identifying an associated variable of the nucleic acid variant, and determining a quantitative value from the sequencing information generated from the sample.

[0095] In another aspect, the disclosure provides a method of identifying somatic or germline origin of a nucleic acid variant from a sample of cell-free nucleic acid (e.g., cfDNA), comprising: (a) determining a mutant allele count (A) and a total molecular count (B) of the nucleic acid variant from a cfDNA sample; (b) identifying at least one germline heterozygous single nucleotide polymorphism (SNP) within a defined genomic region for the nucleic acid variant; (c) determining a total molecular count (y) and mutant allele count of the germline heterozygous SNP; and (d) determining a total molecular count (y) and mutant allele count (y) of the germline heterozygous SNP from a sample of cell-free nucleic acid (e.g., cfDNA) comprising: bin determining estimates of ρ and ρ from a beta binomial distribution, (x,y)~Beta binomial (μ bin ,ρ) where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each germline heterozygous SNP; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each germline heterozygous SNP; μ bin = an estimate of the mutant allele count of a germline heterozygous SNP within a bin, where the bin is a defined genomic region for the nucleic acid variant; and ρ = an estimate of the dispersion parameter; and (ii) calculating a two-sided p-value using: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) In the formula, Pr bbwhere x'=beta binomial probability, x'=a random variable distributed with a beta binomial distribution, B=the total molecular count of the nucleic acid variant, and A=the mutant allele count of the nucleic acid variant; and (e) classifying the nucleic acid variant as (i) of somatic origin when the p-value is below a predetermined threshold, or (ii) of germline origin when the p-value is at or above a predetermined threshold.

[0096] In some embodiments, p comprises a median value of at least one set of p values ​​from a past sample set. bin In some embodiments, the method includes determining a maximum likelihood estimate of μ bin In some embodiments, the method includes determining a mean estimate of p. In some embodiments, the method includes determining a maximum likelihood estimate of p. In some embodiments, the method includes determining a variance estimate of p.

[0097] FIG. 2 illustrates an embodiment of a method for distinguishing somatic and germline variants in a sample of cfDNA using a beta-binomial model. In operation 202, a mutant allele count (A) and a total molecular count (B) of a nucleic acid variant are determined from a cfDNA sample. In operation 204, at least one germline heterozygous single nucleotide polymorphism (SNP) within a genomic region defined for a nucleic acid variant may be identified. In operation 206, a total molecular count (y) and a mutant allele count of the germline heterozygous SNP may be determined. In operation 208, μ is calculated from the beta-binomial distribution. bin and ρ can be estimated using: (x,y)~Beta binomial (μ bin ,ρ) where y=a vector of total molecular counts of at least one germline heterozygous SNP, with one entry for each germline heterozygous SNP considered; x=a vector of min(mutant allele counts of at least one germline heterozygous SNP, y-mutant allele counts of at least one germline heterozygous SNP), with one entry for each germline heterozygous SNP considered; μ bin = estimate of the mutant allele count of the germline heterozygous SNP within a bin, where the bin is a genomic region defined for the nucleic acid variant, and ρ = estimate of the dispersion parameter. In operation 210, a two-sided p-value may be calculated using: p-value=2*min(Pr bb (x'>A|μ bin ,ρ,B),Pr bb (x' <A|μ bin ,ρ,B) In the formula, Pr bb = beta binomial probability, x' = a random variable distributed with a beta binomial distribution, B = the total molecular count of the nucleic acid variant, and A = the mutant allele count of the nucleic acid variant.

[0098] Current solutions for identifying somatic or germline origin of variants in cfDNA may include thresholding on mutant allele fraction (MAF) or applying a Poisson statistical model to determine germline or somatic status. However, such approaches face challenges in accurately modeling the variance found in cfDNA sequencing molecular counts, and thus may result in inaccurate germline / somatic distinction. Furthermore, these methods cannot adjust their somatic threshold in response to evidence from neighboring variables or other covariates to the nucleic acid variant. The beta-binomial model may overcome these problems by modeling the distribution of expected germline mutant allele counts using mean and variance estimates and coverage at the genomic locus of the nucleic acid variant. The mean and variance estimates of expected germline heterozygous SNPs may be used in calculating the p-value of the nucleic acid variant, which may in turn be used to classify the variant as somatic or germline in origin.

[0099] In operation 212, the nucleic acid variant may be classified as (i) of somatic origin when the p-value is below a predetermined threshold, or (ii) of germline origin when the p-value is at or above a predetermined threshold.

[0100] FIG. 3 shows an example of a decision boundary for germline / somatic variant discrimination using a beta-binomial distribution. The beta-binomial decision boundary for a nucleic acid variant MAF may be a function of the MAF of the germline heterozygous SNP, the total count of molecules observed at the variant position, and an adjustable p-value threshold. As an example, a gene with allelic imbalance due to copy number variation (CNV) or loss of heterozygosity (LOH) may have germline MAF in both the 10-30% and 70-90% ranges. Referring back to FIG. 3, 302 (outer solid line), 304 (center solid line), and 306 (inner solid line) represent decision boundaries for germline / somatic discrimination using a beta-binomial model, with a threshold for p-values ​​of 10 -16 where the variant total molecular count (B) is 700, 1,500, and 3,000, respectively. Additionally, 308 (outer dashed line), 310 (center dashed line), and 312 (inner dashed line) represent decision boundaries for germline / somatic discrimination using a beta-binomial model, where the threshold for p-value is 0.01, and the variant total molecular count (B) is 700, 1,500, and 3,000, respectively.

[0101] In some embodiments, sequence information is obtained from targeted sections of nucleic acids. Essentially any number of genomic regions may be targeted, at will. Targeted sections may be at least 10, at least 50, at least 100, at least 500, at least 1,000, at least 2,000, at least 5,000, at least 10,000, at least 20,000, at least 50,000, or at least 100,000 (e.g., 25, 50, 75, 100, 200, 300, 400, 500, 60, 75, 80, 90, 100, 120, 140, 160, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 700, 710, 720, 730, 740, 750, 750, 800, 850, 900, 950, 960, 970, 9 The invention can include 0, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 25,000, 30,000, 35,000, 40,000, 45,000, 50,000, or 100,000) different and / or overlapping genomic regions.

[0102] In some embodiments, the identified germline and / or somatic variants are used as input to generate a report in electronic and / or paper format that provides an indication of the classification of these genetic variants in the polynucleotides as either somatic or germline in origin.

[0103] The various steps of the method may be performed at the same or different times, in the same or different geographic locations, eg countries, by the same or different people or entities. II. General Features of the Method A. Sample

[0104] The sample can be any biological sample isolated from a subject. Samples can include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (white cells or leucocytes), endothelial cells, tissue biopsies (e.g., biopsies from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluids (e.g., fluids from cell gaps), gingival exudate, gingival crevicular fluid, bone marrow, pleural exudate, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples can be bodily fluids such as blood and its fractions, and urine. Such samples include nucleic acids shed from tumors. Nucleic acids can include DNA and RNA, and can be in double-stranded and single-stranded forms. The sample may be in the form originally isolated from the subject, or may be further processed to remove or add components such as cells, enrich one component for another, or convert one form of nucleic acid to another, such as RNA to DNA or single-stranded to double-stranded nucleic acid. Thus, for example, the body fluid for analysis may be plasma or serum, which contains cell-free nucleic acid, e.g., cell-free DNA (cfDNA).

[0105] In some embodiments, the sample volume of bodily fluid taken from the subject depends on the desired read depth for the region to be sequenced. Example volumes are about 0.4 to 40 milliliters (mL), about 5 to 20 mL, about 10 to 20 mL. For example, the volume can be about 0.5 mL, about 1 mL, about 5 mL, about 10 mL, about 20 mL, about 30 mL, about 40 mL, or more milliliters. The volume of sampled plasma is typically about 5 mL to about 20 mL.

[0106] Samples can contain various amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample will correspond to multiple genome equivalents. For example, a sample of about 30 nanograms (ng) of DNA will contain approximately 10,000 (10 4 ) haploid human genome equivalents, approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, or about 600 billion individual molecules for cfDNA.

[0107] In some embodiments, the sample contains nucleic acid from different sources, e.g., from cells and from acellular sources (e.g., blood samples, etc.). Typically, the sample contains nucleic acid carrying mutants. For example, the sample optionally contains DNA carrying germline mutants and / or somatic mutants. Typically, the sample contains DNA carrying cancer-associated mutants (e.g., cancer-associated somatic mutants).

[0108] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, about 10 ng to about 1,000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In some embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In some embodiments, the method includes obtaining from about 1 fg to about 200 ng of cell-free nucleic acid molecules from a sample.

[0109] Cell-free nucleic acids typically have a size distribution from about 100 nucleotides to about 500 nucleotides in length, with molecules from about 110 nucleotides to about 230 nucleotides in length representing about 90% of the molecules in the sample, with about 168 nucleotides in length (in samples from human subjects) being the mode, and a second minor peak in the range of about 240 nucleotides to about 440 nucleotides in length. In some embodiments, the cell-free nucleic acids are about 160 nucleotides to about 180 nucleotides in length, or about 320 nucleotides to about 360 nucleotides in length, or about 440 nucleotides to about 480 nucleotides in length.

[0110] In some embodiments, cell-free nucleic acids can be isolated from bodily fluids through a partitioning step, in which cell-free nucleic acids as found in solution are separated from intact cells and other non-soluble components of the bodily fluid. In some embodiments, partitioning includes techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid can be lysed and the cell-free and cellular nucleic acids can be processed together. Generally, after addition of buffer and washing steps, the cell-free nucleic acids can be precipitated, for example with alcohol. In some embodiments, further cleaning steps, such as silica-based columns to remove contaminants or salts, are used. Non-specific bulk carrier nucleic acids are added, for example, as needed throughout the reaction to optimize exemplary aspects of the procedure, such as yield. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. Optionally, the single-stranded DNA and / or single-stranded RNA are converted to double-stranded form so that they can be included in subsequent processing and analysis steps. B. Tagging

[0111] In some embodiments, the nucleic acid molecules may be tagged with a sample index and / or a molecular barcode (generally referred to as a "tag"). The tag may be incorporated into or otherwise joined to an adaptor by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR), among other methods. Such an adaptor may ultimately be joined to a target nucleic acid molecule. In other embodiments, one or more amplification cycles (e.g., PCR amplification) are applied to introduce the molecular barcode and / or sample index into the nucleic acid molecule, generally using conventional nucleic acid amplification methods. The amplification may be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index may be introduced simultaneously or in any sequential order. In some embodiments, the molecular barcode and / or sample index are introduced prior to and / or after the sequence capture step is performed. In some embodiments, only the molecular barcode is introduced prior to probe capture, and the sample index is introduced after the sequence capture step is performed. In some embodiments, both the molecular barcode and the sample index are introduced prior to performing a probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step is performed. Typically, a sequence capture protocol involves introducing a targeted nucleic acid sequence, e.g., a single stranded nucleic acid molecule complementary to a coding sequence of a genomic region, where mutations in such region are associated with a cancer type.

[0112] In some embodiments, the tag may be located at one or both ends of the sample nucleic acid molecule. In some embodiments, the tag is a predetermined or random or semi-random sequence oligonucleotide. In some embodiments, the tag may be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide in length. The tag may be ligated to the sample nucleic acid either randomly or non-randomly.

[0113] In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely tagged with a molecular barcode or combination of molecular barcodes. In other embodiments, multiple barcodes may be used such that the barcodes are not necessarily unique to each other among the plurality (e.g., non-unique molecular barcodes). In these embodiments, the barcodes are typically attached to individual molecules (e.g., by ligation or PCR amplification) such that the combination of barcode and sequence creates a unique sequence that can be individually tracked. Detection of non-uniquely tagged barcodes, in combination with endogenous sequence information (e.g., the sequence of the original nucleic acid molecule in the sample, subsequences of sequence reads at one or both ends, the length of the sequence reads, and / or the start (initiation) and / or end (termination) portions corresponding to the length of the original nucleic acid molecule in the sample), typically allows for the assignment of a unique identification to a particular molecule. The length or number of base pairs of an individual sequence read may also optionally be used to assign a unique identification to a given molecule. As described herein, a fragment from a single strand of a nucleic acid to which a unique identification has been assigned may thereby allow for subsequent identification of fragments from the parental and / or complementary strands.

[0114] In some embodiments, molecular barcodes are introduced to molecules in a sample at expected ratios of identifiers (e.g., combinations of unique or non-unique barcodes). One exemplary format uses about 2 to about 1,000,000 different molecular barcodes, or about 5 to about 150 different molecular barcodes, or about 20 to about 50 different molecular barcodes ligated to both ends of a target molecule. Alternatively, about 25 to about 1,000,000 different barcodes may be used. For example, for 20-50 x 20-50 tags, a total of 400-2,500 identifiers are created. Such a number of identifiers is typically sufficient for different molecules with the same start and stop points to have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes.

[0115] In some embodiments, the assignment of unique or non-unique molecular barcodes in reactions is performed using methods and systems described, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is incorporated by reference in its entirety. C. Amplification

[0116] The sample nucleic acid may be flanked by adaptors and amplified by PCR and other amplification methods using nucleic acid primers binding to primer binding sites in the adaptors that flank the DNA molecule to be amplified. In some embodiments, the amplification method involves cycles of extension, denaturation, and annealing resulting from thermal cycling, or can be isothermal, for example, as in transcription-mediated amplification. Other examples of amplification methods that may be optionally utilized include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.

[0117] Typically, the amplification reaction produces multiple non-uniquely or uniquely tagged nucleic acid amplicons with molecular barcodes and sample indices with sizes ranging from about 150 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicons have a size of about 180 nt. In some embodiments, the amplicons have a size of about 200 nt. D. enrichment

[0118] In some embodiments, the sequence is enriched prior to sequencing the nucleic acid. Enrichment is optionally performed for specific target regions or non-specifically ("target sequence"). In some embodiments, the target region of interest may be enriched with nucleic acid capture probes ("baits") selected for one or more bait set panels using a discriminatory tiling and capture scheme. A discriminatory tiling and capture scheme generally uses different relative concentrations of bait sets to differentially tile (e.g., at different "resolutions") across the genomic regions associated with the baits according to a set of constraints (e.g., sequencing device constraints such as sequencing load, availability of each bait, etc.) to capture the targeted nucleic acid at a desired level for downstream sequencing. These targeted genomic regions of interest optionally include natural or synthetic nucleotide sequences of nucleic acid structures. In some embodiments, biotin-labeled beads with probes to one or more regions of interest can be used to enrich for the regions of interest after capturing target sequences, optionally followed by amplification of those regions.

[0119] Sequence capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In some embodiments, the probe set strategy involves tiling probes across a region of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. Sets can have a depth (e.g., depth of coverage) of about 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, 15X, 20X, 50X, or greater than 50X. The effectiveness of sequence capture generally depends, in part, on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe. E. Sequencing

[0120] The sample nucleic acid, optionally flanked by the adaptor, is generally subjected to sequencing, with or without prior amplification. Sequencing methods or optionally used commercially available formats include, for example, Sanger sequencing, high throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), digital gene expression (Helicos), next generation sequencing (NGS), single molecule sequencing by synthesis, and the like. Examples of sequencing methods include sequencing using a 3D nanopore platform, such as 3D sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, PacBio, SOLiD, Ion Torrent, or nanopore platform. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. Sample processing units may also include multiple sample chambers, allowing multiple runs to be processed simultaneously.

[0121] The sequencing reaction can be performed on one or more nucleic acid fragment types or regions that are known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequencing reaction can be performed on at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, the sequencing reaction can be performed on less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.

[0122] Simultaneous sequencing reaction may be carried out using multiplex sequencing technology.In some embodiments, at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions are used to sequence cell-free polynucleotides.In other embodiments, less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions are used to sequence cell-free polynucleotides.Sequencing reactions are typically carried out sequentially or simultaneously.Subsequent data analysis is generally carried out on all or part of sequencing reactions. In some embodiments, data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis is performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An example of read depth is about 1000 to about 50,000 reads per locus (e.g., base position). F. Analysis

[0123] Sequencing may generate multiple sequencing reads or reads. The sequencing reads or reads may include a sequence of nucleotide data less than about 150 bases long or less than about 90 bases long. In some embodiments, the reads are about 80 bases to about 90 bases long, for example, about 85 bases long. In some embodiments, the methods of the present disclosure are applied to very short reads, for example, less than about 50 bases long or less than about 30 bases long. The sequencing read data may include sequence data as well as meta information. The sequence read data may be stored in any suitable file format, including, for example, a VCF file, a FASTA file, or a FASTQ file.

[0124] FASTA may refer to a computer program for searching sequence databases, and the name FASTA may also refer to a standard file format. For example, FASTA is described by, for example, Pearson & Lipman, 1988, Improved tools for biological sequence comparison, PNAS85:2444-2448, which is incorporated herein by reference in its entirety. A sequence in FASTA format begins with a single line of description, followed by lines of sequence data. The description line is separated from the sequence data by a greater than (">") symbol in the first column. The word following the ">" symbol is the identifier of the sequence, and the rest of the line is the description (both are optional). There should be no space between the ">" and the first character of the identifier. It is recommended that all lines of text be less than 80 characters long. A sequence ends when another line begins with ">", indicating the start of another sequence.

[0125] The FASTQ format is a text-based format for storing both biological sequences (usually nucleotide sequences) and their corresponding quality scores. It is similar to the FASTA format, but the quality scores follow the sequence data. Both the sequence letters and the quality scores are encoded in a single ASCII character for brevity. The FASTQ format is described, for example, in Cock et al. ("The Sanger FASTQ file format for sequences with quality scores, and the Solexa / Illumina FASTQ variants,” Nucleic Acids Res 38(6):1767-1771, 2009) (incorporated herein by reference in its entirety).

[0126] For FASTA and FASTQ files, the meta information includes a description line and does not include a line of sequence data. In some embodiments, for FASTQ files, the meta information includes a quality score. For FASTA and FASTQ files, the sequence data begins after the description line and is typically present using a subset of the IUPAC ambiguous codes, optionally with a "-". In an embodiment, the sequence data may use the letters A, T, C, G, and N, optionally including a "-" or U (e.g., to represent a gap or uracil) as needed.

[0127] In some embodiments, at least one of the master sequence read file and the output file are stored as plain text files (e.g., using an encoding such as ASCII; ISO / IEC646; EBCDIC; UTF-8, or UTF-16). Computer systems provided by the present disclosure may include a text editor program capable of opening plain text files. A text editor program may refer to a computer program capable of presenting the contents of a text file (such as a plain text file) on a computer screen and allowing a human to edit the text (e.g., using a monitor, keyboard, and mouse). Examples of text editors include, but are not limited to, the Microsoft Text editor programs include Word, emacs, pico, vi, BBEdit, and TextWrangler. Text editor programs may be capable of displaying plain text files on a computer screen and showing meta-information and sequence reads in a human-readable format (e.g., not binary encoded, but instead using alphanumeric characters such as might be used when printed or handwritten).

[0128] Although the methods have been discussed with reference to FASTA or FASTQ files, the methods and systems of the present disclosure may be used to compress any suitable sequence file format, including, for example, files in the variant call format (VCF) format. A typical VCF file may include a header section and a data section. The header contains an arbitrary number of meta-information lines, each beginning with the characters "##", and tab-bounded field definition lines, each beginning with a single "#" character. The field definition lines specify eight required columns, and the body section contains lines of data that fill the columns defined by the field definition lines. The VCF format is described, for example, by Danecek et al. ("The variant call format and VCFtools," Bioinformatics 27(15):2156-2158, 2011) (incorporated herein by reference in its entirety). The header section may be treated as meta-information for writing to the compressed file, and the data section may be treated as lines, each of which will be stored in the master file only if unique.

[0129] Some embodiments provide for the assembly of sequencing reads. In the assembly, by alignment, for example, the sequencing reads are aligned with each other or with a reference sequence. By aligning each read, in turn, with a reference genome, all the reads are positioned relative to each other to create an assembly. In addition, aligning or mapping the sequencing reads to the reference sequence can also be used to identify variant sequences in the sequencing reads. Identifying variant sequences can be used in combination with the methods and systems described herein to further aid in the diagnosis or prognosis of a disease or condition, or to guide treatment decisions.

[0130] In some embodiments, any or all of the steps are automated. Alternatively, the method of the present disclosure may be embodied, in whole or in part, in one or more dedicated programs, for example, each optionally written in a compiled language such as C++, and then compiled and distributed as a binary. The method of the present disclosure may be implemented, in whole or in part, as a module by invoking functionality within an existing sequence analysis platform. In some embodiments, the method of the present disclosure includes several steps that are all automatically invoked in response to a single initiating queue (e.g., a trigger event or combinations thereof that originate from a human activity, another computer program, or a machine). Thus, the present disclosure provides a method in which any step or any combination of steps may occur automatically in response to a queue. "Automatically" generally means without any intervening human input, influence, or interaction (e.g., only in response to the original or pre-queue human activity).

[0131] The disclosed methods may also include various forms of output, including accurate and sensitive interpretation of the subject's nucleic acid sample. The readout output may be provided in a computer file format. In some embodiments, the output is a FASTA file, a FASTQ file, or a VCF file. The output may be processed to produce a text file or an XML file containing sequence data, such as a sequence of the nucleic acid aligned to a sequence of the reference genome. In other embodiments, the processing results in an output containing coordinates or strings that describe one or more mutations in the subject nucleic acid relative to the reference genome. The alignment strings may include Simple UnGapped Alignment Report (SUGAR), Verbose Useful Labeled Gapped Alignment Report (VULGAR), and Compact Idiosyncratic Gapped Alignment Report (CIGAR) (see, e.g., Ning et al., J. Am. Soc. 2007, 143:131-132). al., Genome Research 11(10):1725-9, 2001, which is incorporated herein by reference in its entirety. These strings can be found, for example, in the European Bioinformatics This may be implemented within the Exonerate sequence alignment software from the Institute (Hinxton, UK).

[0132] In some embodiments, a sequence alignment is produced, such as a sequence alignment map (SAM) or binary alignment map (BAM) file, including a CIGAR string (SAM format is described, for example, by Li et al., "The Sequence Alignment / Map format and SAMtools," Bioinformatics, 25(16):2078-9, 2009, incorporated herein by reference in its entirety). In some embodiments, CIGAR displays or includes gapped alignments, one per line. CIGAR is a condensed pairwise alignment format reported as a CIGAR string. CIGAR strings can be useful for representing long (e.g., genomic) pairwise alignments. CIGAR strings may be used in the SAM format to represent alignments of reads to a reference genome sequence.

[0133] CIGAR strings may follow established motifs. Each letter is preceded by a number, giving the base count of the event. Letters used can include M, I, D, N, and S (M=match, I=insertion, D=deletion, N=gap, S=substitution). CIGAR strings define a sequence of matches / mismatches and deletions (or gaps). For example, the CIGAR string 2MD3M2D2M may indicate that the alignment contains 2 matches, 1 deletion (the number 1 is omitted to save some space), 3 matches, 2 deletions, and 2 matches.

[0134] In some embodiments, a nucleic acid population is prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids with single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments thereof that can be used optionally include Klenow large fragment and T4 polymerase. For 5' overhangs, the enzyme typically extends the recessed 3' end on the opposing strand until it is flush with the 5' end and produces a blunt end. For 3' overhangs, the enzyme generally digests from the 3' end up to, and sometimes beyond, the 5' end of the opposing strand. If the digestion proceeds beyond the 5' end of the opposing strand, the gap can be filled by an enzyme with the same polymerase activity used for the 5' overhang. The formation of blunt ends on double-stranded nucleic acids, for example, facilitates attachment of adapters and subsequent amplification.

[0135] In some embodiments, the population of nucleic acids undergoes additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA (e.g., complementary DNA or cDNA). These forms of nucleic acids are also optionally ligated to adaptors and amplified.

[0136] With or without previous amplification, the nucleic acid can be subjected to the process of forming blunt ends described above, and optionally other nucleic acids in the sample can also be sequenced to produce sequenced nucleic acids. A sequenced nucleic acid can refer to either the sequence (e.g., sequence information) of a nucleic acid or the nucleic acid whose sequence has been determined. Sequencing can be performed to provide sequence data of individual nucleic acid molecules in a sample, either directly or indirectly, from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.

[0137] In some embodiments, the double-stranded nucleic acid with single-stranded overhangs in the sample after blunt end formation is ligated to an adaptor containing a barcode at both ends, and sequencing determines the nucleic acid sequence as well as the in-line barcode introduced by the adaptor. The blunt-ended DNA molecule is optionally ligated to the blunt end of an at least partially double-stranded adaptor (e.g., a Y-shaped or bell-shaped adaptor). Alternatively, the blunt ends of the sample nucleic acid and the adaptor can be terminated with complementary nucleotides to facilitate ligation (e.g., for sticky end ligation).

[0138] A nucleic acid sample is typically contacted with a sufficient number of adapters such that there is a low probability (e.g., less than about 1 or 0.1%) that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from the adapters ligated at both ends. The use of adapters can thus allow for the identification of a family of nucleic acid sequences with the same start and stop points on the reference nucleic acid and ligated to the same combination of barcodes. Such a family can represent the sequence of the amplification product of the nucleic acid in the sample before amplification. The sequences of the family members can be compiled to derive a consensus nucleotide or a complete consensus sequence for the nucleic acid molecules in the original sample as modified by blunt end formation and adapter attachment. In other words, a nucleotide occupying a defined position of the nucleic acid in the sample can be determined to be the consensus of the nucleotide occupying its corresponding position in the family member sequence. A family can include sequences of one or both strands of a double-stranded nucleic acid. When a family member includes sequences of both strands from a double-stranded nucleic acid, the sequence of one strand may be converted to its complement for the purpose of compiling the sequence and deriving a consensus nucleotide or sequence. Some families include only a single member sequence. In this case, this sequence may be considered as the sequence of the nucleic acid in the sample before amplification. Alternatively, families with only a single member sequence may be excluded from further analysis.

[0139] Nucleotide variants (e.g., SNVs or indels) in the sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is often a known sequence, such as a known full or partial genome sequence from a subject (e.g., a full genome sequence of a human subject). The reference sequence can be, for example, hG19 or hG38. The sequenced nucleic acid can represent a consensus of sequences determined directly for nucleic acids in a sample or sequences of amplification products of such nucleic acids, as described above. The comparison can be performed at one or more designated positions on the reference sequence. When the individual sequences are maximally aligned, a subset of the sequenced nucleic acids can be identified that includes positions corresponding to the designated positions of the reference sequence. Within such a subset, the sequenced nucleic acids can be determined that include nucleotide variants at designated positions, if applicable, and optionally, if applicable, include reference nucleotides (e.g., identical to those in the reference sequence). If the number of sequenced nucleic acids in the subset that contain a nucleotide variant exceeds a selected threshold, the variant nucleotide may be considered to be at the designated position. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids in the subset that contain a nucleotide variant, among other possibilities, or it can be a ratio of at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, etc., of sequenced nucleic acids in the subset that contain a nucleotide variant. The comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, the comparison can be performed for designated positions that occupy at least about 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, about 20-500 or about 50-300 consecutive positions.

[0140] Additional details regarding nucleic acid sequencing, including the formats and uses described herein, can also be found in, for example, Levy et al., Annual Review of Genomics and Human Genetics,17:95-115(2016), Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11(2012), Voelkerding et al., Clinical Chem., 55:641-658(2009), MacLean et al., Nature Rev. Microbiol., 7:287-296(2009), Astier et al., J Am Chem Soc.,128(5):1705-10(2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Nos. 7,501,245, 6,818,395, 6,911,345, 7,501,245, 7,329,492, 7,170,050, 7,302,146, 7,313,308, and 7,476,503, each of which is incorporated by reference in its entirety. III. Computer Systems

[0141] The methods of the present disclosure may be implemented using or with the aid of a computer system. For example, such a method may include, and may be implemented on a computer processor, (a) determining a plurality of quantitative measurements for a nucleic acid variant from a sample of nucleic acid molecules (e.g., a sample of cfDNA), the plurality of quantitative measurements including a total allele count and a minor allele count for the nucleic acid variant; (b) identifying an associated variable of the nucleic acid variant from the sample; (c) determining a quantitative value for the associated variable of the nucleic acid variant; (d) generating a statistical model for expected germline mutation allele count at the genomic locus at which the nucleic acid variant lies; (e) generating a probability value (p-value) for the nucleic acid variant based, at least in part, on the statistical model for expected germline mutation allele count, the quantitative value for the associated variable of the nucleic acid variant, and at least one of the plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0142] 4 illustrates a computer system 401 that is programmed or otherwise configured to implement the methods of the present disclosure. The computer system 401 can coordinate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 401 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.

[0143] The computer system 401 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 405, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 401 also includes a memory or memory location 410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 415 (e.g., hard disk), a communication interface 420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 425, such as cache, other memory, data storage devices, and / or electronic display adapters. The memory 410, the storage unit 415, the interface 420, and the peripheral devices 425 communicate with the CPU 405 through a communication network or bus (solid lines), such as a motherboard. The storage unit 415 may be a data storage unit (or data repository) for storing data. The computer system 401 may be operatively coupled to a computer network 430 with the aid of the communication interface 420. The computer network 430 may be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The computer network 430 may in some cases be a telecommunications and / or data network. The computer network 430 may include one or more computer servers, which may enable distributed computing such as cloud computing. The network 430 may in some cases implement a peer-to-peer network, which may enable devices coupled to the computer system 401 to behave as clients or servers with the aid of the computer system 401.

[0144] The CPU 405 may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 410. Examples of operations performed by the CPU 405 may include fetch, decode, execute, and write back.

[0145] The storage unit 415 can store files such as drivers, libraries, and saved programs. The storage unit 415 can store programs and recorded sessions generated by a user, as well as output associated with the programs. The storage unit 415 can store user data, such as user preferences and user programs. The computer system 401 can include one or more additional data storage units that are external to the computer system 401, in some cases, such as those located on a remote server in communication with the computer system 401 through an intranet or the Internet. Data may be transferred from one location to another, for example, using a communications network or physical data transfer (e.g., using a hard drive, thumb drive, or other data storage mechanism).

[0146] Computer system 401 can communicate with one or more remote computer systems through network 430. For example, computer system 401 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad®, Samsung® Galaxy Tab), a phone, a smartphone (e.g., Apple® iPhone®, Android-enabled devices, Blackberry®), or a personal digital assistant. A user can access computer system 401 via network 430.

[0147] Methods as described herein can be implemented via machine (e.g., computer processor) executable code stored on electronic storage locations of computer system 401, such as, for example, memory 410 or electronic storage unit 415. Machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by processor 405. In some cases, the code can be read from storage unit 415 and stored on memory 410 for easy access by processor 405. In some situations, electronic storage unit 415 can be eliminated and machine executable instructions are stored on memory 410.

[0148] In one aspect, the disclosure provides a method for determining a genetic variant comprising the steps of: (a) determining a plurality of quantitative measurements for nucleic acid variants from a cfDNA sample, the plurality of quantitative measurements including total and minor allele counts for the nucleic acid variants; (b) identifying associated variables of the nucleic acid variants from the cfDNA sample; (c) determining quantitative values ​​for the associated variables of the nucleic acid variants; (d) generating a statistical model for expected germline mutation allele counts at genomic loci where the nucleic acid variants reside; and (e) determining a plurality of quantitative measurements for the nucleic acid variants from the cfDNA sample. generating a probability value (p-value) for the nucleic acid variant based, at least in part, on at least one of a statistical model for expected germline mutant allele counts, a quantitative value for an associated variable of the nucleic acid variant, and a plurality of quantitative measurements for the nucleic acid variant; and (f) classifying the nucleic acid variant as (i) of somatic origin when the p-value for the nucleic acid variant is below a predetermined threshold, or (ii) of germline origin when the p-value for the nucleic acid variant is at or above a predetermined threshold.

[0149] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled during run-time. The code can be provided in a programming language that can be selected to enable the code to be executed in a pre-compiled or compiled-at-time manner.

[0150] Aspects of the systems and methods provided herein, such as the computer system 401, can be embodied in programming. Various aspects of the technology may be considered as a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data, typically carried on or embodied in a type of machine-readable medium. The machine executable code can be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of the tangible memory of a computer, processor, or equivalent, or their associated modules, such as various semiconductor memories, tape drives, hard drives, and the like, that can provide a non-transitory storage device at any time for software programming.

[0151] All or portions of the software may at times be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical landline networks, and via various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, or the like, may also be considered media bearing the software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0152] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer or equivalent, such as those used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include copper wire and optical fiber, including coaxial cables, i.e., the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data transmission. Common forms of computer readable media thus include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or a DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with a pattern of holes, a RAM, a ROM, a PROM and an EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0153] The computer system 401 can include or communicate with an electronic display, including, for example, a user interface (UI) for providing one or more results of the sample analysis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0154] Additional details relating to computer systems and networks, databases, and computer program products can also be found in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed.(2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed.(2010), Coronel, Database Systems:Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated by reference in its entirety. IV.Applications A. Cancer and Other Diseases

[0155] In some embodiments, the methods and systems disclosed herein may be used to identify customized or targeted therapies to treat a given disease or condition in a patient based on the classification of a nucleic acid variant as being of somatic or germline origin. Typically, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), chronic myelomonocytic leukemia (CMM), chronic myelomonocytic ... L), liver cancer, hepatoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B cell lymphoma, non-Hodgkin's lymphoma, diffuse large B cell lymphoma, mantle cell lymphoma, T cell lymphoma, non-Hodgkin's lymphoma, precursor T lymphoblastic lymphoma / leukemia, peripheral T cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.

[0156] Non-limiting examples of other genetically-based diseases, disorders, or conditions that may optionally be evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth disease (CMT), Crimson-Cry syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duane's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, G-CSF, and GABAergic nephropathy. including Aucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibromatosis, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland's syndrome, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency syndrome (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylamine, Turner's syndrome, palatocardiofacial syndrome, WAGR syndrome, Wilson's disease, or equivalent. B. Therapy and Related Administration

[0157] In an embodiment, the methods disclosed herein relate to identifying and administering customized therapy to a patient, given the status of the nucleic acid variant as somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of improving immune response against a given cancer type. In an embodiment, immunotherapy refers to a method of improving T cell response against a tumor or cancer.

[0158] In an embodiment, the status of nucleic acid variants from a sample from a subject as somatic or germline origin may be compared with a database of comparator results from a reference population to identify customized or targeted therapy for the subject.Typically, the reference population includes patients suffering from the same cancer or disease type as the test subject, and / or patients undergoing or having undergone the same therapy as the test subject.Customized or targeted therapy (or therapy) may be identified when the nucleic acid variants and comparator results meet certain classification criteria (e.g., substantially or approximately match).

[0159] In some embodiments, the customized therapy described herein is typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Some therapeutic agents are administered orally. However, the customized therapy (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, including, for example, oral, sublingual, rectal, vaginal, urethral, ​​topical, ocular, nasal, and / or auricular, and administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like. EXAMPLES

[0160] Example 1 Using a beta-binomial model versus a threshold approach to determine whether EGFRT790M mutations are of germline or somatic origin A set of samples was purchased from Guardant Health, Inc. (Redwood City, CA). The samples were processed and analyzed using a blood-based DNA assay developed by the National Institute of Genetics and Oncology (NIH). One of the analyzed samples harbored a T790M mutation (single nucleotide variant) within the EGFR gene at genomic location 55249071 on chromosome 7. The mutant allele count (A) and total allele count (B) of the variant were estimated to be 1,855 and 10,806, respectively, using bioinformatics analysis. The mutant allele fraction (MAF) of the variant was estimated to be 0.177 (MAF=A / B).

[0161] To determine the origin of the variants, the EGFR gene was used as a bin in the beta-binomial model. Six common germline heterozygous SNPs were found in the EGFR gene that were either (i) listed in the ExAC database with a population allele frequency greater than 0.001 or (ii) listed as known germline heterozygous SNPs in the database of historical sample sets with a MAF less than 0.9. The mutant allele counts and total allele counts of these six common germline heterozygous SNPs were used in the beta-binomial model to estimate the μ EGFR The maximum likelihood estimate (MLE) of the parameters was estimated to be 0.3971 using a beta-binomial model. Figure 5A shows a plot of MAF vs. genomic location for the T790M (●) variant and six common germline heterozygous SNPs (▲). Figure 5B shows a plot of min(MAF,1-MAF) vs. genomic location for the T790M (●) variant and six common germline heterozygous SNPs (▲). The μ of 0.3971 estimated by the beta-binomial model EGFR is shown as a solid line in both Figures 5A and 5B. The ρ parameter was estimated as the median ρ value for germline SNPs in the historical sample set, 9.2 × 10 -5 It was calculated that μ EGFR Using these values ​​for the ρ value, the two-sided p value for the T790M variant was 2.8 × 10 -302It was calculated that the p-value was 10 -16 A predefined threshold of 0.01 was used to distinguish the origin of the variant (e.g., germline or somatic). Since the p-value for the T790M variant is below the predefined threshold, the T790M variant is determined to be of somatic origin.

[0162] In comparison with the use of the beta-binomial model, the origin of any variant can be determined based on the MAF threshold method, such as by using a MAF of 0.15 as a threshold (e.g., classifying variants with MAF less than 0.15 as somatic variants, or variants with MAF greater than or equal to 0.15 as germline variants). As described herein, the T790M variant had a measured MAF of 0.177, which is greater than the MAF threshold of 0.15. Thus, the T790M variant would be erroneously identified as germline origin using the MAF threshold method. In contrast, the beta-binomial model accurately models the local genomic context of the EGFR gene by taking into account any allelic imbalance observed within the EGFR gene, and thus correctly identifies the variant as somatic origin.

[0163] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention. It is therefore contemplated that the present invention shall also cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended to cover methods and structures within the scope of these claims and their equivalents.

[0164] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art from a perusal of the present disclosure that various changes in form and detail may be made without departing from the true scope of the present disclosure and may be practiced within the purview of the appended claims. For example, all methods, systems, computer readable media, and / or component features, steps, elements, or other aspects thereof may be used in various combinations.

[0165] All patents, patent applications, websites, other publications, or documents, and accession numbers, etc. cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual item was expressly and individually indicated to be incorporated by reference. Where different versions of a sequence are associated with accession numbers at different times, the version associated with that accession number on the effective filing date of this application is meant. By effective filing date is meant an earlier date than the actual filing date, or, where applicable, the filing date of a priority application that refers to the accession number. Similarly, where different versions of a publication, or website, etc. are published at different times, the version published closest to the effective filing date of the application is meant, unless otherwise indicated.

Claims

[Claim 1] The invention as depicted in the drawings.